Skip to content

Benchmarks ​

Linnet compiles one checked source to each framework's own fast path. Measured in bf16 on H100s: 24 real checkpoints from the model zoo against the stacks people already run them with, Llama 3.1 8B fine-tuned against TRL, and a synthetic Llama that isolates the generated code.

Faster than vLLM at batch 1 ​

Faster on 10 of 10 models: 1.08× to 2.72×, median 1.21×.

the stack Linnet replacesLinnetthe faster of the twodecode speed in tok/s, higher is better. Each card is one model on its own scale, the stack's bar above Linnet's; the faster of the two is red. Each side is its fastest configuration.

GPT-2 (124M)124 M2.72×faster
vLLM: 696 tok/svLLM696 tok/sLinnet torch (CUDA graphs): 1,897 tok/sLinnet, CUDA graphs1,897 tok/s
Qwen2.5 0.5B Instruct494 M1.49×faster
vLLM: 707 tok/svLLM707 tok/sLinnet JAX (XLA, generated source): 1,055 tok/sLinnet, XLA (generated JAX)1,055 tok/s
TinyLlama 1.1B Chat v1.01.1 B1.21×faster
vLLM: 650 tok/svLLM650 tok/sLinnet JAX (XLA, StableHLO): 788 tok/sLinnet, XLA (StableHLO)788 tok/s
SmolLM2 1.7B Instruct1.7 B1.25×faster
vLLM: 469 tok/svLLM469 tok/sLinnet JAX (XLA, generated source): 585 tok/sLinnet, XLA (generated JAX)585 tok/s
Phi-3 Mini 4K Instruct3.8 B1.20×faster
vLLM: 242 tok/svLLM242 tok/sLinnet JAX (XLA, StableHLO): 292 tok/sLinnet, XLA (StableHLO)292 tok/s
Qwen3 4B4.0 B1.11×faster
vLLM: 230 tok/svLLM230 tok/sLinnet JAX (XLA, generated source): 255 tok/sLinnet, XLA (generated JAX)255 tok/s
Mistral 7B Instruct v0.37.2 B1.10×faster
vLLM: 158 tok/svLLM158 tok/sLinnet JAX (XLA, StableHLO): 174 tok/sLinnet, XLA (StableHLO)174 tok/s
Llama 3.1 8B Instruct8.0 B1.10×faster
vLLM: 152 tok/svLLM152 tok/sLinnet JAX (XLA, StableHLO): 167 tok/sLinnet, XLA (StableHLO)167 tok/s
Qwen3 8B8.2 B1.08×faster
vLLM: 147 tok/svLLM147 tok/sLinnet JAX (XLA, generated source): 159 tok/sLinnet, XLA (generated JAX)159 tok/s
GPT-OSS 20B20.9 B1.23×faster
vLLM: 299 tok/svLLM299 tok/sLinnet torch (CUDA graphs): 368 tok/sLinnet, CUDA graphs368 tok/s

One request at a time, the speed a chat turn or agent step sees, Linnet's fastest path (generated PyTorch as CUDA graphs, or XLA) decodes faster than vLLM on the same checkpoint. Portability costs nothing here: the generated code calls the kernels hand-written code would, and each decoding step runs as one CUDA graph or XLA program, with the host out of the loop.

Faster in PyTorch and JAX ​

Faster on 24 of 24 models: 1.00× to 4.81×, median 2.09×.

the stack Linnet replacesLinnetthe faster of the twoEach model by its own measure: decode speed for a decoder, transcription for Whisper, batch-1 latency for the rest. Each card is one model on its own scale, the stack's bar above Linnet's; the faster of the two is red. Each side is its fastest configuration.

ResNet-1812 Mbatch-1 latency1.77×faster
transformers (torch.compile): 0.59 mstransformers, compiled0.59 msLinnet torch (CUDA graphs): 0.33 msLinnet, CUDA graphs0.33 ms
all-MiniLM-L6-v223 Mbatch-1 latency2.61×faster
transformers (torch.compile): 0.84 mstransformers, compiled0.84 msLinnet torch (CUDA graphs): 0.32 msLinnet, CUDA graphs0.32 ms
ResNet-5026 Mbatch-1 latency1.89×faster
transformers (torch.compile): 1.29 mstransformers, compiled1.29 msLinnet torch (CUDA graphs): 0.68 msLinnet, CUDA graphs0.68 ms
Whisper tiny38 Mtranscription2.32×faster
transformers (torch.compile, static cache): 51.8 mstransformers, compiled51.8 msLinnet torch (CUDA graphs): 22.3 msLinnet, CUDA graphs22.3 ms
SD VAE ft-MSE (decoder)49 Mbatch-1 latency1.00×faster
diffusers (torch.compile): 10.4 msdiffusers, compiled10.4 msLinnet torch (CUDA graphs): 10.4 msLinnet, CUDA graphs10.4 ms
ViT-Base/16 22487 Mbatch-1 latency2.42×faster
transformers (torch.compile): 1.97 mstransformers, compiled1.97 msLinnet torch (CUDA graphs): 0.81 msLinnet, CUDA graphs0.81 ms
DINOv2-Base87 Mbatch-1 latency1.29×faster
transformers (torch.compile): 1.69 mstransformers, compiled1.69 msLinnet torch (CUDA graphs): 1.31 msLinnet, CUDA graphs1.31 ms
SAM ViT-Base94 Mencoder pass1.04×faster
transformers (torch.compile): 28.9 mstransformers, compiled28.9 msLinnet torch (CUDA graphs): 27.7 msLinnet, CUDA graphs27.7 ms
BERT Base (uncased)109 Mbatch-1 latency2.27×faster
transformers (torch.compile): 1.67 mstransformers, compiled1.67 msLinnet torch (CUDA graphs): 0.74 msLinnet, CUDA graphs0.74 ms
RoBERTa Base124 Mbatch-1 latency2.14×faster
transformers (torch.compile): 1.57 mstransformers, compiled1.57 msLinnet torch (CUDA graphs): 0.73 msLinnet, CUDA graphs0.73 ms
GPT-2 (124M)124 Mdecode speed4.53×faster
transformers (torch.compile, static cache): 419 tok/stransformers, compiled419 tok/sLinnet torch (CUDA graphs): 1,897 tok/sLinnet, CUDA graphs1,897 tok/s
ModernBERT-base149 Mbatch-1 latency2.28×faster
transformers (torch.compile): 2.31 mstransformers, compiled2.31 msLinnet torch (CUDA graphs): 1.01 msLinnet, CUDA graphs1.01 ms
SigLIP Base/16 224203 Mbatch-1 latency2.04×faster
transformers (torch.compile): 1.89 mstransformers, compiled1.89 msLinnet torch (CUDA graphs): 0.93 msLinnet, CUDA graphs0.93 ms
Qwen2.5 0.5B Instruct494 Mdecode speed4.81×faster
transformers (torch.compile, static cache): 189 tok/stransformers, compiled189 tok/sLinnet torch (CUDA graphs): 908 tok/sLinnet, CUDA graphs908 tok/s
TinyLlama 1.1B Chat v1.01.1 Bdecode speed4.19×faster
transformers (torch.compile, static cache): 172 tok/stransformers, compiled172 tok/sLinnet torch (CUDA graphs): 720 tok/sLinnet, CUDA graphs720 tok/s
Whisper large-v31.5 Btranscription2.85×faster
transformers (torch.compile, static cache): 192 mstransformers, compiled192 msLinnet torch (CUDA graphs): 67.4 msLinnet, CUDA graphs67.4 ms
SmolLM2 1.7B Instruct1.7 Bdecode speed2.61×faster
transformers (torch.compile, static cache): 188 tok/stransformers, compiled188 tok/sLinnet torch (CUDA graphs): 491 tok/sLinnet, CUDA graphs491 tok/s
SDXL Base UNet2.6 Bone denoising step1.01×faster
diffusers (torch.compile): 35.0 msdiffusers, compiled35.0 msLinnet torch (CUDA graphs): 34.7 msLinnet, CUDA graphs34.7 ms
Phi-3 Mini 4K Instruct3.8 Bdecode speed1.49×faster
transformers (torch.compile, static cache): 177 tok/stransformers, compiled177 tok/sLinnet torch (CUDA graphs): 264 tok/sLinnet, CUDA graphs264 tok/s
Qwen3 4B4.0 Bdecode speed2.00×faster
transformers (torch.compile, static cache): 123 tok/stransformers, compiled123 tok/sLinnet torch (CUDA graphs): 247 tok/sLinnet, CUDA graphs247 tok/s
Mistral 7B Instruct v0.37.2 Bdecode speed1.43×faster
transformers (torch.compile, static cache): 118 tok/stransformers, compiled118 tok/sLinnet torch (CUDA graphs): 169 tok/sLinnet, CUDA graphs169 tok/s
Llama 3.1 8B Instruct8.0 Bdecode speed1.46×faster
transformers (torch.compile, static cache): 110 tok/stransformers, compiled110 tok/sLinnet torch (CUDA graphs): 161 tok/sLinnet, CUDA graphs161 tok/s
Qwen3 8B8.2 Bdecode speed1.43×faster
transformers (torch.compile, static cache): 107 tok/stransformers, compiled107 tok/sLinnet torch (CUDA graphs): 154 tok/sLinnet, CUDA graphs154 tok/s
GPT-OSS 20B20.9 Bdecode speed3.67×faster
transformers (torch.compile, static cache): 100 tok/stransformers, compiled100 tok/sLinnet torch (CUDA graphs): 368 tok/sLinnet, CUDA graphs368 tok/s

Against the implementation people run on each framework today, Linnet's code is faster on every model in PyTorch and JAX. The PyTorch gap is widest where the host sets the pace (small models, single decoding steps, batch-1 encoders), because one CUDA graph removes that overhead. It closes on the large convolutional models: the SD VAE decoder and SDXL UNet come out even with torch.compile, though Linnet's XLA program decodes the VAE fastest (7.4 ms). On ONNX Runtime and behind Triton, Linnet is ahead on most models, but the reference export is faster on MiniLM by 16% and on ResNet-50 by 13%, and 21% ahead on ResNet-50 behind Triton.

Exports run as the originals do ​

Within 5% of the original in 16 of 19 engine and model pairs, 2.9% apart at the median; the same first token in 14 of 14.

the stack Linnet replacesLinnetdecode speed in tok/s, higher is better. Each card is one model on its own scale, the original's bar above the export's. Each side is its fastest configuration.

GPT-2 (124M)124 M+18.0%faster
vLLM: 696 tok/svLLM696 tok/sLinnet -> vLLM: 821 tok/svLLM on Linnet's export821 tok/sSGLang: 897 tok/sSGLang897 tok/sLinnet -> SGLang: 923 tok/sSGLang on Linnet's export923 tok/sText Generation Inference: 548 tok/sTGI548 tok/sLinnet -> Text Generation Inference: 580 tok/sTGI on Linnet's export580 tok/s
Qwen2.5 0.5B Instruct494 M−27.9%slower
vLLM: 707 tok/svLLM707 tok/sLinnet -> vLLM: 510 tok/svLLM on Linnet's export510 tok/s
TinyLlama 1.1B Chat v1.01.1 B+2.9%the same
vLLM: 650 tok/svLLM650 tok/sLinnet -> vLLM: 669 tok/svLLM on Linnet's export669 tok/sSGLang: 632 tok/sSGLang632 tok/sLinnet -> SGLang: 634 tok/sSGLang on Linnet's export634 tok/sText Generation Inference: 274 tok/sTGI274 tok/sLinnet -> Text Generation Inference: 277 tok/sTGI on Linnet's export277 tok/s
SmolLM2 1.7B Instruct1.7 B+3.3%faster
vLLM: 469 tok/svLLM469 tok/sLinnet -> vLLM: 485 tok/svLLM on Linnet's export485 tok/sSGLang: 469 tok/sSGLang469 tok/sLinnet -> SGLang: 472 tok/sSGLang on Linnet's export472 tok/sText Generation Inference: 236 tok/sTGI236 tok/sLinnet -> Text Generation Inference: 242 tok/sTGI on Linnet's export242 tok/s
Phi-3 Mini 4K Instruct3.8 B+3.5%faster
vLLM: 242 tok/svLLM242 tok/sLinnet -> vLLM: 251 tok/svLLM on Linnet's export251 tok/s
Qwen3 4B4.0 B+2.4%the same
vLLM: 230 tok/svLLM230 tok/sLinnet -> vLLM: 236 tok/svLLM on Linnet's export236 tok/s
Mistral 7B Instruct v0.37.2 B+4.5%faster
vLLM: 158 tok/svLLM158 tok/sLinnet -> vLLM: 166 tok/svLLM on Linnet's export166 tok/sSGLang: 165 tok/sSGLang165 tok/sLinnet -> SGLang: 165 tok/sSGLang on Linnet's export165 tok/sText Generation Inference: 118 tok/sTGI118 tok/sLinnet -> Text Generation Inference: 121 tok/sTGI on Linnet's export121 tok/s
Llama 3.1 8B Instruct8.0 B+3.3%faster
vLLM: 152 tok/svLLM152 tok/sLinnet -> vLLM: 157 tok/svLLM on Linnet's export157 tok/sSGLang: 158 tok/sSGLang158 tok/sLinnet -> SGLang: 158 tok/sSGLang on Linnet's export158 tok/sText Generation Inference: 112 tok/sTGI112 tok/sLinnet -> Text Generation Inference: 115 tok/sTGI on Linnet's export115 tok/s
Qwen3 8B8.2 B+4.4%faster
vLLM: 147 tok/svLLM147 tok/sLinnet -> vLLM: 153 tok/svLLM on Linnet's export153 tok/s

A model exported with linnet.hf or linnet.gguf runs in the engines people deploy like the original checkpoint, one request at a time and serving: the engine cannot tell the difference. The smallest models' serving pairs vary the most, since their runs last a second or two.

Serving many requests ​

Faster on 10 of 10 models: 1.02× to 1.84×, median 1.09×.

the stack Linnet replacesLinnetthe faster of the twoserving throughput in tok/s, higher is better. Each card is one model on its own scale, the stack's bar above Linnet's; the faster of the two is red. Each side is its fastest configuration.

GPT-2 (124M)124 M1.48×faster
vLLM (offline, continuous batching): 25,858 tok/svLLM25,858 tok/sLinnet torch (linnet.serve, CUDA graphs): 38,317 tok/slinnet.serve, CUDA graphs38,317 tok/s
Qwen2.5 0.5B Instruct494 M1.84×faster
vLLM (offline, continuous batching): 18,828 tok/svLLM18,828 tok/sLinnet torch (linnet.serve, CUDA graphs): 34,643 tok/slinnet.serve, CUDA graphs34,643 tok/s
TinyLlama 1.1B Chat v1.01.1 B1.11×faster
vLLM (offline, continuous batching): 22,674 tok/svLLM22,674 tok/sLinnet torch (linnet.serve, CUDA graphs): 25,091 tok/slinnet.serve, CUDA graphs25,091 tok/s
SmolLM2 1.7B Instruct1.7 B1.08×faster
vLLM (offline, continuous batching): 12,046 tok/svLLM12,046 tok/sLinnet torch (linnet.serve, CUDA graphs): 12,997 tok/slinnet.serve, CUDA graphs12,997 tok/s
Phi-3 Mini 4K Instruct3.8 B1.19×faster
vLLM (offline, continuous batching): 5,583 tok/svLLM5,583 tok/sLinnet torch (linnet.serve, CUDA graphs): 6,668 tok/slinnet.serve, CUDA graphs6,668 tok/s
Qwen3 4B4.0 B1.02×faster
vLLM (offline, continuous batching): 8,048 tok/svLLM8,048 tok/sLinnet torch (linnet.serve, CUDA graphs): 8,196 tok/slinnet.serve, CUDA graphs8,196 tok/s
Mistral 7B Instruct v0.37.2 B1.04×faster
vLLM (offline, continuous batching): 5,546 tok/svLLM5,546 tok/sLinnet torch (linnet.serve, CUDA graphs): 5,764 tok/slinnet.serve, CUDA graphs5,764 tok/s
Llama 3.1 8B Instruct8.0 B1.03×faster
vLLM (offline, continuous batching): 5,449 tok/svLLM5,449 tok/sLinnet torch (linnet.serve, CUDA graphs): 5,600 tok/slinnet.serve, CUDA graphs5,600 tok/s
Qwen3 8B8.2 B1.02×faster
vLLM (offline, continuous batching): 5,295 tok/svLLM5,295 tok/sLinnet torch (linnet.serve, CUDA graphs): 5,399 tok/slinnet.serve, CUDA graphs5,399 tok/s
GPT-OSS 20B20.9 B1.40×faster
vLLM (offline, continuous batching): 4,313 tok/svLLM4,313 tok/sLinnet torch (linnet.serve, CUDA graphs): 6,031 tok/slinnet.serve, CUDA graphs6,031 tok/s

linnet.serve serves faster than vLLM on every decoder. The lead is widest on small models, where a step is bound by kernel launches that one CUDA graph per step removes. From 4 B up it is 2-4%, and on gpt-oss 40%. Behind Triton Inference Server, linnet.serve outpaces Triton's own vLLM backend on every decoder.

Training faster than TRL ​

The same runs in Linnet and in TRL, step for step. On one GPU, Linnet's PyTorch step is 1.76 times TRL's for LoRA SFT, 2.4 times for DPO and 2.7 times for GRPO. Fully fine-tuned on four GPUs, the three stacks are within 5% of each other, with JAX the fastest. Each pair of runs reaches the same held-out loss, accuracy or reward. Training lists the conditions that differ.

Llama 3.1 8B Instruct on NVIDIA H100 80GB HBM3, 2026-10-04. Same data, 4096-token packed rows, hyperparameters and step count for every stack. Step time is the median after the first two steps; tokens/s counts real tokens, not padding.

fastest step Linnet TRL seconds a step, lower is better
SFT, LoRA (rank 16), one GPU: Alpaca, 4 rows a step1.76×faster than TRL
StackStepTokens/sPeak a GPUheld-out loss after 30 steps, one evaluator
Linnet PyTorch1.04 s15.6K36.4 GiB1.369
Linnet JAX1.32 s12.2K38.0 GiB1.374
TRL + PEFT1.82 s8.9K42.8 GiB1.373
SFT in full, four GPUs: fully sharded, f32 masters, AdamW1.05×faster than TRL
StackStepTokens/sPeak a GPUheld-out loss after 30 steps
Linnet PyTorch0.55 s29.4K47.5 GiB1.342
Linnet JAX0.54 s30.0K65.5 GiB1.345
Linnet JAX, remat0.67 s24.0K42.7 GiB1.344
TRL, FSDP20.57 s28.6K51.5 GiBlast-5 training loss 1.26 (Linnet 1.24)
DPO, LoRA, one GPU: UltraFeedback, 32 pairs a step2.42×faster than TRL
StackStepPeak a GPUaccuracy over the last five steps
Linnet PyTorch1.91 s35.6 GiB0.68
Linnet JAX2.41 s38.3 GiB0.65
TRL + PEFT (2 pairs at a time; 4 and 8 ran out of memory)4.63 s47.9 GiB0.70
GRPO, LoRA, one GPU: 16 prompts × 8 completions of up to 128 tokens2.66×faster than TRL
StackStepmean reward, first three steps to last three
Linnet PyTorch1.18 s0.44 → 0.39
Linnet JAX1.44 s0.44 → 0.40
TRL + vLLM (8 completions at a time; vLLM held 35% of the GPU outside this peak)3.14 s0.43 → 0.41

Two GPUs and offloading ​

Faster on 1 of 2 models: 0.96× to 1.02×, median 0.99×.

the stack Linnet replacesLinnetthe faster of the twodecode speed in tok/s, higher is better. Each card is one model on its own scale, the stack's bar above Linnet's; the faster of the two is red. Each side is its fastest configuration.

Llama 3.1 8B Instruct8.0 B1.02×faster
vLLM (tensor parallel on 2 GPUs): 233 tok/svLLM, tensor parallel233 tok/sLinnet torch (tensor parallel on 2 GPUs, shards): 238 tok/sLinnet, tensor parallel (PyTorch)238 tok/s
Qwen3 8B8.2 B0.96×slower
vLLM (tensor parallel on 2 GPUs): 220 tok/svLLM, tensor parallel220 tok/sLinnet torch (tensor parallel on 2 GPUs, shards): 212 tok/sLinnet, tensor parallel (PyTorch)212 tok/s

Split across two GPUs at batch 1, the two are within 4%: Linnet is ahead on Llama 3.1 8B, vLLM on Qwen3 8B. The offloaded rows run Llama 3.1 8B and Qwen3 8B on a GPU capped at 8 GiB, streaming half the layers from host memory: Llama peaks at 8.9 GiB and decodes 5.7 tokens per second, where otherwise it would not load at all. These rows show what a small GPU can do, so they are never marked best.

One source, every runtime ​

ModelPyTorchXLAONNX RuntimeTensorRTTritonlinnet.servevLLMSGLangTGIllama.cpp
all-MiniLM-L6-v2✓✓✓✓✓·····
BERT Base (uncased)✓✓✓✓✓·····
RoBERTa Base✓✓✓✓✓·····
ModernBERT-base✓✓✓✓✓·····
ResNet-18✓✓✓✓✓·····
ResNet-50✓✓✓✓✓·····
ViT-Base/16 224✓✓✓✓✓·····
DINOv2-Base✓✓✓✓✓·····
SAM ViT-Base✓✓✓✓······
SigLIP Base/16 224✓✓✓✓✓·····
Whisper tiny✓✓✓✓······
Whisper large-v3✓✓✓✓······
SD VAE ft-MSE (decoder)✓✓✓✓······
SDXL Base UNet✓✓✓✓······
GPT-2 (124M)✓✓✓✓✓✓✓✓✓✓
Qwen2.5 0.5B Instruct✓✓✓✓✓✓✓··✓
TinyLlama 1.1B Chat v1.0✓✓✓✓✓✓✓✓✓✓
SmolLM2 1.7B Instruct✓✓✓✓✓✓✓✓✓✓
Phi-3 Mini 4K Instruct✓✓✓✓✓✓✓··✓
Qwen3 4B✓✓✓✓✓✓✓··✓
Mistral 7B Instruct v0.3✓✓✓✓✓✓✓✓✓✓
Llama 3.1 8B Instruct✓✓✓✓ 2 of 3✓✓✓✓✓✓
Qwen3 8B✓✓✓✓ 2 of 3✓✓✓··✓
GPT-OSS 20B✓✓✓ 2 of 3✕ failed✓✓ 2 of 3····

✓ ran · ✓ n of m some of its configurations ran · ✕ failed · · not a runtime this model takes part in. Hover a cell for each configuration.

Every model in the zoo is one .linnet source. The runs that failed:

  • TensorRT cannot build an engine for gpt-oss (its Myelin compiler fails inside NVRTC) or for the 8 B decoders in f32.
  • gpt-oss in f32 is 84 GB of weights, more than the GPU holds, and its ONNX serving row runs out of memory.
  • transformers' generate_batch returns no tokens for gpt-oss, and KerasHub's batched gpt-oss fails in XLA's autotuner.

Every measurement ​

Hover a bar for its notes and how it compares with the stack it replaces. Each model's page in the zoo has its samples and every number.

Decoders, one request at a time ​

best in the chart Linnet existing stacks
GPT-2 (124M)124 M
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 6.70 decode_tok_s: 175 load_s: 2.01 peak_vram_mib: 1097 max |diff| vs reference: 0transformers175 tok/stransformers (torch.compile, static cache) ttft_ms: 4.07 decode_tok_s: 419 load_s: 1.47 peak_vram_mib: 1280 max |diff| vs reference: 1.25transformers, compiled419 tok/sLinnet torch (generated source) 0.69× the speed of transformers, compiled ttft_ms: 3.09 decode_tok_s: 290 load_s: 1.02 peak_vram_mib: 1020 max |diff| vs reference: 0.75 KV cache compiled for 768 positionsLinnet, generated PyTorch290 tok/sLinnet torch (torch.compile, inductor) 1.32× the speed of transformers, compiled ttft_ms: 1.59 decode_tok_s: 553 load_s: 0.47 peak_vram_mib: 1023 max |diff| vs reference: 1.25 KV cache compiled for 768 positionsLinnet, inductor553 tok/sLinnet torch (CUDA graphs) 4.53× the speed of transformers, compiled ttft_ms: 0.77 decode_tok_s: 1897 load_s: 0.46 peak_vram_mib: 1145 max |diff| vs reference: 1.25 KV cache compiled for 768 positionsLinnet, CUDA graphs1897 tok/sKerasHub (JAX) ttft_ms: 6.84 decode_tok_s: 1452 load_s: 6.16 peak_vram_mib: 660 driver_vram_mib: 1654 max |diff| vs reference: 0.25 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub1452 tok/sLinnet JAX (XLA, StableHLO) 1.27× the speed of KerasHub ttft_ms: 1.23 decode_tok_s: 1848 load_s: 0.52 peak_vram_mib: 514 driver_vram_mib: 1141 max |diff| vs reference: 0.75 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1848 tok/sLinnet JAX (XLA, generated source) 1.18× the speed of KerasHub ttft_ms: 1.31 decode_tok_s: 1715 load_s: 0.48 peak_vram_mib: 514 driver_vram_mib: 1141 max |diff| vs reference: 1.25 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1715 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 4.21 decode_tok_s: 427 load_s: 0.29 peak_vram_mib: 2606 max |diff| vs reference: 1.3825111389160156 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f32427 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 3.41 decode_tok_s: 468 load_s: 0.29 peak_vram_mib: 1714 max |diff| vs reference: 1.34375 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f16468 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 3.76 decode_tok_s: 421 load_s: 0.30 peak_vram_mib: 2216 max |diff| vs reference: 0.5 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf16421 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 2.69 decode_tok_s: 542 load_s: 0.29 peak_vram_mib: 4098 max |diff| vs reference: 1.3637847900390625 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f32542 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 2.04 decode_tok_s: 557 load_s: 0.29 peak_vram_mib: 3334 max |diff| vs reference: 1.78125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f16557 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 2.20 decode_tok_s: 510 load_s: 0.28 peak_vram_mib: 3166 max |diff| vs reference: 0.25 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf16510 tok/svLLM ttft_ms: 5.94 decode_tok_s: 696 load_s: 46.96 peak_vram_mib: 69721 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM696 tok/sLinnet -> vLLM 1.18× the speed of vLLM ttft_ms: 6.38 decode_tok_s: 821 load_s: 20.47 peak_vram_mib: 69721 driver_vram_mib: 69721 linnet.hf.export (1 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export821 tok/sSGLang ttft_ms: 10.37 decode_tok_s: 897 load_s: 30.86 peak_vram_mib: 69352 reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang897 tok/sLinnet -> SGLang 1.03× the speed of SGLang ttft_ms: 9.25 decode_tok_s: 923 load_s: 24.53 peak_vram_mib: 69352 driver_vram_mib: 69352 linnet.hf.export (2 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang on Linnet's export923 tok/sText Generation Inference ttft_ms: 14.19 decode_tok_s: 548 load_s: 30.08 peak_vram_mib: 61907 over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI548 tok/sLinnet -> Text Generation Inference 1.06× the speed of TGI ttft_ms: 14.09 decode_tok_s: 580 load_s: 26.07 peak_vram_mib: 61907 driver_vram_mib: 61907 linnet.hf.export (2 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI on Linnet's export580 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 3.16 decode_tok_s: 1651 load_s: 8.61 peak_vram_mib: 1057 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF1651 tok/s
Qwen2.5 0.5B Instruct494 M
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 12.40 decode_tok_s: 93.19 load_s: 1.96 peak_vram_mib: 1910 max |diff| vs reference: 0transformers93 tok/stransformers (torch.compile, static cache) ttft_ms: 7.03 decode_tok_s: 189 load_s: 1.55 peak_vram_mib: 1986 max |diff| vs reference: 0.19140625transformers, compiled189 tok/sLinnet torch (generated source) 0.57× the speed of transformers, compiled ttft_ms: 7.76 decode_tok_s: 108 load_s: 0.83 peak_vram_mib: 2010 max |diff| vs reference: 0.1875 KV cache compiled for 768 positionsLinnet, generated PyTorch108 tok/sLinnet torch (torch.compile, inductor) 1.71× the speed of transformers, compiled ttft_ms: 3.14 decode_tok_s: 322 load_s: 0.80 peak_vram_mib: 2648 max |diff| vs reference: 0.140625 KV cache compiled for 768 positionsLinnet, inductor322 tok/sLinnet torch (CUDA graphs) 4.81× the speed of transformers, compiled ttft_ms: 1.82 decode_tok_s: 908 load_s: 0.68 peak_vram_mib: 2365 max |diff| vs reference: 0.140625 KV cache compiled for 768 positionsLinnet, CUDA graphs908 tok/sKerasHub (JAX) ttft_ms: 10.04 decode_tok_s: 717 load_s: 7.22 peak_vram_mib: 2086 driver_vram_mib: 4728 max |diff| vs reference: 0.15625 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub717 tok/sLinnet JAX (XLA, StableHLO) 1.46× the speed of KerasHub ttft_ms: 3.23 decode_tok_s: 1047 load_s: 0.77 peak_vram_mib: 1572 driver_vram_mib: 2681 max |diff| vs reference: 0.3125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1047 tok/sLinnet JAX (XLA, generated source) 1.47× the speed of KerasHub ttft_ms: 3.01 decode_tok_s: 1055 load_s: 0.77 peak_vram_mib: 1572 driver_vram_mib: 2681 max |diff| vs reference: 0.2734375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1055 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 11.06 decode_tok_s: 153 load_s: 0.32 peak_vram_mib: 4120 max |diff| vs reference: 0.14293169975280762 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f32153 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 9.94 decode_tok_s: 138 load_s: 0.34 peak_vram_mib: 2998 max |diff| vs reference: 0.142578125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f16138 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 10.01 decode_tok_s: 146 load_s: 0.31 peak_vram_mib: 2286 max |diff| vs reference: 0.38671875 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf16146 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 7.83 decode_tok_s: 210 load_s: 0.29 peak_vram_mib: 8448 max |diff| vs reference: 0.13897037506103516 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f32210 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 5.74 decode_tok_s: 215 load_s: 0.29 peak_vram_mib: 5886 max |diff| vs reference: 5.984375 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f16215 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 5.58 decode_tok_s: 218 load_s: 0.29 peak_vram_mib: 5168 max |diff| vs reference: 0.25 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf16218 tok/svLLM ttft_ms: 7.58 decode_tok_s: 707 load_s: 57.83 peak_vram_mib: 70741 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM707 tok/sLinnet -> vLLM 0.72× the speed of vLLM ttft_ms: 12.19 decode_tok_s: 510 load_s: 39.15 peak_vram_mib: 70741 driver_vram_mib: 70741 linnet.hf.export (2 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export510 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 5.95 decode_tok_s: 767 load_s: 11.98 peak_vram_mib: 1951 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF767 tok/s
TinyLlama 1.1B Chat v1.01.1 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 15.95 decode_tok_s: 68.59 load_s: 2.59 peak_vram_mib: 3001 max |diff| vs reference: 0transformers69 tok/stransformers (torch.compile, static cache) ttft_ms: 8.21 decode_tok_s: 172 load_s: 1.60 peak_vram_mib: 3042 max |diff| vs reference: 0.09375transformers, compiled172 tok/sLinnet torch (generated source) 0.64× the speed of transformers, compiled ttft_ms: 8.06 decode_tok_s: 110 load_s: 0.86 peak_vram_mib: 2954 max |diff| vs reference: 0.125 KV cache compiled for 768 positionsLinnet, generated PyTorch110 tok/sLinnet torch (torch.compile, inductor) 1.92× the speed of transformers, compiled ttft_ms: 3.39 decode_tok_s: 331 load_s: 0.81 peak_vram_mib: 3901 max |diff| vs reference: 0.125 KV cache compiled for 768 positionsLinnet, inductor331 tok/sLinnet torch (CUDA graphs) 4.19× the speed of transformers, compiled ttft_ms: 2.53 decode_tok_s: 720 load_s: 0.78 peak_vram_mib: 3977 max |diff| vs reference: 0.125 KV cache compiled for 768 positionsLinnet, CUDA graphs720 tok/sKerasHub (JAX) ttft_ms: 12.06 decode_tok_s: 600 load_s: 9.21 peak_vram_mib: 2647 driver_vram_mib: 4728 max |diff| vs reference: 0.1875 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub600 tok/sLinnet JAX (XLA, StableHLO) 1.31× the speed of KerasHub ttft_ms: 5.42 decode_tok_s: 788 load_s: 1.47 peak_vram_mib: 2441 driver_vram_mib: 4729 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)788 tok/sLinnet JAX (XLA, generated source) 1.27× the speed of KerasHub ttft_ms: 4.76 decode_tok_s: 765 load_s: 1.44 peak_vram_mib: 2432 driver_vram_mib: 4729 max |diff| vs reference: 0.125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)765 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 14.57 decode_tok_s: 158 load_s: 0.31 peak_vram_mib: 6366 max |diff| vs reference: 0.093536376953125 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f32158 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 11.25 decode_tok_s: 158 load_s: 0.29 peak_vram_mib: 4130 max |diff| vs reference: 0.08984375 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f16158 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 11.55 decode_tok_s: 165 load_s: 0.30 peak_vram_mib: 3176 max |diff| vs reference: 0.125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf16165 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 12.25 decode_tok_s: 146 load_s: 0.29 peak_vram_mib: 13602 max |diff| vs reference: 0.09106159210205078 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f32146 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 7.22 decode_tok_s: 195 load_s: 0.34 peak_vram_mib: 8430 max |diff| vs reference: 0.11328125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f16195 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 7.33 decode_tok_s: 181 load_s: 0.31 peak_vram_mib: 7566 max |diff| vs reference: 0.15625 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf16181 tok/svLLM ttft_ms: 7.25 decode_tok_s: 650 load_s: 46.37 peak_vram_mib: 68727 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM650 tok/sLinnet -> vLLM 1.03× the speed of vLLM ttft_ms: 8.02 decode_tok_s: 669 load_s: 26.11 peak_vram_mib: 68727 driver_vram_mib: 68727 linnet.hf.export (2 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export669 tok/sSGLang ttft_ms: 8.53 decode_tok_s: 632 load_s: 35.88 peak_vram_mib: 69610 reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang632 tok/sLinnet -> SGLang 1.00× the speed of SGLang ttft_ms: 9.89 decode_tok_s: 634 load_s: 28.01 peak_vram_mib: 69610 driver_vram_mib: 69610 linnet.hf.export (4 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang on Linnet's export634 tok/sText Generation Inference ttft_ms: 21.75 decode_tok_s: 274 load_s: 32.08 peak_vram_mib: 61589 over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI274 tok/sLinnet -> Text Generation Inference 1.01× the speed of TGI ttft_ms: 23.95 decode_tok_s: 277 load_s: 28.07 peak_vram_mib: 61589 driver_vram_mib: 61589 linnet.hf.export (3 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI on Linnet's export277 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 7.19 decode_tok_s: 702 load_s: 14.96 peak_vram_mib: 2757 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF702 tok/s
SmolLM2 1.7B Instruct1.7 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 15.84 decode_tok_s: 63.96 load_s: 2.66 peak_vram_mib: 4287 max |diff| vs reference: 0transformers64 tok/stransformers (torch.compile, static cache) ttft_ms: 8.50 decode_tok_s: 188 load_s: 1.96 peak_vram_mib: 4340 max |diff| vs reference: 0.1015625transformers, compiled188 tok/sLinnet torch (generated source) 0.64× the speed of transformers, compiled ttft_ms: 7.44 decode_tok_s: 120 load_s: 1.06 peak_vram_mib: 4414 max |diff| vs reference: 0.15625 KV cache compiled for 768 positionsLinnet, generated PyTorch120 tok/sLinnet torch (torch.compile, inductor) 1.55× the speed of transformers, compiled ttft_ms: 4.50 decode_tok_s: 291 load_s: 0.97 peak_vram_mib: 5879 max |diff| vs reference: 0.1171875 KV cache compiled for 768 positionsLinnet, inductor291 tok/sLinnet torch (CUDA graphs) 2.61× the speed of transformers, compiled ttft_ms: 3.61 decode_tok_s: 491 load_s: 0.99 peak_vram_mib: 5929 max |diff| vs reference: 0.1171875 KV cache compiled for 768 positionsLinnet, CUDA graphs491 tok/sKerasHub (JAX) ttft_ms: 14.23 decode_tok_s: 395 load_s: 10.05 peak_vram_mib: 4102 driver_vram_mib: 8824 max |diff| vs reference: 0.125 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub395 tok/sLinnet JAX (XLA, StableHLO) 1.47× the speed of KerasHub ttft_ms: 7.08 decode_tok_s: 579 load_s: 2.16 peak_vram_mib: 4096 driver_vram_mib: 8825 max |diff| vs reference: 0.140625 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)579 tok/sLinnet JAX (XLA, generated source) 1.48× the speed of KerasHub ttft_ms: 6.11 decode_tok_s: 585 load_s: 2.13 peak_vram_mib: 4096 driver_vram_mib: 8823 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)585 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 17.34 decode_tok_s: 121 load_s: 0.31 peak_vram_mib: 10688 max |diff| vs reference: 0.09243297576904297 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f32121 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 13.00 decode_tok_s: 139 load_s: 0.31 peak_vram_mib: 5896 max |diff| vs reference: 0.085693359375 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f16139 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 13.62 decode_tok_s: 140 load_s: 0.35 peak_vram_mib: 5260 max |diff| vs reference: 0.1484375 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf16140 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 16.13 decode_tok_s: 101 load_s: 0.32 peak_vram_mib: 21484 max |diff| vs reference: 0.09415864944458008 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f32101 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 9.70 decode_tok_s: 141 load_s: 0.30 peak_vram_mib: 12000 max |diff| vs reference: 9.6875 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f16141 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 9.62 decode_tok_s: 138 load_s: 0.30 peak_vram_mib: 11552 max |diff| vs reference: 0.359375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf16138 tok/svLLM ttft_ms: 7.68 decode_tok_s: 469 load_s: 49.03 peak_vram_mib: 68883 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM469 tok/sLinnet -> vLLM 1.03× the speed of vLLM ttft_ms: 9.95 decode_tok_s: 485 load_s: 27.25 peak_vram_mib: 68883 driver_vram_mib: 68883 linnet.hf.export (4 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export485 tok/sSGLang ttft_ms: 9.29 decode_tok_s: 469 load_s: 42.68 peak_vram_mib: 69520 reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang469 tok/sLinnet -> SGLang 1.01× the speed of SGLang ttft_ms: 9.74 decode_tok_s: 472 load_s: 29.48 peak_vram_mib: 69520 driver_vram_mib: 69520 linnet.hf.export (6 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang on Linnet's export472 tok/sText Generation Inference ttft_ms: 23.54 decode_tok_s: 236 load_s: 34.07 peak_vram_mib: 61425 over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI236 tok/sLinnet -> Text Generation Inference 1.02× the speed of TGI ttft_ms: 20.83 decode_tok_s: 242 load_s: 28.07 peak_vram_mib: 61425 driver_vram_mib: 61425 linnet.hf.export (6 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI on Linnet's export242 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 10.44 decode_tok_s: 513 load_s: 21.85 peak_vram_mib: 4167 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF513 tok/s
Phi-3 Mini 4K Instruct3.8 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 19.26 decode_tok_s: 70.94 load_s: 4.61 peak_vram_mib: 8336 max |diff| vs reference: 0transformers71 tok/stransformers (torch.compile, static cache) ttft_ms: 12.43 decode_tok_s: 177 load_s: 2.68 peak_vram_mib: 8592 max |diff| vs reference: 0.15625transformers, compiled177 tok/sLinnet torch (generated source) 0.50× the speed of transformers, compiled ttft_ms: 12.25 decode_tok_s: 88.89 load_s: 1.91 peak_vram_mib: 8396 max |diff| vs reference: 0.265625 KV cache compiled for 768 positionsLinnet, generated PyTorch89 tok/sLinnet torch (torch.compile, inductor) 1.19× the speed of transformers, compiled ttft_ms: 7.90 decode_tok_s: 211 load_s: 1.52 peak_vram_mib: 8389 max |diff| vs reference: 0.140625 KV cache compiled for 768 positionsLinnet, inductor211 tok/sLinnet torch (CUDA graphs) 1.49× the speed of transformers, compiled ttft_ms: 7.05 decode_tok_s: 264 load_s: 1.51 peak_vram_mib: 8509 max |diff| vs reference: 0.140625 KV cache compiled for 768 positionsLinnet, CUDA graphs264 tok/sLinnet JAX (XLA, StableHLO) ttft_ms: 9.76 decode_tok_s: 292 load_s: 4.99 peak_vram_mib: 8469 driver_vram_mib: 17019 max |diff| vs reference: 0.125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)292 tok/sLinnet JAX (XLA, generated source) ttft_ms: 9.27 decode_tok_s: 292 load_s: 4.94 peak_vram_mib: 8469 driver_vram_mib: 17019 max |diff| vs reference: 0.3125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)292 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 31.67 decode_tok_s: 75.16 load_s: 0.37 peak_vram_mib: 20596 max |diff| vs reference: 0.13383817672729492 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f3275 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 22.20 decode_tok_s: 95.13 load_s: 0.36 peak_vram_mib: 10918 max |diff| vs reference: 0.15625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f1695 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 22.81 decode_tok_s: 93.82 load_s: 0.41 peak_vram_mib: 10020 max |diff| vs reference: 0.125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf1694 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 30.34 decode_tok_s: 50.12 load_s: 0.39 peak_vram_mib: 43010 max |diff| vs reference: 0.13777589797973633 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f3250 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 15.99 decode_tok_s: 84.37 load_s: 0.36 peak_vram_mib: 22674 max |diff| vs reference: 14.5 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f1684 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 16.09 decode_tok_s: 83.33 load_s: 0.37 peak_vram_mib: 22154 max |diff| vs reference: 0.1875 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf1683 tok/svLLM ttft_ms: 11.91 decode_tok_s: 242 load_s: 67.52 peak_vram_mib: 68419 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM242 tok/sLinnet -> vLLM 1.03× the speed of vLLM ttft_ms: 11.99 decode_tok_s: 251 load_s: 50.27 peak_vram_mib: 68492 driver_vram_mib: 68492 linnet.hf.export (8 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export251 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 16.75 decode_tok_s: 231 load_s: 27.90 peak_vram_mib: 8091 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF231 tok/s
Qwen3 4B4.0 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 20.68 decode_tok_s: 49.97 load_s: 5.74 peak_vram_mib: 8788 max |diff| vs reference: 0transformers50 tok/stransformers (torch.compile, static cache) ttft_ms: 13.16 decode_tok_s: 123 load_s: 2.71 peak_vram_mib: 8900 max |diff| vs reference: 0.125transformers, compiled123 tok/sLinnet torch (generated source) 0.59× the speed of transformers, compiled ttft_ms: 14.52 decode_tok_s: 72.21 load_s: 1.90 peak_vram_mib: 9358 max |diff| vs reference: 0.09375 KV cache compiled for 768 positionsLinnet, generated PyTorch72 tok/sLinnet torch (torch.compile, inductor) 1.36× the speed of transformers, compiled ttft_ms: 8.71 decode_tok_s: 168 load_s: 1.79 peak_vram_mib: 12741 max |diff| vs reference: 0.1875 KV cache compiled for 768 positionsLinnet, inductor168 tok/sLinnet torch (CUDA graphs) 2.00× the speed of transformers, compiled ttft_ms: 7.64 decode_tok_s: 247 load_s: 1.83 peak_vram_mib: 12813 max |diff| vs reference: 0.1875 KV cache compiled for 768 positionsLinnet, CUDA graphs247 tok/sKerasHub (JAX) ttft_ms: 24.59 decode_tok_s: 227 load_s: 16.97 peak_vram_mib: 10289 driver_vram_mib: 17024 max |diff| vs reference: 0.125 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub227 tok/sLinnet JAX (XLA, StableHLO) 1.12× the speed of KerasHub ttft_ms: 14.59 decode_tok_s: 254 load_s: 5.04 peak_vram_mib: 9408 driver_vram_mib: 17025 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)254 tok/sLinnet JAX (XLA, generated source) 1.12× the speed of KerasHub ttft_ms: 13.16 decode_tok_s: 255 load_s: 4.85 peak_vram_mib: 9408 driver_vram_mib: 17023 max |diff| vs reference: 0.1875 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)255 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 38.58 decode_tok_s: 59.43 load_s: 0.42 peak_vram_mib: 20806 max |diff| vs reference: 0.15512847900390625 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f3259 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 30.14 decode_tok_s: 68.74 load_s: 0.45 peak_vram_mib: 10870 max |diff| vs reference: 0.671875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f1669 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 32.10 decode_tok_s: 67.41 load_s: 0.42 peak_vram_mib: 10236 max |diff| vs reference: 0.5 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf1667 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 35.90 decode_tok_s: 45.59 load_s: 0.42 peak_vram_mib: 45882 max |diff| vs reference: 0.15331459045410156 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f3246 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 19.46 decode_tok_s: 72.17 load_s: 0.41 peak_vram_mib: 24196 max |diff| vs reference: 11.4375 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f1672 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 19.33 decode_tok_s: 69.74 load_s: 0.43 peak_vram_mib: 23654 max |diff| vs reference: 0.40625 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf1670 tok/svLLM ttft_ms: 11.50 decode_tok_s: 230 load_s: 95.50 peak_vram_mib: 69633 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM230 tok/sLinnet -> vLLM 1.02× the speed of vLLM ttft_ms: 7.47 decode_tok_s: 236 load_s: 68.39 peak_vram_mib: 69633 driver_vram_mib: 69633 linnet.hf.export (8 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export236 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 17.53 decode_tok_s: 250 load_s: 29.83 peak_vram_mib: 8755 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF250 tok/s
Mistral 7B Instruct v0.37.2 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 21.29 decode_tok_s: 48.70 load_s: 5.07 peak_vram_mib: 14845 max |diff| vs reference: 0transformers49 tok/stransformers (torch.compile, static cache) ttft_ms: 17.62 decode_tok_s: 118 load_s: 3.81 peak_vram_mib: 14858 max |diff| vs reference: 0.078125transformers, compiled118 tok/sLinnet torch (generated source) 0.64× the speed of transformers, compiled ttft_ms: 17.19 decode_tok_s: 75.25 load_s: 3.17 peak_vram_mib: 14752 max |diff| vs reference: 0.078125 KV cache compiled for 768 positionsLinnet, generated PyTorch75 tok/sLinnet torch (torch.compile, inductor) 1.21× the speed of transformers, compiled ttft_ms: 12.87 decode_tok_s: 143 load_s: 3.18 peak_vram_mib: 21817 max |diff| vs reference: 0.0859375 KV cache compiled for 768 positionsLinnet, inductor143 tok/sLinnet torch (CUDA graphs) 1.43× the speed of transformers, compiled ttft_ms: 12.72 decode_tok_s: 169 load_s: 2.94 peak_vram_mib: 21817 max |diff| vs reference: 0.0859375 KV cache compiled for 768 positionsLinnet, CUDA graphs169 tok/sKerasHub (JAX) ttft_ms: 27.39 decode_tok_s: 166 load_s: 35.10 peak_vram_mib: 14905 driver_vram_mib: 17020 max |diff| vs reference: 0.09375 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub166 tok/sLinnet JAX (XLA, StableHLO) 1.05× the speed of KerasHub ttft_ms: 16.00 decode_tok_s: 174 load_s: 10.19 peak_vram_mib: 14529 driver_vram_mib: 17021 max |diff| vs reference: 0.13037109375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)174 tok/sLinnet JAX (XLA, generated source) 1.05× the speed of KerasHub ttft_ms: 15.30 decode_tok_s: 174 load_s: 10.91 peak_vram_mib: 14529 driver_vram_mib: 17019 max |diff| vs reference: 0.1279296875 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)174 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 41.49 decode_tok_s: 55.87 load_s: 0.44 peak_vram_mib: 32940 max |diff| vs reference: 0.07344776391983032 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f3256 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 29.92 decode_tok_s: 75.15 load_s: 0.48 peak_vram_mib: 16906 max |diff| vs reference: 0.07275390625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f1675 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 30.52 decode_tok_s: 73.98 load_s: 0.45 peak_vram_mib: 15426 max |diff| vs reference: 0.15625 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf1674 tok/sLinnet ONNX f32 -> ONNX Runtime (TensorRT) ttft_ms: 50.23 decode_tok_s: 29.66 load_s: 0.45 peak_vram_mib: 75228 max |diff| vs reference: 0.06902849674224854 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f3230 tok/sLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 25.96 decode_tok_s: 51.88 load_s: 0.42 peak_vram_mib: 38750 max |diff| vs reference: 1.81640625 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f1652 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 26.96 decode_tok_s: 50.29 load_s: 0.45 peak_vram_mib: 37580 max |diff| vs reference: 0.1484375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf1650 tok/svLLM ttft_ms: 14.85 decode_tok_s: 158 load_s: 73.47 peak_vram_mib: 67805 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM158 tok/sLinnet -> vLLM 1.04× the speed of vLLM ttft_ms: 16.07 decode_tok_s: 166 load_s: 34.13 peak_vram_mib: 67805 driver_vram_mib: 67805 linnet.hf.export (14 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export166 tok/sSGLang ttft_ms: 19.37 decode_tok_s: 165 load_s: 59.07 peak_vram_mib: 69636 reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang165 tok/sLinnet -> SGLang 1.00× the speed of SGLang ttft_ms: 19.41 decode_tok_s: 165 load_s: 31.49 peak_vram_mib: 69636 driver_vram_mib: 69636 linnet.hf.export (15 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang on Linnet's export165 tok/sText Generation Inference ttft_ms: 23.00 decode_tok_s: 118 load_s: 52.11 peak_vram_mib: 60643 over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI118 tok/sLinnet -> Text Generation Inference 1.03× the speed of TGI ttft_ms: 22.63 decode_tok_s: 121 load_s: 36.07 peak_vram_mib: 60643 driver_vram_mib: 60643 linnet.hf.export (17 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI on Linnet's export121 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 21.11 decode_tok_s: 179 load_s: 108 peak_vram_mib: 14457 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF179 tok/s
Llama 3.1 8B Instruct8.0 B
PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUtransformers (eager) ttft_ms: 22.16 decode_tok_s: 51.15 load_s: 5.29 peak_vram_mib: 16431 max |diff| vs reference: 0transformers51 tok/stransformers (torch.compile, static cache) ttft_ms: 18.61 decode_tok_s: 110 load_s: 4.29 peak_vram_mib: 16476 max |diff| vs reference: 0.0625transformers, compiled110 tok/sLinnet torch (generated source) 0.72× the speed of transformers, compiled ttft_ms: 17.40 decode_tok_s: 79.04 load_s: 3.59 peak_vram_mib: 16244 max |diff| vs reference: 0.09375 KV cache compiled for 768 positionsLinnet, generated PyTorch79 tok/sLinnet torch (torch.compile, inductor) 1.28× the speed of transformers, compiled ttft_ms: 13.12 decode_tok_s: 141 load_s: 3.35 peak_vram_mib: 23311 max |diff| vs reference: 0.0625 KV cache compiled for 768 positionsLinnet, inductor141 tok/sLinnet torch (CUDA graphs) 1.46× the speed of transformers, compiled ttft_ms: 13.04 decode_tok_s: 161 load_s: 3.10 peak_vram_mib: 23315 max |diff| vs reference: 0.0625 KV cache compiled for 768 positionsLinnet, CUDA graphs161 tok/sKerasHub (JAX) ttft_ms: 27.49 decode_tok_s: 160 load_s: 31.00 peak_vram_mib: 17765 driver_vram_mib: 33404 max |diff| vs reference: 0.09375 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub160 tok/sLinnet JAX (XLA, StableHLO) 1.04× the speed of KerasHub ttft_ms: 16.60 decode_tok_s: 167 load_s: 10.67 peak_vram_mib: 16703 driver_vram_mib: 33409 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)167 tok/sLinnet JAX (XLA, generated source) 1.04× the speed of KerasHub ttft_ms: 15.71 decode_tok_s: 166 load_s: 10.58 peak_vram_mib: 16703 driver_vram_mib: 33407 max |diff| vs reference: 0.09619140625 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)166 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 42.03 decode_tok_s: 53.67 load_s: 0.50 peak_vram_mib: 36012 max |diff| vs reference: 0.06537771224975586 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f3254 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 30.48 decode_tok_s: 71.00 load_s: 0.51 peak_vram_mib: 17930 max |diff| vs reference: 0.0703125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f1671 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 31.13 decode_tok_s: 71.79 load_s: 0.48 peak_vram_mib: 16920 max |diff| vs reference: 0.09375 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf1672 tok/slinnet-onnx-f32-trt peak_vram_mib: 67034 failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_3492707504Linnet, TensorRT f32failedLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 28.63 decode_tok_s: 47.46 load_s: 0.49 peak_vram_mib: 42964 max |diff| vs reference: 9.7734375 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f1647 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 28.49 decode_tok_s: 49.49 load_s: 0.49 peak_vram_mib: 40818 max |diff| vs reference: 0.125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf1649 tok/svLLM ttft_ms: 16.50 decode_tok_s: 152 load_s: 74.59 peak_vram_mib: 69061 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM152 tok/sLinnet -> vLLM 1.03× the speed of vLLM ttft_ms: 12.24 decode_tok_s: 157 load_s: 39.27 peak_vram_mib: 69061 driver_vram_mib: 69061 linnet.hf.export (16 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export157 tok/sSGLang ttft_ms: 21.04 decode_tok_s: 158 load_s: 55.73 peak_vram_mib: 69640 reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang158 tok/sLinnet -> SGLang 1.00× the speed of SGLang ttft_ms: 19.89 decode_tok_s: 158 load_s: 38.44 peak_vram_mib: 69576 driver_vram_mib: 69576 linnet.hf.export (18 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a needSGLang on Linnet's export158 tok/sText Generation Inference ttft_ms: 27.90 decode_tok_s: 112 load_s: 44.11 peak_vram_mib: 60613 over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI112 tok/sLinnet -> Text Generation Inference 1.02× the speed of TGI ttft_ms: 25.69 decode_tok_s: 115 load_s: 40.11 peak_vram_mib: 60613 driver_vram_mib: 60613 linnet.hf.export (19 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a needTGI on Linnet's export115 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 21.95 decode_tok_s: 171 load_s: 75.53 peak_vram_mib: 15353 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF171 tok/svLLM (tensor parallel on 2 GPUs) ttft_ms: 11.87 decode_tok_s: 233 load_s: 286 peak_vram_mib: 142714 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM, tensor parallel233 tok/sLinnet torch (tensor parallel on 2 GPUs, shards) 1.02× the speed of vLLM, tensor parallel ttft_ms: 10.85 decode_tok_s: 238 load_s: 4.77 peak_vram_mib: 27757 one process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positionsLinnet, tensor parallel (PyTorch)238 tok/sLinnet JAX (XLA, tensor parallel on 2 GPUs) 0.78× the speed of vLLM, tensor parallel ttft_ms: 16.22 decode_tok_s: 181 load_s: 16.07 peak_vram_mib: 18801 driver_vram_mib: 38734 max |diff| vs reference: 0.125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, tensor parallel (XLA)181 tok/sLinnet torch (all GPUs) ttft_ms: 22.51 decode_tok_s: 87.83 load_s: 7.00 peak_vram_mib: 23878 max |diff| vs reference: 0.09375 KV cache compiled for 768 positionsLinnet, layers on two GPUs88 tok/sLinnet torch (offloaded) ttft_ms: 180 decode_tok_s: 5.73 load_s: 19.61 peak_vram_mib: 9150 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; cuda:0: embedding, layers.0-13; host, streamed in: layers.14-31, norm, lm_headLinnet, offloaded to host6 tok/s
Qwen3 8B8.2 B
PyTorchJAXONNX Runtime and TensorRTLLM enginesBeyond one GPUtransformers (eager) ttft_ms: 23.71 decode_tok_s: 48.62 load_s: 3.85 peak_vram_mib: 16774 max |diff| vs reference: 0transformers49 tok/stransformers (torch.compile, static cache) ttft_ms: 16.63 decode_tok_s: 107 load_s: 3.66 peak_vram_mib: 16832 max |diff| vs reference: 0.09375transformers, compiled107 tok/sLinnet torch (generated source) 0.64× the speed of transformers, compiled ttft_ms: 19.22 decode_tok_s: 68.79 load_s: 3.06 peak_vram_mib: 16564 max |diff| vs reference: 0.1171875 KV cache compiled for 768 positionsLinnet, generated PyTorch69 tok/sLinnet torch (torch.compile, inductor) 1.23× the speed of transformers, compiled ttft_ms: 13.18 decode_tok_s: 132 load_s: 3.44 peak_vram_mib: 23379 max |diff| vs reference: 0.125 KV cache compiled for 768 positionsLinnet, inductor132 tok/sLinnet torch (CUDA graphs) 1.43× the speed of transformers, compiled ttft_ms: 12.72 decode_tok_s: 154 load_s: 3.46 peak_vram_mib: 23379 max |diff| vs reference: 0.125 KV cache compiled for 768 positionsLinnet, CUDA graphs154 tok/sKerasHub (JAX) ttft_ms: 29.06 decode_tok_s: 149 load_s: 30.46 peak_vram_mib: 18055 driver_vram_mib: 33408 max |diff| vs reference: 0.09375 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub149 tok/sLinnet JAX (XLA, StableHLO) 1.06× the speed of KerasHub ttft_ms: 17.26 decode_tok_s: 158 load_s: 10.57 peak_vram_mib: 17205 driver_vram_mib: 33407 max |diff| vs reference: 0.125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)158 tok/sLinnet JAX (XLA, generated source) 1.06× the speed of KerasHub ttft_ms: 16.23 decode_tok_s: 159 load_s: 10.38 peak_vram_mib: 17205 driver_vram_mib: 33405 max |diff| vs reference: 0.09375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)159 tok/sLinnet ONNX f32 -> ONNX Runtime (CUDA) ttft_ms: 49.84 decode_tok_s: 46.71 load_s: 0.58 peak_vram_mib: 37192 max |diff| vs reference: 0.08458328247070312 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f3247 tok/sLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 36.01 decode_tok_s: 59.29 load_s: 0.58 peak_vram_mib: 19064 max |diff| vs reference: 1.3671875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime f1659 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 37.05 decode_tok_s: 58.40 load_s: 0.61 peak_vram_mib: 17350 max |diff| vs reference: 2.5 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, ONNX Runtime bf1658 tok/slinnet-onnx-f32-trt peak_vram_mib: 68692 failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_1765021298Linnet, TensorRT f32failedLinnet ONNX f16 -> ONNX Runtime (TensorRT) ttft_ms: 28.89 decode_tok_s: 46.79 load_s: 0.56 peak_vram_mib: 42760 max |diff| vs reference: 11.3125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT f1647 tok/sLinnet ONNX bf16 -> ONNX Runtime (TensorRT) ttft_ms: 27.86 decode_tok_s: 47.33 load_s: 0.56 peak_vram_mib: 41426 max |diff| vs reference: 1.59375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmaxLinnet, TensorRT bf1647 tok/svLLM ttft_ms: 16.31 decode_tok_s: 147 load_s: 88.96 peak_vram_mib: 69535 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM147 tok/sLinnet -> vLLM 1.04× the speed of vLLM ttft_ms: 13.28 decode_tok_s: 153 load_s: 67.16 peak_vram_mib: 69535 driver_vram_mib: 69535 linnet.hf.export (17 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM on Linnet's export153 tok/sLinnet -> llama.cpp (GGUF, f16) ttft_ms: 25.15 decode_tok_s: 165 load_s: 60.43 peak_vram_mib: 15527 linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt ratellama.cpp on Linnet's GGUF165 tok/svLLM (tensor parallel on 2 GPUs) ttft_ms: 14.15 decode_tok_s: 220 load_s: 78.78 peak_vram_mib: 143660 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM, tensor parallel220 tok/sLinnet torch (tensor parallel on 2 GPUs, shards) 0.96× the speed of vLLM, tensor parallel ttft_ms: 11.70 decode_tok_s: 212 load_s: 4.97 peak_vram_mib: 27039 one process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positionsLinnet, tensor parallel (PyTorch)212 tok/sLinnet JAX (XLA, tensor parallel on 2 GPUs) 0.59× the speed of vLLM, tensor parallel ttft_ms: 18.37 decode_tok_s: 131 load_s: 16.26 peak_vram_mib: 19832 driver_vram_mib: 39210 max |diff| vs reference: 0.125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, tensor parallel (XLA)131 tok/sLinnet torch (all GPUs) ttft_ms: 19.60 decode_tok_s: 68.18 load_s: 7.17 peak_vram_mib: 17082 max |diff| vs reference: 0.1171875 KV cache compiled for 768 positionsLinnet, layers on two GPUs68 tok/sLinnet torch (offloaded) ttft_ms: 191 decode_tok_s: 5.44 load_s: 19.87 peak_vram_mib: 9174 max |diff| vs reference: 0.1171875 KV cache compiled for 768 positions; cuda:0: embedding, layers.0-14; host, streamed in: layers.15-35, norm, lm_headLinnet, offloaded to host5 tok/s
GPT-OSS 20B20.9 B
PyTorchJAXONNX Runtime and TensorRTLLM enginestransformers (eager) ttft_ms: 41.88 decode_tok_s: 44.90 load_s: 115 peak_vram_mib: 41386 max |diff| vs reference: 0 eager attention: no SDPA for this architecturetransformers45 tok/stransformers (torch.compile, static cache) ttft_ms: 25.57 decode_tok_s: 100 load_s: 126 peak_vram_mib: 41286 max |diff| vs reference: 0.1875 eager attention: no SDPA for this architecturetransformers, compiled100 tok/sLinnet torch (generated source) 0.25× the speed of transformers, compiled ttft_ms: 89.99 decode_tok_s: 25.48 load_s: 2.54 peak_vram_mib: 54578 max |diff| vs reference: 0.1953125 KV cache compiled for 768 positionsLinnet, generated PyTorch25 tok/sLinnet torch (torch.compile, inductor) 1.99× the speed of transformers, compiled ttft_ms: 34.82 decode_tok_s: 200 load_s: 3.11 peak_vram_mib: 29393 max |diff| vs reference: 0.1875 KV cache compiled for 768 positionsLinnet, inductor200 tok/sLinnet torch (CUDA graphs) 3.67× the speed of transformers, compiled ttft_ms: 14.11 decode_tok_s: 368 load_s: 3.21 peak_vram_mib: 29443 max |diff| vs reference: 0.1875 KV cache compiled for 768 positionsLinnet, CUDA graphs368 tok/sKerasHub (JAX) ttft_ms: 72.84 decode_tok_s: 67.37 load_s: 2105 peak_vram_mib: 45163 driver_vram_mib: 61452 max |diff| vs reference: 8.9140625 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub67 tok/sLinnet JAX (XLA, StableHLO) 2.09× the speed of KerasHub ttft_ms: 122 decode_tok_s: 141 load_s: 8.61 peak_vram_mib: 15212 driver_vram_mib: 33407 max |diff| vs reference: 0.203125 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)141 tok/sLinnet JAX (XLA, generated source) 4.43× the speed of KerasHub ttft_ms: 25.16 decode_tok_s: 298 load_s: 9.20 peak_vram_mib: 51298 driver_vram_mib: 66714 max |diff| vs reference: 0.109375 KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)298 tok/slinnet-onnx-f32 peak_vram_mib: 79578 failed: RuntimeError: Error in execution: Non-zero status code returned while running Mul node. Name:'' Status Message: /onnxruntime_src/onnxruntime/core/framework/bfc_Linnet, ONNX Runtime f32failedLinnet ONNX f16 -> ONNX Runtime (CUDA) ttft_ms: 107 decode_tok_s: 21.42 load_s: 0.24 peak_vram_mib: 70019 max |diff| vs reference: 0.1923828125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the argmax taken in the graph, each step replayed as a CUDA graph on the CUDA providerLinnet, ONNX Runtime f1621 tok/sLinnet ONNX bf16 -> ONNX Runtime (CUDA) ttft_ms: 113 decode_tok_s: 21.22 load_s: 0.21 peak_vram_mib: 69495 max |diff| vs reference: 0.1875 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the argmax taken in the graph, each step replayed as a CUDA graph on the CUDA providerLinnet, ONNX Runtime bf1621 tok/slinnet-onnx-f32-trt failed: exit -9: [5] Failed to import initializer: In node -1 with name: and operator: (parseGraph): UNSUPPORTED_NODE: Assertion failed: ctx->getWeightsContext().convLinnet, TensorRT f32failedlinnet-onnx-f16-trt peak_vram_mib: 70006 failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_8376132635Linnet, TensorRT f16failedlinnet-onnx-bf16-trt peak_vram_mib: 69482 failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_1106551924Linnet, TensorRT bf16failedvLLM ttft_ms: 17.84 decode_tok_s: 299 load_s: 50.23 peak_vram_mib: 70831 reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a needvLLM299 tok/s

Serving ​

best in the chart Linnet existing stacks
GPT-2 (124M)124 M
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 25858 serve_s: 1.27 load_s: 22.21 peak_vram_mib: 69407 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM25858 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.48× the speed of vLLM serve_tok_s: 38317 serve_s: 0.86 load_s: 4.37 serve_ttft_ms: 364 peak_vram_mib: 2708 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs38317 tok/sLinnet jax (linnet.serve, XLA) 1.36× the speed of vLLM serve_tok_s: 35104 serve_s: 0.93 load_s: 2.51 serve_ttft_ms: 469 peak_vram_mib: 2056 driver_vram_mib: 4756 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA35104 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.87× the speed of vLLM serve_tok_s: 22510 serve_s: 1.46 load_s: 0.24 serve_ttft_ms: 640 peak_vram_mib: 11512 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime22510 tok/sLinnet -> vLLM (offline, continuous batching) 1.08× the speed of vLLM serve_tok_s: 27904 serve_s: 1.17 load_s: 19.21 peak_vram_mib: 69415 driver_vram_mib: 69415 linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export27904 tok/sSGLang (offline, continuous batching) serve_tok_s: 28788 serve_s: 1.14 load_s: 31.08 peak_vram_mib: 70022 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang28788 tok/sLinnet -> SGLang (offline, continuous batching) 0.94× the speed of SGLang serve_tok_s: 27129 serve_s: 1.21 load_s: 25.91 peak_vram_mib: 70022 driver_vram_mib: 70022 linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang on Linnet's export27129 tok/sText Generation Inference (continuous batching) serve_tok_s: 6189 serve_s: 5.13 load_s: 32.07 serve_ttft_ms: 1721 peak_vram_mib: 62205 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI6189 tok/sLinnet -> Text Generation Inference (continuous batching) 0.98× the speed of TGI serve_tok_s: 6066 serve_s: 5.25 load_s: 28.07 serve_ttft_ms: 1882 peak_vram_mib: 62321 driver_vram_mib: 62321 linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI on Linnet's export6066 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 3857 serve_s: 8.50 load_s: 43.05 serve_ttft_ms: 4635 peak_vram_mib: 69970 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend3857 tok/sTriton Inference Server (Python backend, linnet.serve) 3.67× the speed of Triton, vLLM backend serve_tok_s: 14166 serve_s: 2.31 load_s: 19.05 serve_ttft_ms: 1303 peak_vram_mib: 3416 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend14166 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 2951 serve_s: 11.11 load_s: 1.28 peak_vram_mib: 80574 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched2951 tok/sKerasHub (JAX, static batches) serve_tok_s: 2605 serve_s: 12.58 load_s: 6.15 peak_vram_mib: 3192 driver_vram_mib: 8818 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches2605 tok/s
Qwen2.5 0.5B Instruct494 M
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 18828 serve_s: 1.74 load_s: 18.31 peak_vram_mib: 69679 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM18828 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.84× the speed of vLLM serve_tok_s: 34643 serve_s: 0.95 load_s: 3.90 serve_ttft_ms: 403 peak_vram_mib: 3065 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs34643 tok/sLinnet jax (linnet.serve, XLA) 1.54× the speed of vLLM serve_tok_s: 29018 serve_s: 1.13 load_s: 2.75 serve_ttft_ms: 571 peak_vram_mib: 2149 driver_vram_mib: 4792 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA29018 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.70× the speed of vLLM serve_tok_s: 13227 serve_s: 2.48 load_s: 0.25 serve_ttft_ms: 1105 peak_vram_mib: 10372 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime13227 tok/sLinnet -> vLLM (offline, continuous batching) 1.39× the speed of vLLM serve_tok_s: 26248 serve_s: 1.25 load_s: 29.65 peak_vram_mib: 69351 driver_vram_mib: 69351 linnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export26248 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2157 serve_s: 15.19 load_s: 52.06 serve_ttft_ms: 7317 peak_vram_mib: 69888 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2157 tok/sTriton Inference Server (Python backend, linnet.serve) 11× the speed of Triton, vLLM backend serve_tok_s: 24211 serve_s: 1.35 load_s: 621 serve_ttft_ms: 559 peak_vram_mib: 4092 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend24211 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 1865 serve_s: 17.57 load_s: 1.34 peak_vram_mib: 78632 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched1865 tok/sKerasHub (JAX, static batches) serve_tok_s: 2482 serve_s: 13.20 load_s: 6.51 peak_vram_mib: 2751 driver_vram_mib: 4722 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches2482 tok/s
TinyLlama 1.1B Chat v1.01.1 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 22674 serve_s: 1.45 load_s: 24.20 peak_vram_mib: 69769 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM22674 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.11× the speed of vLLM serve_tok_s: 25091 serve_s: 1.31 load_s: 4.11 serve_ttft_ms: 583 peak_vram_mib: 5348 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs25091 tok/sLinnet jax (linnet.serve, XLA) 0.93× the speed of vLLM serve_tok_s: 21035 serve_s: 1.56 load_s: 3.85 serve_ttft_ms: 751 peak_vram_mib: 3669 driver_vram_mib: 4776 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA21035 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.47× the speed of vLLM serve_tok_s: 10580 serve_s: 3.10 load_s: 0.25 serve_ttft_ms: 1467 peak_vram_mib: 13962 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime10580 tok/sLinnet -> vLLM (offline, continuous batching) 1.03× the speed of vLLM serve_tok_s: 23257 serve_s: 1.41 load_s: 22.84 peak_vram_mib: 69921 driver_vram_mib: 69921 linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export23257 tok/sSGLang (offline, continuous batching) serve_tok_s: 20305 serve_s: 1.61 load_s: 36.16 peak_vram_mib: 70446 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang20305 tok/sLinnet -> SGLang (offline, continuous batching) 0.99× the speed of SGLang serve_tok_s: 20053 serve_s: 1.63 load_s: 29.93 peak_vram_mib: 70446 driver_vram_mib: 70446 linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang on Linnet's export20053 tok/sText Generation Inference (continuous batching) serve_tok_s: 3767 serve_s: 8.16 load_s: 32.08 serve_ttft_ms: 2482 peak_vram_mib: 62371 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI3767 tok/sLinnet -> Text Generation Inference (continuous batching) 1.03× the speed of TGI serve_tok_s: 3880 serve_s: 7.92 load_s: 30.07 serve_ttft_ms: 2260 peak_vram_mib: 62647 driver_vram_mib: 62647 linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI on Linnet's export3880 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2113 serve_s: 15.51 load_s: 47.06 serve_ttft_ms: 6590 peak_vram_mib: 69758 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2113 tok/sTriton Inference Server (Python backend, linnet.serve) 8.78× the speed of Triton, vLLM backend serve_tok_s: 18549 serve_s: 1.77 load_s: 652 serve_ttft_ms: 710 peak_vram_mib: 5714 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend18549 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 1672 serve_s: 19.59 load_s: 1.42 peak_vram_mib: 80886 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched1672 tok/sKerasHub (JAX, static batches) serve_tok_s: 1387 serve_s: 23.62 load_s: 8.64 peak_vram_mib: 4808 driver_vram_mib: 8816 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches1387 tok/s
SmolLM2 1.7B Instruct1.7 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 12046 serve_s: 2.72 load_s: 24.71 peak_vram_mib: 69793 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM12046 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.08× the speed of vLLM serve_tok_s: 12997 serve_s: 2.52 load_s: 4.32 serve_ttft_ms: 1082 peak_vram_mib: 14217 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs12997 tok/sLinnet jax (linnet.serve, XLA) 0.82× the speed of vLLM serve_tok_s: 9922 serve_s: 3.30 load_s: 4.09 serve_ttft_ms: 1438 peak_vram_mib: 13345 driver_vram_mib: 17066 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA9922 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.53× the speed of vLLM serve_tok_s: 6333 serve_s: 5.17 load_s: 0.26 serve_ttft_ms: 2382 peak_vram_mib: 28756 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime6333 tok/sLinnet -> vLLM (offline, continuous batching) 1.01× the speed of vLLM serve_tok_s: 12149 serve_s: 2.70 load_s: 22.07 peak_vram_mib: 70301 driver_vram_mib: 70301 linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export12149 tok/sSGLang (offline, continuous batching) serve_tok_s: 10524 serve_s: 3.11 load_s: 38.54 peak_vram_mib: 70782 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang10524 tok/sLinnet -> SGLang (offline, continuous batching) 1.02× the speed of SGLang serve_tok_s: 10734 serve_s: 3.05 load_s: 32.91 peak_vram_mib: 70782 driver_vram_mib: 70782 linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang on Linnet's export10734 tok/sText Generation Inference (continuous batching) serve_tok_s: 3574 serve_s: 8.35 load_s: 32.07 serve_ttft_ms: 2395 peak_vram_mib: 62357 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI3574 tok/sLinnet -> Text Generation Inference (continuous batching) 1.00× the speed of TGI serve_tok_s: 3578 serve_s: 8.37 load_s: 28.07 serve_ttft_ms: 2438 peak_vram_mib: 62475 driver_vram_mib: 62475 linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI on Linnet's export3578 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2199 serve_s: 14.90 load_s: 49.06 serve_ttft_ms: 8071 peak_vram_mib: 69584 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2199 tok/sTriton Inference Server (Python backend, linnet.serve) 5.09× the speed of Triton, vLLM backend serve_tok_s: 11195 serve_s: 2.93 load_s: 813 serve_ttft_ms: 1172 peak_vram_mib: 14536 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend11195 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 1557 serve_s: 21.05 load_s: 1.67 peak_vram_mib: 81002 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched1557 tok/sKerasHub (JAX, static batches) serve_tok_s: 544 serve_s: 60.27 load_s: 9.12 peak_vram_mib: 18647 driver_vram_mib: 25202 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches544 tok/s
Phi-3 Mini 4K Instruct3.8 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 5583 serve_s: 5.87 load_s: 34.47 peak_vram_mib: 70001 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM5583 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.19× the speed of vLLM serve_tok_s: 6668 serve_s: 4.91 load_s: 4.80 serve_ttft_ms: 2131 peak_vram_mib: 24599 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs6668 tok/sLinnet jax (linnet.serve, XLA) 0.93× the speed of vLLM serve_tok_s: 5178 serve_s: 6.33 load_s: 7.27 serve_ttft_ms: 2688 peak_vram_mib: 26643 driver_vram_mib: 33452 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA5178 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.63× the speed of vLLM serve_tok_s: 3526 serve_s: 9.29 load_s: 0.26 serve_ttft_ms: 4344 peak_vram_mib: 49972 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime3526 tok/sLinnet -> vLLM (offline, continuous batching) 1.00× the speed of vLLM serve_tok_s: 5587 serve_s: 5.86 load_s: 32.60 peak_vram_mib: 70001 driver_vram_mib: 70001 linnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export5587 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 1968 serve_s: 16.65 load_s: 51.06 serve_ttft_ms: 7875 peak_vram_mib: 70432 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend1968 tok/sTriton Inference Server (Python backend, linnet.serve) 3.22× the speed of Triton, vLLM backend serve_tok_s: 6328 serve_s: 5.18 load_s: 290 serve_ttft_ms: 2124 peak_vram_mib: 25540 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend6328 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 521 serve_s: 62.90 load_s: 2.37 peak_vram_mib: 80980 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched521 tok/s
Qwen3 4B4.0 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 8048 serve_s: 4.07 load_s: 58.22 peak_vram_mib: 70063 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM8048 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.02× the speed of vLLM serve_tok_s: 8196 serve_s: 4.00 load_s: 5.11 serve_ttft_ms: 1666 peak_vram_mib: 18541 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs8196 tok/sLinnet jax (linnet.serve, XLA) 0.84× the speed of vLLM serve_tok_s: 6786 serve_s: 4.83 load_s: 8.40 serve_ttft_ms: 2136 peak_vram_mib: 15612 driver_vram_mib: 17104 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA6786 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.52× the speed of vLLM serve_tok_s: 4203 serve_s: 7.80 load_s: 0.26 serve_ttft_ms: 3644 peak_vram_mib: 31654 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime4203 tok/sLinnet -> vLLM (offline, continuous batching) 1.01× the speed of vLLM serve_tok_s: 8158 serve_s: 4.02 load_s: 50.15 peak_vram_mib: 70063 driver_vram_mib: 70063 linnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export8158 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2082 serve_s: 15.74 load_s: 72.15 serve_ttft_ms: 5981 peak_vram_mib: 69532 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2082 tok/sTriton Inference Server (Python backend, linnet.serve) 3.74× the speed of Triton, vLLM backend serve_tok_s: 7786 serve_s: 4.21 load_s: 659 serve_ttft_ms: 1747 peak_vram_mib: 18998 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend7786 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 850 serve_s: 38.57 load_s: 2.31 peak_vram_mib: 81030 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched850 tok/sKerasHub (JAX, static batches) serve_tok_s: 384 serve_s: 85.41 load_s: 17.62 peak_vram_mib: 21641 driver_vram_mib: 33398 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches384 tok/s
Mistral 7B Instruct v0.37.2 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 5546 serve_s: 5.91 load_s: 78.80 peak_vram_mib: 70410 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM5546 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.04× the speed of vLLM serve_tok_s: 5764 serve_s: 5.68 load_s: 8.01 serve_ttft_ms: 2447 peak_vram_mib: 27929 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs5764 tok/sLinnet jax (linnet.serve, XLA) 0.86× the speed of vLLM serve_tok_s: 4796 serve_s: 6.83 load_s: 10.90 serve_ttft_ms: 3064 peak_vram_mib: 21097 driver_vram_mib: 33466 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA4796 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.65× the speed of vLLM serve_tok_s: 3617 serve_s: 9.06 load_s: 0.25 serve_ttft_ms: 4283 peak_vram_mib: 36408 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime3617 tok/sLinnet -> vLLM (offline, continuous batching) 1.05× the speed of vLLM serve_tok_s: 5849 serve_s: 5.60 load_s: 29.35 peak_vram_mib: 70339 driver_vram_mib: 70339 linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export5849 tok/sSGLang (offline, continuous batching) serve_tok_s: 4906 serve_s: 6.68 load_s: 49.68 peak_vram_mib: 71674 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang4906 tok/sLinnet -> SGLang (offline, continuous batching) 1.00× the speed of SGLang serve_tok_s: 4927 serve_s: 6.65 load_s: 36.52 peak_vram_mib: 71674 driver_vram_mib: 71674 linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang on Linnet's export4927 tok/sText Generation Inference (continuous batching) serve_tok_s: 1664 serve_s: 6.84 load_s: 42.09 serve_ttft_ms: 1650 peak_vram_mib: 62521 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI1664 tok/sLinnet -> Text Generation Inference (continuous batching) 1.10× the speed of TGI serve_tok_s: 1826 serve_s: 6.32 load_s: 36.08 serve_ttft_ms: 1408 peak_vram_mib: 62647 driver_vram_mib: 62647 linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI on Linnet's export1826 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2071 serve_s: 15.83 load_s: 58.07 serve_ttft_ms: 5191 peak_vram_mib: 70840 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2071 tok/sTriton Inference Server (Python backend, linnet.serve) 2.65× the speed of Triton, vLLM backend serve_tok_s: 5487 serve_s: 5.97 load_s: 652 serve_ttft_ms: 2513 peak_vram_mib: 27538 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend5487 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 915 serve_s: 35.83 load_s: 3.75 peak_vram_mib: 81042 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched915 tok/sKerasHub (JAX, static batches) serve_tok_s: 394 serve_s: 83.23 load_s: 37.68 peak_vram_mib: 24669 driver_vram_mib: 33394 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches394 tok/s
Llama 3.1 8B Instruct8.0 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 5449 serve_s: 6.01 load_s: 135 peak_vram_mib: 70326 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM5449 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.03× the speed of vLLM serve_tok_s: 5600 serve_s: 5.85 load_s: 6.99 serve_ttft_ms: 2508 peak_vram_mib: 28440 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs5600 tok/sLinnet jax (linnet.serve, XLA) 0.87× the speed of vLLM serve_tok_s: 4734 serve_s: 6.92 load_s: 12.02 serve_ttft_ms: 3089 peak_vram_mib: 22581 driver_vram_mib: 33466 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA4734 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.65× the speed of vLLM serve_tok_s: 3558 serve_s: 9.21 load_s: 0.26 serve_ttft_ms: 4341 peak_vram_mib: 38456 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime3558 tok/sLinnet -> vLLM (offline, continuous batching) 1.03× the speed of vLLM serve_tok_s: 5607 serve_s: 5.84 load_s: 34.70 peak_vram_mib: 70317 driver_vram_mib: 70317 linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export5607 tok/sSGLang (offline, continuous batching) serve_tok_s: 4776 serve_s: 6.86 load_s: 55.80 peak_vram_mib: 71690 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang4776 tok/sLinnet -> SGLang (offline, continuous batching) 1.00× the speed of SGLang serve_tok_s: 4757 serve_s: 6.89 load_s: 44.91 peak_vram_mib: 71626 driver_vram_mib: 71626 linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85SGLang on Linnet's export4757 tok/sText Generation Inference (continuous batching) serve_tok_s: 2621 serve_s: 11.87 load_s: 42.14 serve_ttft_ms: 3609 peak_vram_mib: 62157 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI2621 tok/sLinnet -> Text Generation Inference (continuous batching) 0.98× the speed of TGI serve_tok_s: 2566 serve_s: 12.13 load_s: 40.07 serve_ttft_ms: 4158 peak_vram_mib: 62707 driver_vram_mib: 62707 linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85TGI on Linnet's export2566 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 2998 serve_s: 10.93 load_s: 61.10 serve_ttft_ms: 4545 peak_vram_mib: 69134 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend2998 tok/sTriton Inference Server (Python backend, linnet.serve) 1.79× the speed of Triton, vLLM backend serve_tok_s: 5377 serve_s: 6.09 load_s: 645 serve_ttft_ms: 2559 peak_vram_mib: 29020 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend5377 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 881 serve_s: 37.18 load_s: 3.98 peak_vram_mib: 81028 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched881 tok/sKerasHub (JAX, static batches) serve_tok_s: 391 serve_s: 83.81 load_s: 30.10 peak_vram_mib: 26425 driver_vram_mib: 33396 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches391 tok/s
Qwen3 8B8.2 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 5295 serve_s: 6.19 load_s: 78.52 peak_vram_mib: 70218 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM5295 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.02× the speed of vLLM serve_tok_s: 5399 serve_s: 6.07 load_s: 8.73 serve_ttft_ms: 2610 peak_vram_mib: 29190 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs5399 tok/sLinnet jax (linnet.serve, XLA) 0.85× the speed of vLLM serve_tok_s: 4485 serve_s: 7.31 load_s: 13.35 serve_ttft_ms: 3270 peak_vram_mib: 23622 driver_vram_mib: 33488 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA4485 tok/sLinnet onnx (linnet.serve, ONNX Runtime, f16) 0.61× the speed of vLLM serve_tok_s: 3245 serve_s: 10.10 load_s: 0.27 serve_ttft_ms: 4746 peak_vram_mib: 38498 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, ONNX Runtime3245 tok/sLinnet -> vLLM (offline, continuous batching) 1.04× the speed of vLLM serve_tok_s: 5524 serve_s: 5.93 load_s: 61.39 peak_vram_mib: 70221 driver_vram_mib: 70221 linnet.hf.export (0 s), then vLLM (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM on Linnet's export5524 tok/sTriton Inference Server (vLLM backend) serve_tok_s: 1970 serve_s: 16.63 load_s: 72.08 serve_ttft_ms: 6765 peak_vram_mib: 68762 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend1970 tok/sTriton Inference Server (Python backend, linnet.serve) 2.60× the speed of Triton, vLLM backend serve_tok_s: 5118 serve_s: 6.40 load_s: 853 serve_ttft_ms: 2685 peak_vram_mib: 30234 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend5118 tok/stransformers (generate_batch, continuous batching) serve_tok_s: 816 serve_s: 40.16 load_s: 3.51 peak_vram_mib: 80992 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestampstransformers, batched816 tok/sKerasHub (JAX, static batches) serve_tok_s: 353 serve_s: 92.76 load_s: 29.68 peak_vram_mib: 28195 driver_vram_mib: 33398 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub, static batches353 tok/s
GPT-OSS 20B20.9 B
Serving many requestsvLLM (offline, continuous batching) serve_tok_s: 4313 serve_s: 7.60 load_s: 79.97 peak_vram_mib: 70153 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85vLLM4313 tok/sLinnet torch (linnet.serve, CUDA graphs) 1.40× the speed of vLLM serve_tok_s: 6031 serve_s: 5.43 load_s: 7.28 serve_ttft_ms: 2302 peak_vram_mib: 31330 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positionslinnet.serve, CUDA graphs6031 tok/sLinnet jax (linnet.serve, XLA) 0.72× the speed of vLLM serve_tok_s: 3117 serve_s: 10.51 load_s: 11.52 serve_ttft_ms: 4654 peak_vram_mib: 53389 driver_vram_mib: 67891 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionslinnet.serve, XLA3117 tok/sserve-linnet-onnx peak_vram_mib: 81028 failed: RuntimeError: Error in execution: Non-zero status code returned while running Einsum node. Name:'' Status Message: /onnxruntime_src/onnxruntime/core/framework/blinnet.serve, ONNX RuntimefailedTriton Inference Server (vLLM backend) serve_tok_s: 1701 serve_s: 19.26 load_s: 106 serve_ttft_ms: 8389 peak_vram_mib: 69816 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, vLLM backend1701 tok/sTriton Inference Server (Python backend, linnet.serve) 3.43× the speed of Triton, vLLM backend serve_tok_s: 5839 serve_s: 5.61 load_s: 705 serve_ttft_ms: 2378 peak_vram_mib: 31914 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the clientTriton, linnet.serve backend5839 tok/sserve-transformers peak_vram_mib: 79444 failed: RuntimeError: generate_batch returned no tokens for any requesttransformers, batchedfailedserve-keras-hub peak_vram_mib: 61426 failed: JaxRuntimeError: NOT_FOUND: Failed to get configs for: 2 out of 100 instructions. See logs for all failures. Example failure: KerasHub, static batchesfailed

Encoders, vision, audio, and diffusion ​

best in the chart Linnet existing stacks
ResNet-1812 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 1.21 throughput_per_s: 18541 load_s: 3.95 peak_vram_mib: 928 max |diff| vs reference: 0transformers1.2 mstransformers (torch.compile) latency_ms: 0.59 throughput_per_s: 23490 load_s: 3.80 peak_vram_mib: 946 max |diff| vs reference: 0.03125transformers, compiled0.6 msLinnet torch (generated source) 0.57× the speed of transformers, compiled latency_ms: 1.04 throughput_per_s: 18580 load_s: 0.51 peak_vram_mib: 928 max |diff| vs reference: 0Linnet, generated PyTorch1.0 msLinnet torch (torch.compile, inductor) 0.94× the speed of transformers, compiled latency_ms: 0.63 throughput_per_s: 24100 load_s: 0.63 peak_vram_mib: 878 max |diff| vs reference: 0.03125Linnet, inductor0.6 msLinnet torch (CUDA graphs) 1.77× the speed of transformers, compiled latency_ms: 0.33 throughput_per_s: 25051 load_s: 0.62 peak_vram_mib: 968 max |diff| vs reference: 0.03125Linnet, CUDA graphs0.3 msLinnet JAX (XLA, StableHLO) latency_ms: 0.64 throughput_per_s: 32230 load_s: 0.37 peak_vram_mib: 444 driver_vram_mib: 1142 max |diff| vs reference: 0.03125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)0.6 msLinnet JAX (XLA, generated source) latency_ms: 1.10 throughput_per_s: 31053 load_s: 0.43 peak_vram_mib: 513 driver_vram_mib: 1142 max |diff| vs reference: 0.03125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1.1 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 0.65 throughput_per_s: 15125 load_s: 11.92 peak_vram_mib: 1312 max |diff| vs reference: 0.08936309814453125 the reference model's own ONNX export, f32 like Linnet's; torch.onnx export0.6 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.02× the speed of torch.onnx export latency_ms: 0.64 throughput_per_s: 14894 load_s: 0.30 peak_vram_mib: 1308 max |diff| vs reference: 0.08936357498168945 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f320.6 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 0.53 throughput_per_s: 19847 load_s: 0.29 peak_vram_mib: 1080 max |diff| vs reference: 0.1015625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f160.5 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.05 throughput_per_s: 11198 load_s: 0.31 peak_vram_mib: 1552 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf161.0 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.27 throughput_per_s: 36114 load_s: 0.29 peak_vram_mib: 2580 max |diff| vs reference: 0.08950281143188477 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f320.3 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.21 throughput_per_s: 59978 load_s: 0.29 peak_vram_mib: 2454 max |diff| vs reference: 0.103515625 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f160.2 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.45 throughput_per_s: 18561 load_s: 0.33 peak_vram_mib: 2638 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf160.5 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 3.10 throughput_per_s: 1454 load_s: 4.06 peak_vram_mib: 1541 max |diff| vs reference: 0.08936309814453125 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, torch.onnx export3.1 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 0.98× the speed of Triton, torch.onnx export latency_ms: 3.18 throughput_per_s: 1243 load_s: 9.08 peak_vram_mib: 1539 max |diff| vs reference: 0.08936357498168945 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, Linnet's ONNX3.2 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 2.76 throughput_per_s: 1220 load_s: 6.02 peak_vram_mib: 2436 max |diff| vs reference: 0.015625 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flightTriton, Linnet Python backend2.8 ms
all-MiniLM-L6-v223 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 1.91 throughput_per_s: 26734 load_s: 2.19 peak_vram_mib: 880 max |diff| vs reference: 0 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers1.9 mstransformers (torch.compile) latency_ms: 0.84 throughput_per_s: 50395 load_s: 1.49 peak_vram_mib: 892 max |diff| vs reference: 0.046875 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers, compiled0.8 mssentence-transformers (SentenceTransformer.encode) latency_ms: 3.21 throughput_per_s: 6431 load_s: 2.44 peak_vram_mib: 824 natural text tokenizing to 65 tokens, identical across the batch (no padding wasted); not the fixed-length random-token workload the other rows use, since this is the library's own string-in interfacesentence-transformers3.2 msLinnet torch (generated source) 0.82× the speed of transformers, compiled latency_ms: 1.03 throughput_per_s: 49772 load_s: 0.62 peak_vram_mib: 864 max |diff| vs reference: 0.0625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, generated PyTorch1.0 msLinnet torch (torch.compile, inductor) 1.07× the speed of transformers, compiled latency_ms: 0.79 throughput_per_s: 70709 load_s: 0.63 peak_vram_mib: 840 max |diff| vs reference: 0.0625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, inductor0.8 msLinnet torch (CUDA graphs) 2.61× the speed of transformers, compiled latency_ms: 0.32 throughput_per_s: 74093 load_s: 0.57 peak_vram_mib: 968 max |diff| vs reference: 0.0625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, CUDA graphs0.3 msKerasHub (JAX) latency_ms: 3.96 throughput_per_s: 11661 load_s: 6.72 peak_vram_mib: 255 driver_vram_mib: 880 max |diff| vs reference: 0.046875 bf16, under jax.jit; unpadded batches of exactly 128 tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub4.0 msLinnet JAX (XLA, StableHLO) 7.58× the speed of KerasHub latency_ms: 0.52 throughput_per_s: 38961 load_s: 0.33 peak_vram_mib: 347 driver_vram_mib: 1134 max |diff| vs reference: 0.046875 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)0.5 msLinnet JAX (XLA, generated source) 8.91× the speed of KerasHub latency_ms: 0.44 throughput_per_s: 39720 load_s: 0.29 peak_vram_mib: 313 driver_vram_mib: 1132 max |diff| vs reference: 0.046875 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)0.4 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 1.02 throughput_per_s: 19156 load_s: 12.43 peak_vram_mib: 1486 max |diff| vs reference: 0.06286907196044922 the reference model's own ONNX export, f32 like Linnet's; unpadded batches of exactly 128 tokenstorch.onnx export1.0 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 0.84× the speed of torch.onnx export latency_ms: 1.21 throughput_per_s: 20250 load_s: 0.29 peak_vram_mib: 1492 max |diff| vs reference: 0.06287145614624023 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f321.2 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.05 throughput_per_s: 34378 load_s: 0.29 peak_vram_mib: 1140 max |diff| vs reference: 0.064453125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f161.1 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.05 throughput_per_s: 34780 load_s: 0.28 peak_vram_mib: 1140 max |diff| vs reference: 0.03515625 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime bf161.0 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.35 throughput_per_s: 34138 load_s: 0.28 peak_vram_mib: 2556 max |diff| vs reference: 0.06383466720581055 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f320.3 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.23 throughput_per_s: 80632 load_s: 0.28 peak_vram_mib: 2464 max |diff| vs reference: 0.05859375 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f160.2 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.24 throughput_per_s: 74959 load_s: 0.27 peak_vram_mib: 2464 max |diff| vs reference: 0.046875 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT bf160.2 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 2.15 throughput_per_s: 1107 load_s: 6.18 peak_vram_mib: 1559 max |diff| vs reference: 0.06286907196044922 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, torch.onnx export2.1 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.00× the speed of Triton, torch.onnx export latency_ms: 2.15 throughput_per_s: 1111 load_s: 10.04 peak_vram_mib: 1591 max |diff| vs reference: 0.06287145614624023 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, Linnet's ONNX2.2 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 1.39 throughput_per_s: 1100 load_s: 6.03 peak_vram_mib: 2434 max |diff| vs reference: 0.0625 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 64 in flightTriton, Linnet Python backend1.4 ms
ResNet-5026 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 2.94 throughput_per_s: 7141 load_s: 7.20 peak_vram_mib: 1082 max |diff| vs reference: 0transformers2.9 mstransformers (torch.compile) latency_ms: 1.29 throughput_per_s: 11054 load_s: 4.09 peak_vram_mib: 1100 max |diff| vs reference: 0.0625transformers, compiled1.3 msLinnet torch (generated source) 0.61× the speed of transformers, compiled latency_ms: 2.13 throughput_per_s: 7066 load_s: 1.22 peak_vram_mib: 1032 max |diff| vs reference: 0Linnet, generated PyTorch2.1 msLinnet torch (torch.compile, inductor) 1.00× the speed of transformers, compiled latency_ms: 1.29 throughput_per_s: 11138 load_s: 1.39 peak_vram_mib: 1030 max |diff| vs reference: 0.0625Linnet, inductor1.3 msLinnet torch (CUDA graphs) 1.89× the speed of transformers, compiled latency_ms: 0.68 throughput_per_s: 11583 load_s: 1.44 peak_vram_mib: 1070 max |diff| vs reference: 0.0625Linnet, CUDA graphs0.7 msLinnet JAX (XLA, StableHLO) latency_ms: 0.93 throughput_per_s: 15014 load_s: 0.64 peak_vram_mib: 514 driver_vram_mib: 1134 max |diff| vs reference: 0.0625 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)0.9 msLinnet JAX (XLA, generated source) latency_ms: 1.17 throughput_per_s: 13563 load_s: 0.86 peak_vram_mib: 505 driver_vram_mib: 1134 max |diff| vs reference: 0.0625 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1.2 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 1.11 throughput_per_s: 5416 load_s: 15.54 peak_vram_mib: 2128 max |diff| vs reference: 0.2690114974975586 the reference model's own ONNX export, f32 like Linnet's; torch.onnx export1.1 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 0.87× the speed of torch.onnx export latency_ms: 1.28 throughput_per_s: 4510 load_s: 0.29 peak_vram_mib: 1852 max |diff| vs reference: 0.2690119743347168 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f321.3 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.00 throughput_per_s: 7354 load_s: 0.30 peak_vram_mib: 1420 max |diff| vs reference: 0.25390625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f161.0 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.66 throughput_per_s: 3698 load_s: 0.28 peak_vram_mib: 1962 max |diff| vs reference: 0.0625 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf161.7 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.53 throughput_per_s: 14519 load_s: 0.29 peak_vram_mib: 2816 max |diff| vs reference: 0.26560449600219727 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f320.5 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.41 throughput_per_s: 25153 load_s: 0.30 peak_vram_mib: 2572 max |diff| vs reference: 0.2578125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f160.4 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 1.00 throughput_per_s: 6078 load_s: 0.29 peak_vram_mib: 2896 max |diff| vs reference: 0.09375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf161.0 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 2.62 throughput_per_s: 1354 load_s: 5.07 peak_vram_mib: 2586 max |diff| vs reference: 0.2690114974975586 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, torch.onnx export2.6 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 0.79× the speed of Triton, torch.onnx export latency_ms: 3.33 throughput_per_s: 1308 load_s: 13.11 peak_vram_mib: 2171 max |diff| vs reference: 0.2690119743347168 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, Linnet's ONNX3.3 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 3.04 throughput_per_s: 1258 load_s: 6.02 peak_vram_mib: 2594 max |diff| vs reference: 0.0625 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flightTriton, Linnet Python backend3.0 ms
Whisper tiny38 Mtranscribe_ms, batch 1
PyTorchJAXONNX Runtime and TensorRTtransformers (eager) encode_ms: 2.53 transcribe_ms: 63.13 load_s: 1.91 peak_vram_mib: 934 max |diff| vs reference: 0transformers63.1 mstransformers (torch.compile, static cache) encode_ms: 2.18 transcribe_ms: 51.77 load_s: 0.67 peak_vram_mib: 1120 max |diff| vs reference: 0.0008020401000976562transformers, compiled51.8 msLinnet torch (generated source) 0.88× the speed of transformers, compiled encode_ms: 2.36 transcribe_ms: 58.75 load_s: 0.31 peak_vram_mib: 1064 max |diff| vs reference: 0.07889080047607422 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, generated PyTorch58.7 msLinnet torch (torch.compile, inductor) 1.51× the speed of transformers, compiled encode_ms: 1.99 transcribe_ms: 34.28 load_s: 0.30 peak_vram_mib: 1062 max |diff| vs reference: 0.07913780212402344 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, inductor34.3 msLinnet torch (CUDA graphs) 2.32× the speed of transformers, compiled encode_ms: 1.81 transcribe_ms: 22.35 load_s: 0.30 peak_vram_mib: 1184 max |diff| vs reference: 0.07913780212402344 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, CUDA graphs22.3 msLinnet JAX (XLA, StableHLO) encode_ms: 1.02 transcribe_ms: 16.40 load_s: 0.22 peak_vram_mib: 432 driver_vram_mib: 1207 max |diff| vs reference: 0.6507987976074219 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)16.4 msLinnet JAX (XLA, generated source) encode_ms: 1.01 transcribe_ms: 16.60 load_s: 0.22 peak_vram_mib: 432 driver_vram_mib: 1207 max |diff| vs reference: 0.6820688247680664 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)16.6 msLinnet ONNX f32 -> ONNX Runtime (CUDA) encode_ms: 1.70 transcribe_ms: 16.18 load_s: 0.24 peak_vram_mib: 1846 max |diff| vs reference: 0.7507038116455078 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime f3216.2 msLinnet ONNX f16 -> ONNX Runtime (CUDA) encode_ms: 1.64 transcribe_ms: 16.54 load_s: 0.23 peak_vram_mib: 1674 max |diff| vs reference: 0.7178497314453125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime f1616.5 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) encode_ms: 1.26 transcribe_ms: 17.27 load_s: 0.24 peak_vram_mib: 1552 max |diff| vs reference: 8.531454086303711 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime bf1617.3 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) encode_ms: 0.79 transcribe_ms: 24.30 load_s: 0.15 peak_vram_mib: 3418 max |diff| vs reference: 0.7840394973754883 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT f3224.3 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) encode_ms: 0.41 transcribe_ms: 20.40 load_s: 0.25 peak_vram_mib: 3142 max |diff| vs reference: 1.3493905067443848 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT f1620.4 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) encode_ms: 0.45 transcribe_ms: 20.37 load_s: 0.15 peak_vram_mib: 3080 max |diff| vs reference: 5.794703006744385 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT bf1620.4 ms
SD VAE ft-MSE (decoder)49 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTdiffusers (eager) latency_ms: 21.97 throughput_per_s: 66.23 load_s: 6.71 peak_vram_mib: 9160 max |diff| vs reference: 0 64x64 latent (512-pixel image), bf16; throughput at batch 8diffusers22.0 msdiffusers (torch.compile) latency_ms: 10.37 throughput_per_s: 110 load_s: 4.89 peak_vram_mib: 7574 max |diff| vs reference: 0.033203125 64x64 latent (512-pixel image), bf16; throughput at batch 8diffusers, compiled10.4 msLinnet torch (generated source) 0.38× the speed of diffusers, compiled latency_ms: 27.46 load_s: 0.87 throughput_per_s: 39.39 peak_vram_mib: 12516 max |diff| vs reference: 0.03125 64x64 latent (512-pixel image), bf16; throughput at batch 8Linnet, generated PyTorch27.5 msLinnet torch (torch.compile, inductor) 0.95× the speed of diffusers, compiled latency_ms: 10.90 load_s: 0.47 throughput_per_s: 111 peak_vram_mib: 4939 max |diff| vs reference: 0.033203125 64x64 latent (512-pixel image), bf16; throughput at batch 8Linnet, inductor10.9 msLinnet torch (CUDA graphs) 1.00× the speed of diffusers, compiled latency_ms: 10.35 load_s: 0.45 throughput_per_s: 112 peak_vram_mib: 5667 max |diff| vs reference: 0.033203125 64x64 latent (512-pixel image), bf16; throughput at batch 8Linnet, CUDA graphs10.4 msLinnet JAX (XLA, StableHLO) latency_ms: 7.47 throughput_per_s: 175 load_s: 0.66 peak_vram_mib: 3377 driver_vram_mib: 8822 max |diff| vs reference: 0.044921875 64x64 latent (512-pixel image), bf16; throughput at batch 8; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)7.5 msLinnet JAX (XLA, generated source) latency_ms: 7.41 throughput_per_s: 173 load_s: 0.50 peak_vram_mib: 3377 driver_vram_mib: 8822 max |diff| vs reference: 0.052734375 64x64 latent (512-pixel image), bf16; throughput at batch 8; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)7.4 msLinnet ONNX f32 -> ONNX Runtime (CUDA) latency_ms: 24.70 throughput_per_s: 49.05 load_s: 0.27 peak_vram_mib: 26942 max |diff| vs reference: 0.03966502845287323 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f3224.7 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 20.96 throughput_per_s: 68.99 load_s: 0.24 peak_vram_mib: 13338 max |diff| vs reference: 0.0379638671875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f1621.0 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 28.11 throughput_per_s: 42.06 load_s: 0.24 peak_vram_mib: 21742 max |diff| vs reference: 0.056640625 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf1628.1 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 13.72 throughput_per_s: 71.17 load_s: 0.24 peak_vram_mib: 11810 max |diff| vs reference: 0.03982146084308624 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f3213.7 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 5.41 throughput_per_s: 199 load_s: 0.24 peak_vram_mib: 8788 max |diff| vs reference: 0.043212890625 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f165.4 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 12.52 throughput_per_s: 78.14 load_s: 0.24 peak_vram_mib: 10708 max |diff| vs reference: 0.03515625 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf1612.5 ms
ViT-Base/16 22487 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 2.31 throughput_per_s: 8811 load_s: 9.36 peak_vram_mib: 1060 max |diff| vs reference: 0transformers2.3 mstransformers (torch.compile) latency_ms: 1.97 throughput_per_s: 4916 load_s: 3.83 peak_vram_mib: 1166 max |diff| vs reference: 0.0234375transformers, compiled2.0 msLinnet torch (generated source) 0.91× the speed of transformers, compiled latency_ms: 2.16 throughput_per_s: 8215 load_s: 1.06 peak_vram_mib: 1042 max |diff| vs reference: 0.03125Linnet, generated PyTorch2.2 msLinnet torch (torch.compile, inductor) 1.17× the speed of transformers, compiled latency_ms: 1.68 throughput_per_s: 9646 load_s: 1.70 peak_vram_mib: 1062 max |diff| vs reference: 0.01953125Linnet, inductor1.7 msLinnet torch (CUDA graphs) 2.42× the speed of transformers, compiled latency_ms: 0.81 throughput_per_s: 10007 load_s: 1.17 peak_vram_mib: 1190 max |diff| vs reference: 0.01953125Linnet, CUDA graphs0.8 msKerasHub (JAX) latency_ms: 5.61 throughput_per_s: 1814 load_s: 5.27 peak_vram_mib: 415 driver_vram_mib: 1132 max |diff| vs reference: 0.0625 bf16, under jax.jit; ; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub5.6 msLinnet JAX (XLA, StableHLO) 4.09× the speed of KerasHub latency_ms: 1.37 throughput_per_s: 6835 load_s: 0.56 peak_vram_mib: 640 driver_vram_mib: 1648 max |diff| vs reference: 0.03125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1.4 msLinnet JAX (XLA, generated source) 5.63× the speed of KerasHub latency_ms: 1.00 throughput_per_s: 7063 load_s: 0.51 peak_vram_mib: 624 driver_vram_mib: 1646 max |diff| vs reference: 0.03125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1.0 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 2.43 throughput_per_s: 2296 load_s: 17.23 peak_vram_mib: 2392 max |diff| vs reference: 0.03959846496582031 the reference model's own ONNX export, f32 like Linnet's; torch.onnx export2.4 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.14× the speed of torch.onnx export latency_ms: 2.13 throughput_per_s: 2657 load_s: 0.30 peak_vram_mib: 2474 max |diff| vs reference: 0.03966021537780762 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f322.1 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.91 throughput_per_s: 4406 load_s: 0.28 peak_vram_mib: 1692 max |diff| vs reference: 0.041015625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f161.9 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.76 throughput_per_s: 4090 load_s: 0.29 peak_vram_mib: 1708 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf161.8 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.95 throughput_per_s: 4340 load_s: 0.29 peak_vram_mib: 3140 max |diff| vs reference: 0.04030036926269531 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f320.9 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.61 throughput_per_s: 10975 load_s: 0.28 peak_vram_mib: 2940 max |diff| vs reference: 0.041015625 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f160.6 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.64 throughput_per_s: 10282 load_s: 0.29 peak_vram_mib: 2940 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf160.6 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 3.87 throughput_per_s: 1391 load_s: 10.03 peak_vram_mib: 2435 max |diff| vs reference: 0.041262149810791016 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, torch.onnx export3.9 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.02× the speed of Triton, torch.onnx export latency_ms: 3.81 throughput_per_s: 1241 load_s: 16.07 peak_vram_mib: 2429 max |diff| vs reference: 0.041286468505859375 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, Linnet's ONNX3.8 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 3.21 throughput_per_s: 1272 load_s: 6.02 peak_vram_mib: 2880 max |diff| vs reference: 0.03125 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flightTriton, Linnet Python backend3.2 ms
DINOv2-Base87 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 2.95 throughput_per_s: 1170 load_s: 6.78 peak_vram_mib: 2164 max |diff| vs reference: 0transformers3.0 mstransformers (torch.compile) latency_ms: 1.69 throughput_per_s: 1333 load_s: 3.67 peak_vram_mib: 1870 max |diff| vs reference: 1.53125transformers, compiled1.7 msLinnet torch (generated source) 0.79× the speed of transformers, compiled latency_ms: 2.13 throughput_per_s: 1064 load_s: 1.13 peak_vram_mib: 1870 max |diff| vs reference: 0.625Linnet, generated PyTorch2.1 msLinnet torch (torch.compile, inductor) 1.12× the speed of transformers, compiled latency_ms: 1.51 throughput_per_s: 1415 load_s: 1.29 peak_vram_mib: 1662 max |diff| vs reference: 2.234375Linnet, inductor1.5 msLinnet torch (CUDA graphs) 1.29× the speed of transformers, compiled latency_ms: 1.31 throughput_per_s: 1405 load_s: 1.18 peak_vram_mib: 1816 max |diff| vs reference: 2.234375Linnet, CUDA graphs1.3 msLinnet JAX (XLA, StableHLO) latency_ms: 2.50 throughput_per_s: 603 load_s: 0.46 peak_vram_mib: 3419 driver_vram_mib: 9844 max |diff| vs reference: 4.1640625 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)2.5 msLinnet JAX (XLA, generated source) latency_ms: 2.26 throughput_per_s: 1137 load_s: 0.54 peak_vram_mib: 1288 driver_vram_mib: 2674 max |diff| vs reference: 1.78125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)2.3 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 6.38 throughput_per_s: 253 load_s: 17.33 peak_vram_mib: 20406 max |diff| vs reference: 1.0888938903808594 the reference model's own ONNX export, f32 like Linnet's; torch.onnx export6.4 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.01× the speed of torch.onnx export latency_ms: 6.29 throughput_per_s: 254 load_s: 0.30 peak_vram_mib: 16330 max |diff| vs reference: 1.0886340141296387 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f326.3 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 3.95 throughput_per_s: 452 load_s: 0.31 peak_vram_mib: 8620 max |diff| vs reference: 1.2421875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f164.0 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 4.92 throughput_per_s: 355 load_s: 0.29 peak_vram_mib: 8672 max |diff| vs reference: 6.875 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf164.9 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 4.32 throughput_per_s: 300 load_s: 0.28 peak_vram_mib: 9724 max |diff| vs reference: 1.0885581970214844 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f324.3 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 1.16 throughput_per_s: 1471 load_s: 0.34 peak_vram_mib: 9372 max |diff| vs reference: 1.30078125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f161.2 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 1.26 throughput_per_s: 1412 load_s: 0.28 peak_vram_mib: 9372 max |diff| vs reference: 1.03125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf161.3 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 31.56 throughput_per_s: 48.76 load_s: 12.04 peak_vram_mib: 20351 max |diff| vs reference: 1.0265884399414062 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, torch.onnx export31.6 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.17× the speed of Triton, torch.onnx export latency_ms: 27.06 throughput_per_s: 47.27 load_s: 15.05 peak_vram_mib: 20355 max |diff| vs reference: 1.0261340141296387 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, Linnet's ONNX27.1 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 25.61 throughput_per_s: 47.73 load_s: 6.02 peak_vram_mib: 3948 max |diff| vs reference: 1.046875 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flightTriton, Linnet Python backend25.6 ms
SAM ViT-Base94 Mencoder, batch 1
PyTorchJAXONNX Runtime and TensorRTtransformers (eager) encode_ms: 39.02 load_s: 3.17 peak_vram_mib: 3064 max |diff| vs reference: 0transformers39.0 mstransformers (torch.compile) encode_ms: 28.94 load_s: 0.65 peak_vram_mib: 2394 max |diff| vs reference: 0.00010275840759277344transformers, compiled28.9 msLinnet torch (generated source) 0.70× the speed of transformers, compiled encode_ms: 41.58 load_s: 0.50 peak_vram_mib: 2968 max |diff| vs reference: 0.00008186697959899902Linnet, generated PyTorch41.6 msLinnet torch (torch.compile, inductor) 1.03× the speed of transformers, compiled encode_ms: 28.16 load_s: 0.45 peak_vram_mib: 2250 max |diff| vs reference: 0.00009685754776000977Linnet, inductor28.2 msLinnet torch (CUDA graphs) 1.04× the speed of transformers, compiled encode_ms: 27.74 load_s: 0.50 peak_vram_mib: 2392 max |diff| vs reference: 0.00009685754776000977Linnet, CUDA graphs27.7 msLinnet JAX (XLA, StableHLO) encode_ms: 13.72 load_s: 0.57 peak_vram_mib: 2848 driver_vram_mib: 4728 max |diff| vs reference: 0.0012342222034931183 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)13.7 msLinnet JAX (XLA, generated source) encode_ms: 13.77 load_s: 0.82 peak_vram_mib: 2848 driver_vram_mib: 4728 max |diff| vs reference: 0.001276012510061264 peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)13.8 msLinnet ONNX f32 -> ONNX Runtime (CUDA) encode_ms: 22.84 load_s: 0.31 peak_vram_mib: 8782 max |diff| vs reference: 0.0023171473294496536 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; the image encoderLinnet, ONNX Runtime f3222.8 msLinnet ONNX f16 -> ONNX Runtime (CUDA) encode_ms: 20.94 load_s: 0.36 peak_vram_mib: 8002 max |diff| vs reference: 0.002835877239704132 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the image encoderLinnet, ONNX Runtime f1620.9 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) encode_ms: 22.86 load_s: 0.30 peak_vram_mib: 7854 max |diff| vs reference: 0.02760845422744751 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; the image encoderLinnet, ONNX Runtime bf1622.9 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) encode_ms: 14.24 load_s: 0.30 peak_vram_mib: 4492 max |diff| vs reference: 0.0012321211397647858 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; the image encoderLinnet, TensorRT f3214.2 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) encode_ms: 9.64 load_s: 0.32 peak_vram_mib: 4206 max |diff| vs reference: 0.007983803749084473 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; the image encoderLinnet, TensorRT f169.6 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) encode_ms: 9.76 load_s: 0.32 peak_vram_mib: 3904 max |diff| vs reference: 0.029810786247253418 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; the image encoderLinnet, TensorRT bf169.8 ms
BERT Base (uncased)109 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 3.60 throughput_per_s: 11043 load_s: 1.44 peak_vram_mib: 1162 max |diff| vs reference: 0 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers3.6 mstransformers (torch.compile) latency_ms: 1.67 throughput_per_s: 14663 load_s: 1.63 peak_vram_mib: 1142 max |diff| vs reference: 0.109375 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers, compiled1.7 msLinnet torch (generated source) 0.60× the speed of transformers, compiled latency_ms: 2.79 throughput_per_s: 13867 load_s: 1.09 peak_vram_mib: 1108 max |diff| vs reference: 0.1171875 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, generated PyTorch2.8 msLinnet torch (torch.compile, inductor) 0.95× the speed of transformers, compiled latency_ms: 1.75 throughput_per_s: 17002 load_s: 1.11 peak_vram_mib: 1060 max |diff| vs reference: 0.140625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, inductor1.8 msLinnet torch (CUDA graphs) 2.27× the speed of transformers, compiled latency_ms: 0.74 throughput_per_s: 17770 load_s: 1.12 peak_vram_mib: 1148 max |diff| vs reference: 0.140625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, CUDA graphs0.7 msKerasHub (JAX) latency_ms: 4.67 throughput_per_s: 5159 load_s: 6.31 peak_vram_mib: 525 driver_vram_mib: 1650 max |diff| vs reference: 0.2265625 bf16, under jax.jit; unpadded batches of exactly 128 tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub4.7 msLinnet JAX (XLA, StableHLO) 3.01× the speed of KerasHub latency_ms: 1.56 throughput_per_s: 13223 load_s: 0.52 peak_vram_mib: 734 driver_vram_mib: 1646 max |diff| vs reference: 0.109375 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1.6 msLinnet JAX (XLA, generated source) 5.05× the speed of KerasHub latency_ms: 0.93 throughput_per_s: 11749 load_s: 0.58 peak_vram_mib: 727 driver_vram_mib: 1646 max |diff| vs reference: 0.25 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)0.9 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 2.06 throughput_per_s: 4943 load_s: 17.54 peak_vram_mib: 3382 max |diff| vs reference: 0.11434412002563477 the reference model's own ONNX export, f32 like Linnet's; unpadded batches of exactly 128 tokenstorch.onnx export2.1 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.36× the speed of torch.onnx export latency_ms: 1.52 throughput_per_s: 5274 load_s: 0.33 peak_vram_mib: 2586 max |diff| vs reference: 0.1145319938659668 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f321.5 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.25 throughput_per_s: 9557 load_s: 0.56 peak_vram_mib: 1748 max |diff| vs reference: 0.11328125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f161.3 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.25 throughput_per_s: 9728 load_s: 0.29 peak_vram_mib: 1748 max |diff| vs reference: 0.125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime bf161.3 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.74 throughput_per_s: 9036 load_s: 0.28 peak_vram_mib: 3282 max |diff| vs reference: 0.11189699172973633 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f320.7 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.53 throughput_per_s: 19916 load_s: 0.28 peak_vram_mib: 3042 max |diff| vs reference: 0.11328125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f160.5 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.55 throughput_per_s: 18310 load_s: 0.31 peak_vram_mib: 3038 max |diff| vs reference: 0.109375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT bf160.5 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 3.66 throughput_per_s: 579 load_s: 9.10 peak_vram_mib: 2429 max |diff| vs reference: 0.11434412002563477 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, torch.onnx export3.7 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.13× the speed of Triton, torch.onnx export latency_ms: 3.24 throughput_per_s: 594 load_s: 16.06 peak_vram_mib: 2425 max |diff| vs reference: 0.1145319938659668 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, Linnet's ONNX3.2 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 2.71 throughput_per_s: 558 load_s: 6.03 peak_vram_mib: 2880 max |diff| vs reference: 0.140625 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 64 in flightTriton, Linnet Python backend2.7 ms
RoBERTa Base124 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 3.52 throughput_per_s: 10857 load_s: 15.56 peak_vram_mib: 1190 max |diff| vs reference: 0 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers3.5 mstransformers (torch.compile) latency_ms: 1.57 throughput_per_s: 14513 load_s: 1.55 peak_vram_mib: 1170 max |diff| vs reference: 0.125 unpadded batches of exactly 128 tokens (this card takes no padding mask)transformers, compiled1.6 msLinnet torch (generated source) 0.82× the speed of transformers, compiled latency_ms: 1.92 throughput_per_s: 13836 load_s: 1.14 peak_vram_mib: 1136 max |diff| vs reference: 0.15625 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, generated PyTorch1.9 msLinnet torch (torch.compile, inductor) 1.14× the speed of transformers, compiled latency_ms: 1.38 throughput_per_s: 17223 load_s: 1.11 peak_vram_mib: 1088 max |diff| vs reference: 0.125 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, inductor1.4 msLinnet torch (CUDA graphs) 2.14× the speed of transformers, compiled latency_ms: 0.73 throughput_per_s: 18181 load_s: 1.14 peak_vram_mib: 1176 max |diff| vs reference: 0.125 unpadded batches of exactly 128 tokens; bf16 on cudaLinnet, CUDA graphs0.7 msKerasHub (JAX) latency_ms: 4.95 throughput_per_s: 5082 load_s: 5.99 peak_vram_mib: 567 driver_vram_mib: 1650 max |diff| vs reference: 0.1171875 bf16, under jax.jit; unpadded batches of exactly 128 tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsKerasHub4.9 msLinnet JAX (XLA, StableHLO) 4.20× the speed of KerasHub latency_ms: 1.18 throughput_per_s: 12433 load_s: 0.54 peak_vram_mib: 785 driver_vram_mib: 1646 max |diff| vs reference: 0.15625 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1.2 msLinnet JAX (XLA, generated source) 5.32× the speed of KerasHub latency_ms: 0.93 throughput_per_s: 11985 load_s: 0.57 peak_vram_mib: 771 driver_vram_mib: 1646 max |diff| vs reference: 0.21875 unpadded batches of exactly 128 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)0.9 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 2.44 throughput_per_s: 4955 load_s: 18.56 peak_vram_mib: 2874 max |diff| vs reference: 0.08825016021728516 the reference model's own ONNX export, f32 like Linnet's; unpadded batches of exactly 128 tokenstorch.onnx export2.4 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.60× the speed of torch.onnx export latency_ms: 1.53 throughput_per_s: 5276 load_s: 0.28 peak_vram_mib: 2702 max |diff| vs reference: 0.08825016021728516 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f321.5 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.23 throughput_per_s: 9664 load_s: 0.29 peak_vram_mib: 1804 max |diff| vs reference: 0.087890625 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f161.2 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.23 throughput_per_s: 9497 load_s: 0.37 peak_vram_mib: 1804 max |diff| vs reference: 0.125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime bf161.2 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.75 throughput_per_s: 9050 load_s: 0.27 peak_vram_mib: 3410 max |diff| vs reference: 0.08764839172363281 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f320.7 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.53 throughput_per_s: 19752 load_s: 0.27 peak_vram_mib: 3130 max |diff| vs reference: 0.09765625 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f160.5 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.56 throughput_per_s: 18619 load_s: 0.27 peak_vram_mib: 3128 max |diff| vs reference: 0.125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT bf160.6 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 3.73 throughput_per_s: 566 load_s: 10.18 peak_vram_mib: 3585 max |diff| vs reference: 0.08825016021728516 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, torch.onnx export3.7 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.12× the speed of Triton, torch.onnx export latency_ms: 3.32 throughput_per_s: 502 load_s: 17.03 peak_vram_mib: 3961 max |diff| vs reference: 0.08825016021728516 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, Linnet's ONNX3.3 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 2.68 throughput_per_s: 540 load_s: 6.02 peak_vram_mib: 2936 max |diff| vs reference: 0.125 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 64 in flightTriton, Linnet Python backend2.7 ms
ModernBERT-base149 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 9.19 throughput_per_s: 1358 load_s: 3.18 peak_vram_mib: 1922 max |diff| vs reference: 0 unpadded batches of exactly 256 tokens (this card takes no padding mask)transformers9.2 mstransformers (torch.compile) latency_ms: 2.31 throughput_per_s: 4488 load_s: 1.39 peak_vram_mib: 1374 max |diff| vs reference: 2.127197265625 unpadded batches of exactly 256 tokens (this card takes no padding mask)transformers, compiled2.3 msLinnet torch (generated source) 0.40× the speed of transformers, compiled latency_ms: 5.78 throughput_per_s: 2780 load_s: 1.39 peak_vram_mib: 1304 max |diff| vs reference: 1.625 unpadded batches of exactly 256 tokens; bf16 on cudaLinnet, generated PyTorch5.8 msLinnet torch (torch.compile, inductor) 1.07× the speed of transformers, compiled latency_ms: 2.15 throughput_per_s: 6042 load_s: 1.03 peak_vram_mib: 1306 max |diff| vs reference: 1.25390625 unpadded batches of exactly 256 tokens; bf16 on cudaLinnet, inductor2.2 msLinnet torch (CUDA graphs) 2.28× the speed of transformers, compiled latency_ms: 1.01 throughput_per_s: 6153 load_s: 1.36 peak_vram_mib: 1442 max |diff| vs reference: 1.25390625 unpadded batches of exactly 256 tokens; bf16 on cudaLinnet, CUDA graphs1.0 msLinnet JAX (XLA, StableHLO) latency_ms: 1.99 throughput_per_s: 3252 load_s: 0.58 peak_vram_mib: 927 driver_vram_mib: 1648 max |diff| vs reference: 1.46875 unpadded batches of exactly 256 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)2.0 msLinnet JAX (XLA, generated source) latency_ms: 1.81 throughput_per_s: 3506 load_s: 0.63 peak_vram_mib: 1234 driver_vram_mib: 2672 max |diff| vs reference: 2.21142578125 unpadded batches of exactly 256 tokens; bf16 on cuda; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1.8 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 3.59 throughput_per_s: 1302 load_s: 31.85 peak_vram_mib: 6960 max |diff| vs reference: 2.3369617462158203 the reference model's own ONNX export, f32 like Linnet's; unpadded batches of exactly 256 tokenstorch.onnx export3.6 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 0.99× the speed of torch.onnx export latency_ms: 3.63 throughput_per_s: 1347 load_s: 0.29 peak_vram_mib: 4822 max |diff| vs reference: 2.393662452697754 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f323.6 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 2.79 throughput_per_s: 2096 load_s: 0.30 peak_vram_mib: 2812 max |diff| vs reference: 2.188232421875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime f162.8 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 3.24 throughput_per_s: 2040 load_s: 0.28 peak_vram_mib: 2812 max |diff| vs reference: 2.25 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, ONNX Runtime bf163.2 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 1.71 throughput_per_s: 2035 load_s: 0.30 peak_vram_mib: 4116 max |diff| vs reference: 2.3549375534057617 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f321.7 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 1.25 throughput_per_s: 5759 load_s: 0.28 peak_vram_mib: 3768 max |diff| vs reference: 2.162841796875 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT f161.3 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 1.36 throughput_per_s: 5456 load_s: 0.28 peak_vram_mib: 3770 max |diff| vs reference: 1.9375 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; unpadded batches of exactly the sequence lengthLinnet, TensorRT bf161.4 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 6.79 throughput_per_s: 267 load_s: 12.03 peak_vram_mib: 7105 max |diff| vs reference: 2.3369617462158203 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, torch.onnx export6.8 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.03× the speed of Triton, torch.onnx export latency_ms: 6.60 throughput_per_s: 263 load_s: 18.10 peak_vram_mib: 7065 max |diff| vs reference: 2.393662452697754 over HTTP; f32; throughput with 4 requests of batch 64 in flightTriton, Linnet's ONNX6.6 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 3.97 throughput_per_s: 270 load_s: 6.02 peak_vram_mib: 3202 max |diff| vs reference: 1.25390625 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 64 in flightTriton, Linnet Python backend4.0 ms
SigLIP Base/16 224203 Mforward pass, batch 1
PyTorchJAXONNX Runtime and TensorRTTriton Inference Servertransformers (eager) latency_ms: 2.87 throughput_per_s: 6030 load_s: 7.93 peak_vram_mib: 1290 max |diff| vs reference: 0transformers2.9 mstransformers (torch.compile) latency_ms: 1.89 throughput_per_s: 8317 load_s: 4.07 peak_vram_mib: 1312 max |diff| vs reference: 0.03125transformers, compiled1.9 msLinnet torch (generated source) 0.86× the speed of transformers, compiled latency_ms: 2.21 throughput_per_s: 7997 load_s: 2.38 peak_vram_mib: 1270 max |diff| vs reference: 0.015625Linnet, generated PyTorch2.2 msLinnet torch (torch.compile, inductor) 1.15× the speed of transformers, compiled latency_ms: 1.64 throughput_per_s: 9536 load_s: 2.33 peak_vram_mib: 1252 max |diff| vs reference: 0.03125Linnet, inductor1.6 msLinnet torch (CUDA graphs) 2.04× the speed of transformers, compiled latency_ms: 0.93 throughput_per_s: 9937 load_s: 2.28 peak_vram_mib: 1452 max |diff| vs reference: 0.03125Linnet, CUDA graphs0.9 msLinnet JAX (XLA, StableHLO) latency_ms: 1.15 throughput_per_s: 6716 load_s: 0.73 peak_vram_mib: 1085 driver_vram_mib: 2672 max |diff| vs reference: 0.03125 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)1.2 msLinnet JAX (XLA, generated source) latency_ms: 1.11 throughput_per_s: 9054 load_s: 1.02 peak_vram_mib: 1067 driver_vram_mib: 2670 max |diff| vs reference: 0.0234375 bf16, like the torch rows; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)1.1 mstransformers -> torch.onnx -> ONNX Runtime latency_ms: 2.01 throughput_per_s: 2664 load_s: 19.73 peak_vram_mib: 2394 max |diff| vs reference: 0.04458141326904297 the reference model's own ONNX export, f32 like Linnet's; torch.onnx export2.0 msLinnet ONNX f32 -> ONNX Runtime (CUDA) 1.01× the speed of torch.onnx export latency_ms: 2.00 throughput_per_s: 2752 load_s: 0.29 peak_vram_mib: 2444 max |diff| vs reference: 0.04524040222167969 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f322.0 msLinnet ONNX f16 -> ONNX Runtime (CUDA) latency_ms: 1.94 throughput_per_s: 4923 load_s: 0.31 peak_vram_mib: 1684 max |diff| vs reference: 0.046875 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f161.9 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) latency_ms: 1.74 throughput_per_s: 4443 load_s: 0.34 peak_vram_mib: 1694 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf161.7 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) latency_ms: 0.90 throughput_per_s: 4968 load_s: 0.31 peak_vram_mib: 3202 max |diff| vs reference: 0.04555702209472656 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f320.9 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) latency_ms: 0.64 throughput_per_s: 11447 load_s: 0.32 peak_vram_mib: 2960 max |diff| vs reference: 0.04296875 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f160.6 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) latency_ms: 0.67 throughput_per_s: 10896 load_s: 0.34 peak_vram_mib: 2962 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf160.7 msTriton Inference Server (ONNX Runtime backend, torch.onnx export) latency_ms: 4.04 throughput_per_s: 1314 load_s: 9.03 peak_vram_mib: 2437 max |diff| vs reference: 0.029298901557922363 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, torch.onnx export4.0 msTriton Inference Server (ONNX Runtime backend, Linnet ONNX) 1.01× the speed of Triton, torch.onnx export latency_ms: 4.00 throughput_per_s: 1235 load_s: 17.07 peak_vram_mib: 2433 max |diff| vs reference: 0.02989780902862549 over HTTP; f32; throughput with 4 requests of batch 32 in flightTriton, Linnet's ONNX4.0 msTriton Inference Server (Python backend, Linnet torch) latency_ms: 3.19 throughput_per_s: 1225 load_s: 6.03 peak_vram_mib: 3302 max |diff| vs reference: 0.03125 over HTTP; bf16, CUDA graphs; throughput with 4 requests of batch 32 in flightTriton, Linnet Python backend3.2 ms
Whisper large-v31.5 Btranscribe_ms, batch 1
PyTorchJAXONNX Runtime and TensorRTtransformers (eager) encode_ms: 9.21 transcribe_ms: 307 load_s: 8.98 peak_vram_mib: 4196 max |diff| vs reference: 0transformers307 mstransformers (torch.compile, static cache) encode_ms: 7.27 transcribe_ms: 192 load_s: 1.84 peak_vram_mib: 4472 max |diff| vs reference: 4.90625transformers, compiled192 msLinnet torch (generated source) 0.57× the speed of transformers, compiled encode_ms: 7.34 transcribe_ms: 338 load_s: 2.07 peak_vram_mib: 5448 max |diff| vs reference: 1.125 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, generated PyTorch338 msLinnet torch (torch.compile, inductor) 1.21× the speed of transformers, compiled encode_ms: 6.58 transcribe_ms: 159 load_s: 2.01 peak_vram_mib: 5492 max |diff| vs reference: 4.5703125 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, inductor159 msLinnet torch (CUDA graphs) 2.85× the speed of transformers, compiled encode_ms: 5.63 transcribe_ms: 67.42 load_s: 2.00 peak_vram_mib: 5632 max |diff| vs reference: 4.5703125 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, CUDA graphs67.4 msLinnet JAX (XLA, StableHLO) encode_ms: 11.70 transcribe_ms: 96.54 load_s: 2.30 peak_vram_mib: 3699 driver_vram_mib: 4749 max |diff| vs reference: 3.2890625 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)96.5 msLinnet JAX (XLA, generated source) encode_ms: 7.32 transcribe_ms: 108 load_s: 2.13 peak_vram_mib: 3616 driver_vram_mib: 4745 max |diff| vs reference: 16.263671875 transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token each; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)108 msLinnet ONNX f32 -> ONNX Runtime (CUDA) encode_ms: 25.25 transcribe_ms: 158 load_s: 0.25 peak_vram_mib: 12022 max |diff| vs reference: 4.8056793212890625 f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime f32158 msLinnet ONNX f16 -> ONNX Runtime (CUDA) encode_ms: 28.82 transcribe_ms: 154 load_s: 0.24 peak_vram_mib: 7756 max |diff| vs reference: 0.636474609375 f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime f16154 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) encode_ms: 21.01 transcribe_ms: 160 load_s: 0.24 peak_vram_mib: 8864 max |diff| vs reference: 15.03515625 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, ONNX Runtime bf16160 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) encode_ms: 25.00 transcribe_ms: 250 load_s: 0.25 peak_vram_mib: 18440 max |diff| vs reference: 4.8056793212890625 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT f32250 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) encode_ms: 5.42 transcribe_ms: 194 load_s: 0.24 peak_vram_mib: 11168 max |diff| vs reference: 3.9921875 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT f16194 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) encode_ms: 6.13 transcribe_ms: 202 load_s: 0.24 peak_vram_mib: 10936 max |diff| vs reference: 14.21875 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; transcribed with the card's KV caches: `listen` encodes and fills the cross-attention caches, `prefill` feeds the prompt, `step` one token eachLinnet, TensorRT bf16202 ms
SDXL Base UNet2.6 Bone denoising step, batch 1
PyTorchJAXONNX Runtime and TensorRTdiffusers (eager) step_ms: 48.76 load_s: 31.05 peak_vram_mib: 6444 max |diff| vs reference: 0 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16diffusers48.8 msdiffusers (torch.compile) step_ms: 34.96 load_s: 8.13 peak_vram_mib: 6450 max |diff| vs reference: 0.03125 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16diffusers, compiled35.0 msLinnet torch (generated source) 0.62× the speed of diffusers, compiled step_ms: 56.35 load_s: 16.95 peak_vram_mib: 6660 max |diff| vs reference: 0.015625 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16Linnet, generated PyTorch56.3 msLinnet torch (torch.compile, inductor) 0.98× the speed of diffusers, compiled step_ms: 35.71 load_s: 17.24 peak_vram_mib: 6430 max |diff| vs reference: 0.03125 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16Linnet, inductor35.7 msLinnet torch (CUDA graphs) 1.01× the speed of diffusers, compiled step_ms: 34.70 load_s: 19.83 peak_vram_mib: 6578 max |diff| vs reference: 0.03125 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16Linnet, CUDA graphs34.7 msLinnet JAX (XLA, StableHLO) step_ms: 53.69 load_s: 22.19 peak_vram_mib: 6463 driver_vram_mib: 8830 max |diff| vs reference: 0.03125 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (StableHLO)53.7 msLinnet JAX (XLA, generated source) step_ms: 39.27 load_s: 25.02 peak_vram_mib: 5363 driver_vram_mib: 8822 max |diff| vs reference: 0.03125 one step, 128x128 latent (1024-pixel image), batch 2, 77-token context, bf16; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regionsLinnet, XLA (generated JAX)39.3 msLinnet ONNX f32 -> ONNX Runtime (CUDA) step_ms: 141 load_s: 0.30 peak_vram_mib: 24896 max |diff| vs reference: 0.030341625213623047 f32; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f32141 msLinnet ONNX f16 -> ONNX Runtime (CUDA) step_ms: 90.85 load_s: 0.29 peak_vram_mib: 14408 max |diff| vs reference: 0.03125 f16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime f1690.8 msLinnet ONNX bf16 -> ONNX Runtime (CUDA) step_ms: 100 load_s: 0.29 peak_vram_mib: 14944 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's CUDA execution provider; first calls build the sessionsLinnet, ONNX Runtime bf16100 msLinnet ONNX f32 -> ONNX Runtime (TensorRT) step_ms: 141 load_s: 0.30 peak_vram_mib: 24896 max |diff| vs reference: 0.030341625213623047 f32; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f32141 msLinnet ONNX f16 -> ONNX Runtime (TensorRT) step_ms: 91.36 load_s: 0.29 peak_vram_mib: 14408 max |diff| vs reference: 0.03125 f16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT f1691.4 msLinnet ONNX bf16 -> ONNX Runtime (TensorRT) step_ms: 100 load_s: 0.29 peak_vram_mib: 14944 max |diff| vs reference: 0.03125 bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessionsLinnet, TensorRT bf16100 ms

Models with no KerasHub implementation have no JAX baseline.

Every row ​

MethodTTFT mstok/speak GiBload smax |diff|Notes
PyTorch
transformers22.251.216.05.290
transformers, compiled18.611016.14.290.063
Linnet, generated PyTorch17.479.015.93.590.094KV cache compiled for 768 positions
Linnet, inductor13.114122.83.350.063KV cache compiled for 768 positions
Linnet, CUDA graphs13.016122.83.100.063KV cache compiled for 768 positions
JAX
KerasHub27.516017.331.00.094peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (StableHLO)16.616716.310.70.094KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, XLA (generated JAX)15.716616.310.60.096KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
ONNX Runtime and TensorRT
Linnet, ONNX Runtime f3242.053.735.20.500.065f32; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime f1630.571.017.50.510.070f16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, ONNX Runtime bf1631.171.816.50.480.094bf16; ONNX Runtime's CUDA execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT f3265.5failed: Fail: [ONNXRuntimeError] : 1 : FAIL : TensorRT EP failed to create engine from network for fused node: TensorrtExecutionProvider_TRTKernel_graph_main_3492707504
Linnet, TensorRT f1628.647.542.00.499.8f16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
Linnet, TensorRT bf1628.549.539.90.490.13bf16; ONNX Runtime's TensorRT execution provider; first calls build the sessions; logits copied to the host each step for the argmax
LLM engines
vLLM16.515267.474.6reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
vLLM on Linnet's export12.215767.439.3linnet.hf.export (16 s), then vLLM on the exported checkpoint; reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
SGLang21.015868.055.7reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
SGLang on Linnet's export19.915867.938.4linnet.hf.export (18 s), then SGLang on the exported checkpoint; reserves a KV-cache pool up front (mem_fraction_static 0.85), so its memory is a setting, not a need
TGI27.911259.244.1over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
TGI on Linnet's export25.711559.240.1linnet.hf.export (19 s), then Text Generation Inference on the exported checkpoint; over HTTP, streamed; the prompt is the token ids decoded and tokenized again; reserves a KV-cache pool up front (cuda-memory-fraction 0.85), so its memory is a setting, not a need
llama.cpp on Linnet's GGUF22.017115.075.5linnet.gguf.export, then llama-bench with every layer on the GPU; the first token is the prompt at llama-bench's prompt rate
Beyond one GPU
vLLM, tensor parallel11.9233139.4286reserves a KV-cache pool up front (gpu_memory_utilization 0.85), so its memory is a setting, not a need
Linnet, tensor parallel (PyTorch)10.823827.14.77one process per GPU under torchrun, NCCL, CUDA graphs; KV cache compiled for 768 positions
Linnet, tensor parallel (XLA)16.218118.416.10.13KV cache compiled for 768 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
Linnet, layers on two GPUs22.587.823.37.000.094KV cache compiled for 768 positions
Linnet, offloaded to host1805.738.919.60.094KV cache compiled for 768 positions; cuda:0: embedding, layers.0-13; host, streamed in: layers.14-31, norm, lm_head
Serving many requests
vLLM68.7135256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
linnet.serve, CUDA graphs27.86.99256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
linnet.serve, XLA22.112.0256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions
linnet.serve, ONNX Runtime37.60.26256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; 64 fixed cache rows of 640 positions
vLLM on Linnet's export68.734.7linnet.hf.export (0 s), then vLLM on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at gpu_memory_utilization 0.85
SGLang70.055.8256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
SGLang on Linnet's export69.944.9linnet.hf.export (0 s), then SGLang (offline, continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; KV-cache pool at mem_fraction_static 0.85
TGI60.742.1256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
TGI on Linnet's export61.240.1linnet.hf.export (0 s), then Text Generation Inference (continuous batching) on the exported checkpoint; 256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed; stops at end-of-sequence (TGI cannot ignore it), so throughput counts the tokens produced; KV-cache pool at cuda-memory-fraction 0.85
Triton, vLLM backend67.561.1256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
Triton, linnet.serve backend28.3645256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; over HTTP, streamed, first tokens timed at the client
transformers, batched79.13.98256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; paged|sdpa attention; reserves a paged KV-cache pool up front, so its memory is a setting, not a need; no per-request timestamps
KerasHub, static batches25.830.1256 requests of 128-512 prompt tokens and 128 new tokens, at most 64 in flight; every batch runs to the longest possible prompt plus the new tokens; peak memory is JAX's allocator peak; the driver shows its pool, which grows in whole regions

The generated code ​

best in the chart Linnet PyTorch reference numerics=fast
forward · H=512, 8 layers
0×0.5×1×1.5×2×2.5×3×PyTorch eager1.00×PyTorch torch.compile2.03×Linnet interpreted0.90×Linnet generated0.81×Linnet generated0.84×Linnet generated + compile1.53×Linnet generated + compile, fast1.77×Linnet generated + CUDA graphs, fast3.36×Linnet XLA2.99×Linnet XLA, fast2.02×
decode · H=512, 8 layers
0×0.5×1×1.5×2×2.5×3×PyTorch eager1.00×Linnet interpreted0.95×Linnet generated0.96×Linnet generated1.02×Linnet generated + compile1.60×Linnet generated + compile, fast1.73×Linnet generated + CUDA graphs, fast3.20×Linnet XLA1.80×Linnet XLA, fast2.05×
forward · H=2048, 22 layers
0×0.5×1×1.5×2×PyTorch eager1.00×PyTorch torch.compile2.35×Linnet interpreted0.76×Linnet generated0.77×Linnet generated1.21×Linnet generated + compile1.40×Linnet generated + compile, fast2.08×Linnet generated + CUDA graphs, fast2.27×Linnet XLA1.35×Linnet XLA, fast1.48×
decode · H=2048, 22 layers
0×0.5×1×1.5×2×2.5×PyTorch eager1.00×Linnet interpreted0.71×Linnet generated0.72×Linnet generated0.71×Linnet generated + compile1.51×Linnet generated + compile, fast1.49×Linnet generated + CUDA graphs, fast2.57×Linnet XLA2.58×Linnet XLA, fast2.40×

Measured 2026-09-24 on device NVIDIA H100 80GB HBM3, torch 2.9.1+cu128, python 3.12.3, os Linux 6.8.0-90-generic, jax 0.11.2 (gpu), linnet linnet 0.1.0.

Llama-style decoder · H=512 L=8 heads=8/8 B=1 S=512 bf16 · forward

VariantLatency (ms)tokens/svs referenceKernelsmax |Δ|
PyTorch reference (eager) 2.71188,9543580
PyTorch reference (torch.compile) 1.33384,2082.03×1174.9e-4
Linnet → PyTorch (numerics=equivalent) Core IR interpreted per call; library ops on native kernels3.02169,3410.90×2462.4e-4
Linnet → PyTorch (generated source) straight-line PyTorch from `linnet torch`, native kernels3.33153,5690.81×2462.4e-4
Linnet → PyTorch (generated source, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels3.24158,0840.84×2210
Linnet → PyTorch (generated source + torch.compile) straight-line PyTorch from `linnet torch`, native kernels1.77288,7791.53×1224.9e-4
Linnet → PyTorch (generated source + torch.compile, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels1.54333,5311.77×1064.9e-4
Linnet → PyTorch (generated source + CUDA graphs, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels0.81635,8233.36×1094.9e-4
Linnet → XLA (jax, gpu) whole-program compile of the StableHLO export0.91565,0592.99×—2.4e-4
Linnet → XLA (jax, gpu, numerics=fast) whole-program compile of the StableHLO export1.34381,4192.02×—2.4e-4

Llama-style decoder · H=512 L=8 heads=8/8 B=1 S=512 bf16 · decode

VariantLatency (ms)steps/svs referenceKernelsmax |Δ|
PyTorch reference (eager, recomputes prefix) no KV cache: full forward over 257 tokens3.49286.53520
Linnet → PyTorch (numerics=equivalent) KV cache at position 256 of 5123.66273.30.95×3172.4e-4
Linnet → PyTorch (generated source) KV cache at position 256 of 5123.64274.50.96×3172.4e-4
Linnet → PyTorch (generated source, numerics=fast) KV cache at position 256 of 5123.42292.81.02×2852.4e-4
Linnet → PyTorch (generated source + torch.compile) KV cache at position 256 of 5122.18459.41.60×1442.9e-4
Linnet → PyTorch (generated source + torch.compile, numerics=fast) KV cache at position 256 of 5122.02495.41.73×1193.1e-4
Linnet → PyTorch (generated source + CUDA graphs, numerics=fast) KV cache at position 256 of 5121.09915.63.20×1553.1e-4
Linnet → XLA (jax, gpu) KV cache threaded as inputs/outputs, position 2561.94516.31.80×—2.4e-4
Linnet → XLA (jax, gpu, numerics=fast) KV cache threaded as inputs/outputs, position 2561.70586.52.05×—2.4e-4

Llama-style decoder · H=2048 L=22 heads=32/8 B=1 S=512 bf16 · forward

VariantLatency (ms)tokens/svs referenceKernelsmax |Δ|
PyTorch reference (eager) 8.6359,3289000
PyTorch reference (torch.compile) 3.68139,2452.35×3251.1e-3
Linnet → PyTorch (numerics=equivalent) Core IR interpreted per call; library ops on native kernels11.4244,8220.76×10309.8e-4
Linnet → PyTorch (generated source) straight-line PyTorch from `linnet torch`, native kernels11.2445,5710.77×10309.8e-4
Linnet → PyTorch (generated source, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels7.1571,5871.21×5684.9e-4
Linnet → PyTorch (generated source + torch.compile) straight-line PyTorch from `linnet torch`, native kernels6.1882,8431.40×3881.1e-3
Linnet → PyTorch (generated source + torch.compile, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels4.15123,2892.08×2781.1e-3
Linnet → PyTorch (generated source + CUDA graphs, numerics=fast) straight-line PyTorch from `linnet torch`, native kernels3.79134,9472.27×2821.1e-3
Linnet → XLA (jax, gpu) whole-program compile of the StableHLO export6.3980,1061.35×—9.8e-4
Linnet → XLA (jax, gpu, numerics=fast) whole-program compile of the StableHLO export5.8587,5751.48×—9.8e-4

Llama-style decoder · H=2048 L=22 heads=32/8 B=1 S=512 bf16 · decode

VariantLatency (ms)steps/svs referenceKernelsmax |Δ|
PyTorch reference (eager, recomputes prefix) no KV cache: full forward over 257 tokens9.06110.48950
Linnet → PyTorch (numerics=equivalent) KV cache at position 256 of 51212.7178.60.71×11196.1e-4
Linnet → PyTorch (generated source) KV cache at position 256 of 51212.5779.60.72×11186.1e-4
Linnet → PyTorch (generated source, numerics=fast) KV cache at position 256 of 51212.7478.50.71×11406.1e-4
Linnet → PyTorch (generated source + torch.compile) KV cache at position 256 of 5126.01166.31.51×5147.3e-4
Linnet → PyTorch (generated source + torch.compile, numerics=fast) KV cache at position 256 of 5126.07164.71.49×5137.3e-4
Linnet → PyTorch (generated source + CUDA graphs, numerics=fast) KV cache at position 256 of 5123.52283.92.57×5987.3e-4
Linnet → XLA (jax, gpu) KV cache threaded as inputs/outputs, position 2563.51285.02.58×—7.3e-4
Linnet → XLA (jax, gpu, numerics=fast) KV cache threaded as inputs/outputs, position 2563.78264.82.40×—8.1e-4

Compile and load

StepSeconds
linnet check examples/05-llama1.85
linnet.torch.load (small)2.31
torch.compile (small)1.17
linnet torch + first call (small, generated source)4.82
linnet torch + first call (small, generated source, numerics=fast)4.62
linnet torch + first call (small, generated source + torch.compile)6.85
linnet torch + first call (small, generated source + torch.compile, numerics=fast)4.94
linnet torch + first call (small, generated source + CUDA graphs, numerics=fast)4.77
linnet stablehlo + XLA compile (small, forward, numerics=equivalent)6.53
linnet stablehlo + XLA compile (small, forward, numerics=fast)4.24
linnet stablehlo + XLA compile (small, decode, numerics=equivalent) incl. 256 steps5.63
linnet stablehlo + XLA compile (small, decode, numerics=fast) incl. 256 steps5.46
linnet.torch.load (medium)2.22
torch.compile (medium)2.75
linnet torch + first call (medium, generated source)2.92
linnet torch + first call (medium, generated source, numerics=fast)4.82
linnet torch + first call (medium, generated source + torch.compile)8.24
linnet torch + first call (medium, generated source + torch.compile, numerics=fast)4.86
linnet torch + first call (medium, generated source + CUDA graphs, numerics=fast)5.58
linnet stablehlo + XLA compile (medium, forward, numerics=equivalent)8.62
linnet stablehlo + XLA compile (medium, forward, numerics=fast)5.79
linnet stablehlo + XLA compile (medium, decode, numerics=equivalent) incl. 256 steps7.76
linnet stablehlo + XLA compile (medium, decode, numerics=fast) incl. 256 steps6.89

A Llama-shaped model with random weights, against a hand-written PyTorch implementation, shows what the generated code costs on its own. On the medium model, plain generated PyTorch runs a forward pass in 7.2 ms against 8.6 ms for the eager reference, and CUDA graphs bring it to 3.8 ms against 3.7 ms for the compiled reference. The small model's plain generated code is slower than eager (3.2 ms against 2.7 ms) until compiled (1.5 ms) or replayed as CUDA graphs (0.8 ms). A decode step is launch-bound: the medium model's spends most of its 12.7 ms in Python dispatch, so use CUDA graphs in a decoding loop, which bring it to 3.5 ms.

How it was measured ​

  • Hardware: one H100 80GB (SXM), each method in its own process.
  • Dtype: bf16.
  • Workloads: one request at a time is a 512-token prompt, then 128 greedy tokens. Serving is 256 requests of 128 to 512 prompt tokens, each wanting 128 new tokens, all waiting from the start, at most 64 in flight. Encoders, vision, audio and diffusion run one forward pass at batch 1 (latency) and at a large batch (throughput).
  • Baselines: transformers or diffusers (eager or torch.compile, whichever is faster), sentence-transformers, KerasHub in JAX, the model's own torch.onnx export on ONNX Runtime and behind Triton Inference Server, vLLM (prefix cache off, since every timed call sends the same prompt), SGLang, Text Generation Inference, and llama.cpp.
  • Linnet's row: its fastest configuration across generated PyTorch (as is, under torch.compile, or as CUDA graphs), XLA, ONNX Runtime (CUDA or TensorRT), and Triton. Every output is checked against the reference stack's.
  • Memory: peak GPU memory from the driver, or from JAX's allocator for JAX rows. The memory views leave out the offloaded rows and every row that starts vLLM, which reserves 85% of the GPU for its cache pool.
  • Generated code: examples/01-llama with random weights, small (about 60 M parameters) and medium (TinyLlama shape, about 1.1 B), timing forward over B=1, S=512 and one decode step at position 256. The variants are the linnet.torch.load modes and linnet.jax.load under jax.jit.
  • Training: Llama 3.1 8B Instruct on H100s, packed into 4096-token rows. SFT uses Alpaca, DPO UltraFeedback pairs, and GRPO two-number multiplications with an exact-answer reward. The reference is TRL with PEFT, FSDP2 for the four-GPU run and vLLM colocated for GRPO. Every stack uses the same hyperparameters and step count.
  • Data: the charts and tables come from bench/results/zoo.json, bench/results/latest.json and bench/results/training.json; the prose quotes them. The zoo's bench/compare.json pairs the rows, and its bench/setup-pod.sh and bench/run-all.sh reproduce them on NVIDIA's Triton Inference Server image. python bench/zoo.py <zoo checkout> refreshes this page's copy.

To rerun the generated-code benchmark (bench/setup-pod.sh prepares a fresh GPU machine; bench/README.md lists the published environment):

bash
cd python/linnet && uv sync --all-extras
cd ../..
LINNET_BIN=build/release/linnet python bench/run.py --device cuda --out bench/results/latest.json

Released under the MIT License.