Skip to content

Serving and tools ​

Serve a Linnet model with Linnet's own batching engine, or export it to Triton Inference Server, vLLM, SGLang, TGI, llama.cpp, Ollama, or ComfyUI. The framework adapters have their own pages: PyTorch, JAX, ONNX.

Continuous batching ​

python
from linnet import nest
from linnet.serve import Engine, Request

model = nest.load("llama-3.1-8b-instruct", device="cuda",
                  generics={"Batch": 64, "MaxSeq": 2048})
engine = Engine(model)
done, stats = engine.run([Request(prompt=ids, max_new_tokens=256) for ids in prompts])
stats.tokens_per_second, done[0].tokens, done[0].ttft

linnet.serve.Engine decodes many requests together, each at its own length. A request takes a free cache row and leaves at its token budget, an end-of-sequence token, or the end of the cache. The model's Batch generic sets the most requests in flight, and MaxSeq the longest prompt plus completion. With linnet.torch the step replays as a CUDA graph; with linnet.jax.load_model (nest.load(..., backend="jax_model")) it is an XLA program.

Entries ​

Every decoder in the zoo has the two required entries.

EntryDoes
prefill_slots<M, S>(tokens: [M, S], slots: [M], lengths: [M]) -> [M, Vocab]writes M padded prompts into cache rows slots in one pass; lengths says where each ends
decode_rows(tokens: [Batch, 1], positions: [Batch]) -> [Batch, Vocab]one token for every row, each at its own position
prefill_packed<P>(tokens: [P], rows: [P], positions: [P], segments: [P], last: [Batch]) -> [Batch, Vocab]optional: prompts packed end to end into one pass of P tokens, each token with its cache row, position, and prompt; returns the logits after each prompt's last token, last[m]
step_packed<P>(tokens, rows, positions, segments, last, step_tokens: [Batch, 1], step_positions: [Batch]) -> [2 * Batch, Vocab]optional, with prefill_packed: that pass and a decode_rows step in one; returns the logits after each prompt, then each row's step
prefill_paged<P, Rows, Pages>(tokens: [P], positions: [P], rows: [P], slots: [P], last: [Rows], table: [Rows, Pages]) -> [Rows, Vocab]for pages: prompts packed end to end, each token written at its place in the pool (slots) and attending over its row's pages up to its position
decode_paged<Rows, Pages>(tokens: [Rows, 1], positions: [Rows], table: [Rows, Pages]) -> [Rows, Vocab]for pages: one token for every row, its positions in the pages its row of table lists
step_paged<P, Rows, Pages>(tokens, positions, rows, slots, last, step_tokens, step_positions, table) -> [2 * Rows, Vocab]for pages: prefill_paged and decode_paged in one pass

Build them from std.nn.cache::write_slots, write_rows, write_tokens, and page_slots, and std.nn.attention::grouped_attention_rows, paged_attention, and paged_prefill_attention. Every zoo decoder but Phi-3 and gpt-oss has the paged entries.

Engine ​

Call or optionDoes
Engine(model, pack=4096)with prefill_packed, packs waiting prompts into passes of up to pack tokens; otherwise prompts go in passes of up to 8, each padded to its longest
Engine(model, mix=True)with step_packed, runs the decoding rows' step inside a prompt pass that fits the second compiled size (1024 tokens by default); mix=False keeps them apart
Engine(model, share=True)with packed passes, identical prompts admitted together pass once: the rest copy its cache rows and draw their own tokens; with pages, prompts share the pages of a common start instead (see below); share=False passes each
engine.run(requests)submit for every request, then step until nothing is left
engine.submit(request)queues a request; returns its Completion, which fills in as it runs
engine.step()admits waiting requests into free rows, runs one decoding step, and returns the requests that finished, one step after they do
engine.busywhether anything is left
engine.cancel(completion)ends a request early (at a stop string, or when its client has gone) and frees its row
engine.load_weights(model)copies a PyTorch model's weights (adapters merged in) into the served model, in place, between runs; compiled passes and CUDA graphs stay

Pages ​

python
model = nest.load("llama-3.1-8b-instruct", device="cuda",
                  generics={"Batch": 1, "MaxSeq": 131072})
engine = Engine(model, rows=128, max_len=8192)

Loaded with Batch 1, a card with the paged entries keeps its caches as one pool of MaxSeq positions, in pages of PageSize (64 by default). A request takes pages as it grows and gives them back when it ends, so it holds the cache it uses, not a whole row. Requests shorter than the longest allowed fit more of them in the same memory.

OptionDoes
rows=64the most requests in flight
max_lenthe longest prompt plus completion; by default every page but one, up to 8192
reservepages kept free for the rows already decoding; 1% by default
chunkthe most prompt tokens a step passes while rows decode; by default the second compiled size
pagedTrue or False to choose; by default pages when the card has the entries and Batch is 1

Prompts pass in chunks. They fill passes in the order they arrive, and a prompt that does not fit goes on in the next pass, so a pass is full whatever the lengths and a prompt can be longer than pack. While rows decode, a step passes chunk prompt tokens at most, so a long prompt does not stall them.

Full pages are kept by the tokens in them and before them. A prompt that starts like an earlier one (the same system prompt, a conversation's next turn, n samples of one prompt) reads those pages instead of passing them again, up to its last page (completion.cached, stats.cached). Pages nobody uses stay kept until the pool needs them, the oldest first. load_weights drops them.

When the pool runs short, the request admitted last gives its pages back and waits. Once pages are free it passes its prompt and its tokens so far as one longer prompt, most of it found among the pages it gave back, and continues with the same tokens it would have drawn (completion.preempted, stats.preempted). On CUDA, FlexAttention reads the pages where they lie in the pool, a prompt's tokens 128 at a time.

Llama 3.1 8B on one H100, 256 requests of 128 to 512 prompt tokens and 128 new tokens each, greedy:

CacheRequests in flightTokens per secondMedian first token
rows of 640 positions645,5742.52 s
pages of one pool, the same 40,960 positions645,5652.52 s
the same pool1285,7122.52 s
rows of 8,192 positions, 48 GiB484,1983.03 s
pages, max_len=8192, the same 48 GiB1287,6451.69 s

Pages cost nothing when every row fits its requests. They pay off once a server allows long requests: rows then hold max_len each.

linnet serve MODEL --pool 131072 --batch 128 --max-seq 8192 serves the same way.

Requests ​

python
Request(prompt=ids, max_new_tokens=256, temperature=0.8, top_p=0.95, top_k=50, seed=7)
FieldDoes
temperaturesamples on the device; without it, decoding is greedy
top_k, top_pkeep the top_k most likely tokens, then the fewest whose probabilities reach top_p (the order Transformers filters in)
seedthe same tokens alone or in any batch, on PyTorch, JAX, or ONNX Runtime; without it the engine picks one (completion.sampling.seed)
logprobs=keach token's log probability and the k most likely alternatives
on_tokencalled with the completion as each token is read, reason set on the last; use it to stream

For Llama 3.1 8B at 64 rows, a temperature costs under 1% of throughput and top-p about 6% (PyTorch) and 8% (JAX). On ONNX Runtime, sampling runs on the host in NumPy.

Prepared weights ​

linnet torch --prepare and linnet jax --prepare move work that reads only weights, such as dequantizing a quantized checkpoint, into a prepare function that runs once at load instead of on every call. A model loaded with trainable=True keeps that work in the graph, where gradients reach it.

Over HTTP ​

bash
uv add "linnet-lang[serve]"
linnet serve llama-3.1-8b-instruct --batch 32 --max-seq 4096 --port 8000
linnet serve Qwen/Qwen2.5-7B-Instruct          # a transformers checkpoint, converted

linnet serve runs python -m linnet.serve with the Python it was installed beside (LINNET_PYTHON picks another). It takes anything linnet.nest.load does and serves it with OpenAI's API: GET /v1/models, POST /v1/completions and POST /v1/chat/completions (server-sent events with "stream": true, and stream_options.include_usage), and GET /health. OpenAI's clients work against it as they are:

python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="unused")
reply = client.chat.completions.create(
    model="llama-3.1-8b-instruct",
    messages=[{"role": "user", "content": "Name three primary colors."}],
    max_tokens=64,
)
OptionDoes
--batch, --max-seqthe Batch and MaxSeq generics
--backend jax, --backend onnxserves through that runtime instead of PyTorch
--tokenizerthe tokenizer and chat template's repository (default: the card's checkpoint repository)
--warmup 128 512compiles those prompt lengths before the first request; others compile when first seen
--portthe HTTP port
Request fieldDoes
max_tokens (max_completion_tokens in a chat)the token budget
temperature1 unless given, as OpenAI has it
top_p, top_k, seedas in Request
stopup to four strings, cut from the text and never streamed
nthat many choices, each its own request in the batch; with a seed, choice i draws with seed + i
logprobs (completions), logprobs and top_logprobs (chat)each token's log-probability and up to 20 alternatives, in OpenAI's two formats, from the logits before temperature and filtering
echo (completions)the prompt before the completion's text; refused together with logprobs

Log probabilities cost nothing when no request asks for them. To serve an engine of your own, use linnet.serve.server.Server(engine, tokenizer, name=...).

Triton Inference Server ​

bash
python -m linnet.triton export gpt2 -o model_repository --bind B=1 --bind S=64
tritonserver --model-repository=model_repository

linnet.triton.export writes a model directory the server loads as is: config.pbtxt from the entry's signature and bindings, plus a version directory holding the model. The model is a Nest name or directory, or a .linnet file with --weights and --bindings.

BackendWhat is writtenUse it for
--backend onnx (default)1/model.onnx with the checkpoint as initializers (external data beyond 1 GB), platform: "onnxruntime_onnx"stateless entries such as forward
--backend python1/model.py plus the source and weights, run through linnet.torchentries with state (KV caches), any dtype, numerics="fast"

Shapes are static (max_batch_size: 0), so export one model per shape or batch size. Entry generics such as B and S come from --bind, or from the card's [check] table for a Nest model, which --bind overrides. The Python backend needs linnet-lang[torch] and the compiler in the server's environment (LINNET_BIN, LINNET_STD), and returns bf16 results as f32.

python
from linnet import triton

repository = triton.export("models/tinyllama-1.1b-chat", "model_repository",
                           generics={"S": 128}, backend="python", numerics="fast")
repository.inputs      # (Tensor(name='tokens', dtype='i32', dims=(1, 128)),)

For the ONNX model alone, use linnet.onnx.export_model (ONNX).

vLLM, SGLang, TGI ​

bash
python -m linnet.hf export tinyllama-1.1b-chat -o serve/tinyllama
vllm serve serve/tinyllama

linnet.hf.export writes a Transformers checkpoint directory (config.json, model.safetensors, tokenizer files) for a model whose structure matches one of these families:

FamilyStructure
llamagrouped-query attention, RMS norms, SwiGLU; biases on all four attention projections or none, on all three MLP projections or none
qwen2Llama's layout with biases on the query, key, and value projections only
qwen3an RMS norm over each query and key head, and a head width of its own (HeadDim)
phi3one projection for query, key, and value, one for gate and up; as many key/value heads as query heads
gpt2learned positions, biased LayerNorm, GELU MLP

Any other structure is refused, naming the paths that did not fit; there is no general vLLM model class for arbitrary Linnet programs yet. For a .linnet file, pass --weights, --bind, and --tokenizer <repo>. Define THETA for rope_theta, plus FACTOR, LOW_FREQ_FACTOR, HIGH_FREQ_FACTOR, and ORIGINAL_MAX_POSITION_EMBEDDINGS for Llama 3.1's llama3 rope_scaling, and give every norm the same epsilon. Exports attend over the whole sequence, as the programs do, so Phi-3's export sets sliding_window: null.

Tested:

  • transformers matches the Linnet interpreter to 1e-4 on exported tiny GPT-2, Llama, Llama 3.1, Qwen2, Qwen3, and Phi-3 configurations.
  • Qwen2.5 0.5B, Qwen3 4B and 8B, and Phi-3 mini export to the Hub checkpoints' own tensors: bitwise-identical transformers logits for Qwen3 8B, the same greedy tokens in vLLM 0.30 (Qwen3 8B in eager mode), and the same GGUF tensors from llama.cpp's converter.
  • vLLM 0.30 on an H100 served the tinyllama-1.1b-chat export as is, chat template included.
  • In the benchmarks, each Llama-family and GPT-2 export gives the same first token as the original in vLLM, SGLang, and TGI, at the same speed within run-to-run noise.

llama.cpp and Ollama ​

bash
python -m linnet.gguf export tinyllama-1.1b-chat -o serve/tinyllama \
    --converter ~/llama.cpp/convert_hf_to_gguf.py --outtype q8_0
ollama create tinyllama -f serve/tinyllama/Modelfile

linnet.gguf.export writes the Transformers checkpoint above, converts it with llama.cpp's convert_hf_to_gguf.py (--converter, or LLAMA_CPP pointing at a llama.cpp checkout), and writes an Ollama Modelfile next to the GGUF file. The chat template carries over when Ollama has an equivalent (ChatML, Zephyr, Llama 2). The families are the same as for vLLM.

ComfyUI ​

linnet-comfyui is a custom node pack. Install it into custom_nodes/:

bash
cd ComfyUI/custom_nodes
git clone https://github.com/franknoh/linnet-comfyui
pip install -r linnet-comfyui/requirements.txt
Node
Load Linnet modela .linnet file or package, generics as JSON, a SafeTensors path, numerics, compile, device; outputs a LINNET_MODEL
Load from Nesta registry name; downloads the source and checkpoint
Run entrya model, an entry name, up to four TENSOR inputs, entry generics as JSON; outputs the results
Tensor from list / to stringJSON lists and readable summaries for token ids and small tensors
Image to tensor / Tensor to imageComfyUI IMAGE ([B, H, W, C] float) to [B, C, H, W] in a dtype, and back
Argmaxthe token ids of logits
Model infoprint(model): the block tree with names and shapes
Reset statezeroes the model's state members between generations

Each node checks the entry's signature before it runs, so a wrong shape fails at the node, not inside a kernel. The compiler comes from LINNET_BIN or PATH.

Released under the MIT License.