Prefill/decode disaggregation over NIXL¶
examples/pd-disaggregation/llmd_nixl_pull.sq is llm-d's prefill/decode split on vLLM: the router
that decides which pod prefills and which decodes, the sidecar that sends
the prompt to one and the decode request to the other, and the two
schedulers that hand the KV over with the NIXL connector. Every line is
checked against the source: llm-d at 8a2f37d, its router
(llm-d-inference-scheduler) at 13eebdb, vLLM at 0c87a197 (ref/vllm;
the vLLM citations are hashed by make check, the llm-d ones are not).
It is a specification checked against the code, not yet against a scheduler oracle. What the program cannot say without two new statements, and what it still cannot say, is in The KV transfer.
The router picks a prefill instance and a decode instance, each a box with
its engine, its NIC and its pools. The read crosses between them: it holds
the prefiller's NIC and the decoder's at once, which is the bracket around
P.nic[i] and D.nic[j], and moves the prompt's KV from the
prefiller's leased blocks to the decoder's (P.kv[i] → D.kv[j]) after the
decoder's read latency.
The deployment¶
llm-d's P/D guide (guides/pd-disaggregation, llm-d 8a2f37d) deploys
this:
| Piece | Configuration | Where |
|---|---|---|
| model servers | 8 prefill pods (vllm serve, TP 1) and 2 decode pods (TP 4), gpt-oss-120b, --block-size 128, --kv-transfer-config '{"kv_connector":"NixlConnector","kv_role":"kv_producer" / "kv_consumer","kv_connector_extra_config":{"backends":["UCX"]}}', VLLM_NIXL_SIDE_CHANNEL_HOST the pod IP |
modelserver/gpu/vllm/base/patch-prefill.yaml, patch-decode.yaml |
| router (EPP) | disagg-profile-handler with always-disagg-pd-decider (every request is prefilled remotely); decode profile decode-filter → active-request-scorer → max-score-picker (the least busy decoder); prefill profile prefill-filter → prefix-cache-affinity-filter (approx-prefix-cache-producer, a cache-warm prefiller first) → token-load-scorer → max-score-picker; one EPP replica |
router/pd-disaggregation.values.yaml |
| sidecar | llm-d-routing-sidecar on each decode pod, connector nixlv2: prefill leg (max_tokens = 1, do_remote_decode), wait, decode leg with the prefiller's block ids; serial |
pkg/sidecar/proxy/connector_nixlv2.go (router 13eebdb) |
| alternative decider | prefix-based-pd-decider with nonCachedTokens, promptTokens: remote prefill only when the uncached suffix on the chosen decoder is long enough |
docs/disaggregation.md, profilehandler/disagg/README.md |
The examples below run on one prefiller or two, one decoder or two, with Qwen3-8B and block 16. What differs from the guide is written next to each number.
In serQ¶
// Prefill/decode disaggregation as llm-d runs it on vLLM: the llm-d router
// (llm-d/llm-d at 8a2f37d, llm-d/llm-d-inference-scheduler at 13eebdb) in
// front of NP prefill and ND decode instances that hand the KV over with the
// NIXL connector in its pull mode (ref/vllm at 0c87a197,
// `vllm/distributed/kv_transfer/kv_connector/v1/nixl/`).
//
// One request: the router picks a decode pod (the least busy) and asks its
// prefix-cache estimate how much of the prompt that pod already has; a long
// enough uncached suffix means a remote prefill, for which it picks a prefill
// pod (a cache-warm one, then the least loaded); a short one means the decode
// pod prefills the prompt itself. On a remote prefill the decode pod's sidecar
// sends the prompt to the prefiller with max_tokens = 1, waits for its
// answer, and only then sends the decode request to its own engine. The
// prefiller's scheduler frees the request's slot when the one token is
// sampled but keeps its blocks leased for the decoder's read; the decoder's
// scheduler takes the request when the whole prompt fits in its blocks and
// parks it (WAITING_FOR_REMOTE_KVS) while its worker reads the KV from the
// prefiller; the read done, the prefiller frees the lease (its full blocks
// stay cached), and the decoder's request queues once more for a running slot
// and a token of budget, recomputes the last prompt token for its logits,
// and decodes. docs/use-cases/pd.md has the line-by-line correspondence.
use "../../lib/vllm.sq";
let NP = 2; // prefill instances
let ND = 2; // decode instances
let B = 8192; // max_num_batched_tokens
let max_seqsP = 16; // max_num_seqs on a prefiller
let max_seqsD = 64; // max_num_seqs on a decoder
let blocksP = 10000; // num_gpu_blocks on a prefiller
let blocksD = 10000; // num_gpu_blocks on a decoder
let bs = 16; // block_size
let max_model_len = 16384; // every KV pool holds one request of it: vLLM refuses to start otherwise (kv_cache_utils.py:864-900, called at kv_cache_utils.py:2742)
let a = 2e-5; // compute per scheduled token (s)
let omega = 2e-4; // weights read per iteration (s)
let beta = 2e-9; // KV read per iteration per resident context token (s)
let c0 = 0; // fixed cost per iteration (s)
let BwP = 2e5; // NIXL bandwidth out of a prefiller, tokens/s (128 KB per token at 25 GB/s)
let BwD = 2e5; // and into a decoder (one NIC each: the tensor-parallel fan-out of the read is not modelled)
let x0 = 0.002; // per-read latency: posting the READ and the step that polls its completion (s; nixl/pull_worker.py:545-558)
let thr = 512; // prefix-based-pd-decider nonCachedTokens (1: the guide's always-disagg-pd-decider)
let minp = 0; // prefix-based-pd-decider promptTokens
let Lambda = 0.6; // sessions per second
let p = 0.9; // continue after a turn
let Z = 3.0; // tool time (s)
// The deployment: what the router, the sidecar and the two schedulers do
// with one request. The router and the sidecar are the gateway; each
// scheduler is its queue's entries, and the two NICs of a read are links.
queue gw : gateway {
route {
set t0 = now;
set prompt = K + n;
set o = min(o, max_model_len - prompt); // generation stops at max_model_len tokens (sched/utils.py:114-120)
// The router. Decode profile first: the least busy decode pod
// (active-request-scorer). Then the decider: a remote prefill only when the
// prompt's uncached suffix on that pod is long enough
// (prefix-based-pd-decider: nonCachedTokens, promptTokens); the estimate is
// the router's own prefix-cache picture of the pod.
choose j in ND by (holders(D[j].kv) + queued(D[j].kv));
set hitD = min(cachedin(D[j].kv), reusable(prompt, bs));
set remote = prompt >= minp && prompt - hitD >= thr;
branch (remote) {
// Prefill profile: a pod that has the prefix (prefix-cache-affinity-filter),
// then the least loaded (token-load-scorer).
choose i in NP by (cachedin(P[i].kv) > 0 ? 0 : 1, work(P[i]) + queued(P[i].reqs));
P[i].prefill (prompt); // the sidecar: max_tokens = 1, wait for the answer
set t1 = now;
D[j].decode (prompt) from P[i]; // then the decode request, its KV read from P[i]
observe lease = D.read_done - t1; // how long the prefiller's blocks stayed leased
} else {
D[j].decode (prompt);
}
observe ttft = D.first_token - t0;
observe remote = remote;
observe response = now - t0;
}
}
// A prefill instance: vLLM's engine on one device (examples/multi-turn/vllm.sq)
// with no decode. The scheduler admits a request with a slot and the blocks
// of its first chunk, and room for the whole prompt
// (scheduler_reserve_full_isl). When the one token is sampled the request is
// finished on the prefiller: the slot goes and the blocks stay leased for the
// decoder's read (request_finished: delay_free_blocks), 30 s renewed by the
// decoder's heartbeats while it waits, so in effect until the read.
queue P[NP] : prefill {
pool reqs { cap max_seqsP; admit via P; }
pool kv { cap blocksP * bs; block bs; evict lru; preempt lifo; }
serve step {
budget B;
cost c0 + max(omega + beta * (kv_decode + kv_prefill), tokens * a);
memory kv;
}
nic ps(BwP); // the pod's NIC: the reads out of it share it
prefill (prompt) {
hold reqs (1), kv (min(prompt, hit + budget_left(P))) reserve (prompt)
at admission (hit = min(cachedin(kv), reusable(prompt, bs))) {
set c = cached;
prefill (prompt - c) growing kv; // chunked; one token is sampled and discarded
} cache (prompt) lease kv (inf);
}
}
// A decode instance: vLLM's engine on one device, with two ways in. Its
// scheduler serves its queues in this order: the requests whose KV has
// arrived first (vLLM's `skipped_waiting`, scheduler.py:2383-2385), then the
// new ones, and the first that does not fit stops the step's admissions.
queue D[ND] : decode {
pool reqs { cap max_seqsD; admit via D; }
pool kv { cap blocksD * bs; block bs; evict lru; preempt lifo; admit via D; }
serve step {
budget B;
cost c0 + max(omega + beta * (kv_decode + kv_prefill), tokens * a);
memory kv;
}
nic ps(BwD); // the pod's NIC: the reads into it share it
// Decode only: the pod prefills the prompt itself (examples/multi-turn/vllm.sq).
decode (prompt) {
hold kv (min(known, hit + budget_left(D))) reserve (known), reqs (1)
at admission (known = computed < prompt ? prompt : computed + 1,
hit = min(cachedin(kv), reusable(known, bs))) {
set c = min(cached, reusable(known, bs));
observe local_tokens = known - c;
prefill (known - c) growing kv;
branch (known == prompt) { mark first_token; } // not a resumed request's re-prefill
decode (o - 1 - (known - prompt)) growing kv;
} cache (prompt + o);
}
// With remote KV. The sidecar's decode request is taken when the whole
// prompt fits the blocks, at a step with budget left and a running slot
// free - it takes neither (the request is parked, not running) - and
// reuses the full blocks it has of the prompt.
decode (prompt) from src {
set transferred = 0; // one KV transfer per request (nixl/pull_scheduler.py:187-189)
hold kv (known) reserve (known), reqs (0) reserve (1)
reuse (reusable(known, bs))
at admission (known = computed < prompt ? prompt : computed + 1) {
set c = cached;
branch (!transferred) {
// WAITING_FOR_REMOTE_KVS: every block past the local hit is read
// from the prefiller over NIXL, through its NIC and this pod's at
// once; the read done, the prefiller's lease ends. The last prompt
// token is recomputed for its logits.
observe remote_tokens = prompt - c;
transfer (prompt - c) from src to kv (prompt - 1 - c); // over P's NIC and this pod's
mark read_done;
set transferred = 1;
set c = prompt - 1;
}
// Back in the waiting queue: a running slot and a token of budget.
// A request preempted on the decoder afterwards recomputes what it
// lost here, as any local prefill would (do_remote_prefill is spent).
hold reqs (1) {
prefill (known - c) growing kv;
branch (known == prompt) { mark first_token; }
decode (o - 1 - (known - prompt)) growing kv;
}
} cache (prompt + o);
}
}
// The decoder's worker reads the KV (NIXL pull): it posts the READ and finds
// it done at a later step, a fixed wait per read, and the bytes cross the
// prefiller's NIC and the decoder's at once. How concurrent reads divide the
// two NICs is UCX's and the fabric's, not vLLM's: max-min fairness is the
// model, not a measurement (docs/design/bandwidth-sharing.md).
D pull P latency x0 share maxmin;
stage tool : delay;
// The client: sessions arrive, each turn brings a prompt of n new tokens on
// a context of K and wants o tokens back, and after a turn the session calls
// a tool and comes back with probability p.
workload {
arrive poisson(Lambda);
hidden o; // the scheduler knows max_tokens, not the length (sched/utils.py:98-119)
init { set K = 0; }
turn {
set n = K == 0 ? ~uniform(1000, 3000) : ~exp(500);
set o = ~exp(200) + 1;
set more = ~bernoulli(p);
}
session {
turn;
loop {
// vLLM refuses a prompt of max_model_len tokens or more before it is
// scheduled (input_processor.py:512-536): the session ends there.
observe too_long = K + n >= max_model_len;
branch (K + n >= max_model_len) { end; }
request gw;
set K = prompt + o;
branch (more) { tool (~exp(Z)); turn; } else { end; }
}
}
}
// How evenly the router spreads the decoders, read at every instant: the
// time average of a spread is not the spread of the per-pod averages. The
// KV usage is vLLM's `kv_cache_usage` (block_pool.py:879-890, of the free
// blocks; a cached block nobody holds is one, block_pool.py:42-45) but for
// its null block, which vLLM takes out of
// num_gpu_blocks and the pool here does not; the load is the
// active-request scorer's.
gauge running_spread = max k in ND (used(D[k].reqs)) - min k in ND (used(D[k].reqs));
gauge kv_usage_max = max k in ND (used(D[k].kv) / (used(D[k].kv) + free(D[k].kv)));
gauge kv_usage_mean = sum k in ND (used(D[k].kv) / (used(D[k].kv) + free(D[k].kv))) / ND;
gauge load_spread = max k in ND (holders(D[k].kv) + queued(D[k].kv))
- min k in ND (holders(D[k].kv) + queued(D[k].kv));
run { horizon 2000; warmup 200; seed 1; }
Line by line¶
The router and the sidecar, against llm-d¶
| llm-d | serQ | Where |
|---|---|---|
| the endpoint picker runs the decode profile first and picks a decode pod; the guide's decode profile scores by active requests (the least busy) | choose j in ND by (holders(D[j].kv) + queued(D[j].kv)) |
disagg_profile_handler.go:316-334; guides/pd-disaggregation/router/pd-disaggregation.values.yaml |
the decider: a remote prefill when the prompt's uncached suffix on the chosen decode pod is at least nonCachedTokens, and the prompt at least promptTokens; the cached part is the router's own estimate of the pod's prefix cache |
set hitD = min(cachedin(D[j].kv), reusable(prompt, bs)); set remote = prompt >= minp && prompt - hitD >= thr; |
prefix_based_pd_decider.go:266-303; disagg_profile_handler.go:353-366. The guide runs always-disagg-pd-decider, which is thr = 1 |
the prefill profile: pods that have the prefix (prefix-cache-affinity-filter), then the least loaded (token-load-scorer) |
choose i in NP by (cachedin(P[i].kv) > 0 ? 0 : 1, work(P[i]) + queued(P[i].reqs)) |
the same values file |
the decode pod's sidecar sends the prompt to the prefiller with max_tokens = 1 and do_remote_decode, waits for the answer, then sends the decode request to its own engine with the prefiller's block ids |
P[i].prefill (prompt); then D[j].decode (prompt) from P[i]; — the prefiller's entry, its blocks leased at its end, then the decoder's |
connector_nixlv2.go:69-232 (the prefill leg, CapSingleToken at 145), 261-379 (the decode leg) |
| no prefill header: the request goes to the decode pod's engine as it is | the else branch, D[j].decode (prompt);: vLLM's engine on one device (examples/multi-turn/vllm.sq) |
dispatch.go:196-213 |
| the two legs in parallel, so the decoder allocates while the prefiller works | not written: a session waits at one pool at a time (The KV transfer) | connector_nixlv2.go:60-67 (MoRI-IO write mode only); vLLM's push-mode proxy, disagg_proxy_pushconnector_demo.py:227-270 |
The two schedulers, against vLLM¶
| vLLM | serQ | Where |
|---|---|---|
a prompt of max_model_len tokens or more is refused before it is scheduled; a generation stops at max_model_len tokens; no KV cache smaller than one request of max_model_len is started |
branch (K + n >= max_model_len) { end; } before request gw;; set o = min(o, max_model_len - prompt) in the route; max_model_len = 16384 below every pool |
input_processor.py:512-536; sched/utils.py:114-120; kv_cache_utils.py:864-900, called at kv_cache_utils.py:2742 |
| the prefiller admits like any vLLM engine: a slot, the blocks of the first chunk, room for the whole prompt, the prefix hit looked up when the scheduler takes the request | hold reqs (1), kv (min(prompt, hit + budget_left(P))) reserve (prompt) at admission (hit = …) in P's prefill entry |
the waiting loop, scheduler.py:868-1128; the vLLM use case |
| the prefiller computes the prompt in chunks and samples one token, which the sidecar discards | prefill (prompt - c) growing kv |
scheduler.py:624-823; the truncation for Mamba and MTP only, nixl/base_scheduler.py:409-436 |
the request finishes on the prefiller: its slot is freed, its blocks are not — request_finished returns delay_free_blocks and a lease of kv_lease_duration (30 s), renewed by the decoder's heartbeats while the request waits |
} cache (prompt) lease kv (inf); — the scope ends, the slot goes, the blocks stay the session's |
nixl/pull_scheduler.py:191-292; _free_request, scheduler.py:2628-2657; the renewal, nixl/base_scheduler.py:199-238, nixl/base_worker.py:3010-3030 |
| the prefiller keeps every computed full block of the prompt cached once the lease ends | cache (prompt) on that hold, applied when the lease ends |
_connector_finished, scheduler.py:2929-2982; kv_cache_manager.py:610-619 |
| the decoder's scheduler looks at its waiting queue only at a step with budget left and a running slot free | admit via D on both of the decoder's pools; reqs (0) reserve (1) — a slot must be free, none is taken |
scheduler.py:872-879 |
| the decoder's local prefix hit, then the connector: for a remote prefill every prompt token beyond the local hit is external and loaded asynchronously | kv (known) reserve (known) with reuse (reusable(known, bs)) in D's decode … from entry; c = cached is the local hit |
scheduler.py:932-954; nixl/pull_scheduler.py:34-66 |
blocks are allocated for the whole prompt, and the request is parked, WAITING_FOR_REMOTE_KVS, holding them and no slot; one transfer per request |
the hold on kv; transferred is do_remote_prefill, spent |
scheduler.py:1199-1226, 1264-1294; nixl/pull_scheduler.py:108-189 |
| the worker reads the blocks from the prefiller over NIXL (pull mode: a NIXL READ issued by the decoder) | transfer (prompt - c) from src to kv (prompt - 1 - c) in D's entry, and D pull P latency x0 share maxmin;: the decoder reads; a fixed wait, then the bytes out of the prefiller's NIC and into the decoder's at once, the two shared max-min fairly (a model, not a measurement: Bandwidth sharing) |
nixl/pull_scheduler.py:168-177; the worker's _read_blocks, nixl/pull_worker.py:392-575 |
| the read done, the blocks are cached, the last prompt token is marked uncomputed (its logits are needed), and the request is back in the waiting queue, served before new arrivals | the load D.kv[j] (prompt - 1 - c) the transfer stands for; hold reqs (1), with reqs declared before kv |
_update_waiting_for_remote_kv, scheduler.py:3032-3077; _try_promote_blocked_waiting_request, scheduler.py:3079-3092; scheduler.py:2383-2385 |
| the prefiller frees the leased blocks when the read completes | the release P.kv[i] the transfer stands for takes the lease |
_update_from_kv_xfer_finished, scheduler.py:3113-3138 |
| the decoder recomputes the last prompt token and decodes; a request preempted afterwards is rescheduled without a second transfer, prefilling locally what it lost | prefill (known - c) growing kv; decode (o - 1 - (known - prompt)) growing kv; with known from computed |
scheduler.py:1560-1561; nixl/pull_scheduler.py:187-189 |
a parked request is in no running list and is not preempted; nor is the prefiller's finished one |
preempt lifo takes the last admitted resident of the engine |
scheduler.py:742-813 |
| the decoder keeps the prompt's and the output's full blocks cached | } cache (prompt + o) |
kv_cache_manager.py:602-606 |
Writing xPyD¶
The deployment is queues: NP prefill instances and ND decode instances,
each with its own KV pool, request-slot pool, engine and NIC, and one
gateway selected by request gw; in the workload's session.
queue gw : gateway { route { … } } // the router and the sidecar
queue P[NP] : prefill { pool reqs { … admit via P; } pool kv { … } serve step { … memory kv; } nic ps(BwP); prefill (prompt) { … } }
queue D[ND] : decode { pool reqs { … admit via D; } pool kv { … admit via D; } serve step { … memory kv; } nic ps(BwD); decode (prompt) { … } decode (prompt) from src { … } }
D pull P latency x0 share maxmin; // the decoder reads the KV from the prefiller
stage tool : delay;
A queue's pools are its members': P[i].kv is the KV of the i-th
prefiller, admit via P inside P means the member's own scheduler serves
the pool, and memory kv on its stage counts the member's own blocks as its
residents' memory (kv_decode, kv_prefill). The family's size is the
let the router chooses over (queue D[ND], choose j in ND), so the
number is written once. Each pod's NIC is its own (nic ps(BwD);, the
stage D.nic), and D pull P latency x0 share maxmin; is the transfer's
topology and mode in one line: the decoder reads the KV from the
prefiller, waits x0 before each read, and concurrent reads divide the two
NICs max-min fairly.
The router is the gateway's two chooses, and the instance it picked is
the queue the request is handed to: P[i].prefill (prompt), then
D[j].decode (prompt) from P[i]. Inside an entry nothing is indexed — kv,
reqs, P are the member's own. from P[i] is the pool P's entry
leases, and the read's NICs are the relation's: the source member's
P.nic[i] and the reader's own D.nic[j]. The transfer's load and
release name the pool as the holds did, index included. The admission, the
allocation and the two runs are the entries' and appear once each: the
router's body has no hold.
Going from 2P2D to 4P8D is the two lets; the program does not change
otherwise, which is the point of writing the router as choose over a
family rather than as a branch per instance (examples/multi-turn/routing.sq
still has the branch-per-policy shape the design notes call a smell).
The two modes¶
Pull (NixlConnector, kv_role producer and consumer): the decoder
reads. The prefiller must be done before the decoder can be told where the
blocks are, so the sidecar's dispatch is serial: prefill leg, answer, decode
leg. The program above is this mode, on the llm-d guide's deployment
(guides/pd-disaggregation/modelserver/gpu/vllm/base/patch-decode.yaml).
Push (NixlPushConnector): the decoder allocates and registers its
block ids with the prefiller over a NIXL notification
(nixl/push_scheduler.py:128-205); the prefiller writes the KV when it has
both a finished request and its registration (nixl/push_scheduler.py:207-294)
and frees the lease when the write completes (nixl/push_scheduler.py:342-357).
The decoder needs nothing from the prefiller's answer, so the two legs can
be sent at once, and vLLM's push-mode proxy does
(disagg_proxy_pushconnector_demo.py:227-270): the decoder allocates during
the prefill and the write starts the moment it ends. Through the llm-d
sidecar the dispatch is serial for NIXL (parallel for MoRI-IO only,
connector_nixlv2.go:60-67), and then the two modes have the same
lifecycle: the same holds in the same order, the copy moved by the
prefiller's worker instead of the decoder's, one notification more.
The two modes in the program¶
Pull is the program as written. The prefiller's scope ends in a lease, the decoder is admitted, and the decoder's worker reads over both NICs:
P[i].prefill (prompt); // hold reqs (1), kv (…) … { prefill (…) growing kv; } cache (prompt) lease kv (inf);
D[j].decode (prompt) from P[i]; // hold kv (known) reserve (known), reqs (0) reserve (1) … {
// transfer (prompt - c) from src to kv (prompt - 1 - c); // D READs over both NICs; the lease ends
// hold reqs (1) { … }
// } cache (prompt + o);
The relation has no push yet: the push a reader expects is vLLM's proxy,
where the decoder allocates during the prefill, and that needs a
reservation the language does not have (below). Push through the llm-d
sidecar (serial dispatch) can be written with link queues and transfer
on, the same three lines with two differences a reader can see: the copy is the prefiller's WRITE
(nixl/push_worker.py:714-722), which crosses the same two NICs but is
posted by the prefiller's worker, so the latency is egress's; and it
starts one notification after the decoder's admission (the registration,
nixl/push_scheduler.py:128-205):
queue egress[NP] : link { serve ps(BwP) latency x0; } // the prefiller's worker posts the WRITE
queue ingress[ND] : link { serve ps(BwD); }
…
decode (prompt) from src { // D's entry
hold kv (known) reserve (known), reqs (0) reserve (1) … {
run register (x_reg); // the decoder's block ids reach the prefiller
transfer on egress[src], ingress[self] (prompt - c) from src to kv (prompt - 1 - c);
hold reqs (1) { … }
} cache (prompt + o);
}
Which NIC saturates is not the mode's: it is the ratio of prefillers to decoders, where the router sends the reads (the prefill profile prefers the pod that has the prefix, so one prefiller's egress carries most of them), and the sharing policy.
Push with the two legs dispatched at once (vLLM's own push proxy) is the form the language cannot yet write: the decoder's admission would have to be requested when the request arrives, while the session is still queued at the prefiller, so the write can start the moment the prefill ends. A session waits at one pool at a time, so the program above writes the decoder's admission after the prefill. The difference is bounded: under decoder memory pressure both forms lease at the prefiller, since the write cannot start before the decoder has allocated; without it, the concurrent form takes the decoder's blocks a prefill earlier and saves one decoder step of latency. The reservation that would write it exactly is sketched in The KV transfer:
book kvD[j] (prompt) reserve (prompt); // join the decoder's queue now, not written yet
hold reqsP[i] (1), kvP[i] (…) … { … } cache (prompt) lease kvP[i] (inf);
hold kvD[j] { transfer on egress[i], ingress[j] (…) from kvP[i] to kvD[j] (…); … } // open the booking, waiting if it is not granted
Not modelled: the lease's expiry and the decoder's heartbeats (the
lease is granted at nixl/pull_scheduler.py:248-269, reaped at
nixl/base_worker.py:2982-3008, extended by heartbeats the decoder tracks
from nixl/base_scheduler.py:199-238, sends at nixl/base_worker.py:3141-3170
and the prefiller applies at nixl/base_worker.py:3010-3030; they matter
only when a decoder dies);
bidirectional transfer for multi-turn (bidirectional_kv_xfer, the
decoder's blocks read back by the prefiller); the host buffer on
accelerators NIXL cannot read directly; tensor-parallel fan-out of the
read; cross-session prefix sharing on either side (per session here, as in
vllm.sq); the router's approximate prefix cache as a data structure of
its own (its estimate is taken to be the pod's cache).
What the program predicts¶
Two decode pods of 160 000 tokens each, two prefill pods, sessions of
one to several turns (a prompt of 1 000–3 000 new tokens on a growing
context, 200 output tokens, a 3 s tool call between turns, p = 0.9), a
max_model_len of 16 384 tokens, the A100-shaped step cost of
examples/multi-turn/vllm.sq, a 200 000 token/s NIC on every pod. Every number
below is the median over seeds 1–20 of a run of 10 000 s after 1 000 s
of warm-up (--seed N --horizon 10000 --warmup 1000 --set …), with the
range across the seeds where it says more than the median. About 1 % of
turns, and so 11 % of sessions, reach max_model_len and end there.
The decider. The guide's always-disagg-pd-decider sends every prompt
to a prefiller; the prefix-based-pd-decider keeps a follow-up turn whose
context the decoder already has on the decoder. At 0.6 sessions per
second:
thr (nonCachedTokens) |
remote prefills | mean TTFT, median (range) | mean response |
|---|---|---|---|
| 1 (always) | 100 % | 29.5 ms (28.3–30.5) | 72.1 ms |
| 512 | 59 % | 28.7 ms (28.0–29.1) | 71.3 ms |
| 2 048 | 22 % | 29.0 ms (28.5–29.6) | 71.8 ms |
| never (decode only) | 0 % | 29.4 ms (28.2–30.4) | 73.8 ms |
Disaggregating everything costs a transfer per request, disaggregating
nothing costs every decoder a prefill in its decode steps, and at this
load the two cost about the same: the four settings lie within a
millisecond. The decider's thr = 512 sits below both: seed for seed it
is 0.3–1.4 ms under always-disaggregate and 0.2–1.6 ms under decode only,
in all twenty. With the guide's decode profile (the least busy pod,
no prefix affinity) a follow-up turn often lands on the pod that does not
have its context, and the prefiller — which does have it, through the
affinity filter — prefills the new tokens only: 59 % of requests are
remote at thr = 512 although every first turn is a miss.
The lease. A finished prefill's blocks stay allocated on the prefiller
until the decoder has read them, and the decoder reads only once its own
scheduler has room for the whole prompt. Shrinking the decoders' memory
lengthens the lease and grows what the prefillers hold (thr = 512; the
prefill pool is 160 000 tokens; a decoder is never smaller than
max_model_len):
| decoder memory, tokens per pod | Λ, sessions/s | remote prefills | lease, mean | prefiller memory allocated, mean per pod | prefiller queue, mean | mean TTFT |
|---|---|---|---|---|---|---|
| 160 000 | 0.6 | 59 % | 13 ms | 435 | 0.01 | 28.7 ms |
| 32 000 | 0.6 | 86 % | 33 ms | 944 | 0.01 | 49.1 ms |
| 17 600 | 0.6 | 92 % | 37 ms | 1 079 | 0.01 | 54.7 ms |
| 160 000 | 1.2 | 72 % | 25 ms | 1 650 | 0.03 | 47.8 ms |
| 32 000 | 1.2 | 93 % | 42 ms | 2 728 | 0.04 | 70.7 ms |
| 17 600 | 1.2 | 96 % | 49 ms | 3 046 | 0.04 | 78.4 ms |
Smaller decoders send more prompts to the prefillers — they cache less, so the decider's uncached suffix is longer — and each of those prompts waits longer for the decoder's room, so the prefillers hold more memory: 2.2–2.5 times at 0.6 sessions per second, 1.7–1.8 times at 1.2. The reads now share the prefillers' NICs too, and the prefill profile sends most of them to one prefiller, the one that has the prefix: at 1.2 sessions per second with the smallest decoders that costs 5 ms of TTFT and 6 ms of lease over a model that shares the decoders' NICs only. At these loads that memory is under 2 % of the pool and no prefiller queues; the deployment pays for the small decoders in TTFT, not yet in admissions.
The decoder's shortage shows up as memory on the prefiller, which is the
coupling a store-and-forward program gets backwards: its prefiller holds
through the transfer and lets go before the session queues for the decoder,
so a decoder with no room costs the prefiller nothing.
tests/pd_semantics.rs has the deterministic version: six requests, a
decoder with room for one, a prefiller with room for two; the
store-and-forward program prefills all six at once, the NIXL program stops
after the decoder is full and the two leases have taken the prefiller.
Why you should believe it¶
The claim is weaker than the vLLM use case's, and the difference is the point of saying so.
- The source, by line. The tables above name the function for every
statement. The vLLM ranges are hashed against
ref/vllmat0c87a197bymake check(scripts/check_citations.py), so a range that moves fails the build; the llm-d ranges are pinned by commit in the text only. - The statements, deterministically.
tests/pd_semantics.rscheckslease,release,loadandtransfer … from … to …on pools of ten units: a leased pool stays allocated after its scope and is free the moment the link run ends, an untaken lease ends at its bound and keeps its cache, a re-executed hold releases nothing twice, the prefiller stops when the decoder is full. - No machine. No run of a real deployment is compared with the program here: the numbers under "What the program predicts" are the program's, not measurements.
- No scheduler oracle yet.
tools/vllm_oracle.pydrives one real scheduler with a fake model runner. The oracle this program wants drives two — a producer and a consumer with a fake NIXL connector that reports a read complete after a chosen number of steps — and compares, per request, the step the decoder parks it, the step the prefiller frees its blocks and the decoder's first-token step. The one-block cache disagreement of the pull run is the first question for it.
See also: The KV transfer, How serQ is checked.