The pull relation¶
A KV transfer between two pods is one fact of the deployment: which pod's
KV goes to which, who starts the copy, what it waits for, and how
concurrent copies share the wire. examples/pd-disaggregation/llmd_nixl_pull.sq
wrote that fact in three places, none of which said it. This document is
the design of nic and D pull P latency x share s;, issue #200: parse-time
sugar, no IR change.
Before¶
examples/pd-disaggregation/llmd_nixl_pull.sq at 80e55e0, lines 78–87
and 156:
// The NICs: a prefiller's, which the reads out of it share, and a
// decoder's, which the reads into it share. A read holds both at once; how
// concurrent reads divide them is UCX's and the fabric's, not vLLM's, and
// max-min fairness is the model, not a measurement
// (docs/design/bandwidth-sharing.md).
queue egress[NP] : link { serve ps(BwP); }
// The decoder's worker posts the READ and finds it done at a later step:
// a fixed wait per read, not bandwidth.
queue ingress[ND] : link { serve ps(BwD) latency x0; }
share maxmin;
…
transfer on egress[src], ingress[self] (prompt - c) from src to kv (prompt - 1 - c);
- The NICs are families of their own, next to the pods, and the program
joins
egress[i]toP[i]only through the index, withsrcread as a number (egress[src]), a pun the queue design lists in its self-critique. share maxmin;stands alone: it says neither what is shared nor between whom. A reader finds out by reading the transfer.- That the decoder reads (NIXL's pull mode) is in a comment. The latency is
on
ingressbecause the decoder's worker posts the READ, which the program does not say either.
After¶
The program in the repository:
queue P[NP] : prefill {
…
nic ps(BwP); // the pod's NIC: the reads out of it share it
prefill (prompt) { … }
}
queue D[ND] : decode {
…
nic ps(BwD); // the pod's NIC: the reads into it share it
decode (prompt) from src {
…
transfer (prompt - c) from src to kv (prompt - 1 - c); // over P's NIC and this pod's
}
}
D pull P latency x0 share maxmin;
The two programs run the same: at seeds 1 and 7 every count, observable,
pool row and stage row is byte-identical once egress, ingress and
ingress.latency are read as P.nic, D.nic and D.nic.latency.
The design¶
A pod owns its NIC. nic kind; in a queue is the stage Q.nic, one per
member, declared after the serve and before the entries. It is a stage
like any other: transfer on P.nic[i], … still works, and the deployment
figure puts it in the pod's box.
The relation is one line for the topology and the mode. D pull P
latency x share s;, after both queues:
| Part | Says |
|---|---|
D pull P |
the KV goes from P to D, and D starts the copy |
latency x |
D waits x before each read (a number or a constant over lets); optional |
share s |
concurrent reads divide the two NICs by s: the program's share, written here |
A transfer without on is the read. Inside an entry of D called
from P[i], transfer (n) from src to kv (m); names no stage, and the
relation supplies both ends:
| Surface | Kernel |
|---|---|
nic k; in queue Q[N] |
stage Q.nic[N] : k; |
D pull P latency x share s; |
stage D.nic.latency[N] : delay; share s;, x a constant of its own |
transfer (n) from src to kv (m) in D[j], called from P[i] |
run D.nic.latency[j] (x); run P.nic[i], D.nic[j] (n); load D.kv[j] (m); release P.kv[i]; |
The latency's constant has a name of its own (D.pull.time), as a link's
does, so that a parameter or a local of the entry with the same name does
not replace it, and --set still reaches the expression.
Checks. Each is a parse error, at the line:
- either queue is not declared above the relation, or has no
nic; - a reader has two relations, or reads from itself;
- the relation names no
share, or a policy other than the program's; - a
shareon its own as well as the relation's; - an entry of
Dcalledfroma queue other thanP; - a
transferwithoutonin a queue that pulls from none.
What it earns¶
llmd_nixl_pull.sqsays NIXL's pull mode in the program, not in a comment, and the policy says what it divides.- The
srcpun is gone from the transfer:egress[src]was the only use. - No IR change: the relation is the stages and runs the program wrote before.
Self-critique¶
No push. P push D would read as the push a reader expects, vLLM's
proxy, which sends the prefill and the decode request at once
(disagg_proxy_pushconnector_demo.py:227-270), so that the decoder
allocates during the prefill and the prefiller writes as soon as it has
both the finished request and the decoder's registration
(nixl/push_scheduler.py:207-294). That needs a reservation the language does not
have (The KV transfer). The push this model can write
today is llm-d's serial one, which differs from pull only in whose worker
waits; under the word push it would mean less than it says (criterion 0).
Serial push is still written with link queues and transfer on
(the P/D use case, The two modes in the program).
One policy per program. The IR's share is the program's, so every
relation names the same. NVLink one way and RDMA another is a policy per
relation, and an IR change.
Two places for a latency. A link queue's serve … latency (#193) and
the relation's latency. They are not the same fact: the link's is a
property of the wire, waited by any transfer over it; the relation's is the
reader's, waited by the reads it starts. A program uses one or the other,
and a pod's nic takes no latency of its own.
nic is one NIC. A pod with a NIC per tensor-parallel rank, or a
separate NIC for each direction, is not written; the read's fan-out over
ranks is not modelled either.
The relation is declared after both queues. A forward reference would read better at the top of a file, and the stage rule elsewhere is "declared above", which the relation follows.