Serving vocabulary¶
Names for the parts of a request's life. An admission is the kernel's
hold … at admission (…) { … } cache (…); these name what
the request does once it is in.
| Form | The request… | Kernel |
|---|---|---|
prefill W; |
computes its prompt's KV | run prefill (W); or run E prefill (T); |
decode W; |
generates its output, a token per iteration | run decode (W); or run E decode (T); |
tool Z; |
waits outside the engine (a tool call, a person reading) | run tool (Z); |
transfer (X) from P to Q (n); |
has its KV moved to another instance | run link (X); load Q (n); release P; |
Each form is sugar: the parser rewrites it to the kernel statement in the last column, so the AST, the IR and the interpreter know nothing of it.
prefill, decode, tool¶
prefill [ '[' j ']' | on STAGE ] work [growing POOL];
decode [ '[' j ']' | on STAGE ] work [growing POOL];
tool [ '[' j ']' | on STAGE ] work;
| Argument | Type | Description |
|---|---|---|
[j] |
expr |
Index into the role's stage array: prefill[j] W;. |
on STAGE |
stage |
Names the stage explicitly: prefill on P2 (W);. |
work |
expr |
W: clock time on a fifo, ps or delay stage. T: tokens on a step engine. |
growing |
pool |
prefill and decode on a step engine only. Passes through to the run; a form never adds it. |
Which stage. Without [j] or on, the form finds its stage among those
declared above it: the stage named for the role (prefill, link or
transfer, decode, tool); failing that, for prefill and decode, the
step engine. Exactly one must qualify: with none or several the parser stops
at the form. On a step engine the run gets the role's mode; elsewhere it is
plain. transfer and tool on a step engine are link errors, so they take no
growing, which needs one. transfer finds its stage by the same rule and
always says where the KV goes (below); a link that only
takes time is run link (X);.
transfer … from … to¶
| Argument | Type | Description |
|---|---|---|
X |
expr |
The link's work, in its stages' unit: time at a ps(1), tokens at a ps of tokens per second. |
P |
pool |
The session's lease (or hold) the KV comes from. Given back at the end. |
Q |
pool |
The session's hold the KV arrives in. |
n |
expr |
Tokens counted as computed at Q. |
run link (X); load Q (n); release P;. The KV of a prefill/decode split lives
in two pools whose lifetimes overlap without nesting: the decode instance
allocates before the prefill instance frees. on a, b names several stages
the read holds at once, the sender's link and the receiver's
(run): transfer on egress[i], ingress[j] (X) from P to Q (n);
is run egress[i], ingress[j] (X); load Q (n); release P;. A link queue with
a latency (serve ps(BwD) latency x0;) is waited first: each named link
that has one adds run L.latency[k] (x0); before the run, in the order
named.
Example¶
From examples/pd-disaggregation/llmd_nixl_pull.sq, the decoder's read of
the prefiller's leased blocks, over the prefiller's NIC and its own, in the
decoder's decode (prompt) from src entry (src is the prefiller's leased
pool). The NICs are the pods' (nic ps(BwD);), and the relation D pull P
latency x0 share maxmin; makes a transfer without on the read over
P.nic[i] and D.nic[j], after the decoder's wait x0: