Skip to content

6. The cliff

Everything so far degraded gracefully. Turn the load up on the engine of chapter 5 and it does not.

The sweep

for L in 1.5 1.7 1.8 1.9 2.0; do
  serq run docs/tutorial/programs/05-engine.sq --set Lambda=$L --json
done
sessions/s hit rate TTFT preemptions engine utilisation
1.5 0.512 0.364 s 22 043 0.628
1.7 0.490 3.058 s 99 128 0.872
1.8 0.498 40.06 s 169 930 1.000
1.9 0.494 85.44 s 173 225 1.000
2.0 0.486 168.9 s 174 634 1.000

Between 1.7 and 1.8 sessions per second — a 6 % increase in load — the time to first token goes up by a factor of thirteen. There is no knee to plan against here; there is an edge.

Is it just saturation?

The obvious reading is that the engine ran out of compute at \(\rho = 1\). It did, but that is the symptom. Run the same loads with ten times the KV pool and nothing else changed:

serq run docs/tutorial/programs/05-engine.sq --set Lambda=2.0 --set blocks=40000
serq run docs/tutorial/programs/05-engine.sq --set Lambda=3.0 --set blocks=40000
sessions/s KV blocks hit rate TTFT preemptions utilisation
2.0 4 000 0.486 168.9 s 174 634 1.000
2.0 40 000 0.801 0.020 s 0 0.471
3.0 40 000 0.801 0.023 s 0 0.636

Same arrival rate, same engine, same cost model. Ten times the memory, and the time to first token falls by a factor of eight thousand — and at half again the load the big-pool system is still idle.

The engine did not run out of compute. It ran out of compute because it ran out of memory.

The loop

    memory is tight
          │
          ▼
 prefixes get evicted ──────┐
          │                 │
          ▼                 │
   turns miss, and a        │
   miss recomputes the      │
   whole context            │
          │                 │
          ▼                 │
  each turn costs the       │
  engine more              │
          │                 │
          ▼                 │
  turns stay resident       │
  longer, holding KV ───────┘

Nothing in this loop is a bug. Every step is the system behaving exactly as designed. The loop has two stable states — one where prefixes survive and work is cheap, and one where they do not and it is not — and the load at which it falls out of the first is not where utilisation says it should be.

The 174 634 preemptions in the bottom row are the loop running: requests being thrown out of the batch, their blocks freed, and their work recomputed when they come back.

This is not an artefact of the model

examples/replay/vllm_replay.sq is this same structure fitted to an A100 running Qwen3-8B, replaying a measured 333-session trace. The measured replica collapsed between a 3.0 s and a 2.5 s session spacing; the program predicted a mean TTFT of 39.1 s against 34.6 s measured, and a full-hit rate of 0.192 against 0.216.

Better: the program predicted the fix before it was run. Pinning a waiting request's prefix so that it cannot be evicted while it queues — one line, the program without admit via engine — was predicted to take the 2.5 s replay from 39.1 s to 0.888 s. Measured on the real A100 engine afterwards: 34.6 s → 0.878 s.

That pre-registered prediction is what a specification is for. You do not run the experiment to find out what happens; you run it to find out whether you were right.

What to try

The interesting question is not where the cliff is but what moves it:

# more memory
serq run docs/tutorial/programs/05-engine.sq --set Lambda=1.8 --set blocks=8000
# fewer concurrent requests, so each one holds memory for less time
serq run docs/tutorial/programs/05-engine.sq --set Lambda=1.8 --set max_seqs=8
# shorter thinking time, so prefixes are reused before they are evicted
serq run docs/tutorial/programs/05-engine.sq --set Lambda=1.8 --set Z=1

That is the language. Next: a real system written in it, or the reference.