Tutorial¶
Six chapters. Each one is a complete, runnable program that adds exactly one idea to the last, and each ends with something you can measure.
By the end you will have built, from nothing, a colocated LLM engine with continuous batching, chunked prefill, a block-level prefix cache and LIFO preemption — and watched it fall off a cliff.
| Chapter | The idea | New syntax | |
|---|---|---|---|
| 1 | A queue | requests wait for a server | stage, workload, session, run, observe |
| 2 | Memory is a resource | requests also wait for memory | pool, hold |
| 3 | Sessions and turns | a session is many turns with thinking in between | loop, turn, branch, delay |
| 4 | The prefix cache | a finished turn leaves its context behind | cache, cached, evict, drop |
| 5 | The engine | prefill and decode share one iteration | step, budget, growing, preempt, admit via |
| 6 | The cliff | the cache and the queue feed each other | — |
The programs are in
docs/tutorial/programs/.
Every number quoted in these pages came from running them.