Skip to content

Tutorial

Six chapters. Each one is a complete, runnable program that adds exactly one idea to the last, and each ends with something you can measure.

By the end you will have built, from nothing, a colocated LLM engine with continuous batching, chunked prefill, a block-level prefix cache and LIFO preemption — and watched it fall off a cliff.

Chapter The idea New syntax
1 A queue requests wait for a server stage, workload, session, run, observe
2 Memory is a resource requests also wait for memory pool, hold
3 Sessions and turns a session is many turns with thinking in between loop, turn, branch, delay
4 The prefix cache a finished turn leaves its context behind cache, cached, evict, drop
5 The engine prefill and decode share one iteration step, budget, growing, preempt, admit via
6 The cliff the cache and the queue feed each other —

The programs are in docs/tutorial/programs/. Every number quoted in these pages came from running them.

Run as you read

cargo build --release
./target/release/serq run docs/tutorial/programs/01-queue.sq
Every chapter ends with a sweep you can reproduce with --set.