Skip to content

vllm-rbln

The latest tag checked on 2026-09-30 is v0.11.3a21, an alpha. Its dependency is vLLM 0.26.0+cpu; the plugin's version is separate. Only the native RBLNScheduler and RBLNKVCacheManager full-attention paths are considered.

Tagged behavior

RBLNScheduler.schedule constrains a batch to decodes or a lone prefill. Accepting a waiting prefill can discard an already selected decode batch. DecodeBatchBudget separates a PP-dependent hard cap from a soft cap based on current demand. Selection, decode budget

The waiting path calls allocate_slots with full_sequence_must_fit=True: admission tests whether the current full sequence fits even when the allocation is only for the next chunk. Admission capacity and actual allocation are different quantities. Scheduler

Sub-block caching finds a prefix inside a larger physical block and creates copy operations into another block. apply_sub_block_match transfers source references to the copy operations; release_copy_ops releases them. Lookup granularity, allocation granularity and copy lifetime must remain separate. Cache manager

Compiled inputs and padding

The native runner compiles decode shapes up to the PP-stage ceiling max_num_seqs // pipeline_parallel_size, then rounds active decode requests up to the smallest configured bucket. Buckets may be linear, exponential or manual; a missing covering bucket is an error, not an arbitrary new shape. Runner bucket configuration, bucket lookup

For text inputs the runner stages a lone prefill with query dimension max_num_tokens, even when its actual chunk is shorter. Decode uses padded request rows and the step's uniform query length. InputStager fills dummy rows/positions and copies only the actual rectangle. Padding changes device work and buffers; it must not become extra generated/computed request tokens. Input layout, staging

determine_batch_execution_and_padding requires a uniform query length (num_tokens % num_reqs == 0). Its specialized DP path can force decoding peers to the top bucket and prefill token dimension when another rank prefills. This applies even with ordinary one-token full-attention decode; local phase isolation does not remove cross-rank shape coordination. Shape routes

For a local non-speculative model, a cost expression can charge fixed prefill width or bucketed decoders while the logical work stays unchanged. The current example uses illustrative costs and does not emulate these compiled shapes. See input shapes for executable cost expressions and the remaining IR requirements.

Executable serQ approximation

vllm-rbln deployment in serQ

examples/vendors/rbln.sq
// Executable approximation: one prefill or a decode-only batch.
// PP caps, remote-KV admission and sub-block copying are not modeled.
// The native RBLNScheduler keeps vLLM's prefix cache: a waiting request
// looks up its hit and the step's blocks are cached once it is final, and a
// failed allocation preempts the last running request under FCFS
// (docs/use-cases/rbln.md links the lines at tag v0.11.3a21).
use "../../lib/vllm.sq";
pool kv { cap 8192; block 16; evict lru; preempt lifo; }
pool reqs { cap 4; admit via engine; }
stage engine : step {
  budget 256;
  chunk 128;
  serve exclusive prefill;
  cost 0.002 + 0.00001 * tokens;
  memory kv;
}
workload {
  arrive batch(6);
  init { set prompt = 256; set out = 16; }
}
session {
  set t0 = now;
  hold reqs (1), kv (min(known, hit + min(128, left))) reserve (known)
       at admission (known = computed < prompt ? prompt : computed + 1,
                     left = budget_left(engine),
                     hit = min(cachedin(kv), reusable(known, blocksize(kv)))) {
    set c = min(cached, reusable(known, blocksize(kv)));
    prefill (known - c) growing kv;
    branch (known == prompt) { observe ttft = now - t0; }   // not a resumed request's re-prefill
    decode (out - (known - prompt)) growing kv;
  } cache (prompt + out);
  observe response = now - t0;
  end;
}
run { horizon 2; warmup 0; seed 1; }

serve exclusive prefill selects either one prefill or a decode-only batch. A fitting waiting prefill can replace tentative resident decodes and use the full token budget. Cancelled decode work does not advance computed KV; its already acquired allocation stays held. A selected prefill stops further waiting admission. Ordinary capacity, budget-exhaustion and preemption gates still apply. PR #171 implemented this policy in the current interpreter; see the phase-isolation contract for the regression scenarios and exact limits.

reserve (known) tests the current sequence's capacity (the prompt, or what a resumed request had computed), while the hold allocates only the initial chunk and growing kv extends it as computation advances. The live token budget is bound at admission, rather than captured before queuing.

The native scheduler keeps vLLM's prefix cache and preemption. A waiting request looks up its prefix hit (get_computed_blocks), the step's blocks are cached once scheduling is final (cache_blocks), and a failed allocation preempts the last running request under FCFS (running.pop()). The program says the same for a session's own prefix, with the hit bound at admission, cache (prompt + out) and preempt lifo on kv. RBLNScheduler hands a preemption to upstream vLLM (_preempt_request), which keeps the generated tokens and drops their KV, so a resumed request prefills from what it had computed (known, as lib/vllm.sq cites). In this workload neither fires: the six requests are separate sessions, and serQ's cache is keyed by session where vLLM's is shared by content hash, so none hits another's prefix; four slots of at most 272 tokens never fill 8192 tokens (512 blocks) of KV.

The example covers local phase isolation. PP caps, remote-KV decode-ready admission guards and sub-block copy semantics remain outside it. The hand-derived interpreter regressions do not establish native scheduler equivalence.

Remaining gaps

Requirement Needed refinement
PP hard/soft decode caps Per-request selection and batch constraints, separate from resident slot count
Sub-block hit and copy Independent lookup/allocation units and source/destination objects with copy leases

Reducing block to the sub-block size would also reduce physical allocation. It can match hit counts while predicting the wrong memory pressure.

Oracle scenarios

Compare waiting prefill arriving during resident decode; PP hard/soft cap divergence; and a partial-block prefix match requiring a copy. Include source eviction and cancellation while copying. Observe final selected batch, computed/committed progress, physical block count and copy reference acquisition/release.

Validation: this reduced example links and completes its six requests. All 44 iteration assignments were checked to contain either a lone prefill or decodes only. The interpreter policy also has six hand-derived regression tests. No native RBLN differential test has been run.