Skip to content

vllm-ascend

The latest tag checked on 2026-09-30 is v0.27.1rc1, a release candidate. Its release notes specify upstream vLLM v0.27.1. This page considers full-attention request/KV paths; the selected scheduler depends on configuration.

Tagged behavior

ShortRequestFirstRequestQueue classifies requests into immediate, short and long queues. Immediate requests take precedence; sufficiently old long requests are promoted ahead of short requests. This is a class-and-aging policy, not simply sorting by length. Queue implementation

BatchJobAwareRequestQueue uses job-level decode-length predictions, available admission budget, cold-start requests and length buckets. A fixed per-request key cannot reproduce its shared history. Job-aware implementation

In the relevant preemption path, RecomputeScheduler.schedule asks the connector to offload. Success proceeds to normal preemption; failure finishes the local request through _finish_recomputed_request and returns stop_reason="recomputed". DyntraLBPolicyMixin also prefetches remote KV and waits in WAITING_FOR_REMOTE_KVS. Recovery, prefetch

Device inputs and padding

_pad_for_sequence_parallelism rounds scheduled tokens to a TP-size multiple when the relevant SP path is enabled. _determine_batch_execution_and_padding then selects a graph descriptor and checks uniform decode from per-request scheduled lengths and computed state. With DP, metadata synchronization can pad ranks to a common maximum and redispatch the graph mode. Explicit eager execution returns a non-graph descriptor; graph eligibility is separate from logical admission. These are conditional paths, not a universal NPU rule. SP and graph selection, DP coordination

The FIA TND input requires the final query-length boundary to equal the hidden-state token dimension. Padding can therefore insert a dummy request, not merely extend a flat tensor. The full-attention builder also pads sequence length and block-table metadata to that dummy row; dummy outputs are trimmed and KV writes use only actual tokens. Query boundaries, metadata consistency

A single-rank cost can round tokens without changing request progress. Cross-rank graph agreement and per-request boundary validation require more than aggregate cost terms. The current example does not model graph capture or padded device buffers. See input shapes.

Executable serQ approximation

vllm-ascend deployment in serQ

examples/vendors/ascend.sq
// FCFS lanes from vllm-ascend v0.27.1rc1: immediate, aged long, short, long.
// This full-attention example models local admission, not job prediction or remote KV.
let short_tokens = 128;
let max_wait = 0.003;
def waiting_class(urgent, prompt, elapsed) =
  urgent ? 0 : prompt > short_tokens && max_wait > 0 && elapsed >= max_wait ? 1 : prompt <= short_tokens ? 2 : 3;
pool reqs {
  cap 1;
  queue by (waiting_class(immediate, prompt, waited));
  admit via engine;
}
pool kv { cap 4096; block 16; }
stage engine : step { budget 128; chunk 128; cost 0.001; memory kv; }
stage delay : delay;
workload { arrive batch(6); }
session {
  set prompt = serial == 1 || serial == 5 ? 256 : 64;
  set immediate = serial == 4;
  set out = serial == 0 ? 4 : 1;
  run delay (serial == 2 ? 0.001 : serial == 3 ? 0.002 : serial >= 4 ? 0.003 : 0);
  hold reqs (1), kv (prompt + out) {
    observe selected = serial;
    observe admitted = now;
    prefill prompt;
    decode out;
  }
  observe response = now;
  end;
}
run { horizon 1; }

This IR v9 example reevaluates immediate/aged-long/short/long precedence before each admission selection, using waited for elapsed queue time and FIFO ties. Its six requests are selected in order 0, 4, 1, 5, 2, 3; disabling aging with max_wait = 0 produces 0, 4, 2, 3, 1, 5. Request slots, KV capacity, token budget and chunk size are explicit; prompt and output capacity is allocated upfront to avoid recovery in this example.

Remaining gaps

Requirement Current mechanism Remaining gap
Waiting-lane selection Selection-time queue by and waited model FCFS class precedence and aging Priority-lane configuration and lane-specific prepend behavior
Job prediction and cold start Session attributes Shared job history and completion-driven predictor updates
Remote KV arrival lease, load, release, explicit transfer Connector success/failure, cancellation and readiness protocol
Offload or remote recompute preempt lifo, computed-aware local recovery Choose a recovery target and finish/forward to another engine

The waiting-selection design documents the implemented IR v9 semantics and their limits. Aging changes selection at an admission attempt; it does not schedule an independent wakeup timer. Shared job history and connector behavior remain outside this example.

The current IR can retain source memory through a transfer and preserve a local request's known progress on re-admission. Those mechanisms do not by themselves implement Ascend's connector or routing policy.

Oracle scenarios

Compare a long request crossing the aging threshold under continuous short arrivals; predictor updates after cold start; and offload success, failure and remote recompute during decode. Observe selected requests/tokens, held KV, remote readiness and recovery destination.

Validation: this reduced example links and completes its six requests. It has not been compared request by request with the Ascend scheduler or runtime.