Motivation.
Motivation
Issue #2 defines the primary optimization target as faster Ouro inference on one GPU with loop-level batching and last-exited KV semantics. The current runtime already supports chunked prefill, refill scheduling, asynchronous execution, multiple attention backends, CUDA Graphs, prefix caching, and preemption. The scheduler walkthrough documents the existing separation between admission, stage selection, and batch construction, while the module-boundary RFC calls for explicit scheduling tasks and separate ownership of logical request state, KV resources, and device execution.
The current chunked prefill path leaves parallelism inside a request's recurrent-depth dimension unused:
- The scheduler selects a token range for each request.
- The runner embeds those tokens.
- The runner executes all recurrent depths for the selected range before returning.
- The engine advances the prompt frontier and requeues the request.
With LAST_EXITED, the KV written for c0@d0 is sufficient for c1@d0; c1@d0 does not depend on c0@d1. These dependencies allow a scheduler to overlap work across chunks and depths when the required state is ready.
This is useful when active prompt work is sparse or uneven. A conventional chunk batch may contain only a few rows even though several depth/chunk tasks are ready. WCPB can use those tasks to fill the same recurrent invocation. Smaller prompt tasks also produce more scheduler boundaries, allowing refill policies to admit new prompt work or schedule decode work earlier.
Benefits and Chunk-Size Trade-offs
WCPB is motivated by a real trade-off in ordinary chunked prefill. The chunk size controls both execution occupancy and scheduling granularity:
| Chunk size |
Ordinary chunked prefill |
Scheduling consequence |
| Large |
More prompt rows are selected per batch and recurrent kernels are easier to fill |
Fewer scheduler boundaries; a long prefill chunk can delay newly admitted prompt or decode work |
| Small |
More frequent boundaries and smaller units of prompt work |
Each ordinary recurrent invocation has fewer rows, so launch, metadata, and scheduling overhead can dominate and leave the GPU underfilled |
The small-chunk benefit for decode is a scheduling property rather than a claim that every small-chunk workload has lower decode latency. A smaller chunk gives the policy more opportunities to select a decode batch between prompt batches, but the actual ITL/TPOT improvement depends on stage policy, active decode work, and device execution overlap. The static experiments in this RFC did not include a concurrent decode stream, so they measure the occupancy/throughput side of this trade-off rather than directly qualifying decode latency.
What the baseline measurements show
The controlled A100 experiments make the ordinary-chunk trade-off visible. For four 8064-token requests with max_num_batched_tokens=1024, baseline end-to-end throughput fell as the chunk size became smaller:
| Chunk size |
Baseline end-to-end tok/s |
Baseline prefill rows/batch |
| 128 |
443.55 |
512.00 |
| 64 |
417.87 |
256.00 |
| 32 |
367.98 |
128.00 |
| 16 |
275.82 |
64.00 |
The same pattern appears in the mixed workloads. For the 32-request workload, baseline throughput changed from 662.25 tok/s at chunk 128 to 337.51 tok/s at chunk 16, while baseline prefill rows per batch changed from 480.00 to 60.00. For the 128-request workload, baseline throughput changed from 758.72 tok/s to 471.95 tok/s, and rows per batch changed from 451.39 to 97.57. These are paired measurements within each workload; they are not a claim that one chunk size is universally optimal.
Large chunks therefore remain useful when prompt throughput is the primary objective and enough prompt rows are available to fill the recurrent kernels. Their cost is reduced scheduling flexibility: a scheduler sees fewer points at which it can admit work or give decode a turn. Small chunks preserve those boundaries, but ordinary chunking cannot use the recurrent-depth dimension to compensate for the smaller per-chunk batch.
What WCPB adds
WCPB keeps the small chunk as the scheduling unit while allowing ready tasks from different chunks and recurrent depths to share the recurrent-core batch. This directly addresses the underfill mechanism in ordinary small-chunk prefill:
- It does not require increasing
prefill_chunk_size to obtain more tokens in a recurrent invocation.
- It can combine
c0@d1 with c1@d0 and later combinations when their hidden-state and depth-specific KV dependencies are ready.
- It retains more scheduler boundaries than a large ordinary chunk, while using independent ready tasks to improve the effective batch size.
- It is most useful when only a few requests are active or prompt lengths are uneven; in those cases, request-level batching alone may not provide enough rows.
The measurements support this mechanism. In the four-request 8064-token workload, WCPB raised rows per batch from 64.00 to 254.49 at chunk 16 and improved measured throughput from 275.82 to 412.01 tok/s (+49.38%). In the 32-request mixed workload, it raised rows per batch from 60.00 to 238.60 and improved throughput from 337.51 to 628.45 tok/s (+86.20%). In the 128-request mixed workload, it raised rows per batch from 97.57 to 361.96 and improved throughput from 471.95 to 740.29 tok/s (+56.86%). These chunk-16 results used one warmup and one measured run per variant, so they are directional measurements rather than three-run medians.
WCPB does not guarantee that a small chunk reaches the same absolute throughput as a large chunk. In the four-request workload, WCPB measured 465.31 tok/s at chunk 128 and 412.01 tok/s at chunk 16. Task metadata, hidden-state retention, dependency checks, and additional scheduling work remain costs. The intended benefit is to recover part of the occupancy lost by choosing a small chunk, while retaining its scheduling granularity. The appropriate acceptance target is therefore an end-to-end trade-off: higher effective rows and useful throughput at a chosen chunk size, together with bounded TTFT, decode ITL/TPOT, memory, and scheduler overhead.
For the SHARED KV layout, c1@d0 depends on the final-depth KV of the preceding prompt positions. The LAST_EXITED WCPB schedule therefore cannot be directly applied without changing the dependency graph. The first implementation supports only LAST_EXITED and rejects SHARED.
Proposed Change.
Proposed Design
Logical stages and scheduler-visible granularity
Now runner calls _prefill() and _prefill_tokens() for batch in PREFILL, and this batch go through preclude() and full recurrent execution. Thus the prefill operation is implemented as a heavy atomic operation without any inner scheduler room and its execution can be decomposed into the existing PRELUDE and RECURRENT operations so that it exposes more schedule flexibility. For each newly selected prompt chunk, the first operation is:
PRELUDE(chunk)
-> embed prompt tokens and prepare positions/metadata
RECURRENT(chunk, depth=0)
-> execute the first recurrent depth and write depth-0 KV
RECURRENT(chunk, depth=1 ... D-1)
-> continue the same prompt chunk using its retained hidden state
D is model.config.total_ut_steps; The prelude is executed once per chunk. The depth-0 recurrent operation consumes the prelude hidden state, and each later recurrent operation consumes the previous depth's hidden state. In a serialized baseline, the runner may execute the whole sequence as one composite chunk plan. In WCPB, the depth-0 operation and later recurrent operations become separate task transitions.
The decomposition should be available as a common runner primitive, because asynchronous execution, CUDA-graph capture, cancellation, and future layout-specific schedulers all need an explicit boundary between embedding and recurrent work. It should not automatically make every execution mode finer-grained.
With wavefront_prefill=False, the scheduler can lower one selected prompt chunk to an atomic composite plan:
PREFILL(chunk) = PRELUDE(chunk) + RECURRENT(chunk, depth=0 ... D-1)
The runner may still submit this plan asynchronously, but the scheduler does not expose its internal depth boundaries and therefore preserves baseline batching and accounting overhead. With wavefront_prefill=True, the scheduler exposes the depth-0 operation and each later recurrent depth as task transitions, allowing tasks from different chunks and depths to share later recurrent batches. The minimum implementation can keep PRELUDE(chunk) + RECURRENT(chunk, depth=0) together in one submitted item, while depth>0 tasks consume the retained hidden state. A separate scheduler-visible prelude task is a follow-up optimization for cases where embedding, CUDA streams, or prelude batching have enough independent work to amortize the extra task and event overhead.
This separation gives WCPB an interruptible scheduling model without forcing a global change to ordinary prefill. And good for fused prefilling and decoding in the same device batch though we are not going to implement it yet. It also leaves room for a future SHARED implementation to use the same phase vocabulary with a different position-major dependency graph. The phase vocabulary is therefore general, while the cross-chunk/depth wavefront policy is layout-specific and initially applies only to LAST_EXITED.
Reusing the existing PRELUDE and RECURRENT execution semantics
The longer-term implementation plan can replace the dedicated prompt-prefill runner path with a composition of the existing embedding (PRELUDE) and recurrent-core (RECURRENT) operations. In that design, PREFILL remains a scheduler/lifecycle label for a request whose prompt is not complete, while each chunk is lowered to:
prompt chunk
-> prelude / embedding
-> recurrent depth 0
-> recurrent depth 1
-> ...
-> recurrent depth D-1
This removes duplicated model execution logic and makes ordinary prefill, WCPB, and future layout-specific prefill share the same runner primitives. The baseline path can submit the complete lowered plan atomically; WCPB can expose the recurrent operations as ready tasks and retain the prelude output where a later depth needs it.
The current Stage.PRELUDE cannot be reused literally without generalization. It currently means embedding one sampled decode token after CODA, with one request position, a GPU-produced input token, decode-loop reset, and a subsequent RECURRENT enqueue. Prompt prefill instead embeds a contiguous token range, uses a vector of prompt positions, does not consume a sampled token, always executes full depth, and advances num_prefilled_tokens only after final-depth KV completion. The implementation should therefore share the underlying prelude operation and recurrent operation, while carrying an execution kind or phase that distinguishes decode_prelude from prompt_prelude.
This unification is a follow-up implementation phase after the current correctness path is stable. It can be enabled for ordinary prefill as well as WCPB, but only WCPB should expose cross-chunk/depth task boundaries. The initial WCPB change may keep the existing prompt-prefill wrapper if that reduces risk; the RFC's target architecture is the shared primitive path, not direct reuse of the current decode-stage bookkeeping.
The PREFILL -> RECURRENT transition is a per-chunk execution transition, not necessarily a request-level transition. WCPB may have chunk0@depth1 and chunk1@depth0 in flight for the same request. A single request-level Stage cannot represent both phases at once, so the scheduler must keep the request's coarse lifecycle state separately from each task's phase, depth, and completion state. Combining a prompt task with an ordinary decode item in the same device batch is a further heterogeneous-batch extension; it requires per-item execution kinds and result handling, and is not implied merely by unused max_num_seqs or max_num_batched_tokens capacity.
There is two possible ways to break down the PREFILL stage into PRELUDE and RECURRENT phases:
- Maintain
PREFILL as a scheduler-visible stage and runner execution precule and one time recurrent for this batch, then the batch.status turn to RECURRENT to finish the remaining recurrent depths.
- Discard
PREFILL and use PRELUDE and RECURRENT as the scheduler-visible stages, exposing more granularity to the scheduler.
Task representation
Maintain PREFILL Plan
The scheduler represents a prefill task with:
PrefillTask {
task_id
request_id
token_start
token_count
phase: prelude | recurrent | composite_depth0
depth: optional integer
hidden_state_handle
completion_event
state: pending | submitted | completed | cancelled
}
token_start and token_count identify the prompt range. A prelude task has no recurrent depth; a recurrent task's depth selects the depth-specific KV plane. The optional composite_depth0 phase represents the minimum implementation's fused prelude plus depth-zero recurrent operation and has depth=0. The task does not own the request's logical prompt frontier; it reports completion to the scheduler, which advances the contiguous completed frontier only after final-depth chunks have completed.
The scheduler must keep the distinction between:
num_prefilled_tokens: the largest contiguous prompt prefix completed at all depths;
- completed preludes, which only make depth-zero work eligible and do not advance the prompt frontier;
- completed final-depth chunks that may complete out of order;
- pending depth tasks for a chunk;
- hidden-state storage needed by successor depth tasks.
The final prompt frontier must never advance merely because a depth-zero task completed.
Discard PREFILL Plan
The scheduler represents generalized tasks with ExecutionTask for prefilling and decoding:
ExecutionTask {
task_id
request_id
kind: prompt | decode
phase: prelude | recurrent | coda
token_start
token_count
position
depth
exit_policy
hidden_state_handle
completion_event
state
}
For prefilling tasks: ExecutionTask(kind=prompt, phase=prelude/recurrent/,exit_policy=full_depth)
For decoding tasks: ExecutionTask(kind=decode, phase=prelude/recurrent/coda,exit_policy=decode_depth)
Dependency graph
For a chunk c_i and depth d, WCPB uses these dependencies:
task(c_i, d > 0)
depends on hidden_state(task(c_i, d-1))
and KV depth d written through token_start(c_i)
task(c_i, 0)
depends on final depth-0 KV for prompt positions before token_start(c_i)
The second dependency is what preserves causal attention between chunks. A later chunk at depth 0 may run while an earlier chunk is at depth 1 or 2 because those tasks use different depth-specific KV planes under LAST_EXITED.
There are two ways to enforce these dependencies:
- Explicit readiness checking. Keep all pending tasks in a task table and inspect the predecessor hidden-state handle, depth-specific KV prefix, request generation, cancellation state, and storage capacity before every admission. This is the more general option and is useful when tasks can be restored, preempted, imported, or completed out of order by an external worker.
- Implicit readiness through successor-only enqueueing. Do not enqueue a successor until the completion path of its predecessor has observed the required event. The ready queue then contains only tasks that are ready by construction. This has less metadata and scheduler overhead, but it requires every dependency transition to be represented by a single correct completion callback.
The simplest initial scheduler uses the second option and maintains one FIFO of ready prompt tasks. Its transition logic is:
on_request_admitted(request):
enqueue(request, chunk=0, depth=0, phase=composite_depth0)
on_depth0_complete(request, chunk):
if D > 1:
enqueue(request, chunk=chunk, depth=1, phase=recurrent)
else:
mark_final_depth_complete(request, chunk)
advance_contiguous_prompt_frontier(request)
if chunk has a successor:
enqueue(request, chunk=chunk+1, depth=0, phase=composite_depth0)
elif prompt is complete:
enqueue(request, CODA)
on_recurrent_complete(request, chunk, depth):
if depth + 1 < D:
enqueue(request, chunk=chunk, depth=depth+1, phase=recurrent)
else:
mark_final_depth_complete(request, chunk)
advance_contiguous_prompt_frontier(request)
if prompt is complete:
enqueue(request, CODA)
schedule_next_batch():
take ready tasks in queue order until token/row budgets are full
submit one PREFILL-core batch
For LAST_EXITED, this enqueue order is sufficient: a chunk's depth-0 task is created only after the preceding chunk's depth-0 KV prefix has completed, and depth d+1 is created only after depth d has produced its hidden state. It therefore permits chunk0@depth1 and chunk1@depth0 to enter the same later batch without a general readiness scan. The callback must run only after the device completion event is visible to the scheduler; otherwise the implicit invariant is false.
This simple scheme is initially limited to LAST_EXITED, no preemption, and one owner of each task transition. If preemption, PD transfer, external task restoration, or SHARED support is added, the explicit readiness check should be introduced rather than relying solely on queue insertion order.
Wavefront batch construction
Scheduler._take(Stage.PREFILL) changes from selecting request IDs to selecting ready task IDs. It still enforces:
max_num_batched_tokens as the sum of task token counts;
max_num_seqs as the number of selected task rows/owners according to the configured batch contract;
- KV capacity and execution frontier checks;
- selected-request protection during capacity-driven preemption when preemption is supported later.
Under the explicit-readiness variant, the scheduler may skip a task that is not ready and continue scanning other tasks. Under the successor-only variant, a task in the ready queue is ready by construction; the scheduler only needs to discard stale or cancelled entries. If a future implementation can insert blocked tasks, queue rotation must be bounded so a permanently blocked request cannot create an infinite scheduling loop.
The batch output should describe task identity and execution metadata explicitly:
SchedulerOutput {
stage = PREFILL
items = [
{request_id, token_start, token_count, phase, depth?, task_id},
...
]
}
This shape follows issue #32's direction that scheduling outputs should progressively replace interfaces carrying mutable request objects and should distinguish token position from loop depth.
Device execution
For a wavefront batch, ModelRunner:
- Embeds prompt tokens only for depth-0 task rows.
- Gathers retained hidden rows for depth greater than zero.
- Concatenates rows into one hidden tensor with shape
(total_task_tokens, hidden_dim).
- Builds per-row request IDs, positions, depths, block tables, and attention metadata.
- Executes the recurrent core with the per-row depth vector.
- Splits the output hidden rows back into task-owned slices.
- Records a completion event for each task and exposes the hidden-state handle to its successor.
The model recurrent function must accept a depth per row without changing the mathematical operation for any row. The attention backend must continue to read the KV plane selected by that row's depth and position.
The runner must preserve hidden-state lifetime across asynchronous submissions. A task's hidden rows cannot be released or overwritten until all successor submissions that consume them have completed or the request is cancelled.
Engine and scheduler update
The engine update path consumes a completed PrefillTask rather than assuming that a whole request chunk completed at all depths. It:
- removes the completed task from the pending-task table;
- creates successor depth and next-chunk tasks according to the dependency graph;
- records final-depth chunk completion;
- advances the contiguous prompt frontier when adjacent final-depth chunks are complete;
- publishes prefix metadata only for a valid completed prefix;
- enqueues CODA only when the full prompt has completed at all depths.
This keeps the first generated token on the full-depth path. WCPB changes the order and grouping of prefill work, not the full-depth computation required before CODA.
Goals
- Increase recurrent-core batch occupancy for chunked prompt prefill.
- Preserve
LAST_EXITED causal attention and depth-specific KV semantics.
- Keep prompt prefill at full model depth so the first generated token has the same full-depth contract.
- Work with synchronous and asynchronous execution paths.
- Work with the existing Triton/FlashAttention execution and CUDA Graph infrastructure where a stable batch shape is available.
- Keep scheduler, logical KV management, and device execution responsibilities explicit.
- Make correctness and performance claims reproducible under controlled workloads.
- Provide an incremental path from the current minimal implementation to a more general ready-task scheduler.
- Introducing PD prefill/decode separation.
Non-goals
- Fused prefilling and decoding in the same device batch.
- Supporting WCPB with
SHARED KV in the initial implementation.
- Adding preemption or recomputation of partially completed WCPB tasks in the initial implementation.
- Changing Ouro exit policies, gate thresholds, sampling semantics, or decode loop counts.
- Reducing the model's logical recurrent work or claiming fewer FLOPs.
- Replacing the scheduler with a generic task runtime or reproducing the full vLLM scheduler hierarchy.
Feedback Period.
No response
CC List.
No response
Anything else.
No response
Before submitting a new issue...
Motivation.
Motivation
Issue #2 defines the primary optimization target as faster Ouro inference on one GPU with loop-level batching and last-exited KV semantics. The current runtime already supports chunked prefill, refill scheduling, asynchronous execution, multiple attention backends, CUDA Graphs, prefix caching, and preemption. The scheduler walkthrough documents the existing separation between admission, stage selection, and batch construction, while the module-boundary RFC calls for explicit scheduling tasks and separate ownership of logical request state, KV resources, and device execution.
The current chunked prefill path leaves parallelism inside a request's recurrent-depth dimension unused:
With
LAST_EXITED, the KV written forc0@d0is sufficient forc1@d0;c1@d0does not depend onc0@d1. These dependencies allow a scheduler to overlap work across chunks and depths when the required state is ready.This is useful when active prompt work is sparse or uneven. A conventional chunk batch may contain only a few rows even though several depth/chunk tasks are ready. WCPB can use those tasks to fill the same recurrent invocation. Smaller prompt tasks also produce more scheduler boundaries, allowing refill policies to admit new prompt work or schedule decode work earlier.
Benefits and Chunk-Size Trade-offs
WCPB is motivated by a real trade-off in ordinary chunked prefill. The chunk size controls both execution occupancy and scheduling granularity:
The small-chunk benefit for decode is a scheduling property rather than a claim that every small-chunk workload has lower decode latency. A smaller chunk gives the policy more opportunities to select a decode batch between prompt batches, but the actual ITL/TPOT improvement depends on stage policy, active decode work, and device execution overlap. The static experiments in this RFC did not include a concurrent decode stream, so they measure the occupancy/throughput side of this trade-off rather than directly qualifying decode latency.
What the baseline measurements show
The controlled A100 experiments make the ordinary-chunk trade-off visible. For four 8064-token requests with
max_num_batched_tokens=1024, baseline end-to-end throughput fell as the chunk size became smaller:The same pattern appears in the mixed workloads. For the 32-request workload, baseline throughput changed from
662.25 tok/sat chunk 128 to337.51 tok/sat chunk 16, while baseline prefill rows per batch changed from480.00to60.00. For the 128-request workload, baseline throughput changed from758.72 tok/sto471.95 tok/s, and rows per batch changed from451.39to97.57. These are paired measurements within each workload; they are not a claim that one chunk size is universally optimal.Large chunks therefore remain useful when prompt throughput is the primary objective and enough prompt rows are available to fill the recurrent kernels. Their cost is reduced scheduling flexibility: a scheduler sees fewer points at which it can admit work or give decode a turn. Small chunks preserve those boundaries, but ordinary chunking cannot use the recurrent-depth dimension to compensate for the smaller per-chunk batch.
What WCPB adds
WCPB keeps the small chunk as the scheduling unit while allowing ready tasks from different chunks and recurrent depths to share the recurrent-core batch. This directly addresses the underfill mechanism in ordinary small-chunk prefill:
prefill_chunk_sizeto obtain more tokens in a recurrent invocation.c0@d1withc1@d0and later combinations when their hidden-state and depth-specific KV dependencies are ready.The measurements support this mechanism. In the four-request 8064-token workload, WCPB raised rows per batch from
64.00to254.49at chunk 16 and improved measured throughput from275.82to412.01 tok/s(+49.38%). In the 32-request mixed workload, it raised rows per batch from60.00to238.60and improved throughput from337.51to628.45 tok/s(+86.20%). In the 128-request mixed workload, it raised rows per batch from97.57to361.96and improved throughput from471.95to740.29 tok/s(+56.86%). These chunk-16 results used one warmup and one measured run per variant, so they are directional measurements rather than three-run medians.WCPB does not guarantee that a small chunk reaches the same absolute throughput as a large chunk. In the four-request workload, WCPB measured
465.31 tok/sat chunk 128 and412.01 tok/sat chunk 16. Task metadata, hidden-state retention, dependency checks, and additional scheduling work remain costs. The intended benefit is to recover part of the occupancy lost by choosing a small chunk, while retaining its scheduling granularity. The appropriate acceptance target is therefore an end-to-end trade-off: higher effective rows and useful throughput at a chosen chunk size, together with bounded TTFT, decode ITL/TPOT, memory, and scheduler overhead.For the
SHAREDKV layout,c1@d0depends on the final-depth KV of the preceding prompt positions. TheLAST_EXITEDWCPB schedule therefore cannot be directly applied without changing the dependency graph. The first implementation supports onlyLAST_EXITEDand rejectsSHARED.Proposed Change.
Proposed Design
Logical stages and scheduler-visible granularity
Now runner calls
_prefill()and_prefill_tokens()for batch inPREFILL, and this batch go throughpreclude()and full recurrent execution. Thus the prefill operation is implemented as a heavy atomic operation without any inner scheduler room and its execution can be decomposed into the existingPRELUDEandRECURRENToperations so that it exposes more schedule flexibility. For each newly selected prompt chunk, the first operation is:Dismodel.config.total_ut_steps; The prelude is executed once per chunk. The depth-0 recurrent operation consumes the prelude hidden state, and each later recurrent operation consumes the previous depth's hidden state. In a serialized baseline, the runner may execute the whole sequence as one composite chunk plan. In WCPB, the depth-0 operation and later recurrent operations become separate task transitions.The decomposition should be available as a common runner primitive, because asynchronous execution, CUDA-graph capture, cancellation, and future layout-specific schedulers all need an explicit boundary between embedding and recurrent work. It should not automatically make every execution mode finer-grained.
With
wavefront_prefill=False, the scheduler can lower one selected prompt chunk to an atomic composite plan:The runner may still submit this plan asynchronously, but the scheduler does not expose its internal depth boundaries and therefore preserves baseline batching and accounting overhead. With
wavefront_prefill=True, the scheduler exposes the depth-0 operation and each later recurrent depth as task transitions, allowing tasks from different chunks and depths to share later recurrent batches. The minimum implementation can keepPRELUDE(chunk) + RECURRENT(chunk, depth=0)together in one submitted item, whiledepth>0tasks consume the retained hidden state. A separate scheduler-visible prelude task is a follow-up optimization for cases where embedding, CUDA streams, or prelude batching have enough independent work to amortize the extra task and event overhead.This separation gives WCPB an interruptible scheduling model without forcing a global change to ordinary prefill. And good for fused prefilling and decoding in the same device batch though we are not going to implement it yet. It also leaves room for a future
SHAREDimplementation to use the same phase vocabulary with a different position-major dependency graph. The phase vocabulary is therefore general, while the cross-chunk/depth wavefront policy is layout-specific and initially applies only toLAST_EXITED.Reusing the existing PRELUDE and RECURRENT execution semantics
The longer-term implementation plan can replace the dedicated prompt-prefill runner path with a composition of the existing embedding (
PRELUDE) and recurrent-core (RECURRENT) operations. In that design,PREFILLremains a scheduler/lifecycle label for a request whose prompt is not complete, while each chunk is lowered to:This removes duplicated model execution logic and makes ordinary prefill, WCPB, and future layout-specific prefill share the same runner primitives. The baseline path can submit the complete lowered plan atomically; WCPB can expose the recurrent operations as ready tasks and retain the prelude output where a later depth needs it.
The current
Stage.PRELUDEcannot be reused literally without generalization. It currently means embedding one sampled decode token after CODA, with one request position, a GPU-produced input token, decode-loop reset, and a subsequentRECURRENTenqueue. Prompt prefill instead embeds a contiguous token range, uses a vector of prompt positions, does not consume a sampled token, always executes full depth, and advancesnum_prefilled_tokensonly after final-depth KV completion. The implementation should therefore share the underlying prelude operation and recurrent operation, while carrying an execution kind or phase that distinguishesdecode_preludefromprompt_prelude.This unification is a follow-up implementation phase after the current correctness path is stable. It can be enabled for ordinary prefill as well as WCPB, but only WCPB should expose cross-chunk/depth task boundaries. The initial WCPB change may keep the existing prompt-prefill wrapper if that reduces risk; the RFC's target architecture is the shared primitive path, not direct reuse of the current decode-stage bookkeeping.
The
PREFILL -> RECURRENTtransition is a per-chunk execution transition, not necessarily a request-level transition. WCPB may havechunk0@depth1andchunk1@depth0in flight for the same request. A single request-levelStagecannot represent both phases at once, so the scheduler must keep the request's coarse lifecycle state separately from each task'sphase,depth, and completion state. Combining a prompt task with an ordinary decode item in the same device batch is a further heterogeneous-batch extension; it requires per-item execution kinds and result handling, and is not implied merely by unusedmax_num_seqsormax_num_batched_tokenscapacity.There is two possible ways to break down the
PREFILLstage intoPRELUDEandRECURRENTphases:PREFILLas a scheduler-visible stage and runner executionpreculeand one timerecurrentfor this batch, then the batch.status turn toRECURRENTto finish the remaining recurrent depths.PREFILLand usePRELUDEandRECURRENTas the scheduler-visible stages, exposing more granularity to the scheduler.Task representation
Maintain
PREFILLPlanThe scheduler represents a prefill task with:
token_startandtoken_countidentify the prompt range. A prelude task has no recurrent depth; a recurrent task'sdepthselects the depth-specific KV plane. The optionalcomposite_depth0phase represents the minimum implementation's fused prelude plus depth-zero recurrent operation and hasdepth=0. The task does not own the request's logical prompt frontier; it reports completion to the scheduler, which advances the contiguous completed frontier only after final-depth chunks have completed.The scheduler must keep the distinction between:
num_prefilled_tokens: the largest contiguous prompt prefix completed at all depths;The final prompt frontier must never advance merely because a depth-zero task completed.
Discard
PREFILLPlanThe scheduler represents generalized tasks with ExecutionTask for prefilling and decoding:
For prefilling tasks:
ExecutionTask(kind=prompt, phase=prelude/recurrent/,exit_policy=full_depth)For decoding tasks:
ExecutionTask(kind=decode, phase=prelude/recurrent/coda,exit_policy=decode_depth)Dependency graph
For a chunk
c_iand depthd, WCPB uses these dependencies:The second dependency is what preserves causal attention between chunks. A later chunk at depth 0 may run while an earlier chunk is at depth 1 or 2 because those tasks use different depth-specific KV planes under
LAST_EXITED.There are two ways to enforce these dependencies:
The simplest initial scheduler uses the second option and maintains one FIFO of ready prompt tasks. Its transition logic is:
For
LAST_EXITED, this enqueue order is sufficient: a chunk's depth-0 task is created only after the preceding chunk's depth-0 KV prefix has completed, and depthd+1is created only after depthdhas produced its hidden state. It therefore permitschunk0@depth1andchunk1@depth0to enter the same later batch without a general readiness scan. The callback must run only after the device completion event is visible to the scheduler; otherwise the implicit invariant is false.This simple scheme is initially limited to
LAST_EXITED, no preemption, and one owner of each task transition. If preemption, PD transfer, external task restoration, orSHAREDsupport is added, the explicit readiness check should be introduced rather than relying solely on queue insertion order.Wavefront batch construction
Scheduler._take(Stage.PREFILL)changes from selecting request IDs to selecting ready task IDs. It still enforces:max_num_batched_tokensas the sum of task token counts;max_num_seqsas the number of selected task rows/owners according to the configured batch contract;Under the explicit-readiness variant, the scheduler may skip a task that is not ready and continue scanning other tasks. Under the successor-only variant, a task in the ready queue is ready by construction; the scheduler only needs to discard stale or cancelled entries. If a future implementation can insert blocked tasks, queue rotation must be bounded so a permanently blocked request cannot create an infinite scheduling loop.
The batch output should describe task identity and execution metadata explicitly:
This shape follows issue #32's direction that scheduling outputs should progressively replace interfaces carrying mutable request objects and should distinguish token position from loop depth.
Device execution
For a wavefront batch, ModelRunner:
(total_task_tokens, hidden_dim).The model recurrent function must accept a depth per row without changing the mathematical operation for any row. The attention backend must continue to read the KV plane selected by that row's depth and position.
The runner must preserve hidden-state lifetime across asynchronous submissions. A task's hidden rows cannot be released or overwritten until all successor submissions that consume them have completed or the request is cancelled.
Engine and scheduler update
The engine update path consumes a completed
PrefillTaskrather than assuming that a whole request chunk completed at all depths. It:This keeps the first generated token on the full-depth path. WCPB changes the order and grouping of prefill work, not the full-depth computation required before CODA.
Goals
LAST_EXITEDcausal attention and depth-specific KV semantics.Non-goals
SHAREDKV in the initial implementation.Feedback Period.
No response
CC List.
No response
Anything else.
No response
Before submitting a new issue...