You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Maintain one model and feature progress tracker for vllm-rlt, initially limited to Ouro, Nanbeige4.2, and Huginn. Contributors should be able to see which implementations exist, which work is in progress, and which combinations still need integration.
Track architecture implementations rather than duplicating the same backend for every checkpoint size. Keep PD as a core priority alongside loop-level batching, depth-aware KV, asynchronous execution, adaptive depth, and self-speculative decoding.
This issue complements the architecture/refactoring roadmap in #32. Recipes and training/rollout integration will have separate RFCs; this issue links to their results rather than collecting unrelated benchmark scripts or training tasks.
Proposed Change.
Scope and model reuse
Ouro: one model implementation for 1.4B and 2.6B, configured by checkpoint metadata.
Ouro-Thinking: reuse the Ouro core and runtime features. Track Thinking-specific basic integration separately: checkpoint configuration, BOS/EOS, chat-template/enable_thinking behavior, and generation/reference checks. Do not create another recurrent backend solely for the Thinking weights.
Huginn: a separate architecture integration for Huginn-0125, including its prelude/recurrent/coda structure, state initialization/input injection, and cache semantics.
✅ Implemented and merged for the tracked architecture/path.
🚧 In progress: a partial integration, open implementation PR, or explicitly active work item in a linked RFC.
⬜ Not implemented/integrated for this model in the tracked scope.
➖ No separate implementation task: reuse the Ouro implementation and track its progress in the Ouro row/sub-tables.
These are engineering progress indicators, not claims of validation on every checkpoint, hardware platform, workload, or feature combination. A parent feature is 🚧 while some listed subfeatures remain incomplete. An existing shared utility does not by itself complete a new model's integration.
Progress refreshed on 2026-09-29 against main at 9e3d13d0074a63ae9f2a46287544e11c7de22b22, merged PRs, and linked work/claims. Current documented model qualification is Ouro-1.4B; the other Ouro checkpoint names below describe the intended coverage of the shared implementation. Their validation belongs in the corresponding recipes.
Each feature has an Assignee(s) column listing contributors to its claimed work. Multiple contributors may share a feature or claim different subfeatures. — means no assignee is recorded; Ouro-Thinking runtime cells marked ➖ inherit the corresponding Ouro task.
Names prefilled from existing PRs/RFC claims identify contributors only to the linked scope, as shown in the sub-tables. They do not assign all remaining work in the parent feature to those contributors. Record new subfeature claims as status · issue/PR · @username in the corresponding model cell and reflect the contributors in the overall table.
Basic implementation includes configuration and weight loading, tokenizer/chat template, model computation, fixed recurrence, greedy and temperature/top-k/top-p sampling, EOS/length handling, and text generation. These define acceptance for basic model integration; contributors may split the implementation work as needed.
For Ouro-Thinking, only basic integration requires a separate task. All other cells are marked ➖ because they reuse Ouro and require no separate implementation; their progress is maintained in the Ouro row and sub-tables. This does not mean the features are unsupported or that checkpoint validation can be skipped.
Scheduling and batching
Subfeature
Meaning
Ouro, including Thinking runtime
Nanbeige4.2
Huginn
Loop-level continuous batching
Rebuild batches at recurrence boundaries so requests at different positions/depths can execute together and new requests can enter.
Execute compatible prompt and decode recurrent tasks in the same device batch; minimize decode post-processing interference with prefill and keep shared computation and phase-specific state handling clearly separated.
Request arrival, completion, and cancellation are part of continuous batching. WCPB is distinct from mixing prefill and decode in one device batch. Scheduler ownership refactoring is tracked separately in #50; batch-selection protection is addressed by #68.
KV cache
Subfeature
Meaning
Ouro, including Thinking runtime
Nanbeige4.2
Huginn
Depth-aware paged KV
Address cached keys/values by request, token position, layer, and recurrence depth.
Define separate/shared storage across depths and the valid history used by subsequent computation.
✅
⬜
⬜
Incremental allocation
Assign physical KV pages to a request as its execution frontier grows, rather than assigning its full maximum-length footprint immediately.
✅
⬜
⬜
Prefix caching
Reuse valid KV for a shared prompt prefix to avoid repeated prefill computation.
✅
⬜
⬜
Snapshot / restore
Save a suspended request's KV and device state to CPU, release GPU resources, then restore and continue with its logical progress and RNG behavior preserved.
✅
⬜
⬜
Incremental allocation draws pages from the engine's KV pool; it does not imply allocating a new CUDA tensor for every token. Snapshot/restore is request-state preservation, not model-weight checkpointing or recomputation from scratch. Reference tracking, deferred release, and transfer leases are part of the underlying resource lifecycle. KV boundary refactoring is tracked in #57.
GPU execution
Subfeature
Meaning
Ouro, including Thinking runtime
Nanbeige4.2
Huginn
Triton / FlashAttention backends
Execute attention with the selected optimized backend while preserving model and KV semantics.
This table covers ordinary inference. @0z5a has claimed CUDA Graph support for Nanbeige4.2 and Huginn; no model-specific implementation PR is linked yet. Speculative graphs and cross-round asynchronous speculation are tracked below. Backend/hardware constraints and validated configurations must remain explicit in recipes; ✅ does not claim every FlashAttention generation works on every GPU. Attention interface refactoring is tracked in #46.
Hardware and attention backend adaptation
Adapt suitable attention backends for the hardware available to contributors, and track progress per model and hardware family. These are recommended adaptation targets, not claims that every listed backend is already integrated or fastest on every workload. The existing GPU execution table above describes the implemented CUDA paths; broader hardware/backend coverage is tracked here.
For this table: ✅ completed integration for the explicitly identified device/path; 🚧 partial integration or ongoing work/validation; ⬜ adaptation or target-device validation pending. Record the claimed contributor and implementation PR in the relevant cell when work starts. Performance and accuracy evidence belongs in the corresponding recipes.
MindIE-SD is a candidate operator integration. Confirm causal attention, paged KV and incremental decode applicability before claiming a complete language-model attention backend. Retain the existing Triton path as a comparison/fallback where supported.
Fixed recurrence is part of basic support. Adaptive exit is an optional research extension for Nanbeige4.2, not a requirement for its fixed two-loop integration. Huginn's stopping rules require their own implementation/validation rather than reuse of the Ouro gate. #56 also tracks accuracy, mean-depth, and throughput evidence. Its adaptive-exit evaluation harness and depth statistics merged in #58; this does not complete the exact-policy asynchronous exit implementation.
Self-speculative decoding
Subfeature
Meaning
Ouro, including Thinking runtime
Nanbeige4.2
Huginn
Synchronous greedy / sampling speculation
Draft with a shallower recurrence, verify at full depth, and commit a valid suffix; includes correction/bonus tokens, ragged requests, and KV rollback.
Combine wavefront prefill with valid chunk/depth handoff boundaries.
⬜
⬜
⬜
Self-speculation integration
Let the decode worker run draft/verify/rollback after receiving the prompt state from prefill.
⬜
⬜
⬜
PD speculation means combining PD deployment with self-speculative decoding; it does not require putting the draft and target on separate P/D workers. Handoff completion, cancellation, timeouts, and failure cleanup are part of basic PD acceptance. Current PD scope is multiple GPUs on one host. M9 ownership/interface refactoring is tracked in #64 and does not, by itself, complete WCPB or speculation integration.
Independent tasks, dependencies, and inherited support
These relationships describe implementation scope, not contributor ownership. A dependent task can have a different assignee from its prerequisite and can be developed in parallel against an agreed interface.
Relationship
Scope
What to implement / reuse
Shared architecture
Ouro 1.4B / 2.6B; Nanbeige4.2 3B / 3B-Base
Reuse the corresponding architecture implementation; validate checkpoint-specific configuration and behavior without duplicating runtime tasks.
Inherited runtime support
Ouro-Thinking
Claim only Thinking-specific basic integration. Runtime features inherit Ouro's implementation and status; keep their overall cells ➖ and validate Thinking checkpoints in recipes.
Separate model integration
Ouro, Nanbeige4.2, Huginn
Each architecture needs its own computation and state semantics. Shared engine infrastructure can be reused, but each model's feature integration and validation remain tracked separately.
Separately claimable runtime work
Scheduling, KV management, GPU execution, adaptive depth, self-speculation, PD, and their subfeatures
Contributors may claim a whole feature or a bounded subfeature. These are separate work scopes, not a promise that they have no technical dependencies.
Dependent scheduling / state work
WCPB; prefix caching; snapshot/restore
WCPB builds on chunked prefill and recurrence-aware scheduling. Prefix reuse needs valid model/depth-aware KV semantics. Snapshot/restore needs request-state preservation and resource lifecycle handling.
Async exit needs a defined model-specific stopping policy and async execution. Speculative graphs and async speculation build on correct draft/verify/commit/rollback behavior plus compatible execution and buffer-lifetime support.
Each can be claimed separately on top of the speculative path. They must preserve its correctness; preemption/migration additionally needs request-state preservation. None automatically completes the other optimizations.
Explicit cross-feature integration
PD prefix-cache, async/graph, WCPB, and self-speculation integration
Reuse PD and the corresponding feature, then implement and validate their interaction. Completing both standalone features does not automatically complete their PD integration cell.
For each implementation issue, list concrete prerequisites and reused components. A dependency is not a new implementation task if the needed support already exists. Inherited support must be distinguished from integration that still needs model-specific work or validation; do not mark another architecture complete merely because Ouro supports it.
Task assignment and claiming
@Carlos779988 owns the Ouro mixed prefill/decode batching work and should coordinate its design with @LyxWxj before building on #60. @0z5a owns Nanbeige4.2/Huginn CUDA Graph support; high-concurrency graph hit-rate optimization is an additional suggested direction. @CXJorz is getting familiar with the project through the user guide; no implementation task is assigned yet.
Contributors may freely claim and split work at any practical scope, individually or collaboratively. Comment with the intended scope and link the implementation issue/PR; record the contributor(s) under Assignee(s) and note relevant dependencies. Ownership applies only to the claimed scope.
Track subfeature progress independently and mark a parent feature complete once all applicable subfeatures are integrated and validated. Technical dependencies and inherited support describe how implementations relate; they do not restrict task assignment.
Completion and maintenance
Update the relevant cell, assignee(s), and linked implementation issue/PRs when work starts or merges. Ownership applies only to the explicitly claimed scope; parent status aggregates the subfeature results.
For basic model integration, validate the intended checkpoint against its reference computation and generation behavior. Reusing an architecture does not automatically qualify every checkpoint.
For a feature, exercise its real engine path, relevant request/KV/RNG lifecycle, and the combinations being claimed. Unsupported combinations should be rejected explicitly and documented.
For performance work, report matched, repeated before/after results, including high-concurrency/large-batch workloads and low-concurrency controls where relevant. Fewer synchronizations or kernels alone is not a throughput claim.
Every new model or feature must update the README support matrix, relevant documentation, and its recipe. Keep commands, checkpoint/configuration pins, correctness results, and performance measurements in the recipe; avoid adding a separate benchmark script for every tracker cell.
Alternatives and impact
We considered separate implementations/progress matrices for every checkpoint size and Thinking variant. The shared-architecture approach keeps runtime work in one place while preserving a small, explicit Thinking integration task and separate checkpoint validation. A single flat feature list would hide important PD, asynchronous, and speculative integration gaps, so those features have sub-tables.
This RFC itself changes no runtime behavior, public API, or memory usage. Implementation PRs must describe any API/configuration changes and their measured performance and memory effects. It does not expand model scope beyond the three families above, introduce training integration, or declare performance gains without evidence.
Feedback Period.
At least one week after publication. Keep this issue open as a living implementation tracker after the initial scope discussion.
Related architecture and implementation work is linked in the tables. The progress snapshot should be maintained as those PRs evolve; a refactoring PR and a feature-support PR are different kinds of work.
Before submitting a new issue...
Make sure you already searched for relevant issues in the issue tracker, and looked for the answer in the documentation.
Motivation.
Maintain one model and feature progress tracker for vllm-rlt, initially limited to Ouro, Nanbeige4.2, and Huginn. Contributors should be able to see which implementations exist, which work is in progress, and which combinations still need integration.
Track architecture implementations rather than duplicating the same backend for every checkpoint size. Keep PD as a core priority alongside loop-level batching, depth-aware KV, asynchronous execution, adaptive depth, and self-speculative decoding.
This issue complements the architecture/refactoring roadmap in #32. Recipes and training/rollout integration will have separate RFCs; this issue links to their results rather than collecting unrelated benchmark scripts or training tasks.
Proposed Change.
Scope and model reuse
enable_thinkingbehavior, and generation/reference checks. Do not create another recurrent backend solely for the Thinking weights.Official references: Ouro, Ouro-Thinking, Nanbeige4.2, Huginn.
Status convention
These are engineering progress indicators, not claims of validation on every checkpoint, hardware platform, workload, or feature combination. A parent feature is 🚧 while some listed subfeatures remain incomplete. An existing shared utility does not by itself complete a new model's integration.
Progress refreshed on 2026-09-29 against
mainat9e3d13d0074a63ae9f2a46287544e11c7de22b22, merged PRs, and linked work/claims. Current documented model qualification is Ouro-1.4B; the other Ouro checkpoint names below describe the intended coverage of the shared implementation. Their validation belongs in the corresponding recipes.Overall progress
Each feature has an Assignee(s) column listing contributors to its claimed work. Multiple contributors may share a feature or claim different subfeatures.
—means no assignee is recorded; Ouro-Thinking runtime cells marked ➖ inherit the corresponding Ouro task.Names prefilled from existing PRs/RFC claims identify contributors only to the linked scope, as shown in the sub-tables. They do not assign all remaining work in the parent feature to those contributors. Record new subfeature claims as
status · issue/PR · @usernamein the corresponding model cell and reflect the contributors in the overall table.Basic implementation includes configuration and weight loading, tokenizer/chat template, model computation, fixed recurrence, greedy and temperature/top-k/top-p sampling, EOS/length handling, and text generation. These define acceptance for basic model integration; contributors may split the implementation work as needed.
For Ouro-Thinking, only basic integration requires a separate task. All other cells are marked ➖ because they reuse Ouro and require no separate implementation; their progress is maintained in the Ouro row and sub-tables. This does not mean the features are unsupported or that checkpoint validation can be skipped.
Scheduling and batching
Request arrival, completion, and cancellation are part of continuous batching. WCPB is distinct from mixing prefill and decode in one device batch. Scheduler ownership refactoring is tracked separately in #50; batch-selection protection is addressed by #68.
KV cache
Incremental allocation draws pages from the engine's KV pool; it does not imply allocating a new CUDA tensor for every token. Snapshot/restore is request-state preservation, not model-weight checkpointing or recomputation from scratch. Reference tracking, deferred release, and transfer leases are part of the underlying resource lifecycle. KV boundary refactoring is tracked in #57.
GPU execution
This table covers ordinary inference. @0z5a has claimed CUDA Graph support for Nanbeige4.2 and Huginn; no model-specific implementation PR is linked yet. Speculative graphs and cross-round asynchronous speculation are tracked below. Backend/hardware constraints and validated configurations must remain explicit in recipes; ✅ does not claim every FlashAttention generation works on every GPU. Attention interface refactoring is tracked in #46.
Hardware and attention backend adaptation
Adapt suitable attention backends for the hardware available to contributors, and track progress per model and hardware family. These are recommended adaptation targets, not claims that every listed backend is already integrated or fastest on every workload. The existing GPU execution table above describes the implemented CUDA paths; broader hardware/backend coverage is tracked here.
For this table: ✅ completed integration for the explicitly identified device/path; 🚧 partial integration or ongoing work/validation; ⬜ adaptation or target-device validation pending. Record the claimed contributor and implementation PR in the relevant cell when work starts. Performance and accuracy evidence belongs in the corresponding recipes.
Selection references: FlashAttention, FlashInfer hardware support, AMD backend guidance, and MindIE-SD attention APIs.
Adaptive depth
Fixed recurrence is part of basic support. Adaptive exit is an optional research extension for Nanbeige4.2, not a requirement for its fixed two-loop integration. Huginn's stopping rules require their own implementation/validation rather than reuse of the Ouro gate. #56 also tracks accuracy, mean-depth, and throughput evidence. Its adaptive-exit evaluation harness and depth statistics merged in #58; this does not complete the exact-policy asynchronous exit implementation.
Self-speculative decoding
#43 is the algorithm/feature RFC; #47 tracks evaluation. #48 is merged for fixed-depth synchronous speculative CUDA Graph execution; reported device validation covers Ouro-1.4B BF16/Triton on RTX 5090, while packed FA4 graph capture remains unverified. #51, #52 and #66 remain open; #66 targets a restricted greedy/Triton path, and #52 explicitly excludes PD integration.
Prefill/decode disaggregation (PD)
PD speculation means combining PD deployment with self-speculative decoding; it does not require putting the draft and target on separate P/D workers. Handoff completion, cancellation, timeouts, and failure cleanup are part of basic PD acceptance. Current PD scope is multiple GPUs on one host. M9 ownership/interface refactoring is tracked in #64 and does not, by itself, complete WCPB or speculation integration.
Independent tasks, dependencies, and inherited support
These relationships describe implementation scope, not contributor ownership. A dependent task can have a different assignee from its prerequisite and can be developed in parallel against an agreed interface.
For each implementation issue, list concrete prerequisites and reused components. A dependency is not a new implementation task if the needed support already exists. Inherited support must be distinguished from integration that still needs model-specific work or validation; do not mark another architecture complete merely because Ouro supports it.
Task assignment and claiming
@Carlos779988 owns the Ouro mixed prefill/decode batching work and should coordinate its design with @LyxWxj before building on #60. @0z5a owns Nanbeige4.2/Huginn CUDA Graph support; high-concurrency graph hit-rate optimization is an additional suggested direction. @CXJorz is getting familiar with the project through the user guide; no implementation task is assigned yet.
Contributors may freely claim and split work at any practical scope, individually or collaboratively. Comment with the intended scope and link the implementation issue/PR; record the contributor(s) under Assignee(s) and note relevant dependencies. Ownership applies only to the claimed scope.
Track subfeature progress independently and mark a parent feature complete once all applicable subfeatures are integrated and validated. Technical dependencies and inherited support describe how implementations relate; they do not restrict task assignment.
Completion and maintenance
Alternatives and impact
We considered separate implementations/progress matrices for every checkpoint size and Thinking variant. The shared-architecture approach keeps runtime work in one place while preserving a small, explicit Thinking integration task and separate checkpoint validation. A single flat feature list would hide important PD, asynchronous, and speculative integration gaps, so those features have sub-tables.
This RFC itself changes no runtime behavior, public API, or memory usage. Implementation PRs must describe any API/configuration changes and their measured performance and memory effects. It does not expand model scope beyond the three families above, introduce training integration, or declare performance gains without evidence.
Feedback Period.
At least one week after publication. Keep this issue open as a living implementation tracker after the initial scope discussion.
CC List.
@bjf-frz @hsliuustc0106
Anything else.
Related architecture and implementation work is linked in the tables. The progress snapshot should be maintained as those PRs evolve; a refactoring PR and a feature-support PR are different kinds of work.
Before submitting a new issue...