Skip to content

[RFC]: Model support and feature roadmap for Ouro, Nanbeige4.2, and Huginn #69

Description

@bjf-frz

Motivation.

Maintain one model and feature progress tracker for vllm-rlt, initially limited to Ouro, Nanbeige4.2, and Huginn. Contributors should be able to see which implementations exist, which work is in progress, and which combinations still need integration.

Track architecture implementations rather than duplicating the same backend for every checkpoint size. Keep PD as a core priority alongside loop-level batching, depth-aware KV, asynchronous execution, adaptive depth, and self-speculative decoding.

This issue complements the architecture/refactoring roadmap in #32. Recipes and training/rollout integration will have separate RFCs; this issue links to their results rather than collecting unrelated benchmark scripts or training tasks.

Proposed Change.

Scope and model reuse

  • Ouro: one model implementation for 1.4B and 2.6B, configured by checkpoint metadata.
  • Ouro-Thinking: reuse the Ouro core and runtime features. Track Thinking-specific basic integration separately: checkpoint configuration, BOS/EOS, chat-template/enable_thinking behavior, and generation/reference checks. Do not create another recurrent backend solely for the Thinking weights.
  • Nanbeige4.2: one implementation covering 3B and 3B-Base, starting with the official fixed two-loop semantics. Native support is being developed in [MODEL] : support Nanbeige/Nanbeige4.2-3B recurrent causal LM #63.
  • Huginn: a separate architecture integration for Huginn-0125, including its prelude/recurrent/coda structure, state initialization/input injection, and cache semantics.

Official references: Ouro, Ouro-Thinking, Nanbeige4.2, Huginn.

Status convention

  • ✅ Implemented and merged for the tracked architecture/path.
  • 🚧 In progress: a partial integration, open implementation PR, or explicitly active work item in a linked RFC.
  • ⬜ Not implemented/integrated for this model in the tracked scope.
  • ➖ No separate implementation task: reuse the Ouro implementation and track its progress in the Ouro row/sub-tables.

These are engineering progress indicators, not claims of validation on every checkpoint, hardware platform, workload, or feature combination. A parent feature is 🚧 while some listed subfeatures remain incomplete. An existing shared utility does not by itself complete a new model's integration.

Progress refreshed on 2026-09-29 against main at 9e3d13d0074a63ae9f2a46287544e11c7de22b22, merged PRs, and linked work/claims. Current documented model qualification is Ouro-1.4B; the other Ouro checkpoint names below describe the intended coverage of the shared implementation. Their validation belongs in the corresponding recipes.

Overall progress

Model / integration Checkpoint scope Basic implementation Assignee(s) Scheduling / batching Assignee(s) KV cache Assignee(s) GPU execution Assignee(s) Adaptive depth Assignee(s) Self-speculation Assignee(s) PD Assignee(s)
Ouro 1.4B, 2.6B ✅ — 🚧 @LyxWxj, @Carlos779988 ✅ — ✅ — 🚧 @MinhaoLi0318 🚧 @0z5a, @IDEA-V, @liuyao0322, @Dmaner 🚧 —
Ouro-Thinking 1.4B-Thinking, 2.6B-Thinking ⬜ — ➖ — ➖ — ➖ — ➖ — ➖ — ➖ —
Nanbeige4.2 3B, 3B-Base 🚧 #63 @cs-ai-li 🚧 #63 @cs-ai-li 🚧 #63 @cs-ai-li 🚧 @0z5a ⬜ — ⬜ — ⬜ —
Huginn Huginn-0125 ⬜ — ⬜ — ⬜ — 🚧 @0z5a ⬜ — ⬜ — ⬜ —

Each feature has an Assignee(s) column listing contributors to its claimed work. Multiple contributors may share a feature or claim different subfeatures. — means no assignee is recorded; Ouro-Thinking runtime cells marked ➖ inherit the corresponding Ouro task.

Names prefilled from existing PRs/RFC claims identify contributors only to the linked scope, as shown in the sub-tables. They do not assign all remaining work in the parent feature to those contributors. Record new subfeature claims as status · issue/PR · @username in the corresponding model cell and reflect the contributors in the overall table.

Basic implementation includes configuration and weight loading, tokenizer/chat template, model computation, fixed recurrence, greedy and temperature/top-k/top-p sampling, EOS/length handling, and text generation. These define acceptance for basic model integration; contributors may split the implementation work as needed.

For Ouro-Thinking, only basic integration requires a separate task. All other cells are marked ➖ because they reuse Ouro and require no separate implementation; their progress is maintained in the Ouro row and sub-tables. This does not mean the features are unsupported or that checkpoint validation can be skipped.

Scheduling and batching

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
Loop-level continuous batching Rebuild batches at recurrence boundaries so requests at different positions/depths can execute together and new requests can enter. ✅ 🚧 #63 · @cs-ai-li ⬜
Chunked prefill Process a long prompt in smaller chunks to bound how long prefill occupies execution resources. ✅ ⬜ ⬜
Priority scheduling / preemption Prefer higher-priority requests and suspend eligible requests when capacity is needed. ✅ ⬜ ⬜
Wavefront chunked prefill batching (WCPB) Batch ready prompt chunks at different recurrence depths while preserving causal dependencies. 🚧 #60, #65 · @LyxWxj ⬜ ⬜
Mixed prefill/decode batching Execute compatible prompt and decode recurrent tasks in the same device batch; minimize decode post-processing interference with prefill and keep shared computation and phase-specific state handling clearly separated. 🚧 · @Carlos779988 (claimed) ⬜ ⬜

Request arrival, completion, and cancellation are part of continuous batching. WCPB is distinct from mixing prefill and decode in one device batch. Scheduler ownership refactoring is tracked separately in #50; batch-selection protection is addressed by #68.

KV cache

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
Depth-aware paged KV Address cached keys/values by request, token position, layer, and recurrence depth. ✅ 🚧 #63 · @cs-ai-li ⬜
KV layouts Define separate/shared storage across depths and the valid history used by subsequent computation. ✅ ⬜ ⬜
Incremental allocation Assign physical KV pages to a request as its execution frontier grows, rather than assigning its full maximum-length footprint immediately. ✅ ⬜ ⬜
Prefix caching Reuse valid KV for a shared prompt prefix to avoid repeated prefill computation. ✅ ⬜ ⬜
Snapshot / restore Save a suspended request's KV and device state to CPU, release GPU resources, then restore and continue with its logical progress and RNG behavior preserved. ✅ ⬜ ⬜

Incremental allocation draws pages from the engine's KV pool; it does not imply allocating a new CUDA tensor for every token. Snapshot/restore is request-state preservation, not model-weight checkpointing or recomputation from scratch. Reference tracking, deferred release, and transfer leases are part of the underlying resource lifecycle. KV boundary refactoring is tracked in #57.

GPU execution

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
Triton / FlashAttention backends Execute attention with the selected optimized backend while preserving model and KV semantics. ✅ #30 ⬜ ⬜
Asynchronous scheduling Prepare/submit subsequent work while earlier GPU work is still running. ✅ #30 ⬜ ⬜
Multi-stream execution Organize compute, transfers, and readbacks on CUDA streams with explicit event dependencies. ✅ #30 ⬜ ⬜
Device-resident state / buffer reuse Retain ordinary-decode tokens, hidden state, and metadata on GPU and reuse bounded workspaces. ✅ #30 ⬜ ⬜
Recurrent CUDA Graph Capture/replay recurrent-core GPU operations to reduce CPU kernel-submission overhead. ✅ #30 🚧 · @0z5a (claimed) 🚧 · @0z5a (claimed)

This table covers ordinary inference. @0z5a has claimed CUDA Graph support for Nanbeige4.2 and Huginn; no model-specific implementation PR is linked yet. Speculative graphs and cross-round asynchronous speculation are tracked below. Backend/hardware constraints and validated configurations must remain explicit in recipes; ✅ does not claim every FlashAttention generation works on every GPU. Attention interface refactoring is tracked in #46.

Hardware and attention backend adaptation

Adapt suitable attention backends for the hardware available to contributors, and track progress per model and hardware family. These are recommended adaptation targets, not claims that every listed backend is already integrated or fastest on every workload. The existing GPU execution table above describes the implemented CUDA paths; broader hardware/backend coverage is tracked here.

For this table: ✅ completed integration for the explicitly identified device/path; 🚧 partial integration or ongoing work/validation; ⬜ adaptation or target-device validation pending. Record the claimed contributor and implementation PR in the relevant cell when work starts. Performance and accuracy evidence belongs in the corresponding recipes.

Hardware family Recommended attention backend Ouro Nanbeige4.2 Huginn
NVIDIA Ampere: A100 / RTX 30 FA2, FlashInfer FA2 🚧; FlashInfer ⬜ ⬜ ⬜
NVIDIA Ada: L40 / L40S / RTX 40 FA2, FlashInfer FA2 🚧; FlashInfer ⬜ ⬜ ⬜
NVIDIA Hopper: H100 / H800 / H200 FA3 / FA4, FlashInfer FA3 🚧; FA4 🚧 #47; FlashInfer ⬜ ⬜ ⬜
NVIDIA Blackwell: B200 / B300 FA4, FlashInfer FA4 ✅ (B300, #30); FlashInfer ⬜ ⬜ ⬜
NVIDIA Blackwell RTX: RTX 50 / RTX PRO Blackwell FlashInfer with an SM120-compatible path ⬜ ⬜ ⬜
AMD Instinct: MI300 / MI325 / MI350 series AITER ⬜ ⬜ ⬜
AMD Radeon: RX 7000 / RX 9000 series ROCm Triton; HIP paged attention as an optimization candidate ⬜ ⬜ ⬜
Ascend NPU CANN fused inference attention through torch_npu; MindIE-SD attention operators as a candidate ⬜ ⬜ ⬜
  • Ouro-Thinking inherits Ouro. Existing Ouro hardware evidence primarily covers 1.4B; it does not qualify every checkpoint or device in the same family.
  • FA2/FA3 interfaces already exist, but target-device validation remains pending. H800 FA4 results are reported in the still-open [BENCH] Evaluate synchronous self-speculation on H800 #47. Nanbeige basic integration is underway in [MODEL] : support Nanbeige/Nanbeige4.2-3B recurrent causal LM #63; that alone does not establish support for the recommended backends above.
  • MindIE-SD is a candidate operator integration. Confirm causal attention, paged KV and incremental decode applicability before claiming a complete language-model attention backend. Retain the existing Triton path as a comparison/fallback where supported.

Selection references: FlashAttention, FlashInfer hardware support, AMD backend guidance, and MindIE-SD attention APIs.

Adaptive depth

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
Model-specific early exit Use the model's gate or an explicitly defined stopping rule to decide whether a token needs another recurrent traversal. ✅ ⬜ ⬜
Delayed exit Continue submission while an exit signal is pending, with explicitly documented delayed-policy semantics. ✅ #30 ⬜ ⬜
Exact-policy asynchronous exit Preserve the exit depth and hidden state selected by the original policy despite asynchronous execution. 🚧 #56 · @MinhaoLi0318 ⬜ ⬜

Fixed recurrence is part of basic support. Adaptive exit is an optional research extension for Nanbeige4.2, not a requirement for its fixed two-loop integration. Huginn's stopping rules require their own implementation/validation rather than reuse of the Ouro gate. #56 also tracks accuracy, mean-depth, and throughput evidence. Its adaptive-exit evaluation harness and depth statistics merged in #58; this does not complete the exact-policy asynchronous exit implementation.

Self-speculative decoding

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
Synchronous greedy / sampling speculation Draft with a shallower recurrence, verify at full depth, and commit a valid suffix; includes correction/bonus tokens, ragged requests, and KV rollback. ✅ #44 ⬜ ⬜
Speculative CUDA Graph Capture/replay compatible draft and verification operations. ✅ #48 · @0z5a ⬜ ⬜
Token / metadata readback reduction Batch or remove CPU/GPU round trips for draft IDs, verification results, and execution metadata. 🚧 #51 · @IDEA-V ⬜ ⬜
Cross-round asynchronous speculation Submit the next speculative round before the previous result is delivered to CPU. 🚧 #66 · @liuyao0322 ⬜ ⬜
Adaptive K / draft depth Choose the candidate count and draft recurrence depth based on acceptance and execution cost/load. 🚧 #43 · @Dmaner ⬜ ⬜
Speculative preemption / migration Suspend or migrate at safe boundaries while preserving committed outputs, recurrent state, and RNG behavior. 🚧 #52 · @0z5a ⬜ ⬜

#43 is the algorithm/feature RFC; #47 tracks evaluation. #48 is merged for fixed-depth synchronous speculative CUDA Graph execution; reported device validation covers Ouro-1.4B BF16/Triton on RTX 5090, while packed FA4 graph capture remains unverified. #51, #52 and #66 remain open; #66 targets a restricted greedy/Triton path, and #52 explicitly excludes PD integration.

Prefill/decode disaggregation (PD)

Subfeature Meaning Ouro, including Thinking runtime Nanbeige4.2 Huginn
P/D separation and KV/hidden handoff Prefill workers process prompts and transfer the state needed for decode workers to continue generation. ✅ #31 ⬜ ⬜
Chunked transfer / compute overlap Transfer completed prompt chunks while later prefill computation continues. ✅ #31 ⬜ ⬜
Multiple P/D workers / resource credits Route requests across worker pools and reserve receiving/execution capacity. ✅ #31 ⬜ ⬜
Prefix-cache integration Reuse valid cached prefixes within PD and adjust the required computation/transfer ranges. ✅ #31 ⬜ ⬜
Async / CUDA Graph integration Use asynchronous execution and graphs in PD workers while respecting receive completion and buffer lifetimes. ✅ #31 ⬜ ⬜
WCPB integration Combine wavefront prefill with valid chunk/depth handoff boundaries. ⬜ ⬜ ⬜
Self-speculation integration Let the decode worker run draft/verify/rollback after receiving the prompt state from prefill. ⬜ ⬜ ⬜

PD speculation means combining PD deployment with self-speculative decoding; it does not require putting the draft and target on separate P/D workers. Handoff completion, cancellation, timeouts, and failure cleanup are part of basic PD acceptance. Current PD scope is multiple GPUs on one host. M9 ownership/interface refactoring is tracked in #64 and does not, by itself, complete WCPB or speculation integration.

Independent tasks, dependencies, and inherited support

These relationships describe implementation scope, not contributor ownership. A dependent task can have a different assignee from its prerequisite and can be developed in parallel against an agreed interface.

Relationship Scope What to implement / reuse
Shared architecture Ouro 1.4B / 2.6B; Nanbeige4.2 3B / 3B-Base Reuse the corresponding architecture implementation; validate checkpoint-specific configuration and behavior without duplicating runtime tasks.
Inherited runtime support Ouro-Thinking Claim only Thinking-specific basic integration. Runtime features inherit Ouro's implementation and status; keep their overall cells ➖ and validate Thinking checkpoints in recipes.
Separate model integration Ouro, Nanbeige4.2, Huginn Each architecture needs its own computation and state semantics. Shared engine infrastructure can be reused, but each model's feature integration and validation remain tracked separately.
Separately claimable runtime work Scheduling, KV management, GPU execution, adaptive depth, self-speculation, PD, and their subfeatures Contributors may claim a whole feature or a bounded subfeature. These are separate work scopes, not a promise that they have no technical dependencies.
Dependent scheduling / state work WCPB; prefix caching; snapshot/restore WCPB builds on chunked prefill and recurrence-aware scheduling. Prefix reuse needs valid model/depth-aware KV semantics. Snapshot/restore needs request-state preservation and resource lifecycle handling.
Dependent execution work Delayed / exact-policy asynchronous exit; speculative CUDA Graph; cross-round asynchronous speculation Async exit needs a defined model-specific stopping policy and async execution. Speculative graphs and async speculation build on correct draft/verify/commit/rollback behavior plus compatible execution and buffer-lifetime support.
Separate speculative optimizations Token/readback reduction; adaptive K / draft depth; speculative preemption / migration Each can be claimed separately on top of the speculative path. They must preserve its correctness; preemption/migration additionally needs request-state preservation. None automatically completes the other optimizations.
Explicit cross-feature integration PD prefix-cache, async/graph, WCPB, and self-speculation integration Reuse PD and the corresponding feature, then implement and validate their interaction. Completing both standalone features does not automatically complete their PD integration cell.

For each implementation issue, list concrete prerequisites and reused components. A dependency is not a new implementation task if the needed support already exists. Inherited support must be distinguished from integration that still needs model-specific work or validation; do not mark another architecture complete merely because Ouro supports it.

Task assignment and claiming

@Carlos779988 owns the Ouro mixed prefill/decode batching work and should coordinate its design with @LyxWxj before building on #60. @0z5a owns Nanbeige4.2/Huginn CUDA Graph support; high-concurrency graph hit-rate optimization is an additional suggested direction. @CXJorz is getting familiar with the project through the user guide; no implementation task is assigned yet.

Contributors may freely claim and split work at any practical scope, individually or collaboratively. Comment with the intended scope and link the implementation issue/PR; record the contributor(s) under Assignee(s) and note relevant dependencies. Ownership applies only to the claimed scope.

Track subfeature progress independently and mark a parent feature complete once all applicable subfeatures are integrated and validated. Technical dependencies and inherited support describe how implementations relate; they do not restrict task assignment.

Completion and maintenance

  • Update the relevant cell, assignee(s), and linked implementation issue/PRs when work starts or merges. Ownership applies only to the explicitly claimed scope; parent status aggregates the subfeature results.
  • For basic model integration, validate the intended checkpoint against its reference computation and generation behavior. Reusing an architecture does not automatically qualify every checkpoint.
  • For a feature, exercise its real engine path, relevant request/KV/RNG lifecycle, and the combinations being claimed. Unsupported combinations should be rejected explicitly and documented.
  • For performance work, report matched, repeated before/after results, including high-concurrency/large-batch workloads and low-concurrency controls where relevant. Fewer synchronizations or kernels alone is not a throughput claim.
  • Every new model or feature must update the README support matrix, relevant documentation, and its recipe. Keep commands, checkpoint/configuration pins, correctness results, and performance measurements in the recipe; avoid adding a separate benchmark script for every tracker cell.

Alternatives and impact

We considered separate implementations/progress matrices for every checkpoint size and Thinking variant. The shared-architecture approach keeps runtime work in one place while preserving a small, explicit Thinking integration task and separate checkpoint validation. A single flat feature list would hide important PD, asynchronous, and speculative integration gaps, so those features have sub-tables.

This RFC itself changes no runtime behavior, public API, or memory usage. Implementation PRs must describe any API/configuration changes and their measured performance and memory effects. It does not expand model scope beyond the three families above, introduce training integration, or declare performance gains without evidence.

Feedback Period.

At least one week after publication. Keep this issue open as a living implementation tracker after the initial scope discussion.

CC List.

@bjf-frz @hsliuustc0106

Anything else.

Related architecture and implementation work is linked in the tables. The progress snapshot should be maintained as those PRs evolve; a refactoring PR and a feature-support PR are different kinds of work.

Before submitting a new issue...

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

RFCRequests for comments on major architectural changes or design choices

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions