Motivation.
Enable training-driven improvements to vllm-rlt's looped-transformer inference efficiency. Optimize throughput, latency, and memory use under an explicit quality constraint.
Looped transformers provide specific opportunities: predict exits ahead of execution, align shallow draft predictions with deeper target computation, and make better use of recurrent state and shared weights. Exploring these ideas requires reliable rollout and weight-update interfaces before integrating a training framework.
The proposed order is:
- Add framework-independent training support to vllm-rlt through repository PRs.
- Integrate and run training loops on contributor forks of verl/vime.
- Use those forks to train and evaluate loop-specific inference optimizations; upstream useful engine capabilities with evidence.
Proposed Change.
Scope, ownership, and progress
Initial integration target: Ouro-1.4B, BF16, synchronous training iterations, and full weight updates. Training interfaces and framework adapters must preserve existing inference features, including early exit, PD, speculative decoding, and CUDA Graphs, from the start. Do not impose a fixed-depth, single-device, non-speculative restriction on the integration. Raise concrete compatibility issues for discussion; this requirement does not imply that currently unsupported feature combinations are already implemented. Asynchronous training, incremental weight transfer, and colocated training/inference memory switching remain separate extensions.
Start with verl + FSDP as the proposed first integration; add vime with its training stack through a second adapter. Pin framework versions and validate model compatibility in the integration itself. Train/inference consistency and gradient checks are required integration tests.
- ✅ Completed for the stated scope.
- 🚧 Claimed or in progress, with a linked claim, PR, or fork branch; not yet accepted as complete.
- ⬜ Not yet delivered in this tracker; existing components should be reused where applicable.
Contributors may split tasks or collaborate. Record Assignee(s), engine PRs, and fork branch/results links. Phase 1 is a joint development scope for @MinhaoLi0318 and @IDEA-V, coordinated with @0z5a and @Levius-Fubuki on the shared interfaces. @Levius-Fubuki has also expressed follow-on interest in N+2 exit work; that experiment is not yet marked as started. @CXJorz is onboarding and has no implementation task assigned yet.
Implementation references
Reference vLLM's rollout and weight-synchronization implementation, and adapt verl/vime's existing vLLM integrations. Preserve recurrent shared-weight and depth-aware KV semantics, documenting necessary differences and the referenced versions.
Implementation details, experiment procedures, training methods and PR boundaries are left to contributors. The requirements below define the deliverables and acceptance criteria.
Phase 1: training-facing capabilities in vllm-rlt
Provide a shared engine contract based on the relevant vLLM interfaces, usable by both frameworks without adding verl/vime dependencies to normal inference installation. Prefer compatible field and operation names; document required deviations in the implementation PRs.
| Task |
Scope and acceptance |
Status |
Assignee(s) |
PR / plan |
| Training rollout interface |
Token-ID input/output, selected-token logprobs, effective sampling configuration, concurrent requests and request lifecycle. Preserve token alignment and termination reasons; support cancellation/error cleanup; document model-versus-sampling probability conventions. |
🚧 |
@MinhaoLi0318, @IDEA-V |
Plan, co-author coordination |
| Weight update and synchronization |
Support reliable full-weight updates with model-version tracking and correct cache/state handling. Generation must use a consistent weight version, and update failures must be reported safely. |
🚧 |
@MinhaoLi0318, @IDEA-V |
Plan, co-author coordination |
Acceptance requires correct token/probability alignment, concurrent request handling, and consistent generation after weight updates. Document probability conventions, preserve recurrent weight sharing, and prevent reuse of incompatible cached state. Keep the interfaces framework-independent and include relevant correctness and overhead measurements.
The agreed Phase 1 delivery is split into three PRs: rollout outputs; full-weight update and synchronization; and concurrent token-in/token-out API with abort/pause. Build on vLLM and @0z5a’s existing prototype, with shared development and appropriate co-author credit. Existing inference-feature compatibility is part of this delivery, not a later phase.
Phase 2: integration on verl/vime forks
After the engine contract is usable, adapt each framework’s existing vLLM integration to vllm-rlt and run end-to-end experiments on contributor forks. Upstream PRs to verl or vime are not required to complete this phase. Missing generic engine capabilities should be proposed back to vllm-rlt rather than maintained as two incompatible private interfaces.
| Task |
Scope and acceptance |
Status |
Assignee(s) |
Branch / evidence |
| verl/FSDP integration |
Adapt Ouro training and the vllm-rlt rollout backend, map/export shared recurrent weights, coordinate updates, and provide a reproducible multi-iteration training recipe with consistency and save/resume checks. |
🚧 |
@Levius-Fubuki |
Claim |
| vime integration |
Adapt Ouro to the training stack and reuse the engine rollout/update contract; provide parameter conversion and a reproducible multi-iteration recipe with consistency and save/resume checks. |
🚧 |
@0z5a |
vime prototype, engine prototype |
Provide a reproducible end-to-end training run that verifies correct gradients, weight updates, synchronization, and generation with the updated model. Include train/inference probability consistency and checkpoint recovery evidence. Link the runnable fork branch, configuration and results.
Phase 3: loop-specific training experiments for inference efficiency
Use the integrated training environment to optimize recurrent execution.
| Task |
Scope and acceptance |
Status |
Assignee(s) |
Branch / results |
| Lookahead exit optimization |
Train and deploy the N+2 exit predictor, including data/labels, checkpoint loading and runtime decisions. Evaluate task quality, selected versus executed depth, and low/high-concurrency throughput and latency. |
⬜ |
@Levius-Fubuki (follow-on interest) |
Follow-on interest |
| Loop-aware speculative optimization |
Train shallow draft heads/adapters or depth-specific fine-tuning, integrate the resulting draft path, and evaluate quality, acceptance lengths, computation reuse, memory and end-to-end speedup. |
🚧 |
@IDEA-V |
Claim |
Experiment goals and acceptance
- N+2 exit: use information available at round N to predict whether the token can exit at round N+2; execute through N+2 and use that round's hidden state for output. Preserve valid exit/KV semantics and evaluate quality, selected versus actual executed depth, and throughput/latency against relevant existing baselines.
- Loop-aware speculation: improve shallow drafting while preserving valid target verification and identifying reusable recurrent computation. Evaluate output quality, acceptance lengths, extra computation/memory and end-to-end performance against target-only and current speculative inference.
Training architecture, labels, losses, datasets and experimental procedures are chosen by contributors and documented with results. Add reusable trace collection, trained-component loading or other engine support as required by the experiment. Adaptive training/rollout must preserve consistent trajectory and KV semantics.
Code location and PR scope
| Code / artifact |
Destination |
| Generic rollout, probability, versioning and update interfaces |
vllm-rlt PRs |
| Framework adapters, differentiable model adaptation, weight conversion and training configuration |
Contributor verl/vime forks initially |
| Labels, losses, training scripts and experimental results |
Framework forks / linked experiment artifacts |
| Experimental exit/draft execution changes |
Linked engine experiment branches, paired with framework forks |
| Useful trace collection or trained-component loading |
vllm-rlt PRs after validating the shared requirement |
| New runtime exit/speculation policy |
vllm-rlt PR with correctness, quality and performance evidence |
Negative or inconclusive experiments should retain results without requiring a production runtime merge. Keep framework-specific dependencies outside the normal inference installation. Update README/relevant docs for new public engine capabilities and supply reproducible integration/experiment instructions.
Alternatives and impact
Use a shared engine contract for both framework adapters to avoid duplicating rollout and weight-update protocols. Keep framework dependencies in the fork integrations.
Implementation PRs must document API changes, probability-collection overhead, update downtime, memory use and state lifecycle behavior. Experimental heads and policies are opt-in; performance evaluation must include matched quality measurements.
Feedback Period.
At least one week after publication of this draft. Keep the issue open as a living implementation and experiment tracker.
CC List.
@bjf-frz @hsliuustc0106
Anything else.
Related work: #69 for model/backend/runtime support; #56 for adaptive-exit evaluation and execution semantics; #43 for speculative inference.
Training frameworks: verl and vime.
Before submitting a new issue...
Motivation.
Enable training-driven improvements to vllm-rlt's looped-transformer inference efficiency. Optimize throughput, latency, and memory use under an explicit quality constraint.
Looped transformers provide specific opportunities: predict exits ahead of execution, align shallow draft predictions with deeper target computation, and make better use of recurrent state and shared weights. Exploring these ideas requires reliable rollout and weight-update interfaces before integrating a training framework.
The proposed order is:
Proposed Change.
Scope, ownership, and progress
Initial integration target: Ouro-1.4B, BF16, synchronous training iterations, and full weight updates. Training interfaces and framework adapters must preserve existing inference features, including early exit, PD, speculative decoding, and CUDA Graphs, from the start. Do not impose a fixed-depth, single-device, non-speculative restriction on the integration. Raise concrete compatibility issues for discussion; this requirement does not imply that currently unsupported feature combinations are already implemented. Asynchronous training, incremental weight transfer, and colocated training/inference memory switching remain separate extensions.
Start with verl + FSDP as the proposed first integration; add vime with its training stack through a second adapter. Pin framework versions and validate model compatibility in the integration itself. Train/inference consistency and gradient checks are required integration tests.
Contributors may split tasks or collaborate. Record Assignee(s), engine PRs, and fork branch/results links. Phase 1 is a joint development scope for @MinhaoLi0318 and @IDEA-V, coordinated with @0z5a and @Levius-Fubuki on the shared interfaces. @Levius-Fubuki has also expressed follow-on interest in N+2 exit work; that experiment is not yet marked as started. @CXJorz is onboarding and has no implementation task assigned yet.
Implementation references
Reference vLLM's rollout and weight-synchronization implementation, and adapt verl/vime's existing vLLM integrations. Preserve recurrent shared-weight and depth-aware KV semantics, documenting necessary differences and the referenced versions.
Implementation details, experiment procedures, training methods and PR boundaries are left to contributors. The requirements below define the deliverables and acceptance criteria.
Phase 1: training-facing capabilities in vllm-rlt
Provide a shared engine contract based on the relevant vLLM interfaces, usable by both frameworks without adding verl/vime dependencies to normal inference installation. Prefer compatible field and operation names; document required deviations in the implementation PRs.
Acceptance requires correct token/probability alignment, concurrent request handling, and consistent generation after weight updates. Document probability conventions, preserve recurrent weight sharing, and prevent reuse of incompatible cached state. Keep the interfaces framework-independent and include relevant correctness and overhead measurements.
The agreed Phase 1 delivery is split into three PRs: rollout outputs; full-weight update and synchronization; and concurrent token-in/token-out API with abort/pause. Build on vLLM and @0z5a’s existing prototype, with shared development and appropriate co-author credit. Existing inference-feature compatibility is part of this delivery, not a later phase.
Phase 2: integration on verl/vime forks
After the engine contract is usable, adapt each framework’s existing vLLM integration to vllm-rlt and run end-to-end experiments on contributor forks. Upstream PRs to verl or vime are not required to complete this phase. Missing generic engine capabilities should be proposed back to vllm-rlt rather than maintained as two incompatible private interfaces.
Provide a reproducible end-to-end training run that verifies correct gradients, weight updates, synchronization, and generation with the updated model. Include train/inference probability consistency and checkpoint recovery evidence. Link the runnable fork branch, configuration and results.
Phase 3: loop-specific training experiments for inference efficiency
Use the integrated training environment to optimize recurrent execution.
Experiment goals and acceptance
Training architecture, labels, losses, datasets and experimental procedures are chosen by contributors and documented with results. Add reusable trace collection, trained-component loading or other engine support as required by the experiment. Adaptive training/rollout must preserve consistent trajectory and KV semantics.
Code location and PR scope
Negative or inconclusive experiments should retain results without requiring a production runtime merge. Keep framework-specific dependencies outside the normal inference installation. Update README/relevant docs for new public engine capabilities and supply reproducible integration/experiment instructions.
Alternatives and impact
Use a shared engine contract for both framework adapters to avoid duplicating rollout and weight-update protocols. Keep framework dependencies in the fork integrations.
Implementation PRs must document API changes, probability-collection overhead, update downtime, memory use and state lifecycle behavior. Experimental heads and policies are opt-in; performance evaluation must include matched quality measurements.
Feedback Period.
At least one week after publication of this draft. Keep the issue open as a living implementation and experiment tracker.
CC List.
@bjf-frz @hsliuustc0106
Anything else.
Related work: #69 for model/backend/runtime support; #56 for adaptive-exit evaluation and execution semantics; #43 for speculative inference.
Training frameworks: verl and vime.
Before submitting a new issue...