Skip to content

[RFC]: Training support and verl/vime integration for looped-transformer inference optimization #70

Description

@bjf-frz

Motivation.

Enable training-driven improvements to vllm-rlt's looped-transformer inference efficiency. Optimize throughput, latency, and memory use under an explicit quality constraint.

Looped transformers provide specific opportunities: predict exits ahead of execution, align shallow draft predictions with deeper target computation, and make better use of recurrent state and shared weights. Exploring these ideas requires reliable rollout and weight-update interfaces before integrating a training framework.

The proposed order is:

  1. Add framework-independent training support to vllm-rlt through repository PRs.
  2. Integrate and run training loops on contributor forks of verl/vime.
  3. Use those forks to train and evaluate loop-specific inference optimizations; upstream useful engine capabilities with evidence.

Proposed Change.

Scope, ownership, and progress

Initial integration target: Ouro-1.4B, BF16, synchronous training iterations, and full weight updates. Training interfaces and framework adapters must preserve existing inference features, including early exit, PD, speculative decoding, and CUDA Graphs, from the start. Do not impose a fixed-depth, single-device, non-speculative restriction on the integration. Raise concrete compatibility issues for discussion; this requirement does not imply that currently unsupported feature combinations are already implemented. Asynchronous training, incremental weight transfer, and colocated training/inference memory switching remain separate extensions.

Start with verl + FSDP as the proposed first integration; add vime with its training stack through a second adapter. Pin framework versions and validate model compatibility in the integration itself. Train/inference consistency and gradient checks are required integration tests.

  • ✅ Completed for the stated scope.
  • 🚧 Claimed or in progress, with a linked claim, PR, or fork branch; not yet accepted as complete.
  • ⬜ Not yet delivered in this tracker; existing components should be reused where applicable.

Contributors may split tasks or collaborate. Record Assignee(s), engine PRs, and fork branch/results links. Phase 1 is a joint development scope for @MinhaoLi0318 and @IDEA-V, coordinated with @0z5a and @Levius-Fubuki on the shared interfaces. @Levius-Fubuki has also expressed follow-on interest in N+2 exit work; that experiment is not yet marked as started. @CXJorz is onboarding and has no implementation task assigned yet.

Phase Deliverable Code destination Status Assignee(s) PR / branch / evidence
1 Training-facing engine capabilities vllm-rlt PRs 🚧 @MinhaoLi0318, @IDEA-V Plan, co-author coordination
2A Ouro training integration with verl/FSDP Contributor verl fork 🚧 @Levius-Fubuki Claim
2B Equivalent integration with vime Contributor vime fork 🚧 @0z5a vime prototype, engine prototype
3 Loop-specific exit and speculative training experiments Framework forks; experimental engine branches as needed 🚧 @IDEA-V (speculation); @Levius-Fubuki (exit, follow-on interest) Claim, Follow-on interest

Implementation references

Reference vLLM's rollout and weight-synchronization implementation, and adapt verl/vime's existing vLLM integrations. Preserve recurrent shared-weight and depth-aware KV semantics, documenting necessary differences and the referenced versions.

Implementation details, experiment procedures, training methods and PR boundaries are left to contributors. The requirements below define the deliverables and acceptance criteria.

Phase 1: training-facing capabilities in vllm-rlt

Provide a shared engine contract based on the relevant vLLM interfaces, usable by both frameworks without adding verl/vime dependencies to normal inference installation. Prefer compatible field and operation names; document required deviations in the implementation PRs.

Task Scope and acceptance Status Assignee(s) PR / plan
Training rollout interface Token-ID input/output, selected-token logprobs, effective sampling configuration, concurrent requests and request lifecycle. Preserve token alignment and termination reasons; support cancellation/error cleanup; document model-versus-sampling probability conventions. 🚧 @MinhaoLi0318, @IDEA-V Plan, co-author coordination
Weight update and synchronization Support reliable full-weight updates with model-version tracking and correct cache/state handling. Generation must use a consistent weight version, and update failures must be reported safely. 🚧 @MinhaoLi0318, @IDEA-V Plan, co-author coordination

Acceptance requires correct token/probability alignment, concurrent request handling, and consistent generation after weight updates. Document probability conventions, preserve recurrent weight sharing, and prevent reuse of incompatible cached state. Keep the interfaces framework-independent and include relevant correctness and overhead measurements.

The agreed Phase 1 delivery is split into three PRs: rollout outputs; full-weight update and synchronization; and concurrent token-in/token-out API with abort/pause. Build on vLLM and @0z5a’s existing prototype, with shared development and appropriate co-author credit. Existing inference-feature compatibility is part of this delivery, not a later phase.

Phase 2: integration on verl/vime forks

After the engine contract is usable, adapt each framework’s existing vLLM integration to vllm-rlt and run end-to-end experiments on contributor forks. Upstream PRs to verl or vime are not required to complete this phase. Missing generic engine capabilities should be proposed back to vllm-rlt rather than maintained as two incompatible private interfaces.

Task Scope and acceptance Status Assignee(s) Branch / evidence
verl/FSDP integration Adapt Ouro training and the vllm-rlt rollout backend, map/export shared recurrent weights, coordinate updates, and provide a reproducible multi-iteration training recipe with consistency and save/resume checks. 🚧 @Levius-Fubuki Claim
vime integration Adapt Ouro to the training stack and reuse the engine rollout/update contract; provide parameter conversion and a reproducible multi-iteration recipe with consistency and save/resume checks. 🚧 @0z5a vime prototype, engine prototype

Provide a reproducible end-to-end training run that verifies correct gradients, weight updates, synchronization, and generation with the updated model. Include train/inference probability consistency and checkpoint recovery evidence. Link the runnable fork branch, configuration and results.

Phase 3: loop-specific training experiments for inference efficiency

Use the integrated training environment to optimize recurrent execution.

Task Scope and acceptance Status Assignee(s) Branch / results
Lookahead exit optimization Train and deploy the N+2 exit predictor, including data/labels, checkpoint loading and runtime decisions. Evaluate task quality, selected versus executed depth, and low/high-concurrency throughput and latency. ⬜ @Levius-Fubuki (follow-on interest) Follow-on interest
Loop-aware speculative optimization Train shallow draft heads/adapters or depth-specific fine-tuning, integrate the resulting draft path, and evaluate quality, acceptance lengths, computation reuse, memory and end-to-end speedup. 🚧 @IDEA-V Claim

Experiment goals and acceptance

  • N+2 exit: use information available at round N to predict whether the token can exit at round N+2; execute through N+2 and use that round's hidden state for output. Preserve valid exit/KV semantics and evaluate quality, selected versus actual executed depth, and throughput/latency against relevant existing baselines.
  • Loop-aware speculation: improve shallow drafting while preserving valid target verification and identifying reusable recurrent computation. Evaluate output quality, acceptance lengths, extra computation/memory and end-to-end performance against target-only and current speculative inference.

Training architecture, labels, losses, datasets and experimental procedures are chosen by contributors and documented with results. Add reusable trace collection, trained-component loading or other engine support as required by the experiment. Adaptive training/rollout must preserve consistent trajectory and KV semantics.

Code location and PR scope

Code / artifact Destination
Generic rollout, probability, versioning and update interfaces vllm-rlt PRs
Framework adapters, differentiable model adaptation, weight conversion and training configuration Contributor verl/vime forks initially
Labels, losses, training scripts and experimental results Framework forks / linked experiment artifacts
Experimental exit/draft execution changes Linked engine experiment branches, paired with framework forks
Useful trace collection or trained-component loading vllm-rlt PRs after validating the shared requirement
New runtime exit/speculation policy vllm-rlt PR with correctness, quality and performance evidence

Negative or inconclusive experiments should retain results without requiring a production runtime merge. Keep framework-specific dependencies outside the normal inference installation. Update README/relevant docs for new public engine capabilities and supply reproducible integration/experiment instructions.

Alternatives and impact

Use a shared engine contract for both framework adapters to avoid duplicating rollout and weight-update protocols. Keep framework dependencies in the fork integrations.

Implementation PRs must document API changes, probability-collection overhead, update downtime, memory use and state lifecycle behavior. Experimental heads and policies are opt-in; performance evaluation must include matched quality measurements.

Feedback Period.

At least one week after publication of this draft. Keep the issue open as a living implementation and experiment tracker.

CC List.

@bjf-frz @hsliuustc0106

Anything else.

Related work: #69 for model/backend/runtime support; #56 for adaptive-exit evaluation and execution semantics; #43 for speculative inference.

Training frameworks: verl and vime.

Before submitting a new issue...

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    RFCRequests for comments on major architectural changes or design choices

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions