Heng Yu*, David D. Yuan*, Juze Zhang*, Changan Chen, Yao Feng,
Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli
Stanford University
* Equal contribution
Paper (PDF) · Research blog · Technical documentation · Quickstart · Methods · Training and evaluation · Extension SDK · Citation
OpenWAM is the official implementation of our paper and a reusable framework for video-action world models in robot learning. It supports video pretraining, robot policies, and forward and inverse dynamics models within a shared training and evaluation stack.
The research blog introduces the framework, demonstrations, and experimental results. The technical documentation covers installation, training, evaluation, and extension APIs.
Our paper, OpenWAM: An Open Framework for Composable World-Action Models, presents the framework and experimental results.
Trained policy checkpoints and datasets are being prepared for public release. Canonical links and integrity metadata will be added to the artifact documentation as each resource is published.
Train video models, fine-tune robot policies, and evaluate them through shared data and simulator interfaces. The video pretraining guide covers data preparation through prediction. Released pretraining weights are available at OpenWAM-Stanford/OpenWAM-Pretraining on Hugging Face.
OpenWAM separates model topology, video/action conditioning, sequence semantics, visual execution, and action decoding so that controlled experiments share the same trainer and visual stack.
OpenWAM provides:
- one typed train, resume, evaluation, and simulator runtime across policy architectures;
- six standard video/action programs plus generalist joint denoising (GJD);
- full-state checkpoint continuation and versioned run provenance;
- adapters for LIBERO, RoboTwin, CALVIN, heterogeneous LeRobot data, and synthetic fixtures;
- extension APIs for datasets, policies, decoders, attention profiles, and simulators.
| Architecture | Topology | Maintained programs |
|---|---|---|
parallel_stream |
Video and action tokens share one transformer. | Six standard programs, GJD, and standalone conditional FDM/IDM. |
dual_expert |
Video and action use separate transformer experts. | Six standard programs, GJD, and standalone conditional FDM/IDM. |
causal_video_prediction |
The visual model runs without action supervision. | Video-only prediction. |
Programs control how video and action condition one another. GJD combines joint prediction, forward dynamics (FDM), and inverse dynamics (IDM) in one model. See Policy Architectures and Programs for the complete program list and conditioning rules.
Version 0.2.0: OpenWAM remains alpha-stage Linux research software; see the 0.2 migration guide before upgrading. The public CPU lifecycle and synthetic artifacts are self-contained. Large benchmark runs use separately provisioned datasets and checkpoints described by the artifact contract.
OpenWAM is available on PyPI for Linux with Python 3.11 or 3.12. In a virtual environment:
python -m pip install openwamThe base package provides configuration and metadata APIs. For model training and evaluation, install the runtime extras:
python -m pip install 'openwam[train,eval]'Use openwam[train,pretrain] for video data preparation and pretraining, or
openwam[sim] for model-driven simulator rollouts. Python imports use
open_wam. The optional openwam-sdk installation alias provides the same
implementation and extras.
The quickstart covers installed-package usage. For development or exact dependency reproduction, use the frozen source checkout below. Model weights, datasets, and external simulator source trees are provisioned separately.
OpenWAM supports Linux with Python 3.11 or 3.12. Install
uv, then run the public CPU contract:
git clone https://github.com/OpenWAM/OpenWAM.git
cd OpenWAM
uv sync --frozen --group dev --extra train --extra eval
uv run openwam-validate-config \
configs/examples/public_tiny_synthetic_contract.yaml \
configs/evals/public_tiny_synthetic_contract.yaml
uv run --extra train openwam-sanity \
--cfg configs/examples/public_tiny_synthetic_contract.yaml \
--device cpu --max-batches 1 --rollout-steps 1This path requires no private data, checkpoint, GPU, or external simulator. It checks config loading, dataset construction, a train forward pass, batch inference, and recurrent rollout-style inference. The complete CPU first run adds stateful resume and checkpoint-backed evaluation.
Install only the runtime needed for later work:
| Task | Command |
|---|---|
| Config and metadata development | uv sync --group dev |
| Training | uv sync --extra train |
| Offline evaluation | uv sync --extra eval |
| Model-driven simulator rollout | uv sync --extra sim |
| Documentation | uv sync --extra docs |
Benchmark extras supply dependency overlays, not upstream source trees. Follow Benchmarks and Data before a real simulator run.
Real datasets, checkpoints, simulator checkouts, and output directories remain outside versioned experiment YAML. Start with the local path registry:
cp configs/local_paths.sample.yaml configs/local_paths.yaml
uv run openwam-inspect-config \
--cfg configs/experiments/dual_expert_libero_joint.yamlPopulate only the aliases used by the selected config. The local registry is
gitignored; set OPEN_WAM_LOCAL_PATHS=/absolute/path/paths.yaml to keep it
elsewhere.
All architectures use openwam-train. The shipped Parallel Stream and Dual
Expert LIBERO policy programs use the same validated full-trajectory W64 recipe
described in Training and Inference.
The reference 30-layer configs are FSDP workloads characterized with four 48 GB GPUs:
uv run --extra train torchrun --standalone --nproc-per-node=4 \
-m open_wam.cli.train \
--cfg configs/experiments/dual_expert_libero_joint.yaml \
--save-root runs/dual-expert-joint \
--expected-world-size 4For a one-off method ablation, change the public program selector rather than the trainer:
uv run --extra train torchrun --standalone --nproc-per-node=4 \
-m open_wam.cli.train \
--cfg configs/experiments/dual_expert_libero_joint.yaml \
--set policy_variant.program=video_then_action \
--save-root runs/dual-expert-vta \
--expected-world-size 4Resume from a full training-state checkpoint with the same command and
--resume-from:
uv run --extra train torchrun --standalone --nproc-per-node=4 \
-m open_wam.cli.train \
--cfg configs/experiments/dual_expert_libero_joint.yaml \
--save-root runs/dual-expert-joint \
--resume-from runs/dual-expert-joint/checkpoints/checkpoint_step_N \
--expected-world-size 4--resume-from requires full_training_state.pt and restores training state.
Use --initialize-weights-from instead for a fresh run from model weights.
Resume does not guarantee bitwise replay of stochastic data loading; see
initialization and resume
for requirements and limits.
Conditional FDM/IDM uses the dynamics-routing data adapter. The maintained
config mixes real demonstrations with encoded counterfactual train and
validation roots; a real-demo-only ablation is also supported. Read the
data prerequisites before
selecting forward_dynamics or inverse_dynamics.
Run offline metrics through the generic evaluator:
uv run --extra eval openwam-eval \
--cfg configs/evals/dual_expert_robotwin_smoke_eval.yaml \
--checkpoint /path/to/model_state.pt \
--device cuda:0Run a configured environment through the simulator boundary:
uv run --extra sim openwam-sim-rollout \
--cfg configs/experiments/parallel_stream_robotwin_smoke.yaml \
--checkpoint /path/to/model_state.pt \
--benchmark robotwin \
--robotwin-task-name <task-name>Benchmark adapters translate observations and actions. Sequence, attention, cache, and denoising semantics remain owned by the selected policy. Maintained LIBERO and GJD commands are listed in Training and Inference.
Every built-in method follows one composition boundary:
ExperimentConfig -> VariantPipeline -> VisualTower -> PolicyVariant -> ActionDecoder
| Contract | Responsibility |
|---|---|
ExperimentConfig |
Typed architecture, program, data, sequence, runtime, and optimization choices. |
VariantPipeline |
Shared training and inference orchestration. |
VisualTower |
Frontend encoding, visual backbone execution, decode stages, and runtime hooks. |
PolicyVariant |
Parameter topology, architecture-specific packing, conditioning adapters, and recurrent state. |
ActionDecoder |
Final supervised outputs, masks, losses, metrics, and committed actions. |
This boundary keeps the visual stack stable while experiments vary one owned contract at a time. See Architecture and Policy Architectures and Programs.
Out-of-tree packages load through repeatable --extension module[:hook]
arguments. Choose the smallest owning boundary:
| Customization | Extension surface |
|---|---|
| Storage format, camera schema, or action/state representation | Dataset adapter selected by data.dataset_type |
| Learned parameters, conditioning, attention profile, or recurrent state | PolicyVariant |
| Final outputs, loss, sampling, or committed action count | ActionDecoder |
| Environment construction and observation/action translation | Simulator adapter |
| Existing method, geometry, schedule, cache, or optimizer choice | YAML only |
The packaged extension scaffold verifies registration, gradients, inference state, and packaging before custom code is introduced:
uv run --extra train openwam-train \
--cfg templates/extension_method/config.yaml \
--extension open_wam.templates.extension_method \
--save-root runs/extension-method-smoke \
--disable-wandbExtensions import compatibility-managed contracts from the role-specific
open_wam.sdk modules. See the Extension SDK and
cookbooks.
Evaluation, sanity, and simulator commands can emit the same versioned result envelope with source state, exact argv, config hashes, checkpoint identity, dataset metadata, package versions, and device details. Use full provenance to hash a publication checkpoint:
openwam-eval --cfg evaluation.yaml --output-json result.json \
--provenance-mode fullFor numerical comparisons, keep the dependency lock, hardware/software stack, checkpoint, input data, and evaluation settings fixed. See Reproducibility, Compatibility, and Testing.
configs/ typed experiments, evaluations, examples, and path templates
docs/ public guides, experiment cards, and extension cookbooks
scripts/ thin benchmark adapters and checkout-only research tools
src/open_wam/ installable library and role-scoped SDK
tests/ unit, integration, simulator, and numerical parity gates
notes/index/ generated public consortium metadata packaged at runtime
| Topic | Guide |
|---|---|
| Install and first run | Quickstart |
| Runtime ownership | Architecture |
| Architectures and programs | Policy Architectures |
| Train, resume, evaluate, and roll out | Training and Inference |
| Dataset and simulator setup | Benchmarks and Data |
| Custom datasets, policies, decoders, and simulators | Extension SDK |
| Checkpoints and manifests | Artifacts |
| Test and parity tiers | Testing |
Contributions should preserve the typed runtime boundary and add focused tests for every changed contract. Read CONTRIBUTING.md, the Code of Conduct, and the Security Policy before opening a pull request.
If OpenWAM supports your research, please cite our paper:
@article{yu2026openwam,
title = {{OpenWAM}: An Open Framework for Composable World-Action Models},
author = {Yu, Heng and Yuan, David D. and Zhang, Juze and Chen, Changan and
Feng, Yao and Baldonado, Michelle and Cousins, Steve and
Fei-Fei, Li and Wu, Jiajun and Adeli, Ehsan},
journal = {arXiv preprint arXiv:2610.07922},
year = {2026},
url = {https://arxiv.org/pdf/2610.07922}
}The BibTeX entry is also available in CITATION.bib.
OpenWAM is released under the GNU Affero General Public License v3.0
with the redistribution attribution described in NOTICE. Covered
modified versions and network services must provide corresponding source, and
redistributed copies must preserve the OpenWAM attribution notice. Academic
work that uses OpenWAM should cite the paper; see Citation.
Third-party components retain their own terms; the adapted LingBot-VA module
is distributed under Apache License 2.0. Full attributions are listed in
THIRD_PARTY_NOTICES.md and LICENSES/.
Stanford, SAIL, SVL, STAI, and SRC marks are not licensed under AGPL-3.0-only and remain
the property of Stanford University.


