Skip to content
View OpenWAM's full-sized avatar

Block or report OpenWAM

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
OpenWAM/README.md

OpenWAM: An Open Framework
for Composable World-Action Models

Heng Yu*, David D. Yuan*, Juze Zhang*, Changan Chen, Yao Feng,
Michelle Baldonado, Steve Cousins, Li Fei-Fei, Jiajun Wu, Ehsan Adeli

Stanford University
* Equal contribution

Stanford University     Stanford Artificial Intelligence Laboratory     Stanford Vision and Learning Lab     Stanford Translational AI (STAI) Lab     Stanford Robotics Center


arXiv: 2610.07922 Research blog Technical documentation CI Python 3.11 or 3.12 License: AGPL v3

Paper (PDF) · Research blog · Technical documentation · Quickstart · Methods · Training and evaluation · Extension SDK · Citation

OpenWAM is the official implementation of our paper and a reusable framework for video-action world models in robot learning. It supports video pretraining, robot policies, and forward and inverse dynamics models within a shared training and evaluation stack.

The research blog introduces the framework, demonstrations, and experimental results. The technical documentation covers installation, training, evaluation, and extension APIs.

OpenWAM robot manipulation teaser

Research Paper and Releases

Our paper, OpenWAM: An Open Framework for Composable World-Action Models, presents the framework and experimental results.

Trained policy checkpoints and datasets are being prepared for public release. Canonical links and integrity metadata will be added to the artifact documentation as each resource is published.

Research Scope

Train video models, fine-tune robot policies, and evaluate them through shared data and simulator interfaces. The video pretraining guide covers data preparation through prediction. Released pretraining weights are available at OpenWAM-Stanford/OpenWAM-Pretraining on Hugging Face.

OpenWAM separates model topology, video/action conditioning, sequence semantics, visual execution, and action decoding so that controlled experiments share the same trainer and visual stack.

OpenWAM provides:

  • one typed train, resume, evaluation, and simulator runtime across policy architectures;
  • six standard video/action programs plus generalist joint denoising (GJD);
  • full-state checkpoint continuation and versioned run provenance;
  • adapters for LIBERO, RoboTwin, CALVIN, heterogeneous LeRobot data, and synthetic fixtures;
  • extension APIs for datasets, policies, decoders, attention profiles, and simulators.

Maintained Methods

Architecture Topology Maintained programs
parallel_stream Video and action tokens share one transformer. Six standard programs, GJD, and standalone conditional FDM/IDM.
dual_expert Video and action use separate transformer experts. Six standard programs, GJD, and standalone conditional FDM/IDM.
causal_video_prediction The visual model runs without action supervision. Video-only prediction.

Programs control how video and action condition one another. GJD combines joint prediction, forward dynamics (FDM), and inverse dynamics (IDM) in one model. See Policy Architectures and Programs for the complete program list and conditioning rules.

Installation

Version 0.2.0: OpenWAM remains alpha-stage Linux research software; see the 0.2 migration guide before upgrading. The public CPU lifecycle and synthetic artifacts are self-contained. Large benchmark runs use separately provisioned datasets and checkpoints described by the artifact contract.

OpenWAM is available on PyPI for Linux with Python 3.11 or 3.12. In a virtual environment:

python -m pip install openwam

The base package provides configuration and metadata APIs. For model training and evaluation, install the runtime extras:

python -m pip install 'openwam[train,eval]'

Use openwam[train,pretrain] for video data preparation and pretraining, or openwam[sim] for model-driven simulator rollouts. Python imports use open_wam. The optional openwam-sdk installation alias provides the same implementation and extras.

The quickstart covers installed-package usage. For development or exact dependency reproduction, use the frozen source checkout below. Model weights, datasets, and external simulator source trees are provisioned separately.

Quick Start From Source

OpenWAM supports Linux with Python 3.11 or 3.12. Install uv, then run the public CPU contract:

git clone https://github.com/OpenWAM/OpenWAM.git
cd OpenWAM

uv sync --frozen --group dev --extra train --extra eval

uv run openwam-validate-config \
  configs/examples/public_tiny_synthetic_contract.yaml \
  configs/evals/public_tiny_synthetic_contract.yaml

uv run --extra train openwam-sanity \
  --cfg configs/examples/public_tiny_synthetic_contract.yaml \
  --device cpu --max-batches 1 --rollout-steps 1

This path requires no private data, checkpoint, GPU, or external simulator. It checks config loading, dataset construction, a train forward pass, batch inference, and recurrent rollout-style inference. The complete CPU first run adds stateful resume and checkpoint-backed evaluation.

Install only the runtime needed for later work:

Task Command
Config and metadata development uv sync --group dev
Training uv sync --extra train
Offline evaluation uv sync --extra eval
Model-driven simulator rollout uv sync --extra sim
Documentation uv sync --extra docs

Benchmark extras supply dependency overlays, not upstream source trees. Follow Benchmarks and Data before a real simulator run.

Training

Real datasets, checkpoints, simulator checkouts, and output directories remain outside versioned experiment YAML. Start with the local path registry:

cp configs/local_paths.sample.yaml configs/local_paths.yaml
uv run openwam-inspect-config \
  --cfg configs/experiments/dual_expert_libero_joint.yaml

Populate only the aliases used by the selected config. The local registry is gitignored; set OPEN_WAM_LOCAL_PATHS=/absolute/path/paths.yaml to keep it elsewhere.

All architectures use openwam-train. The shipped Parallel Stream and Dual Expert LIBERO policy programs use the same validated full-trajectory W64 recipe described in Training and Inference. The reference 30-layer configs are FSDP workloads characterized with four 48 GB GPUs:

uv run --extra train torchrun --standalone --nproc-per-node=4 \
  -m open_wam.cli.train \
  --cfg configs/experiments/dual_expert_libero_joint.yaml \
  --save-root runs/dual-expert-joint \
  --expected-world-size 4

For a one-off method ablation, change the public program selector rather than the trainer:

uv run --extra train torchrun --standalone --nproc-per-node=4 \
  -m open_wam.cli.train \
  --cfg configs/experiments/dual_expert_libero_joint.yaml \
  --set policy_variant.program=video_then_action \
  --save-root runs/dual-expert-vta \
  --expected-world-size 4

Resume from a full training-state checkpoint with the same command and --resume-from:

uv run --extra train torchrun --standalone --nproc-per-node=4 \
  -m open_wam.cli.train \
  --cfg configs/experiments/dual_expert_libero_joint.yaml \
  --save-root runs/dual-expert-joint \
  --resume-from runs/dual-expert-joint/checkpoints/checkpoint_step_N \
  --expected-world-size 4

--resume-from requires full_training_state.pt and restores training state. Use --initialize-weights-from instead for a fresh run from model weights. Resume does not guarantee bitwise replay of stochastic data loading; see initialization and resume for requirements and limits.

Conditional FDM/IDM uses the dynamics-routing data adapter. The maintained config mixes real demonstrations with encoded counterfactual train and validation roots; a real-demo-only ablation is also supported. Read the data prerequisites before selecting forward_dynamics or inverse_dynamics.

Evaluation And Rollout

Run offline metrics through the generic evaluator:

uv run --extra eval openwam-eval \
  --cfg configs/evals/dual_expert_robotwin_smoke_eval.yaml \
  --checkpoint /path/to/model_state.pt \
  --device cuda:0

Run a configured environment through the simulator boundary:

uv run --extra sim openwam-sim-rollout \
  --cfg configs/experiments/parallel_stream_robotwin_smoke.yaml \
  --checkpoint /path/to/model_state.pt \
  --benchmark robotwin \
  --robotwin-task-name <task-name>

Benchmark adapters translate observations and actions. Sequence, attention, cache, and denoising semantics remain owned by the selected policy. Maintained LIBERO and GJD commands are listed in Training and Inference.

Architecture

Every built-in method follows one composition boundary:

ExperimentConfig -> VariantPipeline -> VisualTower -> PolicyVariant -> ActionDecoder
Contract Responsibility
ExperimentConfig Typed architecture, program, data, sequence, runtime, and optimization choices.
VariantPipeline Shared training and inference orchestration.
VisualTower Frontend encoding, visual backbone execution, decode stages, and runtime hooks.
PolicyVariant Parameter topology, architecture-specific packing, conditioning adapters, and recurrent state.
ActionDecoder Final supervised outputs, masks, losses, metrics, and committed actions.

This boundary keeps the visual stack stable while experiments vary one owned contract at a time. See Architecture and Policy Architectures and Programs.

Use OpenWAM With Your System

Out-of-tree packages load through repeatable --extension module[:hook] arguments. Choose the smallest owning boundary:

Customization Extension surface
Storage format, camera schema, or action/state representation Dataset adapter selected by data.dataset_type
Learned parameters, conditioning, attention profile, or recurrent state PolicyVariant
Final outputs, loss, sampling, or committed action count ActionDecoder
Environment construction and observation/action translation Simulator adapter
Existing method, geometry, schedule, cache, or optimizer choice YAML only

The packaged extension scaffold verifies registration, gradients, inference state, and packaging before custom code is introduced:

uv run --extra train openwam-train \
  --cfg templates/extension_method/config.yaml \
  --extension open_wam.templates.extension_method \
  --save-root runs/extension-method-smoke \
  --disable-wandb

Extensions import compatibility-managed contracts from the role-specific open_wam.sdk modules. See the Extension SDK and cookbooks.

Reproducibility

Evaluation, sanity, and simulator commands can emit the same versioned result envelope with source state, exact argv, config hashes, checkpoint identity, dataset metadata, package versions, and device details. Use full provenance to hash a publication checkpoint:

openwam-eval --cfg evaluation.yaml --output-json result.json \
  --provenance-mode full

For numerical comparisons, keep the dependency lock, hardware/software stack, checkpoint, input data, and evaluation settings fixed. See Reproducibility, Compatibility, and Testing.

Repository Layout

configs/       typed experiments, evaluations, examples, and path templates
docs/          public guides, experiment cards, and extension cookbooks
scripts/       thin benchmark adapters and checkout-only research tools
src/open_wam/  installable library and role-scoped SDK
tests/         unit, integration, simulator, and numerical parity gates
notes/index/   generated public consortium metadata packaged at runtime

Documentation

Topic Guide
Install and first run Quickstart
Runtime ownership Architecture
Architectures and programs Policy Architectures
Train, resume, evaluate, and roll out Training and Inference
Dataset and simulator setup Benchmarks and Data
Custom datasets, policies, decoders, and simulators Extension SDK
Checkpoints and manifests Artifacts
Test and parity tiers Testing

Contributing

Contributions should preserve the typed runtime boundary and add focused tests for every changed contract. Read CONTRIBUTING.md, the Code of Conduct, and the Security Policy before opening a pull request.

Citation

If OpenWAM supports your research, please cite our paper:

@article{yu2026openwam,
  title   = {{OpenWAM}: An Open Framework for Composable World-Action Models},
  author  = {Yu, Heng and Yuan, David D. and Zhang, Juze and Chen, Changan and
             Feng, Yao and Baldonado, Michelle and Cousins, Steve and
             Fei-Fei, Li and Wu, Jiajun and Adeli, Ehsan},
  journal = {arXiv preprint arXiv:2610.07922},
  year    = {2026},
  url     = {https://arxiv.org/pdf/2610.07922}
}

The BibTeX entry is also available in CITATION.bib.

License

OpenWAM is released under the GNU Affero General Public License v3.0 with the redistribution attribution described in NOTICE. Covered modified versions and network services must provide corresponding source, and redistributed copies must preserve the OpenWAM attribution notice. Academic work that uses OpenWAM should cite the paper; see Citation.

Third-party components retain their own terms; the adapted LingBot-VA module is distributed under Apache License 2.0. Full attributions are listed in THIRD_PARTY_NOTICES.md and LICENSES/. Stanford, SAIL, SVL, STAI, and SRC marks are not licensed under AGPL-3.0-only and remain the property of Stanford University.

Popular repositories Loading

  1. OpenWAM OpenWAM Public

    Official code for "OpenWAM: An Open Framework for Composable World-Action Models" (Stanford University)

    Python 111 6