Reinforcement-learning environments for training XgoDuck, a small biped whose geometry and inertia differ from MicroDuck. Policies are trained with PPO on mjlab (MuJoCo Warp) and exported to ONNX.
This repository is a downstream project based on microduck_rl by Pollen Robotics. Training runs on mjlab. Joint actuation uses BAM (Better Actuator Models) from Rhoban.
Thank you to the authors of those projects.
The task code follows the MicroDuck recipes, but the robot does not:
- The MJCF, meshes, masses, and inertias are the XgoDuck model (
src/mjlab_microduck/robot/xgoduck/), scaled to 0.8 kg. Trunk height and sit height are recomputed for that geometry. - Actuators are the HLS1910 BAM model in
src/mjlab_microduck/robot/xgoduck/params/1910_m6.json, not the MicroDuck XL330.
You need a CUDA GPU, Python 3.12, and uv.
cd microduck_rl
uv syncOn ARM machines (DGX Spark / GB10, Jetson), the first uv sync downloads a large CUDA wheel and uv's default 30 s HTTP timeout can abort it. Set UV_HTTP_TIMEOUT=600 for that first sync.
From the repository root:
uv run train Mjlab-Velocity-Flat-XgoDuck --env.scene.num-envs 4096uv run list-envs prints every registered task. XgoDuck checkpoints and TensorBoard logs go under logs/rsl_rl/<experiment>/.
| Task | Terrain | Experiment directory |
|---|---|---|
Mjlab-Velocity-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_velocity/ |
Mjlab-Velocity-Rough-XgoDuck |
rough | logs/rsl_rl/xgoduck_velocity/ |
Mjlab-StandUp-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_standup/ |
Mjlab-StandUp-Rough-XgoDuck |
rough | logs/rsl_rl/xgoduck_standup/ |
Mjlab-GroundPick-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_ground_pick/ |
Mjlab-GroundPick-Rough-XgoDuck |
rough | logs/rsl_rl/xgoduck_ground_pick/ |
Mjlab-SitStand-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_sitstand/ |
Mjlab-SitStand-Rough-XgoDuck |
rough | logs/rsl_rl/xgoduck_sitstand/ |
Mjlab-BallKick-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_ball_kick_right/ |
Mjlab-BallKick-Left-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_ball_kick_left/ |
Mjlab-Roulade-Flat-XgoDuck |
flat | logs/rsl_rl/xgoduck_roulade/ |
play with no checkpoint flag loads the newest model_*.pt in that task's experiment directory:
uv run play Mjlab-Velocity-Flat-XgoDuckPass --checkpoint-file path/to/model_XXXX.pt to choose a specific checkpoint.
Export bakes observation normalization into the graph. Run it from the repository root and point it at a trained checkpoint:
uv run python scripts/export.py Mjlab-Velocity-Flat-XgoDuck \
--checkpoint-file logs/rsl_rl/xgoduck_velocity/<run>/model_XXXX.pt \
--onnx-file logs/rsl_rl/xgoduck_velocity/<run>/<run>.onnxIf you omit --onnx-file, the file is written to output.onnx in the current directory. *.onnx is gitignored. Keep the export next to the run that produced it, for example logs/rsl_rl/xgoduck_velocity/<run>/<run>.onnx.
The code is licensed under Apache 2.0. See LICENSE.