HEX is a whole-body vision-language-action framework for full-sized humanoid robots.
-
Updated
May 26, 2026 - Jupyter Notebook
HEX is a whole-body vision-language-action framework for full-sized humanoid robots.
Towards Versatile Vision-Language-Action Models via Learning Unified Vision-Motion Representations
ACOT-VLA-WM is an advanced Vision-Language-Action robotics framework that uses a predictive world model to generate visual subgoals, improving long-horizon robotic manipulation accuracy and robustness.
Visualize episode embeddings and select maximally diverse training subsets for robotics ML. Train on 10K diverse episodes instead of 50K random ones.
A full-stack Embodied AI simulation suite powered by Genesis World: featuring OpenVLA closed-loop evaluation, procedural scene generation, massively parallel GPU RL (Unitree Go2), and multi-physics coupling (PBD/SPH).
Open-source authoritative technical guides: Cloud Computing, AI Infrastructure & Embodied Intelligence — 云计算、AI Infra、具身智能开源权威指南
Imitation Learning for Surgical Robot Task Automation — Behavioral Cloning, DAgger, Diffusion Policy, and VLA models on JIGSAWS surgical demonstrations
Safety layer for Vision-Language-Action models. Physics guard that catches what VLA misses.
Cross-embodiment visual representation learning using Vision Transformers conditioned on robot kinematic structure via cross-attention.
To associate your repository with the vla-model topic, visit your repo's landing page and select "manage topics."