I'm a Computer Science researcher/engineer focused on Multimodal AI, Vision-Language-Action (VLA) Models, and Embodied AI.
My work explores how multimodal foundation models can reason, act, and generalize in complex environments. Recently, I've been working on improving Chain-of-Thought reasoning in VLAs through temporal action-reasoning alignment.
- Multimodal Foundation Models
- Vision-Language-Action Models
- Reasoning & Agents
- Embodied AI
- Reinforcement Learning
- Generative AI
Python · C++ · PyTorch · HuggingFace · LeRobot · MuJoCo · vLLM · Docker · Linux · Slurm
Making Chain-of-Thought Matter in Vision-Language-Action Models Developed TARA, a method for aligning reasoning with action chunking in VLAs, improving steerability and generalization.
ECoT-Lite Reproduction of embodied Chain-of-Thought reasoning for Vision-Language-Action models.
Moment Retrieval using Video and Audio Extended BLIP-2 with audio-vision-language interleaving for video moment retrieval.
Robotic Athletes Trained a simulated Barrett WAM robot arm to play badminton using MuJoCo and reinforcement learning.
