| 2026 |
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments |
End-to-End VLA |
arXiv |
| 2026 |
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics |
Datasets & Benchmarks |
arXiv |
| 2026 |
Learning Diffusion Policy from Primitive Skills for Robot Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation |
End-to-End VLA |
ICLR |
| 2026 |
A Survey of Language-Conditioned Robot Manipulation |
End-to-End VLA |
IJRR |
| 2026 |
What Breaks Embodied AI Security: LLM Vulnerabilities, CPS Flaws, or Something Else? |
High-Level Planning |
arXiv |
| 2025 |
Generalizable Humanoid Manipulation with 3D Diffusion Policies (iDP3) |
Imitation Learning |
RSS |
| 2025 |
MuJoCo Playground |
Simulation & Sim2Real |
arXiv |
| 2025 |
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks |
High-Level Planning |
arXiv |
| 2025 |
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI |
Datasets & Benchmarks |
arXiv |
| 2025 |
LLaDA-VLA: Vision Language Diffusion Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model |
End-to-End VLA |
arXiv |
| 2025 |
Embodiment Transfer Learning for Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
A Survey on Efficient Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Survey of Vision-Language-Action Models for Embodied Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
MiMo-Embodied: X-Embodied Foundation Model Technical Report |
End-to-End VLA |
arXiv |
| 2025 |
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2025 |
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation |
Diffusion Policy |
ICRA |
| 2025 |
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy |
Datasets & Benchmarks |
ICRA |
| 2024 |
DROID |
Datasets & Benchmarks |
RSS |
| 2024 |
SimplerEnv |
Simulation & Sim2Real |
NeurIPS |
| 2024 |
Affordance-based Robot Manipulation with Flow Matching |
Diffusion Policy |
IROS |
| 2024 |
π₀: A Vision-Language-Action Flow Model for General Robot Control |
Diffusion Policy |
arXiv |
| 2024 |
ALOHA 2 |
Imitation Learning |
Tech Report |
| 2024 |
Universal Manipulation Interface |
Imitation Learning |
RSS |
| 2024 |
GenSim |
High-Level Planning |
ICLR |
| 2024 |
Habitat 3.0 |
Simulation & Sim2Real |
ICLR |
| 2024 |
RDT-1B: Diffusion Foundation Model for Bimanual Manipulation |
End-to-End VLA |
ICLR |
| 2024 |
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2024 |
DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting |
Diffusion Policy |
arXiv |
| 2023 |
LIBERO |
Datasets & Benchmarks |
NeurIPS |
| 2023 |
RH20T |
Datasets & Benchmarks |
RSS Workshop |
| 2023 |
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT/ALOHA) |
Imitation Learning |
RSS |
| 2023 |
Code as Policies: Language Model Programs for Embodied Control |
High-Level Planning |
ICRA |
| 2023 |
ChatGPT for Robotics |
High-Level Planning |
IEEE Access |
| 2023 |
BridgeData V2 |
Datasets & Benchmarks |
CoRL |
| 2023 |
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis |
World Model & Video Policy |
arXiv |
| 2021 |
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation |
Datasets & Benchmarks |
CoRL |
| 2020 |
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning |
Datasets & Benchmarks |
arXiv |
| 2019 |
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning |
Datasets & Benchmarks |
CoRL |