| 2026 |
LeRobot: An Open-Source Library for End-to-End Robot Learning |
Datasets & Benchmarks |
ICLR |
| 2026 |
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments |
End-to-End VLA |
arXiv |
| 2026 |
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots |
End-to-End VLA |
arXiv |
| 2026 |
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models |
End-to-End VLA |
arXiv |
| 2026 |
Membership Inference Attacks on Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2026 |
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics |
Datasets & Benchmarks |
arXiv |
| 2026 |
Learning Diffusion Policy from Primitive Skills for Robot Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation |
End-to-End VLA |
ICLR |
| 2026 |
A Survey of Language-Conditioned Robot Manipulation |
End-to-End VLA |
IJRR |
| 2026 |
What Breaks Embodied AI Security: LLM Vulnerabilities, CPS Flaws, or Something Else? |
High-Level Planning |
arXiv |
| 2025 |
VLAS: VLA Model With Speech Instructions |
Multimodal Ecology |
ICLR |
| 2025 |
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control |
World Model & Video Policy |
arXiv |
| 2025 |
FAST: Efficient Action Tokenization for VLA |
Diffusion Policy |
RSS |
| 2025 |
pi_0.5: VLA with Open-World Generalization |
Diffusion Policy |
arXiv |
| 2025 |
SmolVLA |
Imitation Learning |
arXiv |
| 2025 |
Tactile-VLA |
Multimodal Ecology |
CoRL |
| 2025 |
TLA: Tactile-Language-Action |
Multimodal Ecology |
ICRA |
| 2025 |
DexVLA |
End-to-End VLA |
arXiv |
| 2025 |
OpenHelix |
End-to-End VLA |
arXiv |
| 2025 |
OpenVLA-OFT |
End-to-End VLA |
RSS |
| 2025 |
SpatialVLA |
End-to-End VLA |
arXiv |
| 2025 |
Universal Actions for Enhanced Embodied Foundation Models |
End-to-End VLA |
arXiv |
| 2025 |
LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks |
High-Level Planning |
arXiv |
| 2025 |
EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control |
End-to-End VLA |
arXiv |
| 2025 |
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI |
Datasets & Benchmarks |
arXiv |
| 2025 |
LLaDA-VLA: Vision Language Diffusion Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning |
End-to-End VLA |
arXiv |
| 2025 |
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model |
End-to-End VLA |
arXiv |
| 2025 |
Embodiment Transfer Learning for Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation |
End-to-End VLA |
arXiv |
| 2025 |
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
A Survey on Efficient Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Survey of Vision-Language-Action Models for Embodied Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review |
End-to-End VLA |
arXiv |
| 2025 |
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead |
End-to-End VLA |
arXiv |
| 2025 |
RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI |
End-to-End VLA |
arXiv |
| 2025 |
Embodied Navigation Foundation Model |
End-to-End VLA |
arXiv |
| 2025 |
MiMo-Embodied: X-Embodied Foundation Model Technical Report |
End-to-End VLA |
arXiv |
| 2025 |
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2024 |
OpenVLA: An Open-Source Vision-Language-Action Model |
End-to-End VLA |
CoRL |
| 2024 |
MLA: Multisensory Language-Action Model |
Multimodal Ecology |
arXiv |
| 2024 |
SimplerEnv |
Simulation & Sim2Real |
NeurIPS |
| 2024 |
RoboFlamingo |
High-Level Planning |
ICLR |
| 2024 |
3D-VLA |
End-to-End VLA |
ICML |
| 2024 |
GR-2: Generative Video-Language-Action Model |
End-to-End VLA |
arXiv |
| 2024 |
RoboMamba |
End-to-End VLA |
NeurIPS |
| 2024 |
TinyVLA |
End-to-End VLA |
RA-L |
| 2024 |
TraceVLA: Visual Trace Prompting |
End-to-End VLA |
ICLR |
| 2024 |
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2023 |
LIBERO |
Datasets & Benchmarks |
NeurIPS |
| 2023 |
Open X-Embodiment |
Datasets & Benchmarks |
ICRA |
| 2023 |
VoxPoser |
High-Level Planning |
CoRL |
| 2023 |
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control |
End-to-End VLA |
CoRL |
| 2022 |
RT-1: Robotics Transformer for Real-World Control at Scale |
End-to-End VLA |
RSS |