| 2026 |
Membership Inference Attacks on Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2026 |
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics |
Datasets & Benchmarks |
arXiv |
| 2026 |
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2025 |
Dreamer V3: Mastering Diverse Domains through World Models |
World Model & Video Policy |
Nature |
| 2025 |
Universal Actions for Enhanced Embodied Foundation Models |
End-to-End VLA |
arXiv |
| 2025 |
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI |
Datasets & Benchmarks |
arXiv |
| 2025 |
LLaDA-VLA: Vision Language Diffusion Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning |
End-to-End VLA |
arXiv |
| 2025 |
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation |
End-to-End VLA |
arXiv |
| 2025 |
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review |
End-to-End VLA |
arXiv |
| 2025 |
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead |
End-to-End VLA |
arXiv |
| 2025 |
RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI |
End-to-End VLA |
arXiv |
| 2025 |
MiMo-Embodied: X-Embodied Foundation Model Technical Report |
End-to-End VLA |
arXiv |
| 2025 |
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy |
Datasets & Benchmarks |
ICRA |
| 2024 |
DROID |
Datasets & Benchmarks |
RSS |
| 2024 |
ALOHA 2 |
Imitation Learning |
Tech Report |
| 2024 |
GenSim |
High-Level Planning |
ICLR |
| 2024 |
BEHAVIOR-1K |
Simulation & Sim2Real |
CoRL |
| 2024 |
Habitat 3.0 |
Simulation & Sim2Real |
ICLR |
| 2024 |
TraceVLA: Visual Trace Prompting |
End-to-End VLA |
ICLR |
| 2024 |
DeepSeek-VL: Towards Real-World Vision-Language Understanding |
VLM Foundation |
arXiv |
| 2024 |
Improved Baselines with Visual Instruction Tuning |
VLM Foundation |
CVPR |
| 2024 |
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling |
VLM Foundation |
arXiv |
| 2024 |
LLaVA-NeXT-Interleave |
VLM Foundation |
arXiv |
| 2024 |
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents |
Datasets & Benchmarks |
arXiv |
| 2024 |
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding |
Multimodal Ecology |
arXiv |
| 2024 |
The Essential Role of Causality in Foundation World Models for Embodied AI |
World Model & Video Policy |
ICML |
| 2024 |
A call for embodied AI |
VLM Foundation |
ICML |
| 2023 |
LIBERO |
Datasets & Benchmarks |
NeurIPS |
| 2023 |
BridgeData V2 |
Datasets & Benchmarks |
CoRL |
| 2023 |
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches |
End-to-End VLA |
ICLR |
| 2022 |
CALVIN |
Datasets & Benchmarks |
RA-L |
| 2021 |
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation |
Datasets & Benchmarks |
CoRL |
| 2021 |
Habitat 2.0 |
Simulation & Sim2Real |
NeurIPS |
| 2021 |
ManiSkill |
Simulation & Sim2Real |
NeurIPS |
| 2020 |
Dual-path RNN |
Auditory & Acoustic |
ICASSP |
| 2020 |
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning |
Datasets & Benchmarks |
arXiv |
| 2019 |
Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning |
Datasets & Benchmarks |
CoRL |
| 2019 |
RLBench: The Robot Learning Benchmark & Learning Environment |
Datasets & Benchmarks |
RA-L |