| 2026 |
LeRobot: An Open-Source Library for End-to-End Robot Learning |
Datasets & Benchmarks |
ICLR |
| 2026 |
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots |
End-to-End VLA |
arXiv |
| 2026 |
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
Membership Inference Attacks on Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2026 |
Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics |
Datasets & Benchmarks |
arXiv |
| 2026 |
3D Generation for Embodied AI and Robotic Simulation: A Survey |
Simulation & Sim2Real |
arXiv |
| 2025 |
Diffusion Policy Policy Optimization (DPPO) |
Diffusion Policy |
ICLR |
| 2025 |
FlowPolicy: 3D Flow-based Policy via Consistency Flow Matching |
Diffusion Policy |
AAAI |
| 2025 |
pi_0.5: VLA with Open-World Generalization |
Diffusion Policy |
arXiv |
| 2025 |
Embodiment Transfer Learning for Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
A Survey on Efficient Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation |
Diffusion Policy |
ICRA |
| 2024 |
RoboCasa |
Datasets & Benchmarks |
RSS |
| 2024 |
SimplerEnv |
Simulation & Sim2Real |
NeurIPS |
| 2024 |
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations |
Diffusion Policy |
RSS |
| 2024 |
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation |
Diffusion Policy |
RSS |
| 2024 |
EquiBot: SIM(3)-Equivariant Diffusion Policy |
Diffusion Policy |
CoRL |
| 2024 |
Affordance-based Robot Manipulation with Flow Matching |
Diffusion Policy |
IROS |
| 2024 |
ALOHA 2 |
Imitation Learning |
Tech Report |
| 2024 |
HumanPlus |
Imitation Learning |
CoRL |
| 2024 |
Mobile ALOHA |
Imitation Learning |
CoRL |
| 2024 |
Universal Manipulation Interface |
Imitation Learning |
RSS |
| 2024 |
Behavior Generation with Latent Actions (VQ-BeT) |
Imitation Learning |
ICML |
| 2024 |
GenSim |
High-Level Planning |
ICLR |
| 2024 |
RoboFlamingo |
High-Level Planning |
ICLR |
| 2024 |
3D Diffusion Policy (DP3) |
End-to-End VLA |
RSS |
| 2024 |
RDT-1B: Diffusion Foundation Model for Bimanual Manipulation |
End-to-End VLA |
ICLR |
| 2024 |
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents |
Datasets & Benchmarks |
arXiv |
| 2024 |
DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting |
Diffusion Policy |
arXiv |
| 2024 |
The Essential Role of Causality in Foundation World Models for Embodied AI |
World Model & Video Policy |
ICML |
| 2023 |
LLaVA: Visual Instruction Tuning |
VLM Foundation |
NeurIPS |
| 2023 |
RH20T |
Datasets & Benchmarks |
RSS Workshop |
| 2023 |
Open X-Embodiment |
Datasets & Benchmarks |
ICRA |
| 2023 |
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion |
Diffusion Policy |
RSS |
| 2023 |
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT/ALOHA) |
Imitation Learning |
RSS |
| 2023 |
AnyTeleop |
Imitation Learning |
CoRL |
| 2023 |
Code as Policies: Language Model Programs for Embodied Control |
High-Level Planning |
ICRA |
| 2023 |
LLM+P: Empowering LLMs with Optimal Planning |
High-Level Planning |
arXiv |
| 2023 |
ProgPrompt |
High-Level Planning |
ICRA |
| 2023 |
BridgeData V2 |
Datasets & Benchmarks |
CoRL |
| 2023 |
GAIA-1 |
World Model & Video Policy |
arXiv |
| 2023 |
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis |
World Model & Video Policy |
arXiv |
| 2022 |
SayCan: Do As I Can, Not As I Say |
High-Level Planning |
CoRL |
| 2022 |
CALVIN |
Datasets & Benchmarks |
RA-L |
| 2022 |
Behavior Transformers: Cloning k Modes with One Stone |
Imitation Learning |
NeurIPS |
| 2022 |
DexMV |
Simulation & Sim2Real |
ECCV |
| 2022 |
RT-1: Robotics Transformer for Real-World Control at Scale |
End-to-End VLA |
RSS |
| 2022 |
DayDreamer |
World Model & Video Policy |
CoRL |
| 2021 |
Meta-StyleSpeech |
Auditory & Acoustic |
ICML |
| 2021 |
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation |
Datasets & Benchmarks |
CoRL |
| 2021 |
Implicit Behavioral Cloning |
Imitation Learning |
CoRL |
| 2021 |
ManiSkill |
Simulation & Sim2Real |
NeurIPS |
| 2020 |
robosuite: A Modular Simulation Framework and Benchmark for Robot Learning |
Datasets & Benchmarks |
arXiv |
| 2019 |
Habitat: A Platform for Embodied AI Research |
Simulation & Sim2Real |
ICCV |
| 2016 |
Generative Adversarial Imitation Learning |
Imitation Learning |
NeurIPS |
| 2011 |
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning |
Imitation Learning |
AISTATS |