| 2026 |
LeRobot: An Open-Source Library for End-to-End Robot Learning |
Datasets & Benchmarks |
ICLR |
| 2026 |
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots |
End-to-End VLA |
arXiv |
| 2026 |
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
3D Generation for Embodied AI and Robotic Simulation: A Survey |
Simulation & Sim2Real |
arXiv |
| 2026 |
Learning Diffusion Policy from Primitive Skills for Robot Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
A Survey of Language-Conditioned Robot Manipulation |
End-to-End VLA |
IJRR |
| 2025 |
Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control |
World Model & Video Policy |
arXiv |
| 2025 |
DiT-Policy |
Diffusion Policy |
ICRA |
| 2025 |
Diffusion Policy Policy Optimization (DPPO) |
Diffusion Policy |
ICLR |
| 2025 |
FlowPolicy: 3D Flow-based Policy via Consistency Flow Matching |
Diffusion Policy |
AAAI |
| 2025 |
FAST: Efficient Action Tokenization for VLA |
Diffusion Policy |
RSS |
| 2025 |
Generalizable Humanoid Manipulation with 3D Diffusion Policies (iDP3) |
Imitation Learning |
RSS |
| 2025 |
DexVLA |
End-to-End VLA |
arXiv |
| 2025 |
OpenHelix |
End-to-End VLA |
arXiv |
| 2025 |
Cosmos World Foundation Model |
World Model & Video Policy |
arXiv |
| 2025 |
OpenVLA-OFT |
End-to-End VLA |
RSS |
| 2025 |
1X World Model Challenge |
World Model & Video Policy |
arXiv |
| 2025 |
Navigation World Models |
World Model & Video Policy |
CVPR |
| 2025 |
EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control |
End-to-End VLA |
arXiv |
| 2025 |
RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI |
Datasets & Benchmarks |
arXiv |
| 2025 |
LLaDA-VLA: Vision Language Diffusion Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
Embodiment Transfer Learning for Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation |
End-to-End VLA |
arXiv |
| 2025 |
Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Survey of Vision-Language-Action Models for Embodied Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2025 |
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation |
Diffusion Policy |
ICRA |
| 2024 |
Stable Audio |
Auditory & Acoustic |
ICML |
| 2024 |
DROID |
Datasets & Benchmarks |
RSS |
| 2024 |
3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations |
Diffusion Policy |
RSS |
| 2024 |
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation |
Diffusion Policy |
RSS |
| 2024 |
EquiBot: SIM(3)-Equivariant Diffusion Policy |
Diffusion Policy |
CoRL |
| 2024 |
Affordance-based Robot Manipulation with Flow Matching |
Diffusion Policy |
IROS |
| 2024 |
π₀: A Vision-Language-Action Flow Model for General Robot Control |
Diffusion Policy |
arXiv |
| 2024 |
ALOHA 2 |
Imitation Learning |
Tech Report |
| 2024 |
Mobile ALOHA |
Imitation Learning |
CoRL |
| 2024 |
Universal Manipulation Interface |
Imitation Learning |
RSS |
| 2024 |
Behavior Generation with Latent Actions (VQ-BeT) |
Imitation Learning |
ICML |
| 2024 |
Diffusion Model is a Good Pose Estimator from 3D RF-Vision |
RF Perception & Mapping |
CVPR |
| 2024 |
DexCap |
Imitation Learning |
RSS |
| 2024 |
3D Diffusion Policy (DP3) |
End-to-End VLA |
RSS |
| 2024 |
Octo: An Open-Source Generalist Robot Policy |
End-to-End VLA |
RSS |
| 2024 |
3D-VLA |
End-to-End VLA |
ICML |
| 2024 |
GR-2: Generative Video-Language-Action Model |
End-to-End VLA |
arXiv |
| 2024 |
RDT-1B: Diffusion Foundation Model for Bimanual Manipulation |
End-to-End VLA |
ICLR |
| 2024 |
RoboMamba |
End-to-End VLA |
NeurIPS |
| 2024 |
TinyVLA |
End-to-End VLA |
RA-L |
| 2024 |
Long-CLIP: Unlocking the Long-Text Capability of CLIP |
VLM Foundation |
ECCV |
| 2024 |
Genie: Generative Interactive Environments |
World Model & Video Policy |
ICML |
| 2024 |
UniSim |
World Model & Video Policy |
ICLR |
| 2024 |
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2024 |
DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting |
Diffusion Policy |
arXiv |
| 2024 |
The Essential Role of Causality in Foundation World Models for Embodied AI |
World Model & Video Policy |
ICML |
| 2023 |
3DShape2VecSet: 3D Shape Representation for Diffusion Models |
VLM Foundation |
SIGGRAPH |
| 2023 |
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion |
Diffusion Policy |
RSS |
| 2023 |
GAIA-1 |
World Model & Video Policy |
arXiv |
| 2021 |
What Matters in Learning from Offline Human Demonstrations for Robot Manipulation |
Datasets & Benchmarks |
CoRL |