| 2026 |
LeRobot: An Open-Source Library for End-to-End Robot Learning |
Datasets & Benchmarks |
ICLR |
| 2026 |
VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models |
End-to-End VLA |
arXiv |
| 2026 |
Membership Inference Attacks on Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2026 |
3D Generation for Embodied AI and Robotic Simulation: A Survey |
Simulation & Sim2Real |
arXiv |
| 2026 |
Learning Diffusion Policy from Primitive Skills for Robot Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2026 |
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation |
End-to-End VLA |
ICLR |
| 2026 |
A Survey of Language-Conditioned Robot Manipulation |
End-to-End VLA |
IJRR |
| 2026 |
What Breaks Embodied AI Security: LLM Vulnerabilities, CPS Flaws, or Something Else? |
High-Level Planning |
arXiv |
| 2025 |
VLAS: VLA Model With Speech Instructions |
Multimodal Ecology |
ICLR |
| 2025 |
FAST: Efficient Action Tokenization for VLA |
Diffusion Policy |
RSS |
| 2025 |
pi_0.5: VLA with Open-World Generalization |
Diffusion Policy |
arXiv |
| 2025 |
TLA: Tactile-Language-Action |
Multimodal Ecology |
ICRA |
| 2025 |
OpenHelix |
End-to-End VLA |
arXiv |
| 2025 |
Cosmos World Foundation Model |
World Model & Video Policy |
arXiv |
| 2025 |
OpenVLA-OFT |
End-to-End VLA |
RSS |
| 2025 |
EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning |
End-to-End VLA |
arXiv |
| 2025 |
A Survey on Efficient Vision-Language-Action Models |
End-to-End VLA |
arXiv |
| 2025 |
Survey of Vision-Language-Action Models for Embodied Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review |
End-to-End VLA |
arXiv |
| 2025 |
Toward Embodied AGI: A Review of Embodied AI and the Road Ahead |
End-to-End VLA |
arXiv |
| 2025 |
RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI |
End-to-End VLA |
arXiv |
| 2025 |
Embodied Navigation Foundation Model |
End-to-End VLA |
arXiv |
| 2025 |
MiMo-Embodied: X-Embodied Foundation Model Technical Report |
End-to-End VLA |
arXiv |
| 2025 |
LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2025 |
Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation |
Diffusion Policy |
ICRA |
| 2025 |
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy |
Datasets & Benchmarks |
ICRA |
| 2024 |
OpenVLA: An Open-Source Vision-Language-Action Model |
End-to-End VLA |
CoRL |
| 2024 |
MLA: Multisensory Language-Action Model |
Multimodal Ecology |
arXiv |
| 2024 |
mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment |
RF Perception & Mapping |
SenSys 2024 |
| 2024 |
DROID |
Datasets & Benchmarks |
RSS |
| 2024 |
RoboCasa |
Datasets & Benchmarks |
RSS |
| 2024 |
π₀: A Vision-Language-Action Flow Model for General Robot Control |
Diffusion Policy |
arXiv |
| 2024 |
Behavior Generation with Latent Actions (VQ-BeT) |
Imitation Learning |
ICML |
| 2024 |
OneLLM |
Multimodal Ecology |
CVPR |
| 2024 |
GenSim |
High-Level Planning |
ICLR |
| 2024 |
Tree-Planner |
High-Level Planning |
ICLR |
| 2024 |
Octo: An Open-Source Generalist Robot Policy |
End-to-End VLA |
RSS |
| 2024 |
3D-VLA |
End-to-End VLA |
ICML |
| 2024 |
GR-2: Generative Video-Language-Action Model |
End-to-End VLA |
arXiv |
| 2024 |
RDT-1B: Diffusion Foundation Model for Bimanual Manipulation |
End-to-End VLA |
ICLR |
| 2024 |
RoboMamba |
End-to-End VLA |
NeurIPS |
| 2024 |
TinyVLA |
End-to-End VLA |
RA-L |
| 2024 |
DeepSeek-VL: Towards Real-World Vision-Language Understanding |
VLM Foundation |
arXiv |
| 2024 |
Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks |
VLM Foundation |
CVPR |
| 2024 |
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks |
VLM Foundation |
CVPR |
| 2024 |
Improved Baselines with Visual Instruction Tuning |
VLM Foundation |
CVPR |
| 2024 |
What matters when building vision-language models? |
VLM Foundation |
NeurIPS |
| 2024 |
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling |
VLM Foundation |
arXiv |
| 2024 |
The Llama 3 Herd of Models |
VLM Foundation |
arXiv |
| 2024 |
LLaVA-NeXT-Interleave |
VLM Foundation |
arXiv |
| 2024 |
LLaVA-OneVision: Easy Visual Task Transfer |
VLM Foundation |
arXiv |
| 2024 |
Long-CLIP: Unlocking the Long-Text Capability of CLIP |
VLM Foundation |
ECCV |
| 2024 |
Pixtral 12B |
VLM Foundation |
arXiv |
| 2024 |
Genie: Generative Interactive Environments |
World Model & Video Policy |
ICML |
| 2024 |
AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents |
Datasets & Benchmarks |
arXiv |
| 2024 |
AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding |
Multimodal Ecology |
arXiv |
| 2024 |
DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting |
Diffusion Policy |
arXiv |
| 2024 |
SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems |
High-Level Planning |
arXiv |
| 2024 |
A call for embodied AI |
VLM Foundation |
ICML |
| 2023 |
LLaVA: Visual Instruction Tuning |
VLM Foundation |
NeurIPS |
| 2023 |
AudioLM |
Auditory & Acoustic |
TASLP |
| 2023 |
RH20T |
Datasets & Benchmarks |
RSS Workshop |
| 2023 |
Open X-Embodiment |
Datasets & Benchmarks |
ICRA |
| 2023 |
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model |
Multimodal Ecology |
EACL |
| 2023 |
AudioPaLM |
Multimodal Ecology |
arXiv |
| 2023 |
FROMAGe: Grounding LLMs to Images |
Multimodal Ecology |
ICML |
| 2023 |
Code as Policies: Language Model Programs for Embodied Control |
High-Level Planning |
ICRA |
| 2023 |
LLM+P: Empowering LLMs with Optimal Planning |
High-Level Planning |
arXiv |
| 2023 |
PaLM-E: An Embodied Multimodal Language Model |
High-Level Planning |
ICML |
| 2023 |
ProgPrompt |
High-Level Planning |
ICRA |
| 2023 |
ChatGPT for Robotics |
High-Level Planning |
IEEE Access |
| 2023 |
VoxPoser |
High-Level Planning |
CoRL |
| 2023 |
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control |
End-to-End VLA |
CoRL |
| 2023 |
RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches |
End-to-End VLA |
ICLR |
| 2023 |
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models |
VLM Foundation |
ICML |
| 2023 |
OBELICS |
Multimodal Ecology |
NeurIPS |
| 2023 |
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond |
VLM Foundation |
arXiv |
| 2023 |
Transformers are Sample-Efficient World Models |
World Model & Video Policy |
ICLR |
| 2023 |
TWM: Transformer-based World Models |
World Model & Video Policy |
ICLR |
| 2023 |
GAIA-1 |
World Model & Video Policy |
arXiv |
| 2023 |
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis |
World Model & Video Policy |
arXiv |
| 2022 |
SayCan: Do As I Can, Not As I Say |
High-Level Planning |
CoRL |
| 2022 |
Behavior Transformers: Cloning k Modes with One Stone |
Imitation Learning |
NeurIPS |
| 2022 |
Inner Monologue: Embodied Reasoning through Planning with Language Models |
High-Level Planning |
CoRL |
| 2022 |
RT-1: Robotics Transformer for Real-World Control at Scale |
End-to-End VLA |
RSS |
| 2022 |
Flamingo: a Visual Language Model for Few-Shot Learning |
VLM Foundation |
NeurIPS |
| 2021 |
Learning Transferable Visual Models From Natural Language Supervision |
VLM Foundation |
ICML |