Jason / Works Embodied AIZero to One
Works
没主意?快捷入口
Topic III · 端到端视觉-语言-动作

End-to-End VLA

End-to-End VLA — 端到端视觉-语言-动作
44papers
1founder
5classic
38frontier

VLA = 视觉-语言-动作模型。一个端到端神经网络:左边输入摄像头画面 + 自然语言指令,右边输出关节速度。这是过去三年具身 AI 最热的赛道。


Primer · 入门 3 篇

先读这三篇

RT-1 把动作 token 化 → RT-2 把网络知识带进来 → OpenVLA 把整个范式开源民主化。

  1. 1
    RT-1: Robotics Transformer for Real-World Control at Scale 2022 · RSS · ⭐⭐⭐

    Google 用 13 台机器人花 17 个月录下 13 万段人类示范,训出一个 35M 参数的小 Transformer,让机器人听一句指令就能在真办公室厨房里完成 700 多种操作,并首次证明"大数据 + Transformer"的范式在物理世界同样有效。

  2. 2
    RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control 2023 · CoRL · ⭐⭐⭐⭐

    把机器人动作翻译成一句话,让会看图聊天的 AI 用写句子的方式开口指挥机器人——它会写字,就能动手。

  3. 3
    OpenVLA: An Open-Source Vision-Language-Action Model 2024 · CoRL · ⭐⭐⭐

    把一个会"看图说话"的 AI 改一改,让它学会"看一眼桌面就动手摆东西"——把机械臂动作伪装成"单词",用 7B 参数全面打败闭源 55B 的 RT-2-X,并把全部训练配方开源送出去。


Distribution · 年份分布

2022 到 2026,44 篇怎么排开。

祖师爷 经典 前沿
All papers · 按 era 排

End-to-End VLA 全部 44 篇。

erayeartitlevenue
经典 2024 OpenVLA: An Open-Source Vision-Language-Action Model CoRL
祖师爷 2022 RT-1: Robotics Transformer for Real-World Control at Scale RSS
经典 2023 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control CoRL
经典 2023 RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches ICLR
经典 2024 3D Diffusion Policy (DP3) RSS
经典 2024 Octo: An Open-Source Generalist Robot Policy RSS
前沿 2024 3D-VLA ICML
前沿 2024 GR-2: Generative Video-Language-Action Model arXiv
前沿 2024 RDT-1B: Diffusion Foundation Model for Bimanual Manipulation ICLR
前沿 2024 RoboMamba NeurIPS
前沿 2024 TinyVLA RA-L
前沿 2024 TraceVLA: Visual Trace Prompting ICLR
前沿 2024 CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation arXiv
前沿 2025 DexVLA arXiv
前沿 2025 OpenHelix arXiv
前沿 2025 OpenVLA-OFT RSS
前沿 2025 SpatialVLA arXiv
前沿 2025 Universal Actions for Enhanced Embodied Foundation Models arXiv
前沿 2025 EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control arXiv
前沿 2025 LLaDA-VLA: Vision Language Diffusion Action Models arXiv
前沿 2025 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies arXiv
前沿 2025 Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning arXiv
前沿 2025 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model arXiv
前沿 2025 Embodiment Transfer Learning for Vision-Language-Action Models arXiv
前沿 2025 HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies arXiv
前沿 2025 MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation arXiv
前沿 2025 Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation arXiv
前沿 2025 A Survey on Efficient Vision-Language-Action Models arXiv
前沿 2025 Survey of Vision-Language-Action Models for Embodied Manipulation arXiv
前沿 2025 Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review arXiv
前沿 2025 Toward Embodied AGI: A Review of Embodied AI and the Road Ahead arXiv
前沿 2025 RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI arXiv
前沿 2025 Embodied Navigation Foundation Model arXiv
前沿 2025 MiMo-Embodied: X-Embodied Foundation Model Technical Report arXiv
前沿 2025 LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation arXiv
前沿 2025 villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models arXiv
前沿 2026 Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments arXiv
前沿 2026 Green-VLA: Staged Vision-Language-Action Model for Generalist Robots arXiv
前沿 2026 AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation arXiv
前沿 2026 VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models arXiv
前沿 2026 Membership Inference Attacks on Vision-Language-Action Models arXiv
前沿 2026 Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation arXiv
前沿 2026 InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation ICLR
前沿 2026 A Survey of Language-Conditioned Robot Manipulation IJRR