| 2026 |
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments |
End-to-End VLA |
arXiv |
| 2026 |
Green-VLA: Staged Vision-Language-Action Model for Generalist Robots |
End-to-End VLA |
arXiv |
| 2026 |
AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation |
End-to-End VLA |
arXiv |
| 2026 |
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation |
End-to-End VLA |
ICLR |
| 2025 |
FlowPolicy: 3D Flow-based Policy via Consistency Flow Matching |
Diffusion Policy |
AAAI |
| 2025 |
pi_0.5: VLA with Open-World Generalization |
Diffusion Policy |
arXiv |
| 2025 |
SmolVLA |
Imitation Learning |
arXiv |
| 2025 |
OpenHelix |
End-to-End VLA |
arXiv |
| 2025 |
1X World Model Challenge |
World Model & Video Policy |
arXiv |
| 2025 |
EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control |
End-to-End VLA |
arXiv |
| 2025 |
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning |
End-to-End VLA |
arXiv |
| 2025 |
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model |
End-to-End VLA |
arXiv |
| 2025 |
HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies |
End-to-End VLA |
arXiv |
| 2025 |
MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation |
End-to-End VLA |
arXiv |
| 2025 |
Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation |
Diffusion Policy |
arXiv |
| 2024 |
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation |
Diffusion Policy |
RSS |
| 2024 |
Affordance-based Robot Manipulation with Flow Matching |
Diffusion Policy |
IROS |
| 2024 |
π₀: A Vision-Language-Action Flow Model for General Robot Control |
Diffusion Policy |
arXiv |