Jason / Works Embodied AIZero to One
Works
没主意?快捷入口
Tag

#VLM (57 篇)

yeartitletopicvenue
2026 Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments End-to-End VLA arXiv
2026 Membership Inference Attacks on Vision-Language-Action Models End-to-End VLA arXiv
2026 Learning Diffusion Policy from Primitive Skills for Robot Manipulation Diffusion Policy arXiv
2026 Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation End-to-End VLA arXiv
2026 InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation End-to-End VLA ICLR
2026 A Survey of Language-Conditioned Robot Manipulation End-to-End VLA IJRR
2025 VLAS: VLA Model With Speech Instructions Multimodal Ecology ICLR
2025 pi_0.5: VLA with Open-World Generalization Diffusion Policy arXiv
2025 SmolVLA Imitation Learning arXiv
2025 Tactile-VLA Multimodal Ecology CoRL
2025 DexVLA End-to-End VLA arXiv
2025 OpenHelix End-to-End VLA arXiv
2025 SpatialVLA End-to-End VLA arXiv
2025 LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks High-Level Planning arXiv
2025 EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control End-to-End VLA arXiv
2025 RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI Datasets & Benchmarks arXiv
2025 LLaDA-VLA: Vision Language Diffusion Action Models End-to-End VLA arXiv
2025 Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning End-to-End VLA arXiv
2025 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model End-to-End VLA arXiv
2025 Embodiment Transfer Learning for Vision-Language-Action Models End-to-End VLA arXiv
2025 HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies End-to-End VLA arXiv
2025 Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation End-to-End VLA arXiv
2025 Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review End-to-End VLA arXiv
2025 Toward Embodied AGI: A Review of Embodied AI and the Road Ahead End-to-End VLA arXiv
2025 Embodied Navigation Foundation Model End-to-End VLA arXiv
2025 MiMo-Embodied: X-Embodied Foundation Model Technical Report End-to-End VLA arXiv
2025 LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation End-to-End VLA arXiv
2025 Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy Datasets & Benchmarks ICRA
2024 OpenVLA: An Open-Source Vision-Language-Action Model End-to-End VLA CoRL
2024 RoboFlamingo High-Level Planning ICLR
2024 Octo: An Open-Source Generalist Robot Policy End-to-End VLA RSS
2024 GR-2: Generative Video-Language-Action Model End-to-End VLA arXiv
2024 TinyVLA End-to-End VLA RA-L
2024 TraceVLA: Visual Trace Prompting End-to-End VLA ICLR
2024 DeepSeek-VL: Towards Real-World Vision-Language Understanding VLM Foundation arXiv
2024 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks VLM Foundation CVPR
2024 Improved Baselines with Visual Instruction Tuning VLM Foundation CVPR
2024 What matters when building vision-language models? VLM Foundation NeurIPS
2024 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VLM Foundation arXiv
2024 LLaVA-NeXT-Interleave VLM Foundation arXiv
2024 LLaVA-OneVision: Easy Visual Task Transfer VLM Foundation arXiv
2024 Pixtral 12B VLM Foundation arXiv
2024 UniSim World Model & Video Policy ICLR
2024 CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation End-to-End VLA arXiv
2024 AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents Datasets & Benchmarks arXiv
2024 AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding Multimodal Ecology arXiv
2024 DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting Diffusion Policy arXiv
2024 A call for embodied AI VLM Foundation ICML
2023 LLaVA: Visual Instruction Tuning VLM Foundation NeurIPS
2023 PaLM-E: An Embodied Multimodal Language Model High-Level Planning ICML
2023 VoxPoser High-Level Planning CoRL
2023 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control End-to-End VLA CoRL
2023 BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models VLM Foundation ICML
2023 OBELICS Multimodal Ecology NeurIPS
2023 Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond VLM Foundation arXiv
2023 Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis World Model & Video Policy arXiv
2022 X-VLM: Multi-Grained Vision Language Pre-Training Multimodal Ecology ICML