Jason / Works Embodied AIZero to One
Works
没主意?快捷入口
Tag

#vision (107 篇)

yeartitletopicvenue
2026 LeRobot: An Open-Source Library for End-to-End Robot Learning Datasets & Benchmarks ICLR
2026 AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation End-to-End VLA arXiv
2026 3D Generation for Embodied AI and Robotic Simulation: A Survey Simulation & Sim2Real arXiv
2026 Learning Diffusion Policy from Primitive Skills for Robot Manipulation Diffusion Policy arXiv
2026 Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation Diffusion Policy arXiv
2026 InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation End-to-End VLA ICLR
2025 VLAS: VLA Model With Speech Instructions Multimodal Ecology ICLR
2025 DiT-Policy Diffusion Policy ICRA
2025 Diffusion Policy Policy Optimization (DPPO) Diffusion Policy ICLR
2025 FlowPolicy: 3D Flow-based Policy via Consistency Flow Matching Diffusion Policy AAAI
2025 Generalizable Humanoid Manipulation with 3D Diffusion Policies (iDP3) Imitation Learning RSS
2025 Tactile Beyond Pixels (Sparsh-X) Multimodal Ecology CoRL
2025 TLA: Tactile-Language-Action Multimodal Ecology ICRA
2025 Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completion RF Perception & Mapping arXiv
2025 OpenHelix End-to-End VLA arXiv
2025 Cosmos World Foundation Model World Model & Video Policy arXiv
2025 OpenVLA-OFT End-to-End VLA RSS
2025 SpatialVLA End-to-End VLA arXiv
2025 Universal Actions for Enhanced Embodied Foundation Models End-to-End VLA arXiv
2025 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies End-to-End VLA arXiv
2025 Embodiment Transfer Learning for Vision-Language-Action Models End-to-End VLA arXiv
2025 A Survey on Efficient Vision-Language-Action Models End-to-End VLA arXiv
2025 Survey of Vision-Language-Action Models for Embodied Manipulation End-to-End VLA arXiv
2025 Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review End-to-End VLA arXiv
2025 Embodied Navigation Foundation Model End-to-End VLA arXiv
2024 OpenVLA: An Open-Source Vision-Language-Action Model End-to-End VLA CoRL
2024 MLA: Multisensory Language-Action Model Multimodal Ecology arXiv
2024 mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment RF Perception & Mapping SenSys 2024
2024 Stable Audio Auditory & Acoustic ICML
2024 DROID Datasets & Benchmarks RSS
2024 SimplerEnv Simulation & Sim2Real NeurIPS
2024 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations Diffusion Policy RSS
2024 EquiBot: SIM(3)-Equivariant Diffusion Policy Diffusion Policy CoRL
2024 Affordance-based Robot Manipulation with Flow Matching Diffusion Policy IROS
2024 π₀: A Vision-Language-Action Flow Model for General Robot Control Diffusion Policy arXiv
2024 HumanPlus Imitation Learning CoRL
2024 Universal Manipulation Interface Imitation Learning RSS
2024 OneLLM Multimodal Ecology CVPR
2024 Sparsh: Self-supervised Touch Representations Multimodal Ecology CoRL
2024 GenSim High-Level Planning ICLR
2024 Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on RF Perception & Mapping SenSys
2024 Diffusion Model is a Good Pose Estimator from 3D RF-Vision RF Perception & Mapping CVPR
2024 Enabling Visual Recognition at Radio Frequency (PanoRadar) RF Perception & Mapping MobiCom
2024 DexCap Imitation Learning RSS
2024 3D Diffusion Policy (DP3) End-to-End VLA RSS
2024 Octo: An Open-Source Generalist Robot Policy End-to-End VLA RSS
2024 3D-VLA End-to-End VLA ICML
2024 RDT-1B: Diffusion Foundation Model for Bimanual Manipulation End-to-End VLA ICLR
2024 RoboMamba End-to-End VLA NeurIPS
2024 TinyVLA End-to-End VLA RA-L
2024 TraceVLA: Visual Trace Prompting End-to-End VLA ICLR
2024 DeepSeek-VL: Towards Real-World Vision-Language Understanding VLM Foundation arXiv
2024 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks VLM Foundation CVPR
2024 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks VLM Foundation CVPR
2024 Improved Baselines with Visual Instruction Tuning VLM Foundation CVPR
2024 What matters when building vision-language models? VLM Foundation NeurIPS
2024 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling VLM Foundation arXiv
2024 The Llama 3 Herd of Models VLM Foundation arXiv
2024 LLaVA-NeXT-Interleave VLM Foundation arXiv
2024 LLaVA-OneVision: Easy Visual Task Transfer VLM Foundation arXiv
2024 Long-CLIP: Unlocking the Long-Text Capability of CLIP VLM Foundation ECCV
2024 Pixtral 12B VLM Foundation arXiv
2024 Genie: Generative Interactive Environments World Model & Video Policy ICML
2024 UniSim World Model & Video Policy ICLR
2024 AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents Datasets & Benchmarks arXiv
2024 AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding Multimodal Ecology arXiv
2024 SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems High-Level Planning arXiv
2023 LLaVA: Visual Instruction Tuning VLM Foundation NeurIPS
2023 CartoRadar: RF-Based 3D SLAM Rivaling Vision Approaches RF Perception & Mapping MobiCom 2025 (Best Artifact Award)
2023 MusicLM Auditory & Acoustic arXiv
2023 Robust Speech Recognition via Large-Scale Weak Supervision Auditory & Acoustic ICML
2023 Open X-Embodiment Datasets & Benchmarks ICRA
2023 Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT/ALOHA) Imitation Learning RSS
2023 AnyTeleop Imitation Learning CoRL
2023 RoboCat Imitation Learning TMLR
2023 ImageBind: One Embedding Space To Bind Them All Multimodal Ecology CVPR
2023 AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model Multimodal Ecology EACL
2023 FROMAGe: Grounding LLMs to Images Multimodal Ecology ICML
2023 Code as Policies: Language Model Programs for Embodied Control High-Level Planning ICRA
2023 PaLM-E: An Embodied Multimodal Language Model High-Level Planning ICML
2023 VoxPoser High-Level Planning CoRL
2023 BridgeData V2 Datasets & Benchmarks CoRL
2023 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control End-to-End VLA CoRL
2023 RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches End-to-End VLA ICLR
2023 BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models VLM Foundation ICML
2023 EVA-CLIP: Improved Training Techniques for CLIP at Scale VLM Foundation arXiv
2023 OBELICS Multimodal Ecology NeurIPS
2023 Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond VLM Foundation arXiv
2023 Sigmoid Loss for Language Image Pre-Training VLM Foundation ICCV
2023 GAIA-1 World Model & Video Policy arXiv
2022 CALVIN Datasets & Benchmarks RA-L
2022 X-VLM: Multi-Grained Vision Language Pre-Training Multimodal Ecology ICML
2022 RFMask: A Simple Baseline for Human Silhouette Segmentation with Radio Signals RF Perception & Mapping TMM
2022 ProcTHOR Simulation & Sim2Real NeurIPS
2022 DexMV Simulation & Sim2Real ECCV
2022 Flamingo: a Visual Language Model for Few-Shot Learning VLM Foundation NeurIPS
2022 BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation VLM Foundation ICML
2022 FILIP: Fine-grained Interactive Language-Image Pre-Training VLM Foundation ICLR
2022 DayDreamer World Model & Video Policy CoRL
2021 ManiSkill Simulation & Sim2Real NeurIPS
2021 Learning Transferable Visual Models From Natural Language Supervision VLM Foundation ICML
2020 See Through Smoke: Robust Indoor Mapping with Low-cost mmWave Radar RF Perception & Mapping MobiSys
2020 SAPIEN: A SimulAted Part-based Interactive ENvironment Simulation & Sim2Real CVPR
2019 RLBench: The Robot Learning Benchmark & Learning Environment Datasets & Benchmarks RA-L
2019 Connecting Touch and Vision via Cross-Modal Prediction Multimodal Ecology CVPR
2019 Through-Wall Pose Imaging in Real-Time with a Many-to-Many Encoder/Decoder Paradigm RF Perception & Mapping arXiv
2019 Habitat: A Platform for Embodied AI Research Simulation & Sim2Real ICCV