11 chapters · 202 papers.
视觉-语言基座 · read primer →
- № 01 LLaVA: Visual Instruction Tuning ⭐⭐ deep
- № 02 3DShape2VecSet: 3D Shape Representation for Diffusion Models ⭐⭐⭐⭐ deep
- № 124 Learning Transferable Visual Models From Natural Language Supervision ⭐⭐⭐ deep
- № 125 Flamingo: a Visual Language Model for Few-Shot Learning ⭐⭐⭐⭐ deep
- № 126 BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models ⭐⭐⭐⭐ deep
- № 127 BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation ⭐⭐⭐ deep
- № 128 DeepSeek-VL: Towards Real-World Vision-Language Understanding ⭐⭐⭐ deep
- № 129 EVA-CLIP: Improved Training Techniques for CLIP at Scale ⭐⭐⭐ deep
- № 130 FILIP: Fine-grained Interactive Language-Image Pre-Training ⭐⭐⭐ deep
- № 131 Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks ⭐⭐⭐ deep
- № 132 InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks ⭐⭐⭐⭐ deep
- № 133 Improved Baselines with Visual Instruction Tuning ⭐⭐ deep
- № 135 Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond ⭐⭐⭐ deep
- № 136 Sigmoid Loss for Language Image Pre-Training ⭐⭐⭐ deep
- № 137 What matters when building vision-language models? ⭐⭐⭐ deep
- № 138 Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling ⭐⭐⭐⭐ deep
- № 139 The Llama 3 Herd of Models ⭐⭐⭐⭐ deep
- № 140 LLaVA-NeXT-Interleave ⭐⭐⭐ deep
- № 141 LLaVA-OneVision: Easy Visual Task Transfer ⭐⭐⭐ deep
- № 142 Long-CLIP: Unlocking the Long-Text Capability of CLIP ⭐⭐⭐ deep
- № 143 Pixtral 12B ⭐⭐⭐ deep
- № 200 A call for embodied AI ⭐⭐⭐ deep
高层任务规划 · read primer →
- № 03 SayCan: Do As I Can, Not As I Say ⭐⭐ deep
- № 75 Code as Policies: Language Model Programs for Embodied Control ⭐⭐⭐ deep
- № 76 Inner Monologue: Embodied Reasoning through Planning with Language Models ⭐⭐⭐ deep
- № 77 LLM+P: Empowering LLMs with Optimal Planning ⭐⭐⭐ deep
- № 78 PaLM-E: An Embodied Multimodal Language Model ⭐⭐⭐⭐ deep
- № 79 ProgPrompt ⭐⭐ deep
- № 80 ChatGPT for Robotics ⭐⭐ deep
- № 81 GenSim ⭐⭐⭐ deep
- № 82 RoboFlamingo ⭐⭐⭐⭐ deep
- № 83 Tree-Planner ⭐⭐⭐ deep
- № 84 VoxPoser ⭐⭐⭐⭐ deep
- № 160 LoHoVLA: A Unified Vision-Language-Action Model for Long-Horizon Embodied Tasks ⭐⭐⭐ deep
- № 198 SafeEmbodAI: a Safety Framework for Mobile Robots in Embodied AI Systems ⭐⭐⭐ deep
- № 202 What Breaks Embodied AI Security: LLM Vulnerabilities, CPS Flaws, or Something Else? ⭐⭐⭐⭐ deep
端到端视觉-语言-动作 · read primer →
- № 04 OpenVLA: An Open-Source Vision-Language-Action Model ⭐⭐⭐ deep
- № 109 RT-1: Robotics Transformer for Real-World Control at Scale ⭐⭐⭐ deep
- № 110 3D Diffusion Policy (DP3) ⭐⭐⭐ deep
- № 111 Octo: An Open-Source Generalist Robot Policy ⭐⭐⭐ deep
- № 112 RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control ⭐⭐⭐⭐ deep
- № 113 RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches 3 deep
- № 114 3D-VLA ⭐⭐⭐⭐ deep
- № 116 GR-2: Generative Video-Language-Action Model ⭐⭐⭐⭐ deep
- № 117 DexVLA ⭐⭐⭐⭐ deep
- № 117 OpenHelix ⭐⭐⭐ deep
- № 118 OpenVLA-OFT ⭐⭐⭐ deep
- № 119 RDT-1B: Diffusion Foundation Model for Bimanual Manipulation ⭐⭐⭐⭐ deep
- № 120 RoboMamba ⭐⭐⭐ deep
- № 121 SpatialVLA ⭐⭐⭐⭐ deep
- № 122 TinyVLA ⭐⭐⭐ deep
- № 123 TraceVLA: Visual Trace Prompting ⭐⭐⭐ deep
- № 158 CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation ⭐⭐⭐ deep
- № 159 Universal Actions for Enhanced Embodied Foundation Models ⭐⭐⭐⭐ deep
- № 162 EO-1: Interleaved Vision-Text-Action Pretraining for General Robot Control ⭐⭐⭐⭐ deep
- № 163 Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments ⭐⭐⭐⭐ deep
- № 165 LLaDA-VLA: Vision Language Diffusion Action Models ⭐⭐⭐⭐ deep
- № 166 Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies ⭐⭐⭐⭐ deep
- № 167 Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning ⭐⭐⭐⭐ deep
- № 168 X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model ⭐⭐⭐⭐ deep
- № 169 Embodiment Transfer Learning for Vision-Language-Action Models ⭐⭐⭐⭐ deep
- № 170 HiMoE-VLA: Hierarchical Mixture-of-Experts for Generalist Vision-Language-Action Policies ⭐⭐⭐⭐ deep
- № 171 Green-VLA: Staged Vision-Language-Action Model for Generalist Robots ⭐⭐⭐⭐ deep
- № 172 AC^2-VLA: Action-Context-Aware Adaptive Computation in Vision-Language-Action Models for Efficient Robotic Manipulation ⭐⭐⭐⭐ deep
- № 173 MoS-VLA: A Vision-Language-Action Model with One-Shot Skill Adaptation ⭐⭐⭐⭐ deep
- № 174 Asynchronous Fast-Slow Vision-Language-Action Policies for Whole-Body Robotic Manipulation ⭐⭐⭐⭐ deep
- № 175 VLA-Forget: Vision-Language-Action Unlearning for Embodied Foundation Models ⭐⭐⭐⭐ deep
- № 176 Membership Inference Attacks on Vision-Language-Action Models ⭐⭐⭐⭐ deep
- № 177 A Survey on Efficient Vision-Language-Action Models ⭐⭐⭐ deep
- № 178 Survey of Vision-Language-Action Models for Embodied Manipulation ⭐⭐⭐ deep
- № 179 Embodied AI with Foundation Models for Mobile Service Robots: A Systematic Review ⭐⭐⭐ deep
- № 180 Toward Embodied AGI: A Review of Embodied AI and the Road Ahead ⭐⭐⭐ deep
- № 181 RoboNeuron: A Middle-Layer Infrastructure for Agent-Driven Orchestration in Embodied AI ⭐⭐⭐⭐ deep
- № 182 Embodied Navigation Foundation Model ⭐⭐⭐⭐ deep
- № 183 MiMo-Embodied: X-Embodied Foundation Model Technical Report ⭐⭐⭐⭐ deep
- № 191 Gaze2Act: Gaze-Conditioned Vision-Language-Action Policies for Interactive Robot Manipulation ⭐⭐⭐⭐ deep
- № 192 LACY: A Vision-Language Model-based Language-Action Cycle for Self-Improving Robotic Manipulation ⭐⭐⭐⭐ deep
- № 193 villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models ⭐⭐⭐⭐ deep
- № 194 InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation ⭐⭐⭐⭐ deep
- № 197 A Survey of Language-Conditioned Robot Manipulation ⭐⭐⭐ deep
扩散策略与流匹配 · read primer →
- № 38 Diffusion Policy: Visuomotor Policy Learning via Action Diffusion ⭐⭐⭐ deep
- № 39 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations ⭐⭐⭐ deep
- № 40 Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation ⭐⭐⭐ deep
- № 41 EquiBot: SIM(3)-Equivariant Diffusion Policy ⭐⭐⭐⭐ deep
- № 42 DiT-Policy ⭐⭐⭐⭐ deep
- № 43 Diffusion Policy Policy Optimization (DPPO) ⭐⭐⭐⭐ deep
- № 44 Affordance-based Robot Manipulation with Flow Matching ⭐⭐⭐ deep
- № 45 FlowPolicy: 3D Flow-based Policy via Consistency Flow Matching ⭐⭐⭐⭐ deep
- № 46 FAST: Efficient Action Tokenization for VLA ⭐⭐⭐⭐ deep
- № 47 π₀: A Vision-Language-Action Flow Model for General Robot Control ⭐⭐⭐⭐ deep
- № 48 pi_0.5: VLA with Open-World Generalization ⭐⭐⭐⭐⭐ deep
- № 187 DISCO: Language-Guided Manipulation with Diffusion Policies and Constrained Inpainting ⭐⭐⭐⭐ deep
- № 188 Time-Unified Diffusion Policy with Action Discrimination for Robotic Manipulation ⭐⭐⭐⭐ deep
- № 189 Learning Diffusion Policy from Primitive Skills for Robot Manipulation ⭐⭐⭐⭐ deep
- № 190 Trace-Focused Diffusion Policy for Multi-Modal Action Disambiguation in Long-Horizon Robotic Manipulation ⭐⭐⭐⭐ deep
- № 195 Discrete Policy: Learning Disentangled Action Space for Multi-Task Robotic Manipulation ⭐⭐⭐⭐ deep
模仿学习与遥操作 · read primer →
- № 49 A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning ⭐⭐⭐⭐ deep
- № 50 Generative Adversarial Imitation Learning ⭐⭐⭐⭐ deep
- № 51 Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT/ALOHA) ⭐⭐⭐ deep
- № 52 AnyTeleop ⭐⭐⭐ deep
- № 53 Behavior Transformers: Cloning k Modes with One Stone ⭐⭐⭐ deep
- № 54 Implicit Behavioral Cloning ⭐⭐⭐⭐ deep
- № 55 RoboCat ⭐⭐⭐⭐ deep
- № 56 ALOHA 2 ⭐⭐ deep
- № 58 HumanPlus ⭐⭐⭐⭐ deep
- № 59 Generalizable Humanoid Manipulation with 3D Diffusion Policies (iDP3) ⭐⭐⭐⭐ deep
- № 60 Mobile ALOHA ⭐⭐⭐ deep
- № 61 SmolVLA ⭐⭐⭐ deep
- № 62 Universal Manipulation Interface ⭐⭐⭐ deep
- № 63 Behavior Generation with Latent Actions (VQ-BeT) ⭐⭐⭐⭐ deep
- № 109 DexCap ⭐⭐⭐⭐ deep
世界模型与视频策略 · read primer →
- № 07 Cosmos Policy: Fine-Tuning Video Models for Visuomotor Control ⭐⭐⭐⭐⭐ deep
- № 144 Dream to Control: Learning Behaviors by Latent Imagination ⭐⭐⭐⭐ deep
- № 145 World Models ⭐⭐⭐ deep
- № 146 DayDreamer ⭐⭐⭐ deep
- № 147 Mastering Atari with Discrete World Models ⭐⭐⭐⭐ deep
- № 148 Dreamer V3: Mastering Diverse Domains through World Models ⭐⭐⭐⭐ deep
- № 149 Transformers are Sample-Efficient World Models ⭐⭐⭐⭐ deep
- № 150 TWM: Transformer-based World Models ⭐⭐⭐⭐ deep
- № 118 Cosmos World Foundation Model ⭐⭐⭐⭐ deep
- № 151 1X World Model Challenge ⭐⭐⭐ deep
- № 153 GAIA-1 ⭐⭐⭐⭐ deep
- № 154 Genie: Generative Interactive Environments ⭐⭐⭐⭐ deep
- № 155 Navigation World Models ⭐⭐⭐⭐ deep
- № 156 UniSim ⭐⭐⭐⭐ deep
- № 199 The Essential Role of Causality in Foundation World Models for Embodied AI ⭐⭐⭐⭐ deep
- № 201 Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis ⭐⭐⭐⭐ deep
多模态交互与数据生态 · read primer →
- № 05 VLAS: VLA Model With Speech Instructions ⭐⭐⭐ deep
- № 06 MLA: Multisensory Language-Action Model ⭐⭐⭐⭐ deep
- № 64 ImageBind: One Embedding Space To Bind Them All ⭐⭐⭐ deep
- № 65 Connecting Touch and Vision via Cross-Modal Prediction ⭐⭐⭐ deep
- № 66 AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model ⭐⭐⭐ deep
- № 67 AudioPaLM ⭐⭐⭐⭐ deep
- № 68 FROMAGe: Grounding LLMs to Images ⭐⭐⭐ deep
- № 69 OneLLM ⭐⭐⭐ deep
- № 70 X-VLM: Multi-Grained Vision Language Pre-Training ⭐⭐⭐⭐ deep
- № 134 OBELICS ⭐⭐⭐ deep
- № 71 Tactile Beyond Pixels (Sparsh-X) ⭐⭐⭐⭐ deep
- № 72 Sparsh: Self-supervised Touch Representations ⭐⭐⭐⭐ deep
- № 73 Tactile-VLA ⭐⭐⭐⭐ deep
- № 74 TLA: Tactile-Language-Action ⭐⭐⭐⭐ deep
- № 185 AlanaVLM: A Multimodal Embodied AI Foundation Model for Egocentric Video Understanding ⭐⭐⭐ deep
射频感知与空间建图 · read primer →
- № 08 CartoRadar: RF-Based 3D SLAM Rivaling Vision Approaches ⭐⭐⭐⭐ deep
- № 09 mmCLIP: Boosting mmWave-based Zero-shot HAR via Signal-Text Alignment ⭐⭐⭐⭐ deep
- № 10 mmNorm: Non-Line-of-Sight 3D Object Reconstruction via mmWave Surface Normal Estimation ⭐⭐⭐⭐ deep
- № 85 See Through Smoke: Robust Indoor Mapping with Low-cost mmWave Radar ⭐⭐⭐ deep
- № 86 Can WiFi Estimate Person Pose? ⭐⭐⭐ deep
- № 87 3DRIMR: 3D Reconstruction and Imaging via mmWave Radar based on Deep Learning ⭐⭐⭐ deep
- № 88 milliEgo: Single-chip mmWave Radar Aided Egomotion Estimation via Deep Sensor Fusion ⭐⭐⭐ deep
- № 89 High Resolution Point Clouds from mmWave Radar ⭐⭐⭐ deep
- № 90 RadarSLAM: Radar based Large-Scale SLAM in All Weathers ⭐⭐⭐⭐ deep
- № 91 Through-Wall Pose Imaging in Real-Time with a Many-to-Many Encoder/Decoder Paradigm ⭐⭐⭐⭐ deep
- № 92 RFMask: A Simple Baseline for Human Silhouette Segmentation with Radio Signals ⭐⭐⭐ deep
- № 93 RFPose-OT: RF-Based 3D Human Pose Estimation via Optimal Transport Theory ⭐⭐⭐⭐ deep
- № 94 Argus: Multi-View Egocentric Human Mesh Reconstruction Based on Stripped-Down Wearable mmWave Add-on ⭐⭐⭐⭐ deep
- № 95 Diffusion Model is a Good Pose Estimator from 3D RF-Vision ⭐⭐⭐⭐ deep
- № 96 Enabling Visual Recognition at Radio Frequency (PanoRadar) ⭐⭐⭐⭐ deep
- № 97 Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completion ⭐⭐⭐⭐ deep
听觉智能与声学空间交互 · read primer →
- № 11 Proactive Hearing Assistants that Isolate Egocentric Conversations ⭐⭐⭐ deep
- № 12 NeuralAids: Wireless Hearables With Programmable Speech AI Accelerators ⭐⭐⭐ deep
- № 13 Creating speech zones with self-distributing acoustic swarms ⭐⭐⭐ deep
- № 14 Conv-TasNet: Surpassing Ideal Time-Frequency Magnitude Masking for Speech Separation ⭐⭐⭐ deep
- № 15 SoundStream: An End-to-End Neural Audio Codec ⭐⭐⭐⭐ deep
- № 16 AudioLM ⭐⭐⭐⭐ deep
- № 17 Conformer ⭐⭐⭐ deep
- № 18 Dual-path RNN ⭐⭐⭐⭐ deep
- № 19 EnCodec ⭐⭐⭐⭐ deep
- № 20 Meta-StyleSpeech ⭐⭐⭐ deep
- № 21 MusicLM ⭐⭐⭐⭐ deep
- № 22 Robust Speech Recognition via Large-Scale Weak Supervision ⭐⭐⭐ deep
- № 23 SeamlessM4T ⭐⭐⭐⭐ deep
- № 24 Stable Audio ⭐⭐⭐⭐ deep
- № 25 Universal Source Separation with Weakly Labelled Data ⭐⭐⭐⭐ deep
数据集与评测基准 · read primer →
- № 26 Meta-World: A Benchmark and Evaluation for Multi-Task and Meta Reinforcement Learning ⭐⭐ deep
- № 27 RLBench: The Robot Learning Benchmark & Learning Environment ⭐⭐ deep
- № 28 robosuite: A Modular Simulation Framework and Benchmark for Robot Learning ⭐⭐ deep
- № 30 CALVIN ⭐⭐⭐ deep
- № 31 LIBERO ⭐⭐⭐ deep
- № 32 RH20T ⭐⭐⭐ deep
- № 33 What Matters in Learning from Offline Human Demonstrations for Robot Manipulation ⭐⭐⭐ deep
- № 34 DROID ⭐⭐⭐ deep
- № 35 Open X-Embodiment ⭐⭐⭐ deep
- № 36 RoboCasa ⭐⭐⭐ deep
- № 106 BridgeData V2 ⭐⭐⭐ deep
- № 157 LeRobot: An Open-Source Library for End-to-End Robot Learning ⭐⭐⭐ deep
- № 161 AutoRT: Embodied Foundation Models for Large Scale Orchestration of Robotic Agents ⭐⭐⭐ deep
- № 164 RealMirror: A Comprehensive, Open-Source Vision-Language-Action Platform for Embodied AI ⭐⭐⭐ deep
- № 184 Open-H-Embodiment: A Large-Scale Dataset for Enabling Foundation Models in Medical Robotics ⭐⭐⭐⭐ deep
- № 196 Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy ⭐⭐⭐⭐ deep
仿真与真实迁移 · read primer →
- № 98 Habitat: A Platform for Embodied AI Research ⭐⭐ deep
- № 99 Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning ⭐⭐⭐ deep
- № 101 Habitat 2.0 ⭐⭐⭐ deep
- № 102 ManiSkill ⭐⭐⭐ deep
- № 103 ProcTHOR ⭐⭐⭐ deep
- № 104 SAPIEN: A SimulAted Part-based Interactive ENvironment ⭐⭐⭐ deep
- № 108 DexMV ⭐⭐⭐⭐ deep
- № 37 SimplerEnv ⭐⭐⭐⭐ deep
- № 105 BEHAVIOR-1K ⭐⭐⭐⭐ deep
- № 106 Habitat 3.0 ⭐⭐⭐ deep
- № 107 Isaac Lab ⭐⭐⭐ deep
- № 108 MuJoCo Playground ⭐⭐⭐ deep
- № 186 3D Generation for Embodied AI and Robotic Simulation: A Survey ⭐⭐⭐⭐ deep