这是一篇聚焦“具身基础模型该用什么数据训练”的综述。作者把现有数据来源组织为跨越 real-robot data、UMI-style data、egocentric/exocentric data、simulation data、general vision-language data 五个互补层级的“数据金字塔”,并沿着 Scalability(可扩展性)与 Robot Alignment(与机器人执行的对齐程度)两条主轴,逐层分析各数据源的 quality、diversity、reusability、physical fidelity;随后梳理这些异构数据如何被 embodied brain models、VLA、world-action models 等基础模型选择、对齐与混合使用,最后提炼出六个尚未解决的开放挑战。This is a survey focused on the question of what data embodied foundation models should be trained on. The authors organize existing data sources into a "data pyramid" spanning five complementary layers — real-robot data, UMI-style data, egocentric/exocentric data, simulation data, and general vision-language data — and, along the two main axes of Scalability and Robot Alignment (how closely the data aligns with robot execution), analyze layer by layer the quality, diversity, reusability, and physical fidelity of each data source; they then review how these heterogeneous data are selected, aligned, and mixed by foundation models such as embodied brain models, VLA, and world-action models, and finally distill six unresolved open challenges.
多模态基础模型靠在海量视觉-语言数据上做预训练获得了通用能力,但具身智能体还必须理解物理状态与动力学、推理动作如何改变环境、并在真实物理世界中执行行为——这从根本上改变了预训练所需监督的性质。作者把这一问题归结为一个核心问题:Multimodal foundation models acquire general capabilities by pre-training on massive vision-language data, but an embodied agent must further understand physical states and dynamics, reason about how actions change the environment, and execute behaviors in the real physical world — which fundamentally changes the nature of the supervision that pre-training requires. The authors reduce this to one core question:
“What data should be used to train embodied foundation models with such capabilities?”
论文指出,社区已探索了多种互补的监督来源——物理机器人与仿真中采集的 observation-action 轨迹、人类第一/第三人称交互的 egocentric/exocentric 记录、通用图像-视频-语言数据,以及无需机器人参与采集的 UMI-style 演示——但“their respective roles, trade-offs, relationships, and integration strategies remain insufficiently systematized”。已有的 Motus、GR00T 等模型虽然提出过金字塔式或分层的数据视角,却都是“typically designed around the training recipe of a particular model”,缺乏跨类别的系统分析;VLA、World-Action Model、world modeling 相关综述也多聚焦模型架构而非数据本身。为此,作者提出两个组织性问题:The paper notes that the community has explored many complementary sources of supervision — observation-action trajectories collected on physical robots and in simulation, egocentric/exocentric recordings of first- and third-person human interaction, general image-video-language data, and UMI-style demonstrations collected without any robot in the loop — yet "their respective roles, trade-offs, relationships, and integration strategies remain insufficiently systematized". Existing models such as Motus and GR00T have proposed pyramid-like or layered views of data, but they are all "typically designed around the training recipe of a particular model" and lack systematic cross-category analysis; surveys on VLA, World-Action Model, and world modeling likewise focus mostly on model architectures rather than on the data themselves. The authors therefore pose two organizing questions:
“How can heterogeneous embodied data sources be collected, organized, compared, and integrated despite their differences in scalability, robot alignment, physical fidelity, and transferability?”
“How can these heterogeneous data sources be effectively selected, combined, and utilized by embodied foundation models?”
金字塔沿两条设计原则组织:Scalability(关乎硬件依赖、人力、环境复位、安全监督、边际生成成本等方面的可扩展效率)与 Robot Alignment(观测、表征与监督信号能多直接地支持在物理机器人上学习与执行)。二者常常相互制约:“data that is closely aligned with real-robot execution is typically expensive and difficult to scale, whereas highly scalable data may provide only indirect or imperfect supervision for physical interaction.” 在此之上,作者进一步引入四个类别级维度刻画各数据源的效用:Quality(有效性、一致性、信息量与任务相关性)、Diversity(任务/物体/场景/视角/指令/本体等覆盖面)、Reusability(能否跨任务、环境、本体、传感系统迁移)、Physical fidelity(对接触、摩擦等真实交互动力学的还原程度)。The pyramid is organized along two design principles: Scalability (the efficiency of scaling with respect to hardware dependence, human labor, environment resets, safety supervision, marginal generation cost, and so on) and Robot Alignment (how directly the observations, representations, and supervision signals support learning and execution on physical robots). The two often constrain each other: "data that is closely aligned with real-robot execution is typically expensive and difficult to scale, whereas highly scalable data may provide only indirect or imperfect supervision for physical interaction." On top of this, the authors introduce four category-level dimensions to characterize the utility of each data source: Quality (validity, consistency, informativeness, and task relevance), Diversity (coverage of tasks/objects/scenes/viewpoints/instructions/embodiments), Reusability (whether it transfers across tasks, environments, embodiments, and sensing systems), and Physical fidelity (how faithfully it reproduces real interaction dynamics such as contact and friction).
① Real-Robot Data 位于金字塔顶端,“occupies the apex of the embodied-intelligence data pyramid, representing the layer closest to real-world execution”,直接记录感知-动作-物理结果的闭环,具有最高的质量、物理保真度、多样性与可复用性,但采集成本最高。① Real-Robot Data sits at the top of the pyramid and "occupies the apex of the embodied-intelligence data pyramid, representing the layer closest to real-world execution"; it directly records the closed loop of perception, action, and physical outcome, and offers the highest quality, physical fidelity, diversity, and reusability, but is the most expensive to collect.
② UMI Data(Universal Manipulation Interface)“denotes a family of real-world manipulation datasets collected through portable, robot-independent interfaces rather than directly operating a target robot”,靠便携抓取器/手持接口记录第一人称视觉、末端执行器运动、夹爪状态及部分力/触觉信号,经标定与 retargeting 后可转为可执行的机器人轨迹,实现跨本体部署。② UMI Data (Universal Manipulation Interface) "denotes a family of real-world manipulation datasets collected through portable, robot-independent interfaces rather than directly operating a target robot"; portable grippers / handheld interfaces record first-person vision, end-effector motion, gripper state, and partial force/tactile signals, which after calibration and retargeting can be converted into executable robot trajectories, enabling cross-embodiment deployment.
③ Egocentric & Exocentric Data 处于金字塔中层,“offering more authentic real-world observations and physical interactions than simulation data, but less direct alignment with robot execution than UMI data”;依赖头戴/体戴相机为主,辅以第三人称相机弥补遮挡与视野盲区,并可通过多模态传感(如 RGB-D、gaze、EMG、tactile)、后处理标注(动作分割、手物姿态)与机器人导向的表征转换进一步丰富。③ Egocentric & Exocentric Data occupies the middle of the pyramid, "offering more authentic real-world observations and physical interactions than simulation data, but less direct alignment with robot execution than UMI data"; it relies mainly on head- or body-mounted cameras, complemented by third-person cameras that compensate for occlusion and blind spots, and can be further enriched by multimodal sensing (such as RGB-D, gaze, EMG, tactile), post-hoc annotation (action segmentation, hand-object pose), and robot-oriented representation conversion.
④ Simulation Data 是“a scalable and controllable complement to real-world robot data”:可重复交互条件、可低成本自动获取特权状态与标签,但“compared with the three higher layers of the embodied data pyramid, simulation data generally remains more limited in diversity and physical fidelity”,也存在 sim-to-real gap,需要 digital twin、domain randomization 等手段缓解。④ Simulation Data is "a scalable and controllable complement to real-world robot data": interaction conditions are repeatable, and privileged states and labels can be obtained automatically at low cost; but "compared with the three higher layers of the embodied data pyramid, simulation data generally remains more limited in diversity and physical fidelity", and a sim-to-real gap remains, which must be mitigated by means such as digital twins and domain randomization.
⑤ General Data(通用图像/视频/语言/空间/推理数据)“is characterized by massive scale, broad coverage, and high diversity, but remains weakly aligned with executable robot actions”,其作用不是直接教机器人如何动作,而是为后续动作学习建立感知、语义、空间、时序与规划能力。⑤ General Data (general image / video / language / spatial / reasoning data) "is characterized by massive scale, broad coverage, and high diversity, but remains weakly aligned with executable robot actions"; its role is not to teach the robot how to act directly, but to establish the perceptual, semantic, spatial, temporal, and planning capabilities on which subsequent action learning builds.
这篇综述没有传统意义上的 benchmark 实验,取而代之的是第 7 章对具身基础模型数据配方的系统梳理(Table 7 汇总了 2023.3–2026.7 间发布的代表性 VLA / World-Action Model),并总结出三条趋势。This survey has no benchmark experiments in the traditional sense; in their place, Chapter 7 systematically reviews the data recipes of embodied foundation models (Table 7 summarizes representative VLA / World-Action Models released between 2023.3 and 2026.7) and summarizes three trends.
| 趋势Trend | 论文给出的证据 / 数字Evidence / numbers given in the paper |
|---|---|
| 数据配方从单一来源转向异构混合Data recipes shift from a single source to heterogeneous mixtures | π 系列从 π₀(仅 real-robot data)→ π₀.₅(加入 general-purpose data)→ π₀.₇(进一步引入 egocentric data);LingbotVA(real-robot + UMI + simulation)→ LingbotVA 2.0(扩展到全部五层)The π series goes from π₀ (real-robot data only) → π₀.₅ (adding general-purpose data) → π₀.₇ (further introducing egocentric data); LingbotVA (real-robot + UMI + simulation) → LingbotVA 2.0 (extended to all five layers) |
| 预训练规模快速增长Pre-training scale grows rapidly | Qwen-RobotManip 构建约 38,100 小时多源语料,含约 11.4K 小时 开源机器人数据与由 1,933 小时 egocentric video 合成的 24,808 小时 机器人兼容轨迹;Xiaomi-Robotics-1 在超过 100,000 小时 real-world UMI 轨迹上预训练,随后用约 10,000 小时 跨本体后训练数据(含 7,200+ 小时自采机器人轨迹与 1,000+ 小时带指令标注的 UMI 数据)Qwen-RobotManip builds a multi-source corpus of about 38,100 hours, containing about 11.4K hours of open-source robot data and 24,808 hours of robot-compatible trajectories synthesized from 1,933 hours of egocentric video; Xiaomi-Robotics-1 is pre-trained on more than 100,000 hours of real-world UMI trajectories, followed by about 10,000 hours of cross-embodiment post-training data (including 7,200+ hours of self-collected robot trajectories and 1,000+ hours of instruction-annotated UMI data) |
| egocentric human data 重要性上升egocentric human data grows in importance | HumanScale 用最多 5,000 小时 精选 egocentric video 做受控预训练;EgoScale 把带动作标注的 egocentric 预训练扩展到 20,854 小时,并在 1,000–20,000 小时区间内报告了持续的性能提升HumanScale uses at most 5,000 hours of curated egocentric video for controlled pre-training; EgoScale scales action-annotated egocentric pre-training up to 20,854 hours and reports continued performance gains over the 1,000–20,000 hour range |


作者观察到“current evidence does not establish that incorporating more data sources is inherently preferable”——纯 robot data 预训练的 LingbotVLA、DreamZero 也能取得强性能;但不同数据类型仍能提供互补能力:general-purpose data 支持广泛的视觉/语义/推理能力,egocentric 与 UMI data 扩大场景/物体/视角/人类行为的覆盖面,simulation data 通过可控且可扩展的交互生成强化特定技能。The authors observe that "current evidence does not establish that incorporating more data sources is inherently preferable" — LingbotVLA and DreamZero, pre-trained on pure robot data, also achieve strong performance; yet different data types still provide complementary capabilities: general-purpose data supports broad visual / semantic / reasoning abilities, egocentric and UMI data broaden the coverage of scenes / objects / viewpoints / human behaviors, and simulation data strengthens specific skills through controllable and scalable interaction generation.
现有数据资源仍以 RGB-D 观测、本体感受状态、语言指令与动作轨迹为主,“tactile sensing is still not systematically incorporated as a standard modality”,在数据金字塔中留下了“a missing contact layer”——视觉观测只能间接反映接触力、滑动、局部形变、摩擦与抓取稳定性。Existing data resources are still dominated by RGB-D observations, proprioceptive states, language instructions, and action trajectories; "tactile sensing is still not systematically incorporated as a standard modality", leaving "a missing contact layer" in the data pyramid — visual observations can only indirectly reflect contact force, slippage, local deformation, friction, and grasp stability.
现有机器人学习数据集“remain strongly biased toward successful expert demonstrations”,失败、临界失败与次优轨迹常被丢弃或仅粗略标注,导致策略缺乏对失败感知、原因诊断与恢复行为的监督。作者呼吁未来数据集提供更丰富的失败中心标注(失败前上下文、失败起点、类别与原因、状态变化、恢复动作与结果)。Existing robot learning datasets "remain strongly biased toward successful expert demonstrations"; failures, near-failures, and suboptimal trajectories are often discarded or only coarsely annotated, leaving policies without supervision for failure awareness, cause diagnosis, and recovery behavior. The authors call for future datasets to provide richer failure-centric annotations (pre-failure context, failure onset, category and cause, state changes, recovery actions and outcomes).
可穿戴采集设备(头戴 AR/MR、腕部相机、IMU、数据手套、触觉贴片、EMG 等)尚未成熟到可大规模部署:头/手跟踪易受遮挡、运动模糊、漂移与视野限制影响;手套/触觉/EMG 传感器需要精细佩戴、标定与同步,可能干扰自然操作。作者呼吁未来系统更轻量、无线、模块化、少侵入,并支持自动跨设备标定。Wearable collection devices (head-mounted AR/MR, wrist cameras, IMU, data gloves, tactile patches, EMG, and so on) are not yet mature enough for large-scale deployment: head/hand tracking is easily affected by occlusion, motion blur, drift, and limited field of view; glove / tactile / EMG sensors require careful donning, calibration, and synchronization, and may interfere with natural manipulation. The authors call for future systems that are lighter, wireless, modular, and less intrusive, and that support automatic cross-device calibration.
即便统一采用笛卡尔末端执行器位姿空间以聚合跨本体轨迹,末端位姿仍可能相对机器人基座、相机、末端局部系或世界坐标系定义,“causing the same physical motion to correspond to different numerical representations across datasets”,联合训练时可能引入冲突监督,显著影响跨本体迁移与数据扩展行为。Even when a Cartesian end-effector pose space is uniformly adopted to aggregate cross-embodiment trajectories, the end-effector pose may still be defined relative to the robot base, the camera, a local end-effector frame, or the world frame, "causing the same physical motion to correspond to different numerical representations across datasets", which can introduce conflicting supervision during joint training and significantly affect cross-embodiment transfer and data scaling behavior.
人类与机器人手在关节拓扑、自由度、指尖几何、顺应性、驱动、传感、摩擦与力限制上存在差异,“the human-to-robot gap is not merely visual, but also kinematic, morphological, and physical”,几何上看似合理的 retargeted 动作在真实硬件上仍可能违反本体约束或无法产生稳定接触。Human and robot hands differ in joint topology, degrees of freedom, fingertip geometry, compliance, actuation, sensing, friction, and force limits; "the human-to-robot gap is not merely visual, but also kinematic, morphological, and physical", so retargeted motions that look geometrically plausible may still violate embodiment constraints or fail to produce stable contact on real hardware.
随着数据配方从单一 robot demonstration 走向日益异构的混合,“the optimal data recipe for embodied foundation models remains an open question”——目前证据尚不能证明纳入更多数据源天然更优,如何为不同能力(感知、推理、规划、动作生成、预测)设计原则性的数据组合策略仍待系统研究。As data recipes move from single robot demonstrations toward increasingly heterogeneous mixtures, "the optimal data recipe for embodied foundation models remains an open question" — current evidence cannot yet establish that incorporating more data sources is inherently better, and how to design principled data-combination strategies for different capabilities (perception, reasoning, planning, action generation, prediction) still awaits systematic study.