← 论文海报合集← Paper Notes|
综述 · Survey — Embodied Manipulation DataSurvey — Embodied Manipulation Data

Data Pyramid for Embodied Manipulation

把具身操作的数据生态系统整理成一座“数据金字塔”:five complementary sources,串联收集范式与基础模型的数据配方Organizes the embodied-manipulation data ecosystem into a "data pyramid" of five complementary sources, linking collection paradigms with the data recipes of foundation models
Yifan Ye, Yankai Fu, Yaoxu Lv 等 29 位作者 · 香港大学 (HKU)、北京大学 (PKU)、上海交通大学 (SJTU)、香港中文大学 (CUHK)、新加坡国立大学 (NUS) 等多所机构联合完成Yifan Ye, Yankai Fu, Yaoxu Lv et al., 29 authors · Joint work by the University of Hong Kong (HKU), Peking University (PKU), Shanghai Jiao Tong University (SJTU), the Chinese University of Hong Kong (CUHK), the National University of Singapore (NUS) and other institutions

这是一篇聚焦“具身基础模型该用什么数据训练”的综述。作者把现有数据来源组织为跨越 real-robot data、UMI-style data、egocentric/exocentric data、simulation data、general vision-language data 五个互补层级的“数据金字塔”,并沿着 Scalability(可扩展性)与 Robot Alignment(与机器人执行的对齐程度)两条主轴,逐层分析各数据源的 quality、diversity、reusability、physical fidelity;随后梳理这些异构数据如何被 embodied brain models、VLA、world-action models 等基础模型选择、对齐与混合使用,最后提炼出六个尚未解决的开放挑战。This is a survey focused on the question of what data embodied foundation models should be trained on. The authors organize existing data sources into a "data pyramid" spanning five complementary layers — real-robot data, UMI-style data, egocentric/exocentric data, simulation data, and general vision-language data — and, along the two main axes of Scalability and Robot Alignment (how closely the data aligns with robot execution), analyze layer by layer the quality, diversity, reusability, and physical fidelity of each data source; they then review how these heterogeneous data are selected, aligned, and mixed by foundation models such as embodied brain models, VLA, and world-action models, and finally distill six unresolved open challenges.

类型:综述 (Survey)Type: Survey 领域:cs.RO · cs.CVField: cs.RO · cs.CV 五层数据金字塔 + 六大开放挑战Five-layer data pyramid + six open challenges 📄 arXiv:2607.24744 PDF
embodied data pyramid具身操作数据embodied manipulation datareal-robot dataUMI dataegocentric datasimulation datavision-language-actionworld action model数据配方data recipesurvey

01 动机Motivation

多模态基础模型靠在海量视觉-语言数据上做预训练获得了通用能力,但具身智能体还必须理解物理状态与动力学、推理动作如何改变环境、并在真实物理世界中执行行为——这从根本上改变了预训练所需监督的性质。作者把这一问题归结为一个核心问题:Multimodal foundation models acquire general capabilities by pre-training on massive vision-language data, but an embodied agent must further understand physical states and dynamics, reason about how actions change the environment, and execute behaviors in the real physical world — which fundamentally changes the nature of the supervision that pre-training requires. The authors reduce this to one core question:

“What data should be used to train embodied foundation models with such capabilities?”

论文指出,社区已探索了多种互补的监督来源——物理机器人与仿真中采集的 observation-action 轨迹、人类第一/第三人称交互的 egocentric/exocentric 记录、通用图像-视频-语言数据,以及无需机器人参与采集的 UMI-style 演示——但“their respective roles, trade-offs, relationships, and integration strategies remain insufficiently systematized”。已有的 Motus、GR00T 等模型虽然提出过金字塔式或分层的数据视角,却都是“typically designed around the training recipe of a particular model”,缺乏跨类别的系统分析;VLA、World-Action Model、world modeling 相关综述也多聚焦模型架构而非数据本身。为此,作者提出两个组织性问题:The paper notes that the community has explored many complementary sources of supervision — observation-action trajectories collected on physical robots and in simulation, egocentric/exocentric recordings of first- and third-person human interaction, general image-video-language data, and UMI-style demonstrations collected without any robot in the loop — yet "their respective roles, trade-offs, relationships, and integration strategies remain insufficiently systematized". Existing models such as Motus and GR00T have proposed pyramid-like or layered views of data, but they are all "typically designed around the training recipe of a particular model" and lack systematic cross-category analysis; surveys on VLA, World-Action Model, and world modeling likewise focus mostly on model architectures rather than on the data themselves. The authors therefore pose two organizing questions:

“How can heterogeneous embodied data sources be collected, organized, compared, and integrated despite their differences in scalability, robot alignment, physical fidelity, and transferability?”
“How can these heterogeneous data sources be effectively selected, combined, and utilized by embodied foundation models?”
Overview of Organization and Scope
Figure 1 · Overview of Organization and Scope。全文分三部分:先建立涵盖 real-robot / UMI-style / egocentric-exocentric / simulation / general 五类数据的具身数据金字塔,并沿可扩展性与机器人对齐两条轴回顾各类数据集、采集流程、本体、传感配置、监督方式、规模、多样性与可迁移性;再考察这些异构数据源如何支撑具身基础模型;最后讨论具身数据的未来方向。Figure 1 · Overview of Organization and Scope. The paper has three parts: it first establishes an embodied data pyramid covering the five data categories real-robot / UMI-style / egocentric-exocentric / simulation / general, and reviews, along the two axes of scalability and robot alignment, the datasets, collection pipelines, embodiments, sensor configurations, forms of supervision, scale, diversity, and transferability of each category; it then examines how these heterogeneous data sources support embodied foundation models; and finally discusses future directions for embodied data.
5互补数据源层级(real-robot / UMI / ego-exo / simulation / general)complementary data-source layers (real-robot / UMI / ego-exo / simulation / general)
6论文提炼的开放挑战 (open challenges)open challenges distilled by the paper
38,100 hQwen-RobotManip 多源语料时长Duration of the Qwen-RobotManip multi-source corpus
100,000+ hXiaomi-Robotics-1 real-world UMI 轨迹时长Duration of Xiaomi-Robotics-1 real-world UMI trajectories

02 方法:数据金字塔的组织方式Method: How the Data Pyramid Is Organized

金字塔沿两条设计原则组织:Scalability(关乎硬件依赖、人力、环境复位、安全监督、边际生成成本等方面的可扩展效率)与 Robot Alignment(观测、表征与监督信号能多直接地支持在物理机器人上学习与执行)。二者常常相互制约:“data that is closely aligned with real-robot execution is typically expensive and difficult to scale, whereas highly scalable data may provide only indirect or imperfect supervision for physical interaction.” 在此之上,作者进一步引入四个类别级维度刻画各数据源的效用:Quality(有效性、一致性、信息量与任务相关性)、Diversity(任务/物体/场景/视角/指令/本体等覆盖面)、Reusability(能否跨任务、环境、本体、传感系统迁移)、Physical fidelity(对接触、摩擦等真实交互动力学的还原程度)。The pyramid is organized along two design principles: Scalability (the efficiency of scaling with respect to hardware dependence, human labor, environment resets, safety supervision, marginal generation cost, and so on) and Robot Alignment (how directly the observations, representations, and supervision signals support learning and execution on physical robots). The two often constrain each other: "data that is closely aligned with real-robot execution is typically expensive and difficult to scale, whereas highly scalable data may provide only indirect or imperfect supervision for physical interaction." On top of this, the authors introduce four category-level dimensions to characterize the utility of each data source: Quality (validity, consistency, informativeness, and task relevance), Diversity (coverage of tasks/objects/scenes/viewpoints/instructions/embodiments), Reusability (whether it transfers across tasks, environments, embodiments, and sensing systems), and Physical fidelity (how faithfully it reproduces real interaction dynamics such as contact and friction).

五层数据源The Five Data Layers

① Real-Robot Data 位于金字塔顶端,“occupies the apex of the embodied-intelligence data pyramid, representing the layer closest to real-world execution”,直接记录感知-动作-物理结果的闭环,具有最高的质量、物理保真度、多样性与可复用性,但采集成本最高。① Real-Robot Data sits at the top of the pyramid and "occupies the apex of the embodied-intelligence data pyramid, representing the layer closest to real-world execution"; it directly records the closed loop of perception, action, and physical outcome, and offers the highest quality, physical fidelity, diversity, and reusability, but is the most expensive to collect.

② UMI Data(Universal Manipulation Interface)“denotes a family of real-world manipulation datasets collected through portable, robot-independent interfaces rather than directly operating a target robot”,靠便携抓取器/手持接口记录第一人称视觉、末端执行器运动、夹爪状态及部分力/触觉信号,经标定与 retargeting 后可转为可执行的机器人轨迹,实现跨本体部署。② UMI Data (Universal Manipulation Interface) "denotes a family of real-world manipulation datasets collected through portable, robot-independent interfaces rather than directly operating a target robot"; portable grippers / handheld interfaces record first-person vision, end-effector motion, gripper state, and partial force/tactile signals, which after calibration and retargeting can be converted into executable robot trajectories, enabling cross-embodiment deployment.

③ Egocentric & Exocentric Data 处于金字塔中层,“offering more authentic real-world observations and physical interactions than simulation data, but less direct alignment with robot execution than UMI data”;依赖头戴/体戴相机为主,辅以第三人称相机弥补遮挡与视野盲区,并可通过多模态传感(如 RGB-D、gaze、EMG、tactile)、后处理标注(动作分割、手物姿态)与机器人导向的表征转换进一步丰富。③ Egocentric & Exocentric Data occupies the middle of the pyramid, "offering more authentic real-world observations and physical interactions than simulation data, but less direct alignment with robot execution than UMI data"; it relies mainly on head- or body-mounted cameras, complemented by third-person cameras that compensate for occlusion and blind spots, and can be further enriched by multimodal sensing (such as RGB-D, gaze, EMG, tactile), post-hoc annotation (action segmentation, hand-object pose), and robot-oriented representation conversion.

④ Simulation Data 是“a scalable and controllable complement to real-world robot data”:可重复交互条件、可低成本自动获取特权状态与标签,但“compared with the three higher layers of the embodied data pyramid, simulation data generally remains more limited in diversity and physical fidelity”,也存在 sim-to-real gap,需要 digital twin、domain randomization 等手段缓解。④ Simulation Data is "a scalable and controllable complement to real-world robot data": interaction conditions are repeatable, and privileged states and labels can be obtained automatically at low cost; but "compared with the three higher layers of the embodied data pyramid, simulation data generally remains more limited in diversity and physical fidelity", and a sim-to-real gap remains, which must be mitigated by means such as digital twins and domain randomization.

⑤ General Data(通用图像/视频/语言/空间/推理数据)“is characterized by massive scale, broad coverage, and high diversity, but remains weakly aligned with executable robot actions”,其作用不是直接教机器人如何动作,而是为后续动作学习建立感知、语义、空间、时序与规划能力。⑤ General Data (general image / video / language / spatial / reasoning data) "is characterized by massive scale, broad coverage, and high diversity, but remains weakly aligned with executable robot actions"; its role is not to teach the robot how to act directly, but to establish the perceptual, semantic, spatial, temporal, and planning capabilities on which subsequent action learning builds.

Overview of UMI data and real-robot data
Figure 5 · UMI data 与 real-robot data 的采集范式对比:左侧 UMI 用便携式人操接口强调无机器人参与的野外演示、多模态传感与相对末端执行器轨迹;右侧 real-robot data 在真实本体上用多样遥操作系统与传感配置采集;二者通过 retargeting 连接,把 UMI 演示转为可跨本体部署的机器人动作。Figure 5 · Comparison of the collection paradigms of UMI data and real-robot data: on the left, UMI uses portable human-operated interfaces and emphasizes in-the-wild demonstrations without a robot, multimodal sensing, and relative end-effector trajectories; on the right, real-robot data is collected on real embodiments with diverse teleoperation systems and sensor configurations; the two are connected by retargeting, which turns UMI demonstrations into robot actions deployable across embodiments.
Overview of General Data Categories
Figure 8 · General Data 类别总览:涵盖认知、任务推理、感知、动作四个方向的代表性任务,如 vision-language QA、OCR、视频时序推理、物理推理、规划、分割与定位、空间与 3D 感知、抓取等。Figure 8 · Overview of General Data categories: representative tasks along the four directions of cognition, task reasoning, perception, and action, such as vision-language QA, OCR, video temporal reasoning, physical reasoning, planning, segmentation and grounding, spatial and 3D perception, and grasping.

03 应用与趋势分析Applications and Trend Analysis

这篇综述没有传统意义上的 benchmark 实验,取而代之的是第 7 章对具身基础模型数据配方的系统梳理(Table 7 汇总了 2023.3–2026.7 间发布的代表性 VLA / World-Action Model),并总结出三条趋势。This survey has no benchmark experiments in the traditional sense; in their place, Chapter 7 systematically reviews the data recipes of embodied foundation models (Table 7 summarizes representative VLA / World-Action Models released between 2023.3 and 2026.7) and summarizes three trends.

趋势Trend论文给出的证据 / 数字Evidence / numbers given in the paper
数据配方从单一来源转向异构混合Data recipes shift from a single source to heterogeneous mixturesπ 系列从 π₀(仅 real-robot data)→ π₀.₅(加入 general-purpose data)→ π₀.₇(进一步引入 egocentric data);LingbotVA(real-robot + UMI + simulation)→ LingbotVA 2.0(扩展到全部五层)The π series goes from π₀ (real-robot data only) → π₀.₅ (adding general-purpose data) → π₀.₇ (further introducing egocentric data); LingbotVA (real-robot + UMI + simulation) → LingbotVA 2.0 (extended to all five layers)
预训练规模快速增长Pre-training scale grows rapidlyQwen-RobotManip 构建约 38,100 小时多源语料,含约 11.4K 小时 开源机器人数据与由 1,933 小时 egocentric video 合成的 24,808 小时 机器人兼容轨迹;Xiaomi-Robotics-1 在超过 100,000 小时 real-world UMI 轨迹上预训练,随后用约 10,000 小时 跨本体后训练数据(含 7,200+ 小时自采机器人轨迹与 1,000+ 小时带指令标注的 UMI 数据)Qwen-RobotManip builds a multi-source corpus of about 38,100 hours, containing about 11.4K hours of open-source robot data and 24,808 hours of robot-compatible trajectories synthesized from 1,933 hours of egocentric video; Xiaomi-Robotics-1 is pre-trained on more than 100,000 hours of real-world UMI trajectories, followed by about 10,000 hours of cross-embodiment post-training data (including 7,200+ hours of self-collected robot trajectories and 1,000+ hours of instruction-annotated UMI data)
egocentric human data 重要性上升egocentric human data grows in importanceHumanScale 用最多 5,000 小时 精选 egocentric video 做受控预训练;EgoScale 把带动作标注的 egocentric 预训练扩展到 20,854 小时,并在 1,000–20,000 小时区间内报告了持续的性能提升HumanScale uses at most 5,000 hours of curated egocentric video for controlled pre-training; EgoScale scales action-annotated egocentric pre-training up to 20,854 hours and reports continued performance gains over the 1,000–20,000 hour range
Evolution of Data Utilization
Figure 3 · 代表性具身基础模型的数据使用演化:早期系统主要依赖 real-robot 轨迹,近期系统越来越多地探索与 egocentric、simulation、general、UMI-style 数据的联合训练;模型架构也从以动作为中心的 VLA 系统扩展到 world-model 增强及统一的 world-action model。Figure 3 · Evolution of data utilization in representative embodied foundation models: early systems relied mainly on real-robot trajectories, while recent systems increasingly explore joint training with egocentric, simulation, general, and UMI-style data; model architectures have also expanded from action-centric VLA systems to world-model-augmented and unified world-action models.

数据规模的整体增长Overall Growth of Data Scale

Evolution of Data Scale
Figure 2 · 金字塔各数据源的规模演化:simulation / UMI / robot 数据集按演示数量计,ego 数据集按采集小时数计,general 数据集按问答对数量计;插入曲线展示了对应类别内代表性数据集的规模增长轨迹。Figure 2 · Evolution of scale for each data source in the pyramid: simulation / UMI / robot datasets are measured by number of demonstrations, ego datasets by hours collected, and general datasets by number of QA pairs; the inset curves show the growth trajectories of representative datasets within each category.

论文的结论What the Paper Concludes

作者观察到“current evidence does not establish that incorporating more data sources is inherently preferable”——纯 robot data 预训练的 LingbotVLA、DreamZero 也能取得强性能;但不同数据类型仍能提供互补能力:general-purpose data 支持广泛的视觉/语义/推理能力,egocentric 与 UMI data 扩大场景/物体/视角/人类行为的覆盖面,simulation data 通过可控且可扩展的交互生成强化特定技能。The authors observe that "current evidence does not establish that incorporating more data sources is inherently preferable" — LingbotVLA and DreamZero, pre-trained on pure robot data, also achieve strong performance; yet different data types still provide complementary capabilities: general-purpose data supports broad visual / semantic / reasoning abilities, egocentric and UMI data broaden the coverage of scenes / objects / viewpoints / human behaviors, and simulation data strengthens specific skills through controllable and scalable interaction generation.

04 开放挑战(论文第 8 章,作者明确列出)Open Challenges (Chapter 8 of the paper, explicitly listed by the authors)

说明:以下六点均为论文第 8 章“Challenges and Future Directions”中作者明确提出的开放挑战/未来方向,非本文推测。作者把它们组织为三个问题:what to collect、how to collect it、how to use it。Note: All six points below are open challenges / future directions explicitly put forward by the authors in Chapter 8 of the paper, "Challenges and Future Directions", not speculation by this poster. The authors organize them around three questions: what to collect, how to collect it, and how to use it.
1 · Tactile Data for Contact-Rich Robot Learning

现有数据资源仍以 RGB-D 观测、本体感受状态、语言指令与动作轨迹为主,“tactile sensing is still not systematically incorporated as a standard modality”,在数据金字塔中留下了“a missing contact layer”——视觉观测只能间接反映接触力、滑动、局部形变、摩擦与抓取稳定性。Existing data resources are still dominated by RGB-D observations, proprioceptive states, language instructions, and action trajectories; "tactile sensing is still not systematically incorporated as a standard modality", leaving "a missing contact layer" in the data pyramid — visual observations can only indirectly reflect contact force, slippage, local deformation, friction, and grasp stability.

2 · Failure Data and Recovery-Centric Robot Learning

现有机器人学习数据集“remain strongly biased toward successful expert demonstrations”,失败、临界失败与次优轨迹常被丢弃或仅粗略标注,导致策略缺乏对失败感知、原因诊断与恢复行为的监督。作者呼吁未来数据集提供更丰富的失败中心标注(失败前上下文、失败起点、类别与原因、状态变化、恢复动作与结果)。Existing robot learning datasets "remain strongly biased toward successful expert demonstrations"; failures, near-failures, and suboptimal trajectories are often discarded or only coarsely annotated, leaving policies without supervision for failure awareness, cause diagnosis, and recovery behavior. The authors call for future datasets to provide richer failure-centric annotations (pre-failure context, failure onset, category and cause, state changes, recovery actions and outcomes).

3 · Scalable Data Collection Across Pyramid Layers

可穿戴采集设备(头戴 AR/MR、腕部相机、IMU、数据手套、触觉贴片、EMG 等)尚未成熟到可大规模部署:头/手跟踪易受遮挡、运动模糊、漂移与视野限制影响;手套/触觉/EMG 传感器需要精细佩戴、标定与同步,可能干扰自然操作。作者呼吁未来系统更轻量、无线、模块化、少侵入,并支持自动跨设备标定。Wearable collection devices (head-mounted AR/MR, wrist cameras, IMU, data gloves, tactile patches, EMG, and so on) are not yet mature enough for large-scale deployment: head/hand tracking is easily affected by occlusion, motion blur, drift, and limited field of view; glove / tactile / EMG sensors require careful donning, calibration, and synchronization, and may interfere with natural manipulation. The authors call for future systems that are lighter, wireless, modular, and less intrusive, and that support automatic cross-device calibration.

4 · Cross-Embodiment State-Action Alignment

即便统一采用笛卡尔末端执行器位姿空间以聚合跨本体轨迹,末端位姿仍可能相对机器人基座、相机、末端局部系或世界坐标系定义,“causing the same physical motion to correspond to different numerical representations across datasets”,联合训练时可能引入冲突监督,显著影响跨本体迁移与数据扩展行为。Even when a Cartesian end-effector pose space is uniformly adopted to aggregate cross-embodiment trajectories, the end-effector pose may still be defined relative to the robot base, the camera, a local end-effector frame, or the world frame, "causing the same physical motion to correspond to different numerical representations across datasets", which can introduce conflicting supervision during joint training and significantly affect cross-embodiment transfer and data scaling behavior.

5 · Egocentric Priors for Dexterous Hand Policy Learning

人类与机器人手在关节拓扑、自由度、指尖几何、顺应性、驱动、传感、摩擦与力限制上存在差异,“the human-to-robot gap is not merely visual, but also kinematic, morphological, and physical”,几何上看似合理的 retargeted 动作在真实硬件上仍可能违反本体约束或无法产生稳定接触。Human and robot hands differ in joint topology, degrees of freedom, fingertip geometry, compliance, actuation, sensing, friction, and force limits; "the human-to-robot gap is not merely visual, but also kinematic, morphological, and physical", so retargeted motions that look geometrically plausible may still violate embodiment constraints or fail to produce stable contact on real hardware.

6 · Data Recipes(原则性数据配方设计)6 · Data Recipes (Principled Data Recipe Design)

随着数据配方从单一 robot demonstration 走向日益异构的混合,“the optimal data recipe for embodied foundation models remains an open question”——目前证据尚不能证明纳入更多数据源天然更优,如何为不同能力(感知、推理、规划、动作生成、预测)设计原则性的数据组合策略仍待系统研究。As data recipes move from single robot demonstrations toward increasingly heterogeneous mixtures, "the optimal data recipe for embodied foundation models remains an open question" — current evidence cannot yet establish that incorporating more data sources is inherently better, and how to design principled data-combination strategies for different capabilities (perception, reasoning, planning, action generation, prediction) still awaits systematic study.