← 论文海报合集← Paper Notes|
机器人 · Robotics · ICRA 2024Robotics · ICRA 2024

RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot

首个大规模多模态 contact-rich 机器人操控数据集,支持单样本技能泛化The first large-scale multi-modal contact-rich robotic manipulation dataset, enabling one-shot skill generalization
Hao-Shu Fang, Hongjie Fang, Zhenyu Tang, Jirong Liu, Chenxi Wang, Junbo Wang, Haoyi Zhu, Cewu Lu  ·  上海交通大学Shanghai Jiao Tong University

RH20T(Robot-Human Demonstration in 20TB)是一个包含超过 110,000 条真实世界机器人操控序列的大规模数据集。每条序列同步采集 RGB-D 视觉、6DoF 力/力矩、音频及本体感知信息,并附有对应的人类示范视频与语言描述,覆盖 147 项任务 / 42 种技能,横跨 7 种机器人配置。其设计目标是推动机器人在开放域中实现单样本(one-shot)技能迁移与多模态感知。RH20T (Robot-Human Demonstration in 20TB) is a large-scale dataset containing more than 110,000 real-world robotic manipulation sequences. Each sequence synchronously records RGB-D vision, 6DoF force/torque, audio and proprioceptive information, and is paired with a corresponding human demonstration video and a language description, covering 147 tasks / 42 skills across 7 robot configurations. It is designed to advance one-shot skill transfer and multi-modal perception for robots in open domains.

ICRA 2024 上海交通大学Shanghai Jiao Tong University 2023-07-02 📄 arXiv:2307.00595 🌐 Project Page: rh20t.github.io
机器人操控数据集robotic manipulation dataset one-shot imitation learning multi-modal perception contact-rich manipulation force-torque sensing 机器人技能泛化robot skill generalization 遥操作数据采集teleoperated data collection 多视角标定multi-view calibration

01 动机 MotivationMotivation

现有机器人操控研究主要停留在推送(push)、抓取(pick-place)等简单任务,且几乎仅依赖视觉(visual-only)反馈。然而现实中许多操控技能高度依赖力觉与触觉——如切割、旋转、插拔等 contact-rich 动作。造成这一差距的核心瓶颈有两个:(1)缺乏大规模多样化的真实机器人数据集(2)现有方法忽略了视觉之外的多模态感知Research on robotic manipulation so far has largely remained at simple tasks such as pushing (push) and grasping (pick-place), and relies almost exclusively on visual-only feedback. In reality, however, many manipulation skills depend heavily on force and tactile sensing, as in contact-rich actions such as cutting, rotating, plugging and unplugging. Two bottlenecks account for this gap: (1) the lack of large-scale, diverse real-world robot datasets; (2) existing methods ignore modalities beyond vision.

"In reality, there are many complex skills, some of which may even require both visual and tactile perception to solve. This paper aims to unlock the potential for an agent to generalize to hundreds of real-world skills with multi-modal perception."
RH20T数据集总览
图1:RH20T 数据集总览。使用多种机器人臂与多样化环境配置采集数据。每条机器人操控序列包含多模态视觉、力觉、音频和动作数据,并通过标定的多视角相机记录。数据集涵盖多样化操控技能,每条序列配有对应的人类示范视频与语言描述。共提供超过 110K 条机器人序列和 110K 条人类示范序列,包含超过 5000 万帧图像及 140+ 项任务。Figure 1: Overview of the RH20T dataset. Data are collected with a variety of robot arms under diverse environment configurations. Each robot manipulation sequence contains multi-modal vision, force, audio and action data, recorded by calibrated multi-view cameras. The dataset covers diverse manipulation skills, and every sequence is paired with a corresponding human demonstration video and language description. In total it provides more than 110K robot sequences and 110K human demonstration sequences, comprising over 50000000 image frames and 140+ tasks.
110K+机器人操控序列robot manipulation sequences
110K+对应人类示范视频paired human demonstration videos
50M+总图像帧数image frames in total
147任务 / 42 种技能tasks / 42 skills

与同类公开数据集的对比凸显了 RH20T 的全面性:MIME(8.3K 条)、RoboTurk(2.1K 条)、RoboNet(162K 条,但多为随机游走)、BC-Z(60.1K 条,技能单一)均无法同时覆盖多机器人、力觉感知、相机标定与人类示范等维度。RH20T 是目前社区中规模最大、模态最丰富的真实世界机器人操控数据集。A comparison with comparable public datasets highlights the comprehensiveness of RH20T: MIME (8.3K sequences), RoboTurk (2.1K sequences), RoboNet (162K sequences, but mostly random walks) and BC-Z (60.1K sequences, with a single skill) all fail to cover multiple robots, force sensing, camera calibration and human demonstrations at the same time. RH20T is currently the largest and most modality-rich real-world robotic manipulation dataset in the community.

02 方法 Method(数据集构建)Method (Dataset Construction)

RH20T 的构建核心在于:设计直觉高效的力反馈遥操作平台,建立多模态多视角同步采集流水线,以及完善的数据层级结构(hierarchy),以支持密集的 <human demo, robot manipulation> 配对。The construction of RH20T rests on an intuitive and efficient force-feedback teleoperation platform, a multi-modal multi-view synchronized collection pipeline, and a carefully designed data hierarchy that supports dense <human demo, robot manipulation> pairing.

数据集规模对比与硬件配置
表1(上):与同类数据集对比。RH20T 在序列数量(110K)、机器人种类(12 种)、模态丰富度(RGB-D、力矩、音频、本体感知)、相机外参标定和人类示范等方面全面领先。Table 1 (top): Comparison with comparable datasets. RH20T leads across the board in number of sequences (110K), robot types (12), modality richness (RGB-D, force/torque, audio, proprioception), camera extrinsic calibration and human demonstrations.
表2(下):硬件配置详情。7 种机器人配置涵盖 Flexiv、UR5、Franka、Kuka 等主流机械臂,搭配 ATI、OptoForce 等力矩传感器;表3 给出各模态采样频率(RGB 10Hz、关节力矩 100Hz、触觉 200Hz 等)。Table 2 (bottom): Hardware configuration details. The 7 robot configurations cover mainstream arms such as Flexiv, UR5, Franka and Kuka, paired with force/torque sensors such as ATI and OptoForce; Table 3 lists the sampling rate of each modality (RGB 10Hz, joint torque 100Hz, tactile 200Hz, etc.).

多模态同步采集平台Multi-modal Synchronized Collection Platform

每套采集平台由机械臂(含力矩传感器与夹爪)、8-10 个全局 RGB-D 相机、2 个麦克风、1 个 haptic 设备(提供力反馈)和踏板组成。所有相机在采集前完成外参标定,数据通过时间戳对齐同步保存。人类示范在同一平台上由操作者佩戴第一视角相机完成。平均培训时间不足 1 小时,成功/失败比例约为 10:1。Each collection platform consists of a robot arm (with force/torque sensor and gripper), 8-10 global RGB-D cameras, 2 microphones, 1 haptic device (providing force feedback) and a foot pedal. All cameras are extrinsically calibrated before collection, and the data are saved synchronously with timestamp alignment. Human demonstrations are performed on the same platform by operators wearing a first-person camera. Average operator training takes less than 1 hour, and the success/failure ratio is about 10:1.

关键创新是引入 haptic device 力反馈遥操作代替传统 3D 鼠标或 VR 遥控——后者在 contact-rich 任务中容易引起碰撞和紧急停止。力反馈使操作者能够精确感知接触力,显著提升了 contact-rich 技能(切割、插拔、折叠等)的数据质量。The key innovation is the adoption of force-feedback teleoperation with a haptic device in place of a conventional 3D mouse or VR controller, which in contact-rich tasks easily causes collisions and emergency stops. Force feedback lets the operator perceive contact forces precisely, markedly improving the data quality of contact-rich skills such as cutting, plugging/unplugging and folding.

数据层级结构(Data Hierarchy)Data Hierarchy

RH20T 按照任务内相似度将数据组织成树状层级。叶节点为具体的人类示范(human demo)与机器人操控序列,共同祖先越近则相关性越强。这种层级设计支持为每条机器人序列配对来自不同视角、场景、操作者的多条人类示范,仅一项任务即可构建出数百万条 <human demo, robot manipulation> 配对样本。RH20T organizes the data into a tree-shaped hierarchy according to intra-task similarity. The leaf nodes are individual human demonstrations (human demo) and robot manipulation sequences; the closer their common ancestor, the stronger their correlation. This hierarchical design makes it possible to pair every robot sequence with multiple human demonstrations from different viewpoints, scenes and operators, so that a single task alone yields millions of <human demo, robot manipulation> pairs.

数据多样性设计Design for Data Diversity

数据采集平台与数据统计
图5:数据采集平台示意图。平台配置包括机械臂(含力矩传感器)、手内相机、8-10 个全局相机、麦克风、haptic device 和踏板。图6(右上):多视角 RGBD 融合点云示意——红色锥体表示相机位姿,机器人模型根据关节角度实时渲染,证明所有相机已相对机器人基座坐标系完成标定,且数据在时间域对齐。Figure 5: Schematic of the data collection platform. The platform comprises a robot arm (with force/torque sensor), an in-hand camera, 8-10 global cameras, microphones, a haptic device and a foot pedal. Figure 6 (top right): illustration of the fused multi-view RGBD point cloud, where red cones denote camera poses and the robot model is rendered in real time from joint angles, showing that all cameras have been calibrated with respect to the robot base frame and that the data are aligned in time.

03 实验 ExperimentsExperiments

论文以 ACT(Action Chunking with Transformers)作为 baseline,在真实机器人平台上验证 RH20T 数据集对迁移学习与少样本(few-shot)学习能力的提升效果。实验任务为"抓取方块并放置在砝码上",在与 RH20T 不同摄像头视角、桌布纹理和机器人配置的新环境中评估。Taking ACT (Action Chunking with Transformers) as the baseline, the paper verifies on a real robot platform how much the RH20T dataset improves transfer learning and few-shot learning. The evaluation task is "grasp a block and place it on a weight", assessed in a new environment whose camera viewpoint, tablecloth texture and robot configuration all differ from those in RH20T.

实验配置Experimental Setup

实验结果表格
表4 & 表5:ACT 在不同训练设置下的成功率(%)。上表为原环境评估(20 次),下表为跨环境泛化评估(新物体/桌布,10 次)。"Pretrain Task" 列标注是否使用 RH20T 同任务或跨任务预训练。Tables 4 & 5: Success rate (%) of ACT under different training settings. The upper table reports evaluation in the original environment (20 trials), the lower one cross-environment generalization (new objects/tablecloths, 10 trials). The "Pretrain Task" column indicates whether same-task or cross-task pretraining on RH20T is used.

关键结论Key Findings

实验平台与泛化评估
图7:实验平台与泛化测试物体。(a) 实验平台(Flexiv 臂 + RealSense);(b) 不同砝码(金属、粉色)评估物体泛化;(c) 不同桌布(白色、蓝色)评估场景泛化。这些变量在训练集中均未出现。Figure 7: Experimental platform and generalization test objects. (a) the experimental platform (Flexiv arm + RealSense); (b) different weights (metal, pink) for evaluating object generalization; (c) different tablecloths (white, blue) for evaluating scene generalization. None of these variations appear in the training set.

消融实验要点Highlights of the Ablation Study

实验系统对比了无预训练(不同 epochs)、仅同任务预训练、同任务+跨任务预训练三类设置,在 10 / 40 / 75 条示范规模下分别评估。结论清晰:无论示范数量多少,RH20T 预训练(尤其是多任务预训练)均能一致提升 Reach、Pick、Place 三个阶段成功率。值得注意的是,该实验所用 RH20T 数据与评估环境在相机视角、桌布、机器人配置上均不同,体现了数据集的跨域迁移价值。The experiments systematically compare three settings, namely no pretraining (with different numbers of epochs), same-task pretraining only, and same-task plus cross-task pretraining, each evaluated at 10 / 40 / 75 demonstrations. The conclusion is clear: regardless of the number of demonstrations, RH20T pretraining (especially multi-task pretraining) consistently raises the success rate of all three stages, Reach, Pick and Place. Notably, the RH20T data used here differ from the evaluation environment in camera viewpoint, tablecloth and robot configuration, which demonstrates the cross-domain transfer value of the dataset.

04 局限性 LimitationsLimitations

说明:以下局限性均为论文作者在 Discussion & Conclusion 中明确陈述(stated)。Note: all limitations below are explicitly stated by the authors in the Discussion & Conclusion.
数据采集成本高昂Data collection is expensive

论文明确指出:"the cost of data collection is expensive"。RH20T 涉及多种机器人平台配置、大量人工遥操作、精密传感器标定,难以被资源有限的研究团队复制。这也是当前大规模真实世界机器人数据采集面临的共性挑战。The paper states explicitly: "the cost of data collection is expensive". RH20T involves many robot platform configurations, a large amount of manual teleoperation and precise sensor calibration, which makes it hard for research teams with limited resources to reproduce. This is also a shared challenge for large-scale real-world robot data collection today.

机器人基础模型(foundation model)能力尚未评估The capability of robotic foundation models has not been evaluated

论文明确指出:"the potential of robotic foundation models is not evaluated on our dataset"。作者尝试复现若干近期机器人基础模型的结果,但"haven't succeeded yet due to the limit of computing resources"。因此本文仅以 ACT 为 baseline 进行少样本评估,数据集对大规模基础模型的提升潜力有待后续研究验证。The paper states explicitly: "the potential of robotic foundation models is not evaluated on our dataset". The authors tried to reproduce the results of several recent robotic foundation models, but "haven't succeeded yet due to the limit of computing resources". The paper therefore evaluates only ACT as a baseline in the few-shot setting, leaving the dataset's potential for large-scale foundation models to be verified by future work.

操控范围待扩展The manipulation scope remains to be extended

当前 RH20T 主要覆盖单臂操控场景。论文展望未来工作时指出,希望将数据集扩展至 "broader robotic manipulation, including dual-arm and multi-finger dexterous manipulation"(双臂操控和多指灵巧手操控)。RH20T currently covers mainly single-arm manipulation. Looking ahead to future work, the paper hopes to extend the dataset to "broader robotic manipulation, including dual-arm and multi-finger dexterous manipulation".