← 论文海报合集← Paper Notes|
具身 AI · Embodied World Models · Benchmark · 2026Embodied AI · Embodied World Models · Benchmark · 2026

WorldArena: 具身世界模型感知与功能性统一评测基准WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models

首个同时评估视频感知质量与具身任务功能性的综合 benchmarkThe first comprehensive benchmark that jointly evaluates video perceptual quality and functional utility on embodied tasks
Yu Shang, Zhuohang Li, Yiding Ma, Weikang Su, Xin Jin, Ziyou Wang, Lei Jin, Xin Zhang, Yinzhou Tang, Haisheng Su, Chen Gao, Wei Wu, Xihui Liu, Dhruv Shah, Zhaoxiang Zhang, Zhibo Chen, Jun Zhu, Yonghong Tian, Tat-Seng Chua, Wenwu Zhu, Yong Li

当前具身世界模型(Embodied World Models, EWMs)的评测高度碎片化——主流工作仅关注视频生成质量,却忽视模型在真实具身任务中的实用价值。WorldArena 构建了首个统一评测体系:16 项视频质量指标 + 3 类具身任务评测 + 人工标注,并提出综合评分 EWMScore。对 14 个代表性模型的系统评测揭示了"视觉质量领先≠任务能力领先"的关键鸿沟。Evaluation of current Embodied World Models (EWMs) is highly fragmented — mainstream work attends only to video generation quality while overlooking the practical value of these models on real embodied tasks. WorldArena builds the first unified evaluation system: 16 video quality metrics + 3 categories of embodied task evaluation + human annotation, and proposes the composite score EWMScore. A systematic evaluation of 14 representative models reveals the critical gap that “leading visual quality ≠ leading task capability”.

arXiv · 2026-02 14 models evaluated 16 video metrics 📄 arXiv:2602.08971 🌐 Project / Leaderboard
embodied world models benchmark video generation policy evaluation data engine EWMScore 具身智能embodied intelligence 机器人操作robot manipulation

01 动机Motivation

现有具身世界模型评测存在三大核心问题:只关注视频生成质量(perceptual fidelity),忽略功能性(functional utility);单次评测覆盖模型数量少(通常不超过 10 个);缺乏统一框架将感知质量与下游任务效果关联起来。Existing evaluation of embodied world models suffers from three core problems: it attends only to video generation quality (perceptual fidelity) and ignores functional utility; a single evaluation covers few models (typically no more than 10); and there is no unified framework linking perceptual quality to downstream task effectiveness.

"Current evaluation of embodied world models has largely focused on perceptual fidelity (e.g., video generation quality), overlooking the functional utility of these models in downstream decision-making tasks."
WorldArena 评测框架总览
图 1 (a):14 个代表性具身世界模型的 EWMScore 评分总览。涵盖通用视频模型(CogvideoX、Wan 2.6、Veo 3.1 等)、文本条件具身模型(Genie Envisioner、GigaWorld、TesserAct、WOW 等)、动作条件具身模型(IRASim、Cosmos-Predict 2.5、CtrlWorld)。Figure 1 (a): Overview of the EWMScore ratings of 14 representative embodied world models, covering general-purpose video models (CogvideoX, Wan 2.6, Veo 3.1, etc.), text-conditioned embodied models (Genie Envisioner, GigaWorld, TesserAct, WOW, etc.) and action-conditioned embodied models (IRASim, Cosmos-Predict 2.5, CtrlWorld).
各模型多维度雷达图对比
图 1 (b):14 个模型在六大评测维度(视觉质量、运动质量、内容一致性、物理合理性、三维精度、可控性)的雷达图对比,直观呈现感知质量领先模型与任务能力领先模型之间的差异。Figure 1 (b): Radar-chart comparison of 14 models across the six evaluation dimensions (visual quality, motion quality, content consistency, physics adherence, 3D accuracy, controllability), directly showing the divergence between models leading in perceptual quality and those leading in task capability.
14被评测模型数量(通用 + 具身专用)Models evaluated (general-purpose + embodied-specific)
16视频质量自动化评测指标数量Automated metrics for video quality evaluation
0.825EWMScore 与人工评测相关系数 rCorrelation coefficient r between EWMScore and human evaluation
3,500人工标注视频数(70 位标注员)Human-annotated videos (70 annotators)

02 方法Method

WorldArena 构建三维统一评测体系:(1) 视频质量评测(16 项自动化指标,覆盖 6 个维度);(2) 具身任务评测(数据引擎、策略评估器、动作规划器);(3) 人工评测。最终通过 EWMScore 将多维结果汇聚为单一可解释指标。WorldArena builds a three-part unified evaluation system: (1) video quality evaluation (16 automated metrics covering 6 dimensions); (2) embodied task evaluation (data engine, policy evaluator, action planner); (3) human evaluation. EWMScore finally aggregates the multi-dimensional results into a single interpretable metric.

视频质量评测六维度示意
图 2:视频质量评测的六大维度示意图。每个维度包含 2–3 项自动化指标,从不同角度量化世界模型生成视频的质量。Figure 2: Illustration of the six dimensions of video quality evaluation. Each dimension contains 2–3 automated metrics that quantify, from different angles, the quality of the videos generated by a world model.

维度一:视频质量评测(16 指标 × 6 维度)Dimension I: Video Quality Evaluation (16 metrics × 6 dimensions)

视觉质量 · Visual QualityVisual Quality

  • Image Quality:用 MUSIQ 模型评估技术失真(噪声、压缩伪影、过曝)Image Quality: assesses technical distortion (noise, compression artifacts, overexposure) with the MUSIQ model
  • Aesthetic Quality:用 LAION 预测器评估色彩构图与艺术一致性Aesthetic Quality: assesses color composition and artistic consistency with the LAION predictor
  • JEPA Similarity:用 V-JEPA 编码器的特征最大均值差异(MMD)衡量分布距离JEPA Similarity: measures distributional distance via the maximum mean discrepancy (MMD) of V-JEPA encoder features

运动质量 · Motion QualityMotion Quality

  • Dynamic Degree:光流分析前 5% 活跃像素的运动强度Dynamic Degree: optical-flow analysis of the motion magnitude of the top 5% most active pixels
  • Flow Score:全帧平均光流幅度,反映整体运动动态Flow Score: full-frame mean optical-flow magnitude, reflecting overall motion dynamics
  • Motion Smoothness:帧插值法比较预测与真实中间帧,衡量运动平滑度Motion Smoothness: frame interpolation compares predicted and ground-truth intermediate frames to measure motion smoothness

内容一致性 · Content ConsistencyContent Consistency

  • Subject Consistency:DINO 特征余弦相似度,追踪物体稳定性Subject Consistency: cosine similarity of DINO features, tracking object stability
  • Background Consistency:CLIP 特征相似度,衡量场景稳定性Background Consistency: CLIP feature similarity, measuring scene stability
  • Photometric Consistency:基于光流的纹理稳定性(平均端点误差)Photometric Consistency: optical-flow-based texture stability (mean endpoint error)

物理合理性 · Physics AdherencePhysics Adherence

  • Interaction Quality:用 Qwen3-VL 评估接触行为与力传导合理性Interaction Quality: assesses contact behavior and the plausibility of force transmission with Qwen3-VL
  • Trajectory Accuracy:用 SAM 3 提取边界框,结合 normalized DTW 衡量轨迹精度Trajectory Accuracy: extracts bounding boxes with SAM 3 and measures trajectory accuracy with normalized DTW

三维精度 · 3D Accuracy3D Accuracy

  • Depth Accuracy:单目深度估计 + median-based scaling 策略Depth Accuracy: monocular depth estimation + a median-based scaling strategy
  • Perspectivity:VLM 判断尺度变化、光照一致性与遮挡关系Perspectivity: a VLM judges scale change, lighting consistency and occlusion relations

可控性 · ControllabilityControllability

  • Instruction Following:VLM 评估动作类型、目标物体、任务状态对齐程度Instruction Following: a VLM assesses the alignment of action type, target object and task state
  • Semantic Alignment:Qwen2.5-VL 在生成视频与参考视频描述间的余弦相似度Semantic Alignment: cosine similarity between the captions of the generated video and the reference video, computed with Qwen2.5-VL
  • Action Following:三条不同指令下特征差异的平均成对不相似度Action Following: mean pairwise dissimilarity of feature differences under three distinct instructions

维度二:具身任务评测Dimension II: Embodied Task Evaluation

具身任务评测框架
图 3:具身任务评测体系概览,包含三类角色评测:数据引擎(衡量下游策略的成功率提升)、策略评估器(衡量与真实环境评测结果的相关性)、动作规划器(衡量基于世界模型策略的任务成功率)。Figure 3: Overview of the embodied task evaluation system, comprising three role-based evaluations: data engine (measuring the success-rate gain of the downstream policy), policy evaluator (measuring the correlation with real-environment evaluation results) and action planner (measuring the task success rate of a world-model-based policy).

EWMScore 综合评分The EWMScore Composite Score

将 16 项视频质量指标综合为单一可解释指数:(1) 基于经验定义的边界值线性归一化至 [0,1];(2) 缩放至 [0,100];(3) 跨所有归一化指标的算术均值。通过与人工评测、具身任务性能的相关性验证其有效性。Aggregates the 16 video quality metrics into a single interpretable index: (1) linear normalization to [0,1] against empirically defined bounds; (2) rescaling to [0,100]; (3) the arithmetic mean across all normalized metrics. Its validity is verified through correlation with human evaluation and with embodied task performance.

03 实验Experiments

评测数据集为 RoboTwin 2.0,共 2,500 个视频(2,000 训练 / 500 测试),覆盖 50 个双臂机器人操作场景。评测 14 个模型,涵盖通用视频模型与具身专用模型两大类别。The evaluation dataset is RoboTwin 2.0, with 2,500 videos in total (2,000 train / 500 test) covering 50 bimanual robot manipulation scenarios. 14 models are evaluated, spanning the two categories of general-purpose video models and embodied-specific models.

视频质量排名亮点Highlights of the Video Quality Ranking

模型Model 类别Category 突出指标Standout metrics
Wan 2.6 / Veo 3.1 通用视频General video Image Quality 0.68、Aesthetic Quality 0.46+(最高视觉质量)Image Quality 0.68, Aesthetic Quality 0.46+ (highest visual quality)
CtrlWorld 动作条件具身Action-conditioned embodied 3D 精度最佳(Depth 0.4766)、Subject Consistency 0.9185Best 3D accuracy (Depth 0.4766), Subject Consistency 0.9185
CogvideoX / IRASim 混合Mixed JEPA Similarity 0.93+(内容一致性优秀)JEPA Similarity 0.93+ (excellent content consistency)
WoW 文本条件具身Text-conditioned embodied Action Following 0.0434(最佳动作跟随性)Action Following 0.0434 (best action following)

具身任务性能:数据引擎与动作规划器Embodied Task Performance: Data Engine and Action Planner

合成数据训练下游策略的成功率仅为 1–45%,而真实数据训练的成功率为 66–77%。仅 RoboMaster 和 WoW 在部分任务上超越了真实数据基线。论文指出:"current embodied world models are not yet reliable data sources"。动作规划器任务成功率同样处于 1–45% 范围,表明"world models capture useful predictive structure",但"struggle to reliably support closed-loop task execution"。Downstream policies trained on synthetic data reach a success rate of only 1–45%, whereas those trained on real data reach 66–77%. Only RoboMaster and WoW surpass the real-data baseline on part of the tasks. The paper notes: “current embodied world models are not yet reliable data sources”. Action planner task success rates likewise fall in the 1–45% range, indicating that “world models capture useful predictive structure” but “struggle to reliably support closed-loop task execution”.

策略评估器:与仿真器的相关性Policy Evaluator: Correlation with the Simulator

策略评估相关性
图 4:CtrlWorld 与 RoboTwin 仿真器评测结果呈现强相关性,有效捕捉环境动力学;而 Cosmos-Predict 2.5 相关性较弱,说明其"struggles to accurately model environment dynamics"。Figure 4: CtrlWorld exhibits a strong correlation with the RoboTwin simulator evaluation results, effectively capturing environment dynamics, whereas Cosmos-Predict 2.5 correlates weakly, showing that it “struggles to accurately model environment dynamics”.

EWMScore 与人工评测、任务性能的相关性Correlation of EWMScore with Human Evaluation and Task Performance

EWMScore 相关性分析
图 5:EWMScore 与三类评测维度的相关性。与人工评测相关系数 r = 0.825(强),与数据合成任务 r = 0.600(中等),与动作规划任务 r = 0.360(弱)。说明"perceptual realism is a necessary condition",但不能直接推导出"proportional gains in downstream embodied tasks"。Figure 5: Correlation of EWMScore with three categories of evaluation. The correlation coefficient with human evaluation is r = 0.825 (strong), with the data synthesis task r = 0.600 (moderate) and with the action planning task r = 0.360 (weak). This shows that “perceptual realism is a necessary condition”, but it cannot be directly inferred that there will be “proportional gains in downstream embodied tasks”.
EWMScore 相关性对象EWMScore correlation target 相关系数 rCorrelation coefficient r 解读Interpretation
人工评测(Human Evaluation)Human Evaluation 0.825 强相关,EWMScore 可有效代理人工判断Strong correlation; EWMScore is an effective proxy for human judgment
数据合成任务(Data Engine)Data synthesis task (Data Engine) 0.600 中等相关,感知质量部分迁移到合成数据效果Moderate correlation; perceptual quality partly transfers to the effectiveness of synthetic data
动作规划任务(Action Planner)Action planning task (Action Planner) 0.360 弱相关,视觉质量与规划能力差距显著Weak correlation; a pronounced gap between visual quality and planning capability

核心发现:感知—功能性鸿沟Core Finding: The Perception—Utility Gap

视觉质量最高的 Veo 3.1、Wan 2.6 等通用模型并未在具身任务中取得对应的领先优势。动作条件具身专用模型(如 CtrlWorld)尽管感知评分较低,却在策略评估和物理合理性方面表现更优。通用视频模型展现强视觉保真度,但存在"semantic drift";具身专用模型生成的动作序列"more coherent and goal-consistent";动作条件方法更好地捕捉交互动力学。General-purpose models with the highest visual quality, such as Veo 3.1 and Wan 2.6, do not obtain a corresponding lead on embodied tasks. Action-conditioned embodied-specific models (e.g. CtrlWorld), despite lower perceptual scores, perform better on policy evaluation and physics adherence. General-purpose video models show strong visual fidelity but suffer from “semantic drift”; the action sequences generated by embodied-specific models are “more coherent and goal-consistent”; and action-conditioned approaches capture interaction dynamics better.

04 局限性Limitations

Note: 以下限制部分由作者在论文中明确陈述(标注"stated"),部分从评测设计中推断(标注"inferred")。Some of the limitations below are explicitly stated by the authors in the paper (marked “stated”), others are inferred from the evaluation design (marked “inferred”).
评测域局限于双臂机器人操作(stated)Evaluation domain restricted to bimanual robot manipulation (stated)

当前 WorldArena 全部实验限于 bimanual robotic manipulation 领域。作者承诺将来扩展至更多具身智能场景(导航、移动操作等),但目前结论的泛化性有限。All current WorldArena experiments are confined to the bimanual robotic manipulation domain. The authors promise to extend to more embodied AI settings (navigation, mobile manipulation and so on) in the future, but the generality of the present conclusions is limited.

数据规模相对有限(stated)Relatively limited data scale (stated)

测试数据集仅包含 2,500 个视频,覆盖单一仿真器(RoboTwin 2.0)的 50 个任务场景。人工标注由 70 位标注员完成 3,500 个视频的评测,样本量处于"solid but potentially limited"水平。The test dataset contains only 2,500 videos, covering 50 task scenarios from a single simulator (RoboTwin 2.0). Human annotation was completed by 70 annotators over 3,500 videos, a sample size at a “solid but potentially limited” level.

合成数据质量不足以支撑大规模策略训练(stated)Synthetic data quality is insufficient to support large-scale policy training (stated)

实验显示绝大多数世界模型生成的合成数据训练效果(成功率 1–45%)远低于真实数据(66–77%),当前 EWMs 尚不能作为可靠的数据合成引擎。Experiments show that training on the synthetic data generated by the vast majority of world models (success rate 1–45%) is far worse than on real data (66–77%); current EWMs cannot yet serve as reliable data synthesis engines.

策略评估可能存在系统性高估(inferred)Policy evaluation may be systematically over-optimistic (inferred)

通过世界模型评估策略时,部分模型(如 Cosmos-Predict 2.5)与真实仿真器相关性较弱,说明以世界模型替代真实环境评估存在系统性偏差,可能高估策略的真实成功率。When policies are evaluated through world models, some models (such as Cosmos-Predict 2.5) correlate weakly with the real simulator, showing that substituting world models for real-environment evaluation introduces a systematic bias that may overestimate the true success rate of a policy.

长时序任务执行仍是挑战(inferred)Long-horizon task execution remains a challenge (inferred)

当前世界模型在闭环长时序任务执行中表现受限。作为动作规划器时,任务成功率偏低,说明从短视频生成到多步骤动作序列规划仍有显著差距。Current world models are limited in closed-loop long-horizon task execution. As action planners their task success rates are low, showing that a substantial gap remains between short-video generation and multi-step action sequence planning.