当前具身世界模型(Embodied World Models, EWMs)的评测高度碎片化——主流工作仅关注视频生成质量,却忽视模型在真实具身任务中的实用价值。WorldArena 构建了首个统一评测体系:16 项视频质量指标 + 3 类具身任务评测 + 人工标注,并提出综合评分 EWMScore。对 14 个代表性模型的系统评测揭示了"视觉质量领先≠任务能力领先"的关键鸿沟。Evaluation of current Embodied World Models (EWMs) is highly fragmented — mainstream work attends only to video generation quality while overlooking the practical value of these models on real embodied tasks. WorldArena builds the first unified evaluation system: 16 video quality metrics + 3 categories of embodied task evaluation + human annotation, and proposes the composite score EWMScore. A systematic evaluation of 14 representative models reveals the critical gap that “leading visual quality ≠ leading task capability”.
现有具身世界模型评测存在三大核心问题:只关注视频生成质量(perceptual fidelity),忽略功能性(functional utility);单次评测覆盖模型数量少(通常不超过 10 个);缺乏统一框架将感知质量与下游任务效果关联起来。Existing evaluation of embodied world models suffers from three core problems: it attends only to video generation quality (perceptual fidelity) and ignores functional utility; a single evaluation covers few models (typically no more than 10); and there is no unified framework linking perceptual quality to downstream task effectiveness.
"Current evaluation of embodied world models has largely focused on perceptual fidelity (e.g., video generation quality), overlooking the functional utility of these models in downstream decision-making tasks."
WorldArena 构建三维统一评测体系:(1) 视频质量评测(16 项自动化指标,覆盖 6 个维度);(2) 具身任务评测(数据引擎、策略评估器、动作规划器);(3) 人工评测。最终通过 EWMScore 将多维结果汇聚为单一可解释指标。WorldArena builds a three-part unified evaluation system: (1) video quality evaluation (16 automated metrics covering 6 dimensions); (2) embodied task evaluation (data engine, policy evaluator, action planner); (3) human evaluation. EWMScore finally aggregates the multi-dimensional results into a single interpretable metric.
将 16 项视频质量指标综合为单一可解释指数:(1) 基于经验定义的边界值线性归一化至 [0,1];(2) 缩放至 [0,100];(3) 跨所有归一化指标的算术均值。通过与人工评测、具身任务性能的相关性验证其有效性。Aggregates the 16 video quality metrics into a single interpretable index: (1) linear normalization to [0,1] against empirically defined bounds; (2) rescaling to [0,100]; (3) the arithmetic mean across all normalized metrics. Its validity is verified through correlation with human evaluation and with embodied task performance.
评测数据集为 RoboTwin 2.0,共 2,500 个视频(2,000 训练 / 500 测试),覆盖 50 个双臂机器人操作场景。评测 14 个模型,涵盖通用视频模型与具身专用模型两大类别。The evaluation dataset is RoboTwin 2.0, with 2,500 videos in total (2,000 train / 500 test) covering 50 bimanual robot manipulation scenarios. 14 models are evaluated, spanning the two categories of general-purpose video models and embodied-specific models.
| 模型Model | 类别Category | 突出指标Standout metrics |
|---|---|---|
| Wan 2.6 / Veo 3.1 | 通用视频General video | Image Quality 0.68、Aesthetic Quality 0.46+(最高视觉质量)Image Quality 0.68, Aesthetic Quality 0.46+ (highest visual quality) |
| CtrlWorld | 动作条件具身Action-conditioned embodied | 3D 精度最佳(Depth 0.4766)、Subject Consistency 0.9185Best 3D accuracy (Depth 0.4766), Subject Consistency 0.9185 |
| CogvideoX / IRASim | 混合Mixed | JEPA Similarity 0.93+(内容一致性优秀)JEPA Similarity 0.93+ (excellent content consistency) |
| WoW | 文本条件具身Text-conditioned embodied | Action Following 0.0434(最佳动作跟随性)Action Following 0.0434 (best action following) |
合成数据训练下游策略的成功率仅为 1–45%,而真实数据训练的成功率为 66–77%。仅 RoboMaster 和 WoW 在部分任务上超越了真实数据基线。论文指出:"current embodied world models are not yet reliable data sources"。动作规划器任务成功率同样处于 1–45% 范围,表明"world models capture useful predictive structure",但"struggle to reliably support closed-loop task execution"。Downstream policies trained on synthetic data reach a success rate of only 1–45%, whereas those trained on real data reach 66–77%. Only RoboMaster and WoW surpass the real-data baseline on part of the tasks. The paper notes: “current embodied world models are not yet reliable data sources”. Action planner task success rates likewise fall in the 1–45% range, indicating that “world models capture useful predictive structure” but “struggle to reliably support closed-loop task execution”.
| EWMScore 相关性对象EWMScore correlation target | 相关系数 rCorrelation coefficient r | 解读Interpretation |
|---|---|---|
| 人工评测(Human Evaluation)Human Evaluation | 0.825 | 强相关,EWMScore 可有效代理人工判断Strong correlation; EWMScore is an effective proxy for human judgment |
| 数据合成任务(Data Engine)Data synthesis task (Data Engine) | 0.600 | 中等相关,感知质量部分迁移到合成数据效果Moderate correlation; perceptual quality partly transfers to the effectiveness of synthetic data |
| 动作规划任务(Action Planner)Action planning task (Action Planner) | 0.360 | 弱相关,视觉质量与规划能力差距显著Weak correlation; a pronounced gap between visual quality and planning capability |
视觉质量最高的 Veo 3.1、Wan 2.6 等通用模型并未在具身任务中取得对应的领先优势。动作条件具身专用模型(如 CtrlWorld)尽管感知评分较低,却在策略评估和物理合理性方面表现更优。通用视频模型展现强视觉保真度,但存在"semantic drift";具身专用模型生成的动作序列"more coherent and goal-consistent";动作条件方法更好地捕捉交互动力学。General-purpose models with the highest visual quality, such as Veo 3.1 and Wan 2.6, do not obtain a corresponding lead on embodied tasks. Action-conditioned embodied-specific models (e.g. CtrlWorld), despite lower perceptual scores, perform better on policy evaluation and physics adherence. General-purpose video models show strong visual fidelity but suffer from “semantic drift”; the action sequences generated by embodied-specific models are “more coherent and goal-consistent”; and action-conditioned approaches capture interaction dynamics better.
当前 WorldArena 全部实验限于 bimanual robotic manipulation 领域。作者承诺将来扩展至更多具身智能场景(导航、移动操作等),但目前结论的泛化性有限。All current WorldArena experiments are confined to the bimanual robotic manipulation domain. The authors promise to extend to more embodied AI settings (navigation, mobile manipulation and so on) in the future, but the generality of the present conclusions is limited.
测试数据集仅包含 2,500 个视频,覆盖单一仿真器(RoboTwin 2.0)的 50 个任务场景。人工标注由 70 位标注员完成 3,500 个视频的评测,样本量处于"solid but potentially limited"水平。The test dataset contains only 2,500 videos, covering 50 task scenarios from a single simulator (RoboTwin 2.0). Human annotation was completed by 70 annotators over 3,500 videos, a sample size at a “solid but potentially limited” level.
实验显示绝大多数世界模型生成的合成数据训练效果(成功率 1–45%)远低于真实数据(66–77%),当前 EWMs 尚不能作为可靠的数据合成引擎。Experiments show that training on the synthetic data generated by the vast majority of world models (success rate 1–45%) is far worse than on real data (66–77%); current EWMs cannot yet serve as reliable data synthesis engines.
通过世界模型评估策略时,部分模型(如 Cosmos-Predict 2.5)与真实仿真器相关性较弱,说明以世界模型替代真实环境评估存在系统性偏差,可能高估策略的真实成功率。When policies are evaluated through world models, some models (such as Cosmos-Predict 2.5) correlate weakly with the real simulator, showing that substituting world models for real-environment evaluation introduces a systematic bias that may overestimate the true success rate of a policy.
当前世界模型在闭环长时序任务执行中表现受限。作为动作规划器时,任务成功率偏低,说明从短视频生成到多步骤动作序列规划仍有显著差距。Current world models are limited in closed-loop long-horizon task execution. As action planners their task success rates are low, showing that a substantial gap remains between short-video generation and multi-step action sequence planning.