4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
像让程序员只看一段积木倒塌或布料折叠的视频,再用 Blender 写出可以重播它的三维动画。4DCodeBench 不仅看画面像不像,还检查程序生成的形状与物质运动;但重播成功并不等于找到了真实的物理规律。
Imagine asking a programmer to watch blocks collapse or cloth fold, then write a Blender program that recreates the event. 4DCodeBench checks both the rendered video and the generated geometry and material motion; successfully replaying an event does not establish that the program recovered its physical laws.

为什么普通代码测试和视频相似度不够
Why ordinary code tests and video similarity are insufficient
普通代码题可以用输入输出判断对错,但这里要从像素反推物体形状、遮挡、材质、接触和运动。把每帧画得相似仍可能藏着错误的三维结构;即使几何正确,也可能靠逐帧安排位置来模仿碰撞。液体、柔性物体和断裂还会改变形状或拓扑,不能只靠给刚体设置位姿解决。 [1]
Ordinary coding problems can be checked through inputs and outputs. Here the agent must infer shape, occlusion, materials, contact and motion from pixels. Similar frames can conceal incorrect 3D structure; even correct geometry can use prescribed positions to imitate a collision. Fluids, soft objects and fracture also change shape or topology, exceeding simple rigid-object pose estimation. [1]
核心贡献是把动态视频重建变成可执行、可检查的三维程序任务:真实视频提供复杂观测,合成视频提供已知的三维几何和运动作为标准答案。这是一套 benchmark、数据和评测协议,不是新的视频生成网络或新的物理求解器。 [1]
The central contribution is turning dynamic video reconstruction into an executable, inspectable 3D programming task: real videos supply complex observations, while synthetic videos supply known geometry and motion for evaluation. This is a benchmark, dataset and evaluation protocol, rather than a new video-generation network or physics solver. [1]
| 对照对象 Comparison | 主要任务差别(按本文相关工作) Main task distinction, as characterized by this paper |
|---|---|
| 3DCodeBench | 从文字/图像生成静态三维物体代码。 Generate code for static 3D objects from text or images. |
| PhysCodeBench | 从文字描述编写物理模拟。 Write physical simulations from text descriptions. |
| VisPhyWorld / MPMWorlds | 已有从合成视频恢复模拟代码的任务;本文将 MPMWorlds 描述为二维任务。 Existing tasks recover simulations from synthetic video; this paper characterizes MPMWorlds as a 2D task. |
| 4DCodeBench | 把真实与合成动态场景结合,并在合成部分直接评估三维表面和物质轨迹。 Combine real and synthetic dynamic scenes and directly evaluate 3D surfaces and material trajectories on the synthetic split. |
定位判断:视频转代码的想法本身已有先例。这里更具体的新意,是多种材料与相互作用、统一的显式 4D 导出格式,以及与人类偏好对照的分项评测组合。上表依据本文第 5 节归纳,不代表对各项先前工作的独立复现实验。 [1]
Assessment: video-to-code already has precedents. The more specific contribution is the combination of varied materials and interactions, a unified explicit 4D export format, and decomposed evaluation checked against human preferences. The table summarizes Section 5 of this paper; it does not represent independent replication of the earlier projects. [1]
任务与评测:究竟输入、生成和检查什么
Task and evaluation: what goes in, comes out and gets checked
1. 只给视频,让智能体自行选择场景表示
1. Give only video and let the agent choose a scene representation
每个模型配置在每个场景运行一次。输入是 RGB 视频和统一的任务/输出格式说明,不提供场景名称、物体列表、语义描述或参考三维世界。每次运行使用独立容器和一块 GPU,提供 Blender、数值与模拟库、离线文档以及格式检查器;没有人为步数上限,但停止卡住或超过 6 小时的运行。 [1]
Each model configuration runs once per scene. It receives an RGB video and fixed task/output instructions, without scene names, object lists, semantic descriptions or reference 3D worlds. Each run uses an isolated container and one GPU, with Blender, numerical and simulation libraries, offline documentation and a format checker. There is no artificial step cap, but stalled runs or runs exceeding 6 hours are terminated. [1]
4D 就是三维空间随时间变化。智能体可反复写代码、渲染、比较和修改;运动既可来自 Blender 或自写的模拟器,也可来自解析公式和关键帧。最终程序必须独立重建结果,不能再读取参考视频或把视频像素复制成纹理。渲染分辨率、帧数和帧率必须与输入一致。 [1]
4D means 3D space evolving over time. Agents may repeatedly code, render, compare and revise. Motion can come from Blender or custom simulation, analytic formulas, or keyframes. The final program must regenerate the result independently, without reading the reference video or copying its pixels into textures. Output resolution, frame count and frame rate must match the input. [1]
| 术语 Term | 在这项任务中具体是什么 Concrete meaning in this task |
|---|---|
| Agent | 模型加命令行工具执行框架,负责查看视频、写代码和运行渲染。 A model plus its CLI tool harness, inspecting video, writing code and rendering. |
| World / 4D code | 可执行程序与其导出的相机、逐帧三角网格和物质轨迹,不是单独一段生成视频。 An executable program and its camera, per-frame triangle meshes and material paths, beyond the rendered video alone. |
| Analytic motion | 直接按时间计算位置/形变,例如沿样条曲线移动绳子或按预定时间撕裂物体。 Compute positions/deformations directly from time, such as moving rope along splines or tearing at prescribed times. |
| Simulation | 根据当前状态、力或约束推进下一时刻,例如刚体碰撞或基于位置约束的布料模拟。 Advance a state using forces or constraints, such as rigid collisions or position-based cloth simulation. |
| Lagrangian / Eulerian | 前者追踪同一物质点的完整路径;后者在这里比较每一步的三维位移分布。 The former follows the same material points through time; the latter here compares 3D displacement distributions at each step. |
| VQA / Elo | VQA 是根据重建视频回答场景问题;Elo 是两两偏好比较拟合出的相对评分。 VQA answers scene questions from reconstructed videos; Elo is a relative rating fitted from pairwise preferences. |
2. 导出可检查的几何和运动,而不只是视频
2. Export inspectable geometry and motion alongside video
solution/build.sh 重建 world/:其中 camera.json 描述相机,render.mp4 是渲染视频,meshes/ 保存逐帧表面,dynamics/ 保存持久物质点的位置。网格顶点数量与拓扑允许变化,因此断裂并不要求维持固定网格;无持久粒子身份的 Eulerian 液体还可导出 solver/ 缓存,供评测时恢复轨迹。 [1] [3]
solution/build.sh regenerates world/: camera.json describes the camera, render.mp4 is the video, meshes/ stores per-frame surfaces, and dynamics/ stores persistent material-point positions. Vertex counts and topology may change, so fracture need not preserve a fixed mesh. Eulerian liquids without persistent particle identities can also export solver/ bakes for trajectory extraction at evaluation time. [1] [3]
3. 两类数据提供不同强度的证据
3. Two data sources provide different kinds of evidence
数据集包含 100 个真实视频和 100 个物理模拟场景,涉及刚体/关节系统、可变形固体、布料/绳子等低维结构,以及颗粒/流体。83% 的场景含多个动态物体,66% 含多种材料;类别可重叠。真实片段优先选择相机稳定、遮挡较少、无剪辑的事件,并排除人类和动物;帧率上限为 60 fps,真实视频最多 300 帧。 [1]
The dataset contains 100 real videos and 100 physics-simulated scenes: rigid/articulated systems, deformable solids, cloth/rope and other thin structures, and granular/fluid materials. Multiple dynamic objects appear in 83% of scenes and multiple materials in 66%; categories can overlap. Real clips favor stable cameras, limited occlusion and continuous events, and exclude humans and animals. Frame rates are capped at 60 fps, and real videos at 300 frames. [1]
真实视频没有真实的三维运动答案,所以部分参考量由视觉模型估计;合成视频有已知世界,可直接比较三维表面与轨迹。两者互补,但不同分项并不都覆盖同一组场景。 [1]
Real videos lack ground-truth 3D motion, so vision models estimate some reference quantities. Synthetic videos have known worlds, enabling direct surface and trajectory comparisons. These sources complement each other, but metric families do not all cover the same scenes. [1]
4. 五个分项构成 Overall;VQA 与 Elo 单列
4. Five families form Overall; VQA and Elo remain separate
| 分项 Family | 测什么 Measurement | 适用数据与边界 Coverage and boundary |
|---|---|---|
| Perceptual | DINOv3 逐帧特征相似度。 Framewise DINOv3 feature similarity. | 全部场景;语义外观相似不保证运动正确。 All scenes; semantic resemblance does not ensure correct motion. |
| 2D Dynamics | 运动物体掩码重叠 Dynamic IoU,加光流与二维轨迹。 Moving-object mask overlap (Dynamic IoU), plus flow and 2D paths. | IoU 用于两类数据;Flow / Track2D 用于真实视频,参考由视觉模型估计。 IoU on both splits; Flow / Track2D on real video, with vision-estimated references. |
| 2.5D Geometry | 由表面栅格化得到深度,与估计参考深度比较。 Rasterized surface depth compared with estimated reference depth. | 真实视频;深度做归一化,不等于恢复绝对尺度。 Real video; normalized depth does not establish absolute scale recovery. |
| 3D Geometry | 配准后比较第 0 帧可见表面的 Chamfer 距离。 Chamfer distance on aligned visible surfaces at frame 0. | 合成场景;不直接检查所有时刻或所有隐藏表面。 Synthetic scenes; does not directly test all times or all hidden surfaces. |
| 3D Dynamics | Trajectory DTW 比较完整轨迹;EMD step 比较逐步位移分布。 Trajectory DTW compares full paths; EMD step compares stepwise displacement distributions. | 合成场景;分别测物质路径与整体运动统计。 Synthetic scenes; tests material paths and aggregate motion statistics. |
网站把各指标映射到预先声明的 [0,1] 范围,而非根据参评模型做 min–max 缩放;一个分项内部等权平均,Overall 再对五个分项等权平均。因此 Overall = 0.79 不是“79% 场景通过”。原始 Depth error 的范围是 [0,2],越低越好,不能直接当成表中越高越好的深度分项。 [2] [3]
The site maps metrics to declared [0,1] scales rather than min–max scaling over participating models. Metrics receive equal weight within each family; Overall then equally averages the five families. Thus Overall = 0.79 does not mean “79% of scenes passed.” Raw Depth error ranges from [0,2] with lower being better; it must not be confused with the higher-is-better depth family score in the leaderboard. [2] [3]
失败会被计入:无有效视频会使所有指标失败;视频有效但世界格式无效,则相关几何指标受罚。缺少参考信号的指标会对所有模型统一排除对应场景;另有 6 个含透明主体的真实场景被排除出部分栅格化比较。不能把网站的“覆盖所有 200 场景”理解为每个指标都恰好有 200 个有效参考。 [1]
Failures count: invalid video fails all metrics; valid video with an invalid world incurs penalties on geometry-dependent metrics. Scenes missing a prerequisite reference signal are excluded consistently across models for that metric. Six real scenes with transparent main objects are also excluded from certain rasterization comparisons. The site’s “all 200 scenes” coverage language should not be read as exactly 200 usable references for every metric. [1]
深入:配准、时间容忍度与裁判 Details: alignment, timing tolerance and judges
合成场景先做相似变换配准:分别从初始可见表面和动态轨迹求解,再选更贴合参考可见表面的结果。距离以参考表面的尺度归一化,避免相机导出错误单独摧毁三维评分;代价是分数不完全等于原始世界坐标的准确性。Trajectory DTW 通过最优匹配连接两组物质路径,并允许一定时间拉伸;EMD step 通过逐步位移分布补充即时运动信息。 [1]
Synthetic scenes first undergo similarity-transform alignment, fitted from either initial visible surfaces or dynamic trajectories; the better visible-surface fit is retained. Distances are normalized by reference-surface scale, preventing camera export errors alone from destroying 3D scores. The trade-off is that these scores do not directly measure raw world-coordinate accuracy. Trajectory DTW optimally matches material paths and permits temporal warping; EMD step complements it with instantaneous displacement distributions. [1]
VQA 使用 Gemini 3.8 Flash 查看 16 帧,回答 1,251 个二元问题,涵盖初态、终态、关键事件和接触;空白输出或主要物体缺失会被门控为错误。两两评测则让同一裁判以 3 fps 查看参考和两个重建,隐藏模型身份并平衡 A/B 位置,通过 Bradley–Terry 模型得到 Elo。Elo 的锚点依赖参评群体,跨不同拟合不能直接比较绝对值。 [1] [2]
VQA uses Gemini 3.8 Flash to inspect 16 frames and answer 1,251 binary questions about initial state, final state, key events and contact. Blank outputs or missing main objects are gated as incorrect. Pairwise evaluation gives the same judge the reference and two reconstructions at 3 fps, hides model identities and balances A/B positions, then fits Bradley–Terry Elo ratings. Elo’s anchor depends on the participating field; absolute values are not directly comparable across separate fits. [1] [2]
结果:静态外观更强,动态仍是短板
Results: stronger static reconstruction, weaker dynamics
以下精确数值来自核查当日的官方 leaderboard JSON,与论文图 4 的舍入结果一致;均为作者报告,本研究未重新运行 3,600 次评测。点击列标题排序,窄屏可横向滚动。模型名与推理档位采用论文标注。 [1] [2]
The precise values below come from the official leaderboard JSON on the check date and agree with the rounded results in paper Figure 4. These are author-reported results; this study did not rerun the 3,600 evaluations. Click column headers to sort; scroll horizontally on narrow screens. Model names and reasoning settings follow the paper. [1] [2]
| GPT-6 Astra [Max] | 0.791 | 0.906 | 0.600 | 0.791 | 0.922 | 0.735 | 0.876 | 1732.5 |
| Claude Opus 5.5 [High] | 0.778 | 0.882 | 0.556 | 0.771 | 0.937 | 0.746 | 0.828 | 1556.0 |
| GPT-6 Astra [High] | 0.765 | 0.897 | 0.551 | 0.770 | 0.919 | 0.689 | 0.850 | 1685.1 |
| Claude Fable 5.1 [High] | 0.741 | 0.858 | 0.518 | 0.742 | 0.902 | 0.685 | 0.788 | 1391.4 |
| GPT-6 Astra [Low] | 0.726 | 0.878 | 0.479 | 0.751 | 0.898 | 0.626 | 0.782 | 1432.6 |
| Claude Opus 5 [High] | 0.717 | 0.851 | 0.506 | 0.732 | 0.861 | 0.636 | 0.764 | 1360.8 |
| GPT-5.6 Luna [Max] | 0.669 | 0.834 | 0.420 | 0.700 | 0.861 | 0.533 | 0.673 | 1223.7 |
| GPT-5.6 Sol [High] | 0.645 | 0.833 | 0.357 | 0.631 | 0.846 | 0.560 | 0.674 | 1171.1 |
| Gemini 3.8 Flash [High] | 0.639 | 0.809 | 0.363 | 0.627 | 0.840 | 0.557 | 0.626 | 1037.3 |
| DeepSeek v4.1 Flash [High] | 0.597 | 0.779 | 0.338 | 0.617 | 0.769 | 0.479 | 0.500 | 886.0 |
| Qwen3.8 Flash-Next [XHigh] | 0.563 | 0.707 | 0.343 | 0.554 | 0.714 | 0.495 | 0.580 | 1061.1 |
| GPT-5.6 Terra [High] | 0.550 | 0.794 | 0.229 | 0.523 | 0.772 | 0.435 | 0.484 | 911.1 |
| GLM 5.3 Flash [Max] | 0.473 | 0.691 | 0.196 | 0.514 | 0.601 | 0.364 | 0.295 | 669.8 |
| Muse Glimmer [High] | 0.442 | 0.611 | 0.113 | 0.473 | 0.684 | 0.328 | 0.110 | 345.1 |
| MiMo v2.5 [Def] | 0.435 | 0.690 | 0.116 | 0.483 | 0.571 | 0.315 | 0.246 | 568.9 |
| MiniMax M3 [Def] | 0.339 | 0.525 | 0.114 | 0.355 | 0.445 | 0.259 | 0.178 | 554.3 |
| Mistral Medium 3.5 | 0.315 | 0.461 | 0.024 | 0.411 | 0.422 | 0.259 | 0.013 | 138.0 |
| Gemma-4 31B | 0.188 | 0.512 | 0.004 | 0.132 | 0.166 | 0.126 | 0.056 | 275.3 |
Astra [Max] 综合领先,Opus 5.5 的三维分项更高。前者 Overall 为 0.791,后者为 0.778;Opus 的 3D Geometry / 3D Dynamics 为 0.937 / 0.746,高于 Astra 的 0.922 / 0.735。Overall 与 VLM Elo 的排序并不相同:Astra [High] 的 Elo 高于 Opus 5.5,但 Overall 更低。最强开源权重模型也取决于指标:DeepSeek 的 Overall 为 0.597,Qwen 的 VQA 和 Elo 更高。 [2]
Astra [Max] leads Overall; Opus 5.5 scores higher in the 3D families. Their Overall scores are 0.791 and 0.778. Opus reaches 0.937 / 0.746 on 3D Geometry / 3D Dynamics, above Astra’s 0.922 / 0.735. Overall and VLM Elo also order models differently: Astra [High] has higher Elo but lower Overall than Opus 5.5. The strongest open-weight model depends on the metric too: DeepSeek scores 0.597 Overall, while Qwen has higher VQA and Elo. [2]
0.91 对 0.67 的含义:论文把 Perceptual 与 3D Geometry 平均为“静态”,把 2D Dynamics 与 3D Dynamics 平均为“动态”;这里没有把深度分项计入静态。Astra [Max] 的这两组均值分别约为 0.91 和 0.67。不同指标虽都在 [0,1],其阈值与含义不同,因此这是评测上的差距,不是统一物理单位下的误差比。 [1]
What 0.91 versus 0.67 means: the paper averages Perceptual and 3D Geometry as “static,” and 2D Dynamics and 3D Dynamics as “dynamic.” Depth is not included in this static grouping. Astra [Max] averages approximately 0.91 and 0.67 in these groups. Although the metrics share a [0,1] range, their thresholds and meanings differ; this is an evaluation gap, not an error ratio in a common physical unit. [1]
推理投入的变化:同一模型内有改善,跨模型不成立
Reasoning effort helps within a model, but token volume does not rank models
| Astra 档位 Astra setting | Overall | 2D Dynamics | 3D Dynamics | 输出 token 均值 Mean output tokens | 每任务美元均值 Mean USD / task |
|---|---|---|---|---|---|
| Low | 0.726 | 0.479 | 0.626 | 13,439 | $3.495 |
| High | 0.765 | 0.551 | 0.689 | 29,279 | $6.279 |
| Max | 0.791 | 0.600 | 0.735 | 67,475 | $11.321 |
从 Low 到 Max,Overall 增加约 0.064,报告的每任务费用约为原来的 3.24 倍(本研究根据表值计算)。这是一次推理档位变化实验,不是等预算消融。跨模型的总 token 与步数和 Overall 几乎不相关,论文报告 |ρ| ≤ 0.10;全部 token 中 97% 是缓存读取,因此“几百万 token”不能当作几百万新生成的推理 token。费用是该实验的计费估计,开源权重模型采用 OpenRouter 价格并假定 95% 缓存命中率。 [1] [2]
From Low to Max, Overall rises by about 0.064 while reported cost per task increases approximately 3.24× (calculated here from the table). This varies reasoning settings rather than holding budget fixed. Across models, total tokens and steps barely correlate with Overall: the paper reports |ρ| ≤ 0.10. Cache reads account for 97% of all tokens, so “millions of tokens” does not mean millions of newly generated reasoning tokens. Costs are estimates for this experiment; open-weight models use OpenRouter prices with an assumed 95% cache hit rate. [1] [2]
策略与数据差异,比单一冠军更有信息
Strategies and scene differences reveal more than a single winner
所有解法中,67% 使用解析运动,19% 自写模拟,10% 使用 Blender 物理,3% 使用关键帧;少量其他方法与舍入使总数不一定等于 100%。Opus 5.5 的自写/Blender 模拟合计为 76%;Astra 从 Low 到 High 到 Max 的模拟比例为 13% → 22% → 32%。这些是代码策略的统计,并不能证明使用模拟导致更高分。 [1]
Across solutions, 67% use analytic motion, 19% custom simulation, 10% Blender physics and 3% keyframes. Small residual categories and rounding mean the shares need not sum to 100%. Opus 5.5 uses custom/Blender simulation in 76% of solutions; Astra’s simulation share rises 13% → 22% → 32% from Low to High to Max. These describe implementation choices and do not establish that simulation causes better scores. [1]
真实视频更难:Astra [Max] 的真实/合成 VQA 为 0.853 / 0.898,Opus 5.5 为 0.775 / 0.880。论文还发现多材料场景的外观和 VQA 较弱,布料/绳子类的二维运动和深度更难。应称为跨场景来源的表现差异,不能据此断言未见过的物理干预也能泛化。 [1] [2]
Real footage is harder: Astra [Max] scores 0.853 / 0.898 VQA on real/synthetic scenes, versus Opus 5.5’s 0.775 / 0.880. The paper also finds weaker appearance and VQA on multi-material scenes, and harder 2D motion and depth for cloth/rope. These are performance differences across scene sources, not evidence of generalization to unseen physical interventions. [1] [2]

自动评分与人类偏好的一致性
Agreement between automated scores and human preferences
76 位参与者完成 3,587 次两两判断,覆盖 17 个模型配置;Opus 5.5 因发布时间未参加人类研究。模型层面 Human Elo 与 VLM Elo 的 Spearman ρ = 0.980,与 Overall 的 ρ = 0.96。在 916 个共同比较上,人类与 VLM 的一致率为 89.3%,人类间重复比较一致率为 92.1%。但相差 50 Elo 以内的模型,单次裁判判断接近随机,因此不能用总体高相关来保证相邻名次可靠。 [1]
Seventy-six participants supplied 3,587 pairwise judgments over 17 model configurations; Opus 5.5 was absent from the human study because of its release timing. At the model level, Human Elo correlates with VLM Elo at Spearman ρ = 0.980 and with Overall at ρ = 0.96. Human–VLM agreement is 89.3% on 916 shared comparisons, versus 92.1% agreement between humans on repeated comparisons. However, individual judge decisions are near chance for models within 50 Elo, so high aggregate correlation does not guarantee reliable adjacent ranks. [1]
局限与判断:这是否测到了物理理解?
Limitations and assessment: does this measure physical understanding?
- 最关键的欠约束:同一视频可能由多种几何和运动程序生成。按时间安排位置也能取得高分;测试没有系统改变初态、外力或材料,也没有要求预测观察窗口之后的事件。因此证据支持“动态场景重建能力存在差距”,不支持“高分模型已经学会可迁移的物理规律”,也不能仅凭某模型常用解析运动就判定其完全不懂物理。 [1]
- The central ambiguity: Multiple geometry and motion programs can explain one video. Prescribed positions can score well; the test does not systematically intervene on initial conditions, forces or materials, or demand prediction beyond the observed window. Evidence supports a gap in dynamic reconstruction, not proof of transferable physical laws in high-scoring models. Frequent analytic motion also cannot alone establish a complete absence of physical understanding. [1]
- VQA 的类别偏斜:论文图 5 的始终回答 Yes 基线为 87.5%,Astra [Max] 为 87.6%。这个基线不是一个能生成合格世界的智能体,因此不能替代完整 benchmark;但它说明 87.6% 不能按平衡二元问答解读为强物理推理证据。VQA 应与门控、各类问题、Elo 和几何/运动分项一起阅读。 [1]
- VQA answer imbalance: Paper Figure 5 gives an always-Yes baseline of 87.5%, versus 87.6% for Astra [Max]. This baseline is not an agent producing a valid world and cannot replace the full benchmark. It does show why 87.6% should not be read like accuracy on balanced binary questions or as strong physics-reasoning evidence. Read VQA alongside gates, question categories, Elo and geometry/motion metrics. [1]
- 模型、工具框架和预算混在一起:闭源模型运行各自原生 CLI,开源权重模型使用 Stirrup;单场景仅一次运行,且没有统一 token/费用预算。结果评估的是具体智能体系统,不能纯粹归因于模型权重。场景 bootstrap 的区间也不能替代重复运行来测量生成随机性。 [1]
- Models, harnesses and budgets vary together: Proprietary models use their native CLIs; open-weight models use Stirrup. There is one run per scene and no shared token or dollar budget. Results evaluate particular agent systems rather than model weights alone. Bootstrapping scenes does not replace repeated runs to measure generation variability. [1]
- 参考与指标都有选择:真实数据依赖估计深度/光流,3D Geometry 只比较配准后的初始可见表面,DTW 容忍部分时序差异,五分项平均也隐含权重选择。有限、偏稳定视角的 200 场景不能覆盖全部日常物理。低成本几何有效性检查又可能奖励简单但错误的物体,因此它们未进入 Overall。 [1]
- References and metrics make choices: Real references depend on estimated depth/flow; 3D Geometry checks aligned initial visible surfaces; DTW tolerates some timing differences; equal family averaging embeds a weighting choice. Two hundred scenes favoring stable views do not span everyday physics. Cheap mesh-validity checks may reward simple but inaccurate objects, so they remain outside Overall. [1]

值得保留的结论:这是一个有用的“从视觉到可检查程序”的测试平台,最有价值的信号是分项错误与策略差异。下一步应固定初始程序,改变重力、摩擦、物体位置或施加外力,并检查更长时间与新视角;再用等预算、重复运行比较解析动画和模拟。前一类干预与长程预测是作者提出的方向,后面的控制设计是本研究建议,均不是论文已验证的结果。 [1]
The useful conclusion: this is a valuable testbed for translating vision into inspectable programs, with diagnostic errors and strategy differences providing its strongest signals. Next, keep the inferred program fixed, vary gravity, friction, object positions or external forces, and test longer horizons and new views; compare analytic animation and simulation under matched budgets and repeated runs. Interventions and longer-horizon prediction are author-proposed directions; the additional controls are this study’s suggestions. None are established results of the paper. [1]
如何使用:先检查一个场景的完整交付物
Using it: inspect one scene’s complete deliverable first
- 获取数据与环境:官方仓库提供推理容器、checker、scorer、VLM 裁判和可视化工具;Hugging Face 提供数据。固定代码版本、模型版本、CLI、推理档位和预算后,先运行一个场景。
- Get data and environment: the official repository provides inference containers, checker, scorer, VLM judge and visualization tools; Hugging Face hosts data. Record code/model versions, CLI, reasoning setting and budget, then start with one scene.
- 验证程序与世界:确认 build 程序能独立重建;使用 checker 检查格式,用预览确认导出相机与网格和渲染一致。checker 不判断重建是否真的匹配视频。
- Validate the program and world: check that the build independently regenerates the result. Use the checker for format validity and its preview to check camera/mesh alignment with the render. The checker does not determine whether the reconstruction matches the reference.
- 准备参考,再评分:scorer 的 prepare 阶段生成可复用的参考估计,score 阶段输出各指标。既看具体错误帧和分项,也保留失败状态;需要语义事件或偏好比较时再运行 VQA/Elo。
- Prepare references, then score: the scorer’s prepare stage generates reusable reference estimates; its score stage writes per-metric outputs. Inspect failing frames and individual metrics while retaining failure states; add VQA/Elo for semantic events or preference comparisons.
- 按用途选模型:复现视频外观、重建可编辑几何和推断可迁移模拟是不同目标。先确定关心的输出与分项,再看预算;不要仅凭一个 Overall 数字把模型当作通用物理引擎。
- Select for your objective: recreating video appearance, recovering editable geometry and inferring transferable simulations are different goals. Choose the outputs and metrics that matter, then compare budgets; an Overall score alone does not qualify a model as a general physics engine.
以上流程依据公开 README 和格式规范整理;本次只阅读并核对源码文档与发布结果,未安装或执行完整 GPU 基准。 [3]
This workflow follows the public README and format specification. This study inspected source documentation and reported results; it did not install or execute the full GPU benchmark. [3]
资料、版本与证据边界
Sources, version and evidence boundary
- 4DCodeBench — arXiv:2610.03715v1
2026-10-02 提交。方法与结论见第 2–6 节;指标见附录 D,人工评测见 E,策略与运行统计见 F.4 和 G。PDF 首页另印有 10-5-2026,本页发表日期采用 arXiv 提交记录。
Submitted 2026-10-02. Methods and conclusions: Sections 2–6; metrics: Appendix D; human study: E; strategies and runtime statistics: F.4 and G. The PDF also prints 10-5-2026 on its first page; this study uses the arXiv submission record for publication date.
PDF - 官方项目Official project · 排行榜 JSONLeaderboard JSON
在 2026-10-07 获取。排行榜、费用和图表使用此快照;实时网站后续可能变化。原始数据另存为随页资产,便于复查。
Retrieved 2026-10-07. The table, costs and chart use this snapshot; the live site may subsequently change. The raw data is included as a page asset for inspection.
JSON - 官方仓库Official repository · 评测规范Evaluation specification
在 2026-10-07 阅读 README 与格式/评测规范;数据入口为 4DCodeBench on Hugging Face。
README and format/evaluation specification read on 2026-10-07; data entry point: 4DCodeBench on Hugging Face.
图片来自原论文,保留原图标签;图 2 由公开数值重新绘制。解释、计算与批判性判断已和作者实验结果区分。本页提供完整中英双语正文;右下角语言按钮可切换。
Images come from the paper with original labels preserved; Figure 2 is replotted from public values. Explanations, calculations and critical assessments are distinguished from author-reported experiments. This page contains full Chinese and English versions, switchable using the language button at the lower right.