VISTA: A Visual Harness for Reasoning in an Interactive World
* 前三位作者为共同第一作者。本研究解读以 2026-10-01 的论文 v1 为准;数字均为作者报告,未在本次工作中重跑实验。
* The first three authors contributed equally. This study uses the 2026-10-01 paper, v1. All experimental numbers are author-reported; the experiments were not rerun for this report.
给一个会看图的语言模型配上完整的截图档案、放大镜和笔记本:它可以一边玩陌生游戏,一边回看旧画面、检查细节并修正规则。关键是把原始证据保存在模型上下文之外,等问题明确后再取回来。
Give a vision-capable language model a complete screenshot archive, a magnifying glass, and a notebook. It can learn an unfamiliar game by revisiting earlier images, checking details, and revising its rules; the key is to keep original evidence outside the model context and retrieve it when a question becomes relevant.
这些结果衡量的是指定公开游戏上的完成情况和动作效率,不是实时速度或通用智能。 §4
These results measure completion and action efficiency on the specified public games, rather than real-time speed or general intelligence. §4

研究动机:需要时才能知道该看什么
Motivation: the relevant detail may become clear later
ARC-AGI-3 不直接告诉模型游戏规则或目标。模型必须通过动作的后果,推断哪些物体可以移动、哪些颜色表示危险,以及怎样过关。后面的关卡沿用部分规则,又加入新机制,因此只看当前截图很容易丢失关键证据。
ARC-AGI-3 does not explicitly tell the model the rules or goals. The model must infer which objects move, which colors indicate danger, and how to finish a level from the consequences of its actions. Later levels retain some mechanics and add others, so the current screenshot alone can omit crucial evidence.
文本数字网格保留离散颜色,却增加 token 开销;固定窗口的截图代理会丢掉早期画面;文本摘要则只能保留当时认为重要的信息。VISTA 的核心判断是:不要让首次编码或摘要成为视觉证据的最后一次机会。 原始帧继续存在,模型可以针对新的假设重新查看。 §3
Textual grids retain discrete colors but increase token use; screenshot agents with a fixed window lose older views; textual summaries preserve only what seemed important at the time. VISTA’s central insight is that the initial encoding or summary should not be the last opportunity to examine visual evidence. Original frames remain available for checking new hypotheses. §3
核心方法:存原图,按需重看,继续行动
Method: archive, inspect on demand, and act
术语速查
Terms grounded in this system
| 术语 Term | 在这里具体指什么 Concrete meaning here |
|---|---|
| Harness Harness | 执行工具调用、保存帧并管理上下文的外围软件;没有新增训练模型。 The surrounding software that executes tools, stores frames, and manages context; it adds no trained model. |
| Observation Observation | 环境返回的 PNG 画面;ARC-AGI-3 默认将 64 × 64 帧以最近邻放大到 512 × 512。 A PNG view returned by the environment; ARC-AGI-3 defaults to nearest-neighbor enlargement from 64 × 64 to 512 × 512. |
| Lossless visual memory Lossless visual memory | 按 turn 和 frame 编号保存所有返回帧,包括中间动画帧。 Every returned frame, including intermediate animation frames, stored by turn and frame index. |
| Visual inspection Visual inspection | 调用 inspect 取回历史帧或放大矩形区域;模型决定查哪里。 An inspect call retrieves earlier frames or enlarges rectangular regions; the model chooses where to look. |
| World model World model | 模型在自然语言推理及笔记中形成的可修正规则,而非需要运行的游戏模拟程序。 Revisable rules expressed in the model’s reasoning and notes, rather than an executable game simulator. |
| Trajectory Trajectory | 一局游戏中按时间排列的动作、返回画面、检查调用和笔记更新。 The sequence of actions, returned images, inspection calls, and note updates within a game. |

1. 两层记忆:原始证据与可修改的解释
1. Two memory layers: original evidence and revisable interpretations
所有环境返回帧都进入档案;默认只把最后一帧作为下一步观察。GUIDE.md 记录跨关卡的规则理解,WORKING.md 记录当前关卡的状态和计划。上下文接近上限时,模型保存交接摘要并在新上下文中继续;原始帧与动作历史仍然可用。
Every returned frame enters the archive; by default, only the final frame becomes the next observation. GUIDE.md stores rule knowledge across levels, while WORKING.md holds the current level’s state and plans. Near the context limit, the model saves a handoff summary and continues in a fresh context, with original frames and action history still available.
“无损”描述的是已返回帧的存储,不是模型的视觉编码、理解或笔记。它既不恢复环境未提供的画面,也不保证模型能找到正确帧。放大图像不会创造新像素信息,但可以让局部细节在模型输入中占更大的面积。
“Lossless” describes storage of frames that were returned, not the model’s visual encoding, understanding, or notes. It neither recovers observations the environment never supplied nor guarantees that the model retrieves the right frame. Enlarging an image creates no new pixel information, but can give a local detail more space in the model’s input.
2. 三种主要工具
2. Three main tools
| 工具 Tool | 输入与结果 Input and result | 是否推进游戏 Advances the game? |
|---|---|---|
play |
执行一个允许的动作;返回最后一帧,并存下所有中间帧。 Executes one allowed action, returns the final frame, and archives all intermediate frames. | 是 Yes |
inspect |
用 turn、frame 和区域定位一个或多个视图,裁剪并放大后返回。question 字段说明此次检查要解决的问题。 Selects one or more views by turn, frame, and region, then returns enlarged crops. A question field states what the inspection should resolve. | 否 No |
read_pixels |
读出选定区域的精确颜色或像素数值,帮助区分细小标记和网格边界。 Returns exact color or pixel values in a selected region to resolve small markers or grid boundaries. | 否 No |
这里的“无需 program synthesis”是指不让模型编写并执行求解器或世界模拟器;并不是整个系统没有代码,也不是只允许图像输入。read_pixels 明确提供数值信息,底层控制器负责所有工具的执行。 Appendix A.4
“Without program synthesis” means the model does not write and execute a solver or world simulator. The system still uses software, and its inputs are not exclusively images: read_pixels explicitly supplies numerical information, while the controller executes all tools. Appendix A.4
3. 一次决策怎样进行
3. How one decision unfolds
- 观察当前帧和可用动作,必要时读取笔记。 Observe the current frame and available actions, consulting notes if useful.
- 提出假设;若证据不足,回看旧帧、比较动画阶段或放大细节。检查可以重复多次。 Form a hypothesis; if evidence is insufficient, retrieve earlier frames, compare animation stages, or enlarge details. Inspection may repeat.
- 说明预计会看到的后果,执行一个动作,再比较实际变化与预期。 State the expected consequence, execute one action, and compare actual changes with that prediction.
- 更新规则和当前计划;新返回的全部帧进入档案,最后一帧成为下一次观察。 Update rules and the current plan; archive all newly returned frames and use the last frame as the next observation.

最关键的一句话:保留可以重新检查的证据,比只保留曾经对证据作出的解释更可靠;这是一种工具层面的选择与重看机制,不是新增神经网络 attention 层。
The key idea: keeping evidence available for reinspection is more dependable than retaining only an earlier interpretation. This is selection and reinspection through tools, not an added neural attention layer.
主要实验结果:收益来自哪些改变
Results: which changes account for the gains
公开 ARC-AGI-3 上的系统比较
System comparison on public ARC-AGI-3
| 系统 System | 模型 Model | 推理设置 Reasoning effort | RHAE |
|---|---|---|---|
| 官方实现 Official implementation | GPT-5.6 Sol | max | 13.33 |
| 官方实现 Official implementation | Claude Opus 5.0 | high | 40.68 |
| VISTA | GPT-5.6 Sol | max | 99.00 |
| VISTA | Claude Opus 5.0 | xhigh | 100.00 |
| Tycho(使用 program synthesis) Tycho (uses program synthesis) | GPT-5.6 Sol | max | 100.00 |
VISTA 的两个主模型都完成了 25 个游戏。Opus 使用 7,302 次动作,人类参照为 17,135;减少比例为 1 − 7,302 / 17,135 ≈ 57.4%。GPT 使用 9,126 次动作。人类参照是逐关首次成功玩家动作数的上中位数之和,不是同一个人玩完全部游戏的动作总数。 Table 1 Table 2
Both main VISTA models finish all 25 games. Opus uses 7,302 actions against a human reference of 17,135, a reduction of 1 − 7,302 / 17,135 ≈ 57.4%. GPT uses 9,126 actions. The human reference sums per-level upper-median action counts among successful first-time players; it is not one person’s total across all games. Table 1 Table 2
比较边界:Opus 的官方结果使用 high,VISTA 使用 xhigh;官方基线通过 API 运行,论文实验通过 CLI 运行。因而 40.68 → 100.00 或 13.33 → 99.00 都是系统比较,不能全部归因于视觉记忆。Tycho 同样达到 100.00,所以贡献是展示另一条有效路线,而非独占最高分。
Comparison boundary: the official Opus result uses high effort, while VISTA uses xhigh. Official baselines use API access; the paper’s experiments use CLI access. Thus 40.68 → 100.00 and 13.33 → 99.00 are system comparisons, not isolated effects of visual memory. Tycho also reaches 100.00, so the contribution demonstrates another effective approach rather than exclusive possession of the top score.
RHAE 的 100 到底表示什么?
What does an RHAE of 100 mean?
已完成关卡的分数是 min(1.15, (h/a)²),其中 h 是人类参照动作数、a 是模型动作数;未完成关卡得 0。第 ℓ 关权重为 ℓ,游戏分数既受加权效率限制,也受已完成关卡的加权比例限制。最后对 25 个游戏取平均并乘 100。
A completed level scores min(1.15, (h/a)²), where h is the human reference and a the agent action count; an unfinished level scores 0. Level ℓ has weight ℓ. A game’s score is limited by both weighted efficiency and the weighted fraction of levels completed. The final score averages the 25 games and multiplies by 100.
因此 100 并不要求每个关卡的动作数都少于人类,也不衡量思考时间、token 或存储空间。内部检查不消耗游戏动作,可能用更多计算换取更少探索动作。 Appendix A.3
Consequently, 100 does not require fewer actions than humans on every individual level, and it does not measure thinking time, tokens, or storage. Internal inspection consumes no game actions and may trade additional computation for fewer exploratory actions. Appendix A.3
逐步消融:最有解释力的证据
Incremental ablation: the most informative evidence
以下各阶段使用 GPT-5.6 Sol / max。阶段 (i)–(vi) 使用相同的模型接入方式;“官方文本基线”是外部参照。表中动作数是 25 个游戏的总和,完成率按整局游戏统计。 Figure 5 Appendix A.2
All stages below use GPT-5.6 Sol / max. Stages (i)–(vi) share the same model access interface; the official textual baseline is an external reference. Action counts sum over 25 games, and completion rates count fully completed games. Figure 5 Appendix A.2
| 配置 Configuration | RHAE | 动作数 Actions | 游戏完成率 Games completed |
|---|---|---|---|
| 官方文本基线 Official text baseline | 13.33 | 10,619 | 4% |
| (i) 替换为图像 (i) Image observations | 47.32 | 16,343 | 36% |
| (ii) 增加动作与时间上限 (ii) Larger action and time limits | 51.66 | 22,544 | 44% |
| (iii) 连续对话与上下文压缩 (iii) Continuous conversation and compaction | 65.82 | 28,626 | 60% |
| (iv) 加入笔记与规则建模提示 (iv) Notes and rule-modeling prompt | 70.05 | 23,702 | 76% |
| (v) 无损记忆与主动检查 (v) Lossless memory and inspection | 94.10 | 12,261 | 96% |
| (vi) 精确像素读数:完整 VISTA (vi) Exact pixel readout: full VISTA | 99.00 | 9,126 | 100% |
从 (iv) 到 (v),RHAE 增加 24.05 点,动作数从 23,702 降至 12,261;加入像素读数再增加 4.90 点。但 (v) 同时把自动提供最多七帧改为只提供最后一帧并按需检查,因此此消融验证的是一组协同设计,并未把存储、检索、裁剪和输入调度逐一分离。
From (iv) to (v), RHAE rises by 24.05 points and actions fall from 23,702 to 12,261; pixel readout adds another 4.90 points. However, (v) also replaces automatic delivery of up to seven frames with the final frame plus on-demand inspection. This ablation therefore tests a package of cooperating choices, without separately isolating storage, retrieval, cropping, and input scheduling.
早期配置的低动作数不能直接理解为更高效率:很多运行因达到上限而提早结束。(ii) 把限制从逐段人类动作数的 5× 与 12 小时,调整为每游戏最多 2,000 次动作与 48 小时。
Low action counts in earlier configurations do not necessarily imply efficiency: many runs terminate early at their limits. Stage (ii) changes limits from 5× the human actions per section and 12 hours to at most 2,000 actions and 48 hours per game.
跨环境结果与重复实验
Other environments and repeated runs
| 任务与范围 Task and scope | 指标 Metric | 同模型官方基线 Same-model official baseline | VISTA | 人类参照 Human reference |
|---|---|---|---|---|
| GameWorld:34 游戏 / 170 任务 GameWorld: 34 games / 170 tasks | 成功率 / 进度 Success rate / progress | 40.0% / 62.9% | 63.3% / 79.3% | 55.3% / 64.1% |
| AI GameStore:10 游戏 AI GameStore: 10 games | 人类归一化分数的几何平均 Geometric mean of human-normalized scores | 47.3 | 140.3 | 100 |
| BabyVision:39 道视觉追踪题 BabyVision: 39 visual-tracking questions | 准确率 Accuracy | 41.0% | 63.2% | 92.3% |
以上模型均为 GPT-5.6 Sol。BabyVision 只评估 20 道迷宫与 19 道连线题,不是全部 388 题;它没有交互历史,因此此处收益支持的是局部视觉检查。GameWorld 的人类参照来自一位未接触过这些任务的玩家;AI GameStore 使用人类中位数,140.3 是归一化聚合分数,不是胜率。 §5 Appendix A.6
All model results here use GPT-5.6 Sol. BabyVision evaluates only 20 maze and 19 connecting-lines questions, not all 388 tasks; there is no interaction history, so the gain supports local visual inspection. GameWorld’s human reference is one player without prior exposure to the tasks. AI GameStore uses human medians: 140.3 is a normalized aggregate score, not a win rate. §5 Appendix A.6
三次运行中,GPT 的 ARC-AGI-3 RHAE 为 99.00、98.90、99.23,均值 ± 标准差为 99.04 ± 0.17。GameWorld 成功率为 63.33 ± 2.37%,BabyVision 准确率为 63.23 ± 3.00%。这些是运行间标准差,不是置信区间。AI GameStore 先对每游戏的三次原始分数取平均,再归一化聚合;不能直接平均三个最终总分。 Appendix B.1
Across three runs, GPT’s ARC-AGI-3 RHAE is 99.00, 98.90, and 99.23, giving a mean ± standard deviation of 99.04 ± 0.17. GameWorld success is 63.33 ± 2.37%, and BabyVision accuracy is 63.23 ± 3.00%. These are run-to-run standard deviations, not confidence intervals. AI GameStore first averages each game’s raw scores over three runs and then normalizes and aggregates; directly averaging the three final aggregate scores is not the reported procedure. Appendix B.1
附加实验提供了支持,也显示剩余差距:GLM-5.3 Flash 320B 的 RHAE 从官方实现的 1.89 提高到完整 VISTA 的 66.93;移除 inspect 与 read_pixels 的版本为 13.27。把 ARC 游戏改为 3D 渲染后,GPT 的得分降到 84.12,动作数上升至 20,640。 Figures 8–9
Additional experiments support the approach while exposing remaining gaps: GLM-5.3 Flash 320B rises from 1.89 RHAE with the official implementation to 66.93 with full VISTA; a variant without inspect and read_pixels scores 13.27. With ARC games rendered in 3D, GPT falls to 84.12 and uses 20,640 actions. Figures 8–9
效率与计算成本:动作少不等于计算少
Efficiency and compute: fewer actions need not mean less computation
保持其余 harness 不变时,文本网格仍达到 96.75 RHAE,图像为 99.00;两者分别使用 10,614 与 9,126 次动作。更明确的差别是论文报告的每游戏 token 用量:71.9M 对 30.7M。这个结果支持图像表示的效率优势,同时表明记忆与交互设计并不只对图像有效。 Figure 6
With the rest of the harness unchanged, textual grids still achieve 96.75 RHAE versus 99.00 for images, using 10,614 versus 9,126 actions. The clearer difference is reported token use per game: 71.9M versus 30.7M. This supports the efficiency of image representations while showing that memory and interaction design also benefit textual inputs. Figure 6
| 最大上下文 Maximum context | RHAE | 总动作 Total actions | 每游戏 token Tokens per game |
|---|---|---|---|
| 100K | 98.6 | 8,395 | 17.3M |
| 200K | 99.0 | 9,126 | 30.7M |
| 500K | 94.4 | 12,391 | 84.1M |
| 780K | 93.9 | 12,321 | 105.7M |
上述 token 是论文报告的累计用量,不能当作单次上下文长度,也不能直接换算美元成本。图 7 显示更大上下文没有带来更高分;100K 在接近最高分时显著减少 token。图像放大到 4× 的结果为 99.7 RHAE、28.2M token,而默认 8× 为 99.0、30.7M;16× 则降到 88.3,并消耗 71.1M。放大倍率实验不应被解读为已证明一个跨模型通用的最优值。 Figure 7
These are the paper’s cumulative token counts, not single-request context lengths or directly convertible dollar costs. Figure 7 shows no score benefit from larger contexts; 100K substantially reduces tokens while remaining near the best score. At 4× image enlargement, the result is 99.7 RHAE and 28.2M tokens, versus 99.0 and 30.7M at the default 8×; 16× drops to 88.3 and uses 71.1M. This scale experiment does not establish a universal optimum across models. Figure 7
局限与展望:如何评价真正的贡献
Limitations and outlook: assessing the contribution
怀疑者的归纳与有价值的新点
The skeptical reduction and the useful contribution
怀疑者会说:这就是强多模态模型加截图档案、裁剪工具和持久笔记,并非新的推理算法。这个归纳基本成立。有价值的新点:把完整原始视觉历史、模型自主选择的时空检查以及跨上下文延续组合成一个简单接口,并用逐步消融证明这套输入与记忆管理可以明显改变同一模型的表现。其贡献主要是 harness 设计及实证,而不是训练方法。
A skeptic would say: this is a strong multimodal model plus a screenshot archive, cropping tools, and persistent notes, rather than a new reasoning algorithm. That is broadly fair. The useful contribution: a simple interface combines complete original visual history, model-directed inspection across space and time, and continuity across contexts, with incremental ablations showing that this input and memory management substantially changes the same model’s performance. The contribution is primarily harness design and empirical evidence, rather than training methodology.
需要保留的四个判断边界
Four boundaries on the conclusions
- 训练数据污染尚未排除。作者明确指出模型晚于公开 benchmark 发布,游戏可能出现在训练数据中;需要私有或新生成游戏检验泛化。 Training contamination remains unresolved. The authors explicitly note that the models were released after the public benchmarks, whose games may have appeared in training data. Private or newly generated games are needed to test generalization.
- 外部基线并非完全受控比较。推理设置、接入方式、时限和动作预算都有差异;逐步消融更能定位机制,但仍是累加实验而非完整的独立因素实验。 External baselines are not fully controlled comparisons. Reasoning settings, access interfaces, time limits, and action budgets differ. Incremental ablations locate mechanisms more clearly, but remain cumulative rather than fully factorial experiments.
- “超越人类”只适用于特定聚合指标与参照群体。动作效率不计推理调用开销;GameWorld 的新手参照样本很小,BabyVision 上仍远低于人类。 “Beyond human” applies only to particular aggregate metrics and reference groups. Action efficiency excludes inference-call overhead; GameWorld’s novice reference is very small, and BabyVision remains well below human performance.
- 完整存储不等于完整利用。检索依赖模型知道该查哪一帧,档案会随交互增长;论文未建立存储规模、检索失败率与端到端延迟的系统性成本曲线。 Complete storage does not guarantee complete use. Retrieval depends on the model identifying the right frame, and the archive grows with interaction. The paper does not establish systematic cost curves for archive size, retrieval failures, and end-to-end latency.
前一项是论文明确承认的局限;后面还包括本报告对实验设计和成本的分析。优先后续实验应当在未见游戏上统一模型、推理与预算,分别去掉历史存储、放大、像素读数和笔记,并同时报告成功率、环境动作、推理时间、token 与存储量。 §6
The first point is an explicit limitation acknowledged by the authors; the others also include this report’s assessment of experimental design and cost. Priority follow-ups would hold model, reasoning, and budgets fixed on unseen games, separately remove historical storage, enlargement, pixel readout, and notes, and report success, environment actions, inference time, tokens, and storage together. §6
对实际代理的启发是把“证据保存”与“摘要解释”分开,并允许在不改变环境的情况下核对证据。迁移到 GUI 或机器人是合理的研究方向,但这里的实验还不足以证明真实世界的可靠性或实时控制能力。
The practical lesson for agents is to separate evidence retention from summarized interpretation, and permit checking evidence without changing the environment. Transfer to GUIs or robotics is a reasonable research direction, but these experiments do not establish real-world reliability or real-time control capability.
实现与复核:阅读代码时抓住什么
Implementation and verification: where to look in the code
公开仓库提供 ARC-AGI-3、GameWorld、AI GameStore 与 BabyVision 的运行入口。src/vista/core/tools.py 定义共享的检查、像素读取及笔记工具;src/vista/core/artifacts.py 保存带哈希校验的原始视觉证据;ARC 专用工具位于 src/vista/benchmarks/arc3/。这里检查了公开 README、文件树和上述代码,未运行付费模型或整套 benchmark。
The public repository provides entry points for ARC-AGI-3, GameWorld, AI GameStore, and BabyVision. src/vista/core/tools.py defines shared inspection, pixel readout, and note tools; src/vista/core/artifacts.py stores original visual evidence with hash verification; ARC-specific tools live under src/vista/benchmarks/arc3/. This study inspected the public README, file tree, and those source files; it did not run paid models or full benchmarks.
论文附录说明,模型通过结构化调用使用白名单工具,控制器串行推进游戏;“inspect 中的问题”用于表达取证目的,并不会调用另一个视觉问答模型。复现实验时应固定仓库 commit、模型版本、上下文和图像尺寸,而不能假设最新代码的默认值就是论文配置。
The appendix describes structured calls to allowlisted tools and a controller that advances the game sequentially. The question supplied to inspect expresses the purpose of evidence gathering; it does not invoke another visual question-answering model. Reproduction should pin the repository commit, model version, context limit, and image size rather than assume current code defaults match the paper.
资料来源与版本说明
Sources and version notes
主要依据是论文 v1 的正文、图 5–11 及附录 A–B。项目博客在 2026-08 发布,部分旧摘要使用 56.0% 的动作减少值;本报告使用较新的论文所列 7,302 / 17,135 与 57.4%,不混用不同版本。研究日期为 2026-10-04。
The primary basis is paper v1, including the main text, Figures 5–11, and Appendices A–B. The project blog appeared in 2026-08, and some earlier summaries used a 56.0% action reduction. This report uses the newer paper’s 7,302 / 17,135 and 57.4%, without mixing versions. Research date: 2026-10-04.
- arXiv 摘要与版本;论文 HTML;本地 PDF。 arXiv abstract and version; paper HTML; local PDF.
- 作者项目网站:直观介绍与演示。 Authors’ project site: introduction and demonstrations.
- 作者代码仓库:接口与运行说明。 Authors’ code repository: interfaces and execution instructions.
三幅原图均提取自论文 HTML,图内文字作为共享原始资料保留;每幅图的解读与图注均提供完整中英文版本。
The three original figures were extracted from the paper HTML. Text inside them is retained as shared source material; their explanations and captions are provided fully in both languages.