The model did not change. The test did

On ARC-AGI-3, memory and context management mattered alongside the model name.

The same AI model cube navigating a maze that changes as memory and context layers are added.
Share this article

Subject: The model did not change. The test did

Preview: On ARC-AGI-3, memory and context management mattered alongside the model name.

An analysis reported by TLDR AI says two settings - retained reasoning and context compaction - tripled GPT-5.6 Sol's ARC-AGI-3 score while sharply reducing output. ARC Prize independently confirms the final 7.78% score, but its results page does not document that configuration comparison.

What happened

ARC-AGI-3 places agents in unfamiliar environments where they must observe, form hypotheses, learn rules and plan. The model is only one component. Memory, context management, tools and verification loops all shape the result.

Why it matters

Your organisation does not buy a benchmark; it deploys a complete chain. An average model with sound memory and testing can outperform a stronger model placed inside a weak system.

What is easy to miss

An agent score often measures both the model and the evaluation harness. Comparisons that change several variables deserve caution. The reported tripling remains attributed to the analysis cited by the newsletter.

What to do next

Version prompts, tools, memory rules and stop criteria. Measure quality, cost and failures for the full workflow. When changing models, keep the rest constant first, then optimise one component at a time.

The takeaway

An agent benchmark often measures the environment as much as the intelligence at its centre.

Sources

  1. Primary sourcemail.google.com
  2. Commentaryarcprize.org
  3. Primary sourceopenai.com

Editorial methodology

Last reviewed: · By Arnaud Llamas Bravo

Tell me what your team needs.

Share the essentials and I’ll get back to you with the most useful next step.