The model did not change. The test did
On ARC-AGI-3, memory and context management mattered alongside the model name.

Subject: The model did not change. The test did
Preview: On ARC-AGI-3, memory and context management mattered alongside the model name.
An analysis reported by TLDR AI says two settings - retained reasoning and context compaction - tripled GPT-5.6 Sol's ARC-AGI-3 score while sharply reducing output. ARC Prize independently confirms the final 7.78% score, but its results page does not document that configuration comparison.
What happened
ARC-AGI-3 places agents in unfamiliar environments where they must observe, form hypotheses, learn rules and plan. The model is only one component. Memory, context management, tools and verification loops all shape the result.
Why it matters
Your organisation does not buy a benchmark; it deploys a complete chain. An average model with sound memory and testing can outperform a stronger model placed inside a weak system.
What is easy to miss
An agent score often measures both the model and the evaluation harness. Comparisons that change several variables deserve caution. The reported tripling remains attributed to the analysis cited by the newsletter.
What to do next
Version prompts, tools, memory rules and stop criteria. Measure quality, cost and failures for the full workflow. When changing models, keep the rest constant first, then optimise one component at a time.
The takeaway
An agent benchmark often measures the environment as much as the intelligence at its centre.
Sources
- Primary sourcemail.google.com
- Commentaryarcprize.org
- Primary sourceopenai.com
Last reviewed: · By Arnaud Llamas Bravo


