Sign InOpen Brain
OpenAIOfficial ReleaseOfficial Source

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Two API settings—reasoning retention and compaction—reportedly tripled GPT-5.6’s ARC-AGI-3 score. Agent evals should treat runtime configuration as part of the tested system.

OpenAI · Jul 29, 2026
Open Source Open MarkdownOpen JSON
Source Summary

OpenAI says enabling **reasoning retention** and **compaction** produced **3× ARC-AGI-3 scores** for GPT-5.6 while also improving efficiency.

Practical Implication

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

Agent-Ready Context
OpenAI says enabling **reasoning retention** and **compaction** produced **3× ARC-AGI-3 scores** for GPT-5.6 while also improving efficiency.

Record these settings alongside the model name in agent evaluations. Configuration can materially affect results, so defaults and explicit settings should not be compared as equivalent systems.

The supplied material provides no absolute scores, token usage, latency, or experimental detail, leaving the size and generality of the efficiency gain unclear.
Connected Context · Feed7 Judgment

This turns evaluation configuration into part of the system being measured: the same named model can produce materially different results when reasoning retention and compaction change. It reinforces prior warnings that leaderboard comparisons are invalid without disclosed conditions, while leaving the gain’s efficiency and generality unresolved because absolute scores, resource use, and experimental details are absent.

Context Map
benchmark#agent-evals#benchmark-integrity#context-caching
Uncertainty
The supplied material provides no absolute scores, token usage, latency, or experimental detail, leaving the size and generality of the efficiency gain unclear.