Story thread · 3 reports / 2 sources
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
the-decoder.com · 12d
How the coverage leans
Across 2 sources · syndicated copies counted once
OpenAI counters Anthropic's ARC-AGI-3 record: GPT-5.6 Sol scores 38.3 percent, but only with its own API features instead of the official test setup, where the model landed at 7.8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the comparison with Opus 5. The article OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings appeared first on The Decoder .
First report: How enabling two settings tripled our scores on the ARC-AGI-3 benchmark — openai.com, 13d
The conversation · 0
Sign in to join the conversation.
No comments yet — start the thread.