OpenAI's GPT-6 Astra has posted state-of-the-art scores of 62.7% and 99.9% on ARC-AGI-3 depending on the harness used. The results narrow the measured gap to human-level agentic intelligence on a benchmark designed to track progress toward AGI.

The model used fewer actions than the median tested human on 96% of ARC-AGI-3 levels.
https://x.com/arcprize/status/2095597602545025138
It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions.
Tap a lens to see what this story means for you.
Liked this? The Brief brings you the whole day in tech, verified, every morning.
Two minutes, free forever. What's in The Brief →
See what’s happening right now
The Feed runs all day — short, verified briefs the moment they break.
Open the FeedFollow @thecircuitry_
Every story we publish, as it happens. No noise between.
Reader-supported
The Circuitry is a passion project I've always wanted to build, and I love the work behind it.
Running it costs real money. APIs, hosting, time. To keep improving the site and growing this into something useful for everyone, those costs have to be covered.
Any contribution is appreciated. If not, no pressure. Thanks for reading.
OpenAI released 722 mathematical manuscripts, grouped into 372 result families, produced by an unreleased internal model. The papers are on GitHub, with Lean formalizations for many but not all of them, and OpenAI warns some unformalized results could have issues.
Anthropic has launched the Anthropic Cyber Mission to support defenders of critical infrastructure and open-source software with models, engineers, and tools. The initiative starts with the Critical Infrastructure Defense Program and free OSS Scanner amid ongoing challenges in verifying and fixing vulnerabilities.
Anthropic released Claude Haiku 5.5, its cheapest and fastest small model, at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Anthropic says it costs around 75% less to run than Haiku 4.5 on average and scores far higher on its benchmarks.
Meta and Microsoft are cutting employee use of Anthropic’s Claude AI and directing staff toward their own coding tools, The Information reported. The changes reflect tighter internal AI budgets while customer access to Claude through Microsoft platforms continues to expand.
Microsoft has introduced a local edition of MAI-Code-1.1-Flash that runs on Windows PCs without server access. The model requires substantial memory and is coming to GitHub Copilot in experimental preview by the end of October, starting with Nvidia RTX Spark PCs.