The Circuitry
THE CIRCUITRYYour one-stop source for all tech news
HOMETODAYNEWSFEEDEVENTS
BOOKMARKS
RSS
© 2026 The Circuitry
About UsSourcesContactCorrectionsPrivacy
  • Today
  • Feed
  • Events
  • Saved
Scroll for more
Verification
VERIFIEDConfidence: HIGH
Source identified
Claims cross-referenced
No discrepancies found
Fact-check summary

ARC Prize's official Sep 3 blog and results page confirm GPT-6 Astra's 62.7% Standard and 99.9% Provider Adapter ARC-AGI-3 scores, matching the article exactly.

Sourcing
1source

via Arcprize

From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Home/Tech/OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3
VERIFIEDBy Xavier Rivera· ·1.5 min read

OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3

OpenAI's GPT-6 Astra has posted state-of-the-art scores of 62.7% and 99.9% on ARC-AGI-3 depending on the harness used. The results narrow the measured gap to human-level agentic intelligence on a benchmark designed to track progress toward AGI.

Source:Arcprize
Post
OpenAI GPT-6 Astra Scores 62.7% on ARC-AGI-3
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
TL;DRAI · 60 sec read

OpenAI's GPT-6 Astra scores 62.7 percent on ARC-AGI-3 with a standard harness for 26 thousand dollars and 99.9 percent with a provider adapter harness for 19 thousand dollars. The model exceeds median human action efficiency on 96 percent of levels. This sets a new state of the art on the benchmark that tracks progress toward human-level skill acquisition.

OpenAI's GPT-6 Astra has achieved state-of-the-art results on the ARC-AGI-3 benchmark, scoring 62.7% for $26K using a Standard harness and 99.9% for $19K with a Provider Adapter harness.

GPT-6 Astra surpasses human baseline in action efficiency. The model used fewer actions than the median tested human on 96% of ARC-AGI-3 levels. It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions.
The model used fewer actions than the median tested human on 96% of ARC-AGI-3 levels.

ARC-AGI-3 measures agentic intelligence components. The benchmark tests exploration, modeling, goal-setting, and planning and execution through novel, abstract, turn-based environments. Agents must actively obtain information, build generalizable models, identify target states with sparse rewards, and map paths while course correcting.
POST FROM @arcprize· official announcement and analysis tweet directly referencing the ARC Prize blog post on GPT-6 Astra ARC-AGI-3 results
https://x.com/arcprize/status/2095597602545025138

These environments contain only core knowledge priors and are calibrated through controlled testing with human participants, who can solve 100% of them. ARC-AGI-3 expands on prior generations as frontier AI capabilities advance. The series aims to measure the residual gap to AGI, defined as acquiring any human skill as efficiently as a human.
From The CircuitryThe Feed — live briefs across tech, all day.See what’s happening →

Results vary by reasoning effort and harness type. With the Standard harness, which carries forward notes chosen by the model, Astra (max) scores 62.7% for $26,098 while Astra (high) scores 54.8% for $40,705. Higher reasoning levels generally cost less because Astra solves games in fewer actions, reducing model calls and tokens.

Using the Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction, Astra (high) reaches 99.9% for $18,817 and Astra (max) reaches 98.6% for $17,332. Both harness configurations set new state-of-the-art marks on the semi-private leaderboard. Full results are available on the ARC Prize site.
It turns unfamiliar environments into compact symbolic world models, representing game mechanics as logical rules and developing its own domain-specific language shorthand to track state and plan actions.
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Cost and efficiency improve at higher reasoning. At max effort, Astra requires fewer actions than at lower levels, lowering total cost relative to medium, low, or none configurations. The benchmark's goal remains tracking progress toward human-level efficiency in acquiring new skills.
Why this mattersAI · ~100 words

Tap a lens to see what this story means for you.

Morning Brief

Liked this? The Brief brings you the whole day in tech, verified, every morning.

Two minutes, free forever. What's in The Brief →

Reader-supported
DonateBuy me a coffee →Follow@thecircuitry_ →Follow@thecircuitry.to →
HELP US IMPROVE
From The Circuitry

See what’s happening right now

The Feed runs all day — short, verified briefs the moment they break.

Open the Feed →
From The Circuitry

Follow @thecircuitry_

Every story we publish, as it happens. No noise between.

Follow on X ↗On Bluesky ↗

Reader-supported

The Circuitry is a passion project I've always wanted to build, and I love the work behind it.

Running it costs real money. APIs, hosting, time. To keep improving the site and growing this into something useful for everyone, those costs have to be covered.

Any contribution is appreciated. If not, no pressure. Thanks for reading.

Buy me a coffee
OpenAIAIBenchmark
More inTech
  • GlobalFoundries signs $2B TSMC deal for US silicon interposers

    Tech · 8h
  • SpaceX agrees to buy 800 MHz spectrum for Starlink Mobile

    Tech · 11h
  • Microsoft discloses CVE-2026-83947 in Azure Event Grid

    Tech · 11h
SupportThe Work

The Circuitry is reader-supported. If you find the daily brief useful, you can buy me a coffee to keep it going.

Buy a coffee →
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →

MORE IN THIS BEAT

All Tech →
  • Tech· 

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    OpenAI released 722 mathematical manuscripts, grouped into 372 result families, produced by an unreleased internal model. The papers are on GitHub, with Lean formalizations for many but not all of them, and OpenAI warns some unformalized results could have issues.

  • Tech· 

    Anthropic Launches Cyber Mission to Secure Infrastructure and Open-Source Code

    Anthropic has launched the Anthropic Cyber Mission to support defenders of critical infrastructure and open-source software with models, engineers, and tools. The initiative starts with the Critical Infrastructure Defense Program and free OSS Scanner amid ongoing challenges in verifying and fixing vulnerabilities.

  • Tech· 

    Anthropic launches Claude Haiku 5.5, cutting prices up to 90% from Haiku 4.5

    Anthropic released Claude Haiku 5.5, its cheapest and fastest small model, at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Anthropic says it costs around 75% less to run than Haiku 4.5 on average and scores far higher on its benchmarks.

  • Tech· 

    Meta and Microsoft Slash Employee Claude AI Spending

    Meta and Microsoft are cutting employee use of Anthropic’s Claude AI and directing staff toward their own coding tools, The Information reported. The changes reflect tighter internal AI budgets while customer access to Claude through Microsoft platforms continues to expand.

  • Tech· 

    Microsoft preparing local MAI-Code-1.1-Flash for high-end PCs

    Microsoft has introduced a local edition of MAI-Code-1.1-Flash that runs on Windows PCs without server access. The model requires substantial memory and is coming to GitHub Copilot in experimental preview by the end of October, starting with Nvidia RTX Spark PCs.