The Circuitry
THE CIRCUITRYYour one-stop source for all tech news
HOMETODAYNEWSFEEDEVENTS
BOOKMARKS
RSS
© 2026 The Circuitry
About UsSourcesContactCorrectionsPrivacy
  • Today
  • Feed
  • Events
  • Saved
Scroll for more
Verification
VERIFIEDConfidence: HIGH
Source identified
Claims cross-referenced
No discrepancies found
Fact-check summary

Google's official blog, NVIDIA, The New Stack, and Hugging Face confirm the June 10, 2026 DiffusionGemma release: a 26B MoE diffusion model delivering up to 4x faster local text generation.

Sourcing
1source

via Ars Technica

Ars Technica · track record
30Stories
100%Verified
130d
All sources →
Markets
GOOGL···

Live quote · not investment advice

From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Home/Tech/Google DeepMind releases DiffusionGemma for 4x faster local AI
VERIFIEDBy Xavier Rivera· ·2 min read

Google DeepMind releases DiffusionGemma for 4x faster local AI

Google DeepMind released DiffusionGemma, a parallel text-generation model that produces up to four times more tokens per second than similarly sized autoregressive Gemma models on local GPUs. The approach trades higher error rates for better compute efficiency on non-linear tasks but remains experimental.

Source:Ars Technica
Post
Google DeepMind releases DiffusionGemma for 4x faster local AI
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
TL;DRAI · 60 sec read

Google DeepMind releases DiffusionGemma, a 26-billion-parameter Mixture of Experts model that generates text in parallel via diffusion rather than token by token. It reaches 700-1,000 tokens per second on RTX 5090 and H100 GPUs, four times faster than autoregressive Gemma models, by shifting the bottleneck from memory to compute on local hardware.

Google DeepMind has released DiffusionGemma, a new member of the Gemma 4 open model family that generates text in parallel rather than one token at a time.

DiffusionGemma uses a parallel generation approach borrowed from image models. Unlike autoregressive models that produce text left to right one token at a time, DiffusionGemma starts with a field of placeholder tokens and runs over the canvas multiple times to generate likely tokens. It uses those to improve estimation of others before finalizing outputs in one large block of denoised text. Google says this makes the model faster and more efficient on local hardware such as an Nvidia DGX or a gaming GPU.
In language, a single bad token can render an entire block meaningless and require restarting, unlike image generation where one flawed pixel rarely ruins the result.

The model is a 26-billion-parameter Mixture of Experts design. Only 3.8 billion parameters activate during inference, allowing it to fit in the 18GB RAM of a high-end GPU. In testing on an RTX 5090, DiffusionGemma produces around 700 tokens per second. On a single Nvidia H100, it reaches 1,000-plus tokens per second, roughly four times the output of similarly sized autoregressive Gemma models.
POST FROM @GoogleDeepMind· official announcement tweet matching the article topic and date
https://x.com/GoogleDeepMind/status/2064741061352636762

Parallel generation shifts the performance bottleneck from memory to compute. The model can generate up to 256 tokens at once. Google reports measurable gains on non-linear tasks including in-line editing, molecular sequencing and mathematical graphing. An animation in the release demonstrates how DiffusionGemma solves Sudoku puzzles by continuously self-correcting large sets of tokens, a task that challenges standard autoregressive models because each token depends on future ones.
From The CircuitryThe Feed — live briefs across tech, all day.See what’s happening →

Text diffusion carries trade-offs that limit its use in cloud models. Google has experimented with the technique for its Gemini models but notes a higher error rate. In language, a single bad token can render an entire block meaningless and require restarting, unlike image generation where one flawed pixel rarely ruins the result. Diffusion models also waste resources on short outputs that autoregressive models can complete in just a few steps.
An animation in the release demonstrates how DiffusionGemma solves Sudoku puzzles by continuously self-correcting large sets of tokens, a task that challenges standard autoregressive models because each token depends on future ones.
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Local hardware benefits more from diffusion than cloud systems do. Cloud autoregressive models batch jobs across users and leverage high-bandwidth memory to stay efficient. Local AI often faces idle cycles and lower memory bandwidth. Diffusion makes better use of available compute, outperforming even Google's Multi-Token Prediction drafters that also target wasted cycles. Google describes DiffusionGemma as experimental.
Why this mattersAI · ~100 words

Tap a lens to see what this story means for you.

Morning Brief

Liked this? The Brief brings you the whole day in tech, verified, every morning.

Two minutes, free forever. What's in The Brief →

Reader-supported
DonateBuy me a coffee →Follow@thecircuitry_ →Follow@thecircuitry.to →
HELP US IMPROVE
From The Circuitry

See what’s happening right now

The Feed runs all day — short, verified briefs the moment they break.

Open the Feed →
From The Circuitry

Follow @thecircuitry_

Every story we publish, as it happens. No noise between.

Follow on X ↗On Bluesky ↗

Reader-supported

The Circuitry is a passion project I've always wanted to build, and I love the work behind it.

Running it costs real money. APIs, hosting, time. To keep improving the site and growing this into something useful for everyone, those costs have to be covered.

Any contribution is appreciated. If not, no pressure. Thanks for reading.

Buy me a coffee
AIGoogleDeepMindGemma
More fromArs Technica
  • Vendors object to Spirit Airlines selling operational data to Google in bankruptcy

    Tech · 28d
  • Twitch adds opt-out for Amazon’s generative AI training on user content

    Tech · 1mo
  • Aptoide becomes first rival app store hosted in Google Play

    Tech · 1mo
More inTech
  • GlobalFoundries signs $2B TSMC deal for US silicon interposers

    Tech · 8h
  • SpaceX agrees to buy 800 MHz spectrum for Starlink Mobile

    Tech · 11h
  • Microsoft discloses CVE-2026-83947 in Azure Event Grid

    Tech · 11h
SupportThe Work

The Circuitry is reader-supported. If you find the daily brief useful, you can buy me a coffee to keep it going.

Buy a coffee →
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →

MORE IN THIS BEAT

All Tech →
  • Tech· 

    Google DeepMind and Meta back Biohub virtual cell with $300M

    Google DeepMind, Meta, and Isomorphic Labs are investing $300 million in Biohub to build AI datasets for a virtual cell model. The funding is part of a $1.8 billion initiative that also includes contributions from the Department of Energy and National Institutes of Health.

  • Tech· 

    Anthropic Launches Cyber Mission to Secure Infrastructure and Open-Source Code

    Anthropic has launched the Anthropic Cyber Mission to support defenders of critical infrastructure and open-source software with models, engineers, and tools. The initiative starts with the Critical Infrastructure Defense Program and free OSS Scanner amid ongoing challenges in verifying and fixing vulnerabilities.

  • Tech· 

    Anthropic launches Claude Haiku 5.5, cutting prices up to 90% from Haiku 4.5

    Anthropic released Claude Haiku 5.5, its cheapest and fastest small model, at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens. Anthropic says it costs around 75% less to run than Haiku 4.5 on average and scores far higher on its benchmarks.

  • Tech· 

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    OpenAI released 722 mathematical manuscripts, grouped into 372 result families, produced by an unreleased internal model. The papers are on GitHub, with Lean formalizations for many but not all of them, and OpenAI warns some unformalized results could have issues.

  • Tech· 

    Meta and Microsoft Slash Employee Claude AI Spending

    Meta and Microsoft are cutting employee use of Anthropic’s Claude AI and directing staff toward their own coding tools, The Information reported. The changes reflect tighter internal AI budgets while customer access to Claude through Microsoft platforms continues to expand.