The Circuitry
THE CIRCUITRYYour one-stop source for all tech news
HOMETODAYNEWSFEEDEVENTS
BOOKMARKS
RSS
© 2026 The Circuitry
About UsSourcesContactCorrectionsPrivacy
  • Today
  • Feed
  • Events
  • Saved
Scroll for more
Verification
VERIFIEDConfidence: HIGH
Source identified
Claims cross-referenced
No discrepancies found
Fact-check summary

Frandroid reports on Microsoft's Oct 7 Windows/Surface event demo of hybrid AI running local models (MAI Code 1.1 Flash, Nvidia Nemotron, DeepSeek V4 Flash) on RTX Spark hardware, corroborated by BGR liveblog and previews from Windows Central/PCMag.

Sourcing
1source

via Frandroid

Frandroid · track record
59Stories
100%Verified
1030d
All sources →
Markets
MSFT···

Live quote · not investment advice

From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Home/Tech/Microsoft shows hybrid AI running large models locally on Windows
VERIFIEDBy Xavier Rivera· ·1.5 min read

Microsoft shows hybrid AI running large models locally on Windows

Microsoft demonstrated hybrid intelligence on Windows that runs AI models from 70 billion to 284 billion parameters locally. The approach keeps data on-device, supports offline use, and switches to cloud models for heavier tasks.

Source:Frandroid
Post
Microsoft shows hybrid AI running large models locally on Windows
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
⚐ CORRECTED 

Wrong: said MAI Code 1.1 Flash has 130 billion parameters and hedged Microsoft's own model details as "reportedly." Right: Microsoft says MAI Code 1.1 Flash has 137 billion total parameters (6.8 billion active); the model figures come from Microsoft's Windows blog.

→ All corrections
Microsoft has demonstrated hybrid intelligence on Windows that runs large AI models locally while switching to cloud models when needed.

Microsoft named three models for local use on Windows PCs. MAI Code 1.1 Flash has 137 billion total parameters, 6.8 billion of them active, and was quantized to 3 bits, cutting its size by nearly 80 percent while keeping code quality, Microsoft said. A new Nvidia Nemotron model exceeds 70 billion parameters, uses 2-bit quantization, and occupies a little more than 20 GB of memory. DeepSeek V4 Flash, the largest at 284 billion parameters, was reportedly compressed to 1.6 bits to fit in 60 GB of memory.
Local execution keeps data on the device and removes token costs.

Hybrid intelligence lets the system choose between local and cloud models. The local model handles tasks it can manage, and the system routes more demanding work to cloud services such as GPT, Claude, or Gemini. Microsoft said customers want to stretch token budgets without losing access to the strongest models.

Local execution keeps data on the device and removes token costs. Work can continue offline, and queries no longer draw from paid token quotas. DeepSeek V4 Flash is reportedly a mixture-of-experts model with only 13 billion parameters active per token, which limits compute needs even at its full size.
From The CircuitryThe Feed — live briefs across tech, all day.See what’s happening →

Target hardware centers on Nvidia RTX Spark chips. These chips support up to 128 GB of unified memory. Microsoft highlighted the Surface RTX Spark Dev Box as an example machine built for local AI workloads.
Work can continue offline, and queries no longer draw from paid token quotas.
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →
Microsoft presented the models at its October 7 Windows and Surface event. The session focused on AI agents and showed the local models in action.
Why this mattersAI · ~100 words

Tap a lens to see what this story means for you.

Morning Brief

Liked this? The Brief brings you the whole day in tech, verified, every morning.

Two minutes, free forever. What's in The Brief →

Reader-supported
DonateBuy me a coffee →Follow@thecircuitry_ →Follow@thecircuitry.to →
HELP US IMPROVE
From The Circuitry

See what’s happening right now

The Feed runs all day — short, verified briefs the moment they break.

Open the Feed →
From The Circuitry

Follow @thecircuitry_

Every story we publish, as it happens. No noise between.

Follow on X ↗On Bluesky ↗

Reader-supported

The Circuitry is a passion project I've always wanted to build, and I love the work behind it.

Running it costs real money. APIs, hosting, time. To keep improving the site and growing this into something useful for everyone, those costs have to be covered.

Any contribution is appreciated. If not, no pressure. Thanks for reading.

Buy me a coffee
microsoftwindowsai
More fromFrandroid
  • Microsoft preparing local MAI-Code-1.1-Flash for high-end PCs

    Tech · 14h
  • Musk Says SpaceXAI Will Become SpaceXSI After Trump's Super Intelligence Push

    Tech · 3d
  • Meta GDPR data, Xiaomi 18 Pro tests, PS5 browser jailbreak top tech week

    Tech · 4d
More inTech
  • MacRumors: Gurman Says OLED MacBook Pro Won't Be Significantly Thinner

    Tech · 1h
  • PS5 Crunchyroll Anime Hub Launches Spring 2027

    Tech · 1h
  • Anthropic launches Claude Haiku 5.5, cutting prices up to 90% from Haiku 4.5

    Tech · 11h
SupportThe Work

The Circuitry is reader-supported. If you find the daily brief useful, you can buy me a coffee to keep it going.

Buy a coffee →
From The CircuitryWhy The Circuitry

Verified tech news, cross-checked.

Every story is checked against independent sources before it posts — no rumors dressed up as fact.

How we verify →

MORE IN THIS BEAT

All Tech →
  • Tech· 

    Microsoft opens Surface Laptop Ultra preorders and pushes AI agents deeper into Windows

    At its Windows and Surface event, Microsoft opened preorders for the $2,599 Surface Laptop Ultra and the $5,999 Surface RTX Spark Dev Box, and said Copilot will be able to use files and act across Windows in the coming months.

  • Tech· 

    Microsoft is giving Copilot access to your Windows files and the power to act on them

    Microsoft says Copilot on Copilot+ PCs will be able to use files and recent activity on your PC, take actions across Windows, and tap local AI models, with your permission. The features are expected to start rolling out in the coming months.

  • Tech· 

    Microsoft preparing local MAI-Code-1.1-Flash for high-end PCs

    Microsoft has introduced a local edition of MAI-Code-1.1-Flash that runs on Windows PCs without server access. The model requires substantial memory and is coming to GitHub Copilot in experimental preview by the end of October, starting with Nvidia RTX Spark PCs.

  • Tech· 

    First Nvidia RTX Spark laptops range from $2,599 to nearly $7,000

    The first laptops built on Nvidia's RTX Spark chip, from Microsoft, Asus, Dell, HP, Lenovo and MSI, start at $2,599 and climb to $6,999.99 for a maxed-out Asus ProArt P16 with 128GB of unified memory, The Verge reports. The first models ship as early as October 16.

  • Tech· 

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    OpenAI released 722 mathematical manuscripts, grouped into 372 result families, produced by an unreleased internal model. The papers are on GitHub, with Lean formalizations for many but not all of them, and OpenAI warns some unformalized results could have issues.