The Surface Laptop Ultra starts at $2,599, ships October 16, and is the first machine built around Nvidia’s RTX Spark, a chip that puts a 20-core Arm CPU and a Blackwell GPU with up to 6,144 CUDA cores on one die, sharing up to 128GB of LPDDR5x memory. Microsoft showed it off at a Windows and Surface event in San Francisco on October 7, its first live hardware event in two years, with Nvidia’s Jensen Huang on stage alongside CEO Satya Nadella and Windows chief Pavan Davuluri, as Ars Technica reported.
The number that matters is the memory architecture, not the headline price. A conventional laptop splits its silicon budget between system RAM and a GPU’s own VRAM, and whichever pool runs out first caps what you can do. RTX Spark’s 128GB is unified: CPU and GPU draw from the same pool over Nvidia’s NVLink-C2C interconnect, so a large language model and a game can both claim memory without a hard partition. Windows Central reported the chip’s graphics performance as comparable to a discrete GeForce RTX 5070, while drawing up to 80W total. Nvidia’s own spec sheet puts the desktop RTX 5070’s board power at 250W. If the comparison holds under independent testing, that is roughly a third of the power for similar graphics output — a claim from Nvidia and Microsoft, not yet a benchmark anyone outside either company has published.
RTX Spark itself is five months old. Huang unveiled it May 31 at Computex in Taipei, describing a chip built on a 3-nanometer process with 70 billion transistors, a figure Nvidia’s own announcement and independent trade coverage both carried. MediaTek co-designed the CPU. Dell, Asus, HP, Lenovo, MSI and Microsoft all have RTX Spark laptops shipping this fall; a Morgan Stanley note leaked ahead of launch had pegged pricing between $1,799 and $2,899. Microsoft’s $2,599 entry price sits inside that range, but it is the base configuration — the 128GB, 6,144-core version that can run tens of billions of parameters locally costs more.
Microsoft’s other new box, the Surface RTX Spark Dev Box, starts at $5,999 and ships in November, according to Engadget’s rundown of the event. It is a fanless-looking aluminum unit with 128GB of unified memory, a claimed one petaflop of FP4 AI compute, a 100W thermal envelope, and — per Microsoft’s own press materials — exactly 1,000 chassis air vents, a nod to the 1,000-teraflop figure.
That spec sheet is not new. It matches Nvidia’s own DGX Spark, a Linux-based mini PC built around the same GB10 Grace Blackwell silicon, which launched October 15, 2025 at $3,999 for the Founders Edition. Nvidia raised that price to $4,699 on February 23, 2026 — an 18 percent increase the company attributed, in a post on its developer forums, to “worldwide constraints in memory supply.” Microsoft’s Windows-flavored version of essentially the same one-petaflop, 128GB hardware now costs $1,300 more than Nvidia’s own current listing, a 28 percent premium for a different operating system on top of the same chip family.
The guardrails
The hardware is only half the announcement. Microsoft also pushed Microsoft Execution Containers, or MXC, to general availability on Windows 11 on October 7 — four months after it entered preview at Microsoft’s Build conference on June 2. MXC is not a product but a policy layer: developers and IT administrators write rules for what an agent can touch — which files, which networks, which identity — and Windows enforces those rules at the kernel level, so the agent or any code it generates cannot grant itself more access than it was given, according to Microsoft’s Windows Developer Blog. The containment itself scales along what Microsoft calls a composable sandbox spectrum, from lightweight process isolation up to a disposable, full Windows Sandbox VM or a cloud instance on Windows 365.
| From | To | How |
|---|---|---|
| Agent task | MXC policy layer (declares what the agent can access) | |
| MXC policy layer (declares what the agent can access) | Process container (lightest, used by GitHub Copilot CLI) | |
| MXC policy layer (declares what the agent can access) | Session container | |
| MXC policy layer (declares what the agent can access) | WSL container | |
| MXC policy layer (declares what the agent can access) | MicroVM (experimental) | |
| MXC policy layer (declares what the agent can access) | Windows Sandbox VM | |
| MXC policy layer (declares what the agent can access) | Windows 365 cloud instance (heaviest) |
Based on Microsoft Windows Developer Blog
At general availability, MXC already supports OpenAI’s Codex, GitHub Copilot, the open-source OpenClaw framework, Replit, LM Studio, Unsloth AI, and Nvidia’s OpenShell, according to VentureBeat’s reporting. OpenShell is not a new name: it is the secure agent runtime at the center of Nvidia’s own Open Agent Safety Platform, which Nvidia launched on September 28 alongside a record $150 billion increase to its stock buyback authorization. Anthropic’s Claude Code, Box, Perplexity, Raycast and several other agent tools are listed as “coming soon” rather than supported at launch.
The gap worth watching sits in enterprise management, not the agents. MXC enforces a policy boundary, but the tools to set and audit that boundary across a fleet of company PCs — Intune policy controls, Entra identity separation, and the broader Agent 365 management layer — are also still described by Microsoft as coming later. Until they ship, it is the agent’s developer, not an IT administrator, drawing the line around what it can touch on a given machine.
Hybrid Intelligence
Microsoft is calling the overall pitch Hybrid Intelligence: run an agent locally when the device can handle it, hand off to the cloud when it can’t, and let something decide which is which. The mechanism is GitHub’s HydraFusion router, previously used to route cloud-side model calls, extended to also weigh a device’s local models — arriving October 15, the same day the RTX Spark laptops ship, per Davuluri’s post on the Windows Experience Blog.
As AI models grow, it has become difficult to cover it all with cloud budgets. Customers want to save tokens without giving up frontier-level performance.
That quote, from Davuluri’s blog post, is a cost argument as much as a technical one. Microsoft’s own cloud margins face the same token math Nadella is trying to sell to investors: every inference an RTX Spark laptop runs on its own silicon is one Azure doesn’t have to serve. Pushing compute onto customers’ own hardware lowers Microsoft’s infrastructure bill at the same time it is pitched to users as lower latency and no recurring cloud cost.
Davuluri also said Copilot+ PCs are shipping in “tens of millions” of units a year and running “over 2 trillion inferences per month” locally — both self-reported figures, with no methodology published and no third party cited. The claim arrives three weeks after Microsoft’s own marketing was found to have quietly dropped the “Copilot+ PC” name from its public materials; the brand still appears to be useful inside a keynote even as it recedes from the product pages.
A specific capability claim also doesn’t match across sources. Ars Technica and Windows Central describe the laptop’s top configuration as able to run large language models with “tens of billions” of parameters entirely offline. Davuluri’s own blog post claims the hardware runs local models and agents with “over 120 billion parameters” on-device, delivering “99% of its maximum performance” even unplugged from power. That’s close to an order of magnitude gap between independent hands-on description and the company’s own superlative, and neither side has published a benchmark to settle it.

Ars Technica says it will publish a review before the laptop ships October 16, which is the first chance anyone outside Microsoft and Nvidia gets to check whether 80W of RTX Spark actually plays Gears of War: E-Day like a 250W desktop card, and whether MXC’s “coming soon” admin tools show up before enterprise fleets have agents running loose on them without one.

