metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
Microsoft
Nvidia
Local AI
AI Hardware
Windows 11

Surface RTX Spark: Running 120B AI Models Locally

Microsoft's $5,999 Surface RTX Spark dev box has 128GB unified memory for local models. We break down the math, rivals, and cloud vs local tradeoffs.

Metir AI TeamOctober 8, 20268 min read
Surface RTX Spark: Running 120B AI Models Locally

On October 7, 2026, Microsoft unveiled a new line of Nvidia RTX Spark AI PCs with a revamped Windows 11, according to TechCrunch. The most AI-specific product is the Surface RTX Spark Dev Box, a $5,999 desktop with 128GB of unified memory that Microsoft says can run models above 120 billion parameters locally. This article checks what is confirmed, explains the arithmetic of running large models on a box like this, and compares it with other local AI machines.

Microsoft logoMicrosoft
NVIDIA logoNVIDIA
AMD logoAMD
Apple logoApple
The vendors in the local AI box market covered in this article.
$5,999Surface RTX Spark Dev BoxPreorders opened Oct 7, ships November
128 GBUnified memoryShared by CPU and GPU
120B+ParametersMicrosoft's local-model claim
~1 PFLOPPeak AI computeFP4 with sparsity

What Microsoft announced

Per TechCrunch, the launch covers a Surface Laptop Ultra (two base models starting at $2,600 and $3,700, rising to $5,900 with more memory and storage) and the Surface RTX Spark Dev Box. Dell also announced an XPS 16 Creator Edition on the same chip at $3,800. Microsoft says the highest-end laptop is already sold out.

Details of the dev box come from Unite.AI:

  • Price and timing: $5,999 (listed as "from $5,999.99"), preorders on Microsoft.com in the US, shipping in November. TechCrunch rounds the price to $6,000.
  • Chip: the Nvidia RTX Spark superchip, a Blackwell GPU with 6,144 CUDA cores and a 20-core Grace Arm CPU, linked by NVLink-C2C, at a 100W thermal envelope.
  • Memory: 128GB unified between CPU and GPU. How much the GPU can address depends on configuration and workload.
  • Compute: up to 1 petaflop of AI performance, a theoretical figure at FP4 precision with sparsity.
  • Software: Windows 11 Pro, WSL 2 with GPU passthrough and CUDA, Visual Studio Code, GitHub Copilot and a local model called Aion 1.0 Instruct.

The Windows 11 update adds Execution Containers, which Microsoft says make it easier to sandbox AI agents and will reach all Windows 11 users. CEO Satya Nadella said agents need orchestration and "memory outside of the model," per TechCrunch.

A tree-lined paved plaza between office buildings on the Microsoft Redmond campus
A plaza on the Microsoft Redmond campus in Washington. This is a file photo of the campus, not of the Surface RTX Spark or the October 7 event. Photo by Jonathan Schilling, CC BY-SA 4.0, via Wikimedia Commons.

One figure is missing from Microsoft's published material: memory bandwidth. Neither Microsoft nor Nvidia has published it, and independent testing has not yet measured it. The closest published reference is Nvidia's DGX Spark, which pairs the same 128GB capacity with 273 GB/s, so the examples below use roughly 300 GB/s as an illustrative round number rather than a specification.

The memory math: what fits in 128GB

A model's weights occupy roughly parameters times bytes per parameter. Precision decides the bytes: 16-bit weights use 2 bytes, 8-bit uses 1, and 4-bit uses about 0.5 (slightly more once quantisation scales are stored).

Model size16-bit8-bit4-bit
70B140 GB70 GB35 GB
120B240 GB120 GB60 GB
200B400 GB200 GB100 GB

Only the 4-bit column of the 120B row leaves real headroom in a 128GB machine. At 8-bit, 120B weights alone are 120 GB, leaving almost nothing for the operating system or working memory. That is why the 120B-plus claim implicitly means 4-bit quantisation, the format the chip's FP4 Tensor Cores target. Nvidia makes the same kind of claim for its own box: Nvidia's DGX Spark page lists up to 200B parameters on a single 128GB unit.

Weights are not the whole bill. The key-value (KV) cache stores attention state for every token of context. Its size per token is 2 x layers x KV heads x head dimension x bytes per value. For a hypothetical 80-layer model with 8 KV heads of dimension 128 at 16-bit, that is 327,680 bytes per token, so a 128,000-token context adds about 43 GB. Weights at 4-bit (35 GB for a 70B model) plus that cache already reach 78 GB. Long contexts can consume more memory than the weights themselves, which matters given Microsoft's June reference to a 1 million token context window.

Capacity gets a model in, bandwidth sets its speed

Generating each token requires streaming the active weights from memory to the compute units. For a single user, the speed limit is roughly memory bandwidth divided by bytes read per token. A 35 GB model on a hypothetical 300 GB/s machine tops out near 8.6 tokens per second, before any software overhead. The same model on a 1.2 TB/s machine has a ceiling near 34.

Memory bandwidth sets the speed limit for local models

Peak memory bandwidth (GB/s) and the theoretical ceiling in tokens per second for a 35 GB model, where each generated token reads all the weights once. Real speeds are lower. The Surface RTX Spark Dev Box is omitted because its bandwidth has not been published.

Mac Studio (M5 Ultra)Up to 512 GB memory
1,200 GB/s · ceiling about 34.3 tokens/s
Mac Studio (M5 Max)Up to 128 GB memory
614 GB/s · ceiling about 17.5 tokens/s
Nvidia DGX Spark128 GB
273 GB/s · ceiling about 7.8 tokens/s
AMD Ryzen AI Max+ 395128 GB; up to 96 GB to GPU
256 GB/s · ceiling about 7.3 tokens/s

Ceiling = bandwidth divided by 35 GB. Mixture-of-experts models read only their active parameters per token, so they can run several times faster.

Mixture-of-experts (MoE) models change the picture. They hold many parameters but activate only a few per token, so they need capacity for all weights and bandwidth only for the active slice. A hypothetical 120B MoE with 5B active parameters at 4-bit reads about 2.5 GB per token, a ceiling of roughly 120 tokens per second at 300 GB/s. This is why unified-memory boxes with modest bandwidth suit large MoE models better than large dense ones. Prompt processing, by contrast, is compute-bound, where the FP4 petaflop figure is relevant.

“

Capacity decides whether a model fits. Bandwidth decides whether it feels usable.

Metir AI analysis

How it compares with other local AI boxes

MachineMemoryBandwidthReported starting price
Surface RTX Spark Dev Box128 GBnot published$5,999
Nvidia DGX Spark128 GB273 GB/s$4,699 (reported)
AMD Ryzen AI Max+ 395 systemsup to 128 GB (96 GB to GPU)256 GB/svaries by maker
Mac Studio, M5 Maxup to 128 GB614 GB/s (reported)$2,499 (36 GB base)
Mac Studio, M5 Ultraup to 512 GB1.2 TB/s$5,499 (96 GB base)

Sources: Nvidia's DGX Spark page gives 128GB and 273 GB/s but no price; the $4,699 figure is a reported February 2026 increase and resellers differ. AMD's product page and ApX cite 256 GB/s and a 96 GB GPU ceiling. RuntimeWire reports the M5 Ultra at 1.2 TB/s with up to 512GB, from $5,499 with 96GB, and says Apple has not priced the 512GB configurations. The M5 Max bandwidth and 128GB ceiling come from secondary coverage. Base prices are for different memory sizes, so they are not like-for-like.

Read plainly, the RTX Spark class and DGX Spark share the same capacity-rich, bandwidth-modest design. Apple's higher-bandwidth chips run dense models faster per dollar of memory but sit on a different software stack. The Surface box's distinguishing feature is Windows with CUDA and WSL, which matters if a team's tools are built for Nvidia and Microsoft. We covered Apple's silicon direction in our look at the M6 and M5 Ultra.

Who this is for

  • Developers building and testing agents who want a CUDA-compatible machine on the desk, with sandboxing through Execution Containers.
  • Teams with data-residency or confidentiality constraints where prompts must not leave the device.
  • Heavy, steady users whose token volume would otherwise be billed per call.

It is a poor fit for people who need frontier-model quality. The largest hosted models are far bigger than 120B, and a quantised open model on a desk is a different capability tier.

Cloud vs local: the tradeoffs

FactorLocal boxCloud API
Cost shapeOne-time hardware cost, then powerPay per token, scales to zero
PrivacyData stays on deviceDepends on provider terms
Model qualityOpen-weight models that fit memoryFrontier models
SpeedBandwidth-limited, one userProvider-scaled
UpkeepYou manage models and updatesProvider manages

Break-even depends on usage. At $5,999, a box needs sustained daily load to beat metered pricing, and hosted prices have fallen quickly; see our analysis of the 13x-per-year cost decline. Memory is also costly: Nvidia cited memory supply for its DGX Spark price rise, a theme we explored in the AI memory supercycle. Many teams will end up hybrid, running sensitive or high-volume work locally and routing hard problems to hosted models. Multi-model tools such as Metir make that routing a menu choice rather than a rebuild.

What to watch

  • Independent benchmarks confirming bandwidth and real tokens per second on 70B to 120B models.
  • How much of the 128GB the GPU can actually address.
  • Whether the November shipments arrive on schedule and whether other OEMs price RTX Spark desktops below $5,999.
  • How widely Execution Containers are adopted for agent sandboxing.

Sources:

  • TechCrunch: Microsoft releases new Nvidia chip AI PCs with revamped Windows 11
  • Unite.AI: Surface RTX Spark Dev Box ships in November at $5,999
  • Tom's Hardware: Surface Laptop Ultra with RTX Spark
  • Nvidia: DGX Spark
  • AMD Ryzen AI Max+ 395 (ApX summary)
  • RuntimeWire: Mac Studio M5 Ultra

Image credits

  • Microsoft Redmond campus, hero (intersection) and plaza photos: Jonathan Schilling, CC BY-SA 4.0, via Wikimedia Commons (Intersection on the Microsoft Redmond campus, Plaza on the Microsoft Redmond campus). File photos of the campus, not of the product or event.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational