On October 7, 2026, Microsoft unveiled a new line of Nvidia RTX Spark AI PCs with a revamped Windows 11, according to TechCrunch. The most AI-specific product is the Surface RTX Spark Dev Box, a $5,999 desktop with 128GB of unified memory that Microsoft says can run models above 120 billion parameters locally. This article checks what is confirmed, explains the arithmetic of running large models on a box like this, and compares it with other local AI machines.
NVIDIAWhat Microsoft announced
Per TechCrunch, the launch covers a Surface Laptop Ultra (two base models starting at $2,600 and $3,700, rising to $5,900 with more memory and storage) and the Surface RTX Spark Dev Box. Dell also announced an XPS 16 Creator Edition on the same chip at $3,800. Microsoft says the highest-end laptop is already sold out.
Details of the dev box come from Unite.AI:
- Price and timing: $5,999 (listed as "from $5,999.99"), preorders on Microsoft.com in the US, shipping in November. TechCrunch rounds the price to $6,000.
- Chip: the Nvidia RTX Spark superchip, a Blackwell GPU with 6,144 CUDA cores and a 20-core Grace Arm CPU, linked by NVLink-C2C, at a 100W thermal envelope.
- Memory: 128GB unified between CPU and GPU. How much the GPU can address depends on configuration and workload.
- Compute: up to 1 petaflop of AI performance, a theoretical figure at FP4 precision with sparsity.
- Software: Windows 11 Pro, WSL 2 with GPU passthrough and CUDA, Visual Studio Code, GitHub Copilot and a local model called Aion 1.0 Instruct.
The Windows 11 update adds Execution Containers, which Microsoft says make it easier to sandbox AI agents and will reach all Windows 11 users. CEO Satya Nadella said agents need orchestration and "memory outside of the model," per TechCrunch.

One figure is missing from Microsoft's published material: memory bandwidth. Neither Microsoft nor Nvidia has published it, and independent testing has not yet measured it. The closest published reference is Nvidia's DGX Spark, which pairs the same 128GB capacity with 273 GB/s, so the examples below use roughly 300 GB/s as an illustrative round number rather than a specification.
The memory math: what fits in 128GB
A model's weights occupy roughly parameters times bytes per parameter. Precision decides the bytes: 16-bit weights use 2 bytes, 8-bit uses 1, and 4-bit uses about 0.5 (slightly more once quantisation scales are stored).
| Model size | 16-bit | 8-bit | 4-bit |
|---|---|---|---|
| 70B | 140 GB | 70 GB | 35 GB |
| 120B | 240 GB | 120 GB | 60 GB |
| 200B | 400 GB | 200 GB | 100 GB |
Only the 4-bit column of the 120B row leaves real headroom in a 128GB machine. At 8-bit, 120B weights alone are 120 GB, leaving almost nothing for the operating system or working memory. That is why the 120B-plus claim implicitly means 4-bit quantisation, the format the chip's FP4 Tensor Cores target. Nvidia makes the same kind of claim for its own box: Nvidia's DGX Spark page lists up to 200B parameters on a single 128GB unit.
Weights are not the whole bill. The key-value (KV) cache stores attention state for every token of context. Its size per token is 2 x layers x KV heads x head dimension x bytes per value. For a hypothetical 80-layer model with 8 KV heads of dimension 128 at 16-bit, that is 327,680 bytes per token, so a 128,000-token context adds about 43 GB. Weights at 4-bit (35 GB for a 70B model) plus that cache already reach 78 GB. Long contexts can consume more memory than the weights themselves, which matters given Microsoft's June reference to a 1 million token context window.
Capacity gets a model in, bandwidth sets its speed
Generating each token requires streaming the active weights from memory to the compute units. For a single user, the speed limit is roughly memory bandwidth divided by bytes read per token. A 35 GB model on a hypothetical 300 GB/s machine tops out near 8.6 tokens per second, before any software overhead. The same model on a 1.2 TB/s machine has a ceiling near 34.
Memory bandwidth sets the speed limit for local models
Peak memory bandwidth (GB/s) and the theoretical ceiling in tokens per second for a 35 GB model, where each generated token reads all the weights once. Real speeds are lower. The Surface RTX Spark Dev Box is omitted because its bandwidth has not been published.
Ceiling = bandwidth divided by 35 GB. Mixture-of-experts models read only their active parameters per token, so they can run several times faster.
Mixture-of-experts (MoE) models change the picture. They hold many parameters but activate only a few per token, so they need capacity for all weights and bandwidth only for the active slice. A hypothetical 120B MoE with 5B active parameters at 4-bit reads about 2.5 GB per token, a ceiling of roughly 120 tokens per second at 300 GB/s. This is why unified-memory boxes with modest bandwidth suit large MoE models better than large dense ones. Prompt processing, by contrast, is compute-bound, where the FP4 petaflop figure is relevant.
Capacity decides whether a model fits. Bandwidth decides whether it feels usable.
Metir AI analysis
How it compares with other local AI boxes
| Machine | Memory | Bandwidth | Reported starting price |
|---|---|---|---|
| Surface RTX Spark Dev Box | 128 GB | not published | $5,999 |
| Nvidia DGX Spark | 128 GB | 273 GB/s | $4,699 (reported) |
| AMD Ryzen AI Max+ 395 systems | up to 128 GB (96 GB to GPU) | 256 GB/s | varies by maker |
| Mac Studio, M5 Max | up to 128 GB | 614 GB/s (reported) | $2,499 (36 GB base) |
| Mac Studio, M5 Ultra | up to 512 GB | 1.2 TB/s | $5,499 (96 GB base) |
Sources: Nvidia's DGX Spark page gives 128GB and 273 GB/s but no price; the $4,699 figure is a reported February 2026 increase and resellers differ. AMD's product page and ApX cite 256 GB/s and a 96 GB GPU ceiling. RuntimeWire reports the M5 Ultra at 1.2 TB/s with up to 512GB, from $5,499 with 96GB, and says Apple has not priced the 512GB configurations. The M5 Max bandwidth and 128GB ceiling come from secondary coverage. Base prices are for different memory sizes, so they are not like-for-like.
Read plainly, the RTX Spark class and DGX Spark share the same capacity-rich, bandwidth-modest design. Apple's higher-bandwidth chips run dense models faster per dollar of memory but sit on a different software stack. The Surface box's distinguishing feature is Windows with CUDA and WSL, which matters if a team's tools are built for Nvidia and Microsoft. We covered Apple's silicon direction in our look at the M6 and M5 Ultra.
Who this is for
- Developers building and testing agents who want a CUDA-compatible machine on the desk, with sandboxing through Execution Containers.
- Teams with data-residency or confidentiality constraints where prompts must not leave the device.
- Heavy, steady users whose token volume would otherwise be billed per call.
It is a poor fit for people who need frontier-model quality. The largest hosted models are far bigger than 120B, and a quantised open model on a desk is a different capability tier.
Cloud vs local: the tradeoffs
| Factor | Local box | Cloud API |
|---|---|---|
| Cost shape | One-time hardware cost, then power | Pay per token, scales to zero |
| Privacy | Data stays on device | Depends on provider terms |
| Model quality | Open-weight models that fit memory | Frontier models |
| Speed | Bandwidth-limited, one user | Provider-scaled |
| Upkeep | You manage models and updates | Provider manages |
Break-even depends on usage. At $5,999, a box needs sustained daily load to beat metered pricing, and hosted prices have fallen quickly; see our analysis of the 13x-per-year cost decline. Memory is also costly: Nvidia cited memory supply for its DGX Spark price rise, a theme we explored in the AI memory supercycle. Many teams will end up hybrid, running sensitive or high-volume work locally and routing hard problems to hosted models. Multi-model tools such as Metir make that routing a menu choice rather than a rebuild.
What to watch
- Independent benchmarks confirming bandwidth and real tokens per second on 70B to 120B models.
- How much of the 128GB the GPU can actually address.
- Whether the November shipments arrive on schedule and whether other OEMs price RTX Spark desktops below $5,999.
- How widely Execution Containers are adopted for agent sandboxing.
Sources:
- TechCrunch: Microsoft releases new Nvidia chip AI PCs with revamped Windows 11
- Unite.AI: Surface RTX Spark Dev Box ships in November at $5,999
- Tom's Hardware: Surface Laptop Ultra with RTX Spark
- Nvidia: DGX Spark
- AMD Ryzen AI Max+ 395 (ApX summary)
- RuntimeWire: Mac Studio M5 Ultra
Image credits
- Microsoft Redmond campus, hero (intersection) and plaza photos: Jonathan Schilling, CC BY-SA 4.0, via Wikimedia Commons (Intersection on the Microsoft Redmond campus, Plaza on the Microsoft Redmond campus). File photos of the campus, not of the product or event.
