metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
xAI
NVIDIA
Data Centers
AI Infrastructure
Compute
Colossus

xAI's Colossus 2: Inside the Race to 1.44 Million GPUs

Elon Musk disclosed Colossus 2's exact chip counts and a three-wave GB300 ramp toward 1.44 million combined GPUs by year-end. What the numbers mean, and why power is the real constraint.

Metir AI TeamSeptember 26, 20269 min read
xAI's Colossus 2: Inside the Race to 1.44 Million GPUs

On September 25, 2026, Elon Musk posted the most granular public accounting yet of xAI's compute buildout in Memphis, Tennessee. Replying to a question on X, he gave exact chip counts for both Colossus data centers and laid out a three-wave ramp that, if it lands in full, would put roughly 1.44 million GPUs online at the site by the end of the year. The disclosure is a useful data point in its own right, and it is also a window into a broader pattern playing out across the frontier AI industry: raw NVIDIA chip counts are becoming a marketing metric, while the actual ceiling on how fast any of these clusters can grow is set by something far less glamorous than silicon: electricity.

xAI logoxAI
NVIDIA logoNVIDIA
X logoX
xAI's Colossus 2 buildout, disclosed by Elon Musk on X, runs on NVIDIA's Blackwell-generation GPUs.

What Musk actually disclosed

Musk's post, replying to a user asking about xAI's current cluster size, broke the numbers down by site and by chip generation. Colossus 1, xAI's original Memphis cluster, is running 150,000 NVIDIA H100 GPUs, 50,000 H200s, and 30,000 GB200s, a total of 230,000 chips spanning two hardware generations. Colossus 2, the newer and far larger buildout, already runs 110,000 GB200s and 440,000 GB300s, a combined 550,000 chips, all Blackwell-generation. Added together, the Memphis site currently runs on the order of 780,000 GPUs.

Musk then laid out what comes next: another 220,000 GB300s would be fully operational "next week," a further 220,000 in November, and, in his words, "if we get lucky," a third wave of 220,000 more by late December. If all three waves land, Colossus 2 alone would reach roughly 1.21 million GB200 and GB300 chips, and the combined Memphis site, Colossus 1 plus Colossus 2, would total approximately 1.44 million GPUs by year-end.

230,000Colossus 1 chips150k H100 + 50k H200 + 30k GB200
550,000Colossus 2 chips (current)110k GB200 + 440k GB300
660,000Additional GB300s plannedThree waves of 220k each
~1.44MCombined GPU totalIf all three waves land by year-end

That target lines up with something xAI has said before. The company has previously stated it intends to equip its Memphis facility with 1 million GPUs by the end of 2026, a milestone that, on Musk's own numbers, the site had already effectively cleared with the first GB300 wave alone. The new figure simply pushes the ambition further, from one million chips to closer to one and a half million, contingent on the final, explicitly conditional wave actually shipping.

“

If we get lucky, yet another 220k GB300 by late December.

Elon Musk, on X, September 25, 2026

Three GB300 waves, ~1.44 million combined GPUs by year-end

Cumulative GPU count across xAI's Memphis site (Colossus 1 + Colossus 2), per Elon Musk's September 25, 2026 disclosure. The final wave is explicitly conditional.

Chip counts, not delivered floating-point performance; GB300 racks draw substantially more power per chip than the H100s Colossus 1 started with.

GB200 versus GB300, and why "GPU count" is an imperfect proxy

Both GB200 and GB300 belong to NVIDIA's Blackwell generation, and both ship as part of the NVL72 rack architecture, 72 GPUs paired with Grace CPUs, wired together with NVLink at roughly 130 terabytes per second of interconnect bandwidth. But they are not the same chip. GB300 is the "Blackwell Ultra" refresh: it swaps in a higher-memory GPU with 288 GB of HBM3e per chip against 186 GB on the GB200, which works out to roughly 20 terabytes of pooled memory per rack instead of 13.4 terabytes. It also raises dense low-precision throughput by roughly 50% per rack. None of that comes for free. NVIDIA's reference specifications put a GB200 NVL72 rack's power draw at around 120 kilowatts, and a GB300 NVL72 rack at up to 142 kilowatts, worked out per GPU as a rise from roughly 1,200 watts to roughly 1,400 watts, though NVIDIA itself cautions against multiplying that per-chip figure out to a facility total, since real racks also carry CPUs, networking switches, and conversion losses that don't scale linearly.

Colossus 2 skipped straight to Blackwell

Chip generation mix by site as of September 25, 2026. Colossus 1 carries Hopper-class GPUs (H100/H200) plus a small Blackwell slice; Colossus 2 is Blackwell and Blackwell Ultra only, and already has more chips than Colossus 1 on its own.

GB200 and GB300 are both Blackwell-generation chips; GB300 (Blackwell Ultra) draws more power per GPU and carries more HBM memory.

That difference is exactly why counting GPUs is a blunt way to measure a cluster. A million H100s and a million GB300s are not the same amount of compute, and they are especially not the same amount of power demand: swapping a Hopper-class fleet for a Blackwell Ultra one of equal chip count meaningfully increases the electricity a site needs to draw, even before accounting for the extra compute each chip actually delivers. Colossus 2's makeup underlines the point. Unlike Colossus 1, which still carries a large legacy H100 and H200 footprint from earlier in the buildout, Colossus 2 launched directly on Blackwell and Blackwell Ultra chips, with no older-generation hardware at all. It is, chip for chip, a denser and hungrier cluster than its predecessor from day one.

The real ceiling is gigawatts, not chips

This is where the story becomes less about NVIDIA's shipping schedule and more about the grid. Reporting on the buildout has centered on a roughly 1.2-gigawatt power plant xAI is building to bring the Colossus 2 systems fully online, a permanent facility with 41 gas turbines that is meant to replace a larger set of temporary, unpermitted turbines the company had been running at a site across the state line in Southaven, Mississippi. Industry analysis has separately described Colossus 2 as on track to be the first gigawatt-scale AI data center in the world, tracking a rise from roughly 200 megawatts of capacity in August 2025 toward the multi-gigawatt range as the buildout continues.

The order of magnitude is worth sitting with. Using NVIDIA's own published rack figures, a back-of-envelope estimate of Colossus 2's chip-level power draw at full buildout (roughly 1.21 million GB200 and GB300 GPUs at roughly 1,200 to 1,400 watts each) comes out well above 1.5 gigawatts before accounting for cooling, networking, and other facility overhead, which typically adds a substantial premium on top of raw chip draw. That is a rough illustrative calculation, not an official disclosure, but it helps explain why a single 1.2-gigawatt plant is described as only part of xAI's power stack rather than the whole of it: the company has also relied on temporary gas turbines and Tesla Megapacks for storage and load-smoothing, alongside grid interconnection, to keep the site powered while the permanent plant comes online. Power, not chip supply, is the resource being rationed here. NVIDIA can ship the GPUs; getting enough dependable, permitted electricity to a single site to run over a million of them simultaneously is the harder problem, and it's the one with the longest lead times, since permanent gas turbine and grid-interconnection projects run on a timeline measured in quarters and years, not the weeks Musk's chip waves describe.

How Colossus 2 compares, and what's still unresolved

Downtown Memphis, Tennessee skyline along the Mississippi River
Downtown Memphis, the city xAI has built its Colossus data centers around. Colossus 1 and Colossus 2 themselves sit outside downtown, in South Memphis and across the state line in Southaven, Mississippi, not in this skyline.

Set against the rest of the industry, Colossus 2 is one of a small number of clusters being built and discussed at genuinely gigawatt scale, alongside comparable ambitions from the largest US hyperscalers and other frontier labs racing to secure both chip supply and power capacity simultaneously. What distinguishes xAI's approach so far is less the scale itself than the speed and the method: building dedicated, often off-grid or behind-the-meter power generation directly alongside the compute, rather than waiting on utility-scale grid upgrades that can take years to permit and build.

That speed has drawn scrutiny. The turbines powering the site have been the subject of a NAACP lawsuit alleging operation without required Clean Air Act permits and disproportionate air-quality impacts on nearby communities in Southaven and South Memphis, and local reporting has tracked ongoing tension between xAI's buildout timeline and the pace of environmental permitting. Those are live, unresolved questions, not settled ones, and they sit alongside a second, more mundane uncertainty: Musk's own framing of the final 220,000-chip wave as something that happens only "if we get lucky" is itself a tell. Compute buildouts at this scale routinely slip against announced timelines, whether from chip delivery, construction, or power availability, and a disclosure this specific and this optimistic is worth revisiting in December against what actually ships.

For teams building AI products, the practical lesson in all of this compute-arms-race noise is less about any one company's chip count and more about optionality. As the frontier keeps shifting between whichever lab currently has the most power online, betting a product on a single model or a single vendor's availability is a bet on infrastructure timelines outside your control. Metir AI is built to route across models and providers for that reason, so a workload can move to whichever model is actually available and cost-effective rather than being tied to one lab's capacity constraints.

Sources:

  • Elon Musk on X, September 25, 2026
  • Elon Musk's SpaceXAI to add another 660,000 AI GPUs this year, nearing a total of 1.44 million in operation | Tom's Hardware
  • Elon Musk Says SpaceXAI's Colossus 2 Could Run More Than 1.2 Million Nvidia GB200 And GB300 Chips By Year-End, "If We Get Lucky" | Benzinga
  • Elon Musk says xAI's Colossus 2 could more than double Nvidia chip count by year-end | Invezz
  • Colossus (supercomputer) | Wikipedia
  • xAI's Colossus 2 - First Gigawatt Datacenter In The World | SemiAnalysis
  • SpaceX won't remove all of xAI's unpermitted turbines for another year | TechCrunch
  • xAI announces fourth data center; turbines to be removed from Southaven site | Action News 5
  • GB200 & GB300 NVL72 Power and Cooling Requirements | Pantheon
  • GB300 NVL72: Blackwell Ultra Deployment | Introl
  • xAI targets one million GPUs for Colossus supercomputer in Memphis | Data Center Dynamics

Image credits

Hero: "Elon Musk Colorado 2022 (cropped2)," U.S. Air Force photo by Trevor Cokley, public domain (U.S. federal government work). Source: Wikimedia Commons. Depicts Elon Musk, who made the Colossus 2 chip-count disclosure this post covers; not a photo of Colossus itself. Reviewed before publication.

In-body figure: "Memphis, TN skyline along the Mississippi River" by Quintin Soloviev, licensed under CC BY 4.0. Source: Wikimedia Commons. Shows downtown Memphis; Colossus 1 and Colossus 2 are sited outside downtown, in South Memphis and across the state line in Southaven, Mississippi, not in this skyline. Reviewed before publication.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational