On August 11, 2026, IBM and Together AI announced a multi-year agreement worth $240 million to build and operate a large-scale NVIDIA-powered cluster on IBM Cloud, dedicated to running AI inference. The cluster will use NVIDIA HGX B300 systems connected with NVIDIA Spectrum-X Ethernet networking, is expected to become available in the first quarter of 2027, and will be located in the United States. IBM described it as the first dedicated, large-scale cluster built for inference on IBM Cloud. Together AI, an AI-optimized cloud provider that specializes in open-source model serving, will use the capacity to deliver inference to enterprise customers.
The headline number is modest by the standards of an industry that now measures infrastructure commitments in tens of billions. What makes the deal worth reading closely is not its size but its shape: what each of the three companies puts in, and what the arrangement says about where the economics of AI are moving.
What was actually announced
The agreement is a purchasing-and-hosting arrangement, not an equity investment or a model-training partnership. IBM commits to deploying a cluster of NVIDIA HGX B300 systems, the rack-scale building block of NVIDIA's Blackwell Ultra generation, inside its own cloud. NVIDIA's Spectrum-X Ethernet provides the networking fabric that stitches those systems into a single high-throughput cluster rather than a collection of isolated servers, which is the part that matters for serving large models at scale. Reporting on the deal puts the initial cluster at roughly 2,000 Blackwell B300 chips.
Together AI is the tenant and the reason the cluster exists. The company runs an inference platform built around open-source and open-weight models, and it says it currently serves on the order of 400 trillion tokens per month across that platform. This new IBM Cloud capacity gives it a dedicated pool of the latest NVIDIA silicon to grow into.
How the IBM and Together AI cluster is put together
Three parties, three roles. IBM provides the cloud and the capital, NVIDIA provides the silicon and networking, and Together AI supplies the demand that fills it.
The arrangement lets IBM re-enter frontline AI infrastructure through a specialist inference partner rather than by building or training its own frontier models.
The executives framed the deal around cost and demand rather than raw capability. Vipul Ved Prakash, Together AI's chief executive, said enterprises "want the performance of the best frontier models without the closed-model price tag, and that only works if the infrastructure underneath is fast and reliable at scale." Alan Peacock, general manager of IBM Cloud, said enterprises "are in a race to adopt agentic AI at scale to drive real business outcomes." NVIDIA's Dion Harris, a senior director for HPC and AI infrastructure, offered the widest framing, calling AI factories "essential enterprise infrastructure, like electricity and telecommunications." Those are marketing lines, but the direction they point in is consistent: the pitch is about serving models economically, not about who has the smartest model.
The quiet shift from training to inference
The most analytically interesting word in the announcement is "inference." For most of the current AI cycle, the infrastructure story has been dominated by training: the enormous, one-time cost of building a frontier model by pushing trillions of tokens through a cluster for weeks or months. Training is where the biggest single clusters and the most attention have gone.
Inference is the other half, and it is the half that scales with usage rather than with ambition. Every time a model answers a question, generates code, or runs a step of an agent's workflow, that is inference, and the cost recurs for the life of the product. As AI moves from demonstrations into everyday software, the cumulative cost of serving models starts to rival or exceed the cost of training them. A cluster built and marketed specifically for inference, rather than described in the more prestigious language of training frontier models, is a small signal that the industry's spending is following usage into production.

That distinction also explains the choice of hardware and networking. Serving many concurrent users efficiently is less about a single massive training run and more about throughput, latency, and packing as many requests as possible onto each expensive GPU. NVIDIA's own claim for the Blackwell generation, that an HGX B300-based system can deliver roughly thirty times more AI factory output than the prior generation for certain workloads, is a marketing figure rather than an independent benchmark, but it speaks to why an inference-first buyer would want the newest silicon: the unit economics of serving improve as each chip handles more work.
Why IBM, and why now
IBM is not a name that appears often in frontier-model headlines. It sold its consumer-facing ambitions long ago and has spent the current AI cycle positioned mostly as an enterprise software and consulting company, with its own smaller Granite model family aimed at business use rather than leaderboard supremacy. Its cloud, IBM Cloud, is a distant competitor to Amazon Web Services, Microsoft Azure, and Google Cloud in general-purpose hosting.
This deal is a bid to matter in AI infrastructure without trying to win the model race. Rather than building its own frontier models or a giant training supercluster, IBM is renting the latest NVIDIA hardware into its cloud and handing the demand-generation problem to a specialist. Together AI brings the customers, the software layer, and the token volume; IBM brings the capital, the data-center footprint, and an enterprise sales relationship that many businesses already have. Some coverage of the deal described it as IBM stepping into the "neocloud" business, the term now used for providers whose main product is renting out GPU capacity rather than a broad menu of cloud services.
IBM is buying its way back to relevance in AI infrastructure through a specialist partner, rather than by trying to build a frontier model of its own.
On IBM's positioning in the deal
The logic is defensible on both sides. For IBM, a $240 million commitment is small enough to be a controlled bet and large enough to anchor a credible inference offering, with a named customer already attached so the capacity is not speculative. For Together AI, dedicated access to the newest NVIDIA systems inside a large enterprise cloud gives it room to grow its token volume and a distribution channel into IBM's customer base, without having to finance and operate the data center itself.
The open-source inference thesis
Underneath the deal is a specific bet about how enterprises will buy AI. Together AI's business is built on the premise that open-weight models, the kind whose parameters are published and can be run by anyone, are becoming good enough that many companies will prefer to run them on cost-efficient infrastructure rather than pay per token for the most expensive closed models. The value Together AI adds is the engineering to serve those models quickly and cheaply at scale.
That thesis has become more plausible through 2026 as open-weight models from several labs have narrowed the quality gap with the leading closed systems. It is not a settled question; the closed frontier labs continue to lead on the hardest tasks, and "good enough and much cheaper" is a judgment that varies by use case. But the direction of the deal is clear. IBM and Together AI are wagering that a meaningful share of enterprise inference will run on open models on rented NVIDIA capacity, priced against the closed labs rather than matching their per-token rates.
It is worth naming who benefits regardless of how that bet resolves: NVIDIA. Whether enterprises standardize on closed frontier models or open-weight alternatives, the inference runs on GPUs, and this deal routes more of it onto NVIDIA's newest systems and its networking. The chipmaker supplies the hardware, its performance claims frame the marketing, and it captures value from the growth of inference no matter which model layer wins.
The risks worth naming
A neutral read has to hold several open questions at once. The first is demand timing. The cluster is not expected to be available until the first quarter of 2027, and a purpose-built inference cluster only earns its cost if it fills with paying workloads. That requires enterprise inference demand, and specifically demand for open-model inference on rented capacity, to keep growing through the build-out. If adoption stalls or shifts back toward closed, vertically integrated providers, dedicated capacity is expensive to carry.
The second is concentration. Like most of the AI build-out, this cluster is single-architecture: NVIDIA silicon, NVIDIA networking, NVIDIA performance claims. That is the pragmatic choice today, but it leaves both companies exposed to NVIDIA's pricing, supply, and roadmap, and it is the same concentration that regulators and large buyers have started to scrutinize across the industry.
The third is scale relative to the field. At $240 million and an initial footprint of roughly 2,000 GPUs, this is a real commitment but a small one next to the tens of billions that hyperscalers are spending on AI infrastructure. It buys IBM a credible seat in inference; it does not by itself make IBM Cloud a front-rank AI platform. Whether this becomes a foothold or a one-off depends on demand that will not be visible until well into 2027.
The read-through
For teams building products on top of AI models rather than operating the data centers, the useful takeaway is structural. Inference is becoming its own distinct market, with its own specialized providers, its own economics, and its own capacity deals, separate from the training clusters that get most of the attention. Where a given request is actually served, on which provider's rented GPUs, running which model, is increasingly a supply-and-cost decision rather than a fixed property of the model you chose.
That is an argument for keeping the model and the serving layer loosely coupled. A platform that can route a workload across providers and models, as Metir AI is built to do, treats a new inference cluster like this one as another option in the supply pool rather than a dependency to migrate onto or away from. When capacity, price, and model quality all move independently, the ability to change where and what you run without rewriting your product is worth more than any single provider relationship.
IBM's $240 million agreement with Together AI will not reorder the AI infrastructure landscape on its own. What it captures is a moment in that landscape's evolution: an established enterprise vendor buying back into frontline AI through a specialist inference partner, the newest NVIDIA silicon aimed squarely at serving rather than training, and a bet that open models on rented capacity are becoming a large enough market to build dedicated clusters around. Whether that bet pays off is a 2027 question. That serious money is being placed on it is a fact of 2026.
Sources:
- IBM and Together AI Sign Multi-Year Agreement to Scale Open-Source AI Inference with NVIDIA AI Infrastructure on IBM Cloud | IBM Newsroom
- IBM, Together AI ink $240 million deal for Nvidia-powered AI inference cluster | Reuters via BNN Bloomberg
- IBM inks $240M infrastructure deal with AI-optimized cloud operator Together AI | SiliconANGLE
- IBM Lands $240M Together AI Deal To Build Nvidia-Powered Neocloud Cluster | Stocktwits
Image credits
Header image: Google's data center in The Dalles, Oregon, with its electrical substation in the foreground, shown to illustrate hyperscale data-center facilities and their power infrastructure generally. The IBM and Together AI cluster is a separate, unrelated facility. By Visitor7 via Wikimedia Commons, licensed under CC BY-SA 3.0. In-body photograph: engineers working on a rack-mounted compute cluster at NASA's Glenn Research Center, shown to illustrate cluster hardware generally and unrelated to the systems in this deal. By NASA via Wikimedia Commons, public domain.
