metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Qualcomm
AWS
AI Chips
Inference
Semiconductors

Qualcomm, AWS, and the Race to Own AI Inference Silicon

Qualcomm and Amazon signed a multi-generation deal to co-develop custom AI inference chips, with a warrant tied to up to $60B in spend. Here is why inference silicon is the new battleground.

Metir AI TeamSeptember 8, 20269 min read
Qualcomm, AWS, and the Race to Own AI Inference Silicon

On 8 September 2026, Qualcomm and Amazon Web Services announced a multi-generation strategic collaboration to co-develop custom silicon and high-bandwidth optical interconnects for AI data centers, aimed squarely at inference workloads. Qualcomm's stock rose about 10 percent on the news. The financial mechanics are as interesting as the engineering: Amazon received a warrant to acquire up to 25 million Qualcomm shares, with the full block vesting only if Amazon spends up to 60 billion dollars on Qualcomm chips, networking hardware, and manufacturing services through September 2036.

$60BSpend that fully vests the warrant (through 2036)
25MQualcomm shares in Amazon's warrant
Q1 FY2027When Qualcomm expects first revenue
~10%Qualcomm's stock move on the announcement

The deal is a clean illustration of two shifts reshaping the AI hardware market at once: the move of the center of gravity from training to inference, and the drive by every large cloud to reduce its dependence on a single chip supplier.

Training built the boom, inference will run it

For three years the AI hardware story was about training, the enormously expensive, compute-hungry process of building a model. Training is bursty, runs in giant clusters, and rewards raw throughput, which is why it made one company's top-end accelerators indispensable. But a trained model earns its keep in inference: every time it answers a question, writes code, or processes an image. Inference is continuous rather than bursty, runs at massive aggregate volume, and is dominated by a different cost function, performance per watt and per dollar, rather than peak training speed.

“

Training is a capital expense you pay once per model. Inference is an operating expense you pay on every single request, forever.

On why inference economics differ

As deployment scales, inference becomes the larger and more permanent cost. That is what makes it strategically attractive to a company like Qualcomm, whose entire heritage is efficient computing for battery-powered phones. Power efficiency was a constraint Qualcomm spent two decades optimizing for the smartphone, and it happens to be the exact property that matters most for high-volume inference in a power-constrained data center. The deal is Qualcomm repurposing a mobile-era strength for the data center at the precise moment the market shifted toward it.

A researcher inspecting a silicon wafer in a semiconductor facility
A silicon wafer under inspection. Custom AI chips begin as designs that are fabricated onto wafers like these before being cut into individual processors. Photo: U.S. Department of Energy, public domain, via Wikimedia Commons.

Why a hyperscaler that already makes chips wants another supplier

The subtle part of this deal is that AWS is not short of silicon. Amazon already designs its own AI chips, Trainium for training and Inferentia for inference, and it is one of the more advanced hyperscalers at building custom accelerators in-house. So why bring in Qualcomm?

The answer is that even a company building its own chips wants optionality. Designing leading-edge silicon is hard, slow, and risky; a single in-house roadmap is a single point of failure. Partnering with an outside designer that has deep low-power expertise gives AWS a second source, a competitive benchmark for its internal teams, and access to intellectual property it would otherwise have to build from scratch. The same logic that makes a cloud want multiple chip suppliers is why it is negotiating with more than one.

Every large cloud is building an alternative

The major hyperscalers have all developed or co-developed their own AI accelerators, concentrated first on inference. None has displaced the dominant merchant supplier for frontier training.

Amazon Web Services
Trainium and Inferentia in-house, plus the 2026 Qualcomm inference-chip partnership
Google
Tensor Processing Units (TPUs), now many generations deep
Microsoft
Maia AI accelerator
Meta
MTIA (Meta Training and Inference Accelerator)

Even a cloud that designs its own chips wants a second source: a single silicon roadmap is a single point of failure.

Read across the industry, the pattern is unmistakable. Google has its TPUs, Microsoft has Maia, Meta has MTIA, and Amazon has Trainium and Inferentia, and now a Qualcomm partnership on top. Every large buyer of AI compute is trying to build or co-develop an alternative to depending entirely on one merchant supplier. None of them has displaced that supplier for frontier training, but inference is the softer target, and it is where the diversification is concentrating first.

The warrant is the real design

The financing structure deserves as much attention as the chips. Rather than a simple supply contract, the deal ties Amazon's equity upside in Qualcomm to Amazon's own spending: the warrant vests as Amazon buys, up to the full 25 million shares at 60 billion dollars of cumulative purchases over roughly a decade. This aligns the two companies in a specific way. Amazon is incentivized to route real volume through Qualcomm, because doing so converts into ownership; Qualcomm gets a credible signal of demand it can plan capacity around.

A warrant that vests as Amazon spends

Rather than a plain supply contract, the deal ties Amazon's equity upside in Qualcomm to Amazon's own purchasing.

Multi-generation co-development
Qualcomm and AWS jointly develop custom AI inference silicon and high-bandwidth optical interconnects for data centers
aligned by
↓
Warrant for up to 25M Qualcomm shares
Issued to Amazon, vesting fully only if Amazon spends up to $60B on Qualcomm chips, networking hardware and manufacturing services through September 2036
expected to produce
↓
Revenue from fiscal Q1 2027
Qualcomm says chips are already in production; the stock rose about 10 percent on the announcement

The structure signals real commitment, and it means part of the demand for these chips is contractually arranged rather than open-market.

It is worth being clear-eyed about what a structure like this does and does not prove. It signals genuine commitment and de-risks Qualcomm's investment in a new product line, which is the benign and probably correct reading. It also means some portion of the "demand" for these chips is demand the two parties have contractually manufactured, which makes the deal a weaker independent indicator of open-market pull than a plain purchase order would be. Both things are true at once, and the honest way to read vendor-and-customer equity entanglements, which are now common across AI hardware, is to hold them together rather than pick one.

What it means for everyone buying compute

For organizations that consume AI rather than sell chips, the diversification of inference silicon is quietly good news and a quiet warning. Good, because more suppliers competing on performance per watt should push inference costs down over time. A warning, because the hardware layer is fragmenting into incompatible stacks, each cloud's chips, each with its own software toolchain, and a workload optimized for one can be expensive to move to another.

That is the same dependence problem the whole industry keeps rediscovering at every layer. The defense against it is portability: keeping AI work able to run across providers and, increasingly, across the silicon underneath them, rather than welded to one vendor's stack. Platforms like Metir that stay model-agnostic apply that principle at the software layer, and the scramble among hyperscalers to secure a second and third chip source is the same instinct expressed in silicon and capital.

The measured read on the Qualcomm and AWS deal is that it is a substantive bet on inference becoming the center of the AI hardware market, backed by a financing structure that binds the two companies tightly and signals demand that is partly contracted rather than purely organic. It does not dethrone the dominant training-chip supplier, and it is not meant to. It is a bet on the larger, longer, more power-sensitive half of the market, placed by two companies each trying to reduce how much of that market they have to rent from someone else.

Sources:

  • Qualcomm lands Amazon deal to build custom AI data center chips | Quartz
  • Qualcomm Teams With AWS on AI Chips, Taking On Nvidia | Seoul Economic Daily
  • Qualcomm Wins First Western Hyperscaler: AWS Deal Pays Up to $60B for Inference Silicon | TechTimes
  • Qualcomm Strikes Major Amazon AWS Deal For Custom AI Data Center Chips | HotHardware

Image credits

Hero image: a 12-inch silicon wafer, by user Peellden, via Wikimedia Commons, licensed under CC BY-SA 3.0. In-body photograph: a researcher inspecting a silicon wafer, by the U.S. Department of Energy, public domain, via Wikimedia Commons. Both images are generic illustrations of semiconductor manufacturing and do not depict Qualcomm or AWS facilities.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational