metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
AI Infrastructure
Inference
Modal
Baseten
AI Funding
Cloud

Modal at $15B, Baseten at $26B: The Inference Layer Reprices

Modal and Baseten are reportedly in talks at $15B and $26B as spending shifts from training AI models to running them. A neutral look at why the inference layer is repricing.

Metir AI TeamSeptember 26, 20268 min read
Modal at $15B, Baseten at $26B: The Inference Layer Reprices

Two of the companies that help businesses run AI models are repricing fast. Modal Labs is in talks to raise at a valuation of roughly $15 billion, about triple its worth in a round only four months earlier, while Baseten is discussing a round that could value it near $26 billion, up from $13 billion in June, according to Bloomberg reporting on September 23, 2026. Neither company trains frontier models or owns the bulk of the chips. They sit in the middle, in the inference layer, and that middle is suddenly where a lot of investor attention is landing. This piece explains what these companies do, why the money is moving toward inference, and what the surge does and does not tell us.

~$15BModalvaluation in talks
~$26BBasetenvaluation in talks
~3xModal repricingin about four months
Jun to SepBasetendoubled in a quarter
NVIDIA logoNVIDIA
Anthropic logoAnthropic
OpenAI logoOpenAI
The inference layer runs models from many labs on chips from many clouds, which is the source of its leverage.

What inference infrastructure actually does

In AI, there are two big phases of compute. Training is the one-time, enormous effort of building a model. Inference is everything after: actually running the model to answer a question, generate an image, or take an agentic step, over and over, for every user. Modal and Baseten operate at the inference stage. They take a trained model and make it something a developer can call reliably at scale, handling the unglamorous but essential work of autoscaling, batching requests, minimizing cold starts, and keeping latency and cost under control.

Where the inference layer sits

Serving companies occupy the middle of the stack. They do not train the models or own most of the chips; they make running models fast, cheap, and reliable for everyone above them.

AI applications and agents
Chatbots, copilots, and agentic workflows that call models many times per task.
|
Inference / serving layer
Modal, Baseten, Fireworks, Together and peers: autoscaling, batching, cold-start handling, and routing that turn a model into a reliable, fast API.
|
Models and GPUs
Open and proprietary model weights running on accelerators from clouds and neoclouds.

Illustrative. As spending shifts from training models to running them, more value concentrates in the middle layer.

That position has a particular kind of leverage. A serving company is not betting on which lab wins; it runs whichever models its customers want, on whichever chips are available. As the number of models and the volume of calls both grow, the serving layer sees more traffic almost regardless of how the model race shakes out. That is the structural case for valuing it richly.

Why the money is shifting toward inference

For years the dominant compute story was training: bigger models, bigger clusters, bigger training runs. In 2026 the narrative has tilted. By multiple industry estimates, spending on inference is expected to eclipse spending on training, as adoption broadens and as products move toward agents that make many model calls per task rather than one. An agent that plans, calls tools, checks its work, and retries consumes far more inference than a single chatbot reply.

“

Training is a one-time cost of building a model. Inference is the recurring cost of using it, and usage is what is scaling now.

On the shift in where AI compute spending goes

If inference is where the recurring, growing spend lives, then the companies that make inference cheaper and more reliable are positioned to take a cut of a rising tide. That is the thesis investors appear to be underwriting when they triple Modal's valuation in four months or double Baseten's in a quarter.

Yellow fiber-optic cable management above data-center network racks
Fiber cabling above data-center racks. The inference layer's job is to make the compute beneath it usable, turning raw GPUs and model weights into a fast, scalable API. Photo via Wikimedia Commons, CC BY-SA 3.0.

The valuations, in context

The scale of the repricing is easier to see side by side. Modal was reportedly in talks around a $2.5 billion valuation earlier in 2026 and valued near $5 billion in a round roughly four months before the latest discussions; the new talks put it near $15 billion. Baseten raised a $1.5 billion round at a $13 billion valuation in June and is now discussing roughly $26 billion. Both companies declined to comment on the talks, and it is worth stressing that these are valuations under discussion, not closed rounds. The broader appetite is visible elsewhere too: the Finnish AI cloud startup Verda raised $189 million in late September at a valuation of at least $1 billion.

Inference startups repricing in months, not years

Reported valuations, in billions of dollars, before and during September 2026 funding talks. The later figures are valuations under discussion, not closed rounds.

Modal was valued near $5B about four months before the talks; Baseten was $13B in June. Later figures are reported valuations under discussion.

What the surge does and does not prove

A repricing this fast proves that investor conviction in the inference thesis is strong and that competition for stakes in the category's leaders is intense. It does not, by itself, prove the businesses have grown into these numbers. Valuations set in competitive private talks reflect appetite and scarcity as much as current revenue, and the inference layer faces real questions: hyperscalers and neoclouds offer overlapping serving capabilities, margins can compress as the work commoditizes, and a serving company's economics depend on GPU supply and pricing it does not fully control. The bull case and these pressures are both true at once.

The portability angle

The reason the inference layer can run everyone's models is the same reason it matters to builders: it treats the model as a swappable component. A serving platform that only worked with one lab's models would inherit that lab's pricing and availability; the ones drawing these valuations are valued precisely because they are model-agnostic and can route work to whatever runs best or cheapest.

That principle scales down to the products built on top. A platform like Metir AI applies the same logic at the application layer, staying model-agnostic and routing across providers so a product is not tied to any single model's cost or roadmap. The inference layer's repricing is, in part, the market putting a number on how valuable that portability has become.

The takeaway

Modal and Baseten repricing to $15 billion and $26 billion is the clearest recent sign that AI value is migrating from training models to running them. The inference layer is structurally attractive because it profits from rising usage regardless of which lab leads, and structurally exposed because the work can commoditize and the chips are not its own. Whether these specific valuations hold, the direction is the signal worth tracking: the recurring cost of using AI is now the part that is scaling.

Sources:

  • Startups Modal, Baseten in Funding Talks to Help Businesses Run AI | Bloomberg
  • AI inference startup Modal Labs in talks to raise at $2.5B valuation | TechCrunch
  • AI inference provider Baseten reportedly raising $1.5B in funding | SiliconANGLE
  • Baseten secures $1.5bn Series F funding for AI inference platform | Yahoo Finance
  • AI Cloud Startup Verda Raises $189 Million in Funding Round | Bloomberg
  • Modal and Baseten seek funding at $15B and $26B valuations | RuntimeWire

Image credits

Header image: supercomputer racks with network cabling and a power management module in the NERSC data center, by D Coetzee, via Wikimedia Commons, released under CC0; illustrative of large-scale compute infrastructure. In-body photograph of yellow fiber-optic cable management above data-center racks, by Robert.Harker, via Wikimedia Commons, licensed under CC BY-SA 3.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational