Google is reportedly developing a server chip, informally called Frozen v2, that would etch part of its Gemini model's architecture directly into the silicon. According to a report from The Information, engineers project the chip could deliver 6 to 10 times more tokens per watt than Google's latest TPUs, with a possible deployment as soon as 2028. Google has not confirmed the project, and a spokesperson noted that not every internal effort reaches production.
Even as an unconfirmed report, Frozen v2 is worth understanding because it crystallizes a bet the whole industry is edging toward: that the path to cheaper AI runs through hardware built for one model rather than hardware built for many. That bet carries a real and specific cost, and the tension between the two is the story.
What "hardwiring a model" means
A conventional AI accelerator, whether an Nvidia GPU or a Google TPU, is a general-purpose engine. It can run essentially any neural network, because the model's structure lives in software and the chip simply executes whatever math that software describes. That flexibility is valuable, but it has an energy cost: the chip constantly shuttles data around and makes runtime decisions that a fixed design would not need to make.
Frozen v2 would take a different approach. By fixing parts of Gemini's architecture into the transistors themselves, the chip removes steps and reduces the data it has to move for each query. The name is telling: the model's structure is "frozen" into hardware rather than loaded from software. The reported result, 6 to 10 times the tokens per watt, is the payoff for giving up generality.
Hardwiring one model into silicon: the projected efficiency prize
Reported tokens-per-watt for Google's "Frozen v2" inference chip relative to its latest TPUs, indexed to a baseline of 1x.
These are engineers' internal projections reported by The Information, not benchmarked results. Google has not confirmed the project, and any deployment is targeted for as soon as 2028.
It helps to place this on a spectrum. General-purpose GPUs sit at one end, flexible but less efficient per watt. Google's TPUs and other AI accelerators sit in the middle, tuned for machine learning broadly. A model-specific chip like Frozen v2 sits at the far end: maximally efficient for exactly one architecture, and useless for anything that departs from it. Each step toward specialization trades flexibility for performance per watt.
NVIDIA
GeminiWhy Google would take the risk now
The reported motivation is blunt: compute scarcity. The AI compute shortage has been severe enough that Google Cloud has reportedly turned down deals with outside customers because it needs the capacity for its own products. When you cannot buy or build enough general-purpose compute to meet demand, squeezing far more output from each watt and each square millimeter of silicon becomes an obvious lever, and a model-specific chip is one of the most aggressive ways to pull it.
A chip built for one model is a bet that the model will not change. That is the whole wager, in a single sentence.
Metir AI analysis
There is also an economic logic beyond raw scarcity. Inference, the cost of actually running a model to answer queries, now dominates the lifetime compute bill for a widely used model, far outweighing the one-time cost of training it. Shaving inference energy by a large multiple compounds across billions of queries. For a model deployed at Gemini's scale, even a fraction of the projected efficiency gain would translate into enormous savings and a meaningful edge on price.
The cost hiding inside the efficiency
The catch is structural, and it is the reason most companies have not done this. Because Gemini's architecture is baked into the transistors, Frozen v2 can only serve future Gemini versions if Google keeps its foundational architecture intact. Any significant architectural change, the kind that has repeatedly driven the biggest leaps in model quality, could render the specialized silicon obsolete before it pays back its enormous design and fabrication cost.

That is the wager in full. In a field where the dominant architecture has shifted every couple of years, committing a multi-year, multi-billion-dollar chip program to today's Gemini design is a statement of confidence that the fundamentals are stabilizing. Google's own spokesperson hedged exactly here, framing the work as exploration rather than a committed product and noting that not every project moves into production. The 2028 target leaves ample room for the architecture to move first.
What it signals about the market
Frozen v2 fits a broader 2026 pattern of major AI labs pushing into custom silicon to control cost and supply, from Meta's inference chips to Anthropic's silicon partnerships. The common thread is a desire to escape both the price and the scarcity of merchant GPUs by building hardware matched to a specific workload. Model-specific inference chips are the most extreme version of that instinct.
For everyone building on top of these models rather than designing the chips, the takeaway is subtler. As providers optimize their own stacks in divergent ways, hardwiring silicon here, restructuring pricing there, the performance and cost of a given task can shift meaningfully between vendors and between quarters. That is an argument for keeping workloads portable rather than fused to one provider's roadmap. Platforms like Metir AI apply that logic at the application layer, giving teams access to Gemini, GPT, Claude and other models side by side so a hardware or pricing shift at any one lab can be adopted or routed around without re-platforming.
Looking ahead
Whether Frozen v2 ships at all is genuinely uncertain, and Google has been careful to say so. But the direction it points is not uncertain. As inference becomes the dominant cost of AI and general-purpose compute stays scarce, the pressure to specialize hardware will keep building, and the central question will be how much architectural flexibility labs are willing to surrender for efficiency. Frozen v2 is one lab's early, aggressive answer. The interesting thing to watch is whether the frontier architecture stabilizes enough to make that answer look prescient rather than premature.
One workspace, every leading model
Hardware and pricing at the frontier labs keep shifting in different directions. With Metir AI you get unified access to Gemini, GPT, Claude, Grok and other top models in a single workspace, so you can route each task to whichever performs best and adapt the day a provider changes its stack. Try Metir AI free.
Sources:
- Google reportedly developing 'Frozen v2' chip with Gemini's architecture etched into the silicon | Tom's Hardware
- Google just bet its inference future on a chip built for one model | The New Stack
- Google Reportedly Developing 'Frozen v2' AI Chip to Boost Gemini Efficiency | Quiver Quantitative
- Google develops Frozen v2 AI inference chip, up to 10 times more efficient than TPU | Digital Today
- Google's Frozen v2 Chip Hardwires Gemini Architecture: Up to Tenfold Inference Efficiency | Tech Times
Image credits
Header image: a 12-inch silicon wafer patterned with integrated-circuit dies, by Peellden via Wikimedia Commons, licensed under CC BY-SA 3.0. In-body photograph of server racks inside the BalticServers data center, by BalticServers.com via Wikimedia Commons, licensed under CC BY-SA 3.0. Neither photo depicts the Frozen v2 chip or a Google facility.
