AI inference chip startup Etched is reviewing funding offers that would value it at between $40 billion and $50 billion, according to a TechCrunch report published on October 5, 2026 and citing people familiar with the matter. The low end of that range reportedly comes from top-tier investors and the high end from lesser-known backers. Etched declined to comment, and no round has been announced, so these are offers rather than a closed valuation.
The context is what makes the report notable. Etched raised $700 million at a $21 billion valuation in August, led by Jane Street, and $300 million at $10.3 billion in July, led by Sequoia. Early in 2026 it was valued at around $5 billion. If a deal closes anywhere in the reported range, the company's private valuation would have risen roughly eightfold in about nine months. This piece looks at how that trajectory came together, why inference silicon is attracting this much capital, and the technical bet underneath it: splitting the two phases of running a language model, prefill and decode, onto hardware built for each.
The valuation trajectory, round by round
The run of financings is unusually compressed, and it helps to set the dates side by side.
- January 2026, about $5 billion. Etched raised roughly $500 million in a round led by Stripes, with participation from Peter Thiel, Positive Sum and Ribbit Capital, as reported by Bloomberg on January 13 and covered by Yahoo Finance. That brought total funding to nearly $1 billion. TechCrunch dates this $5 billion mark to December 2025.
- July 23, 2026, $10.3 billion. A $300 million Series C led by Sequoia, with Andreessen Horowitz, Jane Street, Diffusion and SK Hynix participating. Etched said it was the highest valuation ever for a Sequoia-led Series C.
- August 18, 2026, $21 billion. A $700 million round led by Jane Street, which TechCrunch reported had tested the chips and had its own rack running in its data center. We covered that round in more depth in our earlier analysis of the $21 billion raise.
- October 2026, offers at $40 billion to $50 billion. Under review, per TechCrunch's sources.
Etched valuation: closed rounds vs reported offers
Post-money valuation in billions of US dollars. Green bars are closed rounds; grey bars are offers reported on October 5, 2026, not a completed financing.
Sources: Bloomberg and Yahoo Finance (Jan 2026), WMBD and Yahoo Finance (Jul 2026), TechCrunch (Aug 18 and Oct 5, 2026). Hover a bar for details.
TechCrunch frames the July and August rounds as back-to-back financings that effectively work as a single raise split into two tranches, which is a useful lens. Read that way, Etched went from about $5 billion to $21 billion in one extended financing cycle, and the new offers would roughly double or more than double that again. TechCrunch also notes that if Etched raises as much as its last round, the cash could give it as much as 3.5 years of runway. Adding the disclosed rounds, our own arithmetic puts total funding at roughly $2 billion before any new deal.
A private valuation is a negotiated price for a minority stake, often with terms such as liquidation preferences that are not disclosed. A headline number is therefore a signal about investor demand, not an audited measure of what the company is worth. That caveat applies to every bar in the chart above.
Why inference is where the money is going
Training a frontier model is a large but episodic expense. Inference, running a trained model every time a user sends a prompt or an agent takes a step, is continuous and grows with usage. That is the market Etched is built for, and it is crowded with well-funded challengers to Nvidia.
The recent comparables show how differently that market is pricing outcomes:
- Cerebras went public on Nasdaq in May 2026, raising $5.5 billion at $185 per share. TechCrunch reported a fully diluted valuation of $56.4 billion at the IPO price and about $66 billion at the first-day close.
- Groq raised $350 million at a $3.5 billion valuation in August 2026, according to TechCrunch, down from $6.9 billion in September 2025. In between, Nvidia hired founder and CEO Jonathan Ross and other senior staff as part of a $20 billion licensing deal, and Groq shifted toward operating Nvidia systems as a cloud provider.
- Positron AI announced $875 million at a $5 billion post-money valuation on September 10, 2026, for chips that use commodity LPDDR5X memory instead of scarce high-bandwidth memory.
NVIDIAPut together, the comparables suggest the market is not pricing "inference chips" as a single category. It is pricing specific companies on evidence of deployment and customer demand. Cerebras had public-market investors; Groq's independent chip roadmap was effectively absorbed by Nvidia; Positron is pre-scale on its next chip. Etched's case rests on the $1 billion in customer orders it announced in July and on a paying customer, Jane Street, leading its last round.
Prefill and decode: the two halves of inference
To see what Etched says it has built, it helps to understand that generating text with a transformer is two different workloads.
Prefill is when the model reads the whole prompt, including instructions, documents and conversation history, and builds an internal cache of keys and values for every token. All prompt tokens can be processed in parallel, so prefill is dominated by large matrix multiplications. The academic literature describes it as compute-bound: the limit is how much arithmetic the chip can do.
Decode is when the model writes its answer one token at a time. Each new token depends on the previous one, so the work cannot be parallelized across the output in the same way. Each step does relatively little arithmetic but has to read the model weights and the growing key-value cache from memory. That makes decode memory-bound: the limit is how fast data can be moved, not how fast it can be multiplied.
Prefill wants arithmetic. Decode wants memory bandwidth. A chip that is excellent at one is not automatically excellent at the other.
The core tension in inference hardware
Running both phases on the same general-purpose GPU means compromises. The 2024 DistServe paper from academic researchers found that colocating the phases causes strong interference between them, and that assigning them to separate GPUs let a serving system handle up to 7.4 times more requests, or meet latency targets 12.6 times tighter, than the state-of-the-art systems it compared against. A 2025 paper, SPAD, went further and modelled purpose-built prefill and decode chips. Its authors report that their specialized prefill chip delivered 8% higher prefill performance on average at 52% lower hardware cost than a modelled H100, and that the decode chip reached 97% of the decode performance with 28% lower power. Those are simulation results, not shipping products, but they show why the idea is taken seriously.
Nvidia has moved in the same direction. In September 2025 it announced Rubin CPX, a GPU aimed at long-context processing with 128 GB of GDDR7 memory, positioned for the context-heavy first phase of inference. The convergence matters: when the market leader also splits the workload, disaggregation looks less like a startup thesis and more like where inference hardware is heading.
What Etched says it built for each phase
Etched's approach, as co-founder and COO Robert Wachen described it to TechCrunch in August, mirrors that split. For prefill, the company built a chip that operates at low voltage, which it says allows greater transistor density. For decode, it built what it calls cluster-scale memory, which lets many chips connect and share a pool of memory at very low latency. In the October report, Wachen is quoted as saying the company "designed two new components from scratch to speed up inference."

The company has also invested in the physical side. TechCrunch reports that Etched operates a new 10-megawatt data center in Silicon Valley and has established a facility in Taiwan to coordinate production near TSMC, its manufacturing partner. Its July announcement mentioned an 80,000-square-foot facility near its San Jose headquarters. Roughly 15% of its approximately 400 employees previously worked at Nvidia.
None of the published coverage reviewed for this piece includes independent benchmarks of Etched's systems. Performance claims relative to Nvidia hardware are the company's, and the strongest outside signal remains customer behavior: the order book and Jane Street's decision to deploy a rack and then lead a round.
Second-order implications
A few dynamics are worth watching beyond the headline number.
Specialization risk. Etched's chips are built around the transformer. That is a sound bet while transformers dominate deployed models, but a hardwired design has less room to adapt if architectures shift. The disaggregation trend partly hedges this: memory-bound decode exists for any autoregressive model, not only transformers.
Supply and manufacturing. Capital is only useful if it converts into chips. Production at a leading foundry, memory supply and data center power are shared constraints across the whole industry. Positron's choice of commodity memory and Etched's Taiwan facility are both, in different ways, responses to that bottleneck.
Valuation reflexivity. Fast re-marks can help a startup hire and sign customers, since buyers want suppliers that will exist in five years. They also raise the bar for the next round. Groq's reset from $6.9 billion to $3.5 billion shows that marks in this sector can move down as well as up when circumstances change.
Customers stay multi-vendor. For teams building on AI, the hardware layer is fragmenting into GPUs, hyperscaler silicon such as Google TPUs and AWS chips, and specialized startups. Most application builders will consume this through APIs rather than buying racks, which is why model-agnostic platforms such as Metir, which let users switch among models from many providers, sit downstream of these shifts rather than betting on one of them.
What to watch next
- Whether a round closes, at what valuation, and who leads it. TechCrunch's distinction between top-tier and lesser-known bidders implies the investor names may matter as much as the price.
- Conversion of the $1 billion in orders into deployed, revenue-generating systems.
- Independent benchmarks of Etched hardware against Nvidia's current and next-generation parts.
- Whether disaggregated prefill and decode becomes standard in commercial serving stacks, which would broaden the market for specialized chips of every kind.
FAQ
What valuation is Etched being offered? According to TechCrunch's October 5, 2026 report, Etched is reviewing offers ranging from $40 billion from top-tier investors to $50 billion from lesser-known backers. No round had been announced and Etched declined to comment.
How much has Etched raised before? About $500 million at roughly $5 billion (reported in January 2026), $300 million at $10.3 billion in July 2026 led by Sequoia, and $700 million at $21 billion in August 2026 led by Jane Street.
What is the difference between prefill and decode? Prefill processes the prompt in parallel and is compute-bound. Decode generates the response one token at a time and is memory-bandwidth-bound. Etched says it designed separate components for each.
Who are Etched's main competitors? Nvidia is the incumbent. Other inference-focused companies include Cerebras, now public, and Positron AI, while Groq has shifted toward running Nvidia systems as a cloud provider.
Sources:
- Etched fields funding offers at $40B+ valuation, sources say, TechCrunch (Oct 5, 2026)
- Etched's valuation doubles to $21B in a month, TechCrunch (Aug 18, 2026)
- AI chip startup Etched raises $300 million at $10.3 billion valuation, WMBD (Jul 23, 2026)
- AI Chip Startup Etched Raises $500 Million to Take on Nvidia, Bloomberg (Jan 13, 2026)
- AI Chip Startup Etched Raises $500 Million in New Funding Round, Yahoo Finance
- Cerebras raises $5.5B, then stock pops 108%, TechCrunch (May 14, 2026)
- Groq raises $350M to fuel its pivot from AI chips to neocloud, TechCrunch (Aug 17, 2026)
- Positron AI Raises $875 Million at a $5 Billion Valuation, PR Newswire (Sep 10, 2026)
- NVIDIA Unveils Rubin CPX, NVIDIA Newsroom (Sep 9, 2025)
- DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, arXiv
- SPAD: Specialized Prefill and Decode Hardware for Disaggregated LLM Inference, arXiv
Image credits
- Hero: a 300mm silicon photonics wafer patterned with chip dies, by Ehsanshahoseini via Wikimedia Commons, licensed under CC BY-SA 4.0. It is not an Etched wafer.
- In-body: cleanroom at the CNE 300mm wafer foundry, by Abkorshak via Wikimedia Commons, licensed under CC BY-SA 4.0.
