On August 20, 2026, a model called "stealth/ox-alpha" appeared on OpenRouter with no maker attached, a million-token context window, and a free preview period. Within days it was the subject of the AI coding community's favorite recurring ritual: an anonymous model posts a startling benchmark number, independent testers scramble to reproduce it, and amateur forensics teams try to unmask who actually built it. The number this time was a reported 80% on DeepSWE, a closed coding benchmark, which would put OX Alpha ahead of Claude Fable 5 and GPT-5.6 Sol on that specific test. The number is real in the sense that someone measured it. Whether it means what the headlines implied is a separate question, and it is the more useful one to answer.
Z.ai
AnthropicWhat a "stealth model" actually is
OpenRouter, an AI model routing marketplace, periodically lists models under generic codenames instead of a maker's name. "Stealth" listings let a lab collect real-world usage data and community feedback before a public launch, generate hype and press coverage without committing to a release date, and give competitors and journalists a benchmark number to react to without a company's reputation attached if the number does not hold up under scrutiny. It is a low-risk way to test the market's reaction to a model that might still be weeks from a formal release, or might never ship under the tested configuration at all.
OX Alpha fits the pattern closely. Its OpenRouter listing describes a roughly 1,048,576-token context window, a maximum output of 131,072 tokens, and support for text, image, and video input with structured tool calling. It was free to use during its preview window, which closed around August 27, 2026. None of that, on its own, is unusual for a 2026-era frontier-class release. What made OX Alpha a story was the benchmark claim layered on top.
The number that went viral, and its asterisk
Independent researcher Ben Davis ran OX Alpha against DeepSWE, a coding benchmark distinct from the more commonly cited SWE-bench Verified, and reported roughly 80% Pass@1, ahead of Claude Fable 5's approximately 65% and GPT-5.6 Sol's approximately 52% on the same test. That comparison spread quickly across AI news aggregators and social media as evidence that an anonymous, free model had leapfrogged two of the most capable systems on the market.
The number that went viral, and its asterisk
DeepSWE Pass@1 scores from independent researcher Ben Davis's test, reported around August 21-22, 2026. This was a 10-task sample, not an audited leaderboard run, and DeepSWE's public BenchSift leaderboard did not list OX Alpha at the time.
A 10-task sample carries enormous variance. Read this as "frontier-band on this benchmark," not a confirmed win over GPT-5.6 Sol or Claude Fable 5.
The important detail, buried well below most of the headlines, is that Davis's test covered ten tasks. A ten-task sample carries enormous statistical variance; a single additional pass or failure shifts the reported score by ten percentage points. DeepSWE's own public leaderboard, BenchSift, did not list OX Alpha as of August 21, 2026, meaning the number circulating was a private, unaudited test result rather than a verified leaderboard entry. The honest reading of the data is that OX Alpha appears to sit in the frontier band on coding tasks, competitive with the strongest models available. It is not a confirmed, statistically reliable win over GPT-5.6 Sol or Claude Fable 5, and treating a ten-task sample as a settled ranking is the kind of overreach that makes stealth-model hype cycles unreliable in the first place.
A 10-task test can make almost any model look like it topped the leaderboard. The honest reading is frontier-band on coding, not a confirmed win.
Analysis of the OX Alpha benchmark claim
Who is likely behind it
Independent analysts, again led by Davis, used tokenizer alignment, video-encoder token-consumption patterns, and stylistic output fingerprinting to argue with stated confidence in the 90 to 99 percent range that OX Alpha is an unreleased model in Z.ai's (formerly Zhipu AI) GLM-5.x line, pointing specifically at overlap with GLM-5V-Turbo's video handling and GLM-5.3's tokenizer. A separate architectural estimate placed OX Alpha at roughly 744 billion total parameters with about 40 billion active in a mixture-of-experts configuration, a shape that matches the base architecture Z.ai has used across its GLM-5.2 and GLM-5.3 releases. Z.ai has not confirmed the attribution.
That speculation sits inside a pattern rather than standing alone. Every anonymous "stealth/" listing OpenRouter has hosted through 2026 has, once revealed, traced back to a Chinese lab: Pony Alpha turned out to be an early GLM-5 checkpoint from Zhipu, Hunter Alpha was Xiaomi's MiMo-V2-Pro, Elephant Alpha belonged to Ant Group's Inclusion AI line, and Owl Alpha was Meituan's LongCat-2.0. OX Alpha would be the fifth entry in that sequence if the Zhipu attribution holds, which is informed circumstantial reasoning, not a confirmed fact.
Every OpenRouter stealth model so far has been Chinese
Anonymous "stealth/" listings on OpenRouter through 2026, and who each was revealed (or is suspected) to be. OX Alpha is the only one still unconfirmed.
Illustrative timeline based on independent reporting, not an official OpenRouter record. The first four attributions were subsequently acknowledged; OX Alpha's was not, as of publication.

Why this pattern keeps repeating
The incentive structure explains the recurrence better than any single model's specs do. A lab that ships under its own name and underperforms takes a reputational hit that follows it into the next launch cycle. A lab that ships anonymously and underperforms simply lets the listing quietly expire. Anonymous testing also generates a specific kind of coverage, the "who built this" guessing game, that a named launch cannot replicate; the mystery is part of what made OX Alpha a bigger story than a same-quality named release from an established lab would have been. None of that is unique to Chinese labs, but the fact that OpenRouter's five stealth listings have all traced to Chinese developers so far suggests it has become a specific go-to-market tactic in that part of the industry, likely tied to wanting Western developer feedback and usage data ahead of a formal international release.
What to watch next
Three things would meaningfully firm up this story. First, whether OX Alpha or its successor appears on DeepSWE's audited BenchSift leaderboard with a larger task count, which would replace the viral ten-task number with something statistically meaningful. Second, whether Z.ai confirms or denies the attribution once its next GLM release ships. Third, what OX Alpha is priced at once its free preview ends, since stealth-model pricing after launch is often where the real competitive positioning becomes visible.
For teams evaluating coding models, the practical lesson is less about this specific model and more about how fast the field underneath it moves. A model with no confirmed maker held the coding conversation for a week; by the time OX Alpha's origin is settled, there will likely be another release to evaluate. That is the case for keeping the model layer swappable rather than locked to a single provider's roadmap. Platforms like Metir AI that route across models rather than committing to one make it possible to try a new entrant like OX Alpha the week it appears, without re-platforming a product around a model whose backing company will not confirm it exists.
Sources:
- Anonymous AI Model 'OX Alpha' Crushes Coding Benchmarks | Pandaily
- OX Alpha coverage | hyper.ai
- AI News Today, August 22-23, 2026 | Build Fast with AI
- OX Alpha Stealth Model: Comprehensive Analysis | local-ai-zone.github.io
Image credits
Header image: a software developer working across a laptop and monitor, by Jedchela via Wikimedia Commons, licensed under CC BY-SA 4.0. Illustrative of software development generally; does not depict OX Alpha or anyone involved in this story. In-body photograph of Zhongguancun Street in Beijing's Haidian technology district by N509FZ via Wikimedia Commons, licensed under CC BY-SA 4.0. Illustrative of the district, not of any specific company's offices.
