On October 5, 2026, Reflection AI introduced Beam, a sparse mixture-of-experts (MoE) language model with 501 billion total parameters, of which 23 billion are active for any given token. The company says the weights, a technical report and a model card will be released under the Apache 2.0 license later in October, and that the model is currently in final red-teaming with an early access waitlist, according to Reflection's launch post. A day earlier we covered the pre-launch reporting in Reflection AI's first open-weight model: what we know. This piece covers what was actually announced: the specifications, the self-reported benchmarks, how the MoE design shapes the economics, and where Beam sits in the US and China open-weight race.
What Reflection announced
Reflection, founded in 2024 by former Google DeepMind researchers Misha Laskin and Ioannis Antonoglou, describes Beam as a text-only model built for coding, reasoning and agentic work. The key specifications from the launch post:
- Architecture: sparse MoE with fine-grained routed experts (the expert count is not disclosed), 52 layers, and interleaved local and global attention.
- Context: reinforcement learning was run at 256K tokens, and midtraining extended the context window to 1 million tokens.
- Data: 23.8 trillion pretraining tokens drawn from web and proprietary licensed datasets, with an emphasis on source code, technical documentation, STEM and mathematics. Reflection says its curation removed 95% of raw internet tokens.
- Release: Apache 2.0 weights, a technical report, a model card and developer tooling for running, evaluating and fine-tuning, launching with distribution partners.
Reflection's central claim is efficiency rather than leadership. It says Beam is "competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks," and that it reaches reasoning scores comparable to GLM-5.2 "while using 3 to 4x less inference compute." Those are the company's own measurements, and as Startup Fortune and Crypto Briefing both note, they have not yet been independently verified.
How Beam was trained
The training disclosures are unusually detailed for a launch post, and they show how much of the work now happens after pretraining. According to Reflection:
- The base model was pretrained on 6,144 Nvidia GB300 GPUs in NVL72 systems in under four weeks, at a reported goodput of 92.3%.
- The reinforcement learning (RL) run used about 10,500 GB300 GPUs for four weeks and generated more than 100 million rollouts across nearly 1 million environments.
- Grading those rollouts consumed roughly 1.3 billion sandboxes over the run, with up to 170,000 running concurrently.
One derived observation: the RL phase used about 1.7 times as many GPUs as pretraining for a similar duration. That ratio is consistent with a broader industry shift in which post-training on verifiable tasks, especially coding and terminal work, is a large share of the compute budget rather than a short finishing step. It also explains why Reflection lists Beam's RL rollout count next to the counts it attributes to other labs, treating RL scale as a headline specification.
The compute itself is the subject of our earlier post, which documented Reflection's GB300 capacity agreements with SpaceX and Nebius running through 2029.
The benchmarks, read carefully
Reflection published a broad table comparing Beam with open models from China and the United States. The scores below are all as reported by Reflection.
Beam vs leading Chinese open-weight models
Scores as published by Reflection at launch. Higher is better. Not independently verified.
Source: Reflection, Introducing Beam (October 5, 2026). Comparison scores are those Reflection reported. Hover a group for every value.
A few patterns stand out when the full table is read rather than the headline numbers:
- Coding is competitive but not leading. Beam's 80.9 on SWE-Bench Verified and 78.0 on SWE-Bench Multilingual have no listed comparison scores from rivals in Reflection's table, so they cannot be ranked from this source. On Terminal Bench v2.1, where comparisons are listed, Beam's 80.1 sits just below GLM 5.2 at 81.0 and below Kimi K3 (88.3), Qwen 3.8-Max (86.6) and DeepSeek V4.1 (90.6). On SWE Bench Pro v1, Beam's 65.5 is above GLM 5.2 (62.1) and below Qwen 3.8-Max (67.7).
- Hard reasoning trails the larger models. On Humanity's Last Exam without tools, Beam scores 36.2 against 40.5 for GLM 5.2 and 46.9 for Kimi K3. On GPQA Diamond the spread is narrow, from 90.5 for Beam to 93.5 for Kimi K3. AIME 2026 is near saturation, with Beam at 97.8 and GLM 5.2 at 99.2.
- Against US open peers, Beam leads on most shared tests. Compared with Thinking Machines' Inkling and Nvidia's Nemotron 3 Ultra, Beam scores higher on HLE (36.2 vs 29.7 and 26.7), SciCode (49.7 vs 46.1 and 44.6) and GPQA Diamond (90.5 vs 87.2 and 87.0).
- Agentic tool use is mixed. On MCP Atlas Beam scores 78.7, ahead of GLM 5.2 (77.8) but behind Qwen 3.8-Max (84.5). On BrowseComp with context management it posts 77.4 against 91.2 for Kimi K3.
Beam is not pitched as the strongest open model. It is pitched as a strong one that is cheaper to run.
Reading Reflection's launch claims
Two caveats apply to any vendor benchmark table. First, some cells are marked as not reported, which limits like-for-like comparison; a model looks stronger on the rows where its rivals have no score. Second, harness, prompt format and sampling settings move agentic scores by several points, so third-party evaluation after the weights ship will be the more reliable read.
MoE economics: why 23B active matters
Beam's design follows the template that most large open models now share. In a mixture-of-experts model, every layer contains many expert sub-networks and a router selects a few of them for each token. Total parameters determine how much the model can store; active parameters determine how much arithmetic each token costs. Beam activates 23 billion of 501 billion parameters, about 4.6%.
That ratio sits close to its peers. Thinking Machines' Inkling has 975 billion total and 41 billion active parameters, about 4.2%. Moonshot's Kimi K3 has 2.8 trillion parameters and activates 16 of 896 experts per token. The difference is scale: Beam is about half Inkling's size and roughly a sixth of Kimi K3's.

The economics break into two parts, and they pull in different directions:
- Compute per token scales with active parameters. A common rule of thumb is that a forward pass costs about two floating point operations per active parameter per token. On that approximation, a 23B-active model does a little over half the per-token arithmetic of a 41B-active one. This is the basis of Reflection's 3 to 4x inference compute claim against GLM-5.2, though the company has not published the exact method.
- Memory scales with total parameters. All 501 billion weights have to sit in accelerator memory to serve the model, because the router can call any expert. As derived arithmetic, that is about 501 GB at 8 bits per weight or roughly 1 TB at 16 bits, before the memory needed for long contexts. Beam is not a model most teams will run on a single small server, even though each token is relatively cheap to compute.
The practical result is that MoE models reward batching. A provider serving many users at once can keep every expert busy and benefit fully from the low active count, while a single team running Beam in-house pays the full memory cost for a lighter workload. Reflection reports near-uniform expert utilization during training, with the busiest expert at 1.04 times the average load, which matters because uneven routing wastes exactly the hardware that sparsity is meant to save.
The US and China open-weight race
Chinese labs have set much of the pace in open-weight models, and Reflection's own comparison table is effectively a list of them: GLM from Z.ai, Kimi from Moonshot, Qwen from Alibaba and DeepSeek.
NVIDIA
Z.ai
Moonshot AI
Qwen
DeepSeekThe US side has been filling in during 2026. Inkling from Thinking Machines, Nvidia's Nemotron line and now Beam give American enterprises several domestic open-weight options, and Nvidia has been explicit about the policy stakes, as covered in our post on its open-weights letter on American AI leadership. Startup Fortune describes Beam as positioned to compete directly with DeepSeek, Qwen and GLM-5.2, while Crypto Briefing reports that Reflection is aiming at government agencies and regulated industries where model origin matters for compliance. Reflection itself names enterprises, developers and sovereign governments as targets, and frames its offering as customized "AI factories" that connect an organization's own data to locally controlled models.
On capability, the table above suggests the leading Chinese models remain ahead on several of the hardest tests, and Reflection does not claim otherwise. The competitive argument rests on three other factors: a permissive Apache 2.0 license, US provenance for buyers who need it, and a lower serving cost per token. Whether that combination is enough depends on how much a given buyer values origin over the last few points of benchmark accuracy.
On funding, coverage differs on the details. Startup Fortune reports that Reflection has raised about $4.7 billion at a $25 billion valuation, while Crypto Briefing reports more than $4 billion raised at a $25 billion pre-money valuation. Both agree that Nvidia, Sequoia and Lightspeed are among its backers.
What to watch
- The weights drop. Reflection says the Apache 2.0 release comes later in October. The model card and technical report should settle the undisclosed details, including expert count and attention configuration.
- Independent benchmarks. Third-party runs on SWE-Bench Verified, Terminal Bench and HLE will show whether the self-reported numbers hold under neutral harnesses.
- The inference compute claim. A published method for the 3 to 4x comparison with GLM-5.2 would let buyers translate it into serving cost.
- Provider availability. Reflection mentions distribution partners at launch. Hosted pricing per million tokens is where MoE efficiency becomes visible to users.
- Adoption in regulated sectors. Named government or enterprise deployments would test the provenance argument in practice.
For teams that already compare several models, Beam is another entry in an increasingly crowded open-weight field, and the most useful test will be their own tasks rather than any leaderboard. A model-agnostic workspace such as Metir is built for that kind of side-by-side comparison as new models become available.
FAQ
What is Reflection AI Beam? Beam is an open-weight sparse mixture-of-experts language model from Reflection AI, with 501 billion total parameters and 23 billion active per token, built for coding, reasoning and agentic tasks.
Is Beam open source? Reflection says Beam's weights will be released under the Apache 2.0 license later in October 2026, alongside a technical report and model card. The announcement does not say that training data or training code will be released, so "open weights" is the more precise term.
How does Beam compare with DeepSeek, Qwen and Kimi? On Reflection's own numbers, Beam trails Kimi K3, Qwen 3.8-Max and DeepSeek V4.1 on Terminal Bench v2.1 and HLE, and is close to GLM 5.2. It leads US open peers Inkling and Nemotron 3 Ultra on several reasoning tests.
Why does the 23B active parameter count matter? Per-token compute scales with active parameters, so a sparse model can be cheaper to run than its total size implies. Memory still scales with total parameters, so serving Beam requires hardware that can hold all 501 billion weights.
When can developers use Beam? An early access waitlist is open now, with the full weight release and partner availability planned for later in October 2026, according to Reflection.
Sources:
- Introducing Beam: Reflection's 501B open-weight model | Reflection
- Reflection AI unveils Beam, a 501-billion-parameter open model to rival China | Startup Fortune
- Reflection AI unveils Beam, a 501B-parameter open model due this month | Crypto Briefing
- Introducing Inkling | Thinking Machines
- Kimi K3 | Moonshot AI
Image credits
Header image: an ASUS ESC AI POD rack built on Nvidia GB200 NVL72, the predecessor of the GB300 NVL72 systems Reflection says it trained Beam on, at COMPUTEX 2024. Frame from a video by 极客湾Geekerwan via Wikimedia Commons, licensed under CC BY 3.0.
In-body image: Nvidia HGX B200 eight-GPU board, by Pokiiri via Wikimedia Commons, licensed under CC BY-SA 4.0.
