- Beam is a 501B-parameter sparse mixture-of-experts model with 23B active parameters, a 1 million token context window and an Apache 2.0 license, built for coding and agent work.
- On Terminal Bench v2.1, Beam scores 80.1 against 90.6 for DeepSeek V4.1 Flash and 88.2 for Z.ai's GLM 5.3, according to Reflection's own published results.
- Reflection's pitch is efficiency: it says Beam matches GLM 5.2 on reasoning while using three to four times less inference compute. Full weights arrive later in October.
Reflection AI, the New York lab founded by former Google DeepMind researchers and backed by Nvidia, released Beam on October 5, its first open-weight model and the largest American attempt yet to compete with Chinese labs on open AI. Beam has 501 billion parameters, of which only 23 billion activate for any given token, and ships under the permissive Apache 2.0 license. Reflection's own benchmark table, published with the launch, places it behind DeepSeek, Moonshot, Z.ai and Alibaba on most coding tests.
Reflection trained Beam on 23.8 trillion tokens in under two months of compute
Beam is a sparse mixture-of-experts model, an architecture that routes each token through a small subset of specialist sub-networks, which is why a 501B model can run with the serving cost of a much smaller one. Reflection says it pretrained the base model on 23.8 trillion tokens of web and licensed data over four weeks, then ran a second four-week reinforcement learning phase on 10,500 Nvidia GB300 GPUs that generated more than 100 million rollouts and about 1.3 billion sandboxed code evaluations.
| Total parameters | 501 billion |
| Active parameters per token | 23 billion |
| Context window | 1 million tokens |
| Pretraining data | 23.8 trillion tokens over four weeks on 6,144 Nvidia GB300 GPUs |
| Reinforcement learning | Four weeks on 10,500 GB300 GPUs, about 1.3 billion sandbox evaluations |
| License | Apache 2.0, weights due later in October 2026 |
“We are introducing Beam, Reflection's first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads.”
Reflection AI, Introducing Beam, October 5, 2026
The company also built a controllable length penalty into training, which rewards correct answers that use fewer tokens and lets developers trade reasoning depth against compute cost. Early access is open now through Reflection's platform, with weights, a technical report and fine-tuning tools due later this month.
Reflection's own scores place Beam a step behind China's open leaders
Reflection published Beam's results alongside eight rival open-weight models, and the comparison is unusually candid for a launch. Beam beats Nvidia's Nemotron 3 Ultra by more than 20 points on Terminal Bench and edges GLM 5.2 on DeepSWE, yet it sits roughly ten points under DeepSeek V4.1 Flash and eight under GLM 5.3 and Kimi K3 on the same terminal test. On DeepSWE v1.1, DeepSeek's 74.2 nearly doubles Beam's 44.4.
That gap explains the efficiency framing. Reflection is selling Beam as the open model an American bank, agency or defense contractor can run cheaply on its own hardware without depending on Chinese weights, a market that has grown since Washington began scrutinizing Chinese labs over distillation of US models. Reflection was valued at $25 billion in April and agreed in June to rent up to $6.3 billion of Nvidia GB300 capacity at SpaceX's Colossus 2 site, so the company has the compute to keep closing the distance.
Beam makes the US case for open AI with a model that is good enough to deploy and still behind the Chinese frontier on the tests developers check first. When the weights land later this month, independent benchmarks will decide whether the three-to-four-times efficiency claim is enough to win enterprise buyers who have so far defaulted to DeepSeek and Qwen.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.