Reflection's 501 billion parameter Beam is open weight, with trained weights due later this month, and the company claims parity with rival open weight model GLM 5.2 at three to four times lower inference compute.
Reflection's new open-weight model, Beam, is being pitched as the Western open-weight frontier's efficiency play. The numbers under that pitch are the ones worth reading first.
The headline figure is 501 billion parameters. The number that does the work is 23 billion. Beam is a sparse Mixture-of-Experts model: 501 billion total parameters spread across many specialized expert sub-networks, with about 23 billion active on any given input. Sparse MoE routes each token to a subset of experts, so the model can hold a large knowledge base in memory while only paying the compute bill on a fraction of it. In theory, that means more capability per dollar of inference. Reflection's claim is that the theory holds in practice.
"Open-weight" is the next term to pin down. The model is not open source in the way a typical open-source software release is. Reflection is committing to ship the trained weights, technical report, model card, and developer artifacts later this month, but the training data and full pipeline are not part of that release. The package is a downloadable set of numbers that anyone can run, modify, and fine-tune, plus documentation, but not a reproducible recipe. That is the standard open-weight contract in 2026, and it is the one Beam is signing.
Pretraining was on 23.8 trillion tokens drawn from the web and proprietary licensed datasets. The more interesting part of the announcement for builders is the reinforcement-learning phase: more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks of training. That RL scale is large enough to shape the resulting model, and Reflection is treating it as a sustained engineering investment rather than a fine-tuning afterthought. The cluster footprint is the kind of operational fact a builder can plan around, even if the benchmark numbers cannot be.
On benchmarks, Reflection's claims are company-stated and unverified externally. The company says Beam scores comparable to GLM-5.2 on advanced reasoning while using three to four times less inference compute, is competitive with GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks, and trails Kimi K3 on raw capability. The names matter here: GLM 5.2, Qwen 3.8-Max, and Kimi K3 are non-standard public benchmark identifiers, and the community has already flagged a small slip. Reflection described a generalization test as a "puzzle" that was "a few days old"; the Hacker News thread on the post identified the same test as roughly 14 months old. The benchmarks themselves are real; the framing has had a marketing-versus-research wobble.
The pro-building read is that frontier-class coding and reasoning are getting cheaper to access, not that any one lab has cracked the problem. A 23-billion-active model that holds its own on agentic workloads is a more useful tool for a startup than a trillion-parameter model it cannot afford to serve. Reflection is positioning Beam as a "workhorse" for enterprise coding and agentic deployments, and the explicit claim is that the open-weight frontier is now competitive with the closed frontier on capability per token. The agency-expanding version of that claim is real: more teams can run frontier-class coding models without frontier-class compute bills.
What the announcement does not yet answer is the one question worth holding onto. Until the weights land and outside labs reproduce the inference-cost claims, "comparable to GLM-5.2 at 3-4x lower compute" is a Reflection benchmark, not a market one. Reflection says final red-teaming and evaluations are ongoing, and the technical report, model card, and developer artifacts are due later this month. The bet is that the numbers survive contact with reality. The open-weight part of that bet is the part that arrives in a download link, and the verification part is the part that arrives in someone else's reproduction.