Models·3 min read
By BitsMindsSource: Reflection AI

Reflection’s Beam: A 501B Open Model Built to Be Frugal

Reflection AI has previewed Beam, its first open-weight model: a 501-billion-parameter mixture-of-experts with 23 billion active per token, aimed at coding and agents. Weights land later in October under Apache 2.0. Its own table shows it trailing the newest Chinese open models, so the pitch is efficiency.

Reflection Beam: a large core, a narrow active path An original monochrome optical sculpture on a dark reflective laboratory bench. A narrow white beam passes through a hollow silver ring and branches toward a few illuminated cartridges within a large graphite core. Most cartridges remain dark; their paths recombine into a beam that reaches a solid circular receiver. The hollow and solid circles echo Reflection's mark. The labels read BEAM, 501B total parameters, 23B active per token, PREVIEW and WEIGHTS PLANNED — OCT. The illustration represents sparse inference in the model, not physical hardware or its exact expert layout. BitsMinds original editorial vector artwork for reflection-ai-beam-501b-open-weight-model. Facts verified on 6 October 2026 at https://reflection.ai/blog/introducing-beam: 501B total parameters, 23B active, weights planned later in October under Apache 2.0. The module layout is symbolic. No performance ranking or measured-cost claim is depicted. Self-contained SVG with the project Reflection wordmark. BEAM 501B TOTAL PARAMETERS 23B ACTIVE PER TOKEN PREVIEW WEIGHTS PLANNED · OCT BITSMINDS.COM
Share:

Reflection AI, the New York lab that set out to build an American answer to DeepSeek, has shown its first model. Beam, announced on Monday, is a sparse mixture-of-experts with 501 billion total parameters, of which 23 billion are active for any given token, trained with a particular eye on coding and agentic work. It is text-only, has a 256K context window in use (midtraining stretched it to 1M), and comes with a reasoning-effort setting that trades answer length for accuracy.

It is a preview rather than a release. Beam is still going through final red-teaming, and Reflection is admitting a waitlist of early users. The weights, a technical report and a model card are promised for later in October under an Apache 2.0 licence, along with tooling for running, evaluating and fine-tuning the model and a set of distribution partners.

Where it lands

Reflection is unusually frank about its position. The company says Beam “advances the Western open-weight frontier” and is competitive with GLM 5.2 while approaching Qwen 3.8-Max on coding and agentic tasks, but concedes that models like Kimi K3 “remain ahead on raw capability”. Its own table bears that out. Beam scores 80.1 on Terminal Bench 2.1, 80.9 on SWE-bench Verified and 90.5 on GPQA Diamond, yet on every shared benchmark below it sits behind Zhipu’s GLM 5.3 and Moonshot’s Kimi K3, and DeepSeek’s V4.1 Flash beats it on Terminal Bench too.

Beam against the newest open modelsScore, % — higher is better · Reflection, 5 Oct 2026 · n/r = not reportedReflection BeamGLM 5.3Kimi K3DeepSeek V4.1 Flash02040608010080.188.288.390.6Terminal Bench2.177.284.388.2n/rSWE Bench Prov2-Hard78.784.282.3n/rMCP Atlas90.591.793.590.9GPQA Diamond36.242.346.939.1HLE (no tools)
Reflection’s own comparison table puts Beam behind GLM 5.3 and Kimi K3 on every shared benchmark shown here. Data: Reflection AI, citing Artificial Analysis and DataCurve for rival scores.

The argument Reflection makes instead is cost. On advanced reasoning benchmarks it says Beam matches GLM 5.2 while using three to four times less inference compute, estimated from active parameters multiplied by tokens generated, and that the gap widens against two-trillion-parameter models such as Qwen 3.8-Max. The company is careful to call this an approximation that leaves out prefill and serving overhead, so the real bill will only be known once third parties run the weights.

How it was trained

Pretraining covered 23.8 trillion tokens and finished in under four weeks on 6,144 Nvidia GB300 GPUs, with 92.3% goodput by the end. Reinforcement learning was the bigger bet: 10,500 GB300s ran for four weeks, generating more than 100 million rollouts across roughly one million coding, agentic and STEM environments and some 1.3 billion sandboxes. Reflection calls it one of the largest RL runs by any open lab and says scores were still climbing when it stopped. Safety and alignment came from a separately trained teacher model whose behaviour was distilled into Beam alongside the RL teacher.

The demos lean on agentic range. In one, Beam builds a live New York subway map from the MTA’s public feeds, finding the documentation, the line geometry and the authentication rules itself. In another it writes a p5.js game about an astronaut dodging asteroids in free fall, and in a third it prepares a notebook that fine-tunes Google’s smallest Gemma 4 model for text-to-SQL.

Animated map of New York City with subway lines in their route colours and train markers moving between stations, built by Reflection's Beam model from MTA data
Beam’s live NYC subway map, one of four demos Reflection published. Source: Reflection AI.

Reflection raised $2 billion in October 2025, with Nvidia among the backers, on the promise that the US needed a frontier lab that publishes its weights. Beam is the first evidence of what that money bought: a model that does not top the open leaderboard, but that Reflection says is the first in a series, with the next one already training. Whether the efficiency claims hold up will be testable within weeks, once the weights are public.

Want AI news before everyone else?

The morning's most important AI stories, straight to your inbox. No fluff.

Related Articles