Skip to content
reflectionbeam

Model card

Beam specifications

Everything Reflection AI has published about how Beam is built and trained. Numbers come from the official announcement and developer docs; estimates are marked.

At a glance

DeveloperReflection AI (New York)
Announced2026-10-05
Model typeSparse Mixture-of-Experts language model
Total parameters501B
Active parameters per token23B
Layers52
AttentionInterleaved local and global attention
Context window (API beta)256K tokens
Context length reached in midtraining1M tokens
Max output (API)128K tokens
Knowledge cutoff2026-06-30
ModalityText in, text out
ReasoningAlways on, with a reasoning effort setting
API featuresTool calling, structured outputs, streaming
Weights licenseApache 2.0 (weights promised for later in October 2026)
API model IDBeam-501B-A23B

Why 501B total but 23B active?

In a Mixture-of-Experts model, each layer holds many small “expert” networks and a router picks a few of them for every token. Beam stores 501 billion parameters, but each token only passes through about 23 billion of them.

The result: compute per token is close to a 23B dense model, which is what makes Beam cheaper to serve, while total capacity stays large. The catch is memory. All 501B parameters still have to sit in GPU memory to serve the model.

Reflection says its routing is unusually balanced: at the end of pretraining the busiest expert carried only 1.04× the average load, so all experts get used.

501B100%
23B4.6%

How Beam was trained

Pretraining

23.8 trillion tokens from the web, public sources and licensed datasets, with heavy emphasis on code, technical writing and STEM. About 95% of raw web tokens were filtered out. Run on 6,144 NVIDIA GB300 NVL72 GPUs in under four weeks, with 92.3% goodput near the end.

Midtraining

A stage designed to prepare the model for reinforcement learning: long reasoning-rich documents, code repositories and long-horizon tasks. This is where Beam's effective context was extended to 1 million tokens.

Reinforcement learning

More than 100 million rollouts on 10,500 GB300 GPUs over four weeks, across almost one million coding, agentic and STEM environments and about 1.3 billion sandboxes. Reflection calls it one of the largest RL runs by any open lab, and says scores were still rising when it ended.

Safety and alignment

A separate safety-and-alignment model was trained from the same base and merged with the RL model through multi-teacher on-policy distillation. Reflection says it will publish safety evaluations in the technical report and open-source its internal safety evals.

Reasoning effort and token efficiency

Beam was trained with a length penalty that rewards correct answers and discourages unnecessary tokens. In the API you choose a reasoning effort: lower settings answer faster with fewer tokens, higher settings think longer on hard problems.

How much memory does Beam need?

Reflection has not published hardware requirements yet. These are our rough estimates for the weights alone (parameters × bytes per parameter). Real deployments need extra memory for the KV cache and activations.

PrecisionWeights onlyExample hardware
BF16 / FP16≈ 1,002 GB8× 192 GB GPUs (about 1.5 TB) or a multi-node cluster
FP8≈ 501 GB8× 80 GB GPUs is tight; 8× 141 GB or larger is comfortable
INT4 / 4-bit≈ 251 GB4× 80 GB GPUs at minimum, more for long contexts

Estimates, not official requirements. We will replace them with Reflection's numbers when the model card ships.