Skip to content
reflectionbeam

Inside Beam's training: 100 million RL rollouts on 10,500 GPUs

Reflection made reinforcement learning a main scaling axis, and says gains had not plateaued.

For Beam's reinforcement learning phase, Reflection ran 10,500 NVIDIA GB300 GPUs for four weeks, generating more than 100 million rollouts with contexts up to 256K tokens. Training and grading used about 1.3 billion sandboxes across almost one million coding, agentic and STEM environments.

The company says it built new algorithms to keep fully asynchronous RL stable even when learning from data generated more than a day earlier, about 107 weight versions behind the current policy.

A length penalty taught Beam to avoid unnecessary tokens. Reflection also reports that skills transferred: training on coding and terminal tasks improved browsing, and the model learned on its own to query other models and OCR services when given web access.

For comparison, Reflection says Inkling was trained on 30 million rollouts. Scores were still rising when the run ended.

Takeaway

Beam's efficiency story is as much about RL scale as about model size.

All news