Head to head
Beam vs DeepSeek V4.1 Flash
2:7
Beam wins 2 of 9 shared benchmarks
Verdict
DeepSeek V4.1 Flash is ahead on 7 of 9 shared benchmarks, with big leads on DeepSWE and AutomationBench and a 10-point lead on Terminal-Bench. Beam wins on CritPt and has a better AA Omniscience index, which penalizes confident wrong answers. Efficiency does not favor Beam here: DeepSeek V4.1 Flash activates only 16 billion parameters per generated token, fewer than Beam's 23 billion, and its weights are already public under MIT.
Benchmark by benchmark
| Benchmark | Beam | DeepSeek V4.1 Flash | Difference | |
|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 74.2 | -29.8 | |
| Terminal-Bench v2.1 | 80.1 | 90.6 | -10.5 | |
| Humanity's Last Exam (no tools) | 36.2 | 39.1 | -2.9 | |
| SciCode | 49.7 | 52.0 | -2.3 | |
| CritPt (AA) | 16.3 | 14.3 | +2.0 | |
| GPQA Diamond | 90.5 | 90.9 | -0.4 | |
| AutomationBench (public) | 37.0 | 54.8 | -17.8 | |
| AA-LCR (long context) | 79.3 | 84.0 | -4.7 | |
| AA Omniscience index (public split) | 13.0 | 6.6 | +6.4 |
All scores are from Reflection's announcement (October 8, 2026 revision). Scores for DeepSeek V4.1 Flash are as cited by Reflection from Artificial Analysis and DataCurve.
The two models
| Beam | DeepSeek V4.1 Flash | |
|---|---|---|
| Developer | Reflection AI | DeepSeek |
| Based in | United States | China |
| Parameters (total / active) | 501B / 23B | 552B / 16B |
| Weights | Apache 2.0, promised for later in October 2026 | MIT |
| Released | October 5, 2026 | September 10, 2026 |
| How to access | Waitlisted beta API; weights coming | Hugging Face, DeepSeek API, OpenRouter |
Other comparisons
Z.ai (Zhipu AI)
Beam vs GLM 5.2
Beam 8–513 shared benchmarks
Moonshot AI
Beam vs Kimi K3
Beam 1–1314 shared benchmarks
Alibaba (Qwen)
Beam vs Qwen 3.8 Max
Beam 0–1313 shared benchmarks
Z.ai (Zhipu AI)
Beam vs GLM 5.3
Beam 0–1212 shared benchmarks
NVIDIA
Beam vs Nemotron 3 Ultra
Beam 13–115 shared benchmarks
Thinking Machines Lab
Beam vs Inkling
Beam 13–215 shared benchmarks