Head to head
Beam vs GLM 5.2
8:5
Beam wins 8 of 13 shared benchmarks
Verdict
This is the matchup Reflection chose for its headline claim, and Beam holds up: it wins 8 of the 13 shared benchmarks, including SWE-Bench Pro v1, MCP Atlas, AutomationBench and IFBench, while GLM 5.2 is ahead on hard reasoning (AIME, HLE, GPQA, CritPt) and narrowly on Terminal-Bench. Reflection says Beam reaches comparable reasoning scores with three to four times less inference compute.
Benchmark by benchmark
| Benchmark | Beam | GLM 5.2 | Difference | |
|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | +0.4 | |
| SWE-Bench Pro v1 | 65.5 | 62.1 | +3.4 | |
| Terminal-Bench v2.1 | 80.1 | 81.0 | -0.9 | |
| AIME 2026 | 97.8 | 99.2 | -1.4 | |
| Humanity's Last Exam (no tools) | 36.2 | 40.5 | -4.3 | |
| CritPt (AA) | 16.3 | 20.9 | -4.6 | |
| GPQA Diamond | 90.5 | 91.2 | -0.7 | |
| AutomationBench (public) | 37.0 | 26.2 | +10.8 | |
| MCP Atlas | 78.7 | 77.8 | +0.9 | |
| τ³ banking | 38.0 | 37.1 | +0.9 | |
| AA-LCR (long context) | 79.3 | 78.3 | +1.0 | |
| LongBench v2 | 65.5 | 64.0 | +1.5 | |
| IFBench | 79.7 | 73.3 | +6.4 |
All scores are from Reflection's announcement (October 8, 2026 revision). Scores for GLM 5.2 are as cited by Reflection from Artificial Analysis and DataCurve.
The two models
| Beam | GLM 5.2 | |
|---|---|---|
| Developer | Reflection AI | Z.ai (Zhipu AI) |
| Based in | United States | China |
| Parameters (total / active) | 501B / 23B | 753B / 40B |
| Weights | Apache 2.0, promised for later in October 2026 | MIT |
| Released | October 5, 2026 | June 16, 2026 |
| How to access | Waitlisted beta API; weights coming | Hugging Face, Z.ai API, OpenRouter |
Other comparisons
Moonshot AI
Beam vs Kimi K3
Beam 1–1314 shared benchmarks
Alibaba (Qwen)
Beam vs Qwen 3.8 Max
Beam 0–1313 shared benchmarks
DeepSeek
Beam vs DeepSeek V4.1 Flash
Beam 2–79 shared benchmarks
Z.ai (Zhipu AI)
Beam vs GLM 5.3
Beam 0–1212 shared benchmarks
NVIDIA
Beam vs Nemotron 3 Ultra
Beam 13–115 shared benchmarks
Thinking Machines Lab
Beam vs Inkling
Beam 13–215 shared benchmarks