Head to head
Beam vs Qwen 3.8 Max
0:13
Beam wins 0 of 13 shared benchmarks
Verdict
Qwen 3.8 Max leads on every one of the 13 shared benchmarks, though several gaps are small (GPQA 92.6 vs 90.5, AA-LCR 80.3 vs 79.3, LongBench 66.3 vs 65.5). Qwen 3.8 Max is a 2.4-trillion-parameter model with 95 billion active per token, roughly four times Beam's active size, which is the compute gap Reflection points to.
Benchmark by benchmark
| Benchmark | Beam | Qwen 3.8 Max | Difference | |
|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 51.0 | -6.6 | |
| SWE-Bench Pro v1 | 65.5 | 67.7 | -2.2 | |
| Terminal-Bench v2.1 | 80.1 | 86.6 | -6.5 | |
| Humanity's Last Exam (no tools) | 36.2 | 43.6 | -7.4 | |
| SciCode | 49.7 | 52.1 | -2.4 | |
| CritPt (AA) | 16.3 | 20.0 | -3.7 | |
| GPQA Diamond | 90.5 | 92.6 | -2.1 | |
| AutomationBench (public) | 37.0 | 39.8 | -2.8 | |
| MCP Atlas | 78.7 | 84.5 | -5.8 | |
| τ³ banking | 38.0 | 55.2 | -17.2 | |
| AA-LCR (long context) | 79.3 | 80.3 | -1.0 | |
| LongBench v2 | 65.5 | 66.3 | -0.8 | |
| IFBench | 79.7 | 82.8 | -3.1 |
All scores are from Reflection's announcement (October 8, 2026 revision). Scores for Qwen 3.8 Max are as cited by Reflection from Artificial Analysis and DataCurve.
The two models
| Beam | Qwen 3.8 Max | |
|---|---|---|
| Developer | Reflection AI | Alibaba (Qwen) |
| Based in | United States | China |
| Parameters (total / active) | 501B / 23B | 2.4T / 95B |
| Weights | Apache 2.0, promised for later in October 2026 | Qwen3.8-Max License (custom) |
| Released | October 5, 2026 | August 3, 2026 |
| How to access | Waitlisted beta API; weights coming | Hugging Face, Alibaba Cloud Model Studio, OpenRouter |
Other comparisons
Z.ai (Zhipu AI)
Beam vs GLM 5.2
Beam 8–513 shared benchmarks
Moonshot AI
Beam vs Kimi K3
Beam 1–1314 shared benchmarks
DeepSeek
Beam vs DeepSeek V4.1 Flash
Beam 2–79 shared benchmarks
Z.ai (Zhipu AI)
Beam vs GLM 5.3
Beam 0–1212 shared benchmarks
NVIDIA
Beam vs Nemotron 3 Ultra
Beam 13–115 shared benchmarks
Thinking Machines Lab
Beam vs Inkling
Beam 13–215 shared benchmarks