Skip to content
reflectionbeam

Head to head

Beam vs Qwen 3.8 Max

0:13

Beam wins 0 of 13 shared benchmarks

Verdict

Qwen 3.8 Max leads on every one of the 13 shared benchmarks, though several gaps are small (GPQA 92.6 vs 90.5, AA-LCR 80.3 vs 79.3, LongBench 66.3 vs 65.5). Qwen 3.8 Max is a 2.4-trillion-parameter model with 95 billion active per token, roughly four times Beam's active size, which is the compute gap Reflection points to.

Benchmark by benchmark

BenchmarkBeamQwen 3.8 MaxDifference
DeepSWE v1.144.451.0-6.6
SWE-Bench Pro v165.567.7-2.2
Terminal-Bench v2.180.186.6-6.5
Humanity's Last Exam (no tools)36.243.6-7.4
SciCode49.752.1-2.4
CritPt (AA)16.320.0-3.7
GPQA Diamond90.592.6-2.1
AutomationBench (public)37.039.8-2.8
MCP Atlas78.784.5-5.8
τ³ banking38.055.2-17.2
AA-LCR (long context)79.380.3-1.0
LongBench v265.566.3-0.8
IFBench79.782.8-3.1

All scores are from Reflection's announcement (October 8, 2026 revision). Scores for Qwen 3.8 Max are as cited by Reflection from Artificial Analysis and DataCurve.

The two models

BeamQwen 3.8 Max
DeveloperReflection AIAlibaba (Qwen)
Based inUnited StatesChina
Parameters (total / active)501B / 23B2.4T / 95B
WeightsApache 2.0, promised for later in October 2026Qwen3.8-Max License (custom)
ReleasedOctober 5, 2026August 3, 2026
How to accessWaitlisted beta API; weights comingHugging Face, Alibaba Cloud Model Studio, OpenRouter

Other comparisons