Skip to content
reflectionbeam

Head to head

Beam vs GLM 5.2

8:5

Beam wins 8 of 13 shared benchmarks

Verdict

This is the matchup Reflection chose for its headline claim, and Beam holds up: it wins 8 of the 13 shared benchmarks, including SWE-Bench Pro v1, MCP Atlas, AutomationBench and IFBench, while GLM 5.2 is ahead on hard reasoning (AIME, HLE, GPQA, CritPt) and narrowly on Terminal-Bench. Reflection says Beam reaches comparable reasoning scores with three to four times less inference compute.

Benchmark by benchmark

BenchmarkBeamGLM 5.2Difference
DeepSWE v1.144.444.0+0.4
SWE-Bench Pro v165.562.1+3.4
Terminal-Bench v2.180.181.0-0.9
AIME 202697.899.2-1.4
Humanity's Last Exam (no tools)36.240.5-4.3
CritPt (AA)16.320.9-4.6
GPQA Diamond90.591.2-0.7
AutomationBench (public)37.026.2+10.8
MCP Atlas78.777.8+0.9
τ³ banking38.037.1+0.9
AA-LCR (long context)79.378.3+1.0
LongBench v265.564.0+1.5
IFBench79.773.3+6.4

All scores are from Reflection's announcement (October 8, 2026 revision). Scores for GLM 5.2 are as cited by Reflection from Artificial Analysis and DataCurve.

The two models

BeamGLM 5.2
DeveloperReflection AIZ.ai (Zhipu AI)
Based inUnited StatesChina
Parameters (total / active)501B / 23B753B / 40B
WeightsApache 2.0, promised for later in October 2026MIT
ReleasedOctober 5, 2026June 16, 2026
How to accessWaitlisted beta API; weights comingHugging Face, Z.ai API, OpenRouter

Other comparisons