Skip to content
reflectionbeam

Benchmarks

Beam benchmark results

21 benchmarks across agentic coding, reasoning, tool use and general skills, exactly as Reflection published them, next to seven other open models.

These are vendor-reported numbers from Reflection's announcement (table updated October 8, 2026). Scores for other models come from Artificial Analysis and DataCurve, as cited by Reflection. Independent reproductions will follow once the weights are public.

BenchmarkBeamInklingNemotron 3 UltraGLM 5.2GLM 5.3Kimi K3Qwen 3.8 MaxDeepSeek V4.1 Flash
DeepSWE v1.144.4——44.061.068.051.074.2
SWE-Bench Pro v2-Hard77.256.9——84.388.2——
SWE-Bench Pro v165.554.346.462.1——67.7—
Terminal-Bench v2.180.163.856.481.088.288.386.690.6
SWE Atlas Codebase QnA34.6———61.068.0——
SWE-Bench Multilingual78.0—67.7—————
SWE-Bench Verified80.977.670.7—————
AIME 202697.897.1—99.2————
Humanity's Last Exam (no tools)36.229.726.740.542.346.943.639.1
SciCode49.746.144.6—59.058.752.152.0
CritPt (AA)16.35.43.120.919.123.420.014.3
GPQA Diamond90.587.287.091.291.793.592.690.9
AutomationBench (public)37.0——26.248.246.739.854.8
MCP Atlas78.776.063.177.884.282.384.5—
τ³ banking38.025.022.637.1—37.155.2—
BrowseComp (context mgmt)77.477.144.4——91.2——
DeepSearchQA (context mgmt)80.1————95.0——
AA-LCR (long context)79.377.379.378.379.788.780.384.0
LongBench v265.5—61.964.0——66.3—
IFBench79.779.881.773.3——82.8—
AA Omniscience index (public split)13.014.28.6—22.6——6.6

00.0 Best reported score in the row— Not reported

What the numbers say

Beam is strongest on software engineering. It posts 80.9 on SWE-Bench Verified (ahead of Inkling at 77.6) and 78.0 on SWE-Bench Multilingual (ahead of Nemotron 3 Ultra at 67.7), benchmarks the larger Chinese models did not report, and it edges GLM 5.2 on SWE-Bench Pro v1 (65.5 vs 62.1) and on MCP Atlas (78.7 vs 77.8).

On pure reasoning it sits a little behind the leaders: 90.5 on GPQA Diamond against 93.5 for Kimi K3, and 36.2 on Humanity's Last Exam without tools against 46.9.

Reflection's argument is not that Beam wins everything, but that it gets close while spending far less compute per answer. Kimi K3, GLM 5.3 and Qwen 3.8 Max still lead on most rows where they reported results.

What each benchmark measures

Agentic coding & terminal

DeepSWE v1.1
Long-horizon software engineering tasks solved end to end by an agent.
SWE-Bench Pro v2-Hard
The hard split of SWE-Bench Pro: difficult real-world repository issues.
SWE-Bench Pro v1
Real GitHub issues from professional codebases, harder and less contaminated than classic SWE-Bench.
Terminal-Bench v2.1
Completing tasks inside a real Linux terminal: shell, tools, builds and debugging.
SWE Atlas Codebase QnA
Answering questions about large, unfamiliar codebases.
SWE-Bench Multilingual
SWE-Bench style issue fixing across many programming languages.
SWE-Bench Verified
Human-verified subset of SWE-Bench: fix the issue so the hidden tests pass.

Reasoning

AIME 2026
Problems from the 2026 American Invitational Mathematics Examination.
Humanity's Last Exam (no tools)
Humanity's Last Exam: expert-level questions across many fields, answered without tools.
SciCode
Writing code to solve research-level science problems.
CritPt (AA)
Research-level physics reasoning problems, as run by Artificial Analysis.
GPQA Diamond
Graduate-level biology, chemistry and physics questions designed to be hard to look up.

Tool calling & search

AutomationBench (public)
Automating multi-step business workflows across apps.
MCP Atlas
Using tools exposed over the Model Context Protocol (MCP) correctly.
τ³ banking
Customer-service style conversations in a banking domain with tools and policies.
BrowseComp (context mgmt)
Finding hard-to-locate facts by browsing the web, with context management.
DeepSearchQA (context mgmt)
Multi-step research questions that require searching and combining sources.

General capabilities

AA-LCR (long context)
Reasoning over very long documents (Artificial Analysis long-context reasoning).
LongBench v2
Understanding and answering questions over long inputs.
IFBench
Following precise, verifiable formatting and content instructions.
AA Omniscience index (public split)
Factual knowledge with a penalty for confident wrong answers (hallucinations).