Benchmarks
Beam benchmark results
21 benchmarks across agentic coding, reasoning, tool use and general skills, exactly as Reflection published them, next to seven other open models.
These are vendor-reported numbers from Reflection's announcement (table updated October 8, 2026). Scores for other models come from Artificial Analysis and DataCurve, as cited by Reflection. Independent reproductions will follow once the weights are public.
| Benchmark | Beam | Inkling | Nemotron 3 Ultra | GLM 5.2 | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | — | — | 44.0 | 61.0 | 68.0 | 51.0 | 74.2 |
| SWE-Bench Pro v2-Hard | 77.2 | 56.9 | — | — | 84.3 | 88.2 | — | — |
| SWE-Bench Pro v1 | 65.5 | 54.3 | 46.4 | 62.1 | — | — | 67.7 | — |
| Terminal-Bench v2.1 | 80.1 | 63.8 | 56.4 | 81.0 | 88.2 | 88.3 | 86.6 | 90.6 |
| SWE Atlas Codebase QnA | 34.6 | — | — | — | 61.0 | 68.0 | — | — |
| SWE-Bench Multilingual | 78.0 | — | 67.7 | — | — | — | — | — |
| SWE-Bench Verified | 80.9 | 77.6 | 70.7 | — | — | — | — | — |
| AIME 2026 | 97.8 | 97.1 | — | 99.2 | — | — | — | — |
| Humanity's Last Exam (no tools) | 36.2 | 29.7 | 26.7 | 40.5 | 42.3 | 46.9 | 43.6 | 39.1 |
| SciCode | 49.7 | 46.1 | 44.6 | — | 59.0 | 58.7 | 52.1 | 52.0 |
| CritPt (AA) | 16.3 | 5.4 | 3.1 | 20.9 | 19.1 | 23.4 | 20.0 | 14.3 |
| GPQA Diamond | 90.5 | 87.2 | 87.0 | 91.2 | 91.7 | 93.5 | 92.6 | 90.9 |
| AutomationBench (public) | 37.0 | — | — | 26.2 | 48.2 | 46.7 | 39.8 | 54.8 |
| MCP Atlas | 78.7 | 76.0 | 63.1 | 77.8 | 84.2 | 82.3 | 84.5 | — |
| τ³ banking | 38.0 | 25.0 | 22.6 | 37.1 | — | 37.1 | 55.2 | — |
| BrowseComp (context mgmt) | 77.4 | 77.1 | 44.4 | — | — | 91.2 | — | — |
| DeepSearchQA (context mgmt) | 80.1 | — | — | — | — | 95.0 | — | — |
| AA-LCR (long context) | 79.3 | 77.3 | 79.3 | 78.3 | 79.7 | 88.7 | 80.3 | 84.0 |
| LongBench v2 | 65.5 | — | 61.9 | 64.0 | — | — | 66.3 | — |
| IFBench | 79.7 | 79.8 | 81.7 | 73.3 | — | — | 82.8 | — |
| AA Omniscience index (public split) | 13.0 | 14.2 | 8.6 | — | 22.6 | — | — | 6.6 |
00.0 Best reported score in the row— Not reported
What the numbers say
Beam is strongest on software engineering. It posts 80.9 on SWE-Bench Verified (ahead of Inkling at 77.6) and 78.0 on SWE-Bench Multilingual (ahead of Nemotron 3 Ultra at 67.7), benchmarks the larger Chinese models did not report, and it edges GLM 5.2 on SWE-Bench Pro v1 (65.5 vs 62.1) and on MCP Atlas (78.7 vs 77.8).
On pure reasoning it sits a little behind the leaders: 90.5 on GPQA Diamond against 93.5 for Kimi K3, and 36.2 on Humanity's Last Exam without tools against 46.9.
Reflection's argument is not that Beam wins everything, but that it gets close while spending far less compute per answer. Kimi K3, GLM 5.3 and Qwen 3.8 Max still lead on most rows where they reported results.
What each benchmark measures
Agentic coding & terminal
- DeepSWE v1.1
- Long-horizon software engineering tasks solved end to end by an agent.
- SWE-Bench Pro v2-Hard
- The hard split of SWE-Bench Pro: difficult real-world repository issues.
- SWE-Bench Pro v1
- Real GitHub issues from professional codebases, harder and less contaminated than classic SWE-Bench.
- Terminal-Bench v2.1
- Completing tasks inside a real Linux terminal: shell, tools, builds and debugging.
- SWE Atlas Codebase QnA
- Answering questions about large, unfamiliar codebases.
- SWE-Bench Multilingual
- SWE-Bench style issue fixing across many programming languages.
- SWE-Bench Verified
- Human-verified subset of SWE-Bench: fix the issue so the hidden tests pass.
Reasoning
- AIME 2026
- Problems from the 2026 American Invitational Mathematics Examination.
- Humanity's Last Exam (no tools)
- Humanity's Last Exam: expert-level questions across many fields, answered without tools.
- SciCode
- Writing code to solve research-level science problems.
- CritPt (AA)
- Research-level physics reasoning problems, as run by Artificial Analysis.
- GPQA Diamond
- Graduate-level biology, chemistry and physics questions designed to be hard to look up.
Tool calling & search
- AutomationBench (public)
- Automating multi-step business workflows across apps.
- MCP Atlas
- Using tools exposed over the Model Context Protocol (MCP) correctly.
- τ³ banking
- Customer-service style conversations in a banking domain with tools and policies.
- BrowseComp (context mgmt)
- Finding hard-to-locate facts by browsing the web, with context management.
- DeepSearchQA (context mgmt)
- Multi-step research questions that require searching and combining sources.
General capabilities
- AA-LCR (long context)
- Reasoning over very long documents (Artificial Analysis long-context reasoning).
- LongBench v2
- Understanding and answering questions over long inputs.
- IFBench
- Following precise, verifiable formatting and content instructions.
- AA Omniscience index (public split)
- Factual knowledge with a penalty for confident wrong answers (hallucinations).