Skip to content
reflectionbeam

Head to head

Beam vs Inkling

13:2

Beam wins 13 of 15 shared benchmarks

Verdict

Inkling, from Thinking Machines Lab, is the other big US open-weight model of 2026 and already ships under Apache 2.0. Beam beats it on 13 of 15 shared benchmarks, with large margins on SWE-Bench Pro, Terminal-Bench and CritPt. Inkling is marginally ahead on IFBench and the AA Omniscience index. Reflection notes Inkling's RL used 30 million rollouts against more than 100 million for Beam.

Benchmark by benchmark

BenchmarkBeamInklingDifference
SWE-Bench Pro v2-Hard77.256.9+20.3
SWE-Bench Pro v165.554.3+11.2
Terminal-Bench v2.180.163.8+16.3
SWE-Bench Verified80.977.6+3.3
AIME 202697.897.1+0.7
Humanity's Last Exam (no tools)36.229.7+6.5
SciCode49.746.1+3.6
CritPt (AA)16.35.4+10.9
GPQA Diamond90.587.2+3.3
MCP Atlas78.776.0+2.7
τ³ banking38.025.0+13.0
BrowseComp (context mgmt)77.477.1+0.3
AA-LCR (long context)79.377.3+2.0
IFBench79.779.8-0.1
AA Omniscience index (public split)13.014.2-1.2

All scores are from Reflection's announcement (October 8, 2026 revision). Scores for Inkling are as cited by Reflection from Artificial Analysis and DataCurve.

The two models

BeamInkling
DeveloperReflection AIThinking Machines Lab
Based inUnited StatesUnited States
Parameters (total / active)501B / 23B975B / 41B
WeightsApache 2.0, promised for later in October 2026Apache 2.0
ReleasedOctober 5, 2026July 15, 2026
How to accessWaitlisted beta API; weights comingHugging Face, Tinker API

Other comparisons