Skip to content
reflectionbeam

Head to head

Beam vs Kimi K3

1:13

Beam wins 1 of 14 shared benchmarks

Verdict

Reflection itself names Kimi K3 as the model that remains ahead on raw capability. Kimi K3 wins 13 of 14 shared benchmarks, often by 10 points or more on agentic coding and search. Beam's only win is τ³ banking. The trade-off is size: Kimi K3 has 2.8 trillion parameters with 104 billion active per token, about 4.5 times Beam's 23 billion, and its weights take roughly 1.4 TB even in 4-bit form.

Benchmark by benchmark

BenchmarkBeamKimi K3Difference
DeepSWE v1.144.468.0-23.6
SWE-Bench Pro v2-Hard77.288.2-11.0
Terminal-Bench v2.180.188.3-8.2
SWE Atlas Codebase QnA34.668.0-33.4
Humanity's Last Exam (no tools)36.246.9-10.7
SciCode49.758.7-9.0
CritPt (AA)16.323.4-7.1
GPQA Diamond90.593.5-3.0
AutomationBench (public)37.046.7-9.7
MCP Atlas78.782.3-3.6
τ³ banking38.037.1+0.9
BrowseComp (context mgmt)77.491.2-13.8
DeepSearchQA (context mgmt)80.195.0-14.9
AA-LCR (long context)79.388.7-9.4

All scores are from Reflection's announcement (October 8, 2026 revision). Scores for Kimi K3 are as cited by Reflection from Artificial Analysis and DataCurve.

The two models

BeamKimi K3
DeveloperReflection AIMoonshot AI
Based inUnited StatesChina
Parameters (total / active)501B / 23B2.8T / 104B
WeightsApache 2.0, promised for later in October 2026Kimi K3 License (custom)
ReleasedOctober 5, 2026July 16, 2026
How to accessWaitlisted beta API; weights comingHugging Face, Moonshot API, OpenRouter

Other comparisons