Skip to content
reflectionbeam

Head to head

Beam vs DeepSeek V4.1 Flash

2:7

Beam wins 2 of 9 shared benchmarks

Verdict

DeepSeek V4.1 Flash is ahead on 7 of 9 shared benchmarks, with big leads on DeepSWE and AutomationBench and a 10-point lead on Terminal-Bench. Beam wins on CritPt and has a better AA Omniscience index, which penalizes confident wrong answers. Efficiency does not favor Beam here: DeepSeek V4.1 Flash activates only 16 billion parameters per generated token, fewer than Beam's 23 billion, and its weights are already public under MIT.

Benchmark by benchmark

BenchmarkBeamDeepSeek V4.1 FlashDifference
DeepSWE v1.144.474.2-29.8
Terminal-Bench v2.180.190.6-10.5
Humanity's Last Exam (no tools)36.239.1-2.9
SciCode49.752.0-2.3
CritPt (AA)16.314.3+2.0
GPQA Diamond90.590.9-0.4
AutomationBench (public)37.054.8-17.8
AA-LCR (long context)79.384.0-4.7
AA Omniscience index (public split)13.06.6+6.4

All scores are from Reflection's announcement (October 8, 2026 revision). Scores for DeepSeek V4.1 Flash are as cited by Reflection from Artificial Analysis and DataCurve.

The two models

BeamDeepSeek V4.1 Flash
DeveloperReflection AIDeepSeek
Based inUnited StatesChina
Parameters (total / active)501B / 23B552B / 16B
WeightsApache 2.0, promised for later in October 2026MIT
ReleasedOctober 5, 2026September 10, 2026
How to accessWaitlisted beta API; weights comingHugging Face, DeepSeek API, OpenRouter

Other comparisons