跳至内容
reflectionbeam

正面对比

Beam vs DeepSeek V4.1 Flash

2:7

Beam 在 9 项共同基准中赢得 2 项

结论

DeepSeek V4.1 Flash 在双方都参与的 9 项基准测试中赢下 7 项,在 DeepSWE 和 AutomationBench 上优势明显,在 Terminal-Bench 上领先 10 分。Beam 在 CritPt 上胜出,AA Omniscience 指数也更高;该指数会对自信地给出错误答案的模型扣分。这里 Beam 在效率上并无优势:DeepSeek V4.1 Flash 每生成一个 token 仅激活 160 亿参数,少于 Beam 的 230 亿参数,而且其权重已根据 MIT 许可证公开。

逐项基准对比

基准BeamDeepSeek V4.1 Flash差异
DeepSWE v1.144.474.2-29.8
Terminal-Bench v2.180.190.6-10.5
Humanity's Last Exam (no tools)36.239.1-2.9
SciCode49.752.0-2.3
CritPt (AA)16.314.3+2.0
GPQA Diamond90.590.9-0.4
AutomationBench (public)37.054.8-17.8
AA-LCR (long context)79.384.0-4.7
AA Omniscience index (public split)13.06.6+6.4

所有分数均来自 Reflection 的公告(2026 年 10 月 8 日修订版)。DeepSeek V4.1 Flash 的分数是 Reflection 引用 Artificial Analysis 和 DataCurve 的数据。

两款模型

BeamDeepSeek V4.1 Flash
开发者Reflection AIDeepSeek
总部所在地美国中国
参数量(总计 / 激活)501B / 23B552B / 16B
权重Apache 2.0,承诺于 2026 年 10 月晚些时候推出MIT
发布日期2026年10月5日2026年9月10日
获取方式API beta 版需排队申请;模型权重即将推出Hugging Face, DeepSeek API, OpenRouter

其他对比