觉
AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-04-30 2 浏览 免费阅读

趋势解读:AI evals are becoming the new compute bottleneck,提升开发者接入体验

AI 评估成本已跨越临界点,成为新的计算瓶颈。例如,Holistic Agent Leaderboard 运行 21,730 次评估耗资约 4 万美元,单个 GAIA 运行成本高达 2,829 美元。评估成本甚至可能超过预训练,尤其对于小模型,开发者需要更高效的评估策略。

SOURCE / AI技能杠杆 MIN / 4 ACCESS / 免费阅读 POST / 2026-04-30 06:00:06

原贴

查看原文
作者:Hugging Face Blog 来源站点:huggingface.co 原贴时间:
趋势解读:AI evals are becoming the new compute bottleneck,提升开发者接入体验

原文

Summary. AI evaluation has crossed a cost threshold that changes who can do it. The Holistic Agent Leaderboard (HAL) recently spent about $40,000 to run 21,730 agent rollouts across 9 models and 9 benchmarks. A single GAIA run on a frontier model can cost $2,829 before caching. Exgentic 's $22,000 sweep across agent configurations found a 33× cost spread on identical tasks, isolating scaffold choice as a first-order cost driver, and UK-AISI recently scaled agentic steps into the millions to study inference-time compute. In scientific ML, The Well costs about 960 H100-hours to evaluate one new architecture and 3,840 H100-hours for a full four-baseline sweep. While compression techniques have been proposed for static benchmarks, new agent benchmarks are noisy, scaffold-sensitive, and only partly compressible. Training-in-the-loop benchmarks are expensive by construction, and when you try to add reliability to these evals, repeated runs further multiply the cost. The cost problem started before agents. When Stanford's CRFM released HELM in 2022, the paper's own per-model accounting showed API costs ranging from $85 for OpenAI's code-cushman-001 to $10,926 for AI21's J1-Jumbo (178B), and 540 to 4,200 GPU-hours for the open models, with BLOOM (176B) and OPT (175B) at the top end. Perlitz et al. (2023) restate the larger HELM cost pattern, and IBM Research notes that putting Granite-13B through HELM "can consume as many as 1,000 GPU hours." Across HELM's 30 models and 42 scenarios, the aggregate of reported costs and GPU compute came to roughly $100,000. Another shocking observation came from Perlitz et al.'s analysis of EleutherAI's Pythia checkpoints: developers pay for evaluation repeatedly during model development. Pythia released 154 checkpoints for each of 16 models spanning 8 sizes, or 2,464 checkpoints if each model checkpoint is counted separately, so the community could study training dynamics. Running the LM Evaluation Harness across all those checkpoints turns eval into a multiplier on training: Perlitz et al. (2024) noted that evaluation costs "may even surpass those of pretraining when evaluating checkpoints." For small models, evaluation becomes the dominant compute line item across the whole development cycle. When we scale inference-time compute, we scale evaluation costs. Perlitz et al. then asked how much of HELM actually carried the rankings. The result was striking: a 100× to 200× reduction in compute preserved nearly the same ordering, with larger reductions still useful for coarse grouping under the paper's tiered analysis. Flash-HELM turned that finding into a coarse-to-fine procedure: run cheap evaluations first, then spend high-resolution compute only on the top candidates. Much of HELM's compute was confirming rankings that the field could have inferred much more cheaply. Other work reached the same conclusion from different angles. tinyBenchmarks compressed MMLU from 14,000 items to 100 anchor items at about 2% error using Item Response Theory. The Open LLM Leaderboard collapsed from 29,000 examples to 180. Anchor Points showed that as few as 1 to 30 examples could rank-order 87 language-model/prompt pairs on GLUE, and others followed, reducing dataset sizes by 90\%. Static benchmarks had a weakness you could exploit: model differences often concentrate in a small subset of items, so ranking can survive aggressive subsampling. That trick weakened sharply once benchmarks moved from static predictions to agents. A very nice public accounting of agent evaluation comes from the Holistic Agent Leaderboard (Kapoor et al., ICLR 2026). HAL runs standardized agent harnesses across nine benchmarks covering coding, web navigation, science tasks, and customer service, with shared scaffolds and centralized cost tracking. The headline cost: $40,000 for 21,730 rollouts across 9 models and 9 benchmarks. By April 2026, the leaderboard had grown to 26,597 rollouts. Ndzomga's independent reproduction arrives at almost

中文翻译

AI 评估已经跨越了一个成本门槛,改变了谁能够进行它。Holistic Agent Leaderboard (HAL) 最近花费约 4 万美元,在 9 个模型和 9 个基准上运行了 21,730 次智能体 rollout。在缓存前,一个前沿模型的单次 GAIA 运行可能花费 2,829 美元。Exgentic 对智能体配置进行的 2.2 万美元扫描发现,相同任务上成本差异达 33 倍,将 scaffold 选择确定为一级成本驱动因素。UK-AISI 最近将智能体步骤扩展到数百万以研究推理时计算。在科学机器学习中,The Well 评估一个新架构大约需要 960 H100 小时,完整四个基线扫描需要 3,840 H100 小时。虽然已经针对静态基准提出了压缩技术,但新的智能体基准噪声大、对 scaffold 敏感,且仅部分可压缩。循环训练基准本身成本高昂,而当你试图为这些评估增加可靠性时,重复运行会进一步倍增成本。成本问题在智能体出现之前就已经存在。当斯坦福 CRFM 在 2022 年发布 HELM 时,论文自身的每模型核算显示 API 成本从 OpenAI 的 code-cushman-001 的 85 美元到 AI21 的 J1-Jumbo (178B) 的 10,926 美元不等,开源模型则需 540 到 4,200 GPU 小时,BLOOM (176B) 和 OPT (175B) 处于高端。Perlitz 等人 (2023) 重申了更大的 HELM 成本模式,IBM Research 指出将 Granite-13B 通过 HELM 评估“可能消耗多达 1,000 GPU 小时”。在 HELM 的 30 个模型和 42 个场景中,报告的成本和 GPU 计算总量约为 10 万美元。另一个惊人的观察来自 Perlitz 等人对 EleutherAI 的 Pythia 检查点的分析:开发者在模型开发过程中反复为评估付费。Pythia 为 16 个模型(横跨 8 种规模)发布了 154 个检查点,如果分别计算每个模型检查点,则为 2,464 个检查点,以便社区研究训练动态。在所有检查点上运行 LM Evaluation Harness 使得评估成为训练的倍增器:Perlitz 等人 (2024) 指出,评估成本“在评估检查点时甚至可能超过预训练”。对于小模型,评估成为整个开发周期中的主导计算项。当我们扩展推理时计算时,我们也在扩展评估成本。Perlitz 等人接着问道,HELM 中有多少内容实际上支撑了排名。结果令人震惊:计算量减少 100 到 200 倍几乎保持了相同的排序,在论文的分层分析下,更大的减少仍然对粗分组有用。Flash-HELM 将这一发现转化为粗到细的流程:先运行廉价评估,然后将高分辨率计算仅用于顶级候选。HELM 的大部分计算是在确认本可以更廉价推断出的排名。其他工作从不同角度得出了相同结论。tinyBenchmarks 使用项目反应理论将 MMLU 从 14,000 个项目压缩到 100 个锚定项目,误差约 2%。Open LLM Leaderboard 从 29,000 个示例缩减到 180 个。Anchor Points 表明,只需 1 到 30 个示例即可排序 GLUE 上的 87 个语言模型/提示对,其他人则将数据集大小减少了 90%。静态基准有一个可被利用的弱点:模型差异通常集中在一小部分项目上,因此排序能够经受住激进的子采样。一旦基准从静态预测转向智能体,这种技巧就会急剧减弱。一个非常好的智能体评估公开核算来自 Holistic Agent Leaderboard (Kapoor 等人, ICLR 2026)。HAL 在九个涵盖编码、网页导航、科学任务和客户服务的基准上运行标准化智能体工具,配有共享 scaffold 和集中成本跟踪。 headline 成本:在 9 个模型和 9 个基准上进行 21,730 次 rollout 花费 4 万美元。到 2026 年 4 月,排行榜已增长到 26,597 次 rollout。Ndzomga 的独立复现几乎达到了

核心信息

AI 评估成本已跨越临界点,成为新的计算瓶颈。例如,Holistic Agent Leaderboard 运行 21,730 次评估耗资约 4 万美元,单个 GAIA 运行成本高达 2,829 美元。评估成本甚至可能超过预训练,尤其对于小模型,开发者需要更高效的评估策略。

  • AI评估成本激增,成为新计算瓶颈
  • 大量评估费用浪费在确认已知排名上
  • 静态基准可压缩,智能体评估难以压缩
  • 评估成本可能超过预训练,小模型尤甚
  • 开发者应采用粗到细的分层评估策略

详细解读

这是什么信号? AI 评估正在从辅助工具转变为成本中心,甚至成为比训练更昂贵的计算瓶颈。随着智能体系统和推理时计算的普及,评估成本呈指数级增长,静态基准的压缩技巧在动态场景下失效,开发者面临“评估通胀”危机。

为什么重要? 评估成本直接决定了 AI 开发的门槛和速度。小团队可能被挤出前沿评估,而大公司也需要重新优化资源分配。同时,评估结果的可信度受限于重复运行成本,导致模型排名可能失真。理解评估成本结构有助于更理性地设计基准和开发流程。

对谁有价值? 对 AI 开发者:需要重新评估评估策略,采用粗到细的分层方法。对模型供应商:降低评估成本可吸引更多用户。对研究者:探索可压缩评估或新评估范式(如自适应测试)。对投资者:关注评估基础设施和优化工具领域的创业机会。

可以怎么行动? 1. 采用 Flash-HELM 或 tinyBenchmarks 等粗到细评估策略,先进行低成本筛选。2. 对于智能体评估,优先使用共享 scaffold 和集中成本追踪(如 HAL)。3. 在模型开发中,避免对每个检查点进行完整评估,改为间隔采样。4. 关注评估即服务(EaaS)平台,如 Exgentic 等,以降低一次性投入。

风险或限制:压缩方法可能丢失边缘案例的细节,尤其对安全关键场景不适用。智能体评估的压缩仍在探索中,缺乏理论保证。此外,过度优化评估成本可能导致 benchmark hacking,降低评价的真实性。

信息差价值

信息差价值:多数开发者只关注模型训练成本,忽视了评估成本已悄然成为更大的开销。本文揭示的评估成本数据(如4万美元运行2万次rollout)和压缩技巧(如100倍计算量仍保持排名)是圈内少有的量化洞察,可帮助读者建立更完整的成本认知。

业务启发:对AI公司而言,评估成本优化不是锦上添花,而是竞争壁垒。建立内部粗到细评估流水线,可节省50%以上评估预算,同时加速迭代。对于评估平台(如Hugging Face),推出低成本评估套餐或压缩基准服务,将成为吸引中小开发者的利器。

可沉淀动作:1. 立即审计当前评估流程,识别高成本环节(如全量检查点评估)。2. 引入tinyBenchmarks或Flash-HELM思想,建立自己的分层评估库。3. 关注智能体评估工具(如HAL),参与相关开源项目以降低未来成本。4. 在团队内部分享评估成本意识,推动以“每分钱有效信息”为KPI的评估文化。

参考来源

AI SUMMARY

这篇文章回答了什么

趋势解读:AI evals are becoming the new compute bottleneck,提升开发者接入体验主要讲什么?

AI 评估成本已跨越临界点,成为新的计算瓶颈。例如,Holistic Agent Leaderboard 运行 21,730 次评估耗资约 4 万美元,单个 GAIA 运行成本高达 2,829 美元。评估成本甚至可能超过预训练,尤其对于小模型,开发者需要更高效的评估策略。

这篇文章最值得关注的要点是什么?

AI 评估成本已跨越临界点,成为新的计算瓶颈。例如,Holistic Agent Leaderboard 运行 21,730 次评估耗资约 4 万美元,单个 GAIA 运行成本高达 2,829 美元。评估成本甚至可能超过预训练,尤其对于小…;AI评估成本激增,成为新计算瓶颈;大量评估费用浪费在确认已知排名上;静态基准可压缩,智能体评估难以压缩

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI工具、AI超级个体专题里阅读。 关联原因:这篇内容命中「Agent、智能体、工作流」等主题信号。;这篇内容命中「自动化、模型」等主题信号。;这篇内容命中「技能」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 版本更新:langchain-ai/langchain 1.2.16,提升开发者接入体验 下一篇 OpenAI Developers 发布新动态,聚焦 Agent 工作流自动化
北竹游乐场 免费玩小游戏 免费玩