觉
AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-09-25 6 浏览 公开

研究发现:顶尖 AI 专家严重低估了该领域的发展速度

FRI 临时报告称,专家和超级预测者对 AI 在数学、病毒学、网络安全和经济指标上的进展预测系统性偏低;但生物安全和自动驾驶等领域可能被高估。受访者正在上调预期。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-09-25 03:18:33

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

How fast is AI improving? That question usually goes to experts at top universities, heavily cited AI researchers, and seasoned economists. Yet these same specialists significantly underestimated recent progress on benchmarks and some adoption metrics, according to an interim report from the Forecasting Research Institute (FRI). Since mid-2022, FRI has collected forecasts on AI progress across several studies and projects . Its samples include senior specialists. The first round of LEAP (Longitudinal Expert AI Panel) drew 339 experts, including 76 computer scientists, 76 industry experts, 68 economists, and 119 AI policy specialists. The computer scientists included 30 professors at top-20 institutions and 10 of the 200 most-cited AI authors. The panels also included superforecasters, generalists with a proven record of accurate predictions. The widest gap involves math . AI reached gold-medal level at the International Mathematical Olympiad in July 2025 , five years before the median expert forecast and ten years before the median superforecaster forecast. Those predictions were gathered in 2022, before ChatGPT launched, but the pattern held afterward too, according to FRI. AI may also have solved a Millennium Prize Problem , though it's still unclear whether the solution meets the evaluation criteria. In a survey from August and September 2025, experts had put the median odds of such a solution by the end of 2027 at just 10 percent, and superforecasters at 5.4 percent. In a study of AI capabilities in virology, experts predicted AI models wouldn't match a top team of virologists on a troubleshooting benchmark until 2030. Superforecasters said 2034. FRI says that likely happened as early as April 2025 . A cybersecurity benchmark showed similar underestimates. Economic forecasts were also far too conservative. Experts put the median for the highest annual recurring revenue (ARR) of any AI company at the end of 2026 at $20 billion. Economists said $16 billion, and superforecasters said $25 billion. FRI cites roughly $100 billion for Anthropic in September 2026 as a figure that has likely already been reached. Not every forecast ran too low. Biosecurity experts predicted that 22.5 percent of participants using a language model would complete biological lab tasks. Virologists expected 40 percent, superforecasters 16.2 percent. In a controlled trial, only 5.2 percent succeeded with a language model and internet access, compared with 6.6 percent using the internet alone. The language model made no measurable difference, though the trial was small. Experts may also have overshot on self-driving cars. Their median forecast for the share of autonomous US ride-hailing trips in 2027 was 7.3 percent, while an LLM projection puts it at 2.5 percent. FRI says forecasts on economic growth, employment, and major AI harms can't be reliably judged yet. At the same time, respondents are revising their expectations upward. Among those who completed both surveys, the average probability assigned to AI becoming a "technology of the century" rose from 31 to 36 percent for experts and from 28 to 35 percent for superforecasters over nine months. Going forward, FRI will highlight a subsample of respondents who expect very rapid AI progress through 2040 and publish continuously updated LLM forecasts alongside the human ones. According to ForecastBench , some models already match superforecasters on certain question types. FRI also wants to find the most accurate LEAP panelists and feature their forecasts once enough data is in. RI does flag a catch in its own data: underestimates become obvious as soon as reality overtakes a prediction, but overestimates only become clear once a deadline passes. That makes the interim report naturally tilted toward finding cases where forecasters were too cautious. Some of FRI's own assessments also rely on LLM projections that use information the original forecasters didn't have.

中文翻译

AI 进步有多快?这个问题通常会交给顶尖大学的专家、被大量引用的 AI 研究人员以及经验丰富的经济学家。然而,根据 Forecasting Research Institute (FRI) 的一份临时报告,这些专家显著低估了近期在基准测试和一些采用指标上的进展。

自 2022 年中以来,FRI 在多项研究和项目中收集了关于 AI 进展的预测。其样本包括资深专家。第一轮 LEAP(Longitudinal Expert AI Panel)吸引了 339 名专家,包括 76 名计算机科学家、76 名行业专家、68 名经济学家和 119 名 AI 政策专家。计算机科学家包括 30 名来自排名前 20 机构的教授,以及 200 位被引用最多的 AI 作者中的 10 位。这些小组还包括超级预测者,即拥有准确预测记录的通才。

最大的差距涉及数学。AI 在 2025 年 7 月达到国际数学奥林匹克金牌水平,比专家的中位预测早五年,比超级预测者的中位预测早十年。这些预测是在 2022 年 ChatGPT 发布之前收集的,但根据 FRI 的说法,这种模式之后也持续存在。AI 可能还解决了一个千禧年大奖难题,不过目前仍不清楚该解决方案是否符合评估标准。在 2025 年 8 月和 9 月的一项调查中,专家认为到 2027 年底出现此类解决方案的中位概率仅为 10%,超级预测者为 5.4%。

在一项关于病毒学中 AI 能力的研究中,专家预测 AI 模型要到 2030 年才能在故障排除基准上匹敌顶级病毒学家团队。超级预测者说是 2034 年。FRI 称这可能早在 2025 年 4 月就发生了。一个网络安全基准也显示出类似的低估。

经济预测也过于保守。专家对 2026 年底任何 AI 公司最高年度经常性收入(ARR)的中位预测为 200 亿美元。经济学家说是 160 亿美元,超级预测者说是 250 亿美元。FRI 引用 Anthropic 在 2026 年 9 月约 1000 亿美元作为可能已经达到的数字。

并非所有预测都过低。生物安全专家预测,22.5% 的使用语言模型的参与者会完成生物实验室任务。病毒学家预计为 40%,超级预测者为 16.2%。在一项对照试验中,只有 5.2% 的人在使用语言模型和互联网时成功,而仅使用互联网的人为 6.6%。语言模型没有带来可测量的差异,尽管该试验规模很小。专家们也可能高估了自动驾驶汽车。他们对 2027 年美国自动驾驶网约车出行份额的中位预测为 7.3%,而一项 LLM 预测为 2.5%。

FRI 称,关于经济增长、就业和重大 AI 危害的预测目前还无法可靠判断。与此同时,受访者正在上调他们的预期。在完成两次调查的人中,九个月内,AI 成为“世纪技术”的平均概率从专家的 31% 上升到 36%,超级预测者从 28% 上升到 35%。未来,FRI 将突出一个预期到 2040 年 AI 会非常快速进展的受访者子样本,并与人类预测一起发布持续更新的 LLM 预测。根据 ForecastBench,一些模型已经在某些问题类型上匹敌超级预测者。FRI 还想找出最准确的 LEAP 小组成员,并在获得足够数据后展示他们的预测。

FRI 也指出了其自身数据中的一个问题:一旦现实超越预测,低估就会变得明显,但高估只有在截止日期过去后才会清楚。这使得这份临时报告自然偏向于发现预测者过于谨慎的案例。FRI 的一些评估还依赖 LLM 预测,而这些预测使用了原始预测者没有的信息。

核心信息

FRI 临时报告称,专家和超级预测者对 AI 在数学、病毒学、网络安全和经济指标上的进展预测系统性偏低;但生物安全和自动驾驶等领域可能被高估。受访者正在上调预期。

  • FRI 临时报告称,专家和超级预测者对 AI 在数学、病毒学、网络安全和经济指标上的进展预测系统性偏低;但生物安全和自动驾驶等领域可能被高估。受访者正在上调预期。
  • 原贴提到:How fast is AI improving? That question usually goes to experts at top u
  • 来源:the-decoder.com

详细解读

FRI 的临时报告给出的核心信号不是“某个模型又破了纪录”,而是预测者本身存在系统性偏差。339 名专家、包括顶尖计算机科学家和经济学家,以及有战绩的超级预测者,在数学、病毒学、网络安全和 AI 公司 ARR 等指标上普遍低估了实际进展。最典型的是 IMO 金牌:实际发生在 2025 年 7 月,比专家中位预测早五年,比超级预测者早十年。

为什么重要:如果最专业的人群都低估,那么企业预算、投资节奏、政策监管、人才规划和个人职业选择都可能建立在过慢的假设上。尤其当 AI 公司 ARR 被预测为 160-250 亿美元,而 FRI 引用 Anthropic 2026 年 9 月约 1000 亿美元时,说明市场采用速度可能远超传统模型。误差不只影响科技行业,也会影响经济学、生物安全和国家安全讨论。

对谁有价值:AI 创业者与投资者可用它校准融资和产品路线;企业战略与 HR 可重新评估自动化和技能升级时间表;政策制定者需要更快响应的监管框架;个人则可把“AI 能力到 2030 年才如何”的假设调短。超级预测者的存在也提示,可以引入外部校准机制,而不是只依赖领域专家。

可以怎么行动:第一,建立短周期预测复盘,每季度更新关键基准,而不是年度规划。第二,对关键假设做情景规划,同时准备“进展快于预期”和“慢于预期”两套方案。第三,关注 FRI 后续 LEAP 子样本和 LLM 预测,把它们当作辅助信号而非真理。第四,在安全、生物和自动驾驶等有高估案例的领域保持双向校准,避免把“专家低估”简单外推为“所有领域都会更快”。

风险与限制:FRI 自己指出,低估一旦被现实超越就立刻可见,高估要等截止日期才暴露,因此临时报告天然偏向“预测者太保守”的案例。部分评估还使用 LLM 预测,而 LLM 可能拥有原始预测者没有的信息,存在后见之明偏差。生物安全对照试验规模小,自动驾驶预测也显示专家可能高估。因此,结论应是“预测需要更频繁校准”,而不是“AI 一定会在所有方向超预期”。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《研究发现:顶尖 AI 专家严重低估了该领域的发展速度》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

研究发现:顶尖 AI 专家严重低估了该领域的发展速度主要讲什么?

FRI 临时报告称,专家和超级预测者对 AI 在数学、病毒学、网络安全和经济指标上的进展预测系统性偏低;但生物安全和自动驾驶等领域可能被高估。受访者正在上调预期。

这篇文章最值得关注的要点是什么?

FRI 临时报告称,专家和超级预测者对 AI 在数学、病毒学、网络安全和经济指标上的进展预测系统性偏低;但生物安全和自动驾驶等领域可能被高估。受访者正在上调预期。;原贴提到:How fast is AI improving? That question usually goes to experts at top u;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI日报、AI工具、Agent工作流专题里阅读。 关联原因:这篇内容命中「热点解读」等主题信号。;这篇内容来自该专题长期覆盖的栏目。;这篇内容来自该专题长期覆盖的栏目。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI日报、每日AI日报、AI信号、热点解读、BuilderPulse这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 要求对高影响操作提供在场证明 下一篇 安全研究者披露黑客用 GEO 污染 ChatGPT、Gemini 和 Google AI Overview,374 家企业被植入诈骗联系方式