AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-07-16 6 浏览 公开

GPT-5.6 Sol 在90分钟内推翻人类30年未解统计学猜想

宾夕法尼亚大学研究者使用 OpenAI 的 GPT-5.6 Sol Pro 成功推翻了一个存在30年的统计学猜想,该猜想涉及广泛使用的 Benjamini-Hochberg 方法在连续数据上的可靠性。AI 仅用90分钟完成,而前代 GPT-5.5 在超过20小时计算后未能给出有效解。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-07-16 01:35:12

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

A researcher at the University of Pennsylvania has used OpenAI's language model GPT-5.6 Sol Pro to solve a long-standing open problem in statistics. The breakthrough involved disproving the assumption that the widely used Benjamini-Hochberg method for controlling false positives always works reliably, even when applied to continuous data. GPT-5.6 Sol Pro completed the task in roughly 90 minutes, while its predecessor, GPT-5.5, failed to produce a valid solution even after more than 20 hours of computation. A University of Pennsylvania statistics professor used OpenAI's GPT-5.6 to solve one of the central open questions in his field. When researchers test thousands of hypotheses at once, like scanning the human genome for disease-linked genes, they run into a problem: The more tests you run, the more false positives slip through. In 1995, statisticians Yoav Benjamini and Yosef Hochberg developed a method to limit these false positives. It controls the false discovery rate, or FDR, which is the share of reported significant results that are actually false alarms. Ad The Benjamini-Hochberg procedure , or BH, is now widely used in modern statistics and across many scientific fields. According to Edgar Dobriban , an associate professor at the University of Pennsylvania's Wharton School, the original paper has received more than 130,000 citations. Ad DEC_D_Incontent-1 Benjamini and Hochberg originally showed that their method works with independent data. Real-world data points, however, are often linked. Genetic variants can be correlated, for example, when certain locations in the genome are frequently inherited together. For years, experts assumed the BH procedure would also work reliably with correlated, normally distributed data, specifically when testing for deviations in both directions. But nobody had ever proved it. Ad an has now disproven that assumption using OpenAI's GPT-5.6 Sol Pro . In his preprint , he uses the AI to construct a statistical model where the actual false discovery rate provably exceeds the target level. Simulations confirm the result. Dobriban also published the accompanying code . Dobriban writes that the gap above the target level is "relatively small (0.104 vs 0.1)," so the result mainly matters for theory at this point. Practical effects still need further study, and the finding doesn't mean the BH procedure is generally unusable. Ad DEC_D_Incontent-2 The result is still significant for statisticians because AI solved the problem quickly after humans had failed. Dobriban says GPT-5.6 Sol Pro took about 90 minutes. GPT-5.5 couldn't find a solution even after roughly 20 hours of work with several agents. "So the capability improvement is quite real. Exciting times to live in!" he writes. The full chat and prompt are available here . Ad

中文翻译

宾夕法尼亚大学的一位研究人员使用 OpenAI 的语言模型 GPT-5.6 Sol Pro 解决了一个长期未解的统计学开放问题。这一突破涉及推翻一个假设,即广泛使用的用于控制假阳性率的 Benjamini-Hochberg 方法在应用于连续数据时总是可靠。GPT-5.6 Sol Pro 大约在90分钟内完成了任务,而其前代 GPT-5.5 在超过20小时的计算后仍未能产生有效解。宾夕法尼亚大学的一位统计学教授使用 OpenAI 的 GPT-5.6 解决了其领域的核心开放问题之一。当研究人员同时测试数千个假设时,比如扫描人类基因组寻找与疾病相关的基因,他们遇到一个问题:测试越多,假阳性漏网越多。1995年,统计学家 Yoav Benjamini 和 Yosef Hochberg 开发了一种限制这些假阳性的方法。它控制假发现率(FDR),即实际为误报的报告显著结果的比例。Benjamini-Hochberg 程序(BH)现已在现代统计学和许多科学领域广泛使用。据宾夕法尼亚大学沃顿商学院副教授 Edgar Dobriban 称,该原始论文已获得超过13万次引用。Benjamini 和 Hochberg 最初表明他们的方法适用于独立数据。然而,现实世界的数据点通常是相互关联的。例如,当基因组中某些位置经常一起遗传时,遗传变异可能相关。多年来,专家们假设 BH 方法在相关、正态分布的数据上也可靠,特别是在测试两个方向的偏差时。但从未有人证明这一点。Dobriban 现在使用 OpenAI 的 GPT-5.6 Sol Pro 推翻了这一假设。在他的预印本中,他使用 AI 构建了一个统计模型,其中实际假发现率明确超过目标水平。模拟结果证实了这一发现。Dobriban 还发布了配套代码。Dobriban 写道,超出目标水平的差距“相对较小(0.104 vs 0.1)”,因此该结果目前主要对理论有影响。实际效果仍需进一步研究,且该发现并不意味着 BH 方法通常不可用。这一结果对统计学家仍然重要,因为 AI 在人类失败后迅速解决了问题。Dobriban 表示,GPT-5.6 Sol Pro 耗时约90分钟。GPT-5.5 在经过约20小时的工作后仍无法找到解。他写道:“因此能力提升是真实的。生活在一个激动人心的时代!”

核心信息

宾夕法尼亚大学研究者使用 OpenAI 的 GPT-5.6 Sol Pro 成功推翻了一个存在30年的统计学猜想,该猜想涉及广泛使用的 Benjamini-Hochberg 方法在连续数据上的可靠性。AI 仅用90分钟完成,而前代 GPT-5.5 在超过20小时计算后未能给出有效解。

  • 宾夕法尼亚大学研究者使用 OpenAI 的 GPT-5.6 Sol Pro 成功推翻了一个存在30年的统计学猜想,该猜想涉及广泛使用的 Benjamini-Hochberg 方法在连续数据上的可靠性。AI 仅用90分钟完成,而前代 GPT-5.5 在超过20小时计算后未能给出有效解。
  • 原贴提到:A researcher at the University of Pennsylvania has used OpenAI's languag
  • 来源:the-decoder.com

详细解读

这是什么信号:GPT-5.6 Sol Pro 在90分钟内推翻了一个统计学界30年未解的猜想,这是AI在基础科学领域取得突破性进展的明确信号。此前,人类专家长期假设 Benjamini-Hochberg 方法在连续相关数据下依然可靠,但无人能证明或证伪。AI 不仅解决了这一理论难题,还生成了可验证的代码和模型,表明其具备超越单纯模式匹配的推理与构造能力。

为什么重要:Benjamini-Hochberg 方法是控制假阳性率的基石,广泛应用于基因组学、神经科学、经济学等涉及多重假设检验的领域。推翻其连续数据下的无条件可靠性假设,意味着部分基于该方法得出的科学结论可能需要重新审视。更关键的是,AI 在极短时间内(90分钟)完成了人类数十年的未竟工作,凸显了大型语言模型作为科研工具的巨大潜力——尤其是针对需要创造性构造反例的开放问题。

对谁有价值:

  • 统计学家与数据科学家:直接受益,需要重新评估 BH 方法在相关数据下的适用边界,并可能催生新的修正方法。
  • AI 研究者:证明先进 LLM 可参与严肃数学推理,激励开发更专业的科学AI代理。
  • 科研资助机构与决策者:应重新思考AI在学术研究中的角色,加速布局AI辅助科研的基础设施。

可以怎么行动:

  1. 研究者可基于 Dobriban 发布的代码和模型,验证其反例在自己领域数据中的影响程度。
  2. 高校应开设AI辅助科研课程,教授如何利用LLM进行猜想生成、反例构造等高级任务。
  3. 开发专用工具:将类似GPT-5.6的模型集成到统计分析软件中,作为“假设检验顾问”实时发现潜在漏洞。

风险或限制:

  • 当前反例的差距较小(0.104 vs 0.1),实际科学应用中的影响可能有限,不可过度泛化。
  • AI 的推理过程尚不可完全解释,依赖黑箱结果可能引发可重复性危机。
  • 该能力高度依赖模型版本和提示词,GPT-5.5 的失败表明鲁棒性不足,需进一步测试。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《GPT-5.6 Sol 在90分钟内推翻人类30年未解统计学猜想》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

上一篇 前谷歌DeepMind研究员因公司签署无限制军事AI协议而离职 下一篇 【必读】每日AI日报 2026-07-15