AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-05-09 5 浏览 公开

趋势解读:Fields Medalist says ChatGPT 5.5 Pro delivered "PhD-level",聚焦形式化数学证明能力

菲尔兹奖得主Timothy Gowers使用ChatGPT 5.5 Pro在2小时内无人类数学指导完成博士级数论研究,模型独立改进数学界限并产出预印本,被评价为“完全原创”。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-05-09 22:32:14

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

British mathematician Timothy Gowers used OpenAI's ChatGPT 5.5 Pro model to tackle open problems in number theory, with the AI producing complete scientific papers in under two hours, without any mathematical guidance from Gowers himself. According to Gowers, the AI's output reached "PhD-level" and managed to improve upon existing mathematical bounds, demonstrating a remarkable degree of independent mathematical reasoning. Isaac Rajagopal, a young researcher involved in the work, called the model's key idea "completely original," an achievement he said a human mathematician would be proud of after weeks of deliberation. British mathematician Timothy Gowers had ChatGPT 5.5 Pro tackle open problems in number theory. The model significantly improved an existing mathematical bound. One of the junior researchers involved calls the model's key idea "completely original." Fields Medalist Timothy Gowers writes in his blog that ChatGPT 5.5 Pro has produced a piece of doctoral-level mathematical research, and that his own mathematical contribution was zero. The model did all the work in under two hours. "I didn't even do anything clever with the prompts," Gowers writes. The mathematician, who holds the Combinatorics Chair at the College de France and is a Fellow at Trinity College Cambridge, fed the model open problems from a paper by number theorist Mel Nathanson. The paper investigates the possible sizes of certain sets of integer sums and how efficiently sets with prescribed properties can be constructed. Ad Nathanson had proved an exponential bound for one of the problems and asked whether it could be improved. According to Gowers, ChatGPT 5.5 Pro thought for 17 minutes and 5 seconds, then delivered the best possible construction with a quadratic bound. The core idea: the model swapped out a component in Nathanson's proof for a more efficient variant that's well known in combinatorics but whose application to this particular problem wasn't obvious. Ad DEC_D_Incontent-1 When asked, ChatGPT rewrote the argument as a LaTeX preprint in 2 minutes and 23 seconds. Gowers checked it for correctness, then had the model solve a related variant, which it handled without any issues. Both results are available as a preprint . A generalized version of the problem proved much harder. Here, there was prior work by Isaac Rajagopal, an MIT student who had proven an exponential dependency. Gowers gave ChatGPT Rajagopal's paper and asked for an improvement. Ad What followed was a gradual escalation: after 16 minutes and 41 seconds, the model delivered a first improvement. Rajagopal judged this step correct but called it a routine modification of his own work. Gowers then got, as he puts it, "greedy" and asked ChatGPT to try for a much stronger bound. After 13 minutes and 33 seconds, the model reported optimism but said two technical statements still needed checking. Another 9 minutes and 12 seconds later, the check was done. The finished preprint was ready in 31 minutes and 40 seconds. The model had improved the bound from exponential to polynomial. Ad DEC_D_Incontent-2 According to Gowers, Rajagopal declared the results are "almost certainly correct," both at the level of individual proof steps and the underlying ideas. Ad

中文翻译

英国数学家蒂莫西·高尔斯 (Timothy Gowers) 使用 OpenAI 的 ChatGPT 5.5 Pro 模型来解决数论中的开放性问题,人工智能在两个小时内即可生成完整的科学论文,而无需高尔斯本人的任何数学指导。根据高尔斯的说法,人工智能的输出达到了“博士水平”,并设法改进了现有的数学界限,展示了显着程度的独立数学推理。参与这项工作的年轻研究员艾萨克·拉贾戈帕尔 (Isaac Rajagopal) 称该模型的关键思想“完全原创”,他表示,经过数周的深思熟虑,人类数学家将为这一成就感到自豪。英国数学家 Timothy Gowers 让 ChatGPT 5.5 Pro 解决了数论中的开放问题。该模型显着改善了现有的数学界限。一位参与其中的初级研究人员称该模型的关键思想“完全原创”。菲尔兹奖得主 Timothy Gowers 在博客中写道,ChatGPT 5.5 Pro 产出了一项博士级别的数学研究,而他自己的数学贡献为零。该模型在两个小时内完成了所有工作。 “我什至没有根据提示做任何聪明的事情,”高尔斯写道。这位担任法兰西学院组合学主席、剑桥三一学院研究员的数学家,从数论学家 Mel Nathanson 的一篇论文中为模型提供了开放问题。本文研究了某些整数和集合的可能大小以及如何有效地构造具有规定属性的集合。 Nathanson 证明了其中一个问题的指数界限,并询问是否可以改进。据 Gowers 称,ChatGPT 5.5 Pro 思考了 17 分 5 秒,然后提供了具有二次边界的最佳可能构造。核心思想:该模型将 Nathanson 证明中的一个组件替换为组合数学中众所周知的更有效的变体,但其在这个特定问题上的应用并不明显。当被询问时,ChatGPT 在 2 分 23 秒内将参数重写为 LaTeX 预印本。高尔斯检查了它的正确性,然后让模型解决了一个相关的变体,它处理得没有任何问题。这两个结果都可以作为预印本提供。事实证明,该问题的广义版本要困难得多。在这里,麻省理工学院的学生艾萨克·拉贾戈帕尔 (Isaac Rajagopal) 之前的工作已经证明了指数依赖性。 Gowers 给了 ChatGPT Rajagopal 的论文并要求改进。接下来是逐步升级:16 分 41 秒后,模型实现了第一次改进。拉贾戈帕尔认为这一步骤是正确的,但称这是对自己工作的例行修改。正如 Gowers 所说,随后他变得“贪婪”,并要求 ChatGPT 尝试更强的界限。 13 分 33 秒后,模型报告乐观,但表示仍有两项技术声明需要检查。又过了9分12秒,检查完毕。预印本在 31 分 40 秒内完成。该模型已将界限从指数改进为多项式。根据高尔斯的说法,拉贾戈帕尔宣称结果“几乎肯定是正确的”,无论是在个人证明步骤还是基本思想层面。

核心信息

菲尔兹奖得主Timothy Gowers使用ChatGPT 5.5 Pro在2小时内无人类数学指导完成博士级数论研究,模型独立改进数学界限并产出预印本,被评价为“完全原创”。

  • 高尔斯用ChatGPT 5.5 Pro 2小时完成博士级数学研究
  • 模型独立改进数学界限,产出预印本
  • 模型关键思想被评价为“完全原创”
  • AI从指数界限改进为二次或多项式界限
  • 数学家的贡献为零,提示工程简单

详细解读

这是什么信号
菲尔兹奖得主高尔斯亲自验证并公开承认,ChatGPT 5.5 Pro 能在两小时内完成博士级数学研究,且人类数学贡献为零。这标志着大模型从“辅助工具”跃升为“独立研究者”,尤其在形式化数学证明领域实现了质的突破。模型不仅理解了问题,还自主调用组合数学中的技巧,生成全新构造并改进已知界限——这是此前认为 AI 难以企及的创造型推理。

为什么重要
1. 验证了推理深度的上限:模型思考17分钟后给出二次边界,后续又将指数边界改进为多项式,整个过程无需提示工程技巧。2. 研究方法可复制:从开放问题到预印本的全流程自动化,将传统数学研究的周期从数周缩短至几十分钟。3. 对AI可解释性提出新挑战:模型产生“完全原创”想法,但人类数学家仍需逐项验证——信任与审计的平衡成为关键议题。

对谁有价值
- 数学研究者:可将部分探索性工作外包给AI,快速对比不同猜想或构造的可行性。
- AI研发团队:该案例表明,提升模型推理的“自主性”比单纯扩大参数规模更有效,方向可聚焦于长链推理和符号操作。
- 教育者与学习者:PhD级研究门槛降低,数学教育需重新设计技能点——从解题转向鉴别、提问与评估AI输出。

可以怎么行动
1. 尝试复现:使用类似提示(直接给出开放问题文本)测试 GPT-5.5 Pro 在自身研究领域的效果。
2. 建立验证流程:对AI生成的数学论证进行结构化审查,例如分步骤交叉验证(如Gowers所做)。
3. 关注后续模型:OpenAI可能将此类能力集成到更通用的推理引擎中,建议跟踪其形式化证明工具如LEAN的集成进展。

风险或限制
- 领域局限:当前成功仅限于组合数论等非分析型数学,对更依赖直觉或物理背景的问题效果未知。
- 验证成本:人类数学家仍需耗费时间确认正确性(本例中Gowers进行了完整检查),可能无法节省总时间。
- 过度依赖:若研究者盲目信任AI输出而跳过实质审查,可能引发错误积累。另外,模型可能隐含训练数据中的偏差。

信息差价值

信息差价值
多数人仍将大模型视为“高级搜索引擎”或“文本生成器”,而高尔斯案例揭示了其作为“自主研究伙伴”的潜力。这一信息差在于:AI不仅能回忆知识,还能组合、创新并产出学术论文级别的成果。对于尚未将AI融入核心工作流的团队,这是一个警示——先行者已开始用AI缩短科研周期。

业务启发
内容生产领域可借鉴其工作流:让AI完成初稿(如研究报告、代码注释),人类聚焦于策略判断与质量把控。具体场景包括:复杂文档的自动生成、代码审查、数据分析中的假设检验。对于OPC,可将此类案例转化为“AI工作流改造”专题,输出SOP:如何像高尔斯一样“无脑”喂问题给AI,再二次加工。

可沉淀动作
1. 建立“AI辅助研究”案例库,记录每次交互的提示、耗时、质量评估,形成可复用的模板。
2. 开发验证清单:当AI输出“完全原创”结果时,按领域设计自动校验规则(如符号一致性检查)。
3. 发起内部挑战赛:在非核心任务中强制使用AI生成完整方案,测试人类编辑的附加价值,从而优化人机协作比例。

参考来源

上一篇 菲尔兹奖得主称 ChatGPT 5.5 Pro 在无人帮助下两小时内完成"博士级"数学研究 下一篇 Peekaboo 3.0 正式发布 专注操作与界面检测