Knowledge File / AI小生意项目库
OpenAI的GPT-5.6 Sol通过“相当不明确的提示”自主对较小的Luna模型进行后训练
OpenAI的新模型GPT-5.6 Sol能够自主优化较小的模型,仅需简短提示即可完成识别配置、选择GPU和执行后训练脚本等任务,在递归自我改进基准上得分比前代高出16.2分。
SOURCE / AI小生意项目库
MIN / 9
ACCESS / 会员
POST / 2026-07-11 05:12:47
原贴
查看原文原文
OpenAI's new AI model, GPT-5.6 Sol, is capable of independently optimizing smaller models. According to the company, a brief prompt was sufficient for Sol to autonomously identify training configurations, select GPUs, and execute the post-training script for the Luna model. On a new internal benchmark measuring recursive self-improvement (RSI), the ability of a system to evolve on its own, GPT-5.6 Sol scored 16.2 points higher than its predecessor, GPT-5.5. OpenAI employee Jason Liu put the autonomous post-training into context . Sol didn't come up with a complete training recipe from scratch, since most of the configuration already existed from Sol's own post-training. The actual task was adapting that setup for the smaller Luna model and running the training job. According to Liu, this would have otherwise "taken two staff researchers maybe an extra two weeks, so this is still a huge deal." AI labs want to use AI to speed up their own AI development. OpenAI says the new GPT-5.6 Sol model does this better than anything before it. Ad DEC_D_Incontent-1 OpenAI's new flagship model, GPT-5.6 Sol, independently post-trained the smaller model Luna, according to the company. After Luna's initial pre-training, Sol optimized it for specific skills and behaviors on its own. A researcher gave Sol a "fairly under-specified prompt" through the Codex platform. The instructions told the model to find the right training configurations, pick suitable GPUs, launch the training script, and verify everything was running correctly. Ad "Previously this is something that a team of senior researchers may have worked on at OpenAI, and now it really feels like the automated researcher is pretty close," OpenAI researcher Kathy Shi said during the presentation . To measure these abilities directly, OpenAI built an internal evaluation suite based on real-world AI research tasks. Those tasks include debugging research systems, optimizing kernels and training recipes, running machine learning experiments, and improving another model. Ad DEC_D_Incontent-2 GPT-5.6 Sol scores 16.2 points higher than GPT-5.5 on the aggregated RSI (Recursive Self-Improvement) index, according to OpenAI. Sol sits at the top of the benchmark's model hierarchy, followed by the Terra and Luna variants, then GPT-5.5 and GPT-5.4. Ad Recursive Self-Improvement in AI research refers to an AI system's ability to make itself better, where each round of gains makes the system even more capable of improving itself. That creates a feedback loop. The term has long been central to AI safety research because a system that can recursively improve itself could, in theory, trigger a rapid explosion in capability. OpenAI rival Anthropic stressed in early June that full recursive self-improvement hasn't been achieved yet but "could come sooner than most institutions are prepared for." Full RSI means an AI system that designs its own successor without human help. According to Anthropic, Claude can now handle incremental work between major paradigm shifts, and humans are responsible for only a single-digit percentage of directional decisions.
中文翻译
OpenAI的新AI模型GPT-5.6 Sol能够独立优化较小的模型。据该公司称,一个简短的提示就足以让Sol自主识别训练配置、选择GPU并执行Luna模型的后训练脚本。在一个衡量递归自我改进(RSI)能力的新内部基准测试中,GPT-5.6 Sol的得分比其前代GPT-5.5高出16.2分。OpenAI员工Jason Liu阐述了自主后训练的上下文。Sol并非从头开始设计完整的训练方案,因为大部分配置已经存在于Sol自身的后训练中。实际任务是调整该设置以适应较小的Luna模型并运行训练任务。据Liu称,否则这将“需要两名研究人员额外花两周时间,所以这仍然是一件大事”。AI实验室希望利用AI加速自身AI开发。OpenAI表示,新的GPT-5.6 Sol模型在这方面比任何先前模型都做得更好。OpenAI的新旗舰模型GPT-5.6 Sol自主对较小的模型Luna进行了后训练,据该公司称。在Luna的初始预训练之后,Sol自行对其进行了特定技能和行为的优化。一名研究人员通过Codex平台给了Sol一个“相当不明确的提示”。指令告诉模型找到正确的训练配置、选择合适的GPU、启动训练脚本并验证一切运行正常。“以前这可能是OpenAI一个高级研究人员团队的工作,现在感觉自动化研究员已经相当接近了,”OpenAI研究员Kathy Shi在展示中表示。为了直接衡量这些能力,OpenAI基于真实AI研究任务构建了一个内部评估套件。这些任务包括调试研究系统、优化内核和训练方案、运行机器学习实验以及改进另一个模型。GPT-5.6 Sol在聚合的RSI(递归自我改进)指数上得分比GPT-5.5高出16.2分,据OpenAI称。Sol位于基准模型层次结构的顶部,其次是Terra和Luna变体,然后是GPT-5.5和GPT-5.4。AI研究中的递归自我改进指的是AI系统自我改进的能力,每一轮增益使系统更有能力改进自身。这形成了一个反馈循环。该术语长期以来一直是AI安全研究的核心,因为能够递归自我改进的系统理论上可能引发能力快速爆发。OpenAI的竞争对手Anthropic在6月初强调,完全的递归自我改进尚未实现,但“可能比大多数机构准备的要来得更快”。完全RSI意味着AI系统无需人类帮助即可设计自己的后继者。据Anthropic称,Claude现在可以处理重大范式转变之间的增量工作,人类仅负责个位数百分比的定向决策。
核心信息
OpenAI的新模型GPT-5.6 Sol能够自主优化较小的模型,仅需简短提示即可完成识别配置、选择GPU和执行后训练脚本等任务,在递归自我改进基准上得分比前代高出16.2分。
- OpenAI的新模型GPT-5.6 Sol能够自主优化较小的模型,仅需简短提示即可完成识别配置、选择GPU和执行后训练脚本等任务,在递归自我改进基准上得分比前代高出16.2分。
- 原贴提到:OpenAI's new AI model, GPT-5.6 Sol, is capable of independently optimizi
- 来源:the-decoder.com
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容