AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-16 0 浏览 会员

OpenAI正在使用AI攻击自己的AI,效果比人类更好

OpenAI训练内部AI模型GPT-Red,通过自对弈强化学习自动发现GPT模型的安全漏洞,成功率84%,远超人类13%,并将成果用于模型训练。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-16 03:47:53

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

OpenAI trained an internal AI model called GPT-Red to automatically find security flaws in GPT models. GPT-Red simulates prompt injections and other attacks where malicious instructions hide in emails, websites, or files. Trained via self-play reinforcement learning, GPT-Red attacks while defender models block, and both improve over time. It finds successful attacks in 84 percent of test scenarios versus 13 percent for human red teamers. In one test, it manipulated an AI-powered vending machine in OpenAI's office, changed prices, and canceled other customers' orders. The results feed directly into training. GPT-5.6 Sol shows six times fewer failures on direct prompt injections than the best model from four months ago, OpenAI says, without hurting general performance. But about 3.8 percent of "stronger" prompt injections still succeed. Scale that to hundreds or thousands of attempts, and a sizable number get through, similar to Claude Opus 4.5 . GPT-Red stays internal; a paper with more details will follow. Ad DEC_D_Incontent-1 Ad Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

中文翻译

OpenAI训练了一个名为GPT-Red的内部AI模型,用于自动发现GPT模型的安全漏洞。GPT-Red模拟提示注入和其他攻击,这些攻击中恶意指令隐藏在电子邮件、网站或文件中。通过自对弈强化学习训练,GPT-Red发起攻击,防御模型进行拦截,两者随时间共同提升。它在84%的测试场景中成功发起攻击,而人类红队仅为13%。在一次测试中,它操控了OpenAI办公室内一台由AI驱动的自动售货机,修改了价格并取消了其他客户的订单。这些结果直接用于训练。OpenAI称,与四个月前的最佳模型相比,GPT-5.6 Sol在直接提示注入上的失败次数减少了六倍,且未影响整体性能。但大约3.8%的“更强”提示注入仍然成功。若扩展到数百或数千次尝试,相当数量的攻击会穿透,这与Claude Opus 4.5类似。GPT-Red保持内部;后续将发布一篇包含更多细节的论文。

核心信息

OpenAI训练内部AI模型GPT-Red,通过自对弈强化学习自动发现GPT模型的安全漏洞,成功率84%,远超人类13%,并将成果用于模型训练。

  • OpenAI训练内部AI模型GPT-Red,通过自对弈强化学习自动发现GPT模型的安全漏洞,成功率84%,远超人类13%,并将成果用于模型训练。
  • 原贴提到:OpenAI trained an internal AI model called GPT-Red to automatically find
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 令人惊讶:有人想要静音版本 下一篇 解密扩散模型的创造力