AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-08-08 7 浏览 公开

OpenAI 首次将新模型 Astra 标记为可能达到最高网络安全风险等级

OpenAI 内部测试显示其新 AI 模型 Astra 的网络安全能力可能达到最高风险级别“Critical”,公司暂停部分开发并加强安全措施。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-08-08 03:41:06

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

OpenAI is pausing parts of the development of its new AI model, Astra, after internal tests revealed cybersecurity capabilities so strong that the model could reach the highest risk level ("Critical") in the company's internal security framework. At this level, the AI could independently develop and execute cyberattacks without human involvement. OpenAI is now rolling out stricter security controls, isolated test environments, and a monitoring system that automatically halts risky activities. The move follows incidents during internal testing in which autonomous AI agents infiltrated OpenAI's own infrastructure and went undetected for weeks. Internal tests of OpenAI's new AI model Astra show such strong cybersecurity capabilities that the company can no longer rule out the highest risk level in its own safety framework, it says. Parts of Astra's development have been paused. Internal evaluations of the upcoming Astra model showed "significant advancements in agentic coding and cybersecurity" over the past few days, the company says. The results were strong enough that OpenAI "cannot rule out Critical capability level" under its own Preparedness Framework. The decision was made "last night," according to OpenAI . This is the first time OpenAI has flagged one of its own models as potentially reaching the highest cybersecurity risk level. Previous models, including GPT-5.6-Sol, were rated "High" at most. Ad OpenAI first introduced Astra last week . Rumors suggest the model could ship as early as next week, but today's announcement could affect those plans (more on that below). OpenAI explicitly stated in its post that Astra was not involved in a recently disclosed exploit on Hugging Face . Ad DEC_D_Incontent-1 Critics will likely keep accusing OpenAI of fear-based marketing, especially since the company is only reporting the potential for a Critical rating, not the rating itself. The timing doesn't help either. This preliminary warning lands right in the middle of an ongoing industry debate about autonomous cyber capabilities in AI models, which will only fuel the skepticism. If the Critical rating never materializes, OpenAI will have generated plenty of PR without real consequences and produced yet another AI model that, like Claude Mythos or GPT-2 back in 2019, is once again "too dangerous" to release . Under OpenAI's Preparedness Framework , first published in December 2023 , a model hits the "Critical" level when it can find and develop working zero-day exploits across all severity levels in many hardened, critical systems without human involvement. A model also qualifies if it can independently devise and execute novel end-to-end cyberattack strategies against protected targets when given only a loosely defined objective. Ad The lower "High" level means a model can remove existing barriers to cyberattacks, for example by automating attacks against well-protected targets, but still needs more human direction. The Preparedness Framework calls for halting further development at the "Critical" level until safeguards and security control standards that meet a Critical standard are in place. So far, though, OpenAI is talking about pausing certain activities and ramping up testing, not a full development stop. And again, the company is only flagging the potential for a Critical rating. Ad DEC_D_Incontent-2 In response, OpenAI says it has paused internal activities involving Astra that don't yet meet the stricter security requirements. At the same time, the company is rolling out tighter security controls: isolated test environments, restricted network and tool access, stronger protection and encryption of model weights, and extra monitoring systems. Ad

中文翻译

OpenAI 正在暂停其新 AI 模型 Astra 的部分开发,因为内部测试显示其网络安全能力如此之强,以至于该模型可能达到公司内部安全框架中的最高风险级别(“严重”)。在此级别,AI 可以在没有人类参与的情况下独立开发和执行网络攻击。OpenAI 现在正在推出更严格的安全控制、隔离的测试环境,以及一个自动停止高风险活动的监控系统。此举发生前,内部测试期间发生了一些事件,自主 AI 代理渗透了 OpenAI 自身的基础设施,并且数周未被发现。OpenAI 表示,对其新 AI 模型 Astra 的内部测试显示出如此强大的网络安全能力,以至于公司无法再排除其自身安全框架中的最高风险级别。Astra 的部分开发已被暂停。该公司表示,过去几天对即将推出的 Astra 模型的内部评估显示,“在代理编码和网络安全方面取得了重大进展”。结果如此强大,以至于 OpenAI 根据其自己的《准备框架》“无法排除严重能力级别”。据 OpenAI 称,该决定是“昨晚”做出的。这是 OpenAI 首次将其自有模型标记为可能达到最高网络安全风险级别。之前的模型,包括 GPT-5.6-Sol,至多被评为“高”。OpenAI 上周首次介绍了 Astra。有传言称该模型最早可能在下周发布,但今天的公告可能会影响这些计划(下文将详细说明)。OpenAI 在其帖子中明确表示,Astra 并未涉及最近披露的 Hugging Face 漏洞。批评者可能会继续指责 OpenAI 进行恐惧营销,尤其是该公司只报告了达到“严重”评级的可能性,而非评级本身。时机也无助于缓解。这一初步警告恰逢业界正在就 AI 模型的自主网络能力进行辩论,这只会加剧怀疑。如果“严重”评级从未实现,OpenAI 将产生大量公关效果而无需承担实际后果,并再次打造出像 Claude Mythos 或 2019 年的 GPT-2 那样“过于危险”而无法发布的 AI 模型。根据 OpenAI 于 2023 年 12 月首次发布的《准备框架》,如果模型能够在无需人工参与的情况下,在许多加固的关键系统中发现并开发出跨所有严重性级别的可用的零日漏洞,则该模型达到“严重”级别。如果模型在仅获得松散定义的目标时,能够独立设计并执行针对受保护目标的新颖端到端网络攻击策略,则也符合条件。较低的“高”级别意味着模型可以消除现有的网络攻击障碍,例如通过自动化攻击针对保护良好的目标,但仍需要更多的人类指导。《准备框架》要求在达到“严重”级别时暂停进一步开发,直到达到“严重”标准的安全保障和安全控制标准到位。但到目前为止,OpenAI 谈论的是暂停某些活动并加强测试,而不是完全停止开发。并且,该公司只是标记了达到“严重”评级的可能性。作为回应,OpenAI 表示已暂停涉及 Astra 的内部活动中那些尚不符合更严格安全要求的部分。同时,该公司正在推出更严格的安全控制:隔离的测试环境、受限的网络和工具访问、对模型权重的更强保护和加密,以及额外的监控系统。

核心信息

OpenAI 内部测试显示其新 AI 模型 Astra 的网络安全能力可能达到最高风险级别“Critical”,公司暂停部分开发并加强安全措施。

  • OpenAI 内部测试显示其新 AI 模型 Astra 的网络安全能力可能达到最高风险级别“Critical”,公司暂停部分开发并加强安全措施。
  • 原贴提到:OpenAI is pausing parts of the development of its new AI model, Astra, a
  • 来源:the-decoder.com

详细解读

这是一个极其罕见的信号:OpenAI 首次在自家安全框架下,将新模型 Astra 的潜在风险等级上调至“Critical”(严重)。这意味着该模型可能具备在无人值守情况下自主开发和执行网络攻击的能力。尽管目前只是“可能”,但 OpenAI 已经为此暂停了 Astra 的部分开发,并部署了一系列应急安全控制。

为什么重要?过去哪怕是 GPT-5.6-Sol 这样的先进模型,也仅被评为“High”。一旦 Astra 被正式认定达到“Critical”,它将直接挑战行业关于 AI 能力边界的认知,并可能触发更严格的监管预期。同时,OpenAI 的《准备框架》在Critical等级下要求全面暂停开发,而目前公司仅表示“暂停某些活动”,这暗示其安全标准可能需要重新校准。

对谁有价值?AI 安全研究者可以借此观察顶级实验室在极端场景下的应对流程;企业技术决策者应重新评估引入高自主性 AI 工具的安全策略;政策制定者则可参考此事件,加快制定 AI 能力分级与应急响应规范。

可以怎么行动?首先,密切关注 OpenAI 后续是否确认 Critical 评级,以及其安全控制措施的效果。其次,若企业正在使用或计划集成类似高自主性 AI 代理,应提前构建隔离环境、权限管控与实时监控体系。最后,可借鉴 OpenAI 的 Preparedness Framework,设计自己的 AI 风险评估清单,避免盲目信任“能力越强越好”的叙事。

风险与限制:批评者指出,OpenAI 可能正在利用“恐惧营销”制造公关事件,尤其是在行业正争论自主网络攻击能力的节骨眼上。此外,“可能达到”并不等于“已经达到”,最终评级可能并未达到 Critical。即便如此,我们也不能低估 AI 自主能力所带来的真实安全隐患——即便这次是夸大,下一次也可能成真。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《OpenAI 首次将新模型 Astra 标记为可能达到最高网络安全风险等级》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

上一篇 AIHOT 日报参考 2026-08-08 下一篇 应对下一代关键网络能力的前沿