AI觉醒星球
Awakening is here
Knowledge File / 全球热点解读
2026-06-16 13 浏览 公开

趋势解读:The US government may be asking Anthropic the,聚焦形式化数学证明能力

美国政府指责Anthropic无视行政命令发布Fable 5,暴露了自身对AI安全认知的不足。

SOURCE / 全球热点解读 MIN / 9 ACCESS / 公开 POST / 2026-06-16 02:06:33

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Government officials appear to be accusing Anthropic of disregarding Trump's cyber directive and releasing Fable 5 without explicit approval. Discussions are ongoing, but the government's accusation of a "jailbreak" mostly exposes its own gaps in knowledge. "Everybody said Anthropic was a bad actor. Some of us said it was time to give them a chance. Now those people are questioning that. They screwed us." That's how an administration official summed up the conflict between the Trump administration and Anthropic, according to Axios . As I suspected, government officials are accusing Anthropic of ignoring Trump's recently issued cyber executive order . The executive order called for supposedly voluntary government oversight of AI models . Anthropic welcomed the proposal but released Fable 5 without waiting for the designated clearinghouse, which could have signed off on the release, to be set up. Ad A government official also accuses Anthropic of knowing a jailbreak could occur. "They came to every fork in the road and took the wrong fork." The tip about this jailbreak, whose existence and severity haven't been confirmed, reportedly came from Amazon and other tech companies . Ad DEC_D_Incontent-1 Government sources also criticized the communication between the two sides to Axios. "It's like they just speak in different languages." The Department of Commerce and Anthropic employees are reportedly in talks, with more meetings planned involving the CIA and science advisor Michael Kratsios. The accusation that Anthropic knew about the jailbreak risk and stayed silent actually says more about the government's understanding of AI than about Anthropic. Anyone who works closely with AI models knows they can be hacked . OpenAI has warned that prompt injection, a related hacking method, may never be fully solved . There's no fix for LLM security yet. Ad The real question is how severe the breach is and how fast countermeasures kick in. But if the U.S. government insists frontier AI models must be "unhackable" before they ship internationally, tough talks are ahead. Then again, Anthropic isn't in a strong spot either. CEO Dario Amodei said back in 2023 that "a jailbreak could be life or death" if someone managed to bypass safety protocols in science, tech, and biology. Meanwhile, over 100 security experts and tech industry executives have published an open letter to Trade Secretary Lutnick and National Cyber Director Cairncross calling for export controls on Fable and Mythos to be lifted. They argue that while Anthropic's models are good at finding security flaws in software , they aren't uniquely good at it. Other models like GPT-5.5, Opus, Sonnet, and the Chinese Kimi 2.7 can do the same thing. Ad DEC_D_Incontent-2 Anthropic also built several safeguards into Fable that the security community actually dismissed as overkill on launch day. The signatories warn that export controls are stripping defenders of the best tools while Chinese open-weight models are only months behind the top U.S. models . Ad Signatories include Alex Stamos (Corridor), Rachel Tobac (SocialProof Security), Katie Moussouris (Luta Security), Dan Lorenc (Chainguard), and Joe Levy (Sophos). Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

中文翻译

政府官员似乎指责 Anthropic 无视特朗普的网络指令,未经明确批准就发布了《神鬼寓言 5》。讨论仍在进行中,但政府对“越狱”的指控大多暴露了其自身的知识差距。 “每个人都说 Anthropic 是个糟糕的演员。我们中的一些人说是时候给他们一个机会了。现在那些人开始质疑这一点。他们搞砸了我们。”据 Axios 报道,一位政府官员是这样总结特朗普政府与 Anthropic 之间的冲突的。正如我所怀疑的,政府官员指责 Anthropic 无视特朗普最近发布的网络行政命令。该行政命令要求政府对人工智能模型进行自愿监督。 Anthropic 对这一提议表示欢迎,但没有等到指定的信息交换所建立就发布了《神鬼寓言 5》,而指定的信息交换所本来可以签署该版本的。一名政府官员还指责 Anthropic 知道可能会发生越狱。 “他们来到了每一个岔路口,但都走错了岔路口。”据报道,有关此次越狱的消息来自亚马逊和其他科技公司,但其存在和严重程度尚未得到证实。 Ad DEC_D_Incontent-1 政府消息人士也向 Axios 批评了双方之间的沟通。 “就像他们只是说不同的语言一样。”据报道,商务部和 Anthropic 员工正在进行谈判,并计划举行更多会议,中央情报局和科学顾问迈克尔·克拉西奥斯 (Michael Kratsios) 也将参加。指责 Anthropic 明知越狱风险却保持沉默,实际上更多地说明了政府对 AI 的理解,而不是 Anthropic。任何与人工智能模型密切合作的人都知道它们可能会被黑客攻击。 OpenAI 警告称,提示注入这种相关的黑客方法可能永远无法完全解决。目前还没有针对 LLM 安全性的修复。真正的问题是违规行为有多严重以及反制措施的启动速度有多快。但如果美国政府坚持前沿人工智能模型在国际化之前必须“不可破解”,那么艰难的谈判就在眼前。话又说回来,Anthropic 也不处于强势地位。首席执行官达里奥·阿莫迪 (Dario Amodei) 早在 2023 年就表示,如果有人设法绕过科学、技术和生物学领域的安全协议,“越狱可能是生死攸关”。与此同时,100 多名安全专家和科技行业高管向贸易部长卢特尼克和国家网络总监凯恩克罗斯发表了一封公开信,呼吁取消对《神鬼寓言》和《神话》的出口管制。他们认为,虽然 Anthropic 的模型擅长发现软件中的安全缺陷,但它们并不是唯一擅长于此的。其他型号,如 GPT-5.5、Opus、Sonnet 和中国版 Kimi 2.7 也可以做同样的事情。 DEC_D_Incontent-2 Anthropic 还在《Fable》中构建了几项安全措施,但安全社区在发布当天实际上认为这些措施太过分了。签署者警告说,出口管制正在剥夺捍卫者最好的工具,而中国的开放式重量模型仅落后于美国顶级模型几个月。签署者包括 Alex Stamos(Corridor)、Rachel Tobac(SocialProof Security)、Katie Moussouris(Luta Security)、Dan Lorenc(Chainguard)和 Joe Levy(Sophos)。订阅 THE DECODER 即可享受无广告阅读、每周一次的 AI 时事通讯、我们每年六次的独家“AI 雷达”前沿报告、完整的存档访问权限以及我们的评论部分的访问权限。

核心信息

美国政府指责Anthropic无视行政命令发布Fable 5,暴露了自身对AI安全认知的不足。

  • 美政府指责Anthropic发布未授权模型。
  • 政府“越狱”指控暴露AI认知短板。
  • Anthropic面临出口管制压力。
  • 安全专家呼吁取消出口限制。
  • 中美AI竞争加剧监管博弈。

详细解读

这是什么信号

美国政府与前沿AI公司Anthropic之间的冲突升级,焦点在于Anthropic未经明确批准就发布了新模型Fable 5,被指责无视特朗普政府的网络行政命令。政府官员还声称Anthropic明知模型存在“越狱”风险却未及时通报。这反映了监管机构对AI安全问题的认知滞后,以及政策执行与行业发展之间的脱节。

为什么重要

此事件揭示了美国政府试图在AI安全领域加强控制,但缺乏具体有效的技术手段和知识基础。Anthropic作为行业领先公司,其模型安全措施已被社区认为过于严格,仍被指责,说明现有监管框架可能不切实际。同时,出口管制争论凸显了中美AI竞争中的技术博弈,若美国坚持“不可破解”标准,可能阻碍自身AI发展并丧失市场优势。

对谁有价值

对A安全研究人员:理解政府关注的漏洞类型和监管动向;对AI公司:警示与政府沟通的重要性,以及提前预判合规风险;对政策制定者:需与国际合作细化AI安全标准,避免一刀切;对投资人:评估监管不确定性对AI企业估值的影响。

可以怎么行动

AI公司应主动与监管机构建立定期沟通机制,预演安全审查流程;安全团队可针对“越狱”风险提交技术白皮书,帮助政府了解技术限制;行业协会可组织多方讨论,推动务实安全标准;长期关注中美出口管制政策变化,调整研发和市场策略。

风险或限制

报道中的“越狱”指控尚未被证实,可能影响事件判断;政府与企业的认知鸿沟短期内难以弥合,可能引发更多摩擦;若出口管制持续,可能导致美国AI企业将研发重心转移至海外,或催生闭源替代方案。

信息差价值

信息差价值:多数媒体仅报道冲突表面,但本内容揭示了政府内部对AI安全认知的局限,以及安全社区对出口管制的反对声音。这种细节有助于理解政策背后的博弈,而非简单归责。

业务启发:AI公司应预判监管趋势,主动参与标准制定,避免被动应对。同时,模型安全设计需兼顾实用性与合规性,过度谨慎可能适得其反。

可沉淀动作:建立政策跟踪机制,定期输出合规简报;开展与监管机构的闭门交流;将安全案例纳入内部培训,提升团队风险意识。

参考来源

上一篇 趋势解读:OpenRouter新增免费模型gpt-oss-20b和Gemma4 26B,讨论数据集与基础模型 下一篇 下一代投机解码:DFlash 与 Spec V2