Knowledge File / AI技能杠杆
OpenAI发布GPT-5.6 Sol,对标Claude Mythos,但政府准入规则令其直呼不可持续
OpenAI推出GPT-5.6系列,旗舰Sol在代理编程和网络安全方面领先Claude Mythos,但美国政府限制仅限合作方使用,OpenAI批评该政策损害开发者和企业利益。
SOURCE / AI技能杠杆
MIN / 9
ACCESS / 会员
POST / 2026-06-27 02:30:01
原贴
查看原文原文
OpenAI's new GPT-5.6 generation includes the flagship Sol and two cheaper tiers, Terra and Luna. Sol matches or beats Anthropic's Claude Mythos 5 across benchmarks, with a clear lead in agentic coding and better token efficiency in cybersecurity. The US government is restricting access to select partners for now. OpenAI says the policy hurts developers and businesses. OpenAI's new flagship GPT-5.6 Sol claims a lead over Anthropic's Claude Mythos in agentic coding and goes toe to toe with it in cybersecurity. Access stays limited to a handful of partners for now. OpenAI has unveiled GPT-5.6 Sol, a new generation of models built to compete with Claude's Mythos class . The limited preview is only open to select partners through the API and Codex, at the explicit direction of the US government . The same government previously yanked Anthropic's Mythos-class model Fable 5 off the market . OpenAI isn't subtle about its frustration . "We don’t believe this kind of government access process should become the long-term default. It keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them." Ad GPT-5.6 also brings a new layered naming scheme that looks a lot like Claude's. The number (x.6) marks the generation, while Sol, Terra, and Luna are permanent performance tiers that can evolve on their own. Sol is the flagship. Terra matches GPT-5.5 at half the cost. Luna is the budget option. On top of that, there's a "max" mode for deeper reasoning and an "ultra" mode that farms out complex tasks to sub-agents running in parallel. Ad DEC_D_Incontent-1 OpenAI's benchmark numbers put Sol ahead of Anthropic's Claude Mythos 5 in agentic coding. On Terminal-Bench 2.1, Sol scores 88.8 percent. Sol Ultra hits 91.9, Claude Mythos 5 lands at 88 percent, and Fable 5 trails at 84.3. Sol also shows gains in biology. On GeneBench v1, a benchmark for genomics and quantitative biology, it beats GPT-5.5 (30 percent vs. 22 percent best case) while burning fewer tokens. Ad On ExploitBench , which tests how well AI agents can find and exploit real security flaws in Google's V8 JavaScript engine all the way to full code execution, Sol matches Mythos Preview's performance while using roughly a third of the output tokens, OpenAI says. On ExploitGym , a benchmark built by UC Berkeley researchers with OpenAI and other labs, all three GPT-5.6 models get better as reasoning effort goes up. That points to room for scaling with more compute. Claude numbers for this benchmark aren't available yet. Ad DEC_D_Incontent-2 OpenAI calls Sol its most capable cybersecurity model yet but frames it as a defender, not an attacker. The model is better at spotting and fixing flaws than at running full end-to-end attacks on its own, the company says. Mythos pulled that off in a different benchmark . Ad
中文翻译
OpenAI的新GPT-5.6代包括旗舰Sol和两个更便宜的层级Terra与Luna。Sol在基准测试中与Anthropic的Claude Mythos 5持平或超越,在代理编程方面明显领先,网络安全方面token效率更高。美国政府目前限制仅限指定合作伙伴使用。OpenAI表示该政策伤害了开发者和企业。
核心信息
OpenAI推出GPT-5.6系列,旗舰Sol在代理编程和网络安全方面领先Claude Mythos,但美国政府限制仅限合作方使用,OpenAI批评该政策损害开发者和企业利益。
- OpenAI推出GPT-5.6系列,旗舰Sol在代理编程和网络安全方面领先Claude Mythos,但美国政府限制仅限合作方使用,OpenAI批评该政策损害开发者和企业利益。
- 原贴提到:OpenAI's new GPT-5.6 generation includes the flagship Sol and two cheape
- 来源:the-decoder.com
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容