Knowledge File / AI小生意项目库
开源模型现已以更低成本匹配四个月前的前沿网络能力
英国AI安全研究所(AISI)首次公开评估发现,领先开源模型在网络能力上仅落后闭源系统4-7个月,差距较之前缩小,且成本更低,但安全防护易被绕过,给防御者带来新挑战。
SOURCE / AI小生意项目库
MIN / 9
ACCESS / 会员
POST / 2026-07-18 18:16:02
原贴
查看原文原文
An analysis by the British AI Security Institute (AISI) finds that open AI models now trail closed systems in cyber capabilities by four to seven months, down from six to ten months. In tests of specific cyber tasks and simulated networks, open models like GLM-5.2 and DeepSeek V4-Pro matched older closed systems. They're much cheaper to run, and their safeguards are easy to bypass because no one controls access. AISI warns that this leaves cyber defenders less time to prepare for new types of attacks. The British AI Security Institute (AISI) has, for the first time, publicly assessed how far leading open-weight AI models lag behind top proprietary systems in cyber capabilities. According to AISI , that gap is closing. Current open models like GLM-5.2 and DeepSeek V4-Pro have reached a level that closed frontier models hit four to seven months earlier. For most of 2025, the gap was still six to ten months. Critics see a risk in open models, whose weights anyone can download, modify, and run without oversight. Once a model is released, users can remove safety guardrails, share copies freely, and run it on private systems beyond anyone's control. AISI calls this "a persistent and irreversible risk of misuse." Ad But open-weight models also offer clear benefits. Users can host them privately with no data flowing back to providers , customize them, cut costs, and rely on a foundation that providers can't change or shut down. AISI says these competing concerns need to be balanced. Ad DEC_D_Incontent-1 AISI tested the models using two different methods. The "Narrow Cyber Tasks" benchmark includes 70 tasks across four difficulty levels, from nontechnical work to expert-level challenges. It covers vulnerability research, reverse engineering, web exploitation, and cryptography. GLM-5.2, released in June 2026, matched the performance of Opus 4.6 from February 2026 on these tasks. That puts it about four months behind. DeepSeek V4-Pro performed at the level of Opus 4.5, released in November 2025. Ad The second method, called Cyber Ranges , tests autonomous cyber capabilities in simulated networks. "The Last Ones" simulates a 32-step attack on a corporate network with four subnets and about 20 hosts. AISI estimates that a human expert would need roughly 20 hours to complete it. GLM-5.2 performed about as well as Opus 4.5 in this test, while DeepSeek V4-Pro fell below Sonnet 4.5. GPT-5.6-Sol posted the best result, ahead of Claude Mythos 5. Ad DEC_D_Incontent-2 The gap in Cyber Ranges is wider than in the Narrow Cyber Tasks, at around seven months. AISI treats the result as weaker evidence because it comes from fewer test scenarios. The tests also can't show whether a model fails because it lacks cyber capabilities or because it can't sustain planning across a long, complex attack. Ad
中文翻译
英国人工智能安全研究所(AISI)的一项分析发现,开放AI模型在网络能力上现在落后于封闭系统四到七个月,而此前为六到十个月。在特定网络任务和模拟网络的测试中,GLM-5.2和DeepSeek V4-Pro等开放模型匹配了较旧的封闭系统。它们的运行成本要低得多,并且由于无人控制访问,其安全防护很容易被绕过。AISI警告说,这留给网络防御者准备新型攻击的时间更少。AISI首次公开评估了领先的开放权重AI模型在网络能力上与顶级专有系统的差距。根据AISI,这一差距正在缩小。当前的开放模型如GLM-5.2和DeepSeek V4-Pro已经达到了封闭前沿模型四到七个月前的水平。在2025年的大部分时间里,差距仍有六到十个月。批评者认为开放模型存在风险,其权重任何人都可以下载、修改并无监督运行。一旦模型发布,用户就可以移除安全护栏,自由分享副本,并在无法控制的私有系统上运行。AISI称之为“持续且不可逆的滥用风险”。但开放权重模型也带来明显好处:用户可以私有化托管,无数据回流给提供商,进行定制,降低成本,并依赖提供商无法更改或关闭的基础。AISI表示需要平衡这些相互竞争的关注点。AISI使用两种不同方法测试模型。“狭义网络任务”基准包括四个难度级别的70个任务,从非技术工作到专家级挑战,涵盖漏洞研究、逆向工程、网络利用和密码学。2026年6月发布的GLM-5.2在这些任务上匹配了2026年2月的Opus 4.6的性能,落后约四个月。DeepSeek V4-Pro达到了2025年11月发布的Opus 4.5的水平。第二种方法称为“网络靶场”,在模拟网络中测试自主网络能力。“最后之人”模拟了对一个包含四个子网和约20台主机的企业网络的32步攻击。AISI估计人工专家需要约20小时完成。GLM-5.2在此测试中表现与Opus 4.5相当,而DeepSeek V4-Pro低于Sonnet 4.5。GPT-5.6-Sol取得最佳结果,领先于Claude Mythos 5。在网络靶场中差距更大,约为七个月。AISI认为这一结果证据较弱,因为测试场景较少。测试也无法显示模型失败是因为缺乏网络能力,还是因为无法在长期复杂攻击中维持规划。
核心信息
英国AI安全研究所(AISI)首次公开评估发现,领先开源模型在网络能力上仅落后闭源系统4-7个月,差距较之前缩小,且成本更低,但安全防护易被绕过,给防御者带来新挑战。
- 英国AI安全研究所(AISI)首次公开评估发现,领先开源模型在网络能力上仅落后闭源系统4-7个月,差距较之前缩小,且成本更低,但安全防护易被绕过,给防御者带来新挑战。
- 原贴提到:An analysis by the British AI Security Institute (AISI) finds that open
- 来源:the-decoder.com
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容