AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-05-14 0 浏览 会员

趋势解读:Microsoft pits more than 100 AI agents against,提升开发者接入体验

微软推出MDASH,一个由超过100个AI代理组成的多模型系统,用于自动检测软件漏洞。该系统已在Windows中发现16个新漏洞,并在CyberGym基准测试中获得88.45%的分数,创下最高纪录。

SOURCE / AI技能杠杆 MIN / 9 ACCESS / 会员 POST / 2026-05-14 23:35:41

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:
趋势解读:Microsoft pits more than 100 AI agents against,提升开发者接入体验

原文

Microsoft has introduced MDASH, an AI-powered security system that uses more than 100 specialized agents to automatically detect software vulnerabilities. The system has already uncovered 16 new security vulnerabilities in Windows, four of them classified as critical. MDASH scored 88.45 percent on the CyberGym benchmark—the highest result to date—though Microsoft hasn't disclosed which specific AI models power the system. Microsoft has built an agentic multi-model system that uses more than 100 specialized AI agents to detect software vulnerabilities. The security system, called MDASH (Multi-Model Agentic Scanning Harness), is designed to automatically find security vulnerabilities in software. Unlike approaches that rely on a single AI model like Claude Mythos , MDASH orchestrates more than 100 specialized AI agents across an ensemble of frontier and distilled models, according to Microsoft. On Patch Tuesday, May 12, 2026, Microsoft reported 16 new vulnerabilities (CVEs) in the Windows networking and authentication stack that MDASH discovered. The company classifies four of these as critical, including remote code execution vulnerabilities in the tcpip.sys kernel component, the IKEv2 service (ikeext.dll), netlogon.dll, and dnsapi.dll. Ad Ten of the 16 vulnerabilities affect kernel mode, and most are accessible from the network without authentication, Microsoft says. The company points out that its own code base is especially hard to audit: Windows, Hyper-V, and Azure are proprietary and aren't part of public training data. Ad DEC_D_Incontent-1 The system works in a four-stage pipeline. First, it analyzes the source code and maps the attack surface. Specialized auditor agents then scan the code for suspicious areas. In the third stage, a second group of agents, which Microsoft calls "debaters," argue for and against the exploitability of each finding. Duplicates are then merged before Evidence Leader agents attempt to trigger the vulnerability through specific inputs in the final stage. The pipeline is model-agnostic: when a new model comes out, it can be tested against the previous one just by changing the configuration. Plugins let experts feed in domain-specific knowledge, like kernel calling conventions or IPC trust boundaries, that no foundation model knows on its own. Ad On the public CyberGym benchmark with 1,507 real vulnerabilities, the system scored 88.45 percent, the top result on the leaderboard, roughly five points ahead of the next best model. The comparison is misleading, though, since Microsoft is pitting an entire framework against individual models, which would also likely score higher if wrapped in a similar harness . The blog post doesn't reveal which models Microsoft used to achieve this score. The company only refers to "SOTA models" as heavy reasoners, "distilled models" as low-cost debaters, and a "second separate SOTA model" as an independent counterpart. Whether these come from OpenAI, Anthropic, Microsoft's own labs, or third-party providers remains unclear. Ad DEC_D_Incontent-2 MDASH is backed by Microsoft's Autonomous Code Security Team. Some of its members come from Team Atlanta, the winner of the DARPA AI Cyber Challenge , according to Microsoft. For that competition, the team built an autonomous cyber reasoning system that detected and fixed bugs in complex open-source projects. MDASH is currently available in a limited private preview for external customers. A detailed technical report is available on the Microsoft blog . Ad

中文翻译

微软推出了MDASH,这是一个由人工智能驱动的安全系统,使用超过100个专门的代理来自动检测软件漏洞。该系统已经发现了Windows中的16个新安全漏洞,其中四个被归类为严重。

核心信息

微软推出MDASH,一个由超过100个AI代理组成的多模型系统,用于自动检测软件漏洞。该系统已在Windows中发现16个新漏洞,并在CyberGym基准测试中获得88.45%的分数,创下最高纪录。

  • 微软推出MDASH,用100+AI代理自动检测软件漏洞。
  • 已发现16个Windows漏洞,其中4个为严重。
  • CyberGym基准88.45%分数,创最高纪录。
  • 系统采用四阶段流水线:分析、审计、辩论、触发。
  • 目前仅限私人预览,未公开所用模型。
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 趋势解读:x.AI plays catch-up with Grok Build,its first,解读最新 AI 进展 下一篇 趋势解读:开源工具html-anything助力Agent生成高质量HTML,解读最新 AI 进展