AI觉醒星球
Awakening is here
Knowledge File / AI技能杠杆
2026-04-13 0 浏览 会员

论文速读:SAGE,聚焦形式化数学证明能力

本文介绍SAGE,一个用于评估大型语言模型在客服场景中遵循标准化操作流程(SOPs)能力的多智能体基准测试,揭示了执行差距和共情韧性现象。

SOURCE / AI技能杠杆 MIN / 4 ACCESS / 会员 POST / 2026-04-13 12:13:05

原贴

查看原文
作者:arXiv cs.AI 来源站点:arxiv.org 原贴时间:
论文速读:SAGE,聚焦形式化数学证明能力

原文

arXiv:2604.09285v1 Announce Type: new Abstract: The development of Large Language Models (LLMs) has catalyzed automation in customer service, yet benchmarking their performance remains challenging. Existing benchmarks predominantly rely on static paradigms and single-dimensional metrics, failing to account for diverse user behaviors or the strict adherence to structured Standard Operating Procedures (SOPs) required in real-world deployments. To bridge this gap, we propose SAGE (Service Agent Graph-guided Evaluation), a universal multi-agent benchmark for automated, dual-axis assessment. SAGE formalizes unstructured SOPs into Dynamic Dialogue Graphs, enabling precise verification of logical compliance and comprehensive path coverage. We introduce an Adversarial Intent Taxonomy and a modular Extension Mechanism, enabling low-cost deployment across domains and facilitating automated dialogue data synthesis. Evaluation is conducted via a framework where Judge Agents and a Rule Engine analyze interactions between User and Service Agents to generate deterministic ground truth. Extensive experiments on 27 LLMs across 6 industrial scenarios reveal a significant ``Execution Gap'' where models accurately classify intents but fail to derive correct subsequent actions. We also observe ``Empathy Resilience'', a phenomenon where models maintain polite conversational facades despite underlying logical failures under high adversarial intensity. Code and resources are available at https://anonymous.4open.science/r/SAGE-Bench-4CD3/.

中文翻译

大型语言模型(LLMs)的发展促进了客服自动化的进步,但对其性能进行基准测试仍然具有挑战性。现有的基准测试主要依赖静态范式和单维度指标,未能考虑到多样化的用户行为或实际部署中所需严格遵守的结构化标准操作流程(SOPs)。为弥补这一差距,我们提出SAGE(服务代理图引导评估),一个用于自动化双轴评估的通用多智能体基准测试。SAGE将非结构化的SOPs形式化为动态对话图,从而能够精确验证逻辑合规性和实现全面的路径覆盖。我们引入了一种对抗性意图分类法和模块化扩展机制,实现了跨领域的低成本部署,并促进了自动化对话数据合成。评估通过一个框架进行,其中裁判代理和规则引擎分析用户代理与服务代理之间的交互,以生成确定性真实数据。在6个工业场景下对27个LLMs进行的大量实验揭示了一个显著的"执行差距":模型能够准确分类意图,但无法推导出正确的后续动作。我们还观察到"共情韧性"现象,即在高对抗强度下,模型尽管存在潜在逻辑失败,仍能保持礼貌的对话外表。代码和资源可在https://anonymous.4open.science/r/SAGE-Bench-4CD3/获取。

核心信息

本文介绍SAGE,一个用于评估大型语言模型在客服场景中遵循标准化操作流程(SOPs)能力的多智能体基准测试,揭示了执行差距和共情韧性现象。

  • SAGE将SOPs转化为动态对话图进行逻辑合规验证。
  • 发现执行差距:意图分类正确但后续动作错误。
  • 共情韧性:礼貌掩盖逻辑失败,高对抗下仍存在。
  • 支持低成本跨领域部署和自动数据合成。
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 论文速读:DRBENCHER,聚焦形式化数学证明能力 下一篇 论文速读:Constraint-Aware Corrective Memory for Language-Based Drug Discovery Agents