Knowledge File / AI技能杠杆
论文速读:Turing Test on Screen,聚焦形式化数学证明能力
本文提出“屏幕上的图灵测试”概念,将GUI代理与检测器的交互建模为最小化行为差异的MinMax优化问题,建立代理人性化基准(AHB),并通过启发式噪声和数据驱动行为匹配实现高可模仿性而不牺牲性能,推动代理从“能否完成任务”转向“如何以人性化方式执行”。
SOURCE / AI技能杠杆
MIN / 4
ACCESS / 会员
POST / 2026-04-14 12:55:10
原贴
查看原文
原文
arXiv:2604.09574v1 Announce Type: new Abstract: The rise of autonomous GUI agents has triggered adversarial countermeasures from digital platforms, yet existing research prioritizes utility and robustness over the critical dimension of anti-detection. We argue that for agents to survive in human-centric ecosystems, they must evolve Humanization capabilities. We introduce the ``Turing Test on Screen,'' formally modeling the interaction as a MinMax optimization problem between a detector and an agent aiming to minimize behavioral divergence. We then collect a new high-fidelity dataset of mobile touch dynamics, and conduct our analysis that vanilla LMM-based agents are easily detectable due to unnatural kinematics. Consequently, we establish the Agent Humanization Benchmark (AHB) and detection metrics to quantify the trade-off between imitability and utility. Finally, we propose methods ranging from heuristic noise to data-driven behavioral matching, demonstrating that agents can achieve high imitability theoretically and empirically without sacrificing performance. This work shifts the paradigm from whether an agent can perform a task to how it performs it within a human-centric ecosystem, laying the groundwork for seamless coexistence in adversarial digital environments.
中文翻译
arXiv:2604.09574v1 公告类型:新提交。摘要:自主GUI代理的兴起引发了数字平台的对抗性反制措施,但现有研究优先考虑实用性和鲁棒性,忽略了反检测这一关键维度。我们认为,为了使代理在以人为本的生态系统中生存,它们必须进化出人性化能力。我们引入“屏幕上的图灵测试”,将交互正式建模为检测器与旨在最小化行为差异的代理之间的MinMax优化问题。然后我们收集了一个新的高保真移动触摸动态数据集,并进行分析,发现基于普通LMM的代理由于不自然的运动学特性容易被检测。因此,我们建立了代理人性化基准(AHB)和检测指标,以量化可模仿性与实用性之间的权衡。最后,我们提出了从启发式噪声到数据驱动行为匹配的方法,证明代理在理论上和实际上都能在不牺牲性能的情况下实现高可模仿性。这项工作将范式从代理是否能执行任务转变为它如何在以人为本的生态系统中执行任务,为在对抗性数字环境中无缝共存奠定基础。
核心信息
本文提出“屏幕上的图灵测试”概念,将GUI代理与检测器的交互建模为最小化行为差异的MinMax优化问题,建立代理人性化基准(AHB),并通过启发式噪声和数据驱动行为匹配实现高可模仿性而不牺牲性能,推动代理从“能否完成任务”转向“如何以人性化方式执行”。
- 提出屏幕图灵测试,建模代理与检测器对抗博弈
- 建立代理人性化基准AHB,量化可模仿性与实用性权衡
- 发现普通LMM代理因运动不自然易被检测
- 启发式噪声与数据驱动匹配可提升拟真度不损性能
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容