Knowledge File / AI技能杠杆
趋势解读:AI search agents often confirm what they already,评估 LLM Agent 表现
哈尔滨工业大学研究发现,领先的AI搜索agent在基准测试中主要依赖记忆而非实际搜索,新基准LiveBrowseComp暴露了其对新事件的无力,一旦无法依赖记忆,性能崩溃。
SOURCE / AI技能杠杆
MIN / 9
ACCESS / 会员
POST / 2026-05-31 15:48:41
原贴
查看原文原文
Leading AI search agents like GPT-5.4 and Kimi K2.6 don't appear to do much actual research on established benchmarks. They mostly just use the web to confirm what they already learned during training. Researchers at the Harbin Institute of Technology found this using a new time-based benchmark called LiveBrowseComp, which only asks about events from the last 90 days. Once the models can't fall back on memory, performance falls apart and the existing rankings get reshuffled. The article AI search agents often confirm what they already know instead of actually researching the web appeared first on The Decoder .
中文翻译
领先的AI搜索agent,如GPT-5.4和Kimi K2.6,在既定基准测试上似乎并没有进行太多实际研究。它们主要只是利用网络来确认在训练期间已经学到的东西。哈尔滨工业大学的研究人员使用一种名为LiveBrowseComp的基于时间的新基准发现了这一点,该基准只询问最近90天内发生的事件。一旦模型无法依赖记忆,性能就会崩溃,现有的排名也会被打乱。本文《AI搜索agent常常只是确认他们已知的信息,而不是真正进行网络研究》最初发表于The Decoder。
核心信息
哈尔滨工业大学研究发现,领先的AI搜索agent在基准测试中主要依赖记忆而非实际搜索,新基准LiveBrowseComp暴露了其对新事件的无力,一旦无法依赖记忆,性能崩溃。
- GPT-5.4等模型在基准测试中依赖记忆而非真正搜索
- 新基准LiveBrowseComp暴露模型对新事件的无能
- 无法依赖记忆时,模型性能崩溃,排名重排
- 当前评估方法误导了对AI搜索能力的判断
试看内容
成为会员查看完整内容
你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。
详细解读
信息差价值
参考来源
成为会员查看完整内容