觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-09-27 1 浏览 免费阅读

AI 智能体在模型开发中承担更多工作,但人类仍做决策

一项来自复旦等团队的研究记录了人类与 AI 智能体协作开发 Atria Dawn Preview 模型的过程。AI 参与 96.5% 的任务,智能体动作与人类输入比从 11 升至 28.5,但人类仍主导方法、参数和目标等关键决策,最终决策占比在 85% 以上。约三分之一 AI 辅助任务若没有 AI 则无法完成,人类判断而非执行成为瓶颈。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-09-27 23:18:39

原贴

查看原文
作者:Jonathan Kemper 来源站点:the-decoder.com 原贴时间:

原文

A research team documented how humans and AI agents worked together to build a new AI model. The findings challenge some expectations about how independently agents can work. When AI agents help build new AI models, who makes the decisions? A team involving researchers from China's Fudan University studied its own project to find out. It analyzed more than 700 task logs from 56 participants, along with logs from the agents they used. The project centered on developing an agentic language model called Atria Dawn Preview , built on a mixture-of-experts architecture with 744 billion parameters and designed for research and engineering tasks. The model was trained through a pipeline that ties each task to a real execution environment. It calls tools, generates intermediate results, and gets checked against external signals like tests, metrics, or source evidence. The team says it leads on five of 16 benchmarks, including web search and cybersecurity, though it doesn't hold an overall edge over competitors. AI was used in 96.5 percent of the tasks reviewed. Over the course of the project, participants handed off more and more to agents. The median ratio of agent actions to human inputs rose from 11 to 28.5 over four weeks. The team cautions against reading this as growing autonomy. Each human decision led to more agent steps, which didn't mean the agents were making more decisions themselves. Participants were also asked whether they could have completed their share of a task without AI, at the same scope and quality. Of 455 completed AI-assisted tasks, 151 were rated infeasible without AI, roughly a third. These tasks were spread across 27 of the 56 participants, so they didn't come from just a handful of power users. AI didn't speed up existing work in these cases. It made work possible that would never have been started otherwise. For methods and parameters, the most common pattern was "AI proposes, human selects" at 55.4 percent. Overall, humans made 85.5 percent of decisions about methods and parameters, while AI made just 9.2 percent. Humans made the final decision on goals and scope in 93.4 percent of cases. AI's share of proposals ranged from 17 to 55 percent depending on the decision type. Its share of final decisions stayed in the single digits. Who proposed the options varied widely, but humans consistently made most of the final choices. Even among the 151 tasks rated infeasible without AI, humans chose the goal 95.4 percent of the time. The same pattern shows up when things go wrong. Of 588 tasks with a recorded difficulty, 76 percent moved forward through human intervention, and in 23 percent the agent solved the problem on its own. Human help almost always came in the form of information, either by adding context or clarifying requirements (35.2 percent) or by diagnosing issues and switching methods (34.7 percent). Humans rarely did the work themselves. Partial edits accounted for 3.2 percent of cases, and full takeovers just 0.7 percent. When AI outputs needed revision, the AI handled the changes itself 75.4 percent of the time after receiving human feedback. Human judgment, rather than execution, was the bottleneck. The team describes three phases in AI's role, from a subject of research to a tool for individual tasks and now a project partner. In that current role, AI drafts and adjusts plans within goals set by humans. A speculative fourth phase would involve recursive self-improvement, with stronger models producing stronger successors. The authors say a model can improve at its training tasks without getting better at developing its successor. How AI could propose varied research directions and assess their value before results are available remains an open question.

中文翻译

AI 智能体在模型开发中承担了更多工作,但人类仍然做出决策

一个研究团队记录了一个人类和 AI 智能体如何合作构建一个新 AI 模型的过程。这些发现挑战了关于智能体能够多么独立工作的一些预期。

当 AI 智能体帮助构建新 AI 模型时,谁做出决策?一个包括中国复旦大学研究人员的团队研究了自己的项目来找出答案。

核心信息

一项来自复旦等团队的研究记录了人类与 AI 智能体协作开发 Atria Dawn Preview 模型的过程。AI 参与 96.5% 的任务,智能体动作与人类输入比从 11 升至 28.5,但人类仍主导方法、参数和目标等关键决策,最终决策占比在 85% 以上。约三分之一 AI 辅助任务若没有 AI 则无法完成,人类判断而非执行成为瓶颈。

  • 一项来自复旦等团队的研究记录了人类与 AI 智能体协作开发 Atria Dawn Preview 模型的过程。AI 参与 96.5% 的任务,智能体动作与人类输入比从 11 升至 28.5,但人类仍主导方法、参数和目标等关键决策,最终决策占比在 85% 以上。约三分之一 AI 辅助任务若没有 AI 则无法完成,人类判断而非执行成为瓶颈。
  • 原贴提到:A research team documented how humans and AI agents worked together to b
  • 来源:the-decoder.com

详细解读

这是什么信号

这不是“AI 自己写模型”的宣言,而是一份来自复旦等团队的自研项目日志:他们开发了 Atria Dawn Preview,一个基于混合专家架构、7440 亿参数的 agentic 语言模型,并分析了 56 名参与者、700 多条任务日志以及智能体日志。AI 参与了 96.5% 的任务,智能体动作与人类输入的中位数比从 11 升至 28.5。表面看,agent 的“动作量”在四周内翻倍。

但关键信号是:动作量不等于决策权。人类仍做出 85.5% 的方法和参数决策,AI 仅占 9.2%;在目标和范围上,人类做出 93.4% 的最终决策。最常见的协作模式是“AI 提议、人类选择”,占 55.4%。AI 的提议占比在 17% 到 55% 之间,但最终决策占比始终是个位数。

为什么重要

它纠正了一个常见误读:agent 接手更多步骤,并不等于 agent 获得更多自主性。原文明确提醒,“每个类人决策导致更多智能体步骤”,这更像人类指挥下的执行放大,而不是目标自主。

更重要的发现是 AI 的价值形态。455 个完成的 AI 辅助任务中,151 个被参与者评为“没有 AI 则不可行”,约三分之一,且分布在 56 人中的 27 人,不是少数重度用户。AI 在这类任务中没有加速原有工作,而是让原本不会启动的工作成为可能。与此同时,遇到困难时,76% 的任务靠人类介入推进,但人类帮助主要是补充上下文、澄清需求或诊断问题并切换方法,亲自编辑只占 3.2%,完全接管仅 0.7%。人类判断而非执行,才是瓶颈。

对谁有价值

对 AI 实验室和工程团队:可用于设计人机协作流程,把人类放在目标、方法、参数选择和异常诊断环节,把 agent 放在执行、草拟和试错环节。对 agent 工具与平台开发者:需要支持“AI 提议、人类选择”的交互,记录任务日志、执行环境、外部验证信号和人类干预类型。对管理者:评估 AI 产出时,不能只看任务速度,还要看它是否解锁了原本不可行的任务。对个人:把 AI 当作扩大可行工作集合的协作者,而不是替代判断的决策者。

可以怎么行动

  • 在团队内区分“加速现有任务”和“解锁新任务”,分别记录 AI 辅助任务数量与不可行任务占比。
  • 把关键决策显式化:目标与范围由人类定,方法和参数采用“AI 提议、人类选择”,并保留最终决策记录。
  • 当 agent 卡住时,优先提供上下文、澄清需求或协助诊断,而不是直接接管执行。
  • 为 agent 建立真实执行环境和外部验证信号,如测试、指标或来源证据,减少无效动作。
  • 跟踪智能体动作与人类输入的比值,但不要把它当作自主性指标;同时跟踪人类最终决策占比。

风险与限制

这是单项目、单团队的研究,56 名参与者、约四周,结论不一定适用于所有组织或任务类型。Atria Dawn Preview 只在 16 项基准中的 5 项领先,包括网络搜索和网络安全,但整体并未超过竞争对手,因此不能把协作模式的成功等同于模型全面领先。文中提到的“递归自我改进”只是推测性第四阶段,作者也指出模型可以在训练任务上变强,却不一定会更擅长开发后继模型。AI 如何提出多样研究方向并在结果出现前评估价值,仍是开放问题。最后,智能体动作增加可能带来更多审查与验证成本,若外部信号不足,人类判断瓶颈会更严重。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《AI 智能体在模型开发中承担更多工作,但人类仍做决策》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

AI 智能体在模型开发中承担更多工作,但人类仍做决策主要讲什么?

一项来自复旦等团队的研究记录了人类与 AI 智能体协作开发 Atria Dawn Preview 模型的过程。AI 参与 96.5% 的任务,智能体动作与人类输入比从 11 升至 28.5,但人类仍主导方法、参数和目标等关键决策,最终决策占比在 85% 以上。约三分之一 AI 辅助任务若没有 AI 则无法完成,人类判断而非执行成为瓶颈。

这篇文章最值得关注的要点是什么?

一项来自复旦等团队的研究记录了人类与 AI 智能体协作开发 Atria Dawn Preview 模型的过程。AI 参与 96.5% 的任务,智能体动作与人类输入比从 11 升至 28.5,但人类仍主导方法、参数和目标等关键决策,最终决策…;原贴提到:A research team documented how humans and AI agents worked together to b;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在Agent工作流、AI副业、AI工具专题里阅读。 关联原因:这篇内容命中「Agent、智能体」等主题信号。;这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 已经是本栏目第一篇
下一篇 OpenAI 称 80% 到 90% 的研究已瞄准 GPT 7 及以后