AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-06-28 0 浏览 会员

在500天创业生存测试中,仅有三款AI模型在起始资本之上完成

普林斯顿大学构建CEO-Bench测试,模拟AI代理运营公司500天,发现多数模型破产,仅少数盈利,凸显AI在长期战略决策上的不足。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-06-28 18:16:13

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

Researchers at Princeton University built CEO-Bench, a test where AI agents have to run a fictional software company for 500 simulated days. Most current models go broke, and a simple rule-based heuristic with no AI beats nearly all of them. AI agents are getting increasingly good at narrow tasks: fixing a bug, following a service policy in a conversation, or completing a web-based workflow. These tasks share a simple structure, according to the Princeton study: the agent gets a clear goal, acts briefly, and receives quick feedback. Many important real-world tasks look nothing like that. They involve long chains of decisions under uncertainty, where you have to set priorities, allocate limited resources, read noisy signals, and adapt to changing conditions. To test exactly these skills, the researchers developed CEO-Bench . The benchmark simulates a realistic example of this kind of long-horizon task: running a startup for 500 simulated days. The researchers point to a famous example: in 1997, Apple was 90 days from bankruptcy. Steve Jobs drew a simple two-by-two grid—consumer and pro, desktop and portable—and decided Apple would only build products for those four quadrants. The iMac, iPod, and iPhone followed. This type of strategic steering intelligence is fundamentally different from what AI agents do today, the authors argue. Agents are getting better at individual tasks fast. But steering an entire organization toward long-term goals? That's a different problem entirely. CEO-Bench is a first attempt at measuring exactly this "steering intelligence." In CEO-Bench, an agent runs a made-up subscription software company called NovaMind. It starts with zero customers and one million dollars in the bank. Performance is measured by remaining cash at the end. If the balance drops below zero even once, the company is bankrupt and the simulation ends. The agent controls the company through a Python API with 34 tools and a database of 19 tables. Instead of just issuing individual commands, it writes its own code, queries the database with SQL, and builds custom workflows from the results. That puts it in front of the same challenges a human CEO would face, the researchers say. There's a lot to decide: pricing and tiers, ad spend across channels, product quality and R&D, infrastructure capacity and customer support, plus multi-round negotiations with enterprise clients. On top of that, there's a simulated social network where the agent can read complaints, competitor news, and economic trends and post itself. What makes the task hard is time and uncertainty. Decisions play out on realistic business timelines: revenue only arrives at billing dates, R&D projects take days to weeks, and mistakes often don't show up until later through churn or damaged reputation. Costs hit right away. The agent has to spend money whose payoff might not show up for weeks. Much of the company's state stays hidden. The agent can't directly see customer satisfaction, willingness to pay, or minimum quality expectations. It has to piece these together from noisy signals like cancellations, support tickets, or reactions on the social network. The simulation models 26 customer segments and individual customers, each with their own budgets, price sensitivities, and expectations. The world keeps changing, too. Competitors periodically raise customer quality expectations, preferences shift over time, and a simulated business cycle affects demand and willingness to pay, so the agent has to keep adjusting. The researchers deliberately chose fixed, transparent rules rather than a language model as referee. They wanted to avoid a weakness they see in Vending-Bench , a test with a simulated vending machine : there, an AI-simulated supplier can reward an agent for unrealistic verbal promises.

中文翻译

普林斯顿大学的研究人员构建了CEO-Bench测试,AI代理需要在模拟环境中运营一家虚构软件公司500天。大多数当前模型都破产了,而一个简单的、基于规则的启发式方法(无AI)几乎击败了所有模型。

核心信息

普林斯顿大学构建CEO-Bench测试,模拟AI代理运营公司500天,发现多数模型破产,仅少数盈利,凸显AI在长期战略决策上的不足。

  • 普林斯顿大学构建CEO-Bench测试,模拟AI代理运营公司500天,发现多数模型破产,仅少数盈利,凸显AI在长期战略决策上的不足。
  • 原贴提到:Researchers at Princeton University built CEO-Bench, a test where AI age
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Grok 4.5 私测于 SpaceX 和 Tesla,性能接近 Opus 下一篇 新浪开源VibeThinker-3B:推理可压缩,事实知识不能