AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-09 0 浏览 会员

Grok 4.5 相比 Fable 5 和 GPT 5.5 如此便宜,基准测试差距可能没那么重要

xAI 发布 Grok 4.5,性能接近顶尖模型但价格远低于竞品,采用低价策略抢占市场。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-09 03:01:44

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

xAI has released Grok 4.5. The model was trained on tens of thousands of Nvidia GB300 GPUs and targets coding, agentic tasks, and knowledge work. Benchmark results paint a mixed picture. On Terminal Bench 2.1, which tests complex command-line tasks, Grok 4.5 scores 83.3%, nearly matching GPT 5.5 (83.4%) and trailing Anthropic's Fable 5 (84.3%) by just one point. But the gaps widen elsewhere. On DeepSWE 1.1, which measures the ability to resolve real GitHub issues, Grok 4.5 hits 53%, well behind OpenAI's GPT-5.5 at 67% and Fable 5 at 70%. On SWE Bench Pro, a curated set of harder software engineering problems, it scores 64.7%, beating Opus 4.8 (69.2% with max settings) in some configurations but falling short of Fable 5's 80.4%. Ad xAI says it relied on heavy data filtering, deduplication, and domain-specific selection during training to keep data quality high. The reinforcement learning stage covered hundreds of thousands of tasks, mostly from software engineering, with automated scoring. xAI built the training infrastructure for asynchronous learning, so agentic runs could stretch over many hours while training continued in parallel. Ad DEC_D_Incontent-1 Grok 4.5 costs $2 per million input tokens and $6 per million output tokens. That's already far below the competition. Opus 4.8 runs $5 input and $25 output per million tokens. Fable 5 charges $10 input and $50 output per million tokens. GPT-5.5 and GPT-5.6 sit at $5 input and $30 output . xAI also says Grok 4.5 uses 4.2 times fewer tokens than Opus 4.8 on SWE Bench Pro tasks and delivers results at 80 tokens per second. Lower per-token pricing and fewer tokens per task make Grok 4.5 by far the cheapest option in this performance tier, assuming the performance and efficiency gains hold up in practice. Ad The pricing strategy echoes what Chinese vendors like Zhipu and DeepSeek have been doing: get close enough on performance, then win on price. Grok 4.5 is available now through Grok Build, Cursor , and the xAI console . Plugins are live for Word , PowerPoint , and Excel . The model isn't available in the EU yet, with xAI targeting a mid-July launch. xAI trained Grok 4.5 alongside the code editor Cursor, which SpaceX acquired in mid-June for $60 billion in stock . Ad DEC_D_Incontent-2 Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

中文翻译

xAI 发布了 Grok 4.5。该模型在数万个英伟达 GB300 GPU 上训练,针对编程、代理任务和知识工作。基准测试结果喜忧参半。在测试复杂命令行任务的 Terminal Bench 2.1 上,Grok 4.5 得分 83.3%,几乎与 GPT 5.5(83.4%)持平,仅落后 Anthropic 的 Fable 5(84.3%)一个百分点。但在其他方面差距更大。在衡量解决真实 GitHub 问题能力的 DeepSWE 1.1 上,Grok 4.5 达到 53%,远低于 OpenAI 的 GPT-5.5(67%)和 Fable 5(70%)。在 SWE Bench Pro(一套精选的更难软件工程问题)上,其得分为 64.7%,在某些配置下击败了 Opus 4.8(最大设置下为 69.2%),但不及 Fable 5 的 80.4%。xAI 表示,在训练过程中依赖大量数据过滤、去重和领域特定选择来保持数据质量。强化学习阶段涵盖了数十万个任务,大多来自软件工程,并采用自动评分。xAI 构建了异步学习的训练基础设施,因此代理运行可以持续数小时,同时训练并行进行。Grok 4.5 每百万输入令牌收费 2 美元,每百万输出令牌收费 6 美元。这已经远低于竞争对手。Opus 4.8 每百万令牌输入 5 美元、输出 25 美元。Fable 5 每百万令牌输入 10 美元、输出 50 美元。GPT-5.5 和 GPT-5.6 分别为输入 5 美元、输出 30 美元。xAI 还表示,Grok 4.5 在 SWE Bench Pro 任务上使用的令牌数比 Opus 4.8 少 4.2 倍,并以每秒 80 令牌的速度交付结果。更低的单位定价和更少的每任务令牌数使得 Grok 4.5 成为该性能层级中最便宜的选择——前提是性能和效率提升在实践中成立。这一定价策略呼应了中国厂商如智谱和 DeepSeek 的做法:性能足够接近,然后以价格取胜。Grok 4.5 现已通过 Grok Build、Cursor 和 xAI 控制台提供。Word、PowerPoint 和 Excel 的插件也已上线。该模型尚未在欧盟可用,xAI 计划 7 月中旬推出。xAI 与代码编辑器 Cursor(SpaceX 于 6 月中旬以 600 亿美元股票收购)一同训练了 Grok 4.5。订阅 THE DECODER 以获取无广告阅读、每周 AI 新闻简报、每年六次的独家“AI 雷达”前沿报告、完整存档访问以及评论区的访问权限。

核心信息

xAI 发布 Grok 4.5,性能接近顶尖模型但价格远低于竞品,采用低价策略抢占市场。

  • xAI 发布 Grok 4.5,性能接近顶尖模型但价格远低于竞品,采用低价策略抢占市场。
  • 原贴提到:xAI has released Grok 4.5. The model was trained on tens of thousands of
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 通过MDM在VS Code和CLI中部署受管的Copilot设置 下一篇 GitHub Copilot 在 Visual Studio Code 中的更新——2026年6月发布