觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-08-26 2 浏览 免费阅读

OpenAI首款自研芯片“Jalapeño”据称在推理基准测试中超越英伟达Blackwell和Rubin

OpenAI在Hot Chips大会上展示了其自研推理芯片Jalapeño的首批基准测试结果,据称在每瓦吞吐量和token延迟上优于英伟达的Blackwell和Rubin,但该芯片尚未量产,且公平对比下与Vera Rubin各有优劣。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-08-26 02:00:50

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

OpenAI showed off the first benchmarks for its in-house inference chip at the Hot Chips conference. "Jalapeño" reportedly outperforms both Nvidia's Blackwell and Rubin in throughput per watt and token latency. The chip handles inference only, meaning it runs AI models but doesn't train them. Jalapeño isn't tuned to OpenAI's own models either. It's a general-purpose LLM inference accelerator. OpenAI claims Jalapeño delivers 1.5x to 1.9x more AI work per watt at peak throughput across all three tested models, with 1.7x to 3.6x lower end-to-end latency than the best commercially available systems. For interactive workloads, the company says performance is 2.1x to 4.1x higher. Ad The results come from tests using SemiAnalysis's public InferenceX benchmark . OpenAI provided the numbers. SemiAnalysis verified some runs on-site in the lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T. On GPT-OSS, Jalapeño hit about 1,400 tokens per second per user. On Deepseek R1, it topped 700 tokens per second on a single concurrent request. Ad Jalapeño posted these numbers without using techniques like multi-token prediction or speculative decoding , while some of the comparison systems did rely on those optimizations, so there's still room to improve. In its headline performance-per-watt comparison, "Jalapeño smokes every other chip," SemiAnalysis writes . SemiAnalysis CEO Dylan Patel added , "Usually first generation chips aren't competitive, but OpenAI is beating Nvidia Blackwell and even Rubin." Ad SemiAnalysis points out that the fairer comparison isn't Blackwell but Nvidia's newer Vera Rubin platform , since both use HBM4 memory. Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, even though Nvidia's accelerator uses the multi-token prediction optimization that Jalapeño hasn't adopted yet. On total cost of ownership per token, the two come out roughly even. There are caveats, though. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that haven't been tested on Jalapeño yet. And while Rubin systems are already shipping to customers, Jalapeño reportedly hasn't moved beyond engineering samples. Ad OpenAI developed Jalapeño with Broadcom. Design work kicked off in mid-2024, and the final design went to fabrication in November 2025. The full cycle took about 16 months, but OpenAI says only nine months passed between the first chip design and the finished blueprint heading to the factory. The company used its own AI models during development, according to OpenAI . Older model generations helped with chip design, while newer ones sped up programming and optimization. Ad SemiAnalysis sees this as a sign that Nvidia's much-discussed "CUDA moat" may not hold anymore. "The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon," the firm wrote. OpenAI CFO Sarah Friar says the chip fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system. She claims Jalapeño complements OpenAI's existing partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and others rather than replacing them. OpenAI has deep ties with several of these companies. Nvidia , AMD , and AWS are all investors or compute partners, with Nvidia being one of the largest. Each of them is also building its own AI chips, which makes the relationship both cooperative and competitive. That said, all of these companies keep saying the world can't have enough compute , a claim that conveniently supports their own business models.

中文翻译

OpenAI 在 Hot Chips 大会上展示了其自研推理芯片的首批基准测试结果。据称,“Jalapeño”在每瓦吞吐量和 token 延迟方面均优于英伟达的 Blackwell 和 Rubin。该芯片仅处理推理,即运行 AI 模型但不进行训练。Jalapeño 也未针对 OpenAI 自身的模型进行调优。它是一个通用的 LLM 推理加速器。OpenAI 声称,在三个测试模型上,Jalapeño 在峰值吞吐量下每瓦特可完成 1.5 倍至 1.9 倍的 AI 工作量,端到端延迟比最佳商用系统低 1.7 倍至 3.6 倍。对于交互式工作负载,该公司表示性能高出 2.1 倍至 4.1 倍。测试结果使用了 SemiAnalysis 的公开 InferenceX 基准。OpenAI 提供了数据,SemiAnalysis 在实验室现场验证了部分运行。测试模型包括 GPT-OSS 120B、Deepseek R1 670B 和 Kimi K2.5 1T。在 GPT-OSS 上,Jalapeño 每个用户每秒约处理 1400 个 token。在 Deepseek R1 上,单并发请求时每秒超过 700 个 token。Jalapeño 在未使用多 token 预测或投机解码等技术的情况下取得了这些数字,而一些对比系统确实依赖这些优化,因此仍有改进空间。在标题级的每瓦性能对比中,SemiAnalysis 写道:“Jalapeño 碾压其他所有芯片。”SemiAnalysis 首席执行官 Dylan Patel 补充说:“通常第一代芯片不具备竞争力,但 OpenAI 正在击败英伟达 Blackwell,甚至 Rubin。”SemiAnalysis 指出,更公平的比较对象不是 Blackwell,而是英伟达更新的 Vera Rubin 平台,因为两者都使用 HBM4 内存。即便如此,Jalapeño 每兆瓦输出的 token 数仍超过 Vera Rubin,尽管英伟达的加速器使用了 Jalapeño 尚未采用的多 token 预测优化。按每 token 的总拥有成本计算,两者大致相当。不过也有注意事项。英伟达和 AMD 已经发布了更大模型(如 Deepseek V4 Pro 和 Kimi K3)的结果,这些模型尚未在 Jalapeño 上测试。此外,虽然 Rubin 系统已开始向客户发货,但据报道 Jalapeño 仍处于工程样品阶段。OpenAI 与博通合作开发了 Jalapeño。设计工作于 2024 年年中启动,最终设计于 2025 年 11 月提交制造。整个周期约 16 个月,但 OpenAI 表示从首次芯片设计到完成蓝图进厂仅用了 9 个月。据 OpenAI 称,公司在开发过程中使用了自家的 AI 模型。旧一代模型帮助进行芯片设计,新一代模型加速了编程和优化。SemiAnalysis 认为,这标志着英伟达备受讨论的“CUDA 护城河”可能不复存在。该公司写道:“鉴于 OpenAI 能如此快速地在自研芯片上运行新模型,CUDA 护城河可能已经死亡。”OpenAI 首席财务官 Sarah Friar 表示,该芯片契合更广泛的算力战略,数据中心、芯片、模型、开发者平台、产品和设备作为一个集成系统协同工作。她声称 Jalapeño 补充而非取代了 OpenAI 与英伟达、AMD、AWS、Cerebras、CoreWeave 等公司的现有合作伙伴关系。OpenAI 与其中多家公司关系深厚。英伟达、AMD 和 AWS 都是投资者或算力合作伙伴,其中英伟达是最大的之一。它们也都在构建自己的 AI 芯片,这使得关系既合作又竞争。尽管如此,这些公司都不断表示世界不可能拥有足够的计算能力,这一说法恰好支持了它们自身的商业模式。

核心信息

OpenAI在Hot Chips大会上展示了其自研推理芯片Jalapeño的首批基准测试结果,据称在每瓦吞吐量和token延迟上优于英伟达的Blackwell和Rubin,但该芯片尚未量产,且公平对比下与Vera Rubin各有优劣。

  • OpenAI在Hot Chips大会上展示了其自研推理芯片Jalapeño的首批基准测试结果,据称在每瓦吞吐量和token延迟上优于英伟达的Blackwell和Rubin,但该芯片尚未量产,且公平对比下与Vera Rubin各有优劣。
  • 原贴提到:OpenAI showed off the first benchmarks for its in-house inference chip a
  • 来源:the-decoder.com

详细解读

这是什么信号?OpenAI 自研推理芯片 Jalapeño 的首次公开基准测试,标志着 AI 巨头开始自研芯片并试图挑战英伟达的统治地位。该芯片基于博通合作,设计周期短,且在未使用主流优化技术的情况下,性能已超越现有标杆,显示 OpenAI 在算力层面的战略野心。

为什么重要?英伟达的 CUDA 生态长期被视为难以逾越的护城河,但 OpenAI 的自研芯片证明,对于大型 AI 公司而言,快速定制硅片以匹配自家模型已成为可能。其次,Jalapeño 的能效比优势直指 AI 推理成本痛点,可能颠覆当前算力市场的定价逻辑。此外,这一进展发生在全球算力紧缺背景下,意味着算力供应链的格局正在重塑。

对谁有价值?对 AI 基础设施供应商(如云厂商、芯片初创公司)而言,这是竞争加剧的信号;对依赖 GPU 做推理的企业用户而言,未来可能出现更廉价的算力选择;对投资者而言,需要重新评估英伟达的长期护城河,并关注博通等定制芯片伙伴的受益机会。

可以怎么行动?企业可关注 OpenAI 或博通生态的算力接入机会,例如通过云服务获取基于 Jalapeño 的推理能力。技术团队可评估其是否适用于大规模 LLM 推理,并比较与现有 GPU 方案的总拥有成本。投资者可跟踪量产进展及应用落地情况,但需注意当前数据由 OpenAI 自行提供,存在夸大可能。

风险或限制Jalapeño 仍处于工程样品阶段,尚未量产,也没有客户部署验证。基准测试由 OpenAI 提供,SemiAnalysis 仅现场验证部分运行,存在利益冲突。此外,对比系统可能使用了不同优化,而英伟达新款芯片也已发布更优结果,后续竞争力待考。最后,OpenAI 与多家算力伙伴既有合作又有竞争,合作关系存在不确定性。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《OpenAI首款自研芯片“Jalapeño”据称在推理基准测试中超越英伟达Blackwell和Rubin》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

OpenAI首款自研芯片“Jalapeño”据称在推理基准测试中超越英伟达Blackwell和Rubin主要讲什么?

OpenAI在Hot Chips大会上展示了其自研推理芯片Jalapeño的首批基准测试结果,据称在每瓦吞吐量和token延迟上优于英伟达的Blackwell和Rubin,但该芯片尚未量产,且公平对比下与Vera Rubin各有优劣。

这篇文章最值得关注的要点是什么?

OpenAI在Hot Chips大会上展示了其自研推理芯片Jalapeño的首批基准测试结果,据称在每瓦吞吐量和token延迟上优于英伟达的Blackwell和Rubin,但该芯片尚未量产,且公平对比下与Vera Rubin各有优劣。;原贴提到:OpenAI showed off the first benchmarks for its in-house inference chip a;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解模型、副业、创业、项目、小生意这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 Claude 记忆功能全面打通聊天与 Cowork,用户可逐条查看和编辑 下一篇 皮尤研究证实自ChatGPT发布以来网络上AI生成文本急剧增加