AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-17 0 浏览 会员

Kimi的开放模型K3接近GPT-5.6 Sol和Fable 5,同时标志着超便宜中文AI的终结

Kimi发布基于混合专家架构的2.8万亿参数多模态开源模型K3,性能接近顶尖闭源模型,但定价更高且幻觉率上升。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-17 03:49:39

原贴

查看原文
作者:Matthias Bastian 来源站点:the-decoder.com 原贴时间:

原文

Kimi has released K3, a multimodal open-weight model built on a mixture-of-experts architecture with 896 experts, 2.8 trillion parameters, and a context window of one million tokens. Full weights are expected by the end of July. In Kimi's own benchmarks, K3 comes close to Claude Fable 5 and GPT 5.6 Sol but beats all other tested systems by a wide margin. Independent testing by Artificial Analysis largely confirms these results, though K3's hallucination rate increased compared to its predecessor. At $3 per million input tokens and $15 per million output tokens, K3 is much pricier than its predecessor but comparable to Western mid-range models like Sonnet 5. Per-task costs land around $0.94, similar to GPT-5.6 Sol and about half the price of Opus 4.8. Kimi is launching K3, a multimodal model with 2.8 trillion parameters and a context window of one million tokens. In the company's own benchmarks, it performs on par with leading proprietary models. According to Kimi, the new flagship model K3 has 2.8 trillion total parameters, processes images and video natively, and supports a context window of one million tokens. Kimi calls K3 the first open model in the roughly 3 trillion parameter range. Full model weights are scheduled for release by July 27. The model targets long-running programming tasks, knowledge work, and complex reasoning. In Kimi's own benchmarks , K3 still trails the top proprietary models Claude Fable 5 and GPT 5.6 Sol but beats every other system tested, including the Claude Opus models and Chinese rival GLM-5.2. All results come from Kimi and were achieved at maximum or high thinking intensity, according to the company. Ad Across all 35 tests, K3 took first place about seven times and landed second or third in most of the rest. Fable 5 won the most individual tests. In nearly every benchmark, K3 beat Opus 4.8, GPT 5.5, and GLM 5.2 by a wide margin. Depending on the benchmark, one of three agent systems was used: KimiCode, Claude Code, or Codex. That means the results weren't all collected under identical conditions. Ad DEC_D_Incontent-1 Independent testing lab Artificial Analysis has published its first evaluation of Kimi K3. The model scores 57 on the Artificial Analysis Intelligence Index, putting it on par with Opus 4.8 and GPT-5.5 but still behind Fable 5 and GPT-5.6 Sol. That largely lines up with Kimi's own claims. On agentic tasks, K3 reaches an Elo rating of 1,668 on GDPval v2, a big jump from K2.6's 1,190. It beats GLM-5.2 (1,514), GPT-5.5 (1,494), and Claude Opus 4.8 (1,600), though it still falls short of Claude Fable 5 (1,760). K3 also takes the top spot on AutomationBench-AA, Artificial Analysis's version of Zapier's agentic SaaS workflow evaluation, with a score of 53 percent. Ad On AA-Briefcase, a private long-horizon knowledge work evaluation, K3 reaches an overall Elo of 1,547, up 732 points from K2.6. Only Claude Fable 5 scores higher. Artificial Analysis calls K3 well-rounded, with rubric scoring and analytical quality close to Fable 5's level. GPT-5.6 Sol still leads on presentation quality, though. K3's accuracy rate improved from 33 percent to 46 percent on the AA-Omniscience Index, pushing the overall score from +6 to +18. But its hallucination rate climbed from 39 percent to 51 percent, meaning K3 fabricates more answers even as it gets more questions right. Ad DEC_D_Incontent-2 According to Kimi, the model's primary use case is long-running software development with minimal human oversight. K3 is built to analyze large codebases, coordinate terminal tools, and stay focused on a task across many work steps. Ad

中文翻译

Kimi发布了K3,这是一个基于混合专家架构的多模态开放权重模型,拥有896个专家、2.8万亿参数和100万token的上下文窗口。完整权重预计于7月底发布。在Kimi自己的基准测试中,K3接近Claude Fable 5和GPT 5.6 Sol,但大幅超越所有其他测试系统。Artificial Analysis的独立测试基本证实了这些结果,不过K3的幻觉率相比前代有所上升。K3每百万输入token收费3美元,每百万输出token收费15美元,比前代贵得多,但与Sonnet 5等西方中端模型相当。每任务成本约为0.94美元,与GPT-5.6 Sol相近,约为Opus 4.8的一半。Kimi正在推出K3,这是一个具有2.8万亿参数和100万token上下文窗口的多模态模型。在该公司自己的基准测试中,其性能与领先的专有模型持平。据Kimi称,全新旗舰模型K3总参数为2.8万亿,原生处理图像和视频,并支持100万token的上下文窗口。Kimi称K3是首个约3万亿参数范围内的开放模型。完整模型权重计划于7月27日前发布。该模型面向长时间运行的编程任务、知识工作和复杂推理。在Kimi自己的基准测试中,K3仍然落后于顶级专有模型Claude Fable 5和GPT 5.6 Sol,但击败了所有其他测试系统,包括Claude Opus模型和国内竞争对手GLM-5.2。所有结果均来自Kimi,据该公司称,这些结果是在最大或高思考强度下取得的。在所有35项测试中,K3大约有7次获得第一名,其余大部分测试获得第二或第三名。Fable 5赢得了最多单项测试。在几乎所有基准测试中,K3大幅击败了Opus 4.8、GPT 5.5和GLM 5.2。根据基准测试的不同,使用了三种代理系统之一:KimiCode、Claude Code或Codex。这意味着结果并非都在相同条件下收集。Ad DEC_D_Incontent-1 独立测试实验室Artificial Analysis发布了其对Kimi K3的首次评估。该模型在Artificial Analysis Intelligence Index上得分为57,与Opus 4.8和GPT-5.5持平,但仍落后于Fable 5和GPT-5.6 Sol。这与Kimi自身的说法大致相符。在代理任务上,K3在GDPval v2上达到了1,668的Elo评分,较K2.6的1,190大幅跃升。它击败了GLM-5.2(1,514)、GPT-5.5(1,494)和Claude Opus 4.8(1,600),但仍不及Claude Fable 5(1,760)。K3还在Artificial Analysis版本的Zapier代理SaaS工作流评估AutomationBench-AA上以53%的得分位居榜首。在私人长期知识工作评估AA-Briefcase上,K3的总Elo为1,547,较K2.6提升了732分。只有Claude Fable 5得分更高。Artificial Analysis称K3全面均衡,评分量规和分析质量接近Fable 5的水平。不过GPT-5.6 Sol在展示质量上仍领先。K3在AA-Omniscience Index上的准确率从33%提升至46%,总得分从+6提升至+18。但它的幻觉率从39%上升至51%,意味着K3在答对更多问题的同时也编造了更多答案。Ad DEC_D_Incontent-2 据Kimi称,该模型的主要用例是长时间运行的软件开发,人工监督极少。K3旨在分析大型代码库、协调终端工具,并在多个工作步骤中保持专注。

核心信息

Kimi发布基于混合专家架构的2.8万亿参数多模态开源模型K3,性能接近顶尖闭源模型,但定价更高且幻觉率上升。

  • Kimi发布基于混合专家架构的2.8万亿参数多模态开源模型K3,性能接近顶尖闭源模型,但定价更高且幻觉率上升。
  • 原贴提到:Kimi has released K3, a multimodal open-weight model built on a mixture-
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Grok 推出 Automations 功能:定时或邮件触发,自动执行任务并汇报结果 下一篇 GitHub Projects 高级搜索现已全面可用