觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-03 5 浏览 免费阅读

GPT和Claude在桥水基金的金融测试中失败,因为正确答案从未公开

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-07-03 19:16:42

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

Bridgewater and Thinking Machines Lab have trained an open-source AI model for analyzing financial documents that outperforms leading commercial models. The Qwen3-235B model, which has been fine-tuned using internal expert knowledge, achieves nearly 85 percent accuracy in tests and is 14 times cheaper to operate. This demonstrates that companies can develop powerful AI solutions using their own data without having to share sensitive information with large providers. Hedge fund Bridgewater and Thinking Machines Lab say a fine-tuned open-weight model outperforms the strongest AI models at evaluating financial documents, at a fraction of the cost. The numbers come from their own internal evaluation. Investors get buried in news, analysis, corporate filings, and emails every day. According to a report from Bridgewater's AIA Labs and Thinking Machines Lab , the startup founded by former OpenAI CTO Mira Murati, reading isn't the real work. The real work is the constant stream of small, repeated judgment calls about what actually matters. That's the triage the researchers wanted to automate. They defined six tasks drawn from an investor's daily routine. One example: deciding whether a financial article is relevant to an executive. Another: whether a central bank document signals the direction of future rate changes. For investors, these calls are trivial, but they can barely put their reasoning into words. The report gives a telling example. A headline about Trump's claim to Greenland gets flagged as irrelevant, while Trump's threat of new China tariffs is highly relevant. Both touch on geopolitics and finance. Ad Frontier models failed in the authors' tests. Variants of Gemini, Claude, and GPT hit only about 50 percent accuracy with a basic prompt. Expert-written instructions and a three-tier rating system ("relevant and interesting," "relevant but uninteresting," "irrelevant") pushed accuracy into the mid-70s. That still fell short of the 80 percent threshold the authors set for trustworthy deployment. Ad DEC_D_Incontent-1 Newer models barely improve per dollar, the report says. GPT 5.4 costs 43 percent more than 5.2 but is only marginally more accurate. The solution was fine-tuning, retraining an open-weight model on proprietary examples. The key ingredient was the Bridgewater investors' judgment: At first, cheap outside contractors labeled the documents, but many of those labels were wrong. To avoid having expensive professionals review everything, the researchers used a workaround. A first model learned from the flawed labels and re-evaluated the same documents. Wherever the model and the original label disagreed, there was likely an error. Only those disputed cases went to investors for correction. Ad Training ran on the Tinker platform from Thinking Machines Lab, built on top of the open model Qwen3-235B. In the team's own evaluation, the fine-tuned model hit 84.7 percent accuracy versus 78.2 percent for the best frontier model tested. It also cost nearly 14 times less to run. This isn't a truly independent comparison, of course. Both companies have a clear interest in selling their product. Still, the finding beyond the numbers is worth noting. It shows once again that the big labs like OpenAI haven't absorbed all the data out there. Huge pools of proprietary corporate data and untrained human expertise still exist, and they hold real room for improvement. That's especially true where companies deliberately keep their most valuable data private. Anyone who hands that data to a frontier lab risks competing against a product built on top of it. Ad DEC_D_Incontent-2 Fine-tuning open models through tools like Tinker gives companies an alternative. They keep the weights, the data, and, depending on the setup, the GPUs themselves. Ad

中文翻译

桥水基金和Thinking Machines Lab训练了一个用于分析金融文档的开源AI模型,其表现优于领先的商业模型。使用内部专家知识微调的Qwen3-235B模型在测试中达到近85%的准确率,运行成本低14倍。这表明公司可以利用自己的数据开发强大的AI解决方案,而无需与大型提供商分享敏感信息。

核心信息

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。

  • 桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。
  • 原贴提到:Bridgewater and Thinking Machines Lab have trained an open-source AI mod
  • 来源:the-decoder.com

详细解读

这是什么信号?桥水基金和Thinking Machines Lab的联合实验表明,通过微调开源模型(如Qwen3-235B),企业可以在特定领域(如金融文档分析)上以极低成本超越GPT、Claude等通用前沿模型。这打破了“只有大模型厂商才能提供最佳AI”的固有认知,强调专有数据和人类专家知识的关键作用。

为什么重要?金融领域每天处理海量信息,投资者需要快速甄别真正相关的内容。传统上,依赖通用模型效果不佳(准确率仅约50%),而专家标注成本高昂。桥水的方法利用内部专家判断对开源模型进行微调,不仅提升了准确率(84.7% vs 前沿模型78.2%),还将运行成本降低近14倍,揭示了“数据+微调”模式在垂直场景的巨大潜力。

对谁有价值?首先,对冲基金、投资银行等金融机构可借鉴此方法,利用自身数据构建私有AI,避免与大型AI提供商共享敏感信息。其次,拥有专有数据和行业经验的企业(如律所、咨询公司、制药企业)可将微调作为差异化工具。此外,开源模型生态(如阿里巴巴的Qwen系列)获得验证,吸引更多企业投入。

可以怎么行动?1)企业应盘点内部高价值数据和专家经验,作为微调的“燃料”。2)选择合适的基础模型(如Qwen3-235B)和微调平台(如Tinker),小范围试点。3)建立内部标注与模型迭代流程,参考桥水的“争议点优先”策略,降低人工成本。4)关注模型部署和隐私保护,确保数据不外泄。

风险或限制1)实验非独立第三方评估,桥水与Thinking Machines Lab有商业利益,结果可能偏向利好,需谨慎看待具体数字。2)微调依赖高质量专家标注,初期成本高,且专家判断的隐含知识难以完全编码。3)开源模型本身存在漏洞或偏见,微调可能放大特定风险。4)通用任务(如创造性分析)可能仍需前沿模型,此方法适用于规则性强、高频重复的判断场景。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《GPT和Claude在桥水基金的金融测试中失败,因为正确答案从未公开》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

GPT和Claude在桥水基金的金融测试中失败,因为正确答案从未公开主要讲什么?

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。

这篇文章最值得关注的要点是什么?

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。;原贴提到:Bridgewater and Thinking Machines Lab have trained an open-source AI mod;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型、Claude」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 Google DeepMind 与 A24 宣布首次研究合作 下一篇 Meta的AI代理推进速度慢于扎克伯格计划
北竹游乐场 免费玩小游戏 免费玩