AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-07-03 0 浏览 会员

GPT和Claude在桥水基金的金融测试中失败,因为正确答案从未公开

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 会员 POST / 2026-07-03 19:16:42

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

Bridgewater and Thinking Machines Lab have trained an open-source AI model for analyzing financial documents that outperforms leading commercial models. The Qwen3-235B model, which has been fine-tuned using internal expert knowledge, achieves nearly 85 percent accuracy in tests and is 14 times cheaper to operate. This demonstrates that companies can develop powerful AI solutions using their own data without having to share sensitive information with large providers. Hedge fund Bridgewater and Thinking Machines Lab say a fine-tuned open-weight model outperforms the strongest AI models at evaluating financial documents, at a fraction of the cost. The numbers come from their own internal evaluation. Investors get buried in news, analysis, corporate filings, and emails every day. According to a report from Bridgewater's AIA Labs and Thinking Machines Lab , the startup founded by former OpenAI CTO Mira Murati, reading isn't the real work. The real work is the constant stream of small, repeated judgment calls about what actually matters. That's the triage the researchers wanted to automate. They defined six tasks drawn from an investor's daily routine. One example: deciding whether a financial article is relevant to an executive. Another: whether a central bank document signals the direction of future rate changes. For investors, these calls are trivial, but they can barely put their reasoning into words. The report gives a telling example. A headline about Trump's claim to Greenland gets flagged as irrelevant, while Trump's threat of new China tariffs is highly relevant. Both touch on geopolitics and finance. Ad Frontier models failed in the authors' tests. Variants of Gemini, Claude, and GPT hit only about 50 percent accuracy with a basic prompt. Expert-written instructions and a three-tier rating system ("relevant and interesting," "relevant but uninteresting," "irrelevant") pushed accuracy into the mid-70s. That still fell short of the 80 percent threshold the authors set for trustworthy deployment. Ad DEC_D_Incontent-1 Newer models barely improve per dollar, the report says. GPT 5.4 costs 43 percent more than 5.2 but is only marginally more accurate. The solution was fine-tuning, retraining an open-weight model on proprietary examples. The key ingredient was the Bridgewater investors' judgment: At first, cheap outside contractors labeled the documents, but many of those labels were wrong. To avoid having expensive professionals review everything, the researchers used a workaround. A first model learned from the flawed labels and re-evaluated the same documents. Wherever the model and the original label disagreed, there was likely an error. Only those disputed cases went to investors for correction. Ad Training ran on the Tinker platform from Thinking Machines Lab, built on top of the open model Qwen3-235B. In the team's own evaluation, the fine-tuned model hit 84.7 percent accuracy versus 78.2 percent for the best frontier model tested. It also cost nearly 14 times less to run. This isn't a truly independent comparison, of course. Both companies have a clear interest in selling their product. Still, the finding beyond the numbers is worth noting. It shows once again that the big labs like OpenAI haven't absorbed all the data out there. Huge pools of proprietary corporate data and untrained human expertise still exist, and they hold real room for improvement. That's especially true where companies deliberately keep their most valuable data private. Anyone who hands that data to a frontier lab risks competing against a product built on top of it. Ad DEC_D_Incontent-2 Fine-tuning open models through tools like Tinker gives companies an alternative. They keep the weights, the data, and, depending on the setup, the GPUs themselves. Ad

中文翻译

桥水基金和Thinking Machines Lab训练了一个用于分析金融文档的开源AI模型,其表现优于领先的商业模型。使用内部专家知识微调的Qwen3-235B模型在测试中达到近85%的准确率,运行成本低14倍。这表明公司可以利用自己的数据开发强大的AI解决方案,而无需与大型提供商分享敏感信息。

核心信息

桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。

  • 桥水基金和Thinking Machines Lab微调的开源模型Qwen3-235B在金融文档分析上以更低成本超越顶级商业模型,准确率达84.7%,成本降低14倍。
  • 原贴提到:Bridgewater and Thinking Machines Lab have trained an open-source AI mod
  • 来源:the-decoder.com
试看内容

成为会员查看完整内容

你已经看到了这篇内容的前置整理,剩余深度部分仅对会员开放。

详细解读 信息差价值 参考来源
成为会员查看完整内容
上一篇 Google DeepMind 与 A24 宣布首次研究合作 下一篇 Meta的AI代理推进速度慢于扎克伯格计划