觉
AI觉醒星球
Awakening is here
Knowledge File / AI小生意项目库
2026-08-24 2 浏览 免费阅读

汤森路透豪掷4000万美元自建AI,而非租赁OpenAI或Anthropic的模型

汤森路透基于阿里Qwen自建法律AI模型“Thomson”,投入约4000万美元,认为自建比微调外部模型更经济且独立。测试显示,只有接入独家内容时才能勉强超过GPT-5.4,优势微弱。公司计划用于文档审查,并发布小版本非商用。

SOURCE / AI小生意项目库 MIN / 9 ACCESS / 免费阅读 POST / 2026-08-24 20:59:43

原贴

查看原文
作者:Maximilian Schreiner 来源站点:the-decoder.com 原贴时间:

原文

Thomson Reuters has built "Thomson," its own AI language model for legal work, on top of Alibaba's Qwen. The $40 million model was trained on the company's own data and by its domain experts. In testing, though, it only beats rivals like GPT-5.4 when it can access exclusive company content. Building in-house is meant to cut costs and keep the company independent, rather than adapting outside models. "Thomson" will first be used for document review, and a smaller version is being released under a non-commercial license. With "Thomson," the professional information company is launching its first in-house language model, built on Alibaba's Qwen. The model hits top marks when it can tap into the company's own content and tools. Thomson Reuters spent about $40 million on staff and computing power over more than two years, according to the company . The more widely touted figure of $450,000 covers only the final training run of the current version. Even the full sum leaves out the real capital: That's decades of content from Westlaw, Practical Law, Checkpoint, and Reuters, plus the working hours of hundreds of domain experts. The foundation is Alibaba's open Qwen, most recently Qwen3.5-397B, the company says. Working with Imperial College, Thomson Reuters first retrained the Chinese model for safety, ethics, and political neutrality. This intermediate version is called "Snowdon," named after the mountain in Wales. Ad Then came pre-training on the company's own content, post-training with domain experts, and agentic reinforcement learning inside the company's own tool environments. So far, less than 10 percent of the available content has gone into training. Ad CTO Joel Hron says the company has "changed the open source starting point like probably close to a half dozen times already." The bigger finding is "less the individual model and more the model factory we built," adds research chief Jonathan Schwartz. The blog post claims Thomson ranks among the best models in the world. The company's own numbers paint a more sober picture: On Stanford LegalBench, Thomson (0.823) trails Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark it sits just behind Opus 4.8. It leads on instruction following and the tough PrBench Legal. On reasoning, and especially coding, it falls off sharply. The comparison is also skewed by method, as Thomson competes with test-time scaling, while GPT-5.5 runs without a reasoning mode. Ad In the company's in-house Deep Research benchmark with web access alone, Thomson scores 0.53 on factual accuracy, while GPT 5.4 hits 0.65. Only with access to the company's content does Thomson edge past GPT 5.4, 0.83 to 0.82. With web access, Thomson is "within the scope of the other models, but certainly not the leader yet," admits evaluation lead Andrew Bean. According to Bean, there's "a big uplift that comes from being able to train on and practice with your own tools," something outside providers can't do. What's striking is that GPT 5.4 improves just as sharply with that content. So data access does almost as much work as the specialized training. The company didn't test newer models. The in-house model's razor-thin lead could still grow, though, if the company moves to a stronger variant like Qwen3.8 and pushes far past the 10 percent content mark. Ad So why build your own model instead of fine-tuning a frontier model from OpenAI or Anthropic on legal data? Thomson Reuters sees three reasons against it. First, the economics: Standard fine-tuning techniques "tend to have a strong tendency to degrade general capability," Schwartz says. And you stay locked into the provider for inference costs and roadmap. A smaller, in-house model pays off precisely on high-volume work like document review. Ad

中文翻译

汤森路透在阿里巴巴的Qwen之上构建了自己的法律工作AI语言模型“Thomson”。这个耗资4000万美元的模型由公司自己的数据和领域专家训练而成。然而,在测试中,只有当它能够访问公司独家内容时,它才能击败像GPT-5.4这样的对手。内部构建是为了降低成本并保持公司独立,而不是采用外部模型。“Thomson”将首先用于文档审查,并且一个更小的版本将以非商业许可的形式发布。

借助“Thomson”,这家专业信息公司推出了其首个内部语言模型,基于阿里巴巴的Qwen。当模型能够利用公司自身的内容和工具时,它达到了最高水平。据该公司称,汤森路透在两年多的时间里在人员和计算能力上花费了约4000万美元。更广泛宣传的45万美元数字仅涵盖当前版本的最后一次训练运行。即使全额也不包括真正的资本:这是来自Westlaw、Practical Law、Checkpoint和Reuters数十年的内容,加上数百名领域专家的工作时间。基础是阿里巴巴开源的Qwen,最近的是Qwen3.5-397B,公司表示。汤森路透与帝国理工学院合作,首先对中国模型进行了安全、伦理和政治中立方面的重新训练。这个中间版本被称为“Snowdon”,以威尔士的山命名。

随后是在公司自有内容上的预训练、与领域专家的后训练,以及公司自有工具环境内的智能体强化学习。到目前为止,只有不到10%的可用内容被用于训练。

首席技术官Joel Hron表示,公司“已经改变了开源起点,可能接近六次”。研究主管Jonathan Schwartz补充说,更大的发现是“不是单个模型,而是我们建立的模型工厂”。

博客文章声称Thomson跻身世界最佳模型之列。公司自己的数字描绘了更清醒的画面:在斯坦福LegalBench上,Thomson(0.823)落后于Gemini 3.1 Pro和GPT-5.5。在Harvey法律智能体基准上,它略落后于Opus 4.8。它在指令遵循和艰难的PrBench Legal上领先。在推理方面,尤其是编码方面,它大幅下滑。比较也因方法而有偏差,因为Thomson与测试时扩展竞争,而GPT-5.5在无推理模式下运行。

在公司内部的仅限网络访问的Deep Research基准中,Thomson的事实准确性得分0.53,而GPT 5.4达到0.65。只有访问公司内容时,Thomson才以0.83对0.82勉强超过GPT 5.4。评估负责人Andrew Bean承认,在网页访问方面,Thomson“处于其他模型的范围之内,但肯定还不是领先者”。根据Bean的说法,“能够用自己的工具进行训练和练习带来了巨大的提升”,这是外部供应商无法做到的。引人注目的是,GPT 5.4在获得这些内容后也同样大幅提升。因此,数据访问几乎与专门训练一样起到了作用。公司没有测试更新的模型。不过,如果公司转向更强的变体(如Qwen3.8)并大幅超过10%的内容标记,内部模型的微弱领先优势仍可能增长。

那么,为什么构建自己的模型而不是在OpenAI或Anthropic的前沿模型上微调法律数据?汤森路透看到了三个反对理由。首先,经济学:Schwartz说,标准的微调技术“往往具有很强的退化通用能力的倾向”。而且你会因推理成本和路线图而被锁定在提供商上。一个较小的内部模型正是适合文档审查等高容量工作。

核心信息

汤森路透基于阿里Qwen自建法律AI模型“Thomson”,投入约4000万美元,认为自建比微调外部模型更经济且独立。测试显示,只有接入独家内容时才能勉强超过GPT-5.4,优势微弱。公司计划用于文档审查,并发布小版本非商用。

  • 汤森路透基于阿里Qwen自建法律AI模型“Thomson”,投入约4000万美元,认为自建比微调外部模型更经济且独立。测试显示,只有接入独家内容时才能勉强超过GPT-5.4,优势微弱。公司计划用于文档审查,并发布小版本非商用。
  • 原贴提到:Thomson Reuters has built "Thomson," its own AI language model for legal
  • 来源:the-decoder.com

详细解读

这是什么信号?汤森路透基于阿里的开源模型Qwen自建法律专用模型,投入4000万美元,而不是直接租用OpenAI或Anthropic的API。这标志着企业AI应用的一个新方向:在垂直领域,数据壁垒和领域知识可能比模型规模本身更重要。企业不再满足于通用模型的微调,而是开始构建“模型工厂”,将自有数据、专家经验和工具链深度集成到模型中。

为什么重要?如果只是微调,企业会被锁定在外部提供商的推理成本和路线图上,而且微调可能损害通用能力。自建模型让企业掌控AI基础设施,长期成本可控,还能利用独家内容形成差异化优势。汤森路透的案例表明,即使模型基准成绩并不领先,但结合自有数据后能胜过更大的通用模型,这意味着“数据+模型”的组合比单纯的技术领先更具商业价值。

对谁有价值?对法律、金融、医疗等专业信息提供商,它们拥有海量高质量文本,但之前只能通过API使用模型,无法充分发挥数据潜力;对企业AI决策者,这提供了一个评估自建vs租用的具体案例;对开源社区,验证了开源模型在某些场景下可以匹敌甚至超越封闭模型,尤其在垂直领域。

可以怎么行动?企业可以评估自己的数据资产,看看是否有足够独特的内容和专家知识可以支撑自建模型。然后选择合适的基础模型(如Qwen等开源大模型),投入资金进行预训练、后训练和工具环境强化。重点不是在公开基准上领先,而是结合内部工具实现业务效率提升。同时,要考虑分阶段投入,比如先用于高频、低错误容忍度的任务(如文档审查)。

风险或限制?目前Thomson的优势非常微弱,且依赖于独家内容,一旦内容被外部模型获取(如通过API),优势可能缩小。自建模型需要持续投入,如果基础模型迭代慢,可能落后。此外,该公司只用了不到10%的内容,还有提升空间,但也说明前期训练投入巨大。另一个风险是,非商业许可的小版本可能限制了开源生态的贡献,也可能面临法律伦理问题。

信息差价值

这条内容的真正价值,不只是“有人发布了一个新功能”,而是它揭示了 the-decoder.com 背后的产品方向、工作流变化或竞争信号。对 OPC 来说,这种信息可以转化成持续追踪的栏目选题。

如果把《汤森路透豪掷4000万美元自建AI,而非租赁OpenAI或Anthropic的模型》放到你的内容系统里,它最大的价值在于帮助读者更快看懂“为什么值得关注”,而不是只看到一条碎片化动态。

参考来源

AI SUMMARY

这篇文章回答了什么

汤森路透豪掷4000万美元自建AI,而非租赁OpenAI或Anthropic的模型主要讲什么?

汤森路透基于阿里Qwen自建法律AI模型“Thomson”,投入约4000万美元,认为自建比微调外部模型更经济且独立。测试显示,只有接入独家内容时才能勉强超过GPT-5.4,优势微弱。公司计划用于文档审查,并发布小版本非商用。

这篇文章最值得关注的要点是什么?

汤森路透基于阿里Qwen自建法律AI模型“Thomson”,投入约4000万美元,认为自建比微调外部模型更经济且独立。测试显示,只有接入独家内容时才能勉强超过GPT-5.4,优势微弱。公司计划用于文档审查,并发布小版本非商用。;原贴提到:Thomson Reuters has built "Thomson," its own AI language model for legal;来源:the-decoder.com

这篇文章和哪些AI专题相关?

它适合放在AI副业、AI工具、AI内容增长专题里阅读。 关联原因:这篇内容命中「项目、小生意、变现」等主题信号。;这篇内容命中「模型」等主题信号。;这篇内容命中「内容」等主题信号。

阅读这篇文章建议先理解哪些关键词?

建议先理解AI工具、工具、自动化、模型、Cursor这些关键词,再结合正文判断工具、机会或风险是否值得进入自己的工作流。

上一篇 NVIDIA Vera Rubin NVL72 树立 AI 智能体效率新标准:每瓦特工作量提升至 30 倍 下一篇 Kiro 推出 GPT-5.6,为开发者带来更佳性价比