- Thomson Reuters spent roughly $40 million in talent and compute to build Thomson, a model it says performs on par with frontier systems and leads them on factuality, scoring 0.83 against 0.68 and 0.65 for two leading models.
- Thomson was trained on less than 10 percent of the company's 175-year archive of legal, tax, accounting, and news content, leaving the majority of the moat untapped.
- The model went live on August 24 inside CoCounsel Legal, with a small open-weight version on Hugging Face, built by the former Safe Sign Technologies team acquired in 2024.
A 175-year data archive becomes a $40 million model
Thomson Reuters has built a foundation model that it says matches frontier systems on general tasks and beats them where professional accuracy matters, and it did so for about $40 million in talent and compute. The company calls the model Thomson, and its arrival gives the market a working example of an incumbent building its own intelligence on proprietary data rather than renting it from a frontier lab.
The number reframes a build-versus-buy question that most enterprises have already answered in favor of buying. Forty million dollars sits as a rounding error next to the billions frontier labs spend on a single training run, and Thomson Reuters reached competitive quality by starting from a strong open-weight base model and continuing to train it on content the labs cannot license.
The company treats that base model as swappable infrastructure. By its own account it has changed Thomson's root model several times as open-weight systems improved, layering continued pre-training on its Westlaw, Practical Law, Checkpoint, and Reuters archives on top. The durable advantage sits in 175 years of curated professional data and the domain experts who shaped it, the one input a general-purpose lab cannot simply purchase.
On the metric that decides whether a professional will trust an answer, Thomson Reuters reports a factuality score of 0.83 against 0.68 and 0.65 for two leading frontier models, alongside an instruction-following score of 0.914 that it says leads the frontier systems it tested. These are vendor benchmarks and warrant the usual caution, though the direction of the claim is what carries weight: a domain specialist asserting an accuracy edge over generalist systems on its home ground.
Why an incumbent chose to build instead of license frontier AI
For two years the prevailing assumption held that companies would rent intelligence from OpenAI, Anthropic, or Google and wrap their data around it. Thomson Reuters chose the reverse, and its chief technology officer explained the logic plainly.
Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient and entirely under your control.Joel Hron, Chief Technology Officer, Thomson Reuters
The operative phrase is "entirely under your control." A licensed frontier model comes with a landlord, and its price, usage limits, content policies, and release cadence all sit with the lab. For a company whose entire product is a trusted, defensible answer to a legal or tax question, handing the reasoning layer to a supplier that can change its terms plants a strategic dependency at the core of the business. Building in house removes that dependency and turns the archive from a static asset into one that compounds with every additional slice the company trains on.
| Dimension | License a frontier model | Build on proprietary data (Thomson) |
|---|---|---|
| Cost basis | Per-token pricing set by the lab | About $40M once, then marginal training runs |
| Control | Price, limits, policy, cadence sit with the lab | Entirely under the company |
| Data moat | Wrapped around a shared general model | Trained into the model itself |
| Domain accuracy | Generalist baseline | 0.83 factuality vs 0.68 and 0.65 |
| Performance ceiling | Set by the lab's roadmap | Set by how much of the archive is used |
Source: Thomson Reuters, How we built Thomson, and Santage analysis, August 2026.
What Thomson signals for every company sitting on proprietary data
This belongs in a strategy briefing because Thomson Reuters shares its core advantage with a long list of incumbents. Banks, insurers, hospital systems, law firms, and industrial manufacturers all hold decades of proprietary data that no frontier lab can license. The missing ingredient was proof that a non-lab could turn that data into frontier-competitive intelligence for tens of millions of dollars rather than billions, and Thomson now supplies it.
| Cost | Roughly $40 million in talent and compute |
| Data used | Less than 10 percent of a 175-year proprietary archive |
| Factuality | 0.83, against 0.68 and 0.65 for leading frontier models |
| Instruction following | 0.914, ahead of the frontier models tested |
| Live since | August 24, inside CoCounsel Legal, with an open-weight version on Hugging Face |
| Team | The former Safe Sign Technologies startup, acquired in 2024 |
Follow the pattern forward and the most valuable AI asset in many sectors becomes a defensible data corpus paired with a small team that can specialize an open-weight base on top of it. That shifts the center of gravity away from the frontier labs, which have spent two years positioned as the indispensable layer that every enterprise plugs into. The mechanics of doing this are the same ones covered in our guide to large language models.
The number most coverage missed: less than 10 percent
The figure that should concern the labs is the 10 percent. Thomson Reuters reached frontier-competitive quality using less than a tenth of the data it holds, and it plans to keep swapping in stronger open-weight bases as they arrive. The ceiling on Thomson's performance is set by how much of its own archive the company chooses to use next, rather than by a compute budget it must win from investors.
Thomson Reuters has shown that an incumbent with deep proprietary data and a small specialist team can now manufacture frontier-grade intelligence cheaply and on its own terms. That capability threatens the rent-intelligence-from-a-lab model far more directly than any competing chatbot could, because it removes the labs from the center of the value chain for exactly the companies that generate the most valuable data.
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.