AI ALPHA

Thomson Reuters Built Its Own Frontier Model for $40 Million

The Thomson Reuters logo, its dot-cluster emblem and wordmark, on a glass office panel
Thomson Reuters built its own foundation model on 175 years of proprietary data. Source: Reuters
TLDR

A 175-year data archive becomes a $40 million model

Thomson Reuters has built a foundation model that it says matches frontier systems on general tasks and beats them where professional accuracy matters, and it did so for about $40 million in talent and compute. The company calls the model Thomson, and its arrival gives the market a working example of an incumbent building its own intelligence on proprietary data rather than renting it from a frontier lab.

The number reframes a build-versus-buy question that most enterprises have already answered in favor of buying. Forty million dollars sits as a rounding error next to the billions frontier labs spend on a single training run, and Thomson Reuters reached competitive quality by starting from a strong open-weight base model and continuing to train it on content the labs cannot license.

The company treats that base model as swappable infrastructure. By its own account it has changed Thomson's root model several times as open-weight systems improved, layering continued pre-training on its Westlaw, Practical Law, Checkpoint, and Reuters archives on top. The durable advantage sits in 175 years of curated professional data and the domain experts who shaped it, the one input a general-purpose lab cannot simply purchase.

Bar chart showing Thomson scoring 0.83 on factuality versus 0.68 and 0.65 for two leading frontier models, per Thomson Reuters internal benchmarks
Thomson's reported factuality score against two leading frontier models. Source: Thomson Reuters internal benchmarks, How we built Thomson, August 2026.

On the metric that decides whether a professional will trust an answer, Thomson Reuters reports a factuality score of 0.83 against 0.68 and 0.65 for two leading frontier models, alongside an instruction-following score of 0.914 that it says leads the frontier systems it tested. These are vendor benchmarks and warrant the usual caution, though the direction of the claim is what carries weight: a domain specialist asserting an accuracy edge over generalist systems on its home ground.

Why an incumbent chose to build instead of license frontier AI

For two years the prevailing assumption held that companies would rent intelligence from OpenAI, Anthropic, or Google and wrap their data around it. Thomson Reuters chose the reverse, and its chief technology officer explained the logic plainly.

Start with a strong foundation, specialize it deeply for the work that matters, and you can build intelligence that is highly capable, far more efficient and entirely under your control.
Joel Hron, Chief Technology Officer, Thomson Reuters

The operative phrase is "entirely under your control." A licensed frontier model comes with a landlord, and its price, usage limits, content policies, and release cadence all sit with the lab. For a company whose entire product is a trusted, defensible answer to a legal or tax question, handing the reasoning layer to a supplier that can change its terms plants a strategic dependency at the core of the business. Building in house removes that dependency and turns the archive from a static asset into one that compounds with every additional slice the company trains on.

DimensionLicense a frontier modelBuild on proprietary data (Thomson)
Cost basisPer-token pricing set by the labAbout $40M once, then marginal training runs
ControlPrice, limits, policy, cadence sit with the labEntirely under the company
Data moatWrapped around a shared general modelTrained into the model itself
Domain accuracyGeneralist baseline0.83 factuality vs 0.68 and 0.65
Performance ceilingSet by the lab's roadmapSet by how much of the archive is used

Source: Thomson Reuters, How we built Thomson, and Santage analysis, August 2026.

What Thomson signals for every company sitting on proprietary data

This belongs in a strategy briefing because Thomson Reuters shares its core advantage with a long list of incumbents. Banks, insurers, hospital systems, law firms, and industrial manufacturers all hold decades of proprietary data that no frontier lab can license. The missing ingredient was proof that a non-lab could turn that data into frontier-competitive intelligence for tens of millions of dollars rather than billions, and Thomson now supplies it.

The build-versus-buy case, by the numbers
CostRoughly $40 million in talent and compute
Data usedLess than 10 percent of a 175-year proprietary archive
Factuality0.83, against 0.68 and 0.65 for leading frontier models
Instruction following0.914, ahead of the frontier models tested
Live sinceAugust 24, inside CoCounsel Legal, with an open-weight version on Hugging Face
TeamThe former Safe Sign Technologies startup, acquired in 2024
Source: Thomson Reuters press release and How we built Thomson, August 2026.
In short: Thomson Reuters built a frontier-competitive model for about $40 million by training on less than 10 percent of its 175-year data archive, showing that an incumbent with proprietary data can now build rather than rent its core intelligence.

Follow the pattern forward and the most valuable AI asset in many sectors becomes a defensible data corpus paired with a small team that can specialize an open-weight base on top of it. That shifts the center of gravity away from the frontier labs, which have spent two years positioned as the indispensable layer that every enterprise plugs into. The mechanics of doing this are the same ones covered in our guide to large language models.

The number most coverage missed: less than 10 percent

The figure that should concern the labs is the 10 percent. Thomson Reuters reached frontier-competitive quality using less than a tenth of the data it holds, and it plans to keep swapping in stronger open-weight bases as they arrive. The ceiling on Thomson's performance is set by how much of its own archive the company chooses to use next, rather than by a compute budget it must win from investors.

Thomson Reuters has shown that an incumbent with deep proprietary data and a small specialist team can now manufacture frontier-grade intelligence cheaply and on its own terms. That capability threatens the rent-intelligence-from-a-lab model far more directly than any competing chatbot could, because it removes the labs from the center of the value chain for exactly the companies that generate the most valuable data.

Quick quiz
How much of its 175-year data archive did Thomson Reuters use to train Thomson?

Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.