Skip to content
Rush Commerce
AI & Automation3 min read

Thomson Reuters built its own model for $40M — data is the asset

Thomson Reuters trained a proprietary LLM on Westlaw and Reuters content for $40M and owns it outright. The lesson for operators isn't build-your-own — it's what you own.

Thomson Reuters spent $40 million and came out the other side owning a frontier-class model instead of renting one. The company announced Thomson on August 24 — its first proprietary LLM, trained from an open-source foundation on decades of Westlaw, Practical Law, Checkpoint, and Reuters content. You are almost certainly not going to spend $40 million on a model. That's fine. The transferable part of this story is what made the spend possible, and you may already have a smaller version of it.

What actually happened

Thomson Reuters did not train from scratch. It started from a strong open-source base, then applied mid-training and post-training on proprietary content, with hundreds of in-house subject matter experts shaping the training objectives and the evaluations. The $40 million covers talent and compute — a fraction of what a from-zero frontier run costs.

The claims, which are the company's own and worth reading as such: across legal and general benchmarks, Thomson Reuters says the model performs competitively with the strongest frontier models on the market, including Claude Opus 4.8, and ahead of GPT-5.5, Claude Sonnet 5, and Gemini 3.1 Pro. CTO Joel Hron's framing in the release is the honest thesis — build intelligence that is "far more efficient and entirely under your control."

Two details make the strategy legible. First, the model has been trained on less than 10% of the company's proprietary content so far — the moat is nowhere near spent. Second, first deployment is narrow: Tabular Analysis inside CoCounsel Legal, with rollout across the legal and tax portfolio to follow. They are also releasing a small open-weight version on Hugging Face for academic and non-commercial use. SiliconANGLE has the outside read.

Why owning your data matters for your business

The asset was never the model. It was fifty years of structured, edited, citation-linked legal content that nobody else has. The model is what turned that archive into a product. Ask the same question about your business: what do you have that a competitor cannot buy? Service tickets with resolutions attached. Ten years of quotes and which ones closed. Install notes, failure modes, part substitutions. Most small companies have this and store it as exhaust.

Structure now, model later. You do not need a training run to benefit. You need that data extractable, consistently labeled, and out of the vendor's proprietary format — because every option downstream, from retrieval to fine-tuning to just answering a question fast, depends on being able to get at it. The companies that will spend on models in three years are the ones cleaning up their data this year.

Narrow deployment is the tell. A $40 million model shipped into one feature first. That's discipline, and it's the same discipline that makes small-company AI projects work: pick the workflow where the answer is checkable, ship it, measure it, then widen. The projects that die are the ones that launched as a platform.

Key takeaways

  • Thomson Reuters trained a proprietary LLM for $40M starting from an open-source base — not from scratch
  • Training data is decades of Westlaw, Practical Law, Checkpoint, and Reuters content, less than 10% of it used so far
  • Benchmark claims are the company's own internal evaluations — read them as vendor claims, not third-party results
  • First deployment is a single feature (Tabular Analysis in CoCounsel Legal), with wider rollout to follow
  • The lesson for smaller operators is upstream: make your proprietary data structured, labeled, and portable

Your ticket history is a training set nobody else has — if you can get it out of the software it's trapped in. We build the extraction, structure, and retrieval layer that makes your own data usable, in systems you own. See how we do it, or run the numbers on what that workflow is costing you today.

Sources: Thomson Reuters, SiliconANGLE.

  • #llm
  • #proprietary-data
  • #vendor-strategy
  • #thomson-reuters
  • #ai-strategy
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.