Thomson Reuters Launches Its Own Frontier Model, Betting Proprietary Data Beats Renting GPT or Claude
Thomson Reuters unveiled Thomson, its first in-house large language model built on decades of legal, tax and news data, deploying it first inside CoCounsel Legal and releasing a smaller open-weight version on Hugging Face.
Thomson Reuters announced on August 24 the launch of Thomson, its first proprietary large language model built in-house rather than licensed from an outside AI lab, betting that decades of exclusive legal, tax and news data can produce a narrower but more accurate model than general-purpose frontier systems.
Built on Thomson Reuters' own archives — starting from an open base
Rather than training a model from scratch, Thomson Reuters said it built Thomson by starting from an open-weight foundation model and substantially improving it with a mid-training corpus of roughly 200 billion tokens drawn from a larger pool of more than 19 trillion tokens of permissively licensed public and proprietary data. That proprietary data includes Westlaw case law and statutes, Practical Law guidance, contracts, regulatory filings and news content the company has built up over decades. The company said the final training run cost roughly $450,000, part of a total investment of about $40 million in the effort over two years covering talent and compute — a fraction of what training a comparable frontier model from scratch typically costs.
First deployment: high-volume document review
Thomson's first production use is inside Tabular Analysis, a high-volume document-review capability in Thomson Reuters' CoCounsel Legal AI assistant, where the company says a purpose-built model shows a clear advantage on structured, repetitive review tasks. CoCounsel Legal remains multi-model by design, according to the company, meaning Thomson will be used where it performs best while other leading models from outside labs continue to power other parts of the product. Thomson Reuters said early evaluations put Thomson on par with current frontier models across a range of tasks relevant to its legal, tax and news businesses, though it has so far been trained on less than 10% of the company's total content.
A smaller version goes open-weight
Alongside the proprietary flagship, Thomson Reuters published a smaller version, Thomson-1.0-Small, as an open-weight model on Hugging Face for academic and non-commercial use. According to its model card, Thomson-1.0-Small was built by repurposing an open-weight Qwen model and fine-tuning it for high-stakes professional domains — legal, tax and journalism — under what the company describes as a continual-learning approach.
Part of a broader owning-vs-renting trend
The move puts Thomson Reuters alongside a small but growing group of data-rich enterprises building their own models rather than solely licensing frontier systems from labs like OpenAI, Anthropic or Google, aiming to cut inference costs and retain full control over how proprietary content is used to train and run production AI. Thomson Reuters has not said whether it plans to reduce its use of third-party frontier models in CoCounsel over time or expand Thomson's role beyond document review.
Sources
- Thomson Reuters Leverages its World-Class Data Assets to Launch Its Own Frontier Model — Thomson Reuters press release
- Thomson Reuters launches proprietary AI model for legal work — SiliconANGLE
- thomsonreuters/Thomson-1.0-Small — Hugging Face model card
- Thomson Reuters Launches Thomson, Its Own Proprietary LLM Trained on Westlaw and Practical Law Content — LawSites
AI-assisted reporting, overseen by the AgentsAI team. Spotted an error? Let us know.
More ai news
TCS Unit HyperVault to Invest Up to $7.4B in 1GW AI Data Center Campus in India
Tata Consultancy Services' infrastructure arm HyperVault will invest up to $7.4 billion with partners to build a 1-gigawatt AI data center campus in Hyderabad, one of India's largest bets yet on domestic AI compute capacity.
Mistral Raises €3B in Samsung-Led Round, Becomes Europe's Best-Funded AI Startup
French AI lab Mistral raised €3 billion in a Series D round led by Samsung Electronics, pushing its post-money valuation above €21 billion and marking the largest equity round ever completed by a European technology company.
Sanders and Casar Introduce Bill to Ban 'Artificial Superintelligence' Outright
Sen. Bernie Sanders and Rep. Greg Casar introduced the Ban Artificial Superintelligence Act on September 3, which would permanently prohibit superintelligent AI systems, pause frontier development pending federal safety rules, and impose prison terms and a 'corporate death penalty' for violations.
Anthropic Says Claude Produced the First Machine-Checked Proof of Fermat's Last Theorem
Working largely autonomously for 11 days on the open Prove2Me platform, Claude generated a 13-million-line Lean formalization of Fermat's Last Theorem, which mathematician Kevin Buzzard called an 'extraordinary autoformalization achievement.'