Thomson LLM Archives - ¶¶ŇőłÉÄę Blog https://blogs.thomsonreuters.com/general-blog/topic/thomson-llm/ Mon, 28 Sep 2026 14:15:54 +0000 en-US hourly 1 https://wordpress.org/?v=6.8.9 Inside the Thomson technical report https://blogs.thomsonreuters.com/general-blog/inside-the-thomson-llm-technical-report/ Mon, 21 Sep 2026 12:00:29 +0000 https://blogs.thomsonreuters.com/general-blog/?p=22167 What a “frontier model” is, and how we built one for less What the industry calls a “frontier model” is, …

The post appeared first on .

]]>

What a “frontier model” is, and how we built one for less

What the industry calls a “frontier model” is, in plain terms, a system at the leading edge of capability, the bar every other model is measured against. Until now, that bar has been set by a small number of heavily funded labs, each spending billions of dollars and years of infrastructure to reach it.

We took a different path. With Thomson, our first proprietary large language model, we built a system that competes with those frontier models, but did so with fewer than three dozen people, in three months from first experiments, with a final training run estimated at under $450,000 in GPU costs, specialized using ¶¶ŇőłÉÄę’ authoritative professional content and expert-created data, while preserving broad capabilities.

That efficiency is the headline. What matters more is what’s behind it: how do we know any of it is true?

The answer is in the technical report we published alongside the launch of Thomson. It’s built to the standard we call Fiduciary-Grade AI™: AI for professionals with duties of care, where “almost right” is not good enough. The report is worth understanding at a summary level even if you never open the full PDF, because the methodology is what makes the claim credible.

“Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate.”

Joel Hron

Chief Technology Officer, ¶¶ŇőłÉÄę

Thomson: The technical report

Thomson: The technical report

Review our findings and details of the methods, data, and evaluations

Read the full report ↗

 

Continual learning

The most interesting idea in the report isn’t the benchmark scores. It’s how the model was built. Thomson wasn’t trained from scratch, which would have required the billions in compute that frontier labs spend. But it also wasn’t simply fine-tuned on top of an existing model, which typically buys narrow domain gains at the cost of broader capabilities. Instead, we started with open-weight models and applied what the report calls Continual Learning: a training approach that materially reshapes the model across the entire training stack while deliberately preserving the capabilities it already had.

The result is what the report describes as a T-shaped performance profile: pronounced gains in the professional domains we targeted, alongside preserved and frequently improved performance in domains we didn’t. That last part is the surprise. Most domain adaptation trades breadth for depth; Continual Learning, as we applied it, largely avoided that trade-off.

The report makes a broader case that this approach is repeatable: a blueprint for other institutions to build competitive models from open-weight starting points, at a fraction of the cost previously imagined.

Two kinds of evidence

The report doesn’t rest on a single number. It evaluates Thomson two different ways, and the distinction is the most useful thing to take from it: one tests the model in the lab, under identical conditions; the other tests it in the room, against the messy way professionals actually ask questions. Having both is what makes the claim worth taking seriously.

  1. The first is a standard benchmark comparison: Thomson-1.0-Large measured against today’s leading models, under identical conditions, across legal, tax, journalism, safety, and general-purpose tasks. Thomson lands exactly one percentage point behind the top-scoring model on an overall, unweighted average (79.5 vs. 78.5), and it leads every model tested on two specific measures: instruction following and a political-neutrality evaluation. Those two measures matter more than their category names suggest. Specializing a model on dense professional content usually costs something elsewhere. Narrow training tends to buy domain performance by spending general reliability. Instruction following and neutrality holding up, rather than slipping, is evidence the specialization was done with more care than brute force. The report breaks all of this down further in a full comparison table for anyone who wants the underlying numbers.
  2. The second kind of evidence is closer to how the model gets used day to day: a blind preference study comparing complete systems, not just models. Subject-matter experts rated thousands of real conversations without knowing which system had produced which answer. In legal conversations specifically, expert raters preferred Thomson’s answers to those of several named frontier systems a clear majority of the time. This is notable partly because this was a system comparison, not just a model comparison: Thomson had access to our legal tools and Reuters news, while the external models had broader web access. The report lays out the exact win rates against each named system, along with a separate, tighter set of results for general (non-legal) conversation, worth a look if you want the granular comparison. The per-system numbers is what a technical reader will want to scrutinize.

Where it loses, and why that’s the point

The most convincing part of the report is where it says, plainly, that Thomson isn’t ahead everywhere. For a skeptical technical audience, this matters more than any win rate: a launch document that volunteers a loss is one you can trust to report its wins honestly.

There’s one domain where Thomson-1.0-Large scores below its own starting model, attributed to mild forgetting during specialization and explicitly flagged as not a target area for this release. General-purpose reasoning, similarly, trails the strongest proprietary systems rather than leading them.

That’s worth sitting with. The version the data actually supports is narrower and more useful: Thomson is built and tuned for professional, domain-specific work, competitive with frontier systems on the tasks it was built for, and not represented as the best at everything. For the legal, tax, and compliance professionals this model is meant to serve, that’s the claim that should matter: not whether Thomson tops a general leaderboard, but whether it holds up on the specific kind of question they actually ask it.

The report’s constituent-benchmark tables are where that distinction gets precise, if you want to see exactly which tasks it covers and where the boundaries of the claim sit.

An audit, not an advertisement

It’s tempting to treat a technical report as supporting material for the announcement. It’s more accurate to treat it as the actual point.

The Fiduciary-Grade AI standard, the one this report exists to back up, only holds up if someone outside this building can check it, which is why the evaluation methodology carries as much weight here as any single score.

The two kinds of evidence only matter if they can be checked by someone other than us. The reason it holds up is that the evidence is open to outside scrutiny. The evaluations use publicly recognized benchmarks, re-implemented in a common harness so that any model is tested the same way. The report’s training record supports reproducibility and post-hoc analysis, so that for any released checkpoint the constituent datasets and training runs can be recovered exactly. And the smaller model is released as an open weight on Hugging Face for anyone to inspect directly.

Why SovereignAI matters

The report’s argument goes beyond building one strong model efficiently. Its larger stance is that organizations can own and control more of the AI stack than most assume: the model itself, proprietary data and tools, governance and values, infrastructure, and the economics of deployment. The report calls this SovereignAI, and frames it as a spectrum rather than a binary. Thomson doesn’t achieve full sovereignty on every axis, but it demonstrates meaningful progress across all of them, on a budget that makes the path viable for a much wider range of institutions than the current frontier-lab model implies.

If any of this is going to inform how you evaluate Thomson for your own work, read the report. The full benchmark tables, the preference-study breakdowns by system and by domain, the methodology behind each, and the open-weight model on Hugging Face are all there, worth reading firsthand rather than taking on faith.

Thomson

Thomson

The purpose-built, proprietary LLM, engineered for high stakes professional work

Learn more ↗

 

The post appeared first on .

]]>
What a $40 million model says about where enterprise AI is actually headed https://blogs.thomsonreuters.com/general-blog/what-a-40-million-model-says-about-where-enterprise-ai-is-actually-headed/ Wed, 16 Sep 2026 12:05:33 +0000 https://blogs.thomsonreuters.com/general-blog/?p=22318 Every major AI lab has bet that frontier performance requires billions in compute. On a LinkedIn Live event this week, …

The post appeared first on .

]]>
Every major AI lab has bet that frontier performance requires billions in compute. On a this week, ¶¶ŇőłÉÄę challenged that bet directly. The company built Thomson, its first proprietary large language model, in-house — for a reported $40 million, a fraction of what frontier labs spend to get there.

President and CEO Steve Hasker and Alexander Kardos-Nyheim, Senior Director of TR Labs and co-founder of Safe Sign Technologies, joined Emily Colbert, Co-Head of Product for ¶¶ŇőłÉÄę Legal, to defend that bet. What emerged wasn’t a product pitch. It was an argument about where AI competition actually gets won.

Why not just rely on the frontier labs?

Colbert put the obvious question to Kardos-Nyheim directly: why build your own model at all, instead of layering onto a frontier lab’s? His answer was direct. “It’s a good question, and one I’m often asked,” he said. His answer cuts against the industry’s dominant logic: capability was always going to become abundant, because every well-funded lab is racing toward it. That makes capability the wrong thing to bet a company on. Trust, precision, and a defensible chain from answer back to source are the harder problem — and general-purpose models, trained for breadth, were never built to solve it. Hasker framed the urgency behind that bet in blunt terms: “I’ve always thought about three to five year planning cycles. I think we’re in six month increments now.” That pace punishes any company still deciding whether to build.

What $40 million actually buys

The panel met the cost question directly, then reframed it. Instead of defending $40 million as competitive with the billions frontier labs spend, Kardos-Nyheim treated the figure as a signal of efficiency, not a ceiling — and pointed to where the money actually went: not compute, but the human layer underneath the training data. ¶¶ŇőłÉÄę has described that layer elsewhere in detail — partner-level lawyers spent months building evaluation rubrics for the hardest research tasks the company could specify, and qualified attorneys logged thousands of hours choosing between model outputs against criteria no generic benchmark captures. The number makes a pointed argument to the rest of the industry: if compute really sets the ceiling on frontier performance, the price tag should look like the labs spending billions. If it doesn’t, a company holding the right proprietary content and the right in-house expertise can close most of that gap for a fraction of the price.

AI sovereignty: the product, not a feature

The sharpest insight from the conversation was about control. Hasker named exactly what customers fear when they raise “AI sovereignty”: that routing their data through a third-party frontier model, or a startup built as a thin wrapper around one, lets that data train a model their competitors, or even their own clients, can eventually query. “So it’s their IP,” he said. “It’s the source of their competitive advantage.” Owning Thomson outright solves that problem at the root. The model runs behind a customer’s own firewall, alongside the products built on decades of ¶¶ŇőłÉÄę content: Westlaw, Practical Law, Checkpoint, and CoCounsel, instead of sending a firm’s most sensitive work through infrastructure the firm doesn’t control. Owning the model means ¶¶ŇőłÉÄę controls the training data, the behavior, and where it runs, instead of depending on someone else for each.

How Thomson was built

The 2024 acquisition of Safe Sign Technologies, the company Kardos-Nyheim co-founded, planted the seed that grew into TR Labs. What stands out is how little of ¶¶ŇőłÉÄę’ own content the model has needed to get here: Thomson has trained on less than 10% of ¶¶ŇőłÉÄę’ proprietary content to date. The panel treated that not as a limitation but as headroom — the model “only gets better from here” as ¶¶ŇőłÉÄę brings more of that archive to bear. Asked which evaluation result mattered most, Kardos-Nyheim didn’t hesitate: “If there’s one result I want to call out, it would be factuality” not whether the model sounds authoritative, but whether its claims survive being checked against the sources it cites. ¶¶ŇőłÉÄę has said elsewhere that Thomson clears that bar against general-purpose frontier models working over the open web.

Built to a Fiduciary-Grade standard

¶¶ŇőłÉÄę coined a specific term for the bar Thomson has to clear: Fiduciary-Grade AI™. The panel drew the distinction sharply; not about how smart the model is, but about who must survive scrutiny from its output. Strip back who ¶¶ŇőłÉÄę serves, and you get fiduciary professions: lawyers, tax and audit professionals, people whose work needs to be right. A plausible answer that’s nine-tenths correct doesn’t count as close — it counts as a failure. That standard drives Thomson’s design goal directly. As the event’s closing message put it, ¶¶ŇőłÉÄę built the model “so professionals can verify, cite, and defend their work” AI that supports professional judgment rather than replacing it, where “accountability remains human.” It’s a sharper line than most AI vendors draw, and the panel clearly wanted it remembered:

The question is no longer whether AI can generate an answer; it’s whether professionals can verify it and stand behind it.

 

Where the model already earns its keep

already runs Thomson inside Tabular Analysis, reviewing up to 10,000 documents against as many as a hundred questions at once, with every answer traceable back to its source document. That deployment sits inside a platform already operating at real scale: 1 million CoCounsel users across 107 countries and territories, and roughly 2,500 internal domain experts have built the work behind Westlaw, Practical Law, OneSource, Checkpoint, and CoCounsel. The panel described migrating more of the CoCounsel suite onto Thomson over time, extending the same sovereign, firewall-protected environment across a broader set of workflows.

Curious how Thomson holds up against the standard your work demands? Learn more about Thomson.

Why we built Thomson

Why we built Thomson

You shouldn't have to settle for AI built for everyone. Now you don't.

Read the blog ↗

The post appeared first on .

]]>
Why we built Thomson https://blogs.thomsonreuters.com/general-blog/why-we-built-thomson-llm/ Tue, 01 Sep 2026 13:00:41 +0000 https://blogs.thomsonreuters.com/general-blog/?p=21977   There’s a moment every legal professional knows. You’re deep in a document review — hundreds of contracts, thousands of …

The post appeared first on .

]]>

Highlights

  • ¶¶ŇőłÉÄę built Thomson, a proprietary LLM designed specifically for high-stakes legal and compliance work.
  • Thomson powers CoCounsel's Tabular Analysis, extracting structured answers from up to 10,000 documents with verifiable citations.
  • Over one million professionals now use CoCounsel, reflecting a fundamental shift in legal workflows powered by Fiduciary-Grade AI™.

 

There’s a moment every legal professional knows. You’re deep in a document review — hundreds of contracts, thousands of clauses, a deadline closing in — and you ask an AI tool a question that shouldn’t be complicated. The answer comes back confident, well-written, and subtly wrong. Not wrong enough to catch on first read. Just wrong enough to matter.

That’s why we built Thomson. Not just because general models get the details wrong sometimes, but because owning the model outright changes what we can do about it.

¶¶ŇőłÉÄę doesn’t just deploy AI. We build it.

Why general-purpose AI isn’t enough

The past few years have brought extraordinary AI advances. General-purpose large language models (LLM) can write code, summarize articles, and hold nuanced conversations. But legal and compliance professionals work in a different environment. Imprecision has consequences. A misread clause or a missed precedent isn’t an inconvenience; it’s a risk you can’t take.

So what separates a tool that’s merely useful from one you can stand behind?

Not all AI is built for the same stakes. At one end are general-purpose tools — broadly useful, but shallow on any one domain. A step further are professional-grade tools, built for a specific field, in environments where an occasional error is tolerable. And then there’s a third tier: Fiduciary-Grade AI™, built for work where a small error doesn’t just cost time, it can mean a lost case or a client’s trust.

That’s the tier legal and compliance work lives in. It’s the tier Thomson was built for.

The decision to build our own

We’ve always been an expertise company. Grounded by 175 years of intelligence, we’ve combined proprietary data with domain knowledge to give professionals the intelligence they need for their most important work. Thomson is the next chapter of that mission.

Owning more of the AI stack gives us greater control over performance, economics, deployment, and sovereignty — how the model behaves, what it costs us to run, where it runs, and how our clients’ data is protected. Just as importantly, it lets us continuously improve Thomson using the content and expertise only ¶¶ŇőłÉÄę has, rather than depending on what a third party chooses to build next.

That decision started in 2024, when we acquired Safe Sign Technologies, an AI research company founded by leading AI and legal minds. The result is a model that’s built on open-source foundations we can evolve as the AI landscape shifts, while the legal intelligence we put into Thomson only deepens over time. We’re not locked into any single foundational system — and we’re not standing still.

What we built

Thomson is our proprietary LLM, purpose-built by ¶¶ŇőłÉÄę for high-stakes professional work. It was built on our proprietary content, including Westlaw and Practical Law, two of the most authoritative legal data sources in the world, and reflects input from our vast network of human legal experts.

Think of it as a “T-shaped” model. The horizontal bar is broad general capability — language understanding, reasoning, and the fundamentals any strong AI needs. The vertical column is what makes Thomson different: authoritative legal content and deep subject-matter expertise, refined by expert guidance, producing a level of legal comprehension optimized specifically for professional work. We built it to be both broad and deep.

Thomson and CoCounsel aren’t the same thing, and that’s by design. CoCounsel is the product — the agentic workspace, skills, and workflows legal professionals use every day. Thomson is one of the LLMs CoCounsel calls on.

What it’s already doing

Thomson’s first implementation just launched, powering CoCounsel Legal’s Tabular Analysis feature. You can extract structured answers from up to 10,000 documents across up to 100 questions you define — at a scale and consistency that used to be impossible without significant time and cost.

Powering the engine behind Tabular Analysis, Thomson gives you deeper comprehension of complex legal documents, more consistent extraction across question types, lower variance across large document sets, and verifiable, citation-grounded outputs your attorneys can review and rely on. These aren’t incremental improvements — the numbers show it. As of Q1 2026, CoCounsel has crossed one million users. The share of our business that’s Gen AI-enabled has doubled in a little over a year. Monthly CoCounsel SKUs in legal have quadrupled year-over-year. That’s not a pilot program. That’s a shift in how the work gets done.

Independent testers agree. A Washington University law professor pitted Thomson against ChatGPT and Claude using real questions from his own Corporate Tax class and preferred Thomson’s answers overall, pointing to its linked citations to treatises as especially useful. A separate review from Queen’s Conflict Analytics Lab and Cornell Legal AI Lab found Thomson’s citation quality held up against leading frontier models — even on Canadian employment-law questions the model wasn’t specifically tuned for.

We hold Thomson to the same standards you hold your own work to. It’s developed and deployed within enterprise-grade governance frameworks grounded in the principles of Fiduciary-Grade AI™: transparency, traceability, verifiability, structured evaluation, and expert validation. Thomson is built for repeatable performance you can trust.

Built for the people who make change real

Speed is the easy part. Any AI tool can give you an answer fast. What you need is one you can stand behind — in front of a judge, a client, a board. That’s what we mean by Fiduciary-Grade AI™: intelligence built to the standard your profession already holds itself to.

The legal professionals who rely on us every day aren’t passive consumers of information.

You’re the people who protect businesses from harm, hold institutions accountable, and shape the frameworks that govern how society works. You’re changemakers — and you deserve tools built with the same seriousness you bring to your work. don’t just need answers. You need answers you can defend.

You’re in good company: a million professionals like you have already chosen to trust ¶¶ŇőłÉÄę with the work that can’t afford to be wrong.

We built Thomson for you. Not because AI is a trend worth chasing, but because the next phase of professional AI won’t be defined by general capability — it’ll be defined by domain mastery, measurable outcomes, and trust. That means higher-quality outputs and greater confidence when you’re using AI in the workflows that matter most.

You shouldn’t have to settle for AI built for everyone. Now you don’t.

Thomson

Thomson

The purpose-built, proprietary LLM, engineered for high stakes professional work

Learn more ↗

The post appeared first on .

]]>