Thomson 1: The New $40M Legal AI Model | Thomson Reuters Joel Hron
Key takeaways
- Thomson Reuters spent roughly $40 million over several years on people, compute, and training methods to build Thomson, its legal AI model, but the final training run cost about $450,000. Joel Hron’s explanation is that open-source general intelligence tracks the frontier alongside it, if not a couple of months behind, so TR’s post-training rides that tide instead of paying for it.
- The pipeline starts from Qwen 3.5, realigned with Imperial College into a base model TR calls Snowdon, then continuous pre-training on less than 10% of the Westlaw, Practical Law, Checkpoint, and Reuters News Archive corpus, fine-tuning and direct preference optimization with in-house experts, and agentic reinforcement learning with Westlaw and Practical Law as tools.
- The biggest accuracy jump in the episode came from the harness, not the model. Rebuilding CoCounsel around agent-native tools took one-shot accuracy on CoCounsel Bench, an internal benchmark of about 3,000 legal queries, from about 25% to “north of 70%,” which Hron says happened “almost overnight” once the agent had the right tools and basic instructions.
- Hron’s case for owning model weights is about where expert judgment goes. On a rented model, an attorney’s A/B judgment improves a prompt or a skill and is lost at the next model release; on owned weights it compounds. His analogy: renting a house gives you a roof, owning one builds equity.
- Continual learning was an explicit training objective. Hron says the technical paper’s comparison of Qwen, Snowdon, and Thomson across about 10 or 11 dimensions shows gains in legal, tax, and journalism with only small regressions in coding and math, while long context and other general dimensions improved.
- In legal AI the hallucinations hardest to catch are misinterpretations of the law, not fabricated case names, which Hron calls “a very bad thing” but says TR sees at “a pretty low rate.” Because law has no automated verification loop like code, TR relies on in-house former practitioners for verification, a patent-pending citation ledger in the CoCounsel harness, and a Litigation Document Analyzer in Westlaw that checks a brief claim by claim.
Concepts in this episode
AI terms discussed here — each links to a plain-language definition.
AI EvaluationFrontier ModelAccuracyAI AgentFine-TuningAgent MemoryAI BenchmarkFoundation ModelLatencyModel Context Protocol (MCP)
Chapters
- 0:00Cold open: a $40M model and 25% to 70%
- 0:27Why Thomson Reuters built the Thomson model
- 2:46From information services to an AI company
- 5:14The flywheel: compute, data, and expertise
- 8:21The oldest company to ship a model?
- 9:47Training for users without catastrophic forgetting
- 14:37Continuous pre-training on Westlaw and Checkpoint
- 15:29Fine-tuning, DPO, and agentic reinforcement learning
- 17:37Rebuilding CoCounsel: 25% to 70% overnight
- 21:48Capturing expert judgment: own versus rent the model
- 28:59Managing lawyer time and protecting customer IP
- 31:45Eval results and avoiding catastrophic forgetting
- 34:34Tabular analysis and legal deep research
- 36:26Benchmarks, Harvey, and frontier comparisons
- 39:26Verifying legal work with no ground-truth oracle
- 41:54Citation ledgers and the hallucinations that matter
- 45:41Rebuilding the platform and the Trust in AI Alliance
- 48:59Advice to CTOs on open models and owning intelligence
- 51:31The compounding flywheel and what comes next
Show notes
Behind Thomson, the new legal AI model from Thomson Reuters, is a $40 million investment in people, compute, and evaluation methods. The final training run cost just $450,000. CTO Joel Hron, whose teams build Westlaw, Practical Law, and CoCounsel for millions of professionals in more than 100 countries, joined us for the launch to break down why the 175-year-old company chose to own its model layer instead of solely renting frontier intelligence.
We cover:
- Why Thomson Reuters trained its own model instead of relying only on Claude, GPT, or Gemini
- The compute, data, and expertise flywheel behind the Thomson model
- How rebuilding CoCounsel around agent-native tools took one-shot accuracy from 25% to over 70%
- The rent-versus-buy case for owning model weights and compounding expert feedback over time
- The dangers of AI inaccuracies in legal work
- Citation ledgers, deep research, and verifying legal work with no ground-truth oracle
- Joel's advice to CTOs weighing open models and training on their own data
Chapters:
(0:00) Cold open: a $40M model and 25% to 70%
(0:27) Why Thomson Reuters built the Thomson model
(2:46) From information services to an AI company
(5:14) The flywheel: compute, data, and expertise
(8:21) The oldest company to ship a model?
(9:47) Training for users without catastrophic forgetting
(14:37) Continuous pre-training on Westlaw and Checkpoint
(15:29) Fine-tuning, DPO, and agentic reinforcement learning
(17:37) Rebuilding CoCounsel: 25% to 70% overnight
(21:48) Capturing expert judgment: own versus rent the model
(28:59) Managing lawyer time and protecting customer IP
(31:45) Eval results and avoiding catastrophic forgetting
(34:34) Tabular analysis and legal deep research
(36:26) Benchmarks, Harvey, and frontier comparisons
(39:26) Verifying legal work with no ground-truth oracle
(41:54) Citation ledgers and the hallucinations that matter
(45:41) Rebuilding the platform and the Trust in AI Alliance
(48:59) Advice to CTOs on open models and owning intelligence
(51:31) The compounding flywheel and what comes next
Connect with Joel Hron:
- LinkedIn: https://www.linkedin.com/in/joel-hron-90a3421a/
- Thomson Reuters AI: https://www.thomsonreuters.com/en/artificial-intelligence
- The Thomson model and next-gen CoCounsel Legal: https://www.thomsonreuters.com/en/press-releases/2026/august/thomson-reuters-launches-next-generation-of-cocounsel-legal-the-ai-ecosystem-built-for-legal-professionals
Connect with Chain of Thought host Conor Bronsdon:
- Newsletter: https://newsletter.chainofthought.show/
- Twitter/X: https://x.com/ConorBronsdon
- LinkedIn: https://www.linkedin.com/in/conorbronsdon/
- YouTube: https://www.youtube.com/@ConorBronsdon
More episodes: https://chainofthought.show
Our sponsors:
Transcript
40 segmentsJoel Hron 0:00 we probably spent somewhere around $40 million in the last several years. We could answer about 25% of those questions in one shot on the prior version of CoCouncil. And when we built this new version, almost overnight, we got north of 70%. And what we see is that Thompson outperformed some of the best models in the world on deep research. This is not the end for us.
Conor Bronsdon 0:27 Labs spend billions of dollars training frontier models. Thomson Reuters says the final training run for its new Thomson model cost about $450,000, building off of the Imperial College's Snowden open-source model from the QEN 3.5 family and adding the company's proprietary content and expertise. Early evaluations put the results on par with leading frontier models across a range of tasks, with stronger performance in areas where Thomson Reuters' proprietary content and tools come into play. Welcome to Chain of Thought. I'm your host, Conor Bronsdon, here with a model release episode. Make sure you have subscribed and you've turned on downloads or new notifications on YouTube. Joining me today is Joel Rahn, the Chief Technology Officer at Thomson Reuters, where he leads product engineering across legal, tax, audit, and compliance, including Westlaw, Practical Law, and Co-Counsel, which is used by millions of professionals across over 100 countries. Joel and Thompson Reuters are betting that in an era where renting significant intelligence is available to everybody, the advantage has moved to orchestration and data, knowing which intelligence fits which work and which layers of the stack you need to own yourself. Joel, congratulations on the launch of Thompson today. How are you feeling?
Joel Hron 1:44 Thanks. I'm super excited about what's ahead. Feeling good, Conor. Thanks.
Conor Bronsdon 1:48 I'm loving that we are getting these bespoke, highly intelligent model launches, and I think Thompson is a great one to unpack. Before we do that, though, I do want to dive into the details of Thompson, but I want to say a quick thank you to our presenting sponsors for season four of Chain of Thought. Sphix, delivering reliable enterprise webhooks at link.sbix.com slash C-O-T, and Walrus Memory, giving AI agents portable, verifiable memory at walrus.xyz slash C-O-T. All right, Joel, so let's set the stage a bit before we talk in depth about the model, because I suspect a lot of our audience knows Thomson Reuters mainly as a news brand, and that's a small slice of the company. How does an information services company become an AI company? And what does agentic AI change about how you can leverage all the data and the portfolio of Westlaw, et cetera?
Joel Hron 2:46 Yeah, so I'll start at the beginning there. I mean, a lot of people know the brand of Reuters, you know, as a media company, as you would expect. But as you said, you know, the media business is actually a relatively small portion of our overall business. The primary parts of our business, as you said earlier, legal, tax, compliance, audit, risk, across all of those categories, we're one of the effectively top three software providers in the world across each of them. And so we have quite significant businesses across each. And if you look at the common thread that ties them all together, and really the history of Thomson Reuters as a company, it's really been about content and expertise. We have been an information services company for a long time, delivering trusted knowledge and trusted expertise to help practitioners in each of those areas do their job and do their job well. And so that's really, I would say, the thread that ties all of those things together. Now, you know, we have been, I think, under transformation for really the last 10 years as a company, transforming from an information services business into a technology business. And really what that means is that, you know, as you know, print fell off and digital increased, like software became the delivery mechanism for that technology and for that expertise. And as we look at ourselves now, as you say, as an AI technology company, AI is a delivery mechanism for the knowledge and the expertise that TR possesses and serves into the market. And AI, in fact, is like a tremendously effective delivery mechanism for the type of knowledge and the type of expertise that we've built over many decades. And so That is really how we see ourselves evolving, certainly with AI as a central foundational part of what we build, but ultimately the goal and the objective for what and why we build is to deliver expert knowledge and information to the practices that we serve.
Conor Bronsdon 5:03 And this is what drove you to create Thompson. Can you tell the audience a bit about this new launch and what they should expect from this uniquely positioned new model?
Joel Hron 5:14 Yeah, so, you know, when we started the development of Thompson, you know, we really sort of looked at AI and we said, okay, like, what's driving the flywheel of these capabilities? And it really boils down to kind of three main components. It was compute, which everybody talks a lot about, it was data, and it was expertise, and particularly, like, top-tier expertise. And, you know, from a data and expertise standpoint, really what separated those things is not just like more internet data or more of Reddit or more of anything that everybody in the world has access to, but really like unique specialized data and unique specialized expertise that was atypical. Atypical meaning it was out of distribution of everything else that's easy to get, and it moved the frontier. And that was really what was, and still is, I think driving the flywheel of AI capabilities, is unique data, expert data, expert judgment that continues to move that frontier beyond where it is today. And we have those things better than anybody else in the world, in law, in tax, in compliance, audit, et cetera. And so our belief was that if we can really leverage those things, the most, I think, impactful way to leverage those things is through a model itself, like the direct intelligence of the model. And so that's really why we set out on this. The compute side of that equation was certainly, I think, a question that we had in our heads around, like, could we keep up with the compute demands of training models of this size? And I think really the bet we made, if you will, was that the open source community would continue to grow and evolve and keep pace with the frontier. And I think over the last several years, there's certainly been, I think, peaks and troughs of that. But if you plotted kind of the average intelligence, I think you would see pretty consistent tracking of open source capabilities, you know, alongside frontier, if not just a couple months behind. And so that was really the bet we made in that if open source general intelligence continued to increase at that pace, then the specialized intelligence that we could add to it with our content and our data and our expertise would continue to stack on that. And as the tide of open source rose, our ability to post-train our unique knowledge on top of that would rise as well with that tide. And I think we've seen that continue to prove itself out. And that's one of the reasons that our post-training, mid and post-training methodologies around Thomson are quite cost-efficient, because we do benefit quite a lot from all of the great work that's happening in the open-source community.
Conor Bronsdon 8:21 I wonder if Thomson Reuters is the oldest company to release a model, because Thomson Reuters is what, 175 years old?
Joel Hron 8:31 Yeah, yeah. I mean, maybe more actually. We were flying carrier pigeons across the sea to deliver news way back in the day. So yeah, pretty old. I mean, we're probably one of the oldest companies on the planet. So I would venture to say the odds are pretty good we might be the oldest, but I don't want to state something I factually don't know.
Conor Bronsdon 8:54 I think it's really interesting though, because it indicates how willing Thomson Reuters is to transform and to adjust and leverage this massive trove of data that you've built over time. And it also speaks to exactly what you were saying, which is that the open source frontier models are enabling so much innovation for companies that are not open AI, are not anthropic, are not oriented all around, let me build a model. I think that's just such an exciting moment for us to be in where it is realistic for a company that is intentional and has great data and has a great team to train an excellent model for what their users need. Talk to me a bit about how you went down this process and how this came about and the training piece.
Joel Hron 9:47 Yeah. I think he said something important in there, you know, specifically for what their users need. And I think, I think that really grounded our approach to training. You know, our approach to training was to say, Hey, look, we want to take a great open source foundation and add to it what our users really care about. And. ideally not give up anything that that open source foundation brings to the table. And so what we put a lot of focus on was really, I would say, two elements of training. One is bringing certainly our data and expertise to the table and curating very good data sets that are representative of the types of jobs our users and our products are trying to do. good evals that represent those jobs equally well. And then the second element of that was really developing training methodologies that add that data while not regressing or, you know, as the industry would say, catastrophically forgetting what it already knows. So these models are trained and instruction tuned around things that Thomson Reuters fundamentally is not the best in the world at. Things like coding or maths or other like general reasoning or instruction following capabilities. And, you know, the beauty of the open source is we get those capabilities for free. And so when we apply our data to it, we don't want to lose those capabilities. Like we need the model to still be generally fluent in these kinds of other things. And that, I think, is where, you know, I don't want to say the majority, but I would say a significant part of the science and engineering behind our training pipeline comes into play is how to do that effectively. from an algorithm standpoint, from a data mixture and selection standpoint, both in terms of continuous pre-training as well as post-training. And really, what I think I'm most excited about is that All of that work really has built somewhat of a pipeline, from model realignment, to continuous pre-training, to agentic reinforcement learning, Like that pipeline exists and we can continue to add our own data and our own human judgment and our own preferences to it. And at the front of that pipeline is an open source model. And so as both of those things get better and grow over time, like we apply the same pipeline and can really stack and scale on it over time. And so I think that's what's really exciting is that it's quite repeatable. And I think it gives us a flywheel to continue to improve this over time. I guess, you know, talking maybe a little bit more specifically about the training, like I said, there's really a few phases. It starts with an open source model and we sort of, you know, review and select what the best open source models at the time are on the things that we care about. We started training the version of Thompson we're releasing now back in back in June. And at that time, Quen 3.5 was really one of the best models in the world that we could see. I mean, in the last two months, we've had at least a half dozen, if not more, open source models that have been released and surpassed that. So we'll keep looking at that. But at that time, this was Quen 3.5. And the first thing we do is really, for any open source model, realign it to safety and ethics and values that TR represents. And so this is work that we jointly did with Imperial College to develop a practice to realign any open source model to sort of a de-biased standard that we can stand behind. And that was a really important part of work because it gave us confidence to be able to pick up virtually any open source model and get it to a state where we felt good with it being our starting point. And that's what Snowdon is. It's effectively a TR realigned version of Quen. that serves as Thompson's starting point model.
Joel Hron 14:37 The next thing that we do in terms of model training involves continuous pre-training. And so what we do is take the corpus of Thompson-Porter's content from Westlaw practical law, our tax research content, which is called Checkpoint, the Reuters News Archive, and we go through a data selection process to identify what the most critical pieces of that raw content are to perform continuous pre-training on. And to this point, we started small there, but less than 10% of that corpus has been applied for continuous pre-training. So that's an element of the training process that we think will increase over time.
Joel Hron 15:29 The next stage is really around fine-tuning and direct preference optimization. And this is an area where our human experts are very heavily involved in terms of curating datasets that represent the types of tasks that we want the model to perform well at, as well as evals that represent those tasks as well. And then the last phase is really agentic reinforcement learning, where we will put the model into a harness, like a very simple agentic harness, with access to our TR applications and tools. So products like Westlaw or Practical Law. are made available to the agent through the harness, and the model learns how to use these tools effectively to do its work. And this is really important because we're publishing thousands, if not hundreds of thousands of cases a day. There's no practical way that we're going to keep the model current with the ever-changing dynamics of case law. And so we teach the model how to be really good at using Westlaw to stay current on the law or stay current on the tax side of things or stay current on the news. And the model learns how to use those tools effectively through the training process as well. And this is another area we really see us scaling and improving over the course of time as well.
Conor Bronsdon 17:01 This is so interesting because it feels like a rebirth of the Thompson writers portfolio in a way where you're taking all these well-known legacy products and decomposing them into tools that your agents can use and then also enabling your model to grow and learn off of this corpus of information. How did this work in practice? What were the implications of this on things like your new generation of co-counsel legal and everything that is coming alongside the Thompson model?
Joel Hron 17:37 Yeah, so CoCounsel Legal is our product. It is the application that users go to and can work with what we believe is the best legal agent in the world. And Thomson is the model. Now, when you say rebirth of the Thomson Reuters product portfolio, I think that's a good description of what we've done here. But that really happened at the co-counsel stage. So co-counsel is a product today that's really built around the Clawed Agent SDK, and in fact, heavily the Clawed models, but also other models from OpenAI, Google, etc. And that product works very much the same way that I just described the training of Thompson. We have decomposed Westlaw and practical law into agent-native tools that are built into the harness of co-counsel. And co-counsel is sort of, and the harness in particular, is really trained around using those tools effectively to conduct the legal work. And as we did that for co-counsel, we found that it was like night and day. Our prior versions of CoCouncil that were sort of not, I'll say like coding agent native architectures, you know, when we look at, we have a benchmark for like all of legal tasks we call CoCouncil Bench. It's about 3,000 different kind of legal queries we add to it almost every day. But we could answer about 25% of those questions in one shot on the prior version of CoCounsel. And when we built this new version of CoCounsel, almost overnight, we got north of 70% just by giving the agent the tools it needed to sort of do the work and some basic instructions on how to use those tools. And so we were very encouraged by just that. Now, Thompson kind of takes that idea one step further and says, I'm not just going to write some skills and some system prompts on how the harness should use these tools. I'm going to actually train the model to use these tools more directly. And, and I think that's really like where we see the next horizon going. I think, importantly, this doesn't require that Thomson fully replaces Cloud or OpenAI or Gemini, you know, from the stack. Like, you know, the first implementation of Thomson and CoCounsel is actually for bulk document review. So, like, think about large volumes of documents that need to be reviewed for due diligence or something like this. That's where Thomson will ship first next week when we launch it into CoCounsel. And I think we see a progressive path with Thompson to be able to just take on more and more components of that work, like components of search and re-ranking or citation analysis, all of these functions that happen underneath the agent. And really what the frontier of general intelligence is good at is kind of being an orchestrator on top of that. And so I think for us, like we see a path to improve Thompson and continue to funnel it into the larger agent ecosystem of co-counsel rather than like a binary decision of this model versus that model. And I think that's how we'll progress the training of Thompson over time as well.
Conor Bronsdon 21:13 It's interesting to think about how you then apply the human intelligence layer that Thompson has in the, I have no idea how many lawyers and tax professionals, perhaps more than any company in the world or more than most, that's for sure. How did that workforce apply their own feedback to the model and What did that cause over time as you started to bring in extensive human feedback?
Joel Hron 21:48 Yeah, I mean, I think that is the
Joel Hron 21:54 differentiating factor on a model like Thompson at the end of the day. And so we have set up, we've done this for a long time, like we've used our subject matter experts to drive our product development for a long time. And so in terms of like annotation tools and things like this that support like our SMEs building these kinds of tasks, Like, we actually had a lot of systems in place to support that. And it was really about just modeling the appropriate tasks for them to go represent and recreate and standing up sort of RL environments for them to kind of participate in that flow. So, you know, that was a muscle that we had. I think we could certainly get better at it, as any company can, but I think we've done that quite well. I think the real value that I see with owning the model layer, though, is if you think about the development of CoCounsel, the product, what does that look like today? So we have CoCounsel Bench and we evaluate the agent running, you know, whether it's Claude or OpenAI or another model. we evaluate the agent against co-council bench and we say, okay, here are the sort of like, you know, X percent of things that didn't do well and we want to hill climb on these dimensions. And so what happens is our engineering and our science teams go in and they basically refine some system prompting or perhaps they write some new tools for the agent or write some new skills that help the agent use those tools better or make some tweaks to the harness itself. And these are the knobs that they're turning to try to get the agent to do better. Or they hope that OpenAI ships a new model that helps the agent get better at that thing naturally. So these are the knobs that are at our disposal today to hill climb CoCouncil Bench. I think the real, and what happens through that process is effectively a lot of A-B analysis. We turn some knobs, like I just mentioned, and we present that to our SMEs, and they basically adjudicate, hey, these, you know, answer B is better than A for these reasons, and we do that a lot of times. Certainly there's automated evals, but a lot of this comes down to human judgment and preference and feedback and ultimately we reach a candidate that says oh when we turn these knobs it was better and we like this and we're going to deploy that new version of the agent. And so this is the process of development when you don't own a model. But all of those interactions of the SME in the middle are data. Like every time I give like an A-B evaluation and an SME says I like answer B because it contained this and it didn't contain this and it said this this way, That is amazing data that encapsulates the judgment of a very well-practiced attorney in a really, really unique way. And I think what we really see as the opportunity for Thompson is to capture that data and apply it directly to a model and have that compound over time. Versus when you're just doing that with third party models where you don't control the weights directly, all of that intelligence and knowledge that is going into iterating the product is just exhausted. Like you make a change to the product and all of that other work you just did goes away because the next model that comes out changes it and you need to move from there. And so we really see this as a way to capture that knowledge and intelligence throughout those evaluation processes and compound it into the intelligence of a model over time. I've used the analogy before of renting a house versus buying a house. If you rent a house, you're paying a landlord and you've got a roof over your head and the landlord takes care of the house and makes sure everything works and that's great and that works, but you build no long-term equity. Whereas if you own a house, you may pay the same amount, maybe some things are slightly more expensive, you have some headaches that you didn't deal with before. But you are compounding that over time into something that is valuable and arguably increasing in value via that equity. And that's kind of how we think about model ownership. And again, I think model ownership is really focused on what are the things that we do best as a company? What are the things that our experts know best? And those are the things we really need to compound into the model. Everything else, we leave that to the Frontier Labs to solve and for the open source models to solve, potentially. We're not so focused on trying to be the best in the world at all things. We're really focused on being the best in the world at the things we know better than anybody else.
Conor Bronsdon 27:04 Season 4 of Chain of Thought is delivered by Sphix. We spend a lot of time talking about what agents need in production, and one of the least glamorous answers is events. Your customers want agent workflows that react to things happening inside your system, which means your API needs webhooks that actually work. Not just a post request and a prayer, retries, ordering, idempotency, replay protection. Sphix does that as a service, and they wrote standard webhooks, the spec that Anthropic, OpenAI, and Google bailed against. So if your API doesn't have reliable webhooks, that's turning into a lost deal. Join Brex, Dorada, Daytona, and many others on Stix. Get started at link.svix.com slash c-o-t or go to the show notes to grab the link. Qualified startups will get $12,000 in credits. $50,000 for YC companies. I can't recommend Stix enough. I'm a huge fan of their open source project. I've actually contributed a bit myself and they're so easy to integrate with. I think you'll really enjoy it. Check out Sphex at link.svax.com slash C-O-T. And I think West's law is an obvious example here of this incredible decades-long human layer of annotating the law, of citations, notes. And the selective data advantage that I can see you applying to how you're training the Thompson model clearly is built upon this massive corpus of information, not to mention, as you put it, the human hours that you're putting in from lawyers. I can imagine there are people challenges though, because lawyers are busy people with expensive hourlies. How do you manage this in the context of a major model effort while also applying the already annotated data that is being put into these systems?
Joel Hron 28:59 Yeah. So, I mean, I do think it's a mix. Again, this is one, well, there's really two reasons that we lean on our internal experts for this kind of work. A, you know, they're former practicing attorneys or practicing CPAs, things like this,
Joel Hron 29:19 who work directly for us. And so, like, this trade-off of, like, I could bill this hour to a client or I could work on this model doesn't exist for them. Their job is to help TR curate excellent content for its customers, and content comes in the form of maybe a document or a practice note that they write, or it comes in the form of spending a couple hours building a multi-turn eval example for our training pipeline. And our SMEs are becoming fluid at doing both of those tasks as part of their job. That is the expertise that they bring to the table. The second reason it's important that we use our own SMEs is really data and privacy protection for our customers. I think our customers are quite rightly so in tune to protecting their IP as a company and sort of the way that they do things, the way that they practice, and the unique things that they bring to the market. And so we don't use any customer data in sort of informing the process of how Thompson gets trained or anything like that. But we use our internal experts to drive that process for that reason, because I think our customers are rightly so, I think, very conscious about how their IP and their secret sauce is being protected. And we want to be certainly respectful of that in how we train. That's really what we focus on from a human element. And I think our domain experts are probably best asset, even more so than our historical archives of content. The shape and type of data that moves the frontier in many ways is different than the shape of data that our historical content looks like. And it's really important that we continue to leverage those experts to kind of create this new data that continues to move the frontier for us.
Conor Bronsdon 31:28 Let's talk about the results of all this. So I mentioned at the start of the show that the evals for Thompson look very promising in different dimensions. Can you share a bit about the eval methodology and where Thompson is focused as far as evaluations?
Joel Hron 31:45 Yeah. So you'll see, we'll publish a technical paper next week that'll have a lot of detail on these evals and what we're seeing. I think I'll talk about two different dimensions of this. One is kind of the legal specific domains that we've had focus on. And the other is this idea of like catastrophic forgetting and how we prevent against that. In terms of the catastrophic forgetting, there's a really nice graphic that you'll see in this technical paper that's sort of a spider diagram of Quinn, Snowdon, and Thompson models across I think it's about 10 or 11 different general dimensions. So one of those dimensions would be legal, one would be tax, one would be journalism, one would be math, coding, instruction following, general reasoning, and so this is sort of like a a plot that shows what is the delta between Thompson and Quinn on each of those dimensions. And what you see is something very interesting, is we see that Thompson clearly gets better in legal as well as tax and journalism as you would expect. But I think the more important thing is that it does not regress in any of those other areas. We see a very small regression in coding and a very small regression in math, but actually all those other dimensions got better. Long context, all these kind of things. And we see this really nice halo effect through the training of Thompson. in these general dimensions versus just really spiky performance in legal. And that was really an important part of the training objective and the evals that we ran support that. So I would definitely call that out. I think it's a good indication of of how we have, and this paper talks a lot about continual learning. The idea of continual learning is that you don't lose what you learned before, you just learn something new. And that's been like a real core focus of Training Thompson. The second has been sort of on legal specific benchmarks. And there's a lot of public legal benchmarks certainly that we've used to help evaluate Thompson. We have focused a lot on internal benchmarks as well. So two things I'll mention. One is tabular analysis. So we'll be launching tabular analysis into
Joel Hron 34:34 CoCouncil next week. This is an area where you have high volume document review. So there's many, many calls, like speed and latency are quite important here, as well as cost. So this is a natural first place to drop Thompson. And what we see is that Thompson was outperforming the prior model versions that we were using by about five percentage points in terms of accuracy. So we were very encouraged by tabular analysis first and foremost in terms of answer quality, given the volume that is associated with that kind of task, which is why that's sort of the first production launch use case for us. The second thing that we see within legal has really been focused around deep research. So deep research in legal is very similar to what you might think of as deep research on the web, where a model within an agent harness goes out and you know, searches and synthesizes, searches and synthesizes, you know, recursively until it reaches an answer. And so we have built effectively deep research in the same process on top of our legal content. And what we see is that Thompson outperforms some of the best models in the world on deep research when given access to our content, even though these models are much larger in size than Thompson. that reinforcement or that that agentic reinforcement learning really helps support Thompson learning how to use these tools effectively to do these kind of tasks. And so that's another data point that'll be in this paper that I think we're really encouraged about in terms of like training direction for Thompson as well.
Conor Bronsdon 36:26 Yeah, the context harness you provide a model and let it continue to learn on is super important and definitely where I see Thomson Reuters having an advantage. What about some of the broader benchmarks? Because I know, I think the obvious other name that comes up in when people talk about like legal models is Harvey and I know they've just released a model as well. How are you comparing against them and some of the frontier models like Fable 5 or GPT 5.6 Sol when it comes to some of the broader benchmarks that are out there?
Joel Hron 36:58 Yeah, I think, you know, I would say, like, across those broader benchmarks, we see our model, like, in the top 5 to 10 of those best models in the world. Again, like, the broader benchmarks, and also, you know, I don't know too much about Harvey's model, but I think it was a derivative of Kimi K3, you [37:22] Conor Bronsdon: [OVERLAP] It was, yeah. [37:22] Joel Hron: [OVERLAP] know, which is, like, several trillion parameters in size, like, Thompson's base model is a few hundred billion parameters in size. So like we're talking, you know, an order of magnitude smaller in terms of size. And some of that was just driven by when we started training back in June. I think we'll certainly like maybe progress our model sizes up. But again, our goal in training Thompson isn't necessarily to compete with Fable 5 across all dimensions of frontier intelligence either, right? I think our goal with Thompson is to build a model that's on the Pareto frontier of AI, incredibly intelligent and good at doing legal tasks at an extremely competitive price point. And I think when we look at sort of that Pareto frontier and where we want to be with our model, like, I think, I think we feel really, really confident in it. Um, and, and, and we will continue to use, I think, other frontier models, like, like, uh, whether it's five, six or other versions for like frontier intelligence kinds of activities in conjunction with Thompson. [38:35] Conor Bronsdon: [OVERLAP] I know you have some model routing built into a lot of your products too, depending on the task. So [38:40] Joel Hron: [OVERLAP] Right. [38:40] Conor Bronsdon: [OVERLAP] having Thompson here does enable you to simply be much more cost efficient and task efficient depending on your users needs. I also think another interesting element of this is something that you brought up with me when we were chatting before this episode, which is the idea that legal can just be trickier to understand for a model because code verifies itself. Does it build? Does it test? Does it run? It compiles. Law doesn't really have that same standard. It's [39:13] Joel Hron: [OVERLAP] Yeah.
Conor Bronsdon 39:15 much different. It's a very human activity. How do you build the verification loop for this domain that has no real ground truth oracle?
Joel Hron 39:26 Yeah, that's one of the reasons that our domain experts are so important is because there is no automated verification loop for these models like there is with code. Humans are still the primary source of that verification. And I think the more signal that we can get from our SMEs to help the model understand, like, what does verification mean? How do I verify myself? I think the better the model gets at doing that. In fact, within CoCouncil itself, we have a patent pending around really our approach around citation ledgers and citation verification. And so part of this is at the model layer, but part of it is at the application and harness level of how does the model go through the process of verifying the claims that it's making. We even released a product called Litigation Document Analyzer in Westlaw. And what that product does is it will decompose like a legal brief or a document similarly into every claim made in that document. And for every claim in that document, it will go out and it will try to find case law that supports the claim and reference that directly in a table effectively. And if it can't find sources that verify that claim, it'll say so. And so these are like, A, applications that we release to our customers to help them verify, but also these are similar like tools that we build into the agent itself to help it verify claims that it's making and sort of assertions that it's making along the way as well. And I think that that's part of the model layer, certainly in terms of training, but it's also part of like the harness and tool development that we think about as well.
Conor Bronsdon 41:31 And I know part of how you're tracking the success here is around citations and that there's been quite a bit of work that's been done here. Obviously, it's very important to the legal field. What's the citation approach? And can you tell us a bit about how Westlaw and Thompson are working together to provide citation ledgers and some of the patents you filed there as well?
Joel Hron 41:54 Yeah, so this is what I mentioned in terms of the citation ledger, and I think this is a really important part of the way CoCounsel works. Effectively, this is a register for the agent of like every case that it has inspected throughout its research process, and this ledger is used to really ground the agent in terms of the claims it's making. I should back up just a second. Hallucinations are one of the things that plagues the legal industry more than anything right now. And the truth is hallucinations that fabricate case names are really bad. Ideally, we don't want that to happen either. It does happen, but I would say more infrequently than you would expect or than you hear about in the news. generally fabricating case names is a very bad thing, but we see that at a pretty low rate. The bad hallucinations that are really difficult to catch are the things that are misinterpretations of the law. Like, I read this case, or the model read this case, And it took a more, let's say, broad interpretation of what the case implied than what it did. Or it interpreted this statement from the judge as fact versus opinion. These kinds of things might change the way that the agent operates. And these are still hallucinations in a way. And these are the things that I think lead to bad results at the end of the day on our models. And I think this is an area where Thompson can really help. This is one of those things that's really hard to teach a model unless you're really applying a high level of human judgment and guidance to the model as to how it should interpret these things narrowly or broadly and why. And I think that's one of the areas that we're optimistic in terms of the future training of Thompson being able to make it better at.
Conor Bronsdon 44:12 Season four of Chain of Thought is also presented by Walrus Memory. I spent much of this year talking to guests about agent memory, and we're all having challenges that go beyond simple session handoff. You explain your code base to one AI coding tool. You get it tuned to what you need. But when you look to hand off to a new model family, you find yourself having to confront the same memory and context problems. Walrus Memory is solving agent amnesia with a portable memory layer for agents. Walrus Memory takes your agent's memory anywhere you need it. The context your agent builds can be read back later, from a different app, in a different runtime, weeks out, and you set the rules for who reads and writes it. Python and TypeScript SDKs with native MCP support. I'm glad to have them behind Season 4 of Chain of Thought because this is a key problem confronting AI builders in 2026. Learn more at walrus.xyz.cot. Yeah, it's been interesting to look at some of the ways you are rebuilding Thompson Raiders platform to align to Thompson and to these models. You know, CoCounsel, which we've mentioned a couple times, runs on the Cloud Agent SDK. You were early development partners. I know that's been a big impact. I know you've rewritten the APIs for Westlaw, Practical Law, and other Areas of your platform would be more ergonomic for LLMs. How do these different changes you're making to the platform overall and some of the design partnerships you're doing inform the work moving forward?
Joel Hron 45:41 I mean, I think they're critical. We've been phenomenally good partners with Anthropic, as well as OpenAI and Google throughout the last three years. And just by way of example, we formulated what we called the Trust and AI Alliance with those partners, as well as AWS and a few others. really getting to this point of particularly in areas of law and tax where being good is not good enough. We call this fiduciary-grade AI, where these professionals really have a duty to be correct. And that involves Models that are accurate, but knowing that models will not ever be 100% accurate, it really involves transparency and verifiability on the user's behalf. And how do the models represent their uncertainty to the users? And how do the applications themselves represent that uncertainty to the users? So we formulated this alliance with that group of people and partners, which has been extremely, I think, useful and beneficial. I think we continue to collaborate quite closely with those model providers in terms of early access programs to their models and the direction of travel that they're taking, sort of their training and development of their models, the development of their various different agent harnesses, and what works for us, what gaps we're seeing. And so I think those have continued to be extremely valuable in terms of the feedback that we're able to provide to them, but also the insight that they're able to give to us about the direction of travel for their products and how we might shape our products to really ride the exponential with them, rather than every time they release a new model to have to go re-architect your product. kind of build your product in a way that rides the curve that you think that they're on. And so that two-way conversation has been really, really valuable for that.
Conor Bronsdon 47:49 Joel, thank you so much for taking us through the release of Thompson. I am looking forward to sharing the paper with the audience and diving more in depth into it myself as well. I do want to close by asking a couple of questions of you that are more general, because I think what you are doing and what the Thompson Reuters team is doing to create a specialized model for your areas that can help your customers have More efficient answers can provide speed, can also save the company quite a lot of money. We've talked previously to leaders like Intercom slash FinAIs, Chief Scientist Fergal Reed talking about how much money they saved. switching from GPT to Quen for some tasks. And you can check that episode out in our backlog. I think there's plenty of examples of this that are starting to come to fruition. But not everyone has made that switch yet. And for some things like coding tasks, most folks are still focused on frontier models. But what would be your advice to other CTOs and leaders who are thinking about when to use open models and are potentially considering training their own models off of their informational data?
Joel Hron 48:59 I think my advice would be to focus that work on the things that you do best as a business. I think sometimes you can look at the open source capabilities and your eyes can kind of get a little bigger than your stomach and you can try to go eat the world. And the truth is, like, Anthropic OpenAI, like, SpaceX, Google, are making just, like, incredible strides on the general frontier. And, like, I don't know, like, I've thought for three years that at some point you would see it slow down, and I don't think it has. And so I think that benefits everybody. I don't think people should sort of, like, turn a blind eye to that. But there are some things for every business, if you really think about it hard enough, that you know better than anybody else in the world. And for those things, it's not reasonable to expect that a general purpose model or a general foundation lab is going to get better than you at that thing. If they do, then you're probably in trouble. So you need to ask yourself, what is the thing that I know better than anybody else in this world? And how can I encapsulate that intelligence into a model? And does it make sense to do so? And I think that was really the root of our work here was to say, we don't need to be the best at everything. We need to be the best at the things we arguably are the best at and know the best of anybody else in the world. And let's really try to focus and harness our development of intelligence in that direction. And I think if you do that, you can really build a compounding flywheel that benefits you over time and builds real durability in your company. It's not like we did talk about $450,000. That's a trivial amount of money. We probably spent somewhere around $40 million in the last several years between people costs and compute and development of these methods and things like that. It's not a trivial endeavor to go out and do this, but I would say if you stay focused on what matters most to you and what matters most to your customers, you can build a really good muscle around this and have it be a compounding flywheel for you for a long time.
Conor Bronsdon 51:31 Fantastic. Joel, I really appreciate the conversation. Any closing thoughts that we didn't get to that you want to make sure to highlight?
Joel Hron 51:39 Um, you know, I would say, you know, like probably first and foremost is this is not the end for us. I think like we see this model release as a pretty, um, as a pretty impactful moment for TR as a company for sure. And we're like really, really proud of what we've done so far and we're excited to launch it into co-council. But I think, you know, what I personally am more excited about is sort of what I said just now is like, this is a compounding flywheel of benefit for us that I think will build on itself for years to come. And the more that we build the muscle of spinning that flywheel faster and faster with our internal experts, I think the better and better this can become. And I think that's what excites me the most is not sort of like the milestone we hit today, but sort of what that milestone implies about the next two or three or four or five that will knock down over the course of the next year. So excited about it.
Conor Bronsdon 52:40 Fantastic. I'm excited to see what's next for Thomson Reuters. And Joel, thank you for sharing this first look at the new Thomson model. What's the best place for everyone to go find out more information and learn more besides the link in the comments that we were certainly going to add?
Joel Hron 52:56 Yeah, I'd say throw the links in the comments. Check out our website at thompsonrogers.com. I'll be posting some things on LinkedIn, so you can follow me on LinkedIn as well. And then we'll be publishing the technical paper here as well, which I think will be a great read.
Conor Bronsdon 53:16 Fantastic. We will definitely have a link in the show notes and I can recommend Joel's LinkedIn myself as well. He's a good follow. Joel, great to see you. And for everyone who is listening, if you enjoyed this conversation, reach out to Joel and I, let us know what you enjoyed. Drop us a comment or maybe even leave a rating and review on the podcast, which always means the world to us. Like and subscribe, you know, all those good things. Thank you all for joining us at Joel. Thanks again.
Joel Hron 53:44 Thanks for having me, Conor.
Frequently asked questions
- How much did Thomson Reuters spend to train its Thomson model?
- Two numbers, and they measure different things. The final training run cost about $450,000, a figure Hron calls “a trivial amount of money.” The full program came to “somewhere around $40 million in the last several years between people costs and compute and development of these methods.” The run is cheap because Thomson builds on an open-source base rather than being trained from scratch.
- What open-source model is Thomson built on?
- Thomson started from Qwen 3.5, which Hron says “was really one of the best models in the world that we could see” when training began in June. TR first realigns any open-source base to its own safety and values standard in a practice developed with Imperial College; that realigned Qwen is called Snowdon and is Thomson’s starting point. Hron says the base has a few hundred billion parameters and that TR will keep re-evaluating base models as new open releases surpass it.
- Did the Thomson model produce the 25% to 70% accuracy jump in CoCounsel?
- No. That gain came from rebuilding CoCounsel as an agent-native product on the Claude Agent SDK, with Westlaw and Practical Law decomposed into tools the agent can call. On CoCounsel Bench, one-shot accuracy went from about 25% to north of 70% “just by giving the agent the tools it needed.” Thomson takes the idea a step further by training tool use directly into the model, and ships first in bulk document review inside CoCounsel.
- Why does Thomson Reuters want to own its model instead of using Claude or GPT?
- Hron frames it as rent versus buy. CoCounsel is still built on the Claude Agent SDK and runs heavily on Claude models, with frontier general intelligence acting as the orchestrator on top, but every expert judgment made while tuning a rented model is “exhausted” when the next model ships. Owning the weights lets TR compound that judgment over time, focused only on the domains where its experts know more than any lab: “We’re really focused on being the best in the world at the things we know better than anybody else.”
- How does Thomson Reuters verify legal AI output when there is no ground truth?
- Hron says there is no automated verification loop for law the way code has one, so humans remain the primary source of verification, and TR’s in-house subject matter experts feed that signal into training. In the product, a patent-pending citation ledger registers every case the agent inspected and grounds its claims against them, and Westlaw’s Litigation Document Analyzer decomposes a brief into claims and flags any it cannot support with case law.