Thomson Reuters Launches Proprietary AI Model for Enhanced Legal Advice

    Thomson Reuters Launches Proprietary AI Model for Enhanced Legal Advice

    Thomson Reuters Corp., a leading global content provider, has unveiled Thomson, its inaugural proprietary large language model, which integrates the company’s extensive legal knowledge with external LLMs to facilitate legal advice. The model will initially be implemented in Tabular Analysis, enhancing document review capabilities within its CoCounsel Legal AI assistant. CoCounsel will continue to operate as a multimodal product, using Thomson for tasks where it excels, while using third-party models for other functions. Thomson will serve as the default model for Tabular Analysis, with options for administrators to choose alternative models.

    The investment in the Thomson project has reached approximately $40 million over two years, with the final training run costing about $450,000 due to efficiencies achieved. Rather than developing a foundational model from the beginning, the company opted to use an open-weight model as a starting point, supplementing it with its proprietary content, training techniques, and expert knowledge. This strategy has lowered the costs associated with both training and inference when compared to more generic models.

    According to Joel Hron, global head of artificial intelligence and TR Labs at Thomson Reuters, the company does not aim to compete across all areas with the largest AI labs. “Thomson must set the frontier of intelligence for legal,” he stated, emphasizing a focus distinct from that of many leading AI labs.

    The training regimen for Thomson involved aligning the base model with the company’s core values, pre-training on proprietary content, and using professional guidance for targeted post-training. This was supplemented by reinforcement learning to improve its compatibility with key company tools including Westlaw and Practical Law, which boast a vast repository of over 40,000 databases and 150 years of legal publishing expertise. Hron noted that hundreds of legal experts contributed to defining training goals and assessing legal query examples.

    Jonathan Schwartz, head of foundational research at Thomson Reuters, cautioned that excessive specialization might undermine a model’s broader capabilities if not carefully managed. To avoid this, the research team emphasized continual learning, allowing for the addition of domain-specific skills without compromising existing abilities. “Without these enhancements, an open-source model wouldn’t perform as well or reflect our values,” Schwartz explained.

    Internal evaluations indicate that Thomson exhibits competitive performance levels with leading models, particularly when restricted to web-based access. Its performance reportedly improved to match or slightly surpass that of other models when accessing Thomson Reuters content, according to Andrew Bean, a senior research scientist at the company. These evaluations considered both the comprehensiveness of responses and the reliability of citations. A technical report detailing further benchmarks is anticipated in the future.

    While these findings have not yet been extensively validated by independent parties, Thomson Reuters is currently sharing the model with legal professionals and academic institutions for evaluation. Plans are underway to release a more compact version on Hugging Face under a noncommercial academic license, and a portal is being developed to allow external developers to request API keys and directly test the model.

    So far, only about 10 percent of the company’s complete information pool has been used, according to Bean, who highlighted that the next phase will focus not merely on increasing content but on enhancing training signals from the most valuable material and product activities. Ownership of the model supports the company’s stance on sovereign AI, asserting that client data is not used for training while maintaining control over deployment and governance. Discussions are ongoing with large law firms and corporations regarding direct model access, with a potential for clients to customize Thomson according to their specific knowledge and workflows.

    Hron recognized the challenges associated with managing a proprietary model in a rapidly evolving AI landscape, suggesting that enhancements in open models could provide a robust foundation for future iterations while ensuring that the company concentrates on professional applications. “Owning an AI model that encapsulates our expertise is integral to our mission,” Hron concluded, framing AI as a revolutionary tool for delivering expertise.

    Leave a Reply