Thomson Reuters Releases Thomson-1.0-Small, a Specialized Law and Tax LLM
Thomson Reuters has released Thomson-1.0-Small, a 35B Mixture-of-Experts (MoE) large language model optimized specifically for the legal and tax domains. The model was developed in collaboration with Safe Sign, an AI startup acquired by the company, and is trained on a subset of Thomson Reuters' proprietary legal data. This release represents a significant step in domain-specific AI, offering legal professionals a highly specialized tool that outperforms general-purpose models in legal reasoning. It also demonstrates how major publishers can leverage proprietary datasets to build cost-effective, high-performing industry models. The model features a 262,144-token context window, achieves a 74.6 average score on Hugging Face benchmarks, and is released under the Polyform Strict 1.0.0 license. It cost approximately $40 million to train, using less than 10% of Thomson Reuters' legal data, and is first being deployed in CoCounsel Legal's Tabular Analysis.
## BACKGROUND
General-purpose large language models often struggle with the complex terminology, strict reasoning, and compliance requirements of specialized fields like law and finance. Domain-specific LLMs address these limitations by training or fine-tuning models on industry-specific datasets, allowing them to perform specialized tasks with higher accuracy and fewer hallucinations.