US Local Media Outlets Sue Microsoft and OpenAI Over Copyright Infringement
US regional publishers, led by Emmerich Newspapers, have filed a lawsuit against Microsoft and OpenAI for allegedly scraping tens of thousands of copyrighted articles without authorization to train AI models. This lawsuit highlights the mounting legal pressures tech companies face from media publishers of all sizes over training data governance. If courts rule against AI vendors, companies may be forced to alter how they collect web data and enter into costly licensing agreements with content creators. The plaintiffs, including Emmerich Newspapers and Coopwood Media Group, allege that the tech companies bypassed paywalls and violated the Copyright Act and the Digital Millennium Copyright Act (DMCA). In addition to monetary damages, the publishers are seeking an injunction to compel Microsoft and OpenAI to remove the infringing training data from their models.
## BACKGROUND
Training large language models requires vast amounts of text data, which tech companies routinely gather by web scraping news outlets and public websites. While AI developers frequently argue this scraping falls under 'fair use,' news publishers contend that using their proprietary writing to build multi-billion dollar commercial products without permission or financial compensation constitutes unlawful copyright infringement.