~/AI POLICY/unsealed-court-filings-reveal-openai-s-internal-concerns-over-news-content-scraping

Unsealed Court Filings Reveal OpenAI's Internal Concerns Over News Content Scraping

Newly unsealed court documents from The New York Times' lawsuit against OpenAI and Microsoft disclose internal communications showing OpenAI executives were aware their models posed an existential threat to publishers. The filings also allege OpenAI actively used techniques to bypass paywalls to scrape copyrighted content. These internal statements could significantly weaken OpenAI's core legal defense of 'fair use' in high-stakes copyright litigation. If the courts rule against AI developers, tech companies could face billions of dollars in damages and be forced to retrain flagship AI models without proprietary publisher data. The documents quote an OpenAI executive acknowledging that AI tools are 'largely substitutive' of news material, while Microsoft internal data showed AI answer tools cut user click-through rates to original publishers by 83% to 93%. Additionally, Microsoft's CEO Satya Nadella testified he would have demanded model retraining had he known OpenAI was scraping behind paywalls.

## BACKGROUND

Generative AI models require immense volumes of text data from the web to learn patterns and generate human-like responses. Tech companies maintain that training on public web data constitutes fair use under US copyright law, while publishers argue that AI systems infringe on their intellectual property and directly destroy their subscription-based business models.

## KEYWORDS

#AI Policy#OpenAI#Copyright Law#Tech Lawsuits

$ subscribe --daily

Unsealed Court Filings Reveal OpenAI's Internal Concerns Over News Content Scraping | Daily News