Unsealed Court Documents Reveal AI Executives Admitting Mass Unauthorized Content Usage
Unsealed court filings in The New York Times lawsuit against OpenAI and Microsoft reveal quotes from top executives, including Microsoft's Brent Hecht, calling AI training data practices "the largest theft of the fruits of labor in human history." OpenAI President Greg Brockman also admitted in documents that generative AI acts as a direct substitute for publisher content, posing an existential threat to media companies. These executive statements significantly weaken OpenAI and Microsoft's primary defense of "fair use" in AI model training. If courts rule against fair use based on these internal admissions, it could rewrite the legal framework for the entire generative AI industry and force tech companies to license all training data. The motion highlights statements from Hecht noting that winning under a fair use defense would make the legal concept a mockery. The ongoing lawsuit, initiated in 2023 over millions of copyrighted articles, is not expected to receive a decision on proceeding to a formal trial until around 2027.
## BACKGROUND
Generative AI models require vast datasets scraped from the internet, including copyrighted news articles, books, and images, to learn patterns and generate new content. Tech companies historically argue that training AI on publicly available data constitutes "fair use" under copyright law, while publishers argue it infringes on intellectual property and undercuts their market value.