Researcher Accuses OpenAI of Training on User Prompts to Claim Mathematical Breakthroughs
A mathematician claimed on Mastodon that OpenAI trained its models on ChatGPT conversations containing novel mathematical concepts and subsequently announced breakthroughs without giving proper credit. The claim highlights concerns that user-submitted ideas in private prompts are being absorbed directly into model training corpora. This situation underscores growing tensions regarding research attribution, intellectual property, and data contamination in generative AI. As researchers use AI assistants to brainstorm unpublished work, companies risk claiming unearned scientific breakthroughs if training pipelines absorb and reproduce user-provided ideas. The issue stems from posts by mathematician Andreas Thom on Mathstodon detailing instances where proprietary or novel concepts appeared in OpenAI outputs after being shared in ChatGPT sessions. Although ChatGPT enables data sharing for model training by default, critics argue AI developers still bear an academic and ethical responsibility to audit training data for leakage before publishing research claims.
## BACKGROUND
In machine learning, data leakage or data contamination occurs when information that should remain private or separate from training sets unintentionally influences a model's capabilities. Generative AI services often retain user prompt histories by default to train future model iterations unless users explicitly change their privacy settings.