Reddit advances lawsuit accusing Perplexity AI of conspiring with web scraper
Reddit is moving forward with a lawsuit that accuses Perplexity AI of conspiring with a web scraper to bypass its scraping blocks. The lawsuit alleges that this partnership was designed to access Reddit's content despite restrictions. This case could set a significant legal precedent regarding how AI companies access training data and whether bypassing web scraping blocks constitutes a legal violation. It highlights the growing tension between content platforms protecting their data and AI companies seeking information. Reddit claims the web scraper acted in coordination with Perplexity AI to circumvent its blocks, even after Google lost a related legal battle. Perplexity AI has previously faced scrutiny from major media organizations for allegedly using undisclosed crawlers to bypass robots.txt restrictions.
## BACKGROUND
Web scraping is a technique used to extract data from websites, which AI companies often rely on to train large language models. Websites typically use the robots.txt protocol to signal to web crawlers which parts of their site should not be accessed. However, compliance with robots.txt is voluntary, leading to disputes when AI companies or scrapers ignore these directives.