The guidelines address the GDPR implications of web scraping personal data from the internet to train generative AI models, covering both organisations that scrape data themselves and those that reuse already-scraped datasets. They cover the qualification of scraping organisations as controllers, joint controllers, or processors, and set out how the principles of purpose limitation, transparency, data minimisation, and accuracy apply to scraping activities. They also examine legitimate interest as the primary legal basis relied on for such scraping, detailing the three-part balancing test, and address the treatment of special categories of personal data incidentally collected during scraping.
Author: European Data Protection Board
Status: Adopted / Published
Adoption date: 2026-07-07
Last updated: 19 Aug 2026
Category: Guidance
Subcategory: Official guidance