Regulations.ai
AI RegulationsAI GovernanceResearch
AI Chat
Sign In
HomeGlossaryweb scraping
Technical

web scraping

Automated mass collection of web information via bots/crawlers.

Definition

The activity of automated collection—typically via bots and crawlers—of large quantities of information available on the web, with emphasis on repetitive, mass harvesting intended to compile datasets for subsequent analysis or training of AI models.

Provvedimento n.329 (20 May 2024) - Nota informativa su web scraping per finalità di addestramento di intelligenza artificiale generativa (Garante per la protezione dei dati personali)

Related Terms

hybrid scraping

Combination of direct harvesting and third‑party datasets....

indirect scraping

Use of datasets created or redistributed by third parties....

direct scraping

Scraping where the harvester is also the model developer....

Scraping approaches (direct, indirect, hybrid)

Classification of how scraped data are collected and produced....

Machine Learning (ML)

A subset of AI enabling systems to learn from data, identify patterns, and make decisions with minimal human intervention....

Back to Glossary

Regulations.AI

The source for AI regulation. Laws. Governance. Research. Worldwide.

Quick Links

  • Regulation Tracker
  • Research
  • AI Governance
  • Feedback

Resources

  • Glossary
  • Regulation Types
  • Status Guide
  • Topics
  • External Resources

© 2026 Smitteck GmbH. All rights reserved.

AboutPrivacyTerms
Download all regulations database