Training Data
Data used to develop and train AI models to perform their intended functions.
Definitions (9)
Material used to train machine learning models, including open web material and licensed content; encompasses the datasets and content inputs used to develop AI systems' behavior. The report uses this term to frame licensing, transparency and copyright considerations for AI development.
Datasets used to train, retrain, or fine-tune generative AI systems, including data used for reinforcement learning from human feedback (RLHF) and other methodologies; the law requires disclosure of sources, acquisition methods, processing steps, volumes, temporal scope, and whether synthetic data was used.
Described as the datasets used to train AI and generative models (e.g., Danish text corpora); the strategy requires clarifying licensing, rights, privacy and provenance of training data and signals open-release of selected corpora under governed terms to support domestic model development.
All datasets and records used to develop, pre‑train, fine‑tune or update an LLM, including public, proprietary, scraped and third‑party sources; the guidance treats training data as potentially containing personal data and emphasises provenance, retention, minimisation and lawful‑basis considerations.
The dataset or collection of data used to train, fit, or optimize an AI model's parameters during the development phase; it supports the learning process and is distinct from validation or test datasets by its primary role in model learning. Training data may be labelled or unlabelled depending on the learning paradigm.
Datasets and data sources used to train, validate, or fine-tune AI models, including annotations and preprocessing steps, which must meet data governance, quality, and personal data protection requirements to mitigate bias and privacy risks.
Collections of data used to train generative AI models, whose provenance, licensing, and content determine legal, ethical, and technical risks and therefore should be documented, minimized, and managed for compliance and quality.
Collections of data used to train or tune models, templates or algorithms for deep synthesis, including datasets containing personal or biometric information (e.g., faces, voices), which are subject to data-protection, consent and security requirements under the Provisions.
Data sets used to train or fine-tune machine learning models, encompassing sources, provenance, and characteristics of the records used to develop model behaviour; treated as a lifecycle stage that may engage privacy obligations (lawful collection, consent, quality, bias).
Related Terms
General-Purpose AI (GPAI)
AI models trained on broad data capable of performing diverse tasks across domains....
Bias Mitigation
Measures to identify and reduce discriminatory outcomes in AI systems....
Technical Documentation
Detailed records demonstrating AI system compliance with regulatory requirements....