Article-by-article breakdown
California AI Training Data Transparency (AB 2013)
Generative Artificial Intelligence: Training Data Transparency (AB 2013)
Article 1 | Scope and Definitions — Scope of Application and Key Definitions
Applies to
- ›Developers of generative AI systems
Plain English
California AB 2013 mandates transparency for generative AI systems, requiring developers to disclose details about their training data. This law applies to any person, partnership, or corporation that creates, codes, produces, or substantially modifies a generative AI system or service, particularly those made available to California residents.
The law defines 'generative artificial intelligence' as AI capable of producing synthetic content (like text, images, video, audio) that mimics its training data. 'Training data' includes all datasets used for training, retraining, or fine-tuning, even for reinforcement learning. 'Personal information' adopts the broad definition from the California Consumer Privacy Act (CCPA), covering data that identifies or relates to a consumer or household. The law applies prospectively to generative AI systems released after January 1, 2022, meaning current and future models are covered, with disclosures due by January 1, 2026.
Key points
- •Applies to developers of generative AI systems available in California.
- •Covers systems released after January 1, 2022.
- •Defines 'generative AI' as systems creating synthetic content.
- •Defines 'developer' broadly to include creators and those making substantial modifications.
- •Adopts CCPA's definition of 'personal information'.
What you need to do
- 1.Determine if your AI system meets the 'generative AI' definition.
- 2.Assess if your organization qualifies as a 'developer' under the law.
- 3.Identify all generative AI systems released after January 1, 2022, that are subject to this regulation.
- 4.Understand the broad scope of 'personal information' as it relates to your training data.
Article 2 | Disclosure Requirements — Mandatory Public Disclosure of Training Data Information
Applies to
- ›Developers of generative AI systems
Plain English
The core of AB 2013 requires developers to publicly disclose comprehensive information about the datasets used to train their generative AI models. This includes detailing the sources from which data was obtained, such as internet scraping, data brokers, or licensed content, to provide transparency on data provenance.
Developers must also explicitly state whether their training datasets contain copyrighted, trademarked, or patented materials, addressing intellectual property concerns. Furthermore, the law mandates disclosure of whether training data includes 'personal information' as defined by the CCPA, supporting individual privacy rights. Other required disclosures include the general volume and characteristics of the data, the methods used for data acquisition (e.g., purchases, licensing, web scraping), and any processing, cleaning, or transformation applied to the data. Developers must also specify the temporal scope of data collection and training, and whether synthetic data was used in the training process.
Key points
- •Disclose sources of all training datasets.
- •State whether copyrighted, trademarked, or patented materials are included.
- •Indicate if training data contains 'personal information' (CCPA definition).
- •Provide general data volume, types, and characteristics.
- •Detail data acquisition methods (e.g., scraping, licensing, purchasing).
- •Describe data processing, cleaning, and modification steps.
- •Disclose collection timeframes and training dates.
- •Specify any use of synthetic data generation.
What you need to do
- 1.Conduct a thorough audit of all training datasets used for covered AI systems.
- 2.Document the provenance, acquisition method, and processing history for each dataset.
- 3.Identify and categorize all intellectual property and personal information within your training data.
- 4.Prepare a comprehensive, publicly accessible document detailing all required information.
Cross-jurisdiction equivalents
Article 3 | Implementation and Maintenance — Guidelines for Publishing and Updating Disclosures
Applies to
- ›Developers of generative AI systems
Plain English
AB 2013 outlines a straightforward implementation approach focused on public accessibility. Developers are required to publish all mandated training data documentation on their publicly accessible websites. The law does not prescribe a specific format, template, or technical standard for these disclosures, offering flexibility in presentation but potentially leading to varied compliance approaches across the industry.
Crucially, these disclosures are not a one-time requirement. Developers must maintain and update the documentation as their training data practices evolve, especially when 'substantial modifications' to AI systems involve new or different training datasets. While the law takes effect on January 1, 2026, it applies retrospectively to generative AI systems released after January 1, 2022. This means developers must gather and document training data information for existing systems that fall within this timeframe, even if initial development did not anticipate such requirements.
Key points
- •Disclosures must be published on a publicly accessible website.
- •No specific format or template is mandated, allowing developer flexibility.
- •Disclosures must be updated when substantial modifications involve new training data.
- •Applies retrospectively to systems released after January 1, 2022.
- •Compliance deadline for initial disclosures is January 1, 2026.
What you need to do
- 1.Designate a specific section on your company website for AI training data disclosures.
- 2.Develop internal processes for regularly reviewing and updating disclosure documents.
- 3.Ensure historical training data for systems released since January 2022 is adequately documented.
- 4.Train relevant teams (legal, product, engineering) on the ongoing disclosure requirements.
Article 4 | Enforcement and Penalties — Enforcement Mechanisms and Consequences of Non-Compliance
Applies to
- ›Developers of generative AI systems
Plain English
AB 2013 does not specify direct penalties for non-compliance, but enforcement will primarily fall under the existing powers of the California Attorney General's Office. Violations may be pursued under California's Unfair Competition Law (UCL), which prohibits unlawful, unfair, or fraudulent business practices. UCL violations can lead to civil penalties of up to $2,500 per violation, injunctive relief (requiring compliance), and restitution to affected parties.
If inadequate or misleading disclosures are deemed false advertising, the False Advertising Law (Business and Professions Code Section 17500) may also apply, carrying similar penalties. Private plaintiffs, such as copyright holders, may also bring UCL claims. Importantly, AB 2013 disclosures could serve as evidence in separate copyright infringement claims if unauthorized use of protected works is discovered. Furthermore, if training data disclosures intersect with personal information processing, violations could trigger enforcement under the California Consumer Privacy Act (CCPA), which includes penalties up to $7,500 per intentional violation. The Attorney General's office is enhancing its capacity with AI experts to support enforcement.
Key points
- •Enforced by the California Attorney General's Office.
- •Violations may be pursued under the Unfair Competition Law (UCL).
- •UCL penalties include civil fines up to $2,500 per violation, injunctive relief, and restitution.
- •False Advertising Law may apply to misleading disclosures.
- •Disclosures can serve as evidence in copyright infringement lawsuits.
- •Non-compliance related to personal information may trigger CCPA enforcement and penalties.
What you need to do
- 1.Ensure disclosures are accurate and complete to avoid claims of unfair competition or false advertising.
- 2.Be prepared for potential scrutiny from the California Attorney General, privacy advocates, and copyright holders.
- 3.Implement robust data governance to mitigate risks of copyright or privacy violations that could be exposed by disclosures.
- 4.Understand that while no direct penalties are listed, existing laws provide significant enforcement mechanisms.
Article 5 | Interoperability and Global Context — Interaction with Other Regulations and International Alignment
Applies to
- ›Developers of generative AI systems
Plain English
AB 2013 operates within a broader regulatory landscape, both in California and internationally. It complements California's SB 942 (California AI Transparency Act), which focuses on output transparency by requiring detection tools and provenance labels for AI-generated content. Together, these laws create a comprehensive transparency framework for generative AI in California, covering both input (training data) and output.
The law leverages definitions from the California Consumer Privacy Act (CCPA) and California Privacy Rights Act (CPRA) for 'personal information,' indicating an overlap in privacy obligations. Federal copyright law forms the backdrop for intellectual property disclosures. Internationally, AB 2013 aligns with aspects of the European Union AI Act, which also includes training data documentation requirements for general-purpose AI models, particularly those with systemic risks. This convergence may simplify compliance for multinational AI developers. While other jurisdictions like the UK and China have AI regulations, AB 2013's emphasis on public disclosure sets a high standard, potentially influencing global AI governance and requiring developers serving California residents to adapt their practices regardless of their physical location.
Key points
- •Complements California SB 942 (output transparency) for a holistic AI transparency framework.
- •Leverages CCPA/CPRA definitions for personal information, creating privacy compliance overlap.
- •Interacts with federal copyright law regarding intellectual property disclosures.
- •Aligns with the EU AI Act's requirements for general-purpose AI model documentation.
- •Positions California as a leader in AI training data transparency, potentially influencing global standards.
What you need to do
- 1.Develop a unified compliance strategy that addresses both AB 2013 (input) and SB 942 (output) requirements.
- 2.Ensure your privacy compliance programs (e.g., CCPA/CPRA) are integrated with training data disclosure efforts.
- 3.For global operations, leverage compliance efforts for the EU AI Act to inform and streamline AB 2013 compliance.
- 4.Monitor federal AI legislative developments, as they may eventually preempt or harmonize with AB 2013.
Cross-jurisdiction equivalents
Need help applying this to your case?
The wizard takes 60 seconds and tells you which articles you actually need to worry about based on your jurisdictions, use case, and data.
Start the wizard →