United States - California - AI Training Data Transparency (AB 2013)
California AB 2013 — Generative Artificial Intelligence: Training Data Transparency
United States
RAI-US-CA-CA2GAXX-2024California Assembly Bill 2013 requires developers of generative AI systems to publicly disclose detailed information about the datasets used to train their models. Effective January 1, 2026, the law mandates transparency regarding data sources, copyright status, personal information inclusion, and data processing methods, making California the first U.S. state to require comprehensive AI training data disclosure.
Summary
Read full text ↗Plain English
Overview
California Assembly Bill 2013 (Chapter 817, Statutes of 2024) establishes the United States' first comprehensive requirements for transparency in generative artificial intelligence training data. Signed by Governor Gavin Newsom on September 28, 2024, and effective January 1, 2026, the law requires developers of generative AI systems to publicly document and disclose detailed information about the datasets used to train their models. The legislation responds to widespread concerns from creators, artists, journalists, and privacy advocates about the opaque nature of AI training processes and the potential unauthorized use of copyrighted works and personal information in developing large language models, image generators, and other generative AI systems. By mandating training data transparency, AB 2013 aims to empower copyright holders to understand whether their works have been used in AI training, enable consumers to make informed choices about AI products, and create accountability mechanisms for AI developers operating in California's substantial market. The law positions California as a national leader in AI governance, establishing precedents that may influence federal legislation and other state laws addressing artificial intelligence transparency.
Definitions
AB 2013 introduces several key definitions essential to understanding the law's scope and application. Generative artificial intelligence is defined as artificial intelligence that can generate derived synthetic content, including text, images, video, and audio, that emulates the structure and characteristics of the system's training data. This definition encompasses large language models (LLMs), image generation systems, video synthesis tools, and audio generation platforms. Developer means any person, partnership, state or local government agency, or corporation that creates, codes, produces, or substantially modifies a generative AI system or service. Substantial modification includes any update, revision, or other modification that materially changes the functionality or performance of the generative AI system, capturing both initial development and significant updates. Training data refers to datasets used to train, retrain, or fine-tune generative AI systems, including data used for reinforcement learning from human feedback (RLHF) and other training methodologies. Personal information adopts the California Consumer Privacy Act's (CCPA) expansive definition, encompassing any information that identifies, relates to, describes, is reasonably capable of being associated with, or could reasonably be linked with a particular consumer or household. Synthetic data means artificially generated data that mimics the statistical properties of real-world data without directly copying actual records.
Governance and Institutional Framework
AB 2013 establishes a disclosure-based regulatory approach rather than creating new oversight agencies or enforcement bodies. The law places compliance obligations directly on AI developers, requiring them to self-publish training data documentation on their websites without mandating government review or approval. The California Attorney General's Office retains general enforcement authority over the law through its existing consumer protection powers, including the ability to investigate potential violations and pursue enforcement actions under California's Unfair Competition Law (Business and Professions Code Section 17200). The Attorney General has announced plans to hire AI experts and investigative technologists to support enforcement of California's growing body of AI legislation, including AB 2013. The California Privacy Protection Agency (CPPA), which oversees CCPA and CPRA compliance, may play an indirect role where training data disclosures intersect with personal information processing, though AB 2013 does not explicitly delegate authority to the CPPA. The law's reliance on website disclosure means practical enforcement will depend on public scrutiny, investigative journalism, academic research, and complaints from affected parties such as copyright holders whose works may appear in undisclosed training datasets.
Key Focus Areas
- Training Data Source Transparency: Developers must disclose the sources of datasets used for training, including whether data was scraped from the internet, obtained from data brokers, licensed from content providers, or generated synthetically, enabling stakeholders to understand data provenance.
- Copyright and Intellectual Property Disclosure: The law requires developers to state whether training datasets include copyrighted, trademarked, or patented materials, addressing concerns from creators about unauthorized use of their works in AI training.
- Personal Information Identification: Developers must disclose whether training data contains personal information as defined under CCPA, supporting privacy rights and enabling individuals to understand potential data exposure.
- Data Volume and Characteristics: Required disclosures include the number of data points or records in training datasets (general ranges are acceptable) and descriptions of data types and characteristics, providing scale and composition transparency.
- Data Acquisition Methods: The law mandates disclosure of how training data was acquired, including purchases, licensing arrangements, partnerships, and web scraping, revealing commercial relationships and data sourcing practices.
- Data Processing and Modification: Developers must describe any modifications, cleaning, filtering, or transformation processes applied to training data, revealing how raw data was prepared for model training.
- Temporal Scope: Required timeframe disclosures cover when data was collected and when it was used for training, establishing temporal boundaries for training data inclusion.
- Synthetic Data Generation: The law specifically requires disclosure of whether synthetic data generation was used in training, acknowledging the growing use of AI-generated data to train subsequent AI systems.
- Consumer Choice Enablement: By making training data information publicly available, the law empowers consumers to make informed decisions about which AI products to use based on training data practices.
- Accountability and Audit Trail: Website disclosures create a public record that supports accountability, enabling researchers, regulators, and affected parties to evaluate AI training practices over time.
Implementation Framework
AB 2013 employs a straightforward implementation approach centered on public website disclosure rather than registration, certification, or regulatory approval processes. Developers must publish training data documentation on their publicly accessible websites, making the information available to any interested party without requiring account creation or access fees. The law does not prescribe a specific format, template, or technical standard for disclosures, granting developers flexibility in how they present required information while potentially creating inconsistencies that may complicate comparison across different AI systems. Disclosures must be maintained and updated as training data practices change, particularly when substantial modifications to AI systems involve new or different training datasets. The law applies prospectively to generative AI systems released after January 1, 2022, capturing the current generation of large language models and generative AI tools while excluding legacy systems developed before the generative AI boom. Compliance timelines require disclosures to be in place by January 1, 2026, providing developers approximately 15 months from the law's signing to audit their training data practices, document required information, and publish disclosures. For AI systems released between January 2022 and January 2026, developers must retrospectively document training data to the extent this information is available, though practical challenges may arise for systems trained on datasets whose provenance was not carefully documented at the time of development.
Monitoring and Evaluation
AB 2013 does not establish formal monitoring, reporting, or evaluation mechanisms, relying instead on public disclosure and market-based accountability. The law's effectiveness depends on the comprehensiveness and accuracy of developer disclosures, the capacity of stakeholders to review and analyze published information, and the willingness of enforcement authorities to act on potential violations. Third-party monitoring is anticipated from several sources: academic researchers studying AI training practices may systematically analyze published disclosures across the industry; investigative journalists covering AI misinformation may probe disclosures for inconsistencies or omissions; copyright holders and their representatives may compare disclosures against known uses of protected works; and privacy advocates may evaluate personal information disclosures against observed AI system behaviors. The California Attorney General's office has indicated it will monitor AI industry compliance as part of broader AI enforcement initiatives, though specific AB 2013 monitoring programs have not been announced. The law's lack of mandatory reporting to government agencies means no centralized compliance database will exist, potentially complicating systematic evaluation of industry-wide compliance rates. Future legislative amendments may introduce more robust monitoring requirements if initial implementation reveals significant compliance gaps or disclosure quality concerns.
Penalties, Liability, and Appeals
AB 2013 does not specify direct penalties for non-compliance, distinguishing it from regulations like the CCPA that include explicit fine structures. However, enforcement may proceed through several alternative legal mechanisms. The California Attorney General may pursue violations under the Unfair Competition Law (UCL, Business and Professions Code Section 17200 et seq.), which prohibits unlawful, unfair, or fraudulent business acts or practices. UCL violations can result in civil penalties up to $2,500 per violation, injunctive relief requiring compliance, and restitution to affected parties. Where inadequate or misleading training data disclosures constitute false advertising, the False Advertising Law (Business and Professions Code Section 17500) may apply, with similar penalty provisions. Private plaintiffs, including copyright holders, may bring UCL claims seeking injunctive relief and restitution, though UCL does not authorize private damages. Copyright holders who discover their works were used in training without authorization may pursue separate copyright infringement claims, with AB 2013 disclosures potentially serving as evidence of unauthorized use. Privacy violations stemming from training data containing personal information may trigger CCPA enforcement, including penalties up to $7,500 per intentional violation. Appeals from enforcement actions follow standard California administrative and judicial review procedures, with final agency decisions reviewable in Superior Court and appellate courts. The absence of specified penalties may reduce deterrent effects, potentially encouraging legislative amendments if compliance issues emerge.
Relationship to Other Instruments
AB 2013 operates within California's expanding AI regulatory ecosystem and interacts with multiple existing legal frameworks. California SB 942 (California AI Transparency Act) complements AB 2013 by addressing output transparency—requiring detection tools and content provenance labels for AI-generated content—while AB 2013 focuses on input transparency through training data disclosure. Together, these laws create a comprehensive transparency framework covering both sides of the generative AI pipeline. The California Consumer Privacy Act (CCPA) and California Privacy Rights Act (CPRA) provide the definitional framework for personal information disclosures under AB 2013 and may impose independent obligations when training data includes consumer personal information. Federal copyright law, particularly ongoing litigation over AI training practices, establishes the background legal framework for intellectual property disclosures required under AB 2013. The European Union AI Act includes training data transparency requirements for general-purpose AI models that may inform compliance approaches for developers serving both California and EU markets. Proposed federal AI legislation, including training data disclosure provisions in various congressional bills, may eventually preempt or harmonize with AB 2013's requirements.
International Alignment
AB 2013 positions California as a leader in AI training data transparency, establishing requirements that exceed current federal U.S. standards and align with emerging international approaches. The European Union's AI Act includes training data documentation requirements for providers of general-purpose AI models, particularly those presenting systemic risks, creating regulatory convergence between California and EU frameworks that benefits multinational AI developers seeking consistent compliance approaches. The EU requirements include technical documentation covering training data sources, data preparation methods, and validation procedures, though with less emphasis on copyright disclosure than AB 2013. The United Kingdom's AI regulatory framework, based on existing sector regulators applying core AI principles, does not currently mandate comparable training data disclosure, though the UK Intellectual Property Office has published guidance on AI and copyright that may evolve into more specific requirements. China's AI regulations, including the Generative AI Measures effective August 2023, require AI providers to use data from legitimate sources and maintain training data records, but do not mandate public disclosure comparable to AB 2013. Japan and South Korea are developing AI governance frameworks that may address training data transparency as their regulatory approaches mature. AB 2013's public disclosure approach differs from regulatory submission models used in some jurisdictions, reflecting California's emphasis on market transparency and consumer information over government oversight. The law's extraterritorial reach—applying to any AI system made available to California residents regardless of developer location—may influence global AI development practices as companies build compliance capabilities for California's market.
Implementation Timeline
| Date | Milestone |
|---|---|
| February 2024 | AB 2013 introduced in California Assembly by Assemblymember Jacqui Irwin |
| May 2024 | Bill passes Assembly with amendments addressing scope and exemptions |
| August 2024 | Bill passes California Senate |
| September 28, 2024 | Governor Gavin Newsom signs AB 2013 into law (Chapter 817) |
| January 1, 2026 | Law takes effect; all covered AI systems must have training data disclosures published |
| Ongoing from 2026 | Developers must update disclosures when substantial modifications to AI systems involve new training data |
Compliance Checklist
| Requirement | Details |
|---|---|
| Determine Applicability | Assess whether you develop or substantially modify generative AI systems made available to California residents after January 1, 2022 |
| Audit Training Data | Document all datasets used to train, retrain, or fine-tune covered AI systems, including sources, acquisition methods, and processing history |
| Identify Copyright Materials | Determine whether training datasets include copyrighted, trademarked, or patented materials and document accordingly |
| Assess Personal Information | Evaluate training data for personal information as defined under CCPA and document findings |
| Document Data Volume | Calculate or estimate the number of data points in training datasets; general ranges are acceptable |
| Record Acquisition Methods | Document how each dataset was obtained (purchased, licensed, scraped, etc.) |
| Document Processing Steps | Record all modifications, cleaning, filtering, or transformation applied to training data |
| Establish Temporal Records | Document collection timeframes and training dates for all datasets |
| Identify Synthetic Data | Document any use of synthetic data generation in training processes |
| Prepare Website Disclosure | Create comprehensive, publicly accessible documentation addressing all required elements |
| Publish by Deadline | Ensure disclosures are live on company website by January 1, 2026 |
| Establish Update Procedures | Create processes to update disclosures when substantial modifications involve new training data |
Sources and References
| Source | Type |
|---|---|
| AB 2013 Bill Text - California Legislature | Primary Source |
| AB 2013 Bill History - California Legislature | Primary Source |
| AB 2013 Bill Analysis - California Legislature | Legislative Analysis |
| Governor Newsom Signing Announcement | Primary Source |
| California Attorney General Legal Advisory on AI | Government Guidance |
| California Privacy Protection Agency | Government Agency |
California's AB 2013 is a landmark law requiring developers of generative artificial intelligence (AI) systems to publicly disclose detailed information about the data used to train their models. This applies to anyone who creates, codes, produces, or significantly modifies generative AI systems (like large language models or image generators) that are made available to California residents, for systems released after January 1, 2022.
Developers must publish comprehensive documentation on their public websites by the effective date. Key disclosures include: - The specific sources of training data (e.g., scraped from the internet, licensed, purchased, or synthetically generated). - A clear statement on whether the training datasets contain copyrighted, trademarked, or patented materials. - Disclosure of whether the data includes "personal information" as defined by California's Consumer Privacy Act (CCPA). - Details on data volume (general ranges are acceptable) and how the data was acquired, processed, cleaned, or modified.
The law was signed in September 2024 and officially takes effect on January 1, 2026.
AB 2013 itself doesn't list specific fines. However, the California Attorney General can enforce it using existing consumer protection laws like the Unfair Competition Law, which carries civil penalties up to $2,500 per violation. Inaccurate disclosures could also trigger False Advertising Law. Importantly, these disclosures can serve as evidence for copyright holders pursuing infringement claims or for privacy advocates using CCPA enforcement if personal data is mishandled.
A practical challenge for many developers is the retrospective nature of the law. While it takes effect in 2026, it applies to AI systems released as far back as January 2022. This means companies must audit and document training data for models developed years ago, potentially facing difficulties if detailed provenance records weren't meticulously kept at the time.
Plain-English rewrite by Regulations.ai — not legal advice. Verify against the official text.
Read this article-by-article
Plain-English breakdown of 5 key articles, with cross-jurisdiction equivalents where applicable.
What you must do — compliance checklist
0 / 11 marked completePlain-English obligations under United States - California - AI Training Data Transparency (AB 2013). Not legal advice — verify against the official text before relying on it.
- #1Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems made available to California residents.
“developers of generative AI systems to publicly document and disclose detailed information about the datasets used to train their models.”
- #2Critical⏰ Ongoing from 2026-01-01
Applies to: Developers of generative AI systems.
“Disclosures must be maintained and updated as training data practices change, particularly when substantial modifications to AI systems involve new or different training datasets.”
- #3Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“Developers must disclose the sources of datasets used for training, including whether data was scraped from the internet, obtained from data brokers, licensed from content providers, or generated synthetically”
- #4Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“The law requires developers to state whether training datasets include copyrighted, trademarked, or patented materials”
- #5Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“Developers must disclose whether training data contains personal information as defined under CCPA”
- #6Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“Required disclosures include the number of data points or records in training datasets (general ranges are acceptable) and descriptions of data types and characteristics”
- #7Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“The law mandates disclosure of how training data was acquired, including purchases, licensing arrangements, partnerships, and web scraping”
- #8Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“Developers must describe any modifications, cleaning, filtering, or transformation processes applied to training data”
- #9Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“Required timeframe disclosures cover when data was collected and when it was used for training”
- #10Critical⏰ Jan 1, 2026
Applies to: Developers of generative AI systems.
“The law specifically requires disclosure of whether synthetic data generation was used in training”
- #11Important⏰ Jan 1, 2026
Applies to: Developers of generative AI systems released after January 1, 2022.
“For AI systems released between January 2022 and January 2026, developers must retrospectively document training data to the extent this information is available”
Related Regulations
California AI Transparency Act
California, United States94% similar
California SB 942 — California AI Transparency Act
United States93% similar
California SB 53 — Transparency in Frontier Artificial Intelligence Act (TFAIA)
United States91% similar
California AB 2655 - Defending Democracy from Deepfake Deception Act of 2024
United States91% similar
California SB 243 - Companion Chatbot Disclosure Requirements
United States90% similar
© Regulations.AI — created on 06-Jan-2026