> ## Content Index
> Fetch the complete content index at: https://www.thedelatorrereview.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# California’s AI Training Data Transparency Law: What Generative AI Developers Must Disclose
- URL: https://www.thedelatorrereview.com/californias-ai-training-data-transparency-law-what-generative-ai-developers-must-disclose/
- Published: 2026-01-01T02:07:00.000Z
- Updated: 2026-08-28T00:56:25.000Z
- Description: California’s AI Training Data Transparency Law requires generative AI developers to disclose key information about training datasets, including their sources, contents, use of personal information, licensing, modifications, collection periods, and synthetic data.
- Author: Lydia
- Tags: Transparency, California, AI Governance, AI Training, Artificial Intelligence (AI), AI Developer, Synthetic Data, Personal Information, Data Provenance, California Consumer Privacy Act (CCPA)

> **Key Takeaways: (1) California requires training-data transparency for covered generative AI.** Developers must publish high-level information about datasets used to develop covered systems and services. (2) **The requirement reaches systems released on or after January 1, 2022.** Covered systems made publicly available to Californians fall within the statute whether access is paid or free. **(3) January 1, 2026 is not merely a one-time deadline.** Disclosures are also required before subsequent covered releases and substantial modifications. **(4) The required disclosures are extensive.** Developers must address dataset sources, purpose, scale, data types, IP-protected material, licensing, personal information, processing, collection periods, first use, and synthetic data. **(5) Training includes more than initial training.** The statutory definition expressly includes testing, validation, and fine tuning by the developer. **(6) Data provenance is critical.** Developers need sufficient records about where training data came from, what happened to it, and when and why it was used. **(7) The exceptions are narrow.** The statute provides specific exclusions for certain security, aviation, national-security, military, and defense systems.

---

California has adopted a new transparency regime requiring developers of certain generative artificial intelligence systems to publicly disclose information about the data used to develop those systems.

The [**Artificial Intelligence Training Data Transparency Act**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplayText.xhtml?lawCode=CIV&division=3.&title=15.2.&part=4.&chapter=&article=&ref=thedelatorrereview.com), enacted through [**AB 2013**](https://legiscan.com/CA/text/AB2013/id/2910882?ref=thedelatorrereview.com) and codified at California Civil Code §§ 3110–3111, focuses on a question at the center of contemporary AI governance: What data was used to train the generative AI system?

Rather than prohibiting developers from using particular categories of training data, the law primarily takes a transparency-based approach. 

---

> **Covered developers must publish high-level information about their training datasets, including where the data came from, what kinds of data were included, whether datasets contain personal information or intellectual-property-protected material, whether data was purchased or licensed, and whether synthetic data was used.**

---

The law therefore creates an important operational consequence for AI governance: developers cannot easily provide meaningful training-data transparency unless they have established processes for understanding and documenting **data provenance** throughout the AI development lifecycle.

## What Is California’s Artificial Intelligence Training Data Transparency Act?

California's Artificial Intelligence Training Data Transparency provisions were added to the Civil Code by AB 2013 (2024). Section 3110 establishes the statute's principal definitions, while Section 3111 establishes the substantive disclosure obligation.

At a high level, the law requires a covered developer to publish documentation on its website concerning the datasets used to train a covered generative AI system or service. Importantly, the statute does **not require publication of the training datasets themselves**. Instead, it requires a **high-level summary** containing specified information about those datasets.

---

> **California created a transparency obligation, rather than requiring developers to make their proprietary training corpora publicly accessible.**

---

## When Does the Law Apply?

Section 3111 establishes several elements that determine when the disclosure requirement applies.

The obligation concerns a:

- **generative artificial intelligence system or service**;
- or a **substantial modification** to such a system or service;
- released on or after January 1, 2022 that is **made publicly available to Californians for use**.

[**Cal Civil Code §** ](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com)[**3110(a)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) first defines **“artificial intelligence”** (AI) generally as:

> an engineered or machine-based system that varies in its level of autonomy and that can, for explicit or implicit objectives, infer from the input it receives how to generate outputs that can influence physical or virtual environments.

The definition of artificial intelligence in the Artificial Intelligence Training Data Transparency Act is the same substantive definition provided by AB 2885 . It is a broad, technology-neutral definition. For a detailed discussion of this definition—including autonomy, inference, objectives, inputs, and outputs—see [**California Defines AI: AB 2885 and Its Impact*.*](https://www.thedelatorrereview.com/california-defines-ai-ab-2885-and-its-impact/)

[**Cal Civil Code §** ](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com)[**3110(c)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) then defines the narrower category of **“generative artificial intelligence.”** Generative AI is artificial intelligence capable of generating:

> derived synthetic content, such as text, images, video, and audio, that emulates the structure and characteristics of the artificial intelligence’s training data.

Examples expressly identified by the statute include **text, images, video, and audio**. Because the statute uses the phrase “such as,” these examples are illustrative rather than exhaustive.

This distinction between AI generally and generative AI is important because the law does not impose training-data transparency requirements on every AI system. 

---

> **The disclosure obligations apply only to the release of covered generative artificial intelligence systems or services and "substantial modifications"—not to AI systems generally.**

---

[**Cal Civil Code §** ](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com)[**3110(d)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) defines a **“substantial modification”** as a new version, release, or update that **materially changes the functionality or performance** of a generative AI system or service, including through retraining or fine tuning.

Not every update qualifies. The key question is whether it materially changes functionality or performance.

Under [**Cal Civil Code §** ](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com)[**3110(f)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com), training includes **testing, validating, and fine tuning by the developer**, not just initial model training. This broader definition matters because datasets used during these later development stages may also need to be considered when preparing the required training-data disclosures. 

---

> **“Training” Is Broader Than Initial Model Training: Datasets Used for Testing, Validation, and Fine Tuning Should Be Considered as Well**

---

The statute expressly provides that it applies **regardless of whether the terms of use include compensation**. A covered system therefore does not escape the transparency requirement merely because consumers can use it for free.

---

> **A covered system does not escape the transparency requirement merely because consumers can use it for free.**

---

## To Whom Does the Law Apply?

The disclosure obligations apply to **developers** of covered generative AI systems or services.

[**Cal Civil Code § 3110(b)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) defines a **“developer”** as a person, partnership, state or local government agency, or corporation that designs, codes, produces, or substantially modifies an artificial intelligence system or service for use by members of the public.

---

> **The definition of developer is not limited to the original creator of an AI system. An entity that substantially modifies an existing system may also qualify as a developer.**

---

For purposes of the definition, “members of the public” excludes certain affiliates identified by the statute and a hospital’s medical staff members.

> **Practice Tip:** Organizations that build on third-party AI models should not assume that the original model provider is the only relevant developer. Consider whether the organization’s own modifications materially change the system’s functionality or performance and therefore constitute a **substantial modification** under the Act.

# What Must Developers Disclose?

The heart of the law is [**Cal Civil Code § 3111(a)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=3111.&ref=thedelatorrereview.com)**.**

Covered developers must post documentation on their websites concerning the data used to train the generative AI system or service. The documentation must include a high-level summary of the datasets used in development and specified information about those datasets.

> **Practice Tip:** Training-data transparency should begin with **data governance**, not with drafting the eventual public disclosure. Developers should establish data inventories and provenance records as datasets enter and move through the AI development lifecycle. Otherwise, determining what must be disclosed at release may become a costly—and potentially incomplete—reconstruction exercise.

The statutory requirements can be grouped into several practical categories.

### 1\. Where Did the Training Data Come From?

Developers must identify the **sources or owners of the datasets**. 

They must also disclose whether the datasets were:

- **purchased**; or
- **licensed** by the developer.

A developer therefore needs sufficient visibility into its data supply chain to describe where its datasets originated and the circumstances under which certain datasets were obtained.

### 2\. Why Were the Datasets Used?

The disclosure must describe how the datasets further the intended purpose of the AI system or service. In other words, developers must be able to explain, at a high level, why the data is relevant to what the AI system is intended to do.

### 3\. How Much Data Is Included?

Developers must disclose the **number of data points included in the datasets**. 

The statute recognizes that precise counts may not always be practical. The number may therefore be provided in general ranges, and developers may use estimated figures for dynamic datasets.

### 4\. What Types of Data Are Included?

Developers must provide a description of the **types of data points** within the datasets.

The statute distinguishes between labeled and unlabeled datasets:

- for datasets containing labels, the disclosure concerns the types of labels used; and
- for datasets without labeling, it concerns the data's general characteristics.

### 5\. Does the Dataset Contain IP-Protected Material?

Developers must disclose whether datasets contain data protected by:

- **copyright**;
- **trademark**; or
- **patent**;

or whether the datasets are **entirely in the public domain**.

A disclosure that a dataset includes copyright-protected material does not by itself answer whether the developer's particular use of that material is authorized, licensed, subject to an exception or defense, or otherwise lawful.

> **Practice Tip:** Organizations should distinguish between **training-data transparency** and the separate legal analysis governing whether particular data may lawfully be used. Section 3111 requires disclosure about the presence of IP-protected material and whether datasets were purchased or licensed; it does not itself resolve the underlying intellectual-property rights associated with that data.

### 6\. Does the Dataset Include Personal Information or Aggregate Consumer Information?

Developers must disclose whether their datasets include **personal information**, using the definition in the California Consumer Privacy Act (CCPA), [**Cal Civil Code § 1798.140(v)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=1798.140.&ref=thedelatorrereview.com). This definition is **very broad**, extending well beyond information that directly identifies an individual and broadly aligning with the expansive concept of personal data under EU law. (See [*When Everything Is “Personal”: GDPR vs. CCPA*.](https://www.thedelatorrereview.com/when-everything-is-personal-gdpr-vs-ccpa/))

Developers must separately disclose whether the datasets contain **aggregate consumer information**, as defined in the CCPA, [**Cal Civil Code § 1798.140(b).**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=1798.140.&ref=thedelatorrereview.com)

Although the AI transparency provisions do not themselves prohibit the use of personal information or aggregate consumer information to train generative AI, their use may trigger separate obligations under the CCPA and other applicable U.S. privacy laws. 

---

> **Developers should consider AI training-data transparency and CCPA privacy compliance together, rather than treating them as separate compliance exercises.**

---

### 7\. Dataset Cleaning, Processing and Modification

The disclosure must state whether the developer performed any cleaning, processing, or other modification of the datasets.

If so, the developer must also describe the intended purpose of those efforts in relation to the AI system or service.

This requirement provides visibility not merely into where data came from, but also into how developers transformed or prepared it for use.

> **Practice tip:** AI governance documentation should ideally capture important transformations between **data acquisition and model development**, rather than treating the original dataset as the end of the data-governance inquiry.

### 8\. Data Collection Period and First Use

Developers must disclose the time period during which the data in the datasets was collected. If collection remains ongoing, the disclosure must provide notice of that fact.

Developers must also disclose the dates on which the datasets were first used during development of the AI system or service.

---

> Developers need to understand not only **what data they used and where it came from**, but also **when it was collected and when it entered the development process**.

---

### 9\. Synthetic Data Use

[**Cal Civil Code § 3110(e)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3110.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) defines synthetic data generation as a process in which **s**eed data are used to create artificial data having some of the statistical characteristics of the seed data.

Developers must disclose whether a generative AI system or service used synthetic data generation or continuously uses synthetic data generation in its development.

A developer may additionally describe the functional need or desired purpose of the synthetic data in relation to the intended purpose of the system or service.

---

> **Replacing or supplementing original data with artificially generated data does not necessarily remove the dataset from the transparency analysis.**

---

# Which Systems Are Exempt?

[Cal Civil Code Section 3111(b) ](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3111.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com)establishes three specific exclusions from the training-data documentation requirement.

1. **Security and Integrity**: The disclosure requirement does not apply to a generative AI system or service whose **s**ole purpose is to help ensure security and integrity. For this exception, the statute incorporates the meaning of “security and integrity” from the CCPA **(**[**Cal Civil Code § 1798.140(ac)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=1798.140.&ref=thedelatorrereview.com)), with specified adjustments broadening its application to developers and users. The use of “sole purpose” is important. A developer should not assume that a multipurpose generative AI system falls within the exception merely because one of its functions relates to security.
2. **Aircraft:** The law also excludes a generative AI system or service whose sole purpose is the operation of aircraft in the national airspace. Again, the statute expressly uses the narrower “sole purpose” formulation.
3. **National Security, Military and Defense Systems**: Finally, the disclosure obligation does not apply to a generative AI system or service developed for national security, military, or defense purposes; and made available only to a federal entity. Both aspects of this exception matter. The statutory language does not create a general exemption for any AI system that happens to have a defense-related use.

## **Ongoing Transparency Requirements**

[**Cal Civil Code § 3111**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?sectionNum=3111.&nodeTreePath=8.4.91&lawCode=CIV&ref=thedelatorrereview.com) requires covered developers to publish the required documentation on or before January 1, 2026, and before each time thereafter that a covered generative AI system or service—or a substantial modification—is made publicly available to Californians.

The law should therefore be understood as creating an ongoing release-related transparency obligation, rather than simply requiring developers to prepare a single disclosure by January 1, 2026.

> **Practice Tip:** Organizations should integrate the Section 3111 disclosure analysis into their **AI release and change-management processes**. A disclosure prepared for an initial release may not be sufficient when retraining, fine tuning, or another update materially changes the system's functionality or performance.

# Practical Compliance Considerations

Although the statute is framed as a public transparency requirement, complying with it has significant implications for internal AI governance.

### (1) Build a Training-Data Inventory

Organizations developing covered generative AI should be able to identify the datasets used throughout development, including datasets used for testing, validation, and fine tuning where those activities fall within the statutory definition of training.

### (2) Establish Data Provenance

Organizations should document dataset origins, ownership or sources, acquisition methods, collection periods, licensing status, and relevant transformations.

### (3) Classify Dataset Contents

Governance processes should identify whether datasets contain personal information, aggregate consumer information, IP-protected material, or synthetic data.

### (4) Connect Data to Purpose

Because the statute requires an explanation of how datasets further the intended purpose of the AI system, organizations should document why particular datasets are being used, rather than simply maintaining an inventory of what has been collected.

### (5) Integrate Transparency Into Release Management

The requirement applies not merely to initial releases but also to substantial modifications. AI product, engineering, legal, privacy, and governance teams therefore need a process for determining whether retraining, fine tuning, or another update materially changes functionality or performance.

### (6) Maintain Documentation Throughout Development

Many of the required disclosures depend on historical information—such as when data was collected, when datasets were first used, and how datasets were modified.

---

> **Waiting until launch to collect all the information needed will make compliance considerably more difficult.**

---

## Conclusion

California’s Artificial Intelligence Training Data Transparency law requires greater visibility into the data used to develop generative AI systems. Although it does not require developers to publish their datasets or prohibit particular categories of training data, meaningful transparency depends on knowing **where data came from, what it contains, and how and when it was used**.

For developers, **training-data inventories, data provenance, documentation, and release governance** are therefore essential components of compliance.

### Additional Resources

#### California Statutory Provisions

- [**California Civil Code §§ 3110–3111**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplayText.xhtml?division=3.&part=4.&lawCode=CIV&title=15.2.&ref=thedelatorrereview.com) **— Artificial Intelligence Training Data Transparency** — Definitions and training-data transparency requirements applicable to covered generative AI systems and services.
- [**Cal. Civil Code § 1798.140(b), (v)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=1798.140.&ref=thedelatorrereview.com) — CCPA definitions of **aggregate consumer information** and **personal information**, which are incorporated into the training-data disclosure requirements.
- [**Cal. Civil Code § 1798.140(ac)**](https://leginfo.legislature.ca.gov/faces/codes%5FdisplaySection.xhtml?lawCode=CIV§ionNum=1798.140.&ref=thedelatorrereview.com) — CCPA definition of **security and integrity**, incorporated by reference for purposes of the Act's security-and-integrity exception.

#### Legislation

- [**Assembly Bill 2013 (2023–2024), Chapter 817**](https://legiscan.com/CA/text/AB2013/id/3023192?ref=thedelatorrereview.com) — Enacted California's Artificial Intelligence Training Data Transparency provisions, effective January 1, 2025\.
- [**Assembly Bill 1170 (2025), Chapter 67** ](https://legiscan.com/CA/text/AB1170/id/3261463?ref=thedelatorrereview.com)— Amended Civil Code § 3111, effective January 1, 2026\.

#### Related The De La Torre Review AI Resources

- [***California Defines AI: AB 2885 and Its Impact***](https://www.thedelatorrereview.com/california-defines-ai-ab-2885-and-its-impact/) — The de la Torre Review's discussion of California's general AI definition, including the concepts of **AI systems, autonomy, objectives, inference, inputs, and outputs**, and the distinction between AI generally and generative AI.