Ethics

The Consent Crisis: Why AI Training on Human Data Demands a New Legal Framework

The largest AI training datasets were built by scraping billions of images, articles, and videos without explicit consent. As lawsuits mount and regulators demand accountability, a consent-based legal framework is becoming inevitable.

By Patrick T ·

The Consent Crisis: Why AI Training on Human Data Demands a New Legal Framework

The largest AI training datasets—including those powering GPT-5, Claude Opus 4.7, and Gemini 2.0—were built by scraping billions of images, articles, and videos from the internet without explicit consent from creators. As lawsuits mount and regulators demand accountability, the question is no longer whether AI companies violated consent norms, but whether existing legal frameworks can address the scale of the violation.

The Scale of the Problem

In 2024, researchers at UC Berkeley estimated that approximately 1.3 billion copyrighted works were used to train large language models without permission. This includes 400 million books, 200 million academic papers, and over 500 million news articles. The data came from CommonCrawl, a freely available internet archive that indexes the public web—but 'public' does not mean 'free to use for commercial AI training.'

The distinction matters. A photograph published on Instagram is publicly visible, but Instagram's terms of service prohibit automated scraping. A news article on The New York Times' website is publicly accessible, but the Times explicitly forbids machine learning companies from using its content without a license. Yet major AI labs trained their models on these datasets anyway, relying on a legal gray zone that has now become a courtroom battleground.

By May 2026, over 150 lawsuits have been filed against major AI companies by authors, photographers, journalists, and publishers. The cases fall into two categories: copyright infringement (claiming the models memorize and reproduce training data) and right of publicity (claiming the use of personal data violates privacy rights). Neither category has clear precedent in AI law.

Legal frameworks for AI data use are still being developed. The EU's AI Act represents the first major regulatory attempt to require consent for training data.
Legal frameworks for AI data use are still being developed. The EU's AI Act represents the first major regulatory attempt to require consent for training data.

Why Consent Matters More Than Copyright

Copyright law was designed for a different era. When you photocopy a book, you make one copy. When you scan a book into a training dataset, you create a digital representation that can be indexed, searched, and potentially reconstructed by a neural network. The legal question—'Is this fair use?'—has no settled answer.

But the ethical question is simpler: Did the creator agree? Consent is the foundation of trust. When a photographer uploads an image to a portfolio site, they consent to people viewing it. They do not consent to their work being fed into a machine learning system that will generate new images in their style without attribution or compensation. When a journalist publishes an article, they consent to readers reading it. They do not consent to their unique voice and perspective being encoded into a model that will generate text that sounds like them.

The current legal framework treats all uses of published data the same. But AI training is qualitatively different from reading. It is extractive. It is scalable. It is designed to learn from and replicate patterns in human creativity.

The Regulatory Response

The European Union's AI Act, which went into effect in January 2026, requires that AI systems trained on copyrighted material must be disclosed in a registry. Companies must document what data was used, where it came from, and whether creators consented. Violations carry fines up to 6% of global revenue.

The United States has taken a slower approach. The Copyright Office issued guidance in March 2026 stating that AI-generated works cannot be copyrighted unless a human author made 'significant creative choices' in the generation process. But this does not address the legality of training on copyrighted data in the first place.

Japan and South Korea have taken opposite positions. Japan's Ministry of Economy, Trade and Industry issued guidance explicitly permitting AI training on copyrighted material without consent, framing it as necessary for national competitiveness. South Korea's approach is more restrictive, requiring opt-in consent for any commercial use of personal data in AI systems.

What a New Framework Might Look Like

A consent-based legal framework for AI training would require three components: Disclosure (AI companies must disclose exactly what data was used to train a model, at a granular level), Opt-Out Rights (Creators must have the right to request that their work be removed from training datasets), and Compensation (When commercial AI systems are trained on copyrighted material, creators should receive compensation).

These mechanisms already exist in other industries. Music licensing through ASCAP and BMI has worked for decades. Stock photo licensing is a multibillion-dollar market. The technology exists to apply similar models to AI training.

The Path Forward

The consent crisis will be resolved in one of two ways: through legislation or through litigation. Legislation is preferable because it creates certainty and allows for negotiated solutions. Litigation is slower but inevitable if legislation does not materialize. The next 18 months are critical. The EU's AI Act will be the first real test of consent-based regulation. If it works, other jurisdictions will follow. If it fails, then the default will be litigation, which will be more expensive and less predictable for everyone.