Ethics

Sony Music and Warner Publishers Sue Anthropic Over Alleged Pirated Training Data

Sony Music Publishing, Warner Chappell and other publishers sued Anthropic and two co-founders, alleging the company obtained copyrighted lyrics and sheet music through torrenting and scraping to train Claude. Anthropic disputes the claims, extending a legal fight over whether lawful model training can be separated from unlawful acquisition.

By Michael C ·

Sony Music and Warner Publishers Sue Anthropic Over Alleged Pirated Training Data
SUPERBASH_ editorial image.

Sony Music Publishing, Warner Chappell and other music publishers have sued Anthropic and two of its co-founders, alleging that the AI company obtained copyrighted material through torrenting, scraping and unauthorized downloading to train Claude. The complaint expands the copyright conflict around frontier models from books and news into lyrics, compositions and sheet music. Anthropic disputes the allegations and says it will defend itself. At this stage, the new claims are accusations in a civil complaint, not judicial findings.

The case was filed in the U.S. District Court for the Northern District of California and names Dario Amodei and Benjamin Mann alongside the company. According to TechCrunch, the publishers describe a broad campaign to collect protected works, including books that contain lyrics and musical notation. Their argument focuses not only on what a model learned from copyrighted material, but on how the copies used for training were allegedly acquired.

That distinction has become central to AI copyright litigation. A court can find that a particular analytical or training use qualifies as fair use while separately finding that downloading pirated copies was unlawful. Model developers have often emphasized the transformative nature of training, where a system extracts statistical relationships rather than storing a consumer library. Rights holders answer that the transformation argument cannot erase the source and legal status of the copies assembled before training begins.

Training Use and Dataset Acquisition Are Different Questions

Fair use is a fact-specific doctrine that considers purpose, the nature of the work, the amount used and the effect on the market. It is not a general license for acquiring material from any source. A developer may have a defensible argument about using lawfully obtained works for analysis and still face liability if it copied or distributed files through an unlawful channel. The music publishers are trying to keep those stages separate so that a favorable view of model training does not immunize the collection process.

The complaint sharpens the legal distinction between how copyrighted works are used and how training copies are obtained. Image: SUPERBASH_
The complaint sharpens the legal distinction between how copyrighted works are used and how training copies are obtained. Image: SUPERBASH_

Related litigation has already exposed that split. In Bartz v. Anthropic, a court distinguished between permissible training uses and the illegality of obtaining works through piracy, and Anthropic was ordered to pay $1.5 billion. The new publishers' suit seeks to build on that reasoning. It also follows a separate music case involving Concord Music Group and Universal Music Group. Those precedents do not automatically decide the new complaint because the works, evidence, defendants and alleged conduct must still be established.

Music publishing adds complexity because one song can involve several rights. The sound recording, underlying composition, lyrics and arrangement may have different owners and licensing channels. A dataset assembled from books, lyric sites or files can therefore create overlapping claims. Sony Music Publishing and Warner Chappell represent catalogs whose value depends on tracking uses and collecting royalties across many formats. A model-training pipeline built without equivalent provenance records may struggle to show where each text fragment came from and what permission applied.

Warner Chappell and the other plaintiffs also have a strategic interest in establishing that AI developers cannot treat licensing as optional while building commercial systems. Model companies may respond that licensing every item in a web-scale corpus is impractical and that copyright law has long permitted computational analysis. The policy challenge is that both statements can be true: comprehensive licensing can be difficult, and creators can still suffer when valuable catalogs are copied without consent or clear compensation.

Provenance Is Becoming a Product Requirement

The operational lesson for model developers is straightforward even before the case is resolved. Training data needs a chain of custody. Teams should know the source, collection method, applicable license, restrictions and removal history for material entering a corpus. That record can support legal defenses, honor opt-outs and identify contaminated sources before they reach a training run. It is slower than treating the internet as an undifferentiated pool, but litigation has made the absence of records a financial risk.

Anthropic has said it disagrees with the publishers' claims and intends to defend itself robustly. The company may challenge the factual allegations, legal theories, damages or effort to name individual co-founders. A fair account must leave room for that defense. The complaint's language is forceful, but a court will need evidence showing what was downloaded, by whom, when it entered training systems and how the asserted copyrights connect to those copies.

Music publishing rights span lyrics, compositions and licensing records that do not map neatly onto model-training datasets. Image: SUPERBASH_
Music publishing rights span lyrics, compositions and licensing records that do not map neatly onto model-training datasets. Image: SUPERBASH_

The U.S. Copyright Office has been examining artificial intelligence across digital replicas, copyrightability and training. Its work provides a policy venue beyond individual lawsuits, but legislation remains difficult because courts are deciding cases while technology and business models continue to change. A statutory licensing regime could reduce transaction costs, yet it would need rules for valuation, small rights holders, international catalogs and open research. A broad exemption would simplify development while shifting more of the cost to creators.

The commercial response may arrive before a comprehensive law. AI companies can negotiate catalog licenses, use rights-cleared datasets, build tools that let owners exclude works and publish more detailed documentation. Music companies can create machine-readable licensing products and pricing that reflect training, retrieval and generated output separately. Those steps will not resolve every fair-use dispute, but they can reduce the number of cases that begin with uncertainty about whether the source files were lawful at all.

The Evidence Will Matter More Than the Rhetoric

The court will eventually have to move from competing descriptions to records. Publishers will seek logs, datasets, internal messages and training documentation. Anthropic will test ownership, causation and the scope of damages while explaining its collection and model-development practices. The case could settle before those questions receive a final ruling, as much AI copyright litigation has, but discovery alone may influence how every major lab documents future corpora.

Damages could become a separate contest even if some liability is established. Publishers may argue that unauthorized acquisition avoided licensing costs and supported a valuable commercial model. Anthropic may challenge whether each asserted work entered training, whether copies affected a market and how damages should be calculated across a large corpus. Courts will need to avoid both extremes: treating every file as economically identical or assuming scale makes individual rights impossible to value.

Naming individual co-founders increases the pressure but does not establish personal responsibility. The publishers will need to connect each defendant to the conduct and legal duty they allege. Corporate leaders routinely set data strategy without handling files directly. The legal question is whether their decisions, knowledge or supervision satisfy the requirements for the claims pleaded. That issue may affect how executives document approval of future dataset acquisitions.

Output behavior may appear in the evidence even though acquisition is central. If Claude can reproduce protected lyrics or notation, publishers may use that behavior to argue the training data included their works or that safeguards are insufficient. Anthropic can respond with evidence about memorization rates, filtering and the difference between a model generating familiar language and storing a source. Technical testing will have to be carefully designed because prompts can influence what a model appears to recall.

International operations make the compliance task harder. Copyright exceptions, database rights and licensing practices vary across jurisdictions, while a training corpus may be collected and processed across several locations. A global model developer needs provenance detailed enough to apply regional restrictions and respond to removal requests without pretending one legal conclusion governs every market. Rights holders similarly need standardized identifiers that let a developer match claims to the material actually present.

Independent audits could become a practical middle layer between private claims and public litigation. An auditor with controlled access could verify dataset sources, removal procedures and memorization tests without requiring a developer to expose the full corpus to competitors. Rights holders would need confidence in the auditor's methods and a way to challenge omissions. A credible process would not replace courts, but it could identify disputes early enough for licensing or removal to remain possible.

Creators also need remedies that work below the scale of a major publisher. A songwriter or small press cannot finance years of discovery to determine whether a work was copied. Registries, notice systems and collective licensing may give smaller owners a practical route to participate. Any solution should avoid making registration a new condition of copyright while still giving developers identifiers they can match against very large datasets.

The unresolved issue is whether the industry can separate a serious debate about transformative training from the simpler obligation to account for where copies came from. If developers cannot establish provenance, courts may continue treating dataset acquisition as an independent source of liability regardless of how sophisticated the resulting model becomes. The U.S. Copyright Office can help clarify policy, but the immediate legal opening belongs to publishers. For AI labs, it is a warning that data engineering now includes evidence preservation and rights management, not only scale.

Topics: Anthropic, music publishing, copyright, training data, Claude