Articles

Sony, Warner Sue Anthropic Over Song Lyrics in Claude

35 music publishers sued Anthropic over lyrics used to train Claude, seeking up to $150,000 per work. Here's the case and why it matters.

Chisato Chisato · · 6 min read
Rows of metal letterpress type arranged in a printer's tray

The music industry has opened a new front in the fight over AI training data. On August 28, 2026, a coalition of 35 music-publishing entities — led by affiliates of Sony Music Publishing and Warner Chappell Music — sued Anthropic in the U.S. District Court for the Northern District of California, alleging the company pirated song lyrics and sheet music to train its Claude models. The suit names Anthropic alongside CEO Dario Amodei and co-founder Benjamin Mann as defendants.

The publishers are seeking statutory damages of up to $150,000 for each work the court finds Anthropic infringed willfully. With the complaint sweeping in what the plaintiffs describe as “thousands, if not tens of thousands” of compositions, the theoretical exposure runs into the billions of dollars. The filing calls the conduct one of “the largest and most blatant” thefts of intellectual property, escalating a dispute that has moved steadily from books and news into music.

What the publishers allege

The core claim is about sourcing. The complaint alleges Anthropic obtained lyrics and sheet music from pirate repositories — specifically Library Genesis (LibGen) and the Pirate Library Mirror (PiLiMi) — and separately scraped licensed lyric sites including Musixmatch and LyricFind, without permission from or payment to the rightsholders.

The list of works reads like a jukebox of the last half-century: “Eye of the Tiger,” “Hallelujah,” “September,” “Livin’ on a Prayer,” and “Great Balls of Fire” appear among the named compositions, alongside works associated with Mariah Carey and Taylor Swift. The publishers allege that Claude can reproduce copyrighted lyrics verbatim when prompted — the crucial technical claim, because it converts an abstract “you trained on our catalog” grievance into a concrete demonstration that the protected text can be pulled back out of the model.

That verbatim-reproduction argument is the same one that has proven decisive in Europe. A Munich court used precisely this reasoning — that works “retained in a reproducible form” fall outside training exceptions — when it ruled against the AI music startup Suno, a decision we broke down in Suno’s GEMA copyright loss. The publishers suing Anthropic are, in effect, importing that framing into a U.S. courtroom.

Why this case leans on an earlier one

Anthropic does not arrive at this fight with a clean slate on training-data provenance. The publishers explicitly build on findings from Bartz v. Anthropic, the class action brought by book authors that ended in a landmark settlement. In that case, the record established that Anthropic had torrented more than seven million pirated books from LibGen and PiLiMi. Anthropic agreed to pay $1.5 billion to resolve it — the largest copyright settlement in U.S. history — and a federal judge granted final approval on July 20, 2026.

The music publishers’ argument is straightforward: the same pirated libraries that contained those books also contained lyrics and sheet music for hundreds or more of the plaintiffs’ compositions. In other words, they contend the infringement of their catalog was baked into the same acquisition the authors’ case already adjudicated. Rather than re-litigate how the data was obtained, they are pointing at a factual record that, in their telling, has already been established — and asking the court to apply it to a different category of copyrighted work.

That is a deliberately strong opening position. It narrows the dispute from “was this piracy?” toward “does the piracy the record already describes cover our works too, and what are the damages?”

The fair-use question, again

The defense Anthropic and other labs have leaned on is fair use — the argument that training a model on copyrighted material is a transformative act that produces something new rather than a substitute for the original. That doctrine has had a genuinely mixed run in the courts. Some rulings have found training itself can be transformative; but courts have drawn a sharp line at how the data was acquired, treating downloading from pirate sites as a separate wrong from the training that follows.

The verbatim-output allegation is designed to attack the transformative claim head-on. If a plaintiff can show a model reproducing protected lyrics word-for-word, the “we only learned abstract patterns” narrative gets much harder to sustain — the same pivot that has reframed these cases from disputes about analysis into disputes about memorization, where specific training examples end up encoded in the model’s weights. It is the pattern showing up across every modality, and it is why the mechanics of how generative models store and reproduce their inputs — explored in our explainer on how diffusion models work — have become central to the legal fight, not just the technical one.

Content owners, meanwhile, are increasingly choosing to license rather than litigate where they can — the path taken in the Getty Images licensing deal with OpenAI. The publishers suing Anthropic are pursuing the other route: establish liability first, negotiate from strength second.

The timing problem for Anthropic

The lawsuit lands at an awkward moment. Anthropic is in the middle of a historic capital-markets run, having filed confidentially for an IPO and reached a valuation among the highest of any private company in the world. Large, unquantified copyright liabilities are exactly the kind of contingency that public-market investors — and the SEC’s disclosure process — scrutinize closely.

Anthropic has not detailed a substantive response to the specific allegations. The company has previously defended its training practices as consistent with fair use while settling the acquisition-focused claims in Bartz. How it threads that needle here — defending training while having already paid to resolve the piracy that underpins the sourcing claim — will shape the case.

What it means

This is less a novel legal theory than a well-aimed sequel, and its force comes from leverage the plaintiffs did not have to build from scratch.

The Bartz record is the weapon. By anchoring their complaint to facts already established in the authors’ case, the publishers skip the hardest and most expensive part of copyright litigation — proving how the data was obtained — and go straight to scope and damages. That makes this a faster, more dangerous suit for Anthropic than a cold-start filing would be.

Memorization is the industry-wide exposure. The verbatim-reproduction claim is the thread connecting Munich, the authors’ case, and this one. Any model that can be prompted to regurgitate protected text — lyrics, prose, or code — is now vulnerable to the same argument. That is a far narrower and more litigable claim than “training is infringement,” and it does not depend on winning the abstract fair-use debate.

Winners and losers. Music publishers and collecting societies gain a powerful bargaining position and a template other rightsholders will copy. Licensed-data vendors gain a market. AI labs that trained on unlicensed corpora face a mounting, hard-to-quantify liability — and the prospect that “scrape now, settle later” becomes a recurring, budgeted cost rather than a one-time write-off.

What to watch next. First, whether Anthropic tries to move the fight back to acquisition or is forced to defend training and output directly. Second, the damages theory — statutory damages at $150,000 across tens of thousands of works is the kind of number that reshapes settlement math, and a settlement here would set a price for music the way Bartz set one for books. Third, the IPO overhang: how Anthropic discloses and reserves against this liability as it courts public investors. The book case gave the industry its first hard number for training on pirated text. The music case is now testing whether that number was a ceiling or a floor.

Chisato Chisato · · 5 min read

Massachusetts AI Safety Bill: Anthropic vs OpenAI

Anthropic backs strict Massachusetts AI safety rules while OpenAI and Google push a narrower version. Here's what the bill requires and why it matters.

#AI #Policy #Anthropic