Publishers Sue Google Over Gemini AI Training | TLY

AI Regulation Tracker  /  Litigation and enforcement

Publishers and Scott Turow Sue Google, Alleging Gemini Was Trained on Infringing Book Copies

On July 10, 2026, Hachette Book Group, Cengage Learning, Elsevier and bestselling author Scott Turow filed a proposed class action in the Southern District of New York accusing Google of willful copyright infringement to build its Gemini AI models, and of stripping copyright-management information from their works. This is a complaint, not a ruling. Nothing has been decided, and every point below is an allegation the plaintiffs will have to prove.

The Leveraged Years AI Regulation News

On July 10, 2026, three of the largest names in publishing and one of the country's best-known authors took Google to federal court over how it built Gemini. Hachette Book Group, Cengage Learning, and Elsevier, joined by novelist Scott Turow, filed a proposed class action in the Southern District of New York. The Association of American Publishers announced the case on July 13. I want to be clear at the top about what this is: a complaint. It is the plaintiffs' side of a dispute, filed to open a lawsuit. No judge has weighed in. Google has not yet responded. Everything in it is an allegation.

What the suit alleges

The core claim is copyright infringement, and the plaintiffs say it was willful. According to the complaint, Google copied millions of copyrighted books, textbooks, and academic journal articles to train its Gemini models, and did so knowing it lacked permission. The plaintiffs allege some of that material came from pirate sources and from behind paywalls, and that Google used full-length works, not snippets.

A big part of the theory leans on Google's own history with books. The plaintiffs allege that Google originally obtained many of these works for scope-limited programs like Google Books search and Google Play Books, then reused them for a purpose those programs never covered. As the complaint puts it, in their words, "publishers and authors never authorized Google to copy the works they received for Google Books for the completely separate purpose of training its AI models and building a multi-billion dollar competing business."

The plaintiffs also point to what they describe as Google's internal awareness of the risk. The AAP announcement highlights an internal warning, quoted in the complaint, that using copyrighted books from Google Play Books for AI training carried "$10Bs-$100Bs in potential fines." The plaintiffs use that to argue the alleged infringement was not accidental. Google has not responded to the allegations, and an internal risk estimate is not a finding of liability. It is one document the plaintiffs are pointing to.

The DMCA information-stripping claim is the part to watch

Most of the AI training cases so far have turned on one question: was copying for training fair use. This complaint adds a second front. Alongside the infringement counts, the plaintiffs bring a claim under Section 1202 of the Digital Millennium Copyright Act, the provision that deals with copyright-management information.

Copyright-management information is the identifying data attached to a work: the title, the author's name, terms of use, identifying numbers. Section 1202 makes it unlawful, in certain circumstances, to knowingly remove or alter that information, or to distribute works knowing it has been removed. The plaintiffs allege Google stripped this information from their works, removing titles and author names, to conceal where the training data came from.

That matters for counsel because it is a distinct theory with its own statutory-damages structure, separate from the infringement counts. It does not depend on winning the fair-use argument in the same way. If a defendant copied works and removed the identifying information in the process, a 1202 claim can stand on its own footing. Whether it succeeds here is an open question that will turn on the facts and on how the court reads the statute. But the presence of the claim is the escalation. It is why IP and media lawyers should read this filing even if they have been tracking the fair-use fights for a year.

What this is, and what it is not

This is a lawsuit at the starting line. A complaint sets out what the plaintiffs say happened and what they want. It is not evidence, it is not a verdict, and it is not a regulator's finding. The class has not been certified, which means the "class action" label right now is a request, not a status the court has granted. Google has every opportunity to move to dismiss, to answer, and to contest each allegation.

So when you see coverage saying Google trained Gemini on pirated books, read it as the plaintiffs claim Google did. The distinction is not pedantic. It is the difference between an allegation in a pleading and a fact a court has found. As of July 19, 2026, no court has found anything here.

It is also worth keeping this case in its lane. It is distinct from the AI copyright matters people already know. This is a different defendant and, on the 1202 count, a different theory than the fair-use and settlement stories that have dominated the AI training conversation. Treat it as its own filing with its own facts.

What this means for lawyers and rights holders

For litigators and in-house counsel, the practical read is about the theory of the case. The DMCA Section 1202 claim signals a shift in how rights holders are attacking training-data practices. Instead of only arguing that copying was infringement, plaintiffs are arguing that removing identifying information during ingestion was its own violation. If that framing gains traction, provenance and metadata handling inside a training pipeline become live legal exposure, not just engineering hygiene.

For anyone advising an AI developer, the allegations map to concrete questions worth asking now. Where did the training corpus come from, and can you document it. Was any of it obtained under a scope-limited license or program, and was it then used beyond that scope. Was copyright-management information preserved, removed, or altered anywhere in the pipeline, and can you show what happened to it. None of that is legal advice about this case. It is the checklist this complaint implies, and it is the same checklist a plaintiff's lawyer would run.

For rights holders and publishers, this is a filing to watch, not a precedent to rely on. If it survives a motion to dismiss, the 1202 theory becomes more attractive to other plaintiffs. If it does not, that tells you something too. Either way, the outcome is months or years away, and nothing about your own rights or claims changed because this complaint was filed.

What to do now

Read the complaint itself rather than the headlines, because the headlines compress allegations into flat statements of fact. Track the docket for Google's response and for any early motion, since the first real signal will be how the court handles a motion to dismiss, especially on the 1202 count. If you advise AI developers, treat the provenance and metadata questions above as a diligence exercise now, not after a ruling. And in your own writing and client memos, keep the framing precise: these are allegations, the case is at the pleading stage, and nothing has been decided.

Questions professionals are asking

Has a court found that Google infringed copyrights to train Gemini?

No. This is a complaint filed July 10, 2026 in the Southern District of New York. It sets out the plaintiffs' allegations. No court has ruled, the proposed class has not been certified, and Google has not yet answered. Everything in the filing is an allegation the plaintiffs must still prove.

Who filed the suit and against whom?

Hachette Book Group, Cengage Learning, Elsevier, and author Scott Turow filed a proposed class action against Google in the U.S. District Court for the Southern District of New York. The Association of American Publishers announced the case on July 13, 2026.

What is the DMCA Section 1202 claim about?

Section 1202 of the Digital Millennium Copyright Act addresses copyright-management information, meaning identifying data like titles, author names, and terms attached to a work. The plaintiffs allege Google removed or altered this information from their works during training to conceal the source of the data. It is a separate claim from the infringement counts, with its own statutory-damages structure, and it is the part of this case IP and media counsel are watching most closely.

What does the suit ask for?

According to the complaint and the AAP announcement, the plaintiffs seek statutory damages and an injunction to stop the alleged infringement, and ask the court to order infringing copies destroyed. These are requested remedies at the pleading stage, not amounts or orders any court has granted.

Is this the same as the other AI copyright cases in the news?

No. This is a distinct filing with a different defendant, and on the copyright-management-information count it advances a different theory than the fair-use and settlement matters that have dominated the AI training conversation. Read it on its own facts.

RELATED BRIEFINGS

Browse the full AI Regulation News tracker

Informational analysis for working professionals, not legal advice. This briefing describes allegations in a pending lawsuit that has not been decided. Confirm how any claim or development applies to your situation with qualified counsel in the relevant jurisdiction.