Policy

Google Faces New Publisher Lawsuit Over Gemini Training Data

A class action from major publishers and authors accuses Google of using copyrighted books to train Gemini beyond the scope of Google Books and Google Play agreements, widening the legal fight over AI training.

By Michael C ·

Google Faces New Publisher Lawsuit Over Gemini Training Data
Wikimedia Commons / David Nagle, CC BY-SA 4.0.

Google is facing another major legal test over the data used to train Gemini. TechCrunch reported on July 14 that publishers and authors including Hachette, Cengage, Elsevier, Scott Turow and S.C.R.I.B.E. filed a class action accusing Google of using copyrighted works to train its AI platform without authorization. The complaint, filed in the Southern District of New York, also alleges that Google removed or changed copyright information to conceal that Gemini models were trained on the material. Google had not immediately responded to TechCrunch's request for comment.

The case is important because it is not simply another fight over whether public web data can be used for training. The plaintiffs argue that many works reached Google through scope-limited programs such as Google Books and Google Play Books, where authors and publishers had agreed to particular uses, not broad commercial model training. That distinction could matter. Courts may view an open-web scrape differently from a company repurposing material it received under a narrower commercial or archival relationship.

The broader copyright fight has produced mixed signals. TechCrunch noted that early California decisions have favored AI companies on fair-use grounds, while Anthropic's separate litigation produced a large settlement tied to pirated works. Those outcomes do not settle the Google case. The Southern District of New York will have its own facts, its own plaintiffs and a different judge. The question is whether fair use protects model training when the source material entered a company's systems through a relationship that allegedly carried limits.

The publisher case centers on whether books supplied for discovery or distribution were later repurposed for AI training. Image: Wikimedia Commons / Roc0ast3r, CC0.
The publisher case centers on whether books supplied for discovery or distribution were later repurposed for AI training. Image: Wikimedia Commons / Roc0ast3r, CC0.

Publishers are also contesting leverage. Generative AI can summarize, rewrite, draft, translate and answer questions in ways that may substitute for portions of the market that books, textbooks and reference works occupy. A model does not have to reproduce an entire book to create economic pressure. If it can absorb and repackage the value of many copyrighted works, publishers will argue that compensation and consent should be part of the system.

Google's position will likely rely on the same broad defense other AI companies have made, that training is a transformative use and that models do not store books as a simple substitute for copies. The plaintiffs will likely emphasize alleged internal concern about liability and the removal or alteration of copyright information. That makes metadata a key issue. If a court finds that copyright-management information was knowingly stripped or changed, the case could move beyond the usual training-data debate into more damaging territory.

The institutional history raises the stakes. Google Books began as a search and digitization project that promised discovery through snippets and bibliographic information. AI training changes the value exchange. Instead of helping readers find a book, a model may use books to answer a question before the reader reaches the author or publisher. That is why the dispute is becoming a test of consent, not only a test of copying.

Copyright and contract questions around AI training are increasingly moving from product debate to courtrooms. Image: Wikimedia Commons / Krzysztof Golik, CC BY-SA 4.0.
Copyright and contract questions around AI training are increasingly moving from product debate to courtrooms. Image: Wikimedia Commons / Krzysztof Golik, CC BY-SA 4.0.

For AI companies, the practical implication is that provenance is becoming a product risk. Training data can no longer be treated as a hidden input whose legal exposure will be resolved after deployment. Enterprise customers, publishers, regulators and investors are asking which datasets were used, under what rights, and whether opt-outs or licenses were honored. That pressure will push model developers toward cleaner data rooms, more licensing deals and stronger documentation.

For publishers, litigation is only one strategy. Licensing, collective bargaining, technical controls and product partnerships will shape the market faster than courts in some cases. But lawsuits set the fallback rules. The Google case asks whether the search giant's long relationship with books gives it permission to train Gemini, or creates a higher obligation because the trust was narrower. That answer will help define the economics of AI knowledge products.

Topics: Google, Gemini, copyright, AI training data