Connect with us

NEWS

Google’s Gemini Book Lawsuit Echoes a Fight It’s Fought Since 2005

Hachette, Cengage and Elsevier sued Google over Gemini’s book training days after a nearly identical Anthropic case closed for $1.5 billion.

Published

on

Hachette Book Group, Cengage Learning, Elsevier and author Scott Turow sued Google on July 10, accusing it of training Gemini on millions of copyrighted books and journal articles without permission. The case, filed in the Southern District of New York under case number 1:26-cv-05870, calls Google’s conduct “one of the most prolific infringements of copyrighted materials in history.” Ten days later, a nearly identical fight against a Google rival closed in a California courtroom for $1.5 billion.

The plaintiffs did not choose that timing. It still hands them a number Google cannot easily dismiss. This is also the third time in two decades that Turow has taken Google to court over its books.

Four Claims and a Warning From Inside Google

The complaint, brought by Hachette, Cengage Learning, Elsevier, Turow and co-plaintiff S.C.R.I.B.E., Inc., rests on four legal theories: direct copyright infringement, contributory copyright infringement, removal or alteration of copyright management information, and violations of the Digital Millennium Copyright Act. Google has not filed a public response.

The publishers say they gave Google their work for narrow purposes: searchable excerpts through Google Books, ebook distribution through Google Play Books, and academic discovery through Google Scholar. None of those agreements, they argue, authorized Google to copy full works into a dataset for a commercial generative AI product. The complaint also accuses Google of pulling additional material from pirate websites and paywalled academic journals through internet scraping.

One of the four claims, contributory copyright infringement, asks a court to hold Google responsible for infringement carried out across its platforms and partnerships. That is the same legal theory the Supreme Court narrowed last year when it ruled that Cox Communications was not liable for subscribers’ music piracy. How that precedent applies to a company accused of doing the scraping itself, rather than just providing the pipes, is untested.

The sharpest detail in the filing is an internal one. According to the complaint, a Google employee warned colleagues that using books submitted through Google Play’s publishing program for AI development carried serious legal risk. The employee reportedly estimated the potential fines at $10 billion to $100 billion. Other internal documents described in the complaint say Google specifically wanted professionally written, edited books, because models trained only on public domain texts performed worse. Curated facts, organized analysis and professionally edited prose were, according to the filing, the point.

A Rival Just Closed an Almost Identical Fight for $1.5 Billion

On July 20, a day before this case drew wider attention, U.S. District Judge Araceli Martínez-Olguín granted final approval to a $1.5 billion settlement between Anthropic and a class of authors and publishers. The deal resolves Bartz v. Anthropic, filed in 2024 by novelists Andrea Bartz and Charles Graeber and non-fiction writer Kirk Wallace Johnson over Anthropic’s use of pirated books to train its Claude chatbot.

The math behind that settlement matters here. Judge William Alsup, who has since retired, ruled in June 2025 that training an AI model on legally purchased books counted as fair use. Downloading millions of titles from shadow libraries like Library Genesis, Pirate Library Mirror and the Books3 dataset did not. That finding exposed Anthropic to statutory damages that legal observers said could run into the hundreds of billions of dollars had the case reached a jury. Anthropic settled instead, in September 2025, before any trial began.

The final tally: more than 482,000 books, at roughly $3,000 each, adding up to $1.5 billion, described by plaintiffs’ attorney Justin Nelson as the largest copyright recovery in U.S. history. About 91% of eligible authors and publishers have already claimed a payment. Google is far from the only company facing this exact accusation. The AI music startup Suno drew similar claims after leaked code exposed a scraping operation that fair use could not shield. Google’s new complaint describes a similar sourcing pattern: platform content repurposed beyond its licensed scope, plus material pulled from pirate sites.

Because Anthropic settled rather than let a jury rule, its case sets no binding legal precedent. Google can still argue its own training is different. But the publishers suing Google now have a real dollar figure to point toward, not just an internal estimate.

The Books Gemini Allegedly Learned to Imitate

The complaint does not deal in generalities. It names specific titles, tying each plaintiff to concrete, identifiable works it says Google used without permission.

  • Hachette: “The Wild Robot” by Peter Brown, “The Fifth Season” by N.K. Jemisin, “Moon Glacier National Park” by Becky Lomax, “Who Could That Be at This Hour?” by Lemony Snicket, and Turow’s “Innocent”
  • Cengage: “Cognitive Psychology,” “Principles of Economics,” “Milady Standard Barbering,” “Nutrition: Concepts and Controversies” and “Calculus: Early Transcendentals”
  • Turow and S.C.R.I.B.E.: “Presumed Innocent,” “Innocent” and “Testimony”
  • Elsevier: unnamed copyrighted journal articles across its scientific and medical publishing catalog

The filing claims Gemini generated content drawing on “The Fifth Season” and produced answers using characters, events and details from “Who Could That Be at This Hour?” It goes further, arguing Gemini can manufacture cheap substitutes for professionally published work. The plaintiffs estimate the system could produce a 100-page murder mystery set in a quiet seaside town in about 20 minutes for 39 cents. “No publisher or author can compete with that,” the complaint states. Those time and cost figures are the plaintiffs’ own estimate and have not been tested in court.

The proposed class covers owners of registered copyrights in books carrying an ISBN (International Standard Book Number, the identifier printed on nearly every commercially published book), plus journal articles identified by a DOI or an ISSN. To qualify, a rights holder would need to show Google copied their work through one of its services, through scraping, or through Gemini’s training pipeline. No judge has certified that class yet.

Turow’s Third Round With Google Since 2005

Turow’s fight with Google is not new. He led the Authors Guild as its president when the group first sued Google on September 20, 2005, over its plan to scan millions of library books for Google Books. Google agreed to pay $125 million to settle that case in 2008. Judge Denny Chin rejected the deal in 2011, saying it went “too far” and would reward Google for “wholesale copying of copyrighted works without permission.”

The case did not end there. Chin ultimately ruled for Google in 2013, finding its scanning and snippet display transformative under fair use, a finding the Electronic Frontier Foundation tracked through the fair-use proceedings that followed the settlement’s collapse. A Second Circuit panel affirmed that ruling in 2015, and the Supreme Court declined a further appeal in 2016, closing a fight that had run more than a decade. That history is why the new complaint argues this case is different: search snippets are one use, and a generative model competing with the books it trained on is another.

Turow has stayed in the fight since. He is a plaintiff in a 2023 Authors Guild class action against OpenAI, and the same publisher coalition, including Hachette, Cengage, Elsevier, Macmillan and McGraw Hill, sued Meta over its Llama models roughly two months before turning to Google.

A separate, newer fight had already opened over Bard and Gemini training: In re Google Generative AI Copyright Litigation, first filed by illustrators and writers in 2023 and now before Judge Eumi Lee in Oakland. Hachette and Cengage moved to intervene there in January, then withdrew in favor of the fresh New York filing after Google raised the prospect of a three-year statute of limitations defense. The federal docket lists Hachette’s case tied to the French firms Lagardère and Bolloré among its intervenors, a reminder the plaintiffs are not purely American corporate interests.

Case Filed Core Claim Status
Authors Guild v. Google 2005, SDNY Mass book scanning for search snippets Google won on fair use; closed 2016
In re Google Generative AI Copyright Litigation 2023, N.D. Cal. Bard and Gemini trained on authors’ and illustrators’ work Pending before Judge Eumi Lee
Bartz v. Anthropic 2024, N.D. Cal. Claude trained on pirated books Settled for $1.5 billion, approved July 20, 2026
Hachette Book Group et al. v. Google July 10, 2026, SDNY Gemini trained on repurposed and scraped books Pending, no ruling yet

Laid side by side, the pattern is hard to miss. Google has faced some version of this fight for two decades, and the one time a similar case against a rival reached a conclusion, it ended in a nine-figure per-title payout.

Google Already Won This Fight Once, on Narrower Ground

Google’s defense will likely lean on the 2015 ruling that protected Google Books. The appellate opinion that shielded Google’s scanning project found that showing users short snippets and enabling search across millions of books did not substitute for the books themselves. It was transformative, the court said, not competitive.

The new plaintiffs argue that logic breaks down once the same books feed a model built to write fiction, summarize plots or draft textbook explanations on demand. A search index sends readers to a book. Gemini, they claim, can write something that replaces it.

The idea, in short, is that any fair use argument that Gemini has would arguably be mooted by the fact that they allegedly acquired the books unlawfully.

Kirk Sigmon, a technology and intellectual property lawyer at KellDann Law, made that point to Al Jazeera. If the books came from pirate sites or were repurposed outside a licensing deal to begin with, he suggested, the fair use question may never need answering. That is roughly how Anthropic’s case turned, not on whether training itself was legal, but on how the books reached Anthropic’s servers.

What Happens Next in Google’s Courtroom Fight?

Google has not filed a substantive response to the New York complaint and has not addressed the specific allegations publicly. The case now heads toward motions, discovery and a fight over class certification, the same slow track that took the original Google Books case more than a decade to resolve. No trial date exists yet.

  • Confirmed: the lawsuit was filed July 10, 2026, in the Southern District of New York under case number 1:26-cv-05870, resting on four distinct legal claims.
  • Confirmed: Hachette and Cengage withdrew a bid to join the pending California case before filing separately in New York.
  • Unconfirmed: the $10 billion to $100 billion internal estimate has not been authenticated by any court and remains an allegation drawn from internal communications described in the complaint.
  • Unconfirmed: the claim that Gemini can produce a finished 100-page novel in 20 minutes for 39 cents is the plaintiffs’ own estimate, not a tested finding.

For publishers, authors and academic rights holders wondering whether they might belong to the eventual class, the practical steps are straightforward. Confirm which titles are registered with the U.S. Copyright Office, since the class depends on works carrying an ISBN, DOI or ISSN. Revisit existing Google Books, Google Play Books or Google Scholar agreements to see what they actually permit, and keep records of publication dates and any Google correspondence about AI training. Inclusion in any class still depends entirely on a judge’s decision that has not been made.

Google will get its turn to answer in court. Until it does, one number hangs over the case louder than its own staff’s estimate: the $1.5 billion an almost identical fight against a rival just cost, ten days after Google’s own suit was filed.

Frequently Asked Questions

Has Google responded to the lawsuit?

Not to this specific complaint. Google has not issued a public response to the New York allegations. In the related California case, however, Google has argued it does not need individual permission to use content uploaded to its own platforms and has called a proposed rights-holder class overbroad, a position likely to resurface here.

How does this case differ from the Anthropic settlement?

Anthropic’s $1.5 billion deal resolved claims that it downloaded pirated books from shadow libraries, and it settled before a jury ever ruled, so it sets no binding precedent. The Google case adds claims the Anthropic dispute did not centrally feature, including removal of copyright management information, and remains entirely unresolved with no trial date set.

Can independent or self-published authors join the case?

Possibly, once a class is certified. The proposed class covers any registered copyright owner whose book carries an ISBN or whose journal article carries a DOI or ISSN, provided they can show Google obtained the work through its services, through scraping, or through Gemini’s training process.

Will this affect people currently using Gemini?

Not immediately. The lawsuit does not ask a court to shut Gemini down for users. It asks for an injunction limiting future training uses of the disputed works, an accounting of how Google’s datasets were built, and court-supervised destruction of any infringing copies, remedies aimed at training practices rather than current access.

What is In re Google Generative AI Copyright Litigation?

It is a separate, earlier case filed in 2023 by illustrators and writers in federal court in Oakland, now before Judge Eumi Lee, that also accuses Google of training generative AI models on copyrighted work. Hachette and Cengage tried to join it in January 2026 before withdrawing to file the broader New York case instead.

Disclaimer: This article summarizes allegations from a civil complaint that Google has not yet answered in court, is provided for informational purposes only, and is not legal advice; anyone considering a potential copyright claim should consult a qualified attorney.

As the founder of Thunder Tiger Europe Media, Dr. Elias Thornwood brings over 25 years of experience in international journalism, having reported from conflict zones in the Middle East, Asia, and Africa for outlets like BBC World and Reuters. With a PhD in International Relations from Oxford University, his expertise lies in geopolitical analysis and global diplomacy. Elias has authored two bestselling books on European foreign policy and received the Pulitzer Prize for International Reporting in 2015, establishing his authoritativeness in the field. Committed to trustworthiness, he enforces rigorous fact-checking protocols at Thunder Tiger, ensuring unbiased, evidence-based coverage of worldwide news to empower informed global audiences.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending