Major Publishers and Authors File Landmark Class-Action Lawsuit Against Google Over AI Training Data, Alleging Widespread Copyright Infringement

A consortium of prominent publishing houses, academic content providers, and acclaimed authors, including Hachette Book Group, Cengage Learning, Elsevier, and the celebrated novelist Scott Turow, alongside S.C.R.I.B.E., Inc., have initiated a proposed federal class-action lawsuit against technology giant Google LLC. Filed on July 10, 2026, in the U.S. District Court for the Southern District of New York under case number 1:26-cv-05870, the lawsuit accuses Google of systematically copying vast quantities of copyrighted books and journal articles without explicit permission to train its advanced Gemini artificial intelligence models.
The complaint characterizes Google’s alleged actions as "one of the most prolific infringements of copyrighted materials in history." This sweeping accusation sets the stage for a potentially transformative legal battle that could redefine the boundaries of intellectual property rights in the burgeoning era of generative AI. However, it is crucial to note that these are unproven allegations, and no court ruling has yet been issued regarding Google’s liability in this complex and high-stakes case. The legal proceedings are in their early stages, with significant implications for both the technology and publishing industries hanging in the balance.
The Broader Context of AI and Copyright Law
This lawsuit against Google is not an isolated incident but rather a significant development within a rapidly expanding landscape of legal challenges confronting generative AI developers. Over the past few years, numerous artists, authors, and content creators have initiated legal actions against companies behind leading AI models, alleging that their copyrighted works were used without consent to train these systems. Notable cases include Getty Images suing Stability AI for alleged copyright infringement related to its image-generating AI, and a group of authors, including Sarah Silverman, filing similar lawsuits against OpenAI and other AI developers for using their books. More recently, The New York Times initiated a significant lawsuit against OpenAI and Microsoft, alleging infringement of its journalistic content.
The core of these disputes often revolves around the doctrine of "fair use," a critical defense in U.S. copyright law that permits limited use of copyrighted material without acquiring permission from the rights holder. AI companies frequently argue that training their models on publicly available data, even copyrighted material, constitutes fair use because it is "transformative"—meaning it creates something new and different, rather than merely reproducing the original work. This argument posits that AI models learn concepts and patterns from the data, much like a human, rather than storing and regurgitating specific copyrighted works.
However, content creators and rights holders contend that the scale and nature of AI training, which often involves ingesting entire datasets of copyrighted works, goes far beyond traditional notions of fair use. They argue that the output of these AI models can directly compete with, or even devalue, the original works, thus causing significant economic harm. The Google lawsuit is poised to be a pivotal case in clarifying how existing copyright laws apply to the novel capabilities and data demands of generative AI.
Distinguishing from the Google Books Precedent: A Critical Legal Angle
A central element of the plaintiffs’ strategy is to differentiate this current lawsuit from Google’s previous legal victory in the landmark Authors Guild v. Google case. In 2015, a federal appellate court ruled that Google’s digitization of millions of books to create a searchable database and display small "snippets" was a "transformative" use protected under fair use. This decision allowed Google to continue its Google Books project, which aimed to make human knowledge more accessible.
The plaintiffs in the current Gemini lawsuit contend that the 2015 ruling was narrowly tailored to searchable book databases and limited excerpts for discovery purposes. They argue that using entire copyrighted books and articles to train a commercial generative AI model like Gemini serves an entirely different commercial purpose. Unlike Google Books, which provided pathways to the original works, they claim Gemini is designed to generate new content, potentially creating low-cost substitutes that directly compete with the original authors and publishers. The distinction, they assert, lies between search-focused digitization for information access and the training of generative models intended for content creation. This differentiation is critical to their legal strategy, aiming to prevent Google from leveraging its previous fair use victory as a shield for its AI training practices.
The Genesis of Gemini and the Allegations of Unlicensed Data Use
Google officially launched its Gemini AI model in December 2023, positioning it as its most capable and versatile AI yet, capable of understanding and operating across various modalities, including text, code, audio, image, and video. Subsequent updates, such as Gemini 3.1 Flash Live with faster responses and longer conversation contexts, and its integration into products like a new Google Home speaker, underscore its strategic importance to Google’s future. The plaintiffs contend that the development of such an advanced system heavily relied on vast troves of copyrighted material.
The complaint outlines four specific legal claims against Google: direct copyright infringement, contributory copyright infringement, removal or alteration of copyright-management information, and violations of the Digital Millennium Copyright Act (DMCA). These claims collectively assert that Google not only directly copied protected works but also facilitated infringement and deliberately bypassed safeguards.
The plaintiffs allege that Google obtained protected material through several channels. Firstly, they highlight Google’s existing relationships with publishers through services like Google Books, Google Play Books, and Google Scholar. Under agreements for these services, publishers provided books and journal articles for specific uses, such as displaying searchable excerpts, distributing e-books, and aiding academic discovery. The core of the plaintiffs’ argument here is that these agreements did not grant Google permission to copy entire works into massive datasets for commercial generative AI development. They accuse Google of repurposing material obtained through established publishing relationships for uses explicitly not covered by their contracts.
Secondly, the complaint details Google’s alleged "extensive internet scraping" to collect books and other protected material. This practice, often involving automated bots trawling the web, reportedly included sourcing content from illicit "pirate websites" as well as publications hidden behind subscription or paywalls, thereby circumventing legitimate access and licensing mechanisms. This broad-scale collection from diverse and often unauthorized sources further strengthens the plaintiffs’ claims of widespread and systemic infringement.
Internal Warnings and the Quest for Quality Data
Compelling evidence presented in the complaint includes internal Google communications that allegedly reveal awareness of the legal risks involved. According to the lawsuit, a Google employee warned during an internal discussion that utilizing books submitted via Google Play publishing agreements for AI development could expose the company to "serious legal risks." The employee reportedly estimated potential liabilities ranging from a staggering ten to one hundred billion dollars—a clear indication of the perceived magnitude of the infringement risk within the company itself.

Further internal documents cited in the complaint suggest that Google actively sought "professionally written books" to enhance the performance of its Gemini AI system. The plaintiffs claim that internal tests demonstrated a significant performance gap: models trained exclusively on public domain books performed demonstrably worse than those trained on collections that included copyrighted works. This alleged finding motivated Google to incorporate works characterized by "curated facts, organized analyses, fictional narratives, and professionally edited writing," presumably to improve Gemini’s factual accuracy, coherence, and creative capabilities. This detail could be crucial in demonstrating Google’s alleged intent and knowledge of using copyrighted material to achieve superior AI performance.
Specific Examples of Infringement and Gemini’s Alleged Outputs
To substantiate their claims, the plaintiffs have highlighted numerous specific works as examples of allegedly infringed material. Hachette Book Group lists titles such as Peter Brown’s The Wild Robot, N.K. Jemisin’s The Fifth Season (a Hugo Award-winning novel), Becky Lomax’s Moon Glacier National Park, and Lemony Snicket’s Who Could That Be at This Hour? Author Scott Turow, known for his legal thrillers, includes his works Innocent, Presumed Innocent, and Testimony.
Cengage Learning, a major educational publisher, references a range of textbooks, including Cognitive Psychology, Principles of Economics, Milady Standard Barbering, Nutrition: Concepts and Controversies, and Calculus: Early Transcendentals. Elsevier, a global leader in academic publishing, points to numerous copyrighted journal articles within its section of the complaint. The plaintiffs emphasize that these examples represent only a fraction of the vast number of books and articles allegedly copied in connection with Gemini’s training.
The filing also presents examples of Gemini’s alleged output to illustrate how the system can produce content directly related to specific protected books. For instance, it claims that Gemini generated material based on N.K. Jemisin’s The Fifth Season, and produced responses involving characters, events, and intricate details from Lemony Snicket’s Who Could That Be at This Hour? These examples are intended to demonstrate not just that the works were used for training, but that the AI can recall and leverage specific elements, potentially infringing on the original creators’ rights.
A particularly striking claim by the plaintiffs revolves around the economic threat posed by Gemini. They argue that the AI system can generate "low-cost substitutes" for professionally published works. As an illustration, the complaint estimates that Gemini could produce a 100-page murder mystery novel, set in a quiet seaside town, in approximately 20 minutes for a mere 39 cents. The complaint starkly concludes, "No publisher or author can compete with that." These time and cost figures are allegations from the plaintiffs and have not been independently verified by the court, but they underscore the potential for significant market disruption that copyright holders fear.
The Proposed Class and Far-Reaching Remedies Sought
The proposed class in this lawsuit would encompass a broad group of copyright owners. It includes owners of registered copyrights in books identifiable by International Standard Book Numbers (ISBNs), as well as journal articles marked with Digital Object Identifiers (DOIs) or International Standard Serial Numbers (ISSNs). To qualify for inclusion in the class, members would need to demonstrate that Google copied their works from one of its services, obtained them through web scraping, or used them in connection with Gemini’s training. The court has not yet certified this proposed class, a crucial procedural step that will determine the scope and reach of the lawsuit.
The plaintiffs are seeking a comprehensive array of remedies under federal copyright law. These include statutory damages or compensation based on their claimed losses and Google’s alleged profits derived from the infringement, along with legal fees. Beyond monetary compensation, they are seeking an injunction—a court order to stop Google from continuing what they allege to be unauthorized copying. This injunction would specifically restrict the use of protected books and articles in Gemini training and other related AI development activities.
Furthermore, the complaint requests an accounting of the works and methods Google employed to train Gemini. This would entail Google providing detailed information about the specific copyrighted materials obtained, their original sources, and precisely how they were incorporated into Google’s AI systems. Perhaps the most drastic remedy sought is a court-supervised destruction of infringing copies and any datasets derived from their works. The legal feasibility and appropriateness of such an order in the context of complex AI models, where copyrighted material might be deeply embedded within the neural network’s architecture, remains a significant question that the court would have to address.
For publishers, authors, and academic rights holders who might potentially be part of the proposed class, several practical steps are advisable as the case progresses. They should verify which of their titles are officially registered with the U.S. Copyright Office, as the proposed class definition hinges on registered works identified by ISBNs, DOIs, or ISSNs. Reviewing existing agreements with Google Books, Google Play Books, and Google Scholar is also crucial to understand the permitted uses under those contracts. Additionally, maintaining meticulous records of publication dates, distribution channels, and any communications with Google regarding the use of their works in AI training datasets could prove invaluable. Tracking the progress of the class certification process will also be important, as inclusion in any potential class will depend directly on the court’s decision.
The Path Forward: Google’s Response and Broader Implications
This lawsuit marks Hachette and Cengage’s renewed effort to pursue legal action against Google over AI training. They had previously attempted to join a separate copyright lawsuit filed against Google in California in 2023, which involved authors and visual artists. However, Google opposed their participation, leading the publishers to withdraw that effort before submitting the current, more focused case in New York. This history suggests a determined and strategic approach from the plaintiffs.
As of now, Google has not issued a detailed public response specifically addressing the allegations contained within the New York complaint. The company will have the opportunity to respond as the case proceeds through various legal stages, including motions to dismiss, potential discovery processes where evidence is exchanged, and the critical class-certification process. No trial date has been set, and the court has yet to rule on any of the four claims presented by the plaintiffs.
The outcome of Hachette Book Group Inc. et al. v. Google LLC holds profound implications for the future trajectory of both the technology and creative industries. If the plaintiffs prevail, it could establish a precedent that necessitates extensive licensing agreements for AI developers seeking to train their models on copyrighted content. This could create new revenue streams for creators and rights holders but might also lead to increased costs and potential slowdowns in AI innovation, particularly for smaller developers or open-source projects. Conversely, a ruling in Google’s favor could solidify a broader interpretation of fair use in the context of AI training, potentially enabling AI companies to continue training on vast datasets without explicit permission or payment, much to the dismay of content creators.
This case, alongside other ongoing AI copyright lawsuits, is set to significantly shape the legal and economic framework for generative AI, influencing everything from how AI models are built to the availability and cost of AI-generated content in the years to come. The dialogue between technological advancement and intellectual property rights is reaching a critical juncture, and the courts are now tasked with drawing the lines for a new digital frontier.







