NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning, and Elsevier have initiated legal action against Google regarding its Gemini artificial intelligence platform. Author Scott Turow and his company, S.C.R.I.B.E., have joined the class action proposal. The lawsuit was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the class.

According to the complaint, Google sourced material from Google Books, Google Play Books, and Google Scholar. Publishers and authors provided these works for specific functionalities, including search, sales, and research access. The plaintiffs assert that such arrangements did not permit broader commercial AI training. They further allege that Google downloaded extensive web-scraped datasets containing copyrighted content, some of which originated from known piracy sources and paywalled services.
The 57-page complaint details four claims under federal law. Three relate to alleged reproduction via Google services, web scraping, and the development or training of Gemini. The fourth claim references the Digital Millennium Copyright Act, alleging Google removed or altered copyright management information from training materials. The filing also mentions internal discussions about using publisher-supplied books, with one assessment estimating potential fines ranging from $10 billion to $100 billion. These allegations have not yet been tested in court.
Class Includes Registered Works
The proposed class encompasses owners of registered U.S. copyrights in qualifying books and journal articles. Eligible books must bear an International Standard Book Number (ISBN), while eligible articles require a Digital Object Identifier or International Standard Serial Number. The class definition includes works allegedly copied from Google services or obtained through web scraping, as well as those reproduced during Gemini’s development or training.
Restrictions based on registration timing also apply. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another stipulates registration within three months of publication. The lawsuit excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case proceeds on behalf of the larger group.
Damages and Court-Ordered Accounting Sought
The plaintiffs are requesting statutory damages or actual damages for proven infringements, along with damages equivalent to Google’s profits derived from any violations. They also seek an injunction, recovery of legal costs, and a jury trial. The complaint does not specify a total damages amount but demands that Google disclose Gemini training data, collection methods, and known model capabilities via a court-ordered accounting.
This accounting would identify copyrighted books and other works used in Gemini’s training, as well as detail how Google collected, copied, processed, and encoded these materials. The plaintiffs additionally request that the court supervise the destruction of any unauthorized copies under Google’s control. Earlier, Hachette and Cengage sought to join separate litigation involving Google’s generative AI in California. The New York case extends the list of plaintiffs to include Elsevier, Turow, and S.C.R.I.B.E., while pursuing claims related to Google services, web scraping, and Gemini’s training process.
