Publishers open a second front in the AI copyright war: Google
Somewhere inside Google, someone apparently ran the numbers on training AI with copyrighted books and wrote down the answer. An internal document cited in a new lawsuit warned the practice could be “highly problematic for Google,” with exposure of “$10Bs-$100Bs in potential fines,” according to TechCrunch. The plaintiffs would like a jury to check that math.
Hachette Book Group, Cengage, Elsevier, novelist Scott Turow and the author group S.C.R.I.B.E. sued Google in the Southern District of New York on July 14, alleging the company trained its Gemini models on millions of copyrighted books without authorization and stripped copyright management information to cover its tracks, Engadget reports. The complaint also alleges Google downloaded “web scrapes of virtually the entire internet, including from known pirate sources and from behind legitimate paywalls,” Al Jazeera reports.
The sharpest claim in the nearly 60-page complaint is about provenance: Google held books under scope-limited deals, through Google Books scanning and Google Play distribution, and repurposed them anyway. “Google illegally copied works from all these scope-limited programs for AI training, knowing it lacked authorization to do so,” the complaint says, per TechCrunch.
The suit landed in a rough week for AI defendants. Five days earlier, The New York Times and Daily News asked the judge in their case against OpenAI to impose sanctions, alleging the company hid a database of some 78 million de-identified ChatGPT conversations it used internally to assess regurgitation of copyrighted content, along with detection tooling built under the name “Project Giraffe,” TechCrunch reported. OpenAI spokesperson Drew Pusateri called the allegations “blatantly false.” The newspapers’ lead counsel, Ian B. Crosby, put the logic that now runs through both fights in one sentence:
“If OpenAI genuinely believed that copying our clients’ journalism was fair and legal, it wouldn’t have hid the truth about having done it.”
The $1.5 billion floor
Every complaint in this genre is drafted in the shadow of Bartz v. Anthropic. In that case, Judge William Alsup ruled that training on lawfully obtained books was “exceedingly transformative” fair use, but sent Anthropic’s downloading of millions of pirated copies to trial. Anthropic paid $1.5 billion to settle, roughly $3,000 per book across an estimated 500,000 works — what the settlement motion called the largest publicly reported copyright recovery in history. Many eligible authors, per TechCrunch, turned that money down to sue on their own.
Our read: this complaint treats that settlement as the floor, not the ceiling.
Anthropic taught rights holders two things. Claims about how a lab got the books survive the fair-use defense that kills claims about training on them. And a frontier lab, facing statutory damages in front of a jury, will pay ten figures to make the question go away.
Google’s hand is better than Anthropic’s was
Here is the twist the plaintiffs have to route around. Google is the only defendant in this fight with an appellate fair-use win on these exact books: Authors Guild v. Google, where the Second Circuit blessed mass scanning for search snippets as transformative in 2015. If Google can show Gemini’s book corpus came from its own scanned and licensed copies rather than a pirate mirror, its posture looks much stronger than Anthropic’s, whose problem was never training but Library Genesis. That is why the complaint leans on scope, arguing the Books and Play agreements never authorized AI training, and why the “known pirate sources” allegation is the one to watch in discovery. Whether it holds up decides whether this is an Authors Guild rerun or an Anthropic sequel.
Google did not immediately respond to requests for comment from TheWrap or Al Jazeera. The complaint didn’t wait: per TheWrap, it calls the conduct “one of the most prolific infringements of copyrighted materials in history” and says a company “desperate to maintain its online dominance” had “abandoned its early motto of ‘Don’t be evil.'”
Cassandra writes about technology as a cultural force — what it does to how we live, work, and understand ourselves. She has a background in cognitive science and too many browser tabs open. Based in Vancouver.
Leave a Reply