A Tale of Two AI Fair Use Decisions
On June 23 and 25, 2025, two federal judges in the Northern District of California issued decisions on summary judgment motions in two closely watched cases in which book authors claim that AI companies infringed their copyrights. These are the first cases in the U.S. to decide whether training a generative AI model using copyrighted material is fair use. Both judges held that under the facts of the cases before them, using copyrighted books for the purpose of training large language models (LLMs) underlying AI systems was fair use. While at first blush this may seem like slam-dunk victories for the AI companies, the cases revealed ways in which some AI companies have used “pirated” copies of copyrighted books, rather than purchasing or licensing legitimate copies, resulting in other aspects of the decisions that are less favorable to AI companies. Both cases may expose companies to liability for copyright infringement in further proceedings in these cases or in cases brought by other copyright owners.
In the first case decided last week, Bartz v. Anthropic PBC, No. 24-cv-5417-WHA (June 23, 2025), three authors whose books had been used to train Anthropic’s Claude AI model sued for copyright infringement. Anthropic had initially downloaded from certain websites millions of copies of books which it knew had been assembled without authorization of the copyright owners (“pirated”, in the court’s parlance). Subsequently, after inquiring about licensing books from major publishers, Anthropic also “spent many millions of dollars to purchase millions of print books, often in used condition”, from major book distributors and retailers, including some of the books it had previously obtained from the pirate sources without payment. It then scanned the purchased books into digital form, discarded the paper originals, and created a general “research library” which included both the scanned purchased books and the pirated books. Anthropic then used its library to train its Claude AI system.
Anthropic moved for summary judgment, contending that its use of the copyrighted books was a “fair use” under the U.S. Copyright Act. In analyzing the four statutory factors set forth in the law, Judge Alsup reached three major conclusions. First, he held that Anthropic’s copying of the books that it purchased and then used to train LLMs was a fair use. In this regard, he concluded that the purpose and character of the use (the first statutory factor) was “exceedingly” and “spectacularly” transformative, which militated in favor of fair use. He also held that as to the fourth statutory fair use factor (the effect of the use upon the potential market for or value of the copyrighted work), the authors’ contention that “training LLMs will result in an explosion of works competing with their works” is “no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works” and “is not the kind of competitive or creative displacement that concerns the Copyright Act.” In addition, with respect to the authors’ further argument that “training LLMs displaced (or will) an emerging market for licensing their works for the narrow purpose of training LLMs”, Judge Alsup concluded that “such a market for that use is not one the Copyright Act entitles Authors to exploit.”
Second, Judge Alsup held that Anthropic’s digitization of the books it purchased in print form was a fair use. This conclusion seemingly would have “closed the book” on the plaintiffs’ copyright claims relating to Anthropic’s use of the purchased books.
Third, however, Judge Alsup held that Anthropic had no right to use pirated copies for its central library, and that doing so was not a fair use. In this connection, he stated that he would hold “a trial on the pirated copies and the resulting damages, actual or statutory (including for willfulness).” He also indicated that a motion for class certification remains pending (which previously proposed one class related to pirated works, and another class for works that were purchased and scanned). Given that a plaintiff in a copyright infringement action can seek up to $150,000 in statutory damages per work for a willful violation, Anthropic may potentially face significant liability in a trial, even though its summary judgment motion was granted on its fair use defense with respect to its use of books that it purchased to train LLMs.
The second case, which was decided two days later, is Kadrey v. Meta Platforms, Inc., No. 23-dv-03417-VC (N.D. Cal. June 25, 2025). Thirteen published authors (including Sarah Silverman and Andrew Sean Greer) sued Meta for having downloaded hundreds of copies of books whose copyrights they owned and used them to train its AI model, named Llama. The plaintiffs and Meta each filed motions for summary judgment on plaintiffs’ primary claim of copyright infringement arising from Meta’s unauthorized copying of the plaintiffs’ books.
On the first statutory fair use factor, Judge Chhabria held that Meta’s use of the copyrighted books to train its AI system was “highly transformative” (similar to Judge Alsup’s holding in Bartz). He contrasted the purpose of Meta’s copying (“to train its LLMs”) with that of the plaintiffs’ books (“to be read for entertainment or education”), and noted that Meta had “post-trained” its models to prevent them from outputting certain text from their training data (so that Llama could not generate more than 50 words from any of the plaintiffs’ books, even in response to specific prompts).
With respect to the fourth fair use factor, the plaintiffs argued that “Meta’s unauthorized use of their books for LLM training harms the market for licensing books for that purpose.” In this regard, Judge Chhabria agreed with Judge Alsup, holding that “whether such a market exists or is likely to develop is irrelevant, because this market is not one that the plaintiffs are legally entitled to monopolize” and that “harm from the loss of fees paid to license a work for a transformative purpose is not cognizable.” In other respects, however, Judge Chhabria differed from Judge Alsup in his analysis of the fourth factor, which he characterized as the most important. He wrote that “Judge Alsup focused heavily on the transformative nature of generative AI while brushing aside concerns about the harm it can inflict on the market for the works it gets trained on”, and that “using books to teach children to write is not remotely like using books to create a product that a single individual could employ to generate countless competing works with a miniscule fraction of the time and creativity it would otherwise take.”
Judge Chhabria also had a different view of the importance of the source of the books used in the fair use analysis. He rejected the plaintiffs’ contention that “the fact that Meta downloaded the books from shadow libraries and did not start with an ‘authorized copy’ of each book gives them an automatic win. To say that Meta’s downloading was ‘piracy’ and thus cannot be fair use begs the question because the whole point of fair use analysis is to determine whether a given act of copying was unlawful.” He also rejected Meta’s countervailing argument that “its use of shadow libraries is irrelevant to whether its copying was fair use”, and stated that it would be relevant “if it benefitted those who created the libraries and thus supported and perpetuated their unauthorized copying and distribution of copyrighted works.” He concluded, however, that “plaintiffs have not submitted any evidence about this.”
Judge Chhabria ultimately held that in the case before him, Meta’s use of the plaintiffs’ copyrighted works was a fair use. A key consideration in his decision was his conclusion that, as to the fourth fair use factor, “the plaintiffs presented no meaningful evidence on market dilution” and had “fail[ed] to present meaningful evidence on the effect of training LLMs like Llama with their books on the market for those books.” In this regard, he concluded that even if an AI model cannot “regurgitate” a plaintiff’s works or generate substantially similar content, “it can generate works that are similar enough (in subject matter or genre) that they will compete with the originals and thereby indirectly substitute for them.” He further opined that “[i]n cases involving uses like Meta’s, it seems like the plaintiffs will often win, at least where those cases have better-developed records on the market effects of the defendant’s use.”
These cases therefore reached similar conclusions, at least with respect to whether using copyrighted works to train a generative AI software service under the specific facts and evidence presented in the cases constitutes fair use, but by very different routes. And both cases found that generative AI software providers may be liable to the owners of copyrights, again for different reasons. In an echo of the landmark rulings against Napster at the start of the era of digital downloads and streaming music, Judge Alsup concluded that subsequently purchasing a legitimate copy of a copyrighted work will not cure liability for an existing infringing copy. And in some ways reminiscent of the Second Circuit’s rulings concerning Google Books, Judge Chhabria found that the new AI training use for existing copyrighted works was a fair use under the evidence presented, citing the lack of evidence that the copyright owners had suffered significant market harm from the emerging new technology – but forecasting that he might well rule the other way in cases in which “meaningful evidence on market dilution” is presented by copyright owners.
It should be noted that both these decisions are from the Northern District of California, and that decisions in other AI cases are pending there and in many other jurisdictions. In the only other case where there has been a ruling on a fair use defense involving copyrighted content used to train an AI system, a federal judge in Delaware rejected a defense of fair use. The court held that the defendant, which had trained its AI system using content owned by the plaintiff (including Westlaw headnotes), was liable for copyright infringement for having developed an AI tool for finding court opinions. Thomson Reuters Enter. Centre GmbH v. ROSS Intell. Inc., 765 F. Supp. 3d 382 (D. Del. Feb. 11, 2015), appeal docketed, Case No. 25-8018 (3d Cir. Apr. 14, 2025). The court in that case noted that “only non-generative AI” was at issue (instead of generative AI, in which the AI model would create new material as its output instead of just retrieving existing content), and the case was distinguished on that ground by the court in Bartz.
Of course, while these early decisions are important milestones, they are far from the final word on how copyright law will apply to AI. They do, however, set significant precedents and supply critical insights as to how courts may rule in future decisions. In addition, they provide useful guidance to companies that use or seek to deal with AI platforms.

