A Fair Use Holding for the Books
The internet has become an indispensable tool for accessing and exchanging information, and although copyright law has adapted somewhat to the internet age, copyright law's core principles remain relevant and can restrict what content can be shared online.
The recent decision in Hachette v. Internet Archive (March 24, 2023) was a major triumph for book publishers who sought to maintain the status quo copyright protections for their online publications. The plaintiffs in the action, four prominent book publishers, sued defendant Internet Archive for copyright infringement based on its scanning print copies of plaintiffs’ 127 copyrighted books and lending the digital copies through Internet Archive’s website without plaintiffs’ permission.
In this article, we explore how the court applied the fair use analysis to reach its decision, and how the result may impact future cases involving AI systems and copyrighted content.
Hachette v. Internet Archive
Internet Archive is a non-profit that offers free online access to texts, audio, moving images, software, and other cultural artifacts, including its most frequently consulted resource: archived copies that Internet Archive makes available of material this is publicly available on internet websites.
In response to the complaint, Internet Archive advanced a defense of fair use, contending that it was not liable for copyright infringement for the copying and uses it made. On cross-motions for summary judgment, District Judge Koeltl in the Southern District of New York granted the publishers summary judgment and denied Internet Archive’s motion.
In considering Internet Archive’s fair-use defense, Judge Koeltl analyzed the four statutory factors, namely: (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work, 17 U.S.C. § 107, to determine whether a particular use of a work is a fair use.
In recent years, several other cases have “tested the boundaries of fair use” for online uses of published works. Significantly, the Second Circuit Court of Appeals has ruled that fair use protects specific uses that are sufficiently transformative. In Authors Guild v. Google, Inc., 804 F.3d 202 (2d Cir. 2015), the court determined that Google Books’ practice of scanning copyrighted books to create a database that provided the ability to search for the prevalence of specific words in published works, which allowed the sight-impaired to have their computer read text of published books aloud, and for everyone to use a free “snippet view” search function whereby readers could view a few lines of text containing searched-for terms, were all sufficiently transformative. See also Authors Guild, Inc. v. HathiTrust, 755 F.3d 87 (2d Cir. 2014) (holding that defendant’s wholesale scanning of works to create a full-text searchable database was a quintessentially transformative use).
Internet Archive’s use, however, was not afforded such deferential treatment. Instead, the court summarily found the opposite, that Internet Archive’s creation, maintenance and free lending of its copies was “squarely beyond fair use.”
The court found that each fair use factor favored the publishers, concluding that fair use does not allow for the “mass reproduction and distribution of complete copyrighted works in a way that does not transform those works and that creates directly competing substitutes for the originals.” As this sort of mass reproduction and lending of copies was essentially Internet Archive’s business model, the court found that, as a matter of law, its defense of fair use failed.
Implications for Artificial Intelligence and Copyright Law
This case is significant in that it underscores the strength of the Copyright Act and the important protections afforded to authors and copyright holders, despite attempts to blur those lines under the guise of technological advancements. It also emphasizes the courts’ focus on the impact of the actual use made by those who make unauthorized copies. Internet Archive, however, fervently believes that this holding hinders access to information in the digital age, harming all readers. As such, Internet Archive claims that it intends to appeal, making this a matter to watch moving forward.
This case may presage others, which may similarly turn away artificial intelligence systems which create and use vast databases of content, many of which are “scraped” from the internet and copied to servers, along with the metadata about the content and potentially copyright management information that accompanies the content.
The court in this case viewed Internet Archive’s ultimate uses as fatal to Internet Archive’s argument that it was entitled to create and maintain copies of the plaintiffs’ published works on its servers. As a result, this case may be seen as a retreat, albeit in different factual circumstances, from the Second Circuit’s conclusions that all the uses in issue in cases about Google’s “Book Project” were protected by fair use, and that such uses excused any liability for creating infringing copies. Specifically, this court is now ruling that the intended uses by Internet Archive determine whether the copying and storing of copies constitute infringement.
Conclusion
Given that current artificial intelligence models and tools can generate images, writings and even music that may be used to replace the original works, Judge Koeltl’s logic employed in this case about Internet Archive may be informative of how courts will soon treat cases about mass copying of works available on the internet in future cases about AI.
