AI Output – Gold Mine Or Minefield?

Moses Singer Client Alert
Share this page:

With the continued development of artificial intelligence, and as machines become capable of generating remarkable and sophisticated output, a new era of legal complexity unfurls, sparking heightened anticipation for precedent-setting decisions that will address the outstanding questions on the rights and responsibilities surrounding AI.

The federal district court in Oakland recently handed down one of the first decisions calling into question the free creation, distribution, and use of AI output. While it is only a decision on an early motion in the case, plaintiffs’ allegations about the reproduction and distribution of their copyrighted computer programs by AI that is operated by Microsoft and OpenAI may lead to significant changes in AI operation. Doe 1 v. GitHub, Inc., 2023 WL 3449131 (N.D. Calif. May 11, 2023).

Factual Background

Plaintiffs are pseudonymous software developers seeking to represent a class of software developers.1 The two AI programs in question are coding tools called Codex and Copilot, a collaboration between GitHub and Open AI. Codex and Copilot help programmers write code. Copilot provides or fills in blocks of code. Codex, which is “integrated into” Copilot, converts natural language into code. Together, these programs suggest both code and entire functions to programmers.2

Codex and Copilot use machine learning. Their training data is described as “billions of lines of publicly available code” that is used to teach the machine learning model. This publicly available code includes code posted by programmers on a site called GitHub, which is owned by Microsoft. GitHub is a hosting service on which programmers can post the code of their software projects and collaborate on projects. Plaintiffs publicly posted code on GitHub and plaintiffs’ code was copied as training data for Codex and Copilot.

The fact that plaintiffs voluntarily made public their code on GitHub does not end the case. GitHub offers code posters a choice of 13 form licenses that the posters can use to regulate others’ use of their posted code. Eleven of the form licenses “require that any derivative work or copy of the licensed work include attribution to the owner, inclusion of a copyright notice, and inclusion of the license terms.” Plaintiffs used licenses containing these terms.

Plaintiffs’ Claims

Plaintiffs sued GitHub, Microsoft, and OpenAI over Codex’s and Copilot’s AI output. Plaintiffs sued for wrongful alteration or removal of copyright management information, breach of the GitHub licenses, and a variety of other federal and state law claims.

The most important claims in the case relate to defendants’ use of plaintiffs’ code as AI training data for Codex and Copilot. The key allegation in the case is allegedly taken from GitHub’s internal research: “Copilot reproduces code from training data ‘about 1% of the time.’”3 2023 WL 3449131 at *6.

This allegation has grave implications for all generative AI (a type of artificial intelligence that can generate new forms of content, such as audio, code, images, text, simulations, and video), because, if true, a significant amount of training data is simply reproduced in AI output. This could expose Copilot – and all other AI – to copyright infringement claims by the owners of copyright in the training data.

Generative AI receives training data and responds to users’ prompts based on what it has learned from the training data. To get training data into the AI, it has to be copied. Copyright law, however, gives copyright owners the exclusive right to copy. Why isn’t it copyright infringement for an AI owner to input copyrighted works as training data? The likely answer is “fair use.”

“Fair use” is a defense to copyright infringement. If a use of a copyrighted work that would otherwise infringe the copyright owner’s exclusive rights is deemed to be a “fair use,” there is no liability for infringement. What makes a use “fair” is the subject of many judicial decisions dating back almost 200 years, as well as a section of the current U.S. Copyright Act. Each fair use decision depends on the very specific facts of that case; the general principles stated in statute and case law are mere guidelines. However, it can be said that a use that does not injure the copyright owner’s natural markets for the work and that significantly transforms the copyrighted work or puts it to a use different from the typical uses for the copyrighted work has a good chance of being considered a fair use. CAUTION: Do not conclude that any particular use is “fair” from this highly abbreviated description of fair use – consult knowledgeable counsel.

Google Books was held not to infringe copyright despite copying all of the published literature ever written. Why? Because Google did not simply channel copyrighted works to users. Google’s overall use was what is called in fair use cases “transformative”: its copying was for the purpose of facilitating access to the literature rather than giving out free copies of copyrighted works. Google’s use of copyrighted works was limited to exposing relatively short sections of the works to give users an idea of what the works were about.

Determining that the use of computer programs as training data is fair use does not answer the question whether Copilot’s or Codex’s output is a fair use. AI learns from its training data and generally outputs something new and different. If the output is not substantially similar to any of the inputted works, we don’t get to the issue of fair use. Copyright infringement requires substantial similarity to the copyrighted work. Even if a work is to some extent copied, the resulting product must be “substantially similar” to the copied work for the resulting product to constitute an infringement.

However, to the extent that the AI does nothing more than output training data, it is doing nothing more than copying and distributing copyrighted matter – basic copyright infringement. Outputting training data will expose the AI owner or operator to possible infringement liability.

Even if the AI does not output exact copies of copyrighted training data, it may still infringe copyright owners’ “derivative work” rights. Copyright owners have the exclusive right to “recast, transform or adapt” their own works. 17 U.S.C. §101 The resulting works are called “derivative works.” If AI creates derivative works and outputs them, this too exposes AI owners and operators to infringement claims.

Why isn’t all AI output an infringing derivative work of copyrighted works used as training data? Because, as noted above, there must be “substantial similarity” between the original work and the alleged derivative work. Using a work as the starting point for a new work is not necessarily infringement. It is only infringement if the copyist’s end product is substantially similar to the original work. A new work is substantially similar to the original work if an ordinary observer (such as a jury member) would think so.

The plaintiffs in the GitHub case, however, are not claiming copyright infringement. What then are they complaining about?

Copyright Management Information

The Digital Millennium Copyright Act of 1998, which amended the Copyright Act, protects “copyright management information” applied to or associated with copyrighted works. The most familiar copyright management information is the statutory copyright notice, e.g., © Moses & Singer LLP 2023. The DMCA forbids removal or alteration of any copyright management information if the remover knows or should know that the removal or alteration will induce, enable, facilitate, or conceal copyright infringement. 17 U.S.C. §1202(b)(1).

According to plaintiffs, defendants trained their AI programs to ignore or remove the copyright management information that plaintiffs had applied. Together with the fact that defendants knew that their AI sometimes reproduced training data as output, the court in the GitHub case held that these facts permitted the inference that defendants knew or had reasonable grounds to know that stripping out plaintiffs’ copyright management information risked inducing infringing uses of Copilot’s or Codex’s output.

The court also found that plaintiffs had sufficiently alleged a claim that defendants had breached the GitHub licenses that plaintiffs had selected when the plaintiffs uploaded code to GitHub. As noted above, those licenses required “attribution to the owner, inclusion of a copyright notice, and inclusion of the license terms.” Outputting plaintiffs’ code without such information may have breached those licenses.

Impact of the Decision

The GitHub decision addresses many more procedural and substantive points, but plaintiffs’ aforesaid claims and the court’s conclusions that they could give rise to liability for AI owners and operators raises possibly existential questions for at least some kinds of AI. Can AI be designed or trained not to simply output training data? What about the harder question of whether AI can be designed or trained to make sure that its output is not “substantially similar” to training data? Is a legislative solution required, akin to the qualified immunity of internet service providers from copyright infringement claims?

Of one thing we can be certain: There will be plenty of litigation on these and related questions.

 

1 The developers justify proceeding pseudonymously because they claim their attorneys have received many death threats for representing them in this case.

2 The facts reported in this article are taken from the court’s decision. For the purposes of a motion to dismiss, the court takes its facts exclusively from plaintiffs’ Complaint without determining whether they are true or false, so the “facts” recited in this article may or may not be true.

3 It is interesting that, despite this allegation, there is no copyright infringement claim, amounting to a claim of distribution of copyrighted material. Presumably, the rights reserved to GitHub in its terms and conditions immunized it from such infringement liability.