Fair use

RawGraph

Fair use is a limitation on the exclusive rights of copyright owners in United States law, codified at 17 U.S.C. 107. It permits unlicensed use of copyrighted material in certain circumstances, and the statute names criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, and research as illustrative purposes [1]. Whether a particular use qualifies is decided by weighing four statutory factors, and the statute adds that a work being unpublished does not by itself bar a finding of fair use, a sentence Congress inserted by amendment in 1992 [1]. The U.S. Copyright Office describes fair use as a judge-created doctrine dating back to the nineteenth century that was codified in the 1976 Copyright Act, and stresses that there is no formula fixing a safe percentage or word count [2].

Fair use is now the central legal question in litigation over how generative AI systems are built. Training a large language model or an image generator involves copying works at scale: acquiring them, assembling them into corpora, cleaning and tokenizing them, and reproducing them repeatedly during pre-training. Each of those steps implicates the reproduction right, and the Copyright Office has framed the key question as whether those acts of prima facie infringement can be excused as fair use [9]. Because the doctrine is fact-specific and balances four factors that can point in different directions, it has produced rulings that reach opposite conclusions on similar-sounding facts.

Section 107 is also close to unique. Most other jurisdictions have no open-ended fair use provision and address machine learning through narrower text and data mining (TDM) exceptions with conditions attached, so the same training run can be lawful in one country and actionable in another [9]. That asymmetry is one reason the regulation of AI and the copyright status of training data are debated together.

The statute and the four factors

Section 107 instructs courts to consider at least the following four factors, which the Supreme Court has said must be explored and weighed together in light of the purposes of copyright rather than reduced to bright-line rules [1][3].

FactorStatutory textWhat courts ask
One"the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes"Is the use commercial, and does it serve a further purpose or different character than the original?
Two"the nature of the copyrighted work"Is the work creative or factual, published or unpublished?
Three"the amount and substantiality of the portion used in relation to the copyrighted work as a whole"Was the taking reasonable in light of the purpose, and was the heart of the work taken?
Four"the effect of the use upon the potential market for or value of the copyrighted work"Does the use substitute for the original or harm markets the owner would develop or license?

Factors one and four carry the most weight in practice, a proposition courts trace to the Second Circuit's Google Books opinion, and the Supreme Court called the fourth "undoubtedly the single most important element of fair use" in Harper & Row v. Nation Enterprises (1985) [10]. Factor two, by contrast, "has rarely played a significant role in the determination of a fair use dispute" [10].

Statutory damages give the analysis its financial stakes. Under 17 U.S.C. 504(c), an infringer faces between $750 and $30,000 per work, rising to as much as $150,000 per work for willful infringement and falling to as little as $200 for innocent infringement [19]. Multiplied across a corpus of millions of books, those figures make a fair use ruling close to dispositive of a company's exposure.

Origins and the transformative-use turn

Modern fair use analysis turns on the idea of transformative use, which the Supreme Court adopted in Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569 (1994). Writing for the Court on 7 March 1994, Justice Souter framed the first-factor inquiry as whether the new work "merely 'supersede[s] the objects' of the original creation ... or instead adds something new, with a further purpose or different character, altering the first with new expression, meaning, or message; it asks, in other words, whether and to what extent the new work is 'transformative.'" Campbell also refused to treat commercial use as presumptively unfair, and reasoned that a parody and the original usually serve different market functions, so a parody is less likely to substitute for the work it targets [3].

Two Second Circuit decisions extended that reasoning to mass digitization. In Authors Guild, Inc. v. HathiTrust, 755 F.3d 87 (2d Cir. 2014), the court held that building a full-text searchable database from more than ten million digitized works was "quintessentially transformative," because "the result of a word search is different in purpose, character, expression, meaning, and message from the page (and the book) from which it is drawn," and that full-text search caused no harm to any existing or potential traditional market [5]. In Authors Guild, Inc. v. Google Inc., decided 16 October 2015, the court reached the same conclusion for Google Books, a project in which Google digitized more than twenty million books from the collections of research libraries. Copying entire texts was not dispositive because Google limited what users could see, and the snippet results were too brief and disjointed to substitute for buying the book [4].

The Supreme Court applied fair use to software in Google LLC v. Oracle America, Inc., 593 U.S. 1 (2021), decided 5 April 2021. Assuming for argument's sake that the material was copyrightable, Justice Breyer's opinion held that Google's copying of roughly 11,500 lines of declaring code from the Java SE API, only 0.4 percent of the 2.86 million lines in the API at issue, was fair use, because the copied material was a functional interface bound up with uncopyrightable ideas and because Android was not a market substitute for Java SE [6]. Ninth Circuit decisions in Sega Enterprises Ltd. v. Accolade, Inc. (1992) and Sony Computer Entertainment v. Connectix Corp. (2000) had already treated intermediate copying of code as transformative where it was necessary to reach unprotected functional elements [10].

Not every digitization case succeeds. In Hachette Book Group, Inc. v. Internet Archive, argued 28 June 2024 and decided 4 September 2024, the Second Circuit affirmed judgment against the Internet Archive over its scanning and lending of 127 books under a controlled digital lending model, holding that distributing full digital copies for free, even on a one-to-one owned-to-loaned ratio, was not fair use [8].

Warhol v. Goldsmith and the limits of transformation

Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508 (2023), decided 18 May 2023 by a vote of 7 to 2, narrowed how far transformative meaning alone can carry a defendant. Justice Sotomayor's majority opinion held that where an original work and a secondary use "share the same or highly similar purposes" and the secondary use is commercial, the first factor is likely to weigh against fair use unless some other justification for the copying exists [7]. The dispute concerned the foundation's 2016 licensing of Warhol's "Orange Prince" to Condé Nast for a special edition magazine commemorating Prince, the same use Lynn Goldsmith's photograph would have served. The Court was explicit about its narrow scope: only the first factor was before it, and it expressed no opinion on the creation, display, or sale of the original Prince Series works [7]. Justice Kagan dissented, joined by Chief Justice Roberts.

Warhol changed the vocabulary of AI litigation, since a defendant must now identify a further purpose or different character rather than assert that a model is a new kind of thing. Later courts have used the shared-purpose test as the pivot of the first-factor analysis [10][14].

Why AI training raises the question

The U.S. Copyright Office examined the problem in Copyright and Artificial Intelligence, Part 3: Generative AI Training, released in a pre-publication version in May 2025 [9]. The report separates the stages at which copying occurs: data collection and curation, training itself, retrieval-augmented generation, and outputs. Its conclusion declines to announce a rule. "Various uses of copyrighted works in AI training are likely to be transformative," it says, but the extent to which they are fair depends on what works were used, from what source, for what purpose, and with what controls on outputs. It then draws a boundary: "But making commercial use of vast troves of copyrighted works to produce expressive content that competes with them in existing markets, especially where this is accomplished through illegal access, goes beyond established fair use boundaries." On licensing, the Office concluded that voluntary markets were emerging unevenly and that government intervention would be premature, while flagging extended collective licensing as an option if gaps persist [9].

The report identifies the two market theories that have since divided the courts: lost licensing opportunities, meaning the market for selling works as training data, and market dilution, meaning the flood of AI-generated substitutes competing with the works the model learned from [9].

The 2025 district court rulings

Three district court decisions in 2025 supplied the first substantive American case law on fair use for AI training. None is binding beyond its own district, and all three are narrower than the headlines they generated.

CaseCourt and dateSystem at issueFair use outcome
Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc.D. Del., 11 February 2025 (Bibas, Circuit Judge, sitting by designation)Non-generative legal search tool built from Westlaw headnotesDefense rejected on summary judgment; factors one and four to Thomson Reuters [10]
Bartz v. Anthropic PBCN.D. Cal., 23 June 2025 (Alsup, J.)LLM training on purchased and pirated booksTraining copies and print-to-digital conversion fair use; retention of pirated library copies not fair use [12]
Kadrey v. Meta Platforms, Inc.N.D. Cal., 25 June 2025 (Chhabria, J.)LLM training on books from shadow librariesPartial summary judgment for Meta on this record only, because plaintiffs failed to prove market dilution [14]

Thomson Reuters v. Ross Intelligence

Ross Intelligence, denied a Westlaw license, bought roughly 25,000 "Bulk Memos" of legal questions that a contractor had written from Westlaw headnotes, and used them to train a legal search tool. Judge Stephanos Bibas revised his own 2023 opinion, granted partial summary judgment for Thomson Reuters on 2,243 headnotes, and rejected the fair use defense [10].

The first factor went to Thomson Reuters because Ross's use was commercial and lacked a further purpose or different character from Westlaw's. Bibas distinguished the intermediate-copying software cases on the ground that they involved computer code and depended on copying being necessary to reach underlying ideas, neither of which applied. Factors two and three went to Ross, since headnotes are less creative than a novel and no headnote reached the public. Factor four went to Thomson Reuters: Ross meant to build a market substitute, and the effect on a potential market for AI training data was enough even though Thomson Reuters had not itself sold headnotes as training data [10]. Bibas added a warning that is often dropped when the case is cited: "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today" [10].

The Third Circuit accepted an interlocutory appeal under 28 U.S.C. 1292(b), granting the petition (Misc. No. 25-8018) by order entered 17 June 2025 and docketing the merits appeal as No. 25-2153 on 24 June 2025. The appeal was argued on 11 June 2026 before Judges Restrepo, Montgomery-Reeves, and Bove, and no opinion had issued as of the docket entries through July 2026 [11]. It would be the first federal appellate ruling on fair use for AI training data.

Bartz v. Anthropic

Three authors, Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson, joined by two corporate entities that two of them set up to market their works, sued Anthropic over books used to train the models behind Claude. Judge William Alsup's order of 23 June 2025 found that Anthropic had downloaded pirated book collections including Books3 (196,640 books, obtained in early 2021), at least five million copies from Library Genesis in June 2021, and at least two million from the Pirate Library Mirror in July 2022, amounting to more than seven million pirated copies. It separately bought millions of print books, stripped the bindings, scanned them, and discarded the paper [12].

Alsup split the uses. Copies made to train the models were fair use: "the purpose and character of using copyrighted works to train LLMs to generate new text was quintessentially transformative," and the technology "was among the most transformative many of us will see in our lifetimes." Converting purchased print books into digital copies was also fair use, since no new copies were added and the originals were destroyed. But downloading pirated copies to build a permanent general-purpose library was not: "Every factor points against fair use," because a separate justification was required for each use and Anthropic offered none beyond cost and convenience [12].

The order was narrower than it appeared. The authors never alleged that any infringing output reached users, so outputs were not adjudicated, and the court expressly declined to bless copies made from the library for purposes other than training [12].

Kadrey v. Meta

Two days later, Judge Vince Chhabria granted partial summary judgment to Meta Platforms on the fair use defense to claims by thirteen authors over books downloaded from shadow libraries and used to train Llama, the model family built by Meta AI [14]. The reasoning cut against the result. Chhabria wrote that in most cases the answer to whether such copying is illegal "will likely be yes," because generative models can flood the market with substitutes and so undermine the incentive to create. He criticized the Bartz analogy between model training and teaching children to write as an inapt basis for discounting the fourth factor [14].

Meta won because of the record, not the law. The plaintiffs argued that Llama could reproduce snippets and that they had lost an AI-training licensing market; the court rejected both. The potentially winning argument, market dilution, was barely developed and supported by no evidence, so the fourth factor could only favor Meta. Chhabria set out the limits himself: the case was not a class action, and "this ruling does not stand for the proposition that Meta's use of copyrighted materials to train its language models is lawful. It stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record in support of the right one" [14]. The ruling also left untouched a separate claim that Meta unlawfully distributed the books by reuploading them while torrenting, on which neither side had moved for summary judgment [14].

The Anthropic settlement

Because the pirated-library holding survived, Bartz proceeded toward a trial on those copies and on damages, including willfulness [12]. A class was certified covering copyright owners of books in the LibGen and PiLiMi collections that Anthropic had downloaded, and the parties then settled for a non-reversionary fund of $1.5 billion. Judge Alsup granted preliminary approval at a hearing on 25 September 2025 and issued a memorandum opinion explaining that decision on 17 October 2025 [13].

Judge Araceli Martínez-Olguín granted final approval on 20 July 2026 after a hearing on 14 May 2026, entered judgment, and dismissed the case with prejudice [13]. The order records the mechanics: notice was sent to 594,945 potential class members, the Works List contained 482,460 works, and 440,490 of them (91.3 percent) had been claimed as of 16 April 2026. The estimated payment is about $3,000 per work, which the court noted is four times the $750 statutory minimum. Class counsel had asked for 12.5 percent of the fund and were awarded $101,561,111 in fees, nearly 6.8 percent, and the three class representatives received $15,000 each rather than the $50,000 requested [13].

The settlement resolved the piracy claims without any court deciding whether an appellate panel would agree with Alsup that training itself is fair use.

Cases still pending

The most prominent American case remains undecided. The New York Times sued Microsoft and OpenAI in the Southern District of New York on 27 December 2023 [16]. The Judicial Panel on Multidistrict Litigation transferred the OpenAI copyright cases to that court on 3 April 2025, where they were centralized as In re: OpenAI, Inc., Copyright Infringement Litigation, MDL No. 3143, before Judge Sidney H. Stein with Magistrate Judge Ona T. Wang handling pretrial matters [15]. Docket activity through late July 2026 remained pretrial, with no ruling on fair use [15][16]. Related suits consolidated into that proceeding include Authors Guild v. OpenAI.

Getty Images took a different route in its suit against Stability AI. It filed a notice of voluntary dismissal in its Delaware case against Stability on 14 August 2025, and that docket closed four days later. The same day it filed the notice, Getty opened a new action in the Northern District of California, which was still in discovery before Judge Trina L. Thompson in July 2026 [18].

Disputes over model training and outputs have been brought across other media as well, including music publishers' claims in Concord Music Group v. Anthropic, studio claims in Disney and Universal v. Midjourney, recording-industry claims against Suno covered at RIAA v. Suno, and news-publisher claims in the Perplexity copyright lawsuits. The general subject is covered at AI and copyright.

Fair use outside the United States

Most jurisdictions have no open-ended fair use provision, and instead rely on enumerated exceptions [9].

JurisdictionProvisionKey conditions
European UnionDirective (EU) 2019/790, Articles 3 and 4Article 3 covers research organisations and cultural heritage institutions for scientific research; Article 4 covers any actor for any purpose but requires lawful access and respects rightsholder opt-outs expressed in an appropriate manner [9]
European UnionRegulation (EU) 2024/1689 (AI Act), Article 53 and Recital 105General-purpose model providers must have a policy to comply with Union copyright law and to identify and respect Article 4 opt-outs [9]
JapanCopyright Act, Article 30-4Permits use for data analysis and AI development where the purpose is not to personally enjoy the expression, and excludes uses that would unreasonably prejudice the copyright owner [9]
SingaporeCopyright Act 2021, section 244Requires lawful access and limits copies to computational data analysis [9]
United KingdomCopyright, Designs and Patents Act 1988, section 29AComputational analysis for non-commercial research only, with lawful access required [9]
IsraelMinistry of Justice opinion, 18 December 2022Fair use provision modeled on section 107; the ministry concluded machine learning uses are in most but not all cases fair [9]

The practical difference shows in Getty Images (US) Inc v Stability AI Ltd 2025 EWHC 2863 (Ch), decided by Mrs Justice Joanna Smith on 4 November 2025. Getty abandoned its training and development claim at trial for want of evidence that training happened in the United Kingdom, and abandoned its output and database-right claims. What remained was secondary infringement and trade marks. The judge dismissed the secondary infringement claim on the ground that a model such as Stable Diffusion, which does not store or reproduce the copyright works, is not an "infringing copy" under sections 22 and 23 of the 1988 Act. Getty succeeded only in part on trade mark claims about generated watermarks, and the judgment summarizes the result plainly: "In summary, although Getty Images succeed (in part) in their Trade Mark Infringement Claim, my findings are both historic and extremely limited in scope. The Secondary Infringement Claim fails" [17].

Unsettled questions

Several issues remain genuinely open after the 2025 rulings.

The source of the training data may matter more than the training. Alsup held that downloading pirated copies to build a permanent library was not fair use even though training on the same books was, and the Copyright Office concluded that the knowing use of a dataset made up of pirated or illegally accessed works "should weigh against fair use without being determinative" [9][12]. That distinction has pushed developers toward licensed corpora, purchased copies, and provenance work of the kind tracked by the Data Provenance Initiative, and away from undocumented web scraping and shadow-library dumps such as the Books3 collection distributed inside The Pile, which the Copyright Office describes as 196,640 books sourced from an unauthorized BitTorrent tracker [9].

The fourth factor has no settled theory. Chhabria's market dilution framing and Bibas's potential-licensing-market framing identify different harms, and Alsup, reasoning by analogy to the "explosion of competing works" that would follow from training schoolchildren to write well, called that effect "not the kind of competitive or creative displacement that concerns the Copyright Act" [10][12][14]. Which theory prevails will decide most future cases, since factor four is where transformative uses are usually lost or won.

Inputs and outputs are still separate questions. Bartz reserved outputs entirely because the plaintiffs did not allege infringing outputs, and Alsup noted that if the outputs seen by users had been infringing the authors would have had a different case [12]. The Copyright Office likewise treats retrieval-augmented generation and outputs as distinct acts from training [9]. In Kadrey, where the plaintiffs did argue that the model reproduced their text, the court found that Llama was "not capable of generating enough text from the plaintiffs' books to matter" [14].

Finally, no federal appellate court has ruled on fair use for AI training. Until the Third Circuit decides the Ross appeal or a circuit reaches a generative AI case, the governing American authority is a set of district court opinions that reason differently from one another, none of them binding outside its own court, plus a settlement that avoided the question [10][11][12][14].

See also

References

  1. ^17 U.S.C. 107, Limitations on exclusive rights: Fair use. Legal Information Institute, Cornell Law School. law.cornell.edu/...107
  2. ^U.S. Copyright Office, "More Information on Fair Use." copyright.gov/fair-use
  3. ^Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569 (1994), full opinion. Legal Information Institute, Cornell Law School. law.cornell.edu/...569
  4. ^Authors Guild, Inc. v. Google Inc., No. 13-4829-cv (2d Cir. Oct. 16, 2015). U.S. Copyright Office Fair Use Index case summary. copyright.gov/...authorsguild-google-2dcir2015.pdf
  5. ^Authors Guild, Inc. v. HathiTrust, 755 F.3d 87 (2d Cir. 2014). U.S. Copyright Office Fair Use Index case summary. copyright.gov/...orsguild-hathitrust-2dcir2014.pdf
  6. ^Google LLC v. Oracle America, Inc., 593 U.S. 1 (2021), No. 18-956, slip opinion. Supreme Court of the United States. supremecourt.gov/...18-956_d18f.pdf
  7. ^Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508 (2023), No. 21-869, slip opinion. Supreme Court of the United States. supremecourt.gov/...21-869_87ad.pdf
  8. ^Hachette Book Group, Inc. v. Internet Archive, No. 23-1260 (2d Cir. Sept. 4, 2024), slip opinion. storage.courtlistener.com/..._internet_archive.pdf
  9. ^U.S. Copyright Office, Copyright and Artificial Intelligence, Part 3: Generative AI Training, Pre-Publication Version, May 2025. copyright.gov/...eport-Pre-Publication-Version.pdf
  10. ^Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., No. 1:20-cv-613-SB, Memorandum Opinion (D. Del. Feb. 11, 2025), Dkt. 770. storage.courtlistener.com/...s.ded.72109.770.0.pdf
  11. ^Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., No. 25-2153 (3d Cir.), docket. courtlistener.com/...-gmbh-v-ross-intelligence-inc
  12. ^Bartz v. Anthropic PBC, No. C 24-05417 WHA, Order on Fair Use (N.D. Cal. June 23, 2025), Dkt. 231. storage.courtlistener.com/...nd.434709.231.0_5.pdf
  13. ^Bartz v. Anthropic PBC, No. 3:24-cv-05417-AMO, Order Granting Final Approval of Class Action Settlement; Granting in Part Motion for Attorney's Fees, Reimbursement of Expenses, and Plaintiff Service Awards; Judgment (N.D. Cal. July 20, 2026), Dkt. 680, and Memorandum Opinion on Preliminary Approval of Class Action Settlement (N.D. Cal. Oct. 17, 2025), Dkt. 437. storage.courtlistener.com/...nd.434709.680.0_5.pdf and storage.courtlistener.com/...cand.434709.437.0.pdf
  14. ^Kadrey v. Meta Platforms, Inc., No. 23-cv-03417-VC, Order Denying the Plaintiffs' Motion for Partial Summary Judgment and Granting Meta's Cross-Motion for Partial Summary Judgment (N.D. Cal. June 25, 2025), Dkt. 598. storage.courtlistener.com/...cand.415175.598.0.pdf
  15. ^In re: OpenAI, Inc., Copyright Infringement Litigation, No. 1:25-md-03143 (S.D.N.Y.), docket. courtlistener.com/...right-infringement-litigation
  16. ^The New York Times Company v. Microsoft Corporation, No. 1:23-cv-11195 (S.D.N.Y., filed Dec. 27, 2023), docket. courtlistener.com/...mpany-v-microsoft-corporation
  17. ^Getty Images (US) Inc & Ors v Stability AI Ltd 2025 EWHC 2863 (Ch), judgment of Mrs Justice Joanna Smith, 4 November 2025. Find Case Law, The National Archives. caselaw.nationalarchives.gov.uk/...2863
  18. ^Getty Images (US), Inc. v. Stability AI, Inc., No. 1:23-cv-00135 (D. Del.), docket, and Getty Images (US), Inc. v. Stability AI, Ltd., No. 3:25-cv-06891 (N.D. Cal.), docket. courtlistener.com/...ges-us-inc-v-stability-ai-inc and courtlistener.com/...ges-us-inc-v-stability-ai-ltd
  19. ^17 U.S.C. 504, Remedies for infringement: Damages and profits. Legal Information Institute, Cornell Law School. law.cornell.edu/...504

Improve this article

Add missing citations, update stale details, or suggest a clearer explanation. Every suggestion is reviewed for sourcing before it goes live.

v1 · 4,175 words · full history

Fact-checks are independent of edits: a reviewer re-verifies the article against its sources and stamps the date. How we verify

Research and drafting on this wiki are AI-assisted, under named human editorial standards. How AI is used here

Reviewer note: Independent adversarial fact-check at creation (wanted175 campaign, 2026-07-24): every claim verified against primary sources by a dedicated verification agent; corrections applied before publication.

Cite this page: AI Wiki. "Fair use." aiwiki.ai, updated 24 Jul 2026, fact-checked 24 Jul 2026. CC BY 4.0. https://aiwiki.ai/wiki/fair_use

Suggest edit