The DOJ’s AI Fair Use Brief Is Correct, But From A DOJ That Has No Credibility

from the worst-doj-you-know-just-made-a-great-point dept

I am never asking the monkey’s paw for a federal government that defends fair use again.

For decades on Techdirt, I’ve talked about the absolute necessity of strong fair use, and have been disappointed over and over again at how little the federal government has fought for fair use. For decades, the rare times when the federal government has weighed in on copyright cases, it’s often been in support of copyright maximalism. So, when the government finally makes a strong stand for fair use… it’s this White House? With this DOJ? And, in defense of a giant centralized AI provider?

Sigh.

So, look, you’re right to be skeptical about the DOJ weighing in regarding the big OpenAI copyright case filed by the NY Times. You’re right to be skeptical about the DOJ’s motives. You’re right to be skeptical about OpenAI’s motives.

But the DOJ is correct in its legal analysis. AI training absolutely is fair use, and a ruling the other way would blow a hole in fair use protections that have nothing to do with AI, including everything from search engines to book scanning to reverse engineering to (most importantly) text and data mining for research. Also, the NY Times’ case against OpenAI is incredibly weak, and (as I’ve discussed) involves bizarre theories of copyright that would put all sorts of companies (including the NY Times itself!) at real risk of liability for doing basic reporting. So, having the federal government weigh in and make good points could actually be helpful.

However, this is also why it’s so frustrating that the Trump DOJ and Todd Blanche (and Pam Bondi before him) completely burnt through the presumption of regularity by repeatedly filing bullshit briefs in bullshit cases. Because when they actually file a reasonable thing in an important case, judges are still going to be quite skeptical.

But, in this case, the DOJ is correct.

We’ve talked about some of this already, regarding the big case against Anthropic where Judge William Alsup found that AI training was easily fair use.

Much of the coverage of the DOJ’s filing focuses on the “national security” claims which the filing itself spends much of the opening on — the argument that if we don’t let American companies train on everything, China will eat our lunch in AI. That argument is pretty weak, and it’s also unnecessary. The fair use analysis stands on its own without any appeal to beating our adversaries. But much more interesting (and correct) to me is the argument that if training is not fair use, then only a few giant, wealthy companies can afford to create AI tools, and we’d just be recreating the broken “big tech” structures of the last decade, rather than enabling more decentralized, more widely competitive tools:

An erroneous fair use ruling would hamper competition in the market for LLMs, because only the largest technology companies might have the capital necessary to pay licensing fees. And such licensing fees would disproportionately benefit legacy media outlets due to the sheer volume of their written publications. By contrast, if not hindered by a strained understanding of copyright law, LLMs can and should help level the playing field between mainstream and independent publishers, for several reasons. Authors with limited resources can use LLMs to compete (by, for example, using an LLM to generate an image to accompany an article—which otherwise might require a photographer or license). And LLMs can direct users to dissenting sources that offer contrary information or perspectives. It is not in the public’s interest for the largest technology companies to have an oligopoly on LLM training due to licensing entry barriers that function primarily as large subsidies for old mainstream media companies.

To me this is the whole ballgame, and part of what makes it so frustrating that many people insist training can’t be fair use. They often think the end result is somehow punishing the “big” AI companies, but the reverse is true. A finding against fair use locks in the biggest AI companies, and wipes out everyone else, especially decentralized open weight models that actually empower end-users without enabling tech giants.

The DOJ also cites all the right precedents (including some that previous administrations were not happy about) on a point that gets mangled constantly: fair use isn’t just a defense you raise after infringing. It means there was no infringement in the first place.

A copyright owner’s exclusive rights are thus subject to various exceptions and limitations. For example, “copyright assures authors the right to their original expression, but encourages others to build freely upon the ideas and information conveyed by a work.” Feist, 499 U.S. at 349- 50. “This principle, known as the idea/expression or fact/expression dichotomy, applies to all works of authorship.” Id. at 350.

Relatedly, and as most relevant here, the “fair use” doctrine provides that certain secondary uses of a copyrighted work are “not an infringement.” 17 U.S.C. § 107. Although fair use originated as “judge-made,” Congress subsequently codified it. Campbell, 510 U.S. at 576. The statute continues a common-law tradition, which recognized that certain amounts and types of copying must occur to promote the purposes of the Intellectual Property Clause. See id. at 575 (citing U.S. CONST. art. I, § 8, cl.8). Fair use is an “equitable rule of reason that permits courts to avoid rigid application of the copyright statute when, on occasion, it would stifle the very creativity which that law is designed to foster.” Google LLC v. Oracle Am., Inc., 593 U.S. 1, 18 (2021).

It also notes (correctly, though contrary to the belief of many non-copyright lawyers) that fair use was written to be broad and flexible, not just based on a specific set of categories, or if you meet specific rules:

The preamble of section 107 specifically recites six uses likely to result in a finding of fair use: “criticism, comment, news reporting, teaching . . ., scholarship, or research.” But determining a “fair use” is “not to be simplified with bright-line rules, for the statute, like the doctrine it recognizes, calls for case-by-case analysis.” Campbell, 510 U.S. at 576-77. As such, the statutory list is not exhaustive. See 17 U.S.C. § 107 (identifying purposes “such as” the listed set). The statute’s legislative history confirms the same. See Harper & Row Publishers v. Nation Enters., 471 U.S. 539, 562 (1985); Pac. & S. Co., Inc. v. Duncan, 744 F.2d 1490, 1496 (11th Cir. 1984); Cambridge Univ. Press v. Becker, 863 F. Supp. 2d 1190, 1225 (N.D. Ga. 2012). The House Report for section 107 indicates Congress’s intent for a flexible inquiry that can adapt to new technologies and scenarios:

[T]here is no disposition to freeze the doctrine in the statute, especially during a period of rapid technological change. Beyond a very broad statutory explanation of what fair use is and some of the criteria applicable to it, the courts must be free to adapt the doctrine to particular situations on a case-by-case basis.

H.R. Rep. No. 94–1476, 94th Cong., 2d Sess. 66 (1976); see also id. at 65 (“[S]ince the doctrine is an equitable rule of reason, no generally applicable definition is possible, and each case raising the question must be decided on its own facts.”).

The DOJ agrees with Alsup’s analysis that training is quite clearly fair use as transformative.

The first statutory fair-use factor, the “purpose and character of the use,” requires consideration of “whether the new work merely ‘supersede[s] the objects’ of the original creation (‘supplanting’ the original), or instead adds something new, with a further purpose or different character, altering the first with new expression, meaning, or message.” Campbell, 510 U.S. at 578-79 (quoting Folsom v. Marsh, 9 F. Cas. 342, 348 (C.C.D. Mass. 1841) (No. 4,4901) (Story, J.) and Harper & Row, 471 U.S. at 562). The Supreme Court has described the latter type of use as “transformative.” Id. “[T]ransformative uses tend to favor a fair use finding because a transformative use is one that communicates something new and different from the original or expands its utility, thus serving copyright’s overall objective of contributing to public knowledge.” Authors Guild, 804 F.3d at 214.

The copying of protected text articles as part of training an LLM is a use of a different kind or character that is “transformative—spectacularly so.” Bartz v. Anthropic PBC, 787 F. Supp. 3d 1007, 1021 (N.D. Cal. 2025). The New York Times alleges that OpenAI’s training results in a model that “predict[s] words that are likely to follow a given string of text based on the potentially billions of examples used to train” OpenAI’s LLMs, such that the LLMs can subsequently generate original responses to a wide range of user inputs. Microsoft Corp., No. 1:23-cv-11195-SHS-OTW, ECF 1677 ¶ 75 (Aug. 21, 2026). An OpenAI LLM thus uses the copyrighted work not to duplicate the work’s expressive content, but as part of a process to learn and act on statistical patterns in written text, including vocabulary, syntax, and knowledge. The purpose of the copying (to build an intelligent, interactive model) differs in kind from the purpose of the copied work (to use language to directly entertain or educate a reading audience). This use of text-based works to create an LLM engine for “innovative tools” that can “edit an email . . . , translate an excerpt from or into a foreign language, write a skit based on a hypothetical scenario, or do any number of other tasks” is undoubtedly “highly transformative.” Kadrey v. Meta Platforms, Inc., 788 F. Supp. 3d 1026, 1044 (N.D. Cal. 2025).

Notably, the whole concept of “transformative use” being so central to fair use is generally traced back to Judge Pierre Leval’s wonderful 1990 Law Review article “Toward a Fair Use Standard.” At the time he wrote that, he was a federal judge in the Southern District of NY, where this case is being heard (side note: it’s ridiculous that Leval’s “Toward a Fair Use Standard” article seems to mainly only be available behind JSTOR’s paywall…. if ever there were an article that should be available freely…).

Beyond the transformativeness, the DOJ leans heavily (again, correctly) on the other big factor that shows up in fair use cases: the impact on the market. As we said when the NY Times first floated this lawsuit, no one is replacing the NY Times with ChatGPT. They serve very different purposes. And the DOJ filing emphasizes this:

Using a copyrighted work to train an LLM—without more—generally does not result in this sort of substitution because it does not “reveal[]” a significant amount of original “authorial expression.” Authors Guild, 804 F.3d at 224; see also, e.g., Bartz, 787 F. Supp. 3d at 1031 (“[T]raining LLMs did not result in any exact copies nor even infringing knockoffs of their works being provided to the public.”). In fact, training does not reveal anything to the public at all—it simply creates a copy of a protected work in order to teach an LLM to recognize relationships between data and adapt to new information. The potential for future outputs that might cause market harm is simply not relevant to evaluating an LLM training use under the required use-byuse analysis.

To be clear, even LLM outputs that compete with text articles—without reproducing or substantially resembling protected aspects of text articles—would not be substantially similar to, or substitutes for, copyrighted works in the relevant sense. When outputs “copy no protected elements of the original work, much less significant portions,” they cannot cause the relevant form of market harm just because they happen to be “in the same genre or category of works” as the original, given that “a genre is an uncopyrightable idea or method of expression.” Edward Lee, Copyright Dilution Under Constitutional Scrutiny, 25 Chi.-Kent J. Intell. Prop. 1, 6 (2026) (citing Peters v. West, 692 F.3d 629, 636 (7th Cir. 2012) (“[N]o poet can claim copyright protection in the form of a sonnet or a limerick.”)); accord Abdin v. CBS Broad. Inc., 971 F.3d 57, 70 (2d Cir. 2020) (no infringement where “an independent comparison of the works reveals that there is no substantial similarity between the protectible features of [the original]” and the secondary use). Whether a particular output or category of outputs is substantially similar to the copyrighted work, and whether any substantially similar reproduction might be a significantly competing substitute, are distinct questions involving distinct “challenged use[s]” (and potentially additional legal questions). Harper & Row, 471 U.S. at 568. But outputs lacking in substantial similarity cannot cause the sort of market harm that is cognizable in the fair-use analysis.

I’m also happy to see the DOJ make a point that usually gets ignored in these cases: creative people learn by copying. It is natural. Creative people imitate others until they find their own voice. A ruling that training isn’t fair use would turn that whole creative learning trajectory into infringement:

This type of logic would have problematic implications for copyright law generally. To illustrate, when she was a teenager, Joan Didion “would type out” Ernest Hemingway’s “stories to learn how the sentences worked,” and as a result she considered him the greatest influence on her writing. See Linda Kuehl, Joan Didion, the Art of Fiction No. 71, The Paris Review (Issue 74, Fall-Winter 1978).18 By the Kadrey court’s logic, Didion should have incurred liability to Hemingway every time she published a piece, because the process by which she trained herself and the process by which she produced works was all one use, and her works competed with those of other authors in the market for literature. But “to make anyone pay specifically for the use of a book . . . each time they later draw upon it when writing new things in new ways would be unthinkable.” Bartz, 787 F. Supp. 3d at 1021.

I’m also glad to see the DOJ call out the simple fact that, for all of the NY Times’ whining about how awful it is that OpenAI trained on the NY Times (and basically every other published work out there), NY Times reporters regularly rely on LLM tools themselves. And… that it’s helping smaller media providers level up to compete with a media company as large and full of resources as the NY Times.

LLMs can inspire or help someone to write a story, compose a song, write a movie script, or produce any other kind of art. Indeed, authors at the New York Times itself are using LLMs to help them “conceptualize and edit” articles.19 Independent and start-up publications, as well as ordinary people, can too. For example, an independent writer used an LLM and his background as a physics teacher to offer a contrarian perspective about data center water usage and critique the New York Times.

This is a good filing, and it sucks that this DOJ is so untrustworthy that the court may discount it.

And while it’s easy to claim that this was just done because of how the tech oligarchs have lined up behind Donald Trump, there are some suggestions that this is not the case here. This filing looks like the work of a few lawyers at the DOJ who actually understand copyright law — which, these days, is its own kind of shocking. And the best sign of this is that Mike Davis, the MAGA whisperer who appears quite gleeful about how if you pay him, he’ll get your antitrust case to turn out the way you want, is absolutely freaking out about this filing, and went public with his demand that the DOJ withdraw the filing in a Fox News op-ed calling it “the art of the steal” — a phrase that reveals he has no idea what fair use is, since a use that isn’t infringement isn’t theft. (Also necessary: a reminder that Mike Davis became anti-tech only after big tech companies refused to hire him).

The court may decide to ignore it, but the DOJ’s filing is absolutely correct on the issue of fair use. It’s just too bad it’s coming from a DOJ that has spent every last bit of credibility it had on cases that deserved none of it.

Filed Under: , , , , ,
Companies: ny times, openai

Rate this comment as insightful
Rate this comment as funny
You have rated this comment as insightful
You have rated this comment as funny
Flag this comment as abusive/trolling/spam
You have flagged this comment
The first word has already been claimed
The last word has already been claimed
Insightful Lightbulb icon Funny Laughing icon Abusive/trolling/spam Flag icon Insightful badge Lightbulb icon Funny badge Laughing icon Comments icon

Comments on “The DOJ’s AI Fair Use Brief Is Correct, But From A DOJ That Has No Credibility”

Subscribe: RSS Leave a comment
28 Comments
Anonymous Coward says:

I don’t really give a damn how correct it may be in principle; the practical upshot is that the rules are different for the little geese and the big AI gander, and that’s unacceptable. AI companies should get the pointy end of copyright maximalism until it has been durably lifted from the rest of us.

Fair use for AI companies but not for thee is intolerable bullshit and must not be allowed to stand.

TLDR: Fuck em.

Anonymous Coward says:

Re:

IF the randomness on an ML model could be turned to 0 AND the model used open weights AND the government/people weren’t subsidizing the resources needed to build and run the models, I’d have no issues with a bifurcated fair use doctrine in practice (not in principle) because I could easily claim that any content I’ve got that someone’s denying my fair use right to use is actually AI generated and therefore completely free of copyright restrictions.

The problem now is that so many different things are broken/abused that applying copyright and the fair use doctrine correctly is near to impossible, and if done, would expose many other gaps where laws, licenses and conventions developed over the past 60 years got things completely wrong.

Anonymous Coward says:

“Much of the coverage of the DOJ’s filing focuses on the “national security” claims which the filing itself spends much of the opening on — the argument that if we don’t let American companies train on everything, China will eat our lunch in AI. That argument is pretty weak, and it’s also unnecessary.”

It’s also wrong, because China is going to eat our lunch in AI no matter what the US-based AI companies train on, because the Chinese are serious — they’re doing the hard, meticulous, painstaking work necessary to build high-quality models. Meanwhile, the US companies are falling all over themselves to rush slop out the door so that they can make another deal and get another infusion of VC money.

A tell-tale sign of that is their proclivity to feed anything and everything into training regardless of provenance, accuracy, authenticity, or anything else — it’s a sign of desperation, they’re in too big a hurry to bother with input curation. (Imagine handing medical students a special issue of the The New England Journal of Medicine on vaccine efficacy and J. Random Blogger’s wacked-out conspiracy theory on vaccines and reptilian aliens, and expecting them to give both equal credence.)

TL;DR: the US has already lost the AI race. The outcome of this case won’t change that.

KevinQ says:

Fair Use

Look, I went to law school and studied intellectual property law, including specific classes on IT law and copyright, and I have never, never understood the “transformative” prong of the fair use test. If “I heavily transformed your work” weighs towards fair use, then how come “I transformed your book into a movie” is not fair use? In this case, how come “I transformed your book into computer instructions” not just an adaptation? Can somebody please explain it in small words a lawyer would understand?

Anonymous Coward says:

Re:

This is not really an answer to your (interesting) question; it’s something that I hope will clarify the question.

My remarks are rooted in something that Grady Booch wrote about LLMs, something I think really sums it up for people who haven’t done the math: “Such architectures are inherently just next word predictors and any correlation their output has with truth is only due to statistical coincidence.”

With that in mind, consider: what would happen if I fed one of these LLMs Carl Sagan’s excellent “The Dragons of Eden” and only that?

Its output would consist entirely of words/phrases/sentence/paragraphs from that book because it wouldn’t have anything else to work with: no New York Times, no CRC Reference on Physics and Chemistry, no For Whom The Bell Tolls, no Tao Te Ching, nothing.

Is that transformative? My answer is “no” because it’s 100% Sagan’s words/phrases/sentence/paragraphs, albeit emitted in different quantities and order than the original depending on how the model works.

Now if I feed it those other works — today’s New York Times, the CRC reference, etc. — what will its output consist of? It’ll be entirely words/phrases/sentences/paragraphs from those BUT with some combinations because that’s how LLMs work: e.g., there might be a sentence that pulls from Sagan and from Hemingway.

Is that transformative? My answer is still “no”. It’s still ingestion and regurgitation, just on a larger scale. There’s nothing creative about it: it’s just math. I’ll compare it to picking one line at random from 60 Beatles songs and stringing those together into a “new” song: that’s not transformative, it’s just math.

I give these answers because I think “transformative” requires human creativity and intelligence, and of course neither of these are present in any LLM, nor will they ever be. This isn’t comparable to nicking the plot of The Tempest, transplanting it to space, and coming up with Forbidden Planet: that was transformative and creative, because human beings did that. Not the case with LLMs: it’s just math and that’s all it will ever be.

But I didn’t go to law school: you did. What say you?

Stephen T. Stone (profile) says:

Re:

If “I heavily transformed your work” weighs towards fair use, then how come “I transformed your book into a movie” is not fair use?

You’d still be wholesale adapting the story in the book. That means you’d inevitably use lines of dialogue verbatim. Maybe you can get away with quoting one or two lines of the original work if you’re writing fanfiction. You can’t really lift what amounts to an entire script and still expect lawyers not to wreck your shit in court.

Azuaron says:

Re: Transformation!

“I transformed your book into a movie” isn’t very transformative–all these are a sliding scale, not just “yes” or “no”–because both of them are telling a story, and the same story. Further, none of the Fair Use prongs operate in isolation; a key provision of Fair Use is that the potentially infringing work uses as little of the original as possible to accomplish its Fair Use. “Transforming” a book into a movie falls flat in large part because it’s using the whole story, the characters, worldbuilding, dialog, etc. etc. etc.. It also serves as a market replacement–which goes against another prong of Fair Use.

So even if a movie is a technically different thing than a book, turning a book into a movie still reuses a maximal amount of copyrighted material to the detriment of the original work.

Contrast this with something like “I’m using parts of your novel to create educational content about storytelling.” That’s very transformative: it’s not telling a story, it’s teaching. It’s not republishing all of the copyrighted elements of the original novel. The educational materials cannot be used as a “market replacement” for the novel.

impurify (profile) says:

Re: IANAL, so small words are a bar for me. ;)

Because (1.a) “transformative” is, if true, not alone dispositive; (1.b) “transformative” is only one (major) part of the analysis of one of the four statutory factors (plus any other applicable case law factors) in the overall fair use analysis; and, (2) none of these factors is applied mechanistically.

Can somebody please explain it in small words a lawyer would understand?

Ah, so you want an “ELI5”! 😉 Say that judges are children playing with Lego bricks. A “fair use” ruling is a big, scary, complicated tower made of lots of Legos. The supply of Lego bricks is divided into four bins numbered #1, #2, #3, #4. A big sub-bin within bin #1 is labeled “transformative”.

Even if the judge takes many Legos from that sub-bin, the overall result may perhaps be shaped even more by other Legos from all four bins—maybe. In the example you gave, it would be. Indeed, in your example, I think that as applied by the judge, every Lego from every bin except the “transformative” sub-bin would shape the result into “not fair use”.

My first comment on this page subtly misfired on one minor technicality, and it was thus a tiny bit legally incorrect; let’s see if any lawyer catches it.

impurify (profile) says:

“AI” is the Unfairest Possible Use.

The brain damage of copyright law has helped billionaires with vested interests (and their cat’s-paws) to obscure the real issues. I’ve said this before:

The complete works of Shakespeare are in the public domain. Uncopyrighted. Is it okay to take one of Shakespeare’s sonnets, put your name on it, and republish it as your own work? Of course not!

Well, what if you change a few words here and there? Or mix and match passages from a few of Shakespeare’s sonnets? These are classic methods of cheating, which will get you expelled from a university or will incinerate your career in a scandal when you are caught.

—Well, what if you feed all public-domain sonnets ever published in the English language into a machine that mixes them up for you into a statistical predictor that extrudes a sonnet-shaped glob of text? That’s Emily Bender’s stochastic parrot or text extruder, which hype-mongers dishonestly sell to an ignorant public as “artificial intelligence”.

It cannot be copyright infringement: The LLM’s entire training dataset is in the public domain. But it’s wrong. It’s reprehensible. In the large, it is the end of human culture—and in the small, it’s a direct personal threat to every talented writer and artist.

Consider that hypothetical sonnet-parroting LLM. Suppose, quite plausibly, that it can produce writing which, to perhaps ninety-nine per cent of human readers, is indistinguishable from William Shakespeare. Thus does Shakespeare, no longer immortal, die of social grade inflation (my own original term for this). Shakespeare just isn’t special anymore, when anyone can use a mass-plagiarism machine to write (almost) just like Shakespeare.

Copyright law.

Would that be 17 U.S.C. § 107 “fair use” legally, if Shakespeare’s works were copyrighted? I argue firmly that it is not. It flunks multiple prongs of the four-part fair use test—most of all, “the effect of the use upon the potential market for or value of the copyrighted work.” A mass-plagiarism machine that can produce unlimited forgeries statistically mimicking a work reduces the value of the work and its author to nil.

As I have said, this is a process of social inflation and devaluation. I care about that infinitely more than financial market value; this is not about money to me. However, my same argument does apply to market value—squarely within the ambit of copyright law.

“AI” also flunks the “substantiality” test. LLMs substantially rip off a whole work. What’s confusing judges is that the LLMs rip off numerous whole works en masse—together, all at once. Nothing like that has ever existed before. It is a matter of first impression, and it is not being optimally argued.

I think that part of the problem is that most plaintiffs are money-minded. In principle, they would accept a paid licensing deal. A case needs to be fought by someone who will never in any circumstance consent to the “AI” companies’ misappropriation of their works. Fuck money. This is not about money.

I could argue about the other two prongs of the fair use test; I think I need not reach them, when two prongs are so strongly against “AI”.

Most important: Being purely mechanistic and uncreative appropriations of whole works strip-mined en masse, LLM training isn’t even “transformative” in a meaningful sense. That should bar fair use defenses at the threshold.

What can be done?

In principle, I advocate that copyright law should be abolished. On pragmatic grounds, however, I support the use of it as an available weapon, when we have no others, to fight back against the “AI” companies—the only corporations in history that strip-mine for the direct purpose of creating pollution.

However, not only is copyright the wrong instrument in principle: Even if the “fair use” defense is properly defeated, copyright law is inadequate. For starters: Copyright law does not protect works in the public domain! It is not only a question of protecting old cultural treasures: Some people, such as myself, would prefer in principle to renounce copyright.

New laws are needed. To advocate for those laws, we must first disentangle the issues.

Historically, copyright law rose from a religious censorship system of press control which, by the law of unintended consequences, inadvertently created an entrenched lobby of publishers perversely incentivized to rent-seeking. Copyright strangles culture. Plagiarism is a fraud most odious to the mind, which has been condemned by intellectuals for thousands of years. Plagiarism poisons culture. The copyright industry has often conflates these issues (and unfortunately, so does CC-BY).

And meanwhile, as a practical matter?

When LLM trainers are deadset determined to rip off as many works as they can, it may seem like the only viable solution is to drop out: Renounce human society, and contribute nothing to it. Alternatively, one may spend half one’s life researching technologies that may imperfectly help to keep one’s works out of the LLM trainers’ hands.

This issue is what silenced me in my desire to speak out politically since November 2024. I should be writing on my own website. Instead, I was perpetually sidetracked into researching anti-robot defenses—until one day, I couldn’t keep quiet anymore, and I blew up on Techdirt.

I wanted to keep my words out of the LLMs… sigh.

It doesn’t work. Renouncing human contact is not a good solution, anyway.

Most urgent, perhaps, is to understand the corrupting nature of “AI”, and to advocate accordingly. I have called it the Donald Trump of technology: It’s a cult, and it corrupts all what it touches. People who are bedazzled by “AI” stop seeing what is right before their eyes.

Keep that in mind when reading articles about “AI”.

I’m also happy to see the DOJ make a point that usually gets ignored in these cases: creative people learn by copying. It is natural. Creative people imitate others until they find their own voice. A ruling that training isn’t fair use would turn that whole creative learning trajectory into infringement[.]

Creative people are human beings. They think. They are intelligent—and they are creative. LLMs don’t “learn by copying”: Mechanistic copying is all they do. As aforesaid, I even deny that LLMs are transformative within the meaning of copyright law.

I could take this article apart point by point. I’ll cut it for word count reasons—save to note that my argument about value, which I have made before, flatly contradicts the article.

———

I expressly refuse my consent to this text ever being used to train an LLM. The same applies categorically to anything and everything that I write.

Azuaron says:

Re: Not really

I’m with you about the abolishment of copyright law, and I’m with you on how much AI sucks. However, you’re not correctly applying copyright law.

Would that be 17 U.S.C. § 107 “fair use” legally, if Shakespeare’s works were copyrighted? I argue firmly that it is not. It flunks multiple prongs of the four-part fair use test—most of all, “the effect of the use upon the potential market for or value of the copyrighted work.” A mass-plagiarism machine that can produce unlimited forgeries statistically mimicking a work reduces the value of the work and its author to nil.

Before we even get to the Fair Use test, something has to be infringing in the first place. Infringement is based upon two things:

  • Did the infringer have access to the original copyrighted work? (We’ll just assume an obvious “yes” on this point.)
  • Is the infringing work substantially similar to the copyrighted work?

The problem you instantly run into is that sonnets produced by the LLM are not substantially similar to any sonnet in the training corpus. Therefore, these sonnets don’t even need to pass a Fair Use test; they are simply not infringing.

“AI” also flunks the “substantiality” test. LLMs substantially rip off a whole work. What’s confusing judges is that the LLMs rip off numerous whole works en masse—together, all at once. Nothing like that has ever existed before. It is a matter of first impression, and it is not being optimally argued.

This seems to me to be a completely different argument. Now you’re not saying that a sonnet produced by the LLM is an infringement, but that the training of the LLM itself is an infringement. Well, let’s go back through our standards:

For “training”, LLMs ingest a text document, tokenize it, and create vectors of the tokens in relationship to the other tokens, and in relation to other tokens in the training corpus. This creates a statistically-based language model that allows the LLM to respond to the most-likely next word in a sequence.

Is this token-vector map substantially similar to an ingested sonnet? I don’t think that’s even broadly answerable; it really depends on how mixed the vectors are across many works. A “small” LLM might be essentially just encrypting its training data. But the larger ones are going to completely lose individual works in the statistical mass.

But let’s assume we have an instance where the LLM does maintain enough of an ingested work to say that, though functionally “encrypted”, it can be “decrypted” out of the training model into a substantially similar work. Now we’re in “is this Fair Use?” territory.

The four prongs of Fair Use are:

  • the purpose and character of the use
  • the nature of the copyrighted work — creative work, assume maximally in favor of copyrighted sonnet
  • the amount and substantiality of the portion taken, and — For this discussion, we can assume 100% copied here
  • the effect of the use upon the potential market.

The two we are likely to disagree about are “the purpose and character of the use” and “the effect of the use upon the potential market”.

The purpose and character of the use

This is the “transformativeness” test, and it’s entirely in favor of the LLM being Fair Use. A sonnet is not a machine for creating sonnets. Creating a machine that creates sonnets from math is fundamentally and in essentially every way a different thing than a sonnet.

The effect of the use upon the potential market

You’re not going to like this one, but it’s also in favor of the LLM. I know, I know. But, the sonnet-making machine is not in market competition with a sonnet. It’s doing a completely different thing. When you want to read a sonnet, you read a sonnet. When you want to write a sonnet,^1 you go to the sonnet-making machine and tell it what sonnet you want to write. Buying a sonnet and renting a machine that makes sonnets are two completely separate markets.

“But won’t people read the sonnets created by the sonnet-making machine?”

Of course, but now we’re back on the other side of the fence: the output of the machine, not the machine itself. Is the sonnet the sonnet-machine made substantially similar to another sonnet? The sonnet output by the machine is unlikely to be substantially similar to any sonnet the machine trained on, so we don’t even need to get into the Fair Use factors.

Factor weight

When employing the four Fair Use factors, they are weighted against each other, and they’re not all the same weight. For instance, you could show the entirety of a movie as part of a video series closely analyzing every single shot in the film. That video series would likely be hundreds of hours long, and the film within it would be completely unwatchable as a film. I bring this up to demonstrate that even if factors 2 and 3 are 100% in favor of the copyrighted work, that doesn’t necessarily mean the other two factors can still tip the scale in favor of Fair Use. (Also, I am not a lawyer, and I’m definitely not your lawyer, so don’t come calling on me if you try it and Disney disagrees.)

In particular, this being the US-capitalist hellscape that it is, the “market replacement” factor often gets weighed the heaviest. If the potentially infringing work is not a “market replacement” for the copyrighted work, that alone can decide that Fair Use applies.

I will also point out: this has been litigated twice, and both times the judges found AI training to be Fair Use as a matter of law.

I will reiterate: I don’t like AI and I don’t like copyright law–in part because of how completely it clearly fails when it comes to AI. My goodness am I not talking about what I want, here, just what is. The law and judicial precedent are actually pretty clear on these points.

^1 No, I don’t consider this “writing a sonnet” in the creative sense.

Arianity (profile) says:

AI training absolutely is fair use, and a ruling the other way would blow a hole in fair use protections that have nothing to do with AI, including everything from search engines to book scanning to reverse engineering to (most importantly) text and data mining for research.

You’ve mentioned this a few times, but I don’t see why it has to blow a hole. You could very easily make an argument based on parts of the four factors that don’t affect things like search engines. The most obvious being effect on the market, but also purpose/use. You don’t have to touch e.g. search engines or research.

But much more interesting (and correct) to me is the argument that if training is not fair use, then only a few giant, wealthy companies can afford to create AI tools,

“This would be expensive otherwise” isn’t really a legal argument for fair use, nor is it really one fair use is intended to address.

They often think the end result is somehow punishing the “big” AI companies, but the reverse is true. A finding against fair use locks in the biggest AI companies, and wipes out everyone else,

I mean, the big tech companies are arguing for fair use, instead of using it as a moat.

As we said when the NY Times first floated this lawsuit, no one is replacing the NY Times with ChatGPT. They serve very different purposes

I don’t know if that’s true. There are a lot of people who are using them as effectively search engines, which is replacing/competing with the NYT. And there’s certainly been some sites hit with a resulting drop in traffic. To the point where there are some sites that are blocking AI retrieval bots, despite the hit to SEO/searchability. (That said, the quoted text is making a more nuanced argument)

impurify (profile) says:

Re: The Social Contract of Human Discourse

But much more interesting (and correct) to me is the argument that if training is not fair use, then only a few giant, wealthy companies can afford to create AI tools,

“This would be expensive otherwise” isn’t really a legal argument for fair use, nor is it really one fair use is intended to address.

I find this deeply objectionable about the article—and indeed, about the plaintiffs in most of these lawsuits.

It excludes people who are not seeing this in terms of dollars—as if we don’t exist. As if no one in the world cares about anything more than money!

From the dawn of time, the social contract of human discourse has been that if you put words, artworks, or other creative works out in public, you are giving them over to human benefit.

For my part, I refuse consent to my words being abused to mimic my personality, steal my humanity, devalue me by social inflation, and replace me with a robot. It doesn’t matter what you pay me. My words are only and exclusively for humans, and for tools that organize information for humans to access directly (such as old-fashioned, non-“AI” search engines). And I should never need to opt-out of what I never opted into!

LLM training should require affirmative consent from each and every author whose works are used—including dead ones, whose historic immortality should not be posthumously stolen by stochastic forgery-machines.

TJ Ryan says:

Mike, thanks for this.
I don’t know the facts of the case (need to read), but what if OpenAI actually agreed to NYT’s TOS and created a binding contract — one that prohibited automated login, scraping, etc., or login by multiple computers at one corporate entity, etc. — in order to gain digital access to the NYT archive?
Or what if they hacked in?
Wouldn’t this change the analysis — especially if the TOS contract caused OpenAI to voluntarily cede any fair use rights / defenses it might have had?

Add Your Comment

Your email address will not be published. Required fields are marked *

Have a Techdirt Account? Sign in now. Want one? Register here

Comment Options:

Make this the or (get credits or sign in to see balance) what's this?

What's this?

Techdirt community members with Techdirt Credits can spotlight a comment as either the "First Word" or "Last Word" on a particular comment thread. Credits can be purchased at the Techdirt Insider Shop »

Follow Techdirt

Techdirt Daily Newsletter

Subscribe to Our Newsletter

Get all our posts in your inbox with the Techdirt Daily Newsletter!

We don’t spam. Read our privacy policy for more info.

Ctrl-Alt-Speech

A weekly news podcast from
Mike Masnick & Ben Whitelaw

Subscribe now to Ctrl-Alt-Speech »
Techdirt Deals
Techdirt Insider Discord
The latest chatter on the Techdirt Insider Discord channel...
Loading...