On September 2, 2026, the US Department of Justice filed a Statement of Interest in the copyright infringement litigation between The New York Times and OpenAI and Microsoft, et. al., currently underway in the Southern District of New York. That Statement outlines the perspective of the current administration on the case and argues strongly in favor of OpenAI and against the case for infringement of the publisher’s copyright based on the principles of Fair Use, supporting OpenAI’s case that this use is transformative.

Image of a gavel with the letters AI projected on the surface beneath it

Todd Carpenter

The Administration first outlined its perspective on the use of copyrighted materials in its National Policy Framework on Artificial Intelligence, stating: “Although the Administration believes that training of AI models on copyrighted material does not violate copyright laws, it acknowledges arguments to the contrary exist and therefore supports allowing the Courts to resolve this issue. Similarly, Congress should not take any actions that would impact the judiciary’s resolution of whether training on copyrighted material constitutes fair use.” In this context, it is somewhat surprising then that the administration should advance its biased perspective with the Courts in this statement. The Statement is not ambiguous in its perspective, warning that “an erroneous fair use ruling would hamper competition” among other threats.

Last week, announcing the filing of the Statement of Interest on X, Associate Attorney General Stanley Woodward described the DOJ’s brief as historic. While it is not entirely unheard of for the government to weigh in on district court cases, this Administration has significantly increased the frequency of its use and the range of topics it is willing to engage upon. According to the law firm Wiggins and Dana, these types of filings are typically made between 5 and 20 times per year, often in antitrust or civil rights lawsuits. This administration and its previous incarnation more actively used this signaling mechanism to influence court outcomes in support of its political agenda than other administrations.

It is seemingly part of a larger effort to support AI over the interests of content-creation industries. When the White House removed Librarian of Congress Carla Hayden and attempted to remove Shira Perlmutter in May 2025, Association of American Publishers CEO Maria A. Pallante commented on the events by stating: “Is it all related to the Copyright Office’s AI report? It might be, but we’ll have to see. So far, it has been Big Tech, not the White House, making anti-copyright statements…” Given this statement, it should be quite clear on which side the Administration has aligned itself.

By adamantly stating its opinion in this way, the Administration is seeking to influence the very process it purportedly supported. Perceiving the Court as malleable to its opinion, it is simultaneously seeking to avoid any Congressional action that could influence the outcome. It would entirely be within Congress’ power to adjust copyright law to definitively decide the applicability of Fair Use, or licensing, or to build in some other framework within the law that could apply to AI training and AI systems. It seems, for the moment, that relying on Fair Use exemptions is the preferred and simplest strategy. What that pathway will ensure is a very long road ahead. Given the pace of technological change, this is probably in the best interest of the technology companies, who can avoid any costly outcome for years as these cases wind their way, inevitably to the Supreme Court. As this Court and this Administration have proven time and again, the judicial branch is not always interested in intervening while the facts on the ground create their own new reality.

I’ll leave it to the other Chefs to critique the arguments in more detail. While I’m not a regular critic of legal filings from the Department of Justice, this Statement was striking in its dismissiveness toward other perspectives and the way it drips with condescension. This is perhaps at its worst in a footnote (Note 17 on page 16), where it not only rejects the perspective of the Country’s top copyright official, the Register of Copyrights, it belittles her position: “Her understanding does not warrant deference,” and it further criticizes her reasoning as “threadbare” and “ignor[ing] all the caselaw”. Noting my limited context, it would be surprising if this tone would be as persuasive as the DOJ lawyers think it might be.

The legal argumentation seemed at its worst when arguing that “It is not in the public interest for the largest technology companies to have an oligopoly on LLM training due to the licensing barriers.” As if the licensing costs of training a frontier model are anywhere near the costs of the data centers, the power, or the high-performance chips necessary to train the models. The argument hit its nadir when in the footnote (number 13) in which it argued that “American AI companies [were] at a competitive disadvantage” to those in countries that “don’t respect US intellectual property law.”  This argument is fundamentally and oddly like saying “others are stealing your work, so you should let us steal too.”

This action is also aligned with other administration efforts around copyright. In 2015, there was an effort to pull the Copyright Office out of the Library of Congress and set it up as an independent agency. That effort has seen a rebirth in recent months and appears to be gaining traction. This was originally perceived as a strategy to pull copyright management from the hands of librarians to a situation that could be more effectively managed by industry. However, the concerns seem to only be increasing about regulatory capture of the office if it were independent, although the interests of the content creators and media worlds seem to be at odds with the prevailing interests of the technology companies. It seems the biggest checkbooks will win, and all of the media — including scholarly and mass media, films and music combined — aren’t sufficient.

While the result of the case is far from certain, it is an interesting thought experiment to consider a world that might develop if machine ingesting and training on content is deemed Fair Use. One might see a world in which the free exchange of information is massively curtailed. Authors and publishers might rush to protect their works behind a licensing wall. Even structures like Creative Commons licenses are not without limitations on what one can do with works even though reading the content is free. Except for CC-0, each of the CC licenses has an attribution requirement on subsequent reuse, as well as other restrictions based on what license the author chooses. Machines have been routinely ignoring this license restriction without consequence, but CC is seeking to address this with its Signals effort. We could enter a world where ever-more content is licensed rather than sold, and as part of that license, machine interactions might be limited, if not entirely rejected. We’re starting to see this in the application of AI training rights-reservation statements that were initiated by the EU in the Text and Data Mining (TDM) Exception authorized under Article 4 of the DSM Directive and the EU AI Act, in article 53. Without these reservations in the US, I expect we will see ever-increasing reliance on content licensing to limit access and reuse as a matter of contract law, avoiding the issue of Fair Use and protections of copyright.

Roy Kaufman

Later this month, I will be speaking at a NISO Plus panel where my talk is entitled: “AI, Law, and the Elusive Search for Clarity.” Sorry for the shameless plug, but it is for a good cause.

Still Searching

I was preparing for the NISO talk when I received news of the DoJ’s statement of interest. It has been hard to follow the Administration’s views on copyright and IP. On the one hand, a book, Art of the Deal (protected by copyright) led to The Apprentice TV show (protected by copyright), which led — in part — to the presidency. Coupled with the trademark licensing profile of the Trump family, we saw a very pro-IP first Trump administration.

The current Administration has also been strongly pro-copyright and IP in the trade context (until now) while carefully threading the needle in the US. On the one hand, there are pro-IP White House supporters (the Ellisons, Fox, etc.). On the other, the Administration is very tech-friendly. This tension was best displayed in the March 20, 2026 National Policy Framework for AI, which said:

“Although the Administration believes that training of AI models on copyrighted material does not violate copyright laws, it acknowledges arguments to the contrary exist and therefore supports allowing the Courts to resolve this issue. Similarly, Congress should not take any actions that would impact the judiciary’s resolution of whether training on copyrighted material constitutes fair use….” (italics added)

“Congress should consider enabling licensing frameworks or collective rights systems for rights holders to collectively negotiate compensation from AI providers, without incurring antitrust liability. Any such legislation, however, should not address when or whether such licensing is required.”

There have been opportunities for the DoJ to opine in other AI cases. That this case involves OpenAI (to which the administration has been friendly) and The New York Times (well…) may have some bearing, but of course we need to take the facts as they are.

What Did DoJ Say?

In terms of legal arguments, I do not agree with the DoJ’s conclusions. An “Ask the Chefs” contribution is not the place for a line-by-line refutation, so instead I will just annotate parts of one paragraph:

“An erroneous fair use ruling would hamper competition in the market for LLMs, because only the largest technology companies might have the capital necessary to pay licensing fees”

This is an unironically oft-repeated talking point straight out of the big tech playbook: “We can of course afford to pay for the content we rely upon to build our products, but what about those poor defenseless startups who want to compete with us? If we pay, they will have to do so too. Therefore, none of us should have to pay.” I suppose they make the same arguments as to why they should not pay for electricity in data centers. All else aside, start-ups and smaller entities pay lower licensing fees. At CCC, where our collective licenses cover AI use for millions of works (and users), small and medium enterprises pay much less than large multinationals. And they have access to the same rights as larger entities. The only difference is cost.

“And such licensing fees would disproportionately benefit legacy media outlets due to the sheer volume of their written publications.”

True, but kind of a weird argument. Entities with more copyrights benefit more from copyright law. That is a feature, not a bug. Also, the benefit to the LLMs is greater where there is more volume. Why is this a problem?

“By contrast, if not hindered by a strained understanding of copyright law, LLMs can and should help level the playing field between mainstream and independent publishers, for several reasons. Authors with limited resources can use LLMs to compete (by, for example, using an LLM to generate an image to accompany an article—which otherwise might require a photographer or license).” (italics added)

Thanks DoJ! For a second, I thought you were trying to argue in favor of fair use. Given the Warhol decision and the importance of the fourth fair use factor (market harm), this should end the analysis in favor of the Plaintiffs. That said, let’s see if the “independent publishers” whose materials have been used without consent line up to support the DoJ position by joining in amici briefs. They certainly have been joining CCC’s licenses.

What’s Next?

The Statement of Interest seems to be the first step in a change in the Administration’s views on copyright and AI. At the recent G20 meeting, Secretary Lutnick urged other countries to “embrace fair use,” without offering how to do so. (Note, I can count the countries with fair use on my fingers.) This is seemingly at odds with the Administration’s America First Trade Policy in two major ways. First, in terms of exports, sales of select U.S. copyright products in overseas markets amounted to $272.6 billion in 2023, exceeding US exports from pharma, chemical, agriculture, and aerospace. Harming copyright harms US jobs. Second, if you agree with the DoJ that fair use is a free pass for training (and I obviously do not believe this to be the case), then lowering the copyright bar in other countries will encourage outsourcing from the US. It is therefore no surprise that Fox News has come out against the DoJ position.

Will we soon find clarity on AI and the law, or will the search continue?

Rick Anderson

I expect others will be examining important questions about what we can infer from this document (and other administrative behaviors) regarding the current U.S. government’s posture towards copyright holders, or readers, or publishers, or journalism, or authors (who may or may not be copyright holders). Those are all important issues, but in the space I have here I’m going to focus on whether I believe it makes sense, from what I understand of copyright and case law, to consider it to be fair use when a large-language AI model (LLM) is trained on copyrighted material. (My thoughts will all be framed within the context of U.S. law, and may or may not have much relevance to other jurisdictions.)

As we all know, U.S. law (17 U.S.C. § 107) provides a fourfold test that can be applied when a user is contemplating use of a copyrighted work that technically infringes on one or more of the copyright holder’s exclusive rights. If the use meets one or more of those tests sufficiently, then the use is “fair,” and therefore “not an infringement.”

When it comes to the training of LLMs on copyrighted work, the argument provided in favor of fair use is usually based on the doctrine of “transformative use,” as established in case law. The basic idea here is that if the use of the copyrighted material results in a new product that is very different in form and application from the copyrighted work, it’s more likely to be a fair use. It was a transformative use defense that allowed Google to prevail in Authors Guild v. Google, Inc. – they successfully argued that scanning copyrighted books to create a searchable database constituted a transformative use of those books, and was therefore a fair use of them.

In Section C of the statement we’re examining, the U.S. government relies on a similar argument, asserting that “training of LLMs on written works is exceedingly transformative.” The statement cites a variety of legal precedents in support of that assertion, including Bartz v. Anthropic PBC, in which the judge found Anthropic’s use of copyrighted materials to train an LLM “quintessentially transformative — spectacularly so” — while, at the same time, the judge ruled that using illegally acquired texts for that purpose was itself still illegal; in other words, making a fair use of the copyrighted content did not erase the illegality of using pirated copies. (Unsurprisingly, the Copyright Alliance disagreed strongly with that finding.)

I’m not an attorney, so take my opinion for what it’s worth. But it does seem clear to me that copying copyrighted works for the explicit and limited purpose of training an LLM, and then letting the LLM use what it “learns” in that process to create new documents, is quite similar (from an intellectual-property perspective) to what Google did in creating its searchable database, and therefore represents a transformative use of the copyrighted works. Not only does this principle seem to be well established in case law, it also just makes intuitive sense: if the LLM is simply being “trained” on the copyrighted content, and is then producing new content based on what it has “learned,” that seems no more a breach of copyright than it would be if I read ten copyrighted books and then wrote my own book, using my own original expression of my own thoughts, informed by what I learned from my reading.

Of course, it’s entirely possible that an AI agent, thus trained, could subsequently produce infringing documents – the government’s statement acknowledges this, pointing out that an LLM can create infringing products if it “reconstructs and disseminates an original copyrighted work.” It also matters how the human owners of the LLM gain access to the content. Referring back to Bartz, gaining access to that content in a manner that breaches the copyright holder’s exclusive rights would not be justified by a subsequent fair use of the content itself.

Assuming that access to the content is gained legally, I have a hard time imagining a compelling argument against the transformative nature of using that content to train an AI agent, regardless of what we may see as the motivations of the AI companies engaging in these activities — or of those who produced the U.S. government’s statement in support of them.

Haseeb Irfanullah

After reading the Statement of Interest of the United States, one thing came to my mind — how is the “fair use” doctrine of copyright law of the US applied on those copyrighted materials (e.g., newspaper editorials and journal articles) that were also prepared following the fair use principle? Is there any moral dimension to it? Along the same train of thoughts, I was also wondering, if the newspapers’ copyrighted materials are allowed to be used to “teach” and improve LLMs, will that same use be expanded to include copyrighted journal articles on the same grounds? Or are journal articles different from pieces published in newspapers?

Journal articles, and all kinds of science communications in a broader sense, highly depend on the fair use. New knowledge is essentially built on the existing and widely copyrighted data, information, and  knowledge. Since new manuscripts, which are submitted to journals as individual outputs at the end of a research process, are “exceedingly transformative”, thus the fair use doctrine is applied. But, there is another aspect to it. Despite having thousands of academic publishers and tens of thousands of journals, the publishing industry is an extremely closed world. It is often run by norms, practices, and traditions as well as business interests, rather than laws alone. That’s why no two journals of the same discipline are considered to negatively affect each other’s potential market or value — an important aspect of fair use. Rather, the more each cites the other’s articles, the “market value” of the cited journal increases as measured by different citation-based metrics. This further translates into revenues (doesn’t matter if a journal is commercial or not-for-profit) given the ten or so billion dollar market size of academic publishing.

Let me now draw an analogy between the two ecosystems: journal publishing and LLM development. As we design a research project by accessing existing journal articles and other pieces of literature and then conduct the actual research, it is broadly similar to training an LLM — everything happens in a “confined” space, not in the “public” domain, based on the existing (copyrighted or not) materials. Even the peer-review, either double anonymous or open, of research manuscripts is similar to testing of LLMs by stakeholders. Once a journal article is published as a final output, it may be similar to a version of an LLM.

I believe that the collective attitude of publishers within the journal publishing industry on the grounds of fair use needs to be embraced by copyright holders of all materials toward LLM development. Given the current pace of advancement of GenAI writ large and LLMs specifically, it is no longer ‘us versus them’, we are already in it together.

Todd A Carpenter

Todd A Carpenter

Todd Carpenter is Executive Director of the National Information Standards Organization (NISO). He additionally serves in a number of leadership roles of a variety of organizations, including as Chair of the ISO Technical Subcommittee on Identification & Description (ISO TC46/SC9), founding partner of the Coalition for Seamless Access, Past President of FORCE11, Treasurer of the Book Industry Study Group (BISG), and a Director of the Foundation of the Baltimore County Public Library. He also previously served as Treasurer of SSP.

Roy Kaufman

Roy Kaufman

Roy Kaufman is Managing Director of both Business Development and Government Relations for the Copyright Clearance Center (CCC). Prior to CCC, Kaufman served as Legal Director, John Wiley and Sons, Inc. He is a member of, among other things, the Bar of the State of New York, the Author’s Guild, and the editorial board of UKSG Insights. Kaufman also advises the US Government on international trade matters through membership in International Trade Advisory Committee (ITAC) 13 – Intellectual Property and the Library of Congress’s Copyright Public Modernization Committee in addition to serving on the Board of the United States Intellectual Property Alliance (USIPA).

Rick Anderson

Rick Anderson

Rick Anderson is University Librarian at Brigham Young University. He has worked previously as a bibliographer for YBP, Inc., as Head Acquisitions Librarian for the University of North Carolina, Greensboro, as Director of Resource Acquisition at the University of Nevada, Reno, and as Associate Dean for Collections & Scholarly Communication at the University of Utah.

Haseeb Irfanullah

Haseeb Irfanullah

Haseeb Irfanullah is a biologist-turned-development facilitator, who often introduces himself as a research enthusiast. Over the last 26 years, Haseeb has worked for different international development organizations, academic institutions, donors, and the Government of Bangladesh in different capacities. Currently, he is an independent consultant on environment, climate change, and research system. He is also involved with the University of Liberal Arts Bangladesh as a visiting research fellow of its Center for Sustainable Development.

Discussion

2 Thoughts on "Ask the Chefs: AI and Copyright Licensing"

“But it does seem clear to me that copying copyrighted works for the explicit and limited purpose of training an LLM, and then letting the LLM use what it “learns” in that process to create new documents, is quite similar (from an intellectual-property perspective) to what Google did in creating its searchable database, and therefore represents a transformative use of the copyrighted works. ”
Not only that, but that’s exactly what happens in the heads of human scholars who often obtain the texts through other aspects of fair use/fair dealing, including interlibrary loan, learning from what they “processed” to create new works.
Far be it for me to agree with anything coming out of the Trump Regime, but in this case the stopped clock that is correct twice per day accidentally got something right, even if for all of the wrong reasons.
This is about as “transformative” as any use can get.
There is one horrible error in the case law in the US regarding fair use and the “effect on the market” factor which sometimes seems in the US to be the only one that actually matters. That error made by some judges is that the secondary “market” created by the option of paying for what would otherwise have been fair use, that is the publishers licensing the use to the AI companies, is itself an argument against that factor. That is absurdly circular but unfortunately it does sit in the US case law un-overturned, I believe.

Melissa- Reasonable people (and reasonable courts) will differ on how to apply fair use to these cases. Factually, humans read, listen, and watch. Machines “learn” (which is an anthropomorphic term applied by tech similar to how tech companies call their mistakes “hallucinations”) by copying, storing, and making available for perfect retrieval. These acts are covered by copyright law. Also, we should not lose sight of the fact that for hundreds of years, copyright law protected against unauthorized copying (subject to exceptions such as fair use) for human consumers and learners. Even if the machine copying is for the “same purpose,” that does not end the inquiry. I absolutely agree that the Google cases are going to be influential here and will need to be reconciled with TV Eyes, Thomson Reuters v Ross, Warhol, and the Internet Archives cases, for a start. Difficult stuff and I appreciate your engagement.

Leave a Comment