Editors’ note: Today’s post is by Wendy Queen and Jonathan Woahn. Wendy is Chief Transformation Officer at Johns Hopkins University Press and former Director of Project MUSE. Jonathan is a co-founder of Cashmere, a platform designed to help publishers safely and responsibly monetize their content in AI-powered applications. 

Scholarly publishing has a habit most industries would envy: it competes on products while cooperating on infrastructure.

Publishers compete for authors, readers, subscriptions, and prestige. Then those same publishers, along with libraries, scholarly societies, and technology partners, get together and build the shared plumbing that makes the whole ecosystem more valuable at once — persistent identifiers, researcher identity systems, linking standards, and cross-platform usage reporting. These are not products, but community infrastructure no single organization could justify alone, yet that created value for everyone once it existed.

Video wall with many small screens and the letters AI superimposed

Here’s the uncomfortable part: we keep building the pieces that lay the foundation and letting someone else own what gets built on top of it.

The first time, we digitized the scholarly literature, and a handful of commercial aggregators came to own most of the reader or library relationships, along with most of the value that was created. The second time, we built the citation graph, and Google came to own discovery. Twice we built the foundation and ended up a supplier inside a system someone else erected on their terms.

Era Community infrastructure What it standardized
Print Library catalogs, ISBN Description and discovery
Digital/The Web

DOI, Crossref

OpenURL, COUNTER

ORCID

Citation and persistence

Linking, access, and usage measurement

Researcher identity

AI How AI participates

The empty box representing AI community infrastructure is the point of this post. Each previous generation of scholarly infrastructure addressed a coordination problem that required collective action. AI presents the next really big one. The question is not whether an AI participation layer will emerge, but whether it will be defined, built, and run by the scholarly community, or by someone else.

The Question Has Changed

For much of the past two years, the community has framed AI as a question of access: can AI systems reach trustworthy content, and on what terms? That work matters, but it treats AI like the last transition — a better search box, a new front door to content we still control.

It isn’t. Researchers don’t just ask AI to find a paper or a book anymore; they ask it to discover, compare, synthesize, explain, and reason across bodies of scholarship. Search sent the user onward to the source. AI reads and answers — it is no longer simply finding literature, but becoming an active participant in the research and analysis itself. This changes the question from “How can AI access our content?” to “How should AI participate in scholarship?”

And here’s what makes the shift easy to miss: AI already behaves reasonably well. A colleague recently asked a popular AI system a niche literary-history question, logged in through her university library. It returned a credible answer with real citations, and even noticed her library didn’t hold one of the books, showing the abstract instead. Not a bad experience.

Which is exactly the point.

The problem is no longer that AI hallucinates or invents sources, because the models got better. Citations solved yesterday’s problem. The problem is that AI has no obligations. In this case, it provided some citations, but it wasn’t required to. It respected an entitlement, but nothing guaranteed that it would. It could just as easily have blended authoritative sources with a preprint, a retracted paper, or provided a confident summary that never sent the reader back to the record, let alone cited it. Good behavior today is a happy accident. Good behavior tomorrow must become the standard.

Trust Markers Are Only Part of the Answer

The scholarly community has invested heavily in trust markers: peer review, editorial oversight, persistent identifiers, licensing, metadata, and marking the version of record. Together, these communicate to an AI system what a work is — who created it, how it has been evaluated, and where it belongs within the scholarly record. But markers are descriptions, not instructions.

Think about how we trust a human researcher. We’d never hire one on credentials alone and skip asking how they work—whether they cite sources, respect embargoes, acknowledge collaborators, recognize uncertainty, and direct readers back to the original evidence instead of passing off a summary as the thing itself. Credentials tell you what someone has accomplished. Conduct tells you whether to trust them.

That distinction organizes everything that follows.

To trust an AI system like a responsible researcher, we would expect it to authenticate authorized users, retrieve licensed content appropriately, preserve attribution, distinguish authoritative scholarship from preliminary findings, report meaningful usage, and return researchers to the scholarly record rather than standing in for it.

Those are behaviors, not descriptions — and none of them happen simply because metadata exists. There is nothing in metadata that can authenticate a user. Provenance can identify the version of record, but it cannot require an AI system to privilege it. The challenge is no longer describing scholarship, but enabling AI systems to act consistently on what scholarship already knows about itself.

Why Coordination Matters

The reason these behaviors remain difficult is structural. No one organization holds everything an AI system would need to behave well.

Publishers maintain the version of record, editorial oversight, licensing, and provenance. Libraries fund and negotiate access, steward collections, and support researchers through discovery and long-term preservation. Technology providers increasingly mediate how researchers interact with scholarship. That distribution of responsibilities is not a weakness of the scholarly ecosystem. It reflects decades of specialization across organizations, and each performs those distinct roles exceptionally well.

This leads to every AI provider rebuilding the same integrations, one publisher at a time; and every publisher is fielding the same requests from every provider. It’s the pre-Crossref world again: a thicket of one-off, bilateral connections, none interoperable. We’ve seen this shape before — in Jonathan’s series earlier this year, he called these the “token partnerships” that appear right before shared infrastructure forms: bespoke, fragile, unscalable.

Here is the part that should concern us: if AI becomes the primary interface between researchers and scholarship, whoever defines that interface defines discovery, attribution, reporting, licensing, and ultimately the flow of value. The failure mode here is not a villain, but our own bad habit.

Both times before, like with commercial aggregators and Google, when a new layer emerged, we optimized the content for immediate reach and convenience, and let the layer harden above us; either because it was easier than coordinating or because we couldn’t come to a consensus fast enough. And, both times, whoever owned the layer ended up setting the terms our industry had to settle for.

The AI interaction layer is forming now, and it’s far cheaper and more convenient to set the terms before they harden than to renegotiate them after someone else has done it for us — the way we’ve let them happen twice already.

We Don’t Have Three Years

In his June interview, Todd Toler was asked to imagine a coordination layer as something that might exist “three years from now.” This is the precise type of assumption worth challenging directly — we don’t have three years to get this done.

Consider the pace of the thing we are trying to coordinate with.

ChatGPT reached 1 million users in 5 days, and 100 million users in 2 months. The previous record holder was TikTok with 9 months. Anthropic, Claude’s parent company, took Claude Code from $0 to $1 billion in annualized revenue in 6 months. Anthropic reported company annualized revenue of $47 billion in May of this year. It took Google, Amazon, and Facebook each over a decade to reach even $10 billion.

Debate any single figure, but the direction and speed cannot be disputed — and it has already reached our doorstep. On June 30, 2026, Anthropic launched Claude Science, a research workbench pre-loaded with dozens of databases and built, in the company’s own framing, to own the operational layer of research the way Claude Code came to own software development. The starting shot for scholarship has been fired.

So picture taking three years. Someone else builds the layer where AI and scholarship meet; when it’s ready, everyone else uses it on terms they didn’t write — exactly how discovery and the reader relationship slipped away the last two times. A three-year-old standard arriving into a market moving at this speed is just plain too late.

The Opportunity

Someone is going to build that layer. The only question is who.

The answer isn’t for scholarly publishing to build a competitor. The scholarly community has never succeeded by out-competing the technology companies that enter its market; it succeeds by building the infrastructure that lets many organizations participate within a common framework.

The same opportunity exists here. Rather than creating hundreds or thousands of bilateral integrations, the community can define a shared participation layer through which authorized access, attribution, provenance, licensing, and reporting work consistently regardless of publisher, institution, or AI provider. Build it once, and every provider integrates once. This helps everyone in the system.

And it doesn’t replace anyone.

It preserves the roles the ecosystem already depends on, carried into an AI-mediated environment. Researchers continue to benefit from institutional access negotiated by libraries. Publishers continue to steward the integrity of the scholarly record, and ensure that their authors’ works are discovered, used, and have impact. AI systems become better scholarly participants, more effectively serving that slice of the AI companies’ customers, because the ecosystem provides the tools needed to support responsible behavior.

The Work Has Already Begun

The scholarly community has solved this kind of problem before — the rows in the table at the beginning of this article are proof. AI is the next coordination problem, but it is more challenging than anything we have previously faced; primarily because it is taking shape faster than any other layer before it.

The conversations have already started. Publishers, libraries, societies, standards bodies, and technology partners are working through what responsible scholarly AI has to do, and what it takes to build the layer that makes it possible. The goal isn’t another principles document or standard that arrives three years from now. It’s something tangible, live in the market in months.

No single organization can build this alone, and no single organization’s version should become the standard. It will only succeed if the community shapes it together. The terms get set by whoever is at the table. If you and/or your organization believe the scholarly community should set the terms of what responsible AI participation in scholarship should look like, there’s a place for you at the table.

We’ve built the foundation twice and handed away what got built on top; the third time is not the charm. Let’s not build the foundation, yet again, and rent back what we could have owned.

Wendy Queen

Wendy Queen

Wendy Queen is Chief Transformation Officer at Johns Hopkins University Press and former Director of Project MUSE. She led major advances in accessibility, metadata, digital workflows, and open access. At Hopkins, she guides AI integration and transformation efforts, promoting sustainable, collaborative innovation across scholarly publishing.

Jonathan Woahn

Jonathan Woahn

Jonathan Woahn is a co-founder of Cashmere, a platform designed to help publishers safely and responsibly monetize their content in AI-powered applications. He believes human-created content is what connects us and advances shared knowledge, and that creators and publishers need clear incentives to continue producing it. His work focuses on building systems that preserve compensation, credit, and control over how content is used, while enabling AI to responsibly incorporate high-quality, human-generated material at scale.

Discussion

8 Thoughts on "Guest Post — The Third Time is NOT the Charm: Who Will Own How AI Participates in Scholarship?"

Thank you for this thoughtful post. Re: “AI participation layers” to standardise authorised access, provenance, licensing, attribution and reporting across publishers, institutions and AI providers: As an outcome of the Cambridge metrics workshop this spring (https://scholarlykitchen.sspnet.org/2026/06/25/making-ai-use-of-scholarly-content-traceable-measurable-and-trustworthy-a-meeting-report-from-cambridge-scholarly-ai-workshop/) , NISO and Creative Commons will be co-leading a provenance and attribution working group that seeks to do exactly this. Announcements will be coming soon, and colleagues across our space who are interested in this are encouraged to join this effort. Involving standards organisations and the entire ecosystem, not just commercial groups, in the development of this infrastructure is critical.

Monica,

We actually had you all in mind as we wrote this piece.

Let’s grab some time to chat further, we’d love to learn more about what’s on the docket, and we can share some of the things we’re working on as well—it’s moving quickly!

I’ll send you an email.

Sounds great, Jonathan! Todd has the guiding vision for this, so I’ll make sure to loop him in.

Thank you Wendy and Jonathan!

I think of the internet, and the technologies and people that interact through it, as an exchange that inherently favors scale and network effect. Always have, always will. There may be many centers, but scale and funding drive centralization.

There’s a tension between near-perpetual community-owned infrastructure and a capital-driven layer, and it rarely feels obvious that the same body can satisfy both over the long term.

So my question is organizational. What body do you envisage that can be stood up expeditiously?

When I think of centralized community medication in our industry at scale, my head lands on someone like CrossRef. I think Crossref is a near stroke of genius. From their website today “>25,000 members in 167 countries, we drive metadata exchange and support 2.1 billion monthly API queries”.

Someone like Crossref could objectively help define schema and required behaviors, and vendors could sell their own MCPs, gateways, etc using that mediation. CrossRef wouldn’t become a competitive service operator, but provide a floor that we can all move on, but it does not become a service operator. Maybe this just my imagination.

Is it something along those lines that you imagine, or do you have something different in mind?

I ask because the difference between a floor and an operated hub will influence who ends up “owning” the layer, which is the very thing your piece is about.

Best, Neil

Neil,

It’s a great point—the tension is real. It’s one of the things Wendy and I spoke at length about in the writing of this article, as I come from the “capital-driven” layer and Wendy more represents the “community” side of it.

There are a number of inherent challenges here in either direction. If left fully to the community, then there’s the time lag inherent in aligning collaborators, finding funding, and establishing direction. If left fully at the capital-driven layer, then the concerns of funding and direction move faster, but the long term concern is precisely what Wendy and I highlight in this article—with them also capturing the long term economics.

What’s needed then, is initially the speed and direction typically associated with the capital layer, with the right long term community incentives and ownership—the two need to be balanced with each other—and we think there’s a path forward on how this could work.

What do you think?

Thank you Jonathan, I think you and Wendy are onto a very important initiative!

An additional variable that you made me think of in this effort is big-publisher support. They can diminish or amplify this type of initiative. They may be wary of supporting initiatives where they sense that they are needed more than they need the initiative, where it potentially undermines their current or future competitive leverage, and where they are uncertain of future ownership or influence. Just a natural part of business, but that may be another balancing act.

I still ponder, what keeps capital-driven players from hardening into permanent layers where humanly/ AI-possible. Every business wants and needs to be needed more than its competitors. Eventually, and for some time, a few always come out on top in ways no-one quite predicted.

There are very thoughtful perspectives in this article. Thank you for putting it together.
I agree with “It will only succeed if the community shapes it together”, and I found it wise to ask others so openly for participation.

Leave a Comment