Editors’ note: Today’s guest post is by Patrick Hargitt and Steve Smith. Patrick is co-founder and CEO of Olive Turn, enabling researchers to use trusted content in AI workflows; he is also co-chair of NISO’s Journal Article Versions Revision Working Group and previously chaired STM’s CUSAP Task & Finish Group. Steve is founder of STEM Knowledge Partners, advising scholarly publishers and research organizations on AI access, licensing, and strategies for building durable value around trusted content, context, and research communities. Reviewer credit to Chef Alice Meadows.
Crossref’s new position paper on persistent identifiers makes a simple but important point: persistent identifiers (PIDs) are a means to an end. Their value comes from the metadata, relationships, services, and stewardship around them. An identifier gives us a stable reference. The surrounding infrastructure tells us what the object is, how it relates to other research outputs, and whether its record is still current.
AI makes that distinction much more urgent.

Research is no longer found only through journal pages and conventional indexes. The same work may appear on a preprint server, a publisher platform, an institutional repository, a scholarly index, a licensed AI collection, and eventually inside generated summaries or answers. AI systems may retrieve passages from any of these places, split the content into chunks, and reuse claims without taking the reader back through the context shown on the original article page.
Version ambiguity is therefore becoming more than a discovery problem. It is becoming a provenance problem.
Consider three common situations: A preprint may go through several revisions before journal publication. A Version of Record may be corrected or retracted, while repository copies and downstream indexes still hold the earlier text or fail to reflect its changed status. In post-publication peer review, an article may move through several reviewed and revised states before, or instead of, reaching a conventional Version of Record.
In each case, PIDs are critical for keeping the research record connected. But identifiers alone may not tell a researcher, an index, or an AI system which state of the work was actually used.
A PID Is Necessary, but It Is Not Self-Describing
PID assignment practices for outputs vary. Some platforms retain a single PID as a work evolves, often a Digital Object Identifier (DOI). Others assign separate PIDs to distinct versions or manifestations. Both approaches can serve valid purposes. The problem is assuming that a PID will, by itself, communicate what changed, which version is newer, or which version should be used.
A cited DOI may resolve correctly to the maintained record even though an AI system retrieved an older repository copy. Multiple identifiers can identify related outputs without revealing their sequence or differences. A citation can identify the general work while remaining ambiguous about the exact text that supported the claim.
This is consistent with Crossref’s wider argument that the practical value of PIDs does not sit in the identifier string alone. It comes from the metadata and connections that allow people and machines to understand the record.
What the Upcoming NISO JAV Adds
The revision of the NISO Journal Article Versions Recommended Practice is now in its final editorial stage. It is expected to be sent to NISO voting members shortly, with an official release anticipated in the fall if approved.
The revised recommendation reflects a publishing environment in which articles can be shared, reviewed, revised, accepted, published, corrected, and redistributed through several systems. It combines clearer article stages and statuses with semantic versioning, providing a way to identify repeated changes within the same stage.
This matters because article stage alone is not enough. A preprint, manuscript under review, Accepted Manuscript, or Version of Record may each be updated more than once. Semantic versioning can distinguish those states and signal whether a change is limited, meaningful, or substantial, while PIDs and relationship metadata keep the versions connected.
The precise boundaries will depend on publisher and community policy. A minor metadata correction should not necessarily be treated like a change that alters how the research is interpreted. The important point is that the version can be expressed in a structured form that both humans and machines can use.
Three Layers of Identity
It may help to separate three questions:
- What research work or object is this? A PID, such as a DOI, provides a stable identity, while its associated metadata and relationships connect the object to the wider research record.
- What state of the work is this? Article stage, status, and semantic versioning provide specificity.
- What did the AI system actually use? Provenance identifies the exact source and version retrieved to support the output, not merely the version to which the final citation happens to resolve.
These layers should reinforce one another. The goal is not to replace PIDs with version numbers. It is to make the metadata and relationships associated with a PID precise enough to be meaningful in an environment where content is distributed and reused by machines.
In an AI-generated literature review, policy summary, or news report, the user may never see the original source page. A citation including only the DOI may be technically correct but still insufficient to reproduce the evidence path. Version specificity becomes part of citation accuracy, not an optional display enhancement.

Version Metadata Needs to Become Operational
Article version information has often been treated mainly as display metadata. It may appear on a landing page, inside a PDF, or in a publication history. In AI workflows, however, version information needs to inform ingestion and retrieval, not merely appear in the final display.
A retrieval system should be able to recognize whether it has found a preprint, a manuscript under review, an Accepted Manuscript, or a Version of Record, and whether the specific version has subsequently been corrected, retracted, or superseded. Where relevant, it should identify whether a later version exists before producing an answer. If it uses an earlier or non-final version, that should be visible in the citation or supporting evidence.
The same principle applies downstream. When a Version of Record is corrected, information about the correction should travel beyond the publisher page to repositories, indexes, licensed collections, and AI services. Otherwise, AI can amplify a version that the scholarly record has already moved beyond.
The NISO Communication of Retractions, Removals, and Expressions of Concern (CREC) Recommended Practice provides an existing model, setting out how related metadata should be created, transferred, and displayed across the scholarly ecosystem. Version and correction information should have comparable reach into downstream AI services.
This does not require every answer to display a full version history. At minimum, the visible citation can be as simple as the DOI plus version. The system should also retain the supplying source and surface the article stage and any correction, update, or retraction status when relevant.
What Version-Aware AI Access Requires in Practice
For publishers, this is not only a metadata question. It is also a question about how content is delivered, licensed, and maintained once it enters an AI workflow. For institutions, it becomes a procurement question: can the AI services they license preserve version and integrity metadata, update their systems when articles are corrected or retracted, and report the exact source and version used to support an answer?
An article supplied to an AI provider will often not remain a complete object. It may be converted, divided into chunks, embedded in a vector index, and combined with copies obtained from repositories or other sources. In that process, the text can become separated from the version, status, and publication history information that gives it meaning. Supplying accurate metadata at the point of ingestion is therefore necessary, but not sufficient. Publishers also need to know whether that context survives indexing, retrieval, and generation.
The problem becomes more difficult when the scholarly record changes. A licensed collection may contain the Version of Record as it existed on the day it was ingested, but what happens when the article is subsequently corrected, retracted, or replaced by a later valid version? Unless the arrangement includes a reliable update mechanism, an apparently authoritative AI service can continue retrieving content that the publisher’s own platform no longer presents without qualification.
The risk is greatest with static ingestion, where a collection becomes a snapshot unless updates are pushed or regularly reconciled. Live retrieval can reduce that risk by checking the maintained record at the point of use, although it still depends on reliable status and update services. The revised NISO JAV recommendation describes how versions should be identified and related; it does not itself provide an update channel. That requires complementary infrastructure, including the push-based alerting service being developed through STM’s Content-update Signaling and Alerting Protocol (CUSAP) initiative and point-of-use status signals such as the retraction and errata service from Get Full Text Research (GetFTR).
This suggests several practical requirements for AI vendors, platform providers, and licensing partners. Version and integrity metadata should remain bound to the text or passage delivered to the retrieval system. Corrections, retractions, and superseding versions should trigger updates across downstream collections. The system should retain enough provenance to report not only the article DOI, but also the source, version, and status used to support an answer. These capabilities should be tested during a pilot rather than assumed from the presence of metadata in the original content feed.
These update flows need clearly assigned responsibilities and explicit service obligations. Publishers should treat the delivery of corrections, retractions, and version changes as part of the maintained service they provide, whether through subscriptions, licenses or other content-supply arrangements. AI vendors and platform providers should, in turn, commit to receiving, reconciling, and applying those updates within an agreed period, while preserving enough provenance to identify the source and version used to support an answer. If publishers and institutions do not set clear expectations through standards, licenses, and pilots, vendors and aggregators will establish their own “good enough” bar, which may not preserve the maintained scholarly record.
From Persistent Identity to Trusted Context
AI will make PIDs more important, not less. Research output is growing, distribution is widening, and generated derivative content creates more places where the scholarly record must remain connected.
But persistence and precision are different problems. An additional PID may be justified when a genuinely distinct object needs its own identity. It should not be the default response to every change when article stage, status, semantic versioning, and provenance can communicate the evolution more clearly.
There is a strategic opportunity here as well as a responsibility. Much scholarly text can already be found through repositories, indexes, and other distributed sources. What those copies do not necessarily provide is a live, authenticated connection to the maintained scholarly record. Version-aware access could therefore become part of the differentiated value of licensed publisher content: not simply the ability to retrieve an article, but greater confidence that an AI system is using the current, appropriately labeled and verifiable state of that article.
This is one aspect of what it might mean for institutions and AI services to subscribe not only to content, but to context.
AI can already find the article. The next step is making sure it can identify which version it used.
Authors’ note: Generative AI was used to assist with drafting and copyediting this article. The ideas, argument, source selection, and final editorial decisions are entirely the authors’ own, and all factual claims were reviewed by the authors.