As momentum built to move research publishing from a print-based distribution model to one driven by electronic access in the 1990s and early 2000s, a variety of community initiatives helped to formulate and build the infrastructure that we have since come to rely on. The DOI pilot project began in the mid 1990s and led to the formation of Crossref. The Digital Library Federation E-Resource Management Initiative (ERMI) helped to catalyze work on ERM systems, indexed discovery systems, knowledge base information exchange (KBART), and usage data standards (COUNTER). The movement toward open access also began to develop momentum with the meeting organized by the Open Society in Budapest in 2001. While not every digital publishing issue has been solved, a lot of progress has been made, and even more is continuing to advance.

Today, we’re experiencing a similarly rapid transformation around AI tools and how the landscape of research communication will need to shift to accommodate agentic access to content, machine consumption, and the generation of content based on those processes. Much as electronic content distribution upended many of the structures and systems of print-based content sharing, so too will the role of AI systems require a reconsideration of everything from access, citation, usage tracking, and the financial models that support the entire marketplace. Certainly, not everything will necessarily need to be redeveloped, but many new systems and standards will need to be put into place. While the process for the print transformation to electronic took three decades, the pace of our adaptation to the new AI world will need to happen significantly faster. The impacts could be just as profound, even if we can’t fully discern their shape today.
Launching TRACE
As it did in the early days of web-based publishing, the community is coming together to accelerate progress on AI systems. This week, Priya Madina from Springer Nature, Tasha Mellins-Cohen from COUNTER, Todd Toler from Ithaka S+R, and I spoke from the stage at the STM Conference prior to the start of the Frankfurt Book Fair to discuss a new community initiative to coordinate some of this work. TRACE: Trusted Retrieval & Attribution for Content Ecosystems is an initiative aimed at bringing together members of the scholarly publishing and AI communities to benefit everyone in the research communications ecosystem. TRACE aims to accelerate the development of the technical infrastructure needed to improve the trustworthiness of AI systems in scholarly communication. It does so by supporting the development and adoption of content exchange standards, provenance tracking mechanisms, and usage reporting protocols. Beyond funding and incubating new work, TRACE also provides a coordinating forum that helps align and amplify complementary efforts already underway across the ecosystem.
The initiative has launched with funding for three parallel projects that address key challenges in AI-mediated scholarly communication: how AI agents discover and access content, how provenance and attribution are preserved throughout the content lifecycle, and how AI usage can be measured and reported consistently. Together, these infrastructure components aim to ensure that authors and content creators receive appropriate recognition regardless of publisher business model, while also reducing friction and computational overhead in content discovery and access. By establishing a coherent framework for trusted retrieval, attribution, and measurement, TRACE will enable AI developers to build more reliable products and services. In turn, this will strengthen confidence in AI-assisted research workflows and help foster greater trust among researchers, publishers, technology providers, and the broader scholarly communications community.
The First TRACE-related Projects
COUNTER began work on AI usage metrics in 2025. Phase one of that work progressed and, in early 2026, established initial guidance, including a new Access Method, Agent, and new AI-specific metrics. This effort sought to align AI usage metrics with traditional metrics in a way that clearly demonstrates changing usage patterns and budget implications. Feedback on phase one noted that the approach was too technology-dependent and, as a result, phase two is taking a more technology-agnostic approach in the guidance. Most importantly, phase two will incorporate reporting about inference-time usage — tracking content when an AI system retrieves, grounds, or cites publisher material in response to user prompts — outside of publishers’ own platforms.
Participants in the TRACE initiative were engaged in that phase one work and continue to engage in COUNTER’s phase two, which began in the spring of 2026. Funding from TRACE will support the continued development of the COUNTER API and JSON Schema to include the new AI metrics that were advanced earlier this year. COUNTER is also exploring usage models developed by the SPUR Coalition, which is seeking to develop broader AI usage metrics that overlap with the guidance COUNTER has proposed. This includes plans to reference SPUR telemetry in our phase two guidelines. We’ll be mapping citation and presentation telemetry events to the existing AI COUNTER metrics. We are also going to use retrieval telemetry events for a new holistic “AI retrievals” metric.
Like COUNTER, NISO has been tracking AI developments, considering potential standards work and providing training on AI tools well before TRACE was formulated. NISO’s work on provenance tracking had its origins last year during a prioritization discussion NISO hosted in the spring of 2025. Those discussions led NISO to organize sessions during the NISO Plus Conference in February 2026 around AI applications, usage, and discussions of ”How to keep the robots in line.” Following the conference, further conversations grew out of the many ideas around AI during that conference. Those formulations led to a workshop in Cambridge in May, which I reported on earlier this summer. One concrete outcome of the Cambridge Workshop was to launch work on a minimum metadata package that could describe content provenance through the inference process. This project is being supported by the TRACE initiative.
STM has also announced a third component of the TRACE project. A three-month study, led by Ithaka S+R, will investigate how AI agents discover scholarly content, authenticate through a user’s entitlements, receive provenance-bearing information, and create a credible record of content use. Bringing together publishers, repositories, infrastructure providers and AI developers, the research will explore where approaches are converging, where challenges remain, and what requirements there are for future implementation to support trusted scholarly communications.
Further information about each of these initiatives will follow as the work develops. Building on the previous posts about the work on AI provenance, here is a bit more detail about the pilot work NISO is launching.
Deep Dive on NISO’s New Provenance Pilot Work
As I’ve previously highlighted, it is widely known that large language models generate fluent, confident text but routinely produce fabricated or misattributed citations. Research has documented hallucinated references at rates of approximately 20% to 40% in general LLM outputs. Even where genuine sources are cited, many systems cannot always demonstrate which training data, which retrieved document, or which detailed model parameter contributed to a specific claim. This then creates legal, ethical, and factual gaps that threaten trust in AI-assisted research, publishing, and enterprise decision-making. How can we trust a result if we can’t see where it was derived from?
Researchers and other users of scholarly content increasingly rely on generative and agentic AI systems to retrieve, summarize, compare, and synthesize authoritative content. A fundamental risk is not just that systems may produce fabricated or inaccurate citations, but that users often lack a verifiable pathway from a generated output back to the specific source object, component, passage, figure, dataset, or standard that supports the claim. This lack of connectivity between research generation and its use creates challenges for verifiability, but also for attribution, recognition, usage assessment, version control, and retractions awareness. These challenges also create barriers to trust in the generated outputs, but also to the trustworthiness of the entire scholarly record. Community-wide efforts to explore signals of trustworthiness, the importance of peer-review vetting, the values of copyright and its protections for attribution, and controls on redistribution and tracking usage for assessment all require a foundation of embedded provenance.
This is why a community-wide effort is necessary to address these challenges systemically. In May 2026, NISO, COUNTER, and Cambridge University Press co-hosted a workshop on these interconnected issues. Technical experts from publishing, the AI tool developer community, and systems librarianship met in Cambridge to scope potential community efforts on these topics. Work driven by COUNTER to improve usage metrics for AI systems is another lens through which to view these issues. NISO is launching a related pilot project to discern a minimum set of metadata sufficient to describe the source of a piece of content used in the inference process of AI systems.
The pilot is therefore focused on the layer where the research publishing and repository communities have practical agency: publisher/repository-controlled source delivery, repository and content APIs, retrieval pipelines, agentic tool calls, metadata packages, signed content objects, and verification services. The pilot will not attempt to solve the full range of problems around AI use of content, such as deterministic source attribution inside foundation-model weights, enforcement issues around copyright, nor issues around access-control management.
The effort will include a twelve-month (ideally the work won’t take this long) pilot to define, test, evaluate, compare, and refine a minimum viable metadata model for provenance and attribution tracking in AI systems used for research and business R&D applications. The pilot concentrates on inference-time provenance: the information that should travel with scholarly and professional source materials when they are retrieved, synthesized, cited, or transformed by generative and agentic AI systems after initial training of the LLMs has occurred. The pilot will convene members of five stakeholder categories (content creators, publishers, repositories, AI tool developers, and end-users) in a coordinated framework designed to produce actionable, evidence-based guidance for further standards development. While specific technological approaches will be determined by willing pilot participants, the goal is to study how provenance functions and is supported in systems such as Retrieval-Augmented Generation (RAG) with citation grounding, inference functions/training data attribution (TDA), output watermarking, statistical provenance testing, and knowledge graph documentation, to ensure provenance data can be consistently preserved across all architectures.
The rationale for the pilot is that existing provenance standards provide important foundations but do not yet fully specify the attribution context needed for AI systems that retrieve and synthesize scholarly and professional content. W3C PROV can describe entities, activities, and agents; C2PA can support signed content credentials and manifests; DOI and related PID systems identify scholarly objects; and usage reporting frameworks can describe measurable access events. The missing layer is a practical, interoperable descriptive profile that connects these resources to AI-generated outputs at the point of retrieval, synthesis, and display through their source metadata. This metadata set will need to be sufficient to describe any inference content, and work across a range of AI system architectures. The goal of the initiative is to test various approaches and, through an iterative process, discern what the minimum amount of data would be to describe an object’s source.
It is important to distinguish the difference between the approach a service provider (e.g., publisher or repository) may deploy to transmit or verify this provenance metadata payload and the publisher content-delivery endpoints and communication protocols. That is a separate and distinct project, and one that should also be specified, developed, and tested to carry the provenance data. The Ithaka S+R research project is scoped to include what types of MCP-style access protocols and mapping of scholarly-specific attributes in MCP-like protocols might be needed. That next phase of work could adapt the results of the research work led by Ithaka S+R and lead to the launch of subsequent work by TRACE in coordination with NISO or possibly other technical groups depending on its scope.
Engaging the entire community
In many ways, the various elements are both part of the TRACE initiative, while also being broader than TRACE. The three current TRACE projects, Ithaka’s research, NISO provenance metadata, and COUNTER’s usage projects, include members of the TRACE community, but also include other community voices as well. Ithaka’s research will be conducted independently, drawing on the experience of publishers, open repositories, platform providers and AI developers. Its findings will be published openly. COUNTER includes members of the library community and other publishers, as well as AI developers. NISO’s provenance initiative currently includes TRACE participating publishers, but also AI tool developers, content repositories, and other content businesses. For example, NISO is beginning conversations with Creative Commons to explore how a provenance metadata package could be incorporated into work on the CC Signals effort and could find wide adoption across different media types.
The reason for this broad view is that these problems are facing not only research publishers, but organizations across the content world. Open repositories need to manage their traffic to prevent them from crashing regularly. Authors and creators of every media type care about provenance and getting recognized for their work. Usage tracking is important not only to provide cost-per-use metrics, but they are significant signals to funders and administrators regardless of whether a site is subscribed or open. These same motivations in our community also drive interest in usage metrics work by advertising-driven media. There is also significant value in aligning work by the research community with wider commercial and web interests, so that we are engaged with related conversations and are more relevant to the companies at the forefront of AI development. We are long past the time when the research community was sufficiently large to build tools and approaches that suited only our own ecosystem but were distinctly different from others. We need to understand the value that building for our community around our core interests in verifiability, attribution, and citation has outside of our own and leverage those needs to push adoption where we might not be able to alone.