Editors’ note: Today’s post is by Chris Reid, Senior Director with Wiley, leading product development and management of the Atypon platform.

Sometime in the last few years, the amount of data the world produces stopped being a number any person can hold in their head. The annual figure is now measured in zettabytes, and the curve that describes it has gone very nearly vertical. It’s very hard to conceptualize the scale of a zettabyte. Simply saying 1021 bytes is frankly unhelpful; a more fun analogy is about 100 million Libraries of Congress, with a full text collection estimated at around 10 terabytes.

For roughly 10 thousand years, human creative output was essentially flat. We then learned to write, invented the printing press, computers, and the internet. Now we enter the AI age, with sensors, instruments, and software that generate data and record without pause. Building on the growth trends in the most-cited industry estimates, the world will produce on the order of 220 zettabytes of content this year. An entire career spent reading full time would cover the tiniest fraction of a percent of that volume.

Decorative image representing machine learning and data acquisition.

In some ways, this is not a new problem. The literature is littered with references to researchers having too much to read, for instance:

“To keep pace with the advance of medical science, by the perusal of the numerous works which are continually proceeding from the press, is a matter of difficulty even for the man of leisure; for the busy practitioner to do so is next to an impossibility. The latter individual, however, is precisely the one to whom a steady and progressive acquaintance with the practical improvements and discoveries of the day is most necessary, as it is he who is the most frequently placed under circumstances requiring a ready fund of therapeutical resources.”

This example comes from the American Journal of Dental Science, back in 1845, when reviewing the latest half-yearly abstract of the medical sciences. With the massive growth in data, this problem will accelerate. Already, millions of articles, preprints, datasets, protocols, and supplementary files appear every year, as shown by Hao, et al., in Nature.

As scholarly professionals, we have spent a long time worrying about access to this literature, but the emerging challenge will be attention. There is simply more relevant material published in any given field than a working researcher can read, and the gap widens every year. At some point the honest description of what happens next has to be blunt: humans will stop being the primary readers of scholarly content. Software, and specifically AI agents, will read the vast majority of it.

For publishers, that can sound like an existential threat. I’ll argue for the opposite: the scholarly communications community wrote down the blueprint for this world more than a decade ago, and we are ready.

We Have Been Building for Machines All Along

In 2016, a large group of researchers, publishers, and funders published the FAIR Guiding Principles, which stated that scholarly outputs should be Findable, Accessible, Interoperable, and Reusable. The principles are now so embedded in funder mandates and infrastructure roadmaps that it is easy to forget how pointed the original argument was. FAIR was not written only for human readers. In fact, its defining move was an insistence on machine actionability: digital objects described and structured so that computational systems can locate, retrieve, combine, and reuse them with minimal human intervention.

The word used was “agents.” The pivotal 2016 Scientific Data paper was clear about computational agents acting on digital objects. Read it again now, with AI assistants and retrieval systems in mind, and it is less of a data-management standard and more like a design brief for today’s emerging world. When an AI agent reads a hundred papers on a researcher’s behalf and returns a synthesis, it is doing precisely what FAIR asked publishers’ content to support.

This is important because it reveals that, while the AI revolution is transformative, it is also the continuation of a consistent narrative within scholarly publishing. This is the scenario the system was explicitly designed to serve. Content that a machine can find, parse, and reuse is more discoverable, more reproducible, and more useful than a PDF that a human must download, squint at, and try to remember. The decade of work behind persistent identifiers, structured metadata, and shared vocabularies was, whether we envisioned it as such or not, preparation for non-human readers.

The Website Was Always Just One Way to Read Content

If the reader is increasingly a machine, the obvious question is what, and how, we serve it. Here, the useful concept is the headless, or multi-headed, content management system (CMS). The idea is straightforward. Hold the content once in a single structured core, with the full text, metadata, identifiers, and entitlements in one place, and then attach multiple delivery surfaces, providing flexibility for those who want to consume the content. In other words, meeting the readers where they are.

One surface is the publisher website, optimized for human browsing and reading. Another is a set of application programming interfaces (APIs) and data feeds, optimized for machines. A newer one serves AI agents directly, through emerging conventions such as the Model Context Protocol (MCP).

What makes this less radical than it sounds is that publishers have been running multi-surface setups for years without naming them. At one extreme, this is known as print. In the digital space, content already flows from the platform to reading surfaces like PubMed Central (PMC), to abstracting and indexing services, to discovery layers, to ResearchGate, to Google Scholar, and to aggregators. Each of those serves readers: a destination consuming the same underlying content, formatted to its own requirements.

Creating a surface built for AI agents is the next entry in a long list, not a departure from how distribution has always worked. The strategic advantage is that content is authored and managed once, then shaped for each audience, and new audiences can be served without re-platforming the core. This means that readers, be they human or machine, access the Version of Record — the published output that the publisher and the editors care most about, and serves as a defined version of the journal’s output.

Agents Do Not  Like Your Website, and Humans Likely Don’t Either

The uncomfortable truth hits hard for all of us who have invested in crafting reading experiences: not only will humans read less of your content directly, but the machines that read it on their behalf do not want your website at all. An AI agent has no use for navigation, cookie banners, related-article carousels, newsletter prompts, or your carefully curated homepage, honed through many hours of painful stakeholder alignment. To an agent, a beautifully designed article page is a small payload of useful content wrapped in a great deal of noise that it must strip away. This is not to say these things are entirely unnecessary. Humans will still read and there will be massive differences in consumption between subject areas and reader contexts. So, visual brand identity will still matter, but all will evolve.

What the agent wants is closer to what FAIR described: clean, structured content, with stable identifiers, explicit licensing, and a predictable way to request exactly the piece it needs. The design question will shift. Alongside the long-running work of making the website better for humans, there is now a parallel, and arguably more urgent question, of what to serve readers who will never load the website. A web page built for human eyes does not suit machine readers. A documented endpoint with structured content and clear terms is the better option.

Publisher websites serve many purposes, providing access to high-quality journal content, acting as signals of trust and authority, and representing the values and community that a publisher and journals serve. However, publishing professionals are largely cognizant that humans don’t love publisher websites any more than machine readers. The number of researchers coming to browse publisher sites, in the expectation of serendipitous discovery, is getting smaller. If humans don’t love publisher websites, then agents certainly will not.

The Debates Worth Having

I should be clear: none of this is settled, and we are likely not yet through the initial phase of this transformation. The most interesting problems are not technical, but more about value and trust. There will be different answers for different groups, and I welcome disagreement and debate from the community. The key here is that this is not just a technology shift. This is a wider shift, with a need for engagement from everyone. Even if AI turns out to be a “normal technology,” it will have broad implications for every aspect of scholarly publishing.

If everything is changing, where shall we look first? I see three major areas: attribution, compensation, and governance.

  1. Attribution: When an agent answers a researcher’s question using your content, does the researcher know it came from you, and is the original work cited in a way that survives the synthesis? A publisher can be the source of an answer and yet be invisible in it.
  2. Compensation: The close cousin of attribution. If machine consumption becomes the dominant mode of access, the licensing and metering frameworks built for human downloads need to evolve, and the standards work on counting and reporting AI-mediated usage is only beginning.
  3. Governance: Deciding which crawlers and agents to allow, on what terms, and how to signal those terms in a machine-readable way is now a board-level question for many publishers, not a task buried in a robots.txt file.

Underneath all three of these is a single strategic concern: in this new world, how do we ensure publishers are not quietly disintermediated, our content absorbed, and brands lost? The infrastructure choices we make in the next few years will determine what the future looks like. We all have agency in determining the scholarly landscape of the future. Now is the time to engage.

Author’s note: AI (Claude) was used to review this post for accuracy, spelling, and grammar.

Chris Reid

Chris Reid

Chris Reid is Senior Director of Product Management at Wiley, working on the Atypon platform, with a focus on the future of the platform, including how discovery and consumption of scholarly content evolves in the AI world. Prior to joining Wiley, Chris was Director of Publishing and Product Development at Science/ AAAS. 

Discussion

Leave a Comment