Academic and research libraries, as well as public libraries, historical societies, and other organizations, hold extensive collections of rare and unique materials, such as archival and manuscript collections. These materials form an essential part of humanity’s cultural and historical record, and they are the vital source material for many humanistic fields. I believe that a second digital transformation is at hand for distinctive collections. How can library leaders establish a strategy for their distinctive collections that is fit for purpose today?

The first digital transformation
Over several decades, libraries have made substantial efforts to generate greater impact with their distinctive collections, even amid resource constraints.
Archives can constitute a useful example. Backlogs in processing archival collections were mitigated by introducing more efficient approaches to the creation of finding aids, resulting in less granular description. Digitization was typically crafted as a separate step, often in a separate department, where some circumstances yielded more expensive item-level metadata while others yielded less expensive metadata with commensurate limitations on discovery. Every processing and digitization initiative helped get more materials more discoverable, and sometimes these approaches were institutionalized and provided with ongoing funding.
To support digital collections of distinctive materials, each library established a repository-based infrastructure. At first these were run locally “on premises.” An investment in local digital systems was essential to provide digital access.
For many important collections, this first digital transformation expanded access dramatically. Even so, today, the vast majority of archival and other distinctive collections have not been digitized — and notwithstanding creativity, innovation, and investment across many professionals and communities, many of these distinctive collections remain largely undiscoverable in modern workflows.
Today’s status quo
The first digital transformation incorporated bold experimentation and the development of models that incorporated the technologies available at the time. In the role that I have held over the past year, I have engaged extensively with academic and research libraries about digital approaches for their distinctive collections. Many libraries and archives today remain primarily situated in this first digital transformation. Looking broadly across institutions, a fairly common set of approaches and paradigms indicate the limitations of this approach relative to today’s possibilities.
First, is that digitized collections are often incorporated within existing discovery workflows — rather than enabled to achieve their full potential. For example, primarily textual archival collections may be scanned without handwriting or text recognition, creating one PDF for every folder, with no description beyond “Folder 17” for that PDF. A researcher who has consulted the finding aid for that collection can now find what they are looking for without physically visiting the archive — so a real benefit for the user of a traditional discovery workflow. But the opportunity to open up new entry points for modern forms of discovery and analysis — both primarily human and increasingly computational — is foregone.
Second, is that the alternative, which I might call thorough digitization, which fully enables current capabilities, remains expensive. For distinctive collections, it is not the actual scanning or digital conversion that is the primary cost driver of digitization; rather costs are high because of manual processes for discoverability and accessibility, such as item-level metadata generation, handwriting or text transcription, and alt-text creation. Google’s book digitization initiative showed that it is possible to reduce the unit cost of scanning at scale, although that required us to think in system-wide terms that aren’t as directly relevant for distinctive collections. The bottleneck for thorough digitization of distinctive collections is cost; or more specifically, the unit economics of exclusively human description and accessibility work.
Third, is that the access model for digital collections is primarily designed for institutional showcasing — rather than generating visibility and usage within research workflows. For example, the vast majority of libraries prioritize an access front-end with strong institutional identity and branding, counting on library discovery services or websearch for discovery. This institutional model reflects how responsibility for tangible collections is structured institutionally, rather than benefiting from digital flexibility. But just because something is digitally available doesn’t mean it is readily discoverable, let alone that it can achieve its potential for impact. Instead, as a result of these institutional silos, digital collections sit largely outside the research workflows of relevant user communities, and related collections across institutions are alienated from one another. Incredibly important distinctive collections remain underutilized.
Fourth, is that many larger libraries have built a complex web of bespoke platforms and systems — one that library leaders increasingly tell me they suspect does not deliver enough value relative to the costs they consume. Some large research libraries run dozens of digital platforms, many to support digital collections activities. Sometimes, a new platform was added because it was considered optimal for a single project (for example, one funded by a grant), and yet it consumes ongoing maintenance resources without contributing to a broader organizational strategy. In other cases, a commitment to controlling one’s own systems, for example through on-premises operation or open source software, limits the ability to generate the benefits of scale or operate across institutions. To be sure, local technical capability can be a source of institutional strength, but it can also become the rationale to preserve architectures whose opportunity costs are not regularly reassessed.
Finally, but perhaps most consequentially, is that digitization is thought of as a different set of responsibilities than collections processing. The work of processing, for example for archival records, includes selection, organization, stabilization, and description, among other activities. Only after materials are processed — at a collection level — are they digitized. While there may be some redundant activity across these two workflows even today, the larger problem is that processing is not leveraging opportunities for discoverability and accessibility that can be created by modern approaches and tools. This problem must be addressed for born-digital collections, which are an increasingly important part of all this work.
The digital collections paradigms status quo sometimes developed through an aggregation of individual project-level choices rather than through an aligned organizational strategy. And indeed, this status quo may have made sense when cost drivers and platform options were limited by the technology of the day.
Taken together, looking across the sector, these paradigms today limit the impact of library-controlled digital collections activities relative to their potential. Many professionals in the field have wanted to make a break with this status quo, and there has been substantial experimentation but not at the kind of organizational scale or structural change that is required for implementation. I suspect this overall status quo has, in turn, contributed to reduced budget and funding decisions for digitization and related activities.
The second digital transformation
Today, of course, the operating model of this organizational status quo is no longer fit for purpose for two primary reasons.
- Digitization itself stands to become an input into collections processing rather than a separate and typically downstream workflow, driving down the unit cost of both and allowing for greater accessibility and discoverability; and
- Digitized collections can be brought into modern systems and exposed for modern research and learning workflows, allowing for greater discovery, access, and impact;
Ultimately, we are moving away from a model that is primarily organized around digital asset management. Instead, we may wish to think in terms of the activation infrastructure for distinctive collections. Rather than comparatively passive institutional digital custody, an activation infrastructure has as its purpose to maximize the impact of distinctive collections. This activation infrastructure incorporates library and archival professionals, technicians and student workers, distinctive collections, processing workflows, digital systems, and consortial and other cross-institutional tooling, among other elements. Foundationally, it is a kind of fly-wheel, benefiting from, and in turn generating, cross-institutional scale.
Leadership engagement
In many cases, archival professionals and other leaders of distinctive and digital collections activity recognize many of the opportunities in this second digital transformation. But investments, systems, and structures have been misaligned. Library leaders have an opportunity to lean in boldly with their organizations to develop this activation infrastructure and thereby maximize the impact of the investment they are making. Here are a series of questions that library leadership teams can consider:
- What is our vision for distinctive collections? What outcomes are we actually trying to produce? Are we prepared to align our organization around those outcomes?
- What would we design differently if today’s technical capabilities existed when our workflows were created? Is the organizational structure and workflow fit for purpose? For example, should digitization and collections processing be brought together over time into a more efficient single movement that might start with scanning or digital conversion rather than physical processing?
- What would we design differently if today’s technical capabilities existed when our platforms were selected or built? Is the platform infrastructure fit for purpose? For example, can we bring our digital collections more effectively into the research workflows — of our own community or more broadly? Does each component of our infrastructure contribute enough additional value to justify the costs of integration, maintenance, support? Can we reduce the complexity of our platform infrastructure without undue tradeoffs and redirect resources to higher value opportunities?
- Where should we prioritize institutional differentiation and where can cross-institutional scale or collaboration generate more efficient operations or improved outcomes? How does our vision, our infrastructure, and our practice connect with other stewardship organizations? Are we committed enough to researcher and learner needs that we can collaborate more radically? What is the role of our consortial partners in supporting this work? How can platforms and processing align to user needs?
- What is the appropriate role of AI — both in helping us do our work and as an analytical engine into which our collections might be contributed? Where can machine assistance change scale and economics while preserving professional judgment, rights management, and organizational agency? Should our collections be included for LLM training and if so on a proprietary basis? Are there ways other than LLM training through which our collections should be incorporated into AI-assisted research models and workflows?
There are many big questions here, several of which libraries can’t hope to answer overnight, nor can leaders do so just in the leadership suite any more so than individuals can without organizational planning. Building a modern activation infrastructure for distinctive collections requires organizational learning and alignment and a new way of thinking and working.
In Sum
In the first digital transformation, libraries digitized selected distinctive collections and made them available online. The second gives them the opportunity to activate them at scale: how they are processed and digitized, made discoverable and available, and ultimately connected and analyzed. For libraries, this second digital transformation presents a newfound opportunity — which some might therefore see as an imperative — to maximize the impact of distinctive collections.
Ongoing interaction with colleagues, friends, and, more recently in addition, AI services, helps me develop and work through my thinking and expression, including for this piece. I wrote this piece and take full responsibility for its content.