Editors’ note: Today’s guest post is by Jason Hu, Director of Research Integrity Engagement, Taylor & Francis; Chair of United2Act Education Working Group; Member of STM Association Task & Finish Group: Responsible Use of Research Content in GenAI; COPE Trustee. Reviewer credits to Chef Alice Meadows.

When Windows Advertise and Science Persuades

If you ever walk along the Blackfriars Bridge in central London, you will see the iconic OXO Tower on the south bank of the River Thames. Its upper floors are marked by a curious sequence of windows: a circle, a cross and another circle. Illuminated at night, they unmistakably spell “O-X-O.”

Photograph of the OXO Tower in London, England
OXO Tower, as seen from the Thames, London. Photo: Wikimedia Commons, CC-BY, https://commons.wikimedia.org/wiki/File:Oxo_Tower_London_June_2016_002.jpg.

This was not just decorative flair. When the building was completed in the late 1920s, overt outdoor advertising along the riverfront was prohibited. The tower’s designers found an ingenious workaround: rather than mounting a billboard, they embedded the brand identity directly into the building’s structure. Technically, the building complied with the rules; substantively, it achieved exactly what the rules were designed to prevent.

It’s often told as a charming example of ingenuity. But it also offers a metaphor for how systems of governance can be navigated, stretched, and even bypassed without ever being formally broken.

The brand behind the tower, OXO, built its reputation on meat extract products widely marketed as nutritious and restorative. These claims can be traced back to the work of the 19th-century chemist Justus von Liebig, who popularized the belief that concentrated meat extract retained much of the nutritional value of the meat itself.

Drawing on earlier work, Liebig seized on protein as the “only true nutrient” responsible for building the body and powering muscular work. As Julia Belluz and Kevin Hall note in their book Food Intelligence, he rapidly promoted the idea, declaring meat to be “the ultimate replenisher.” In 1865, Liebig launched “Extract of Meat,” a concentrated syrup marketed as a restorative “beef tea.” Backed by his scientific authority, the product was enthusiastically endorsed by the scientific and medical communities, even famously by Florence Nightingale.

Magazine insert advertisement of "Extract of Meat," probably between 1890 and 1899, a concentrated syrup marketed as a restorative “beef tea."
Magazine insert advertisement of “Extract of Meat,” probably between 1890-1899 (Wellcome Collection/Public Domain Work).

Yet these claims rested more on plausible theory than scientific evidence.

In the same year, Adolf Fick and Johannes Wislicenus set out to test them. The two climbed the Faulhorn in the Swiss Alps on a diet stripped of protein, carefully measuring their energy expenditure and nitrogen excretion along the way. Their findings pointed the other way: it was fats and carbohydrates, not protein, that fueled muscular work.

Later studies reinforced this conclusion. Meat extract, it turned out, contained negligible protein and offered little nutritional value. Liebig had never rigorously tested his hypothesis, and when challenged, he defended it aggressively rather than revising it.

Despite mounting evidence to the contrary, Liebig’s ideas persisted, shaping research agendas, dietary beliefs, and food systems for decades. His product remained commercially successful, and the broader association between meat, protein, and strength endured, echoed today in global food culture and the vast market for protein supplements. More than a century later, an idea built on weak evidence continues to shape how we eat, and how we think.

From Persistent Error to Systemic Risk

Liebig’s case illustrates a broader property of knowledge systems: once ideas become embedded across research, practice, and communication, correction is often slow, uneven, and incomplete.

That latency, the gap between error creation and error correction, has always mattered in science.  For much of modern scholarly history, the primary integrity concern has been the presence of flawed or fraudulent individual studies within the literature. The working assumption was that science is self-correcting: errors might enter the record, but replication, critique, and accumulation of evidence would eventually identify and displace them.

That assumption depends, however, on a critical condition: that the rate of error correction keeps pace with the rate of error creation. Historically, errors entered the literature, but diffusion was slow. Reviews took time to compile. Textbooks were updated gradually. The inertia of the system, while frustrating in many contexts, acted as a form of containment.

But that containment no longer works.

From Error Propagation to Synthetic Consensus

Today’s challenge is no longer confined to flawed individual studies. It concerns the integrity of the knowledge system built upon them. Errors now propagate through that system faster than corrections can follow, and they increasingly do so through mechanisms that make the resulting claims look more credible, not less, as they spread.

This condition can be described as synthetic consensus: a condition in which the perceived credibility of a claim arises primarily from its repetition, aggregation, and recursive reproduction across knowledge systems, rather than from independent empirical validation.

It differs from traditional forms of scholarly consensus in one critical respect: the appearance of agreement is system-generated rather than evidence-earned.

A claim enters synthetic consensus when three conditions are met:

  • Propagation without verification: it is widely cited, summarized, or reused without re-evaluation of underlying evidence.
  • Cross-system embedding: it appears across multiple layers (papers, reviews, textbooks, databases, AI systems).
  • Recursive reinforcement: each reuse increases its apparent legitimacy, regardless of evidentiary strength.

Synthetic consensus is not the same as citation bias or herd behavior. Those describe human tendencies within the system. Synthetic consensus is a system-level property, where the infrastructure itself produces the appearance of agreement.

Three Forces Driving Synthetic Consensus

Three interacting dynamics drive the process: persistence, supply, and recursion.

Persistence. Consider a randomized trial of omega-3 supplementation in COPD, retracted in 2008 after data falsification was confirmed. Yet Schneider et al., show that, 11 years later, the trial was still being cited positively in more than 100 papers, reviews, textbooks, and clinical nutrition discussions, almost none of which acknowledged the retraction. Its claims had diffused into thousands of second-generation citing articles. Once a finding is absorbed into the literature, retraction rarely travels with it.

Supply. Paper mills have intensified the problem at its source. Analysis from Richardson et al., shows they are industrialized, resilient, and expanding, with output growing exponentially over the past decade and thousands of suspect articles entering the literature each year. The rate at which potentially flawed works enter the record now structurally outpaces the capacity of editors, publishers, and research institutions to detect and retract them.

Recursion. If paper mills alter the volume of flawed inputs, generative AI fundamentally alters the nature, not merely the magnitude, of this threat. Large language models can synthesize and reproduce existing literature at negligible marginal cost, including claims that are flawed, unsupported, or entirely fabricated. Critically, they can do so recursively. At each pass, the claim’s original uncertainty is diluted, the appearance of corroboration strengthens, and apparent coherence amplifies.  Errors are no longer only persistent; they are compounded each time they pass through a generative system.

Two findings help illuminate what recursion does to a knowledge system. Shumailov et al., show that “indiscriminate use of model-generated content in training causes irreversible defects in the resulting models, in which tails of the original content distribution disappear,” an effect they term “model collapse.” Traberg et al., describe a related consequence: AI systems are increasingly used to generate ideas, synthesize literature, and frame research questions. That accelerates the production and visibility of AI-related research, while driving convergence in topics, methods, and languages. The outcome, they argue, is a “scientific monoculture.”

Together, these dynamics form a pipeline: flawed claims persist, proliferate, and are recursively normalized. With each iteration, the claim appears more established, even as its empirical foundation remains unchanged.

Synthetic consensus is the endpoint of this pipeline.

Why Existing Mechanisms Are No Longer Enough

Traditional mechanisms of research integrity, peer review, editorial oversight, and post-publication correction remain necessary, but they were designed for a different problem: discrete failures within a largely human-mediated system, on the assumption that the literature was the knowledge system. That assumption no longer holds.

Each of the three dynamics identified earlier exposes a limitation.

  • Persistence defeats correction at the reception end: retraction notices do not travel with citations, and once a claim has been absorbed downstream, flagging the original PDF changes little.
  • Supply overwhelms detection at the input end: no combination of editorial scrutiny and post-publication forensics scales to the volume that industrialized paper mills now produce.
  • Recursion evades both ends at once: generative systems ingest, transform, and redistribute claims through channels that no publisher controls and no retraction mechanism currently reaches.

The risk is therefore no longer limited to what is published. It extends to what is subsequently known. Corrections, if they come at all, must now travel not only through the original publication record, but through an expanding network of downstream representations: databases, meta-analyses, educational materials, decision tools, and algorithmic models. The integrity question has moved upstream, into the citation trails an article rests on, and downstream, into the systems that will carry its claims forward.

Building the Propagation Layer

If the problem is systemic, the response must be too.

Some of that response is now emerging, and I should note that I am directly involved in two of the initiatives I describe here, which shapes both my optimism about what is possible and my candor about what is not yet in place.

United2Act brings publishers, funders, institutions, researchers, and infrastructure providers together to coordinate responses to paper mills and systematic manipulation. The STM Association’s Responsible Use of Research Content in Generative AI pushes for cross-sector alignment between publishers, technology providers, and the research community, grounded in the principle that trustworthy AI systems depend on trustworthy inputs, which in turn depend on attribution, transparency, version control, and the integration of corrections and retractions into downstream systems. Both efforts rest on the same recognition: no single actor can address the problem in isolation, and the scale of the threat demands shared standards and shared accountability.

These initiatives are necessary. They are not yet sufficient. The building blocks of correction infrastructure exist — Crossmark flags, the Retraction Watch database, NISO’s recommendations on communicating retractions — but adoption is patchy, and standards are fragmented. Most critically, downstream systems — particularly AI models — remain largely outside the correction loop. As a result,  a retraction often remains, in practice, a flag on a PDF. Its downstream representations continue to circulate with their evidentiary weight intact.

This is why the paramount challenge has shifted from mere detection to effective propagation: ensuring that corrections travel downstream as effectively as the claims they amend.

Building this propagation layer is unglamorous work. It requires coordination, standardization, and sustained collective effort. But without it, the system will continue to amplify what it cannot reliably correct.

What We Are Embedding

The lesson from Liebig’s story is not that science fails. It is that science can take a long time to correct itself, and that during that time, errors can shape reality. Protein supplements line the shelves of every gym in the world partly because a 19th-century chemist was more persuasive than he was rigorous.

Today, the gap between error creation and correction is shrinking in one sense and widening in another. Claims move faster than ever; errors, once embedded in training corpora, retrieval systems, and downstream tools, prove far harder to unwind than any correction notice can currently reach.

Walk across Blackfriars Bridge tonight, and the OXO Tower still spells its name into the London sky, long after the rules it once navigated have changed. What is built into a structure tends to endure.

The question for the scholarly ecosystem is what we are embedding into systems that will carry knowledge forward, and whether those systems will still be able to remember how to correct themselves.

Author note: This post was written in a personal capacity; views are my own. Claude Opus 4.6 and ChatGPT-5.4 mini were used to improve the language and readability of the article.

Jason Hu

Jason Hu

Jason Hu is Director of Research Integrity Engagement at Taylor & Francis, where he leads cross-sector collaboration on research integrity. He is an elected Trustee of COPE, having served on its Council since 2016. He chairs the United2Act Education Working Group and is a member of the STM working group on the responsible use of research content in generative AI. He has previously held roles at ORCID and Wiley, and writes and speaks regularly on research integrity, AI, and the changing shape of scholarly communication.

Discussion

12 Thoughts on "Guest Post — When Errors Become Consensus: Science’s Self-Correction Can No Longer Keep Up"

Yes, “The lesson from Liebig’s story is not that science fails. It is that science can take a long time to correct itself, and that during that time, errors can shape reality.” And in Liebig’s time, Gregor Mendel discovered (1865) what we now refer to as “genes.” But the distraction of the Darwinian revolution meant that there was little trace of it in the literature until 1900. Had AI systems existed the event would have fallen below their radar. Today, Genetic textbooks tell of how three European botanists brought it to light, often neglecting to mention William Bateson’s 6 year battle with Pearson (famed statistician) before it could gain general attention. As coauthor of Bateson’s biography (2nd edition 2022) I know this well. Thank you Jason Hu for a very timely exposition of the issues at stake. We dismiss the “synthetic consensus” at our peril.

Thanks Donald. this is the sharper version of the problem.
My piece worries about errors that persist. Mendel is the mirror case: correct, published, but effectively invisible for thirty-five years.
I will try to get a copy of Bateson’s biography, thank you for pointing me to it.

My pleasure! The biography expands on “errors that persist.” Backed by the authorities of Theodosius Dobzhansky and Ernst Mayr, who convinced the world that they knew the history, the ground Bateson had prepared for furthering our understanding of speciation remained “effectively invisible” for much of the twentieth century. In the 1960s two young historians of science (William B. Provine and Mark B. Adams) were sold on it. For both, around 1990 there was an awakening. Admitting this, Adams wrote a paper in French, with later an expanded English version. And Provine in 2001 rejected the status quo long embedded in his popular 1971 book.

Great piece Jason, thank you. Your thesis that we’re moving from dealing with distinct individual errors to a systemic issue requiring fundamentally different solutions certainly resonates with me. Though I couldn’t help but wonder if there was some nominative determinism at work in von Liebig’s activities!

thanks Rob, glad it resonated.
my first draft was actually titled “Did Dr Liebig Tell a Big Lie?”. I was talked out of it rightly, not least because the answer is no. He was wrong rather than dishonest, which is the more uncomfortable half of the story, since the system amplifies both the same way.

Thank you Jason, this is a through-provoking piece. It makes me wonder. What if Liebig’s claim was sticky precisely because it was being rewarded? Backed by his authority, the product sold, and the claim and the incentive reinforced each other.

If so, the systemic scholarly challenges you describe may persist for a similar reason. Not merely because correction is slow, but because the incentives across authors, funders, institutions, and publishers quietly favor the amplified status quo over the cost of addressing root causes. Each stakeholder optimizes locally and rationally, with the dysfunction built into the structure, not chosen by anyone in it. It may raise an uncomfortable question. Is the scholarly ecosystem, with the tensions you point out, itself an OXO tower of sorts?

Thanks Neil, you have taken the metaphor somewhere better than I did.
Nobody broke the rules at the OXO Tower. Every decision was rational within its own constraints, and the result still defeated what the rules existed for.
The part I would add is that correction is a public good. Everyone benefits from a corrected record, almost nobody is rewarded for producing the correction. Volume, in the meanwhile, is rewarded everywhere…
So, probably yes. My only defence of the piece is that propagation is the layer where incentives are least opposed to doing the right thing, so it is where something can actually move. Though I take the risk you are implying: that treating the symptom well can make the underlying arrangment easier to tolerate.

agreed, and that is rather my point. Consensus is a proxy for verification, useful only while it tracks the thing it stands for, and my worry is that it can now be produced without it. Worth adding the symmetric caution, though: dissent is not science either. The test in both directions is independent verification.

This is a very thoughtful and helpful article. From my own work, I can see how much AI-supported research is starting to crowd out genuinely novel research, especially in the analysis of open health datasets like NHANES or CDC WONDER. These research areas are seeing problems and I think they fit well with your three-part analysis of ‘synthetic consensus’. This is a different form of model collapse (publication inflation, the dilution of the literature, and the loss of signal to noise) but it’s still another example of what you are writing about in this article.

Science will still be self-correcting in the long run, I think. But as Keynes said, in the long run we’re all dead.

Thank you Matt, and your PLOS Biology paper with Suchak and colleagues is the sharpest illustration of this I know.
The Keynes line is the whole argument. Self-correction that runs slower than the decisions it ought to inform is not much of a safeguard, especially in health field. Liebig’s took a century.

Leave a Comment