Tools arriving in the marketplace have a way of reshuffling people’s expectations and norms about what is acceptable. Once calculators became widely available in the 1970s, it led to a robust argument about whether they should be allowed in schools and what an appropriate age might be for their use. Distinguishing whether athletes in professional sports have used of steroids, despite their potential harms, is another recent example. Debate has raged about the use of devices versus printed materials for reading materials in schools. Another long set of debates about the pros and cons of various technologies exists related to film, to radio, and to television in education. Conceptually, these arguments even date back to Plato’s Phaedrus and Socrates thoughts about whether writing itself would have a positive or negative impact.

Enigma Machine gears removed from their case. There are two gears with letters on the outside in black with white letters The dials from a World War 2 'Bombe' checking machine used to crack the German 'Enigma' codes.
The dials from a World War 2 ‘Bombe’ checking machine used to crack the German ‘Enigma’ codes.

As AI use has expanded, it is not surprising a vigorous debate is ongoing about its appropriate use and what guardrails should be employed. Regularly, people decry the apparent use of AI tools in criticizing other’s work. Policies have been deployed by many organizations about acceptable use of AI systems in the authorship process, including COPE, STM, and publishers such as Elsevier and the  National Academies Press, to Springer-Nature, Taylor & Francis, and Wiley among many others. Even this blog has adopted an AI use policy. Now it has become best practice to disclose one’s use of AI systems in your scholarly writing.

AI Tools Begin Disclosing Themselves

The decision whether to discloses AI use is rapidly shifting away from the author. Earlier this month, a clause of the EU AI Act — specifically Article 50 Section 2 — came into force. Regulatory requirements have a way of driving action, even by the largest technology companies. That section requires that, “Providers of AI systems… shall ensure that the outputs of the AI system are marked in a machine-readable format and detectable as artificially generated or manipulated.”  Because of this law, Anthropic released a new AI watermarking system for AI generated content from its system, regardless of its platform and for all of its models released after August 2nd. This applies to both AI-generated text and images. Text is watermarked using a complex token selection algorithm. Information for images and other non-text content is marked with source information based on the C2PA protocol. While identifying AI-generated content can build trust and may potentially be beneficial, this is a case of regulation solving an issue that isn’t the core problem, particularly in scholarly applications. It also creates several other related concerns.

The Challenge of Disclosure

One deeper flaw in accepting any trust markers is the presumption they are definitive statements of fact. Whether an article was peer reviewed is a simple binary question of whether it was or was not. That response does not answer the core question of whether the claims in the article are valid, replicable or trustworthy. The question describes the process the content went through before publication, but not whether that process was rigorous or perfunctory. Similarly, every publication outlet includes comparatively strong works and some others that barely crossed the bar for inclusion. Measures like the Impact Factor provide average quality, based on citation counts, across a period of papers published. The same can be said of every graduating class at a prestigious university, some excelled while others barely passed. We still call all those people doctor, or lawyer, or graduate. Unfortunately, people gravitate toward shorthand rankings and assessments, presuming that averages can be applied to every instance.

Complete AI disclosure of this article would include: Apple voice recognition and speech to text transcription; Microsoft Word AutoCorrect; Microsoft Word Editor; Anthropic Claude for fact-checking and “reading” for critique, Google Gemini search output for a reference search, link inclusion and fact-checking. This is not to say though that AI “wrote” any of the text, in the way we think of a Generative AI system producing content. Even by the certification standards advanced by ProudlyHuman, each of these applications would be considered de minimis use. But if you were being stringent in your rejection of AI tools or pedantic in your reading of Article 50 of the EU AI Act, unless you are working on a manual typewriter or, like George R.R. Martin, who prefers a pre-Internet word processor device, it would be difficult for you to claim AI was not involved in your writing. It’s not even clear whether a red-line version of an edited document, say for grammar or cleaning up citation styles, by an AI tool wouldn’t also include the watermarking.

The challenge is not that disclosure is inherently good or bad, it is that people will make uninformed snap judgments based on these heuristics. We do this the same way we presume every article from a top-ranked journal is accurate and rigorously vetted, or that every graduate of a prestigious university is of the highest skill level. We turn to shorthand signs of quality as proxy for actual vetting. We should certainly build transparency into our ecosystems, particularly as AI systems become more deeply embedded in our social interactions , our creative work, and our business lives. Understanding whether something is a deep fake or not will be critical to our understanding of culture. We should know that bad actors will attempt to circumvent any system built to block them. Unfortunately as well, it will be increasingly difficult to make fine distinctions between where purely human activity ends and machine supported activity begins. The spectrum from ‘purely human’ to ‘purely machine’ will be full of gradients and hard rules will be difficult to apply. If the past is any guide, it will take a very long time to settle on the differences, and a lot of errors will be made along the way.

Solving the Wrong Provenance Problem

Anthropic’s watermark solution solves the wrong provenance problem. Output is the one thing that Anthropic can easily control and track. A bigger provenance concern is not whether the content was produced by a machine, but what human content went into the generation of the machine’s output. We have spent far more energy focused on whether humans are passing off AI-generated content as their own, and not enough on the core question of human attribution. There are number of lawsuits centered around the question of compensation, but not fundamentally centered on recognition. The Bartz vs Anthropic settlement, for example, revolves around copyright violations, infringing use and payment, but not around stipulating the need for attribution of the licensed content in generated outputs. In the settlement, content owners skipped over the challenging technical issues of how Anthropic’s model functions in exchange for a simple payout, leaving the harder technical questions of how models function largely unresolved.

The embedding of Anthropic watermark system will make AI content somewhat easier to identify. I say somewhat because these signals of machine generation — like all watermarking approaches — can be degraded or removed through manipulation of the content after it is generated with the watermark. For example, the C2PA metadata can easily be stripped out by screen capture techniques or several online tools. The steganographic approach can be degraded through editing the text, but one wouldn’t know in advance what editing, or how much rewriting, would be necessary.

The watermarking is brittle in other ways. The content needs to be sufficiently long to function at all, possibly at least hundreds of words, so the solution will only work on some —but not all —outputs (although the specific details aren’t yet public). Knowing this, people might cobble together a variety of outputs across systems to break the encoding.

Research presented at the International Conference on Machine Learning in 2024 showed that it was comparatively easy and inexpensive to defeat watermarking strategies available at the time. Another paper presented the following year at ICML25 showed a similar result that the watermarking could easily be stripped out and at an even cheaper cost. Knowing that editing systems are tracking AI generation though watermarking or other statistical measures of language usage, a number of tools have popped up to mask AI involvement, such as Grammerly’s Humanizer, HumanizeAI, or StealthWriter. There is something of an insane arms race of tools, where some write AI text, others that scan for the AI text, more that rewrite the AI text to avoid detection, and on in an environmental doom loop of AI processing.

The broader usefulness of the watermarking system hinges on intent, which is often indiscernible. We cannot know how the author used the tool or for what purpose. If a non-native English speaker used it to convey their ideas by translating it from their native language, we can all agree this would be an acceptable use. Similarly, using these tools to critique, analyze, or enhance an early draft is similarly a reasonable use of these tools. Without understanding, intent, simply noting the involvement of generative AI tools in the writing process should not be disqualifying on its face. Unfortunately, these distinctions will likely be lost simply by the presence of a binary, Yes/No flag. This is yet another example of building assessment systems that replicate errors of the past.

What is Anthropic Actually Doing?

Anthropic’s system for text is an application of steganography, which is the art of concealing something that is visible in plain sight. This system is not just another simple watermarking approach like those developed in the early days of digital content, such as white text on a white background, or the use of invisible Unicode character strings, both of which can easily be removed with a few lines of python code. Anthropic is instead using a complicated algorithmic system that biases token placement probabilities to embed a string of word elements. These may or may not be the best word choice but they make sense in the context of the query and response. The algorithm biases the word choice to accept words, or parts of words, that fit the purposes of the cryptographic hashing protocol. From a reader’s perspective there is little reason to perceive the subtle difference between the selection of “selection” in this sentence, when one could have used “choice” or “optimal” or “phrasing” or “jargon”. In this structure, the human readers focus on the content and the meaning, while the AI systems focus on the micro-constraints that generated the outputs.

The words themselves are the watermark. It is much like an incredibly complicated application of an acrostic poem. In that style, a pattern is given to the author that they need to follow, such as “begin every sentence in the first two paragraphs with a word that starts in a way that spells out my name”. A reader would have no way of knowing the Scholarly Kitchen post had that embedded meaning without knowing to look for it. They might notice strange phrasing, awkward sentence construction, or perhaps not the best word choice, but the content itself would be coherent and accurate. Since many already accept strange language output from AI systems, this added notice might not even be noticed.

If this watermarking system forces AI tools into specific output choices, rather than the best possible choice, what are we losing in that trade-off and is it worth understanding the content was machine generated. Even knowing that, we’re left with an even more troubling set of nuanced conversations that cover not just whether an AI was involved in the content creation, but why. This is a point that no tool will be able to answer.

Other AI tool providers are also exploring adoption of these approaches. Anthropic has acknowledged that their text-based system is a derivative of SynthID and uses the fundamental Google model. Google launched a project called SynthID last year, which is slowly gaining traction. Earlier this year OpenAI and Nvidia announced they would incorporate it. While it’s functional for images and audio, it hasn’t yet launched this system into production for text content.

Anthropic explicitly claims that statistical watermarking technique “doesn’t change the meaning, quality, or readability of Claude’s responses.” It is hard to see how this is possible if the technique intentionally preferences one output against another, based not on overall quality, but on some other watermarking preference. Anthropic argues that at the scale and pace at which Claude is operating, that these subtleties couldn’t be noticed in the output. Interestingly, the process works better in circumstances where there can be more randomness in word choice.  Unfortunately, this means that the more fact-based and therefore the more deterministic the word choice must be, the more inaccurate the watermarking. This is of particular concern for research publication, since exact word choice can be extremely important. Of course, since the precise statistical approach and embedding algorithm isn’t public, so it’s hard to verify this claim today.

Unintended Consequences of Watermarking

This points to a subtle critique of this approach noted by John Gruber on his Daring Fireball blog where he wrote, “If I ask a tool to generate text, I expect that tool to generate the best possible word choices it can, not corrupt its output for the sake of watermarking.” This almost philosophical argument goes to the heart of what an AI tool is supposed to do. It also highlights the tension between regulatory compliance and product functionality. We see the same issue arise regarding safety and security post-training ‘tuning’.

This process of adding guardrails to systems for safety purpose can have potentially good or bad consequences, depending on your perspective of what the tuning seeks to accomplish. It is what Chinese-based models do to remove references to the Tiananmen Square protests in 1989. Tuning is why Google’s AI image generation model was criticized for inaccurate historical image generation. It is also what created part of the challenge that Hugging Face ran into as it tried to resolve an attack by an AI agent last month. The company was seeking to address the security problems caused by the rogue OpenAI agent that broke containment and began hacking into Hugging Face’s systems. The staff of Hugging Face attempted to use the frontier AI tools it had access to, but couldn’t use them because of security mitigation controls that OpenAI and the other AI companies had built into their frontier models. They instead had to turn to an advanced open-weight model that didn’t have those constraints to solve their problem.

There is another, less obvious, downside to watermarking content: its potential as a surveillance tool. During a recent Intelligent Machines podcast, Jeff Jarvis recalled the case of NSA contractor Reality Winner, who supplied secret documents to a journalist that contained embedded secret watermarks. In June 2017, Winner was tracked down and arrested for leaking a classified report after forensic analysis revealed that the physical document she printed contained invisible, machine-readable tracking dots. Based on that evidence, she pled guilty and was convicted of sharing state secrets, and sentenced to more than five years in prison.

Microsoft had considered embedding all Microsoft Office documents with traceable content identifier to track a documents source back to an original author, or at least authoring device. This service was never broadly deployed, but it can be enabled using Microsoft’s Enterprise compliance tools for corporate customers. Potentially, if AI generated text has embedded watermark information, it might be possible to connect the watermark string to a specific use of the tool. Authorities could then subpoena that information from the AI tool company and trace back the human source, say if a record of the encoded string is retained. To be clear, nothing indicates Anthropic is tracking information that closely, but there is no technical reason this couldn’t be possible. Anthropic, despite its denials, was found to be embedding personalized tracking in its code generation tools. If one cares about social and intellectual freedom, this watermark technology should give one pause.

We need to be clear about what we’re aiming to accomplish with disclosures and content signals around AI tools, what purpose they serve, and exactly how a given solution addresses that concern. As we begin to infer the role of AI tools in the writing process, whether it is through Anthropic’s watermarking, text-analysis tools like Pangram, or the detailed disclosure (or lack thereof) of AI use in the writing process, we need to be careful we don’t fall into the same assessment trap we have so many times before. While implementing solutions, we should work to ensure that we don’t create even bigger problems than the ones we’re aiming to solve.

Todd A Carpenter

Todd A Carpenter

Todd Carpenter is Executive Director of the National Information Standards Organization (NISO). He additionally serves in a number of leadership roles of a variety of organizations, including as Chair of the ISO Technical Subcommittee on Identification & Description (ISO TC46/SC9), founding partner of the Coalition for Seamless Access, Past President of FORCE11, Treasurer of the Book Industry Study Group (BISG), and a Director of the Foundation of the Baltimore County Public Library. He also previously served as Treasurer of SSP.

Discussion

Leave a Comment