Amidst all the uncertainty around the value and the future of using artificial intelligence (AI) in scholarly communications, there is one thing of which I am sure — the people who don’t like AI really don’t like AI and they really want to tell you about it. We’ve been on something of a learning curve with AI in The Scholarly Kitchen (TSK) in recent months and wanted to share what we’ve learned so far and reiterate our current policy.

Several months back, we ran a guest post with an AI-generated image that was submitted by the authors. The feedback came quickly and was largely furious. No one seemed to have any interest in the content of the post, which was largely ignored — all they wanted to do was complain about the “slop” image, which, to be honest, wasn’t particularly extreme or incorrect. I’d guess that most readers wouldn’t have given it a second thought if we hadn’t labeled the image as “AI-generated.” We ended up replacing it with a much less interesting but human-generated image, as we felt the controversy was taking away from the newsworthy content of the post itself. Lesson learned.

Since then, we have received emails from readers — at first maybe once a month, and now multiple emails daily — complaining that a TSK post “seems” like it must have been AI-written, or that they ran it through some random AI “detector,” which proves it was AI-generated.

Illustration of angry human fingers pointing at a robot, meant to represent people blaming or accusing AI

In case you missed it, we worked with Nature News on a story based on one such complaint, which inspired the reporter, Richard Van Noorden, to try to get a better handle on how alleged AI detection tools work.

Essentially, Van Noorden learned that while AI detection tools like Pangram and GPTZero have greatly improved their accuracy, there are still a lot of tools on the market with low-quality results. Perhaps more importantly, even the best tools do not distinguish well between AI-generated and AI-enhanced content. In other words, Pangram and GPTZero will give the same sorts of “AI-generated” labels to content written from scratch by an AI as they do to articles written entirely by a human and then translated to English by an AI or even articles that were human-written and then merely run through an AI for editing and polish. These are very different use cases and represent very different considerations for editorial decision-making.

As the Nature News article puts it:

The inescapable limitation of AI detection, however good it gets, researchers say, is that although software can spot whether AI was involved in a piece, it can’t prove how it was used or judge what’s ethically acceptable.

I know for a fact that I personally edited and added text to many of the posts our readers are accusing of being “100% AI-generated,” as have other members of the TSK editorial team. A TSK Chef recently related running a report they wrote through Pangram — which accurately declared that the report was 9% AI-generated, but misidentified which parts came from AI and which parts were human-written. Clearly, a lot of complaints about potential AI use stem from either using poor analysis tools or misinterpreting the results of those tools.

TSK’s AI Policy

Our policy at The Scholarly Kitchen is that AI tools cannot be credited as the author of a post, and we do not allow AI-generated images to be included in posts unless there is a specific reason for their use (e.g., a post about AI-generated images). In general, I find most AI-generated images to be aesthetically unpleasant, their use takes away from the livelihood of human artists, often perpetuates harmful social biases and injustices, and incurs significant negative environmental consequences.

We do, however, allow the use of AI tools by authors in drafting and editing their posts; same for TSK editors and reviewers. In these cases, we require transparency from authors in documenting their AI use. While this policy has been stated for a while, we didn’t proactively enforce it until recently, relying instead on author discretion as far as disclosure. But, as Avi Staiman accurately noted in a pair of recent posts, authors are generally hesitant to make any such declaration (see also part 2). We are now taking a more direct stance toward disclosure to ensure all authors comply with our policy.

As a community, we’re all grappling with these AI tools and how they should be used in various contexts — what’s okay and what’s not okay. Personally, I see The Scholarly Kitchen as a place where some experimentation should be allowed — given the relatively low stakes of an opinion blog about publishing versus a clinical practice medical journal. As Staiman wrote, if you outright ban the tools, then authors will use them anyway and just lie to you about it. I’d rather we be transparent and honest as we figure things out.

Creating New Biases

One of the big problems with the AI backlash, and all the detection tools that give false positives, is the creation of new biases. Authors’ use of AI is being employed by many readers (well beyond just our publishing blog) as an excuse to dismiss anything with which they disagree. I don’t like what this person is saying, but rather than address and evaluate it, I can just declare it to be “AI,” and then I don’t have to think about it.

We also know that poorly written papers often receive a biased reaction from peer reviewers. Papers that exhibit incorrect spelling and grammar are regularly dismissed in the review process, regardless of the quality of the research results they describe. One of the great promises of AI in our sphere is that it can eliminate that bias for non-native English speakers, and their work can be judged directly on its quality, rather than for the author’s ability to write in a foreign language.

Unfortunately, because so many seem ready (if not eager) to dismiss anything that AI has touched, we have instead introduced a new bias. Now, when a reviewer or reader sees something they think might have some aspect of AI involvement, they immediately dismiss it, just as they would had the piece been poorly written. This is exacerbated by the presence of countless AI detection tools that deliver so many false positives. As with language concerns, this represents an abdication of responsibility on the part of the reader to actually engage with the content.

My Own Concerns

I don’t have easy answers for where this will lead us. AI tools are becoming more commonplace, and I don’t expect them to go away. I personally don’t use them at all and have gone out of my way to disable every AI tool I can find on my phone, laptop, and any online services I use. In some cases, however, AI is unavoidable (e.g., does anyone know how to turn off the AI shopper on Amazon?). As a writer and editor, however, I am not a typical use case.

For me (and for many), writing is thinking. If I want to understand how I feel about something, the easiest route to doing so is to write about it. If I want to explain something to you, the reader, in a clear and convincing manner, that requires me to think it through enough to form a logical argument. And writing is a process that facilitates forming that argument and, with it, my own understanding. Outsourcing writing means outsourcing thinking and that’s not something I’m willing to do.

As far as editing tools, I’ve been writing publicly for more than 20 years, have plenty of good editors/writers I can bounce things off of, and generally feel like I know what I’m doing. I’ve also long lost any fears about being publicly wrong — it’s happened enough to me that I know it’s not the end of the world. I haven’t spent a lot of time experimenting with AI tools because I simply don’t see a need for them in my daily workflows.

I am concerned that the use of AI writing tools will lead to a lot more boring, average writing that all sounds the same. Large language models (LLMs) are word prediction machines — their answers are based on predicting the most likely next word in a sequence. And so, they bring everything to the middle. That’s probably a good thing if you’re a below-average writer and the tool improves the quality of your output. It’s not so good if you’re an above-average writer and it brings you down a few levels, takes away your personal style, and makes you more bland. In writing, as in research, there’s a lot of value in the outliers and the extremes, and forcing everything to the middle of the bell curve means a lot of homogeneity and a lot of missed discovery.

Okay, AI haters, now’s your chance to weigh in below in the comments. I know you have strong feelings about AI, so let’s hear them. And if you’re someone who has seen significant value in the use of AI tools, please do share your results too. We’re all learning here, so the more input, the better.

David Crotty

David Crotty

David Crotty is the Executive Director of Cold Spring Harbor Laboratory Press. Founded in 1933, CSHL Press is an internationally renowned publisher of books, journals, and electronic media, and is a division of Cold Spring Harbor Laboratory, an innovator in life science research and the education of scientists, students, and the public. Previously, David was a Senior Consultant at Clarke & Esposito, a boutique management consulting firm focused on strategic issues related to professional and academic publishing and information services. David was the Editorial Director, Journals Policy for Oxford University Press. He oversaw journal policy across OUP’s journals program, drove technological innovation, and served as an information officer. David acquired and managed a suite of research society-owned journals with OUP, and before that was the Executive Editor for Cold Spring Harbor Laboratory Press, where he created and edited new science books and journals, along with serving as a journal Editor-in-Chief. He has served on the Board of Directors for the STM Association, the Society for Scholarly Publishing and CHOR, Inc., as well as The AAP-PSP Executive Council. David received his PhD in Genetics from Columbia University and did developmental neuroscience research at Caltech before moving from the bench to publishing.

Discussion

Leave a Comment