Amidst all the uncertainty around the value and the future of using artificial intelligence (AI) in scholarly communications, there is one thing of which I am sure — the people who don’t like AI really don’t like AI and they really want to tell you about it. We’ve been on something of a learning curve with AI in The Scholarly Kitchen (TSK) in recent months and wanted to share what we’ve learned so far and reiterate our current policy.
Several months back, we ran an interview post with an AI-generated image that was submitted by the interview subjects. The feedback came quickly and was largely furious. No one seemed to have any interest in the content of the post, which was largely ignored — all they wanted to do was complain about the “slop” image, which, to be honest, wasn’t particularly extreme or incorrect. I’d guess that most readers wouldn’t have given it a second thought if we hadn’t labeled the image as “AI-generated.” We ended up replacing it with a much less interesting but human-generated image, as we felt the controversy was taking away from the newsworthy content of the post itself. Lesson learned.
Since then, we have received emails from readers — at first maybe once a month, and now multiple emails daily — complaining that a TSK post “seems” like it must have been AI-written, or that they ran it through some random AI “detector,” which proves it was AI-generated.

In case you missed it, we worked with Nature News on a story based on one such complaint, which inspired the reporter, Richard Van Noorden, to try to get a better handle on how alleged AI detection tools work.
Essentially, Van Noorden learned that while AI detection tools like Pangram and GPTZero have greatly improved their accuracy, there are still a lot of tools on the market with low-quality results. Perhaps more importantly, even the best tools do not distinguish well between AI-generated and AI-enhanced content. In other words, Pangram and GPTZero will give the same sorts of “AI-generated” labels to content written from scratch by an AI as they do to articles written entirely by a human and then translated to English by an AI or even articles that were human-written and then merely run through an AI for editing and polish. These are very different use cases and represent very different considerations for editorial decision-making.
As the Nature News article puts it:
The inescapable limitation of AI detection, however good it gets, researchers say, is that although software can spot whether AI was involved in a piece, it can’t prove how it was used or judge what’s ethically acceptable.
I know for a fact that I personally edited and added text to many of the posts our readers are accusing of being “100% AI-generated,” as have other members of the TSK editorial team. A TSK Chef recently related running a report they wrote through Pangram — which accurately declared that the report was 9% AI-generated, but misidentified which parts came from AI and which parts were human-written. Clearly, a lot of complaints about potential AI use stem from either using poor analysis tools or misinterpreting the results of those tools.
TSK’s AI Policy
Our policy at The Scholarly Kitchen is that AI tools cannot be credited as the author of a post, and we do not allow AI-generated images to be included in posts unless there is a specific reason for their use (e.g., a post about AI-generated images). In general, I find most AI-generated images to be aesthetically unpleasant, their use takes away from the livelihood of human artists, often perpetuates harmful social biases and injustices, and incurs significant negative environmental consequences.
We do, however, allow the use of AI tools by authors in drafting and editing their posts; same for TSK editors and reviewers. In these cases, we require transparency from authors in documenting their AI use. While this policy has been stated for a while, we didn’t proactively enforce it until recently, relying instead on author discretion as far as disclosure. But, as Avi Staiman accurately noted in a pair of recent posts, authors are generally hesitant to make any such declaration (see also part 2). We are now taking a more direct stance toward disclosure to ensure all authors comply with our policy.
As a community, we’re all grappling with these AI tools and how they should be used in various contexts — what’s okay and what’s not okay. Personally, I see The Scholarly Kitchen as a place where some experimentation should be allowed — given the relatively low stakes of an opinion blog about publishing versus a clinical practice medical journal. As Staiman wrote, if you outright ban the tools, then authors will use them anyway and just lie to you about it. I’d rather we be transparent and honest as we figure things out.
Creating New Biases
One of the big problems with the AI backlash, and all the detection tools that give false positives, is the creation of new biases. Authors’ use of AI is being employed by many readers (well beyond just our publishing blog) as an excuse to dismiss anything with which they disagree. I don’t like what this person is saying, but rather than address and evaluate it, I can just declare it to be “AI,” and then I don’t have to think about it.
We also know that poorly written papers often receive a biased reaction from peer reviewers. Papers that exhibit incorrect spelling and grammar are regularly dismissed in the review process, regardless of the quality of the research results they describe. One of the great promises of AI in our sphere is that it can eliminate that bias for non-native English speakers, and their work can be judged directly on its quality, rather than for the author’s ability to write in a foreign language.
Unfortunately, because so many seem ready (if not eager) to dismiss anything that AI has touched, we have instead introduced a new bias. Now, when a reviewer or reader sees something they think might have some aspect of AI involvement, they immediately dismiss it, just as they would had the piece been poorly written. This is exacerbated by the presence of countless AI detection tools that deliver so many false positives. As with language concerns, this represents an abdication of responsibility on the part of the reader to actually engage with the content.
My Own Concerns
I don’t have easy answers for where this will lead us. AI tools are becoming more commonplace, and I don’t expect them to go away. I personally don’t use them at all and have gone out of my way to disable every AI tool I can find on my phone, laptop, and any online services I use. In some cases, however, AI is unavoidable (e.g., does anyone know how to turn off the AI shopper on Amazon?). As a writer and editor, however, I am not a typical use case.
For me (and for many), writing is thinking. If I want to understand how I feel about something, the easiest route to doing so is to write about it. If I want to explain something to you, the reader, in a clear and convincing manner, that requires me to think it through enough to form a logical argument. And writing is a process that facilitates forming that argument and, with it, my own understanding. Outsourcing writing means outsourcing thinking and that’s not something I’m willing to do.
As far as editing tools, I’ve been writing publicly for more than 20 years, have plenty of good editors/writers I can bounce things off of, and generally feel like I know what I’m doing. I’ve also long lost any fears about being publicly wrong — it’s happened enough to me that I know it’s not the end of the world. I haven’t spent a lot of time experimenting with AI tools because I simply don’t see a need for them in my daily workflows.
I am concerned that the use of AI writing tools will lead to a lot more boring, average writing that all sounds the same. Large language models (LLMs) are word prediction machines — their answers are based on predicting the most likely next word in a sequence. And so, they bring everything to the middle. That’s probably a good thing if you’re a below-average writer and the tool improves the quality of your output. It’s not so good if you’re an above-average writer and it brings you down a few levels, takes away your personal style, and makes you more bland. In writing, as in research, there’s a lot of value in the outliers and the extremes, and forcing everything to the middle of the bell curve means a lot of homogeneity and a lot of missed discovery.
Okay, AI haters, now’s your chance to weigh in below in the comments. I know you have strong feelings about AI, so let’s hear them. And if you’re someone who has seen significant value in the use of AI tools, please do share your results too. We’re all learning here, so the more input, the better.
Discussion
28 Thoughts on "AI Use in The Scholarly Kitchen"
David, super classy as always. It was important for you to write this post. I appreciate you and your leadership and know that you, Lettie, the Chefs, and our guest authors are doing their best in this wild lawless AI world. Well-said and thank you.
“Writing is thinking” – David Crotty, TSK
“Instead of a period at the end of each sentence, there should be a tiny clock that shows you how long it took you to write that sentence” – Laurie Anderson, Homeland
I use AI multiple times every day. I use it for coding, for making my messages easier to digest (including this one), for testing ideas, and for several other tasks. It would be hypocritical for me to criticise anyone for using AI to improve their blog posts.
At the same time, I cannot bring myself to read a blog post that starts with an AI-generated or AI-inspired title. There were three of those back to back a couple of weeks ago. To me, it signals lack of originality and lack of effort. Why should the reader make an effort to read if the author has not made the effort to write?
Interesting. How can you know that a title, perhaps consisting of 5 to 10 words, was AI-generated? Isn’t that too small a sample size for any reasonable detection tool? Is the title of the article more essential than the content of the article? Would you read an article that you knew was entirely human-generated but that had an AI title? I’m thinking of newspapers where the author of the article doesn’t create the headline, that’s done by someone else.
I do understand the sentiment — if you can’t be bothered to put the work into your writing, then why should I put the work into reading it? But I would also ask, given that you admit that you used AI in your comment here and your messages, why should the reader make an effort to read your messages or your comment? Knowing that you, yourself wouldn’t bother to read your own comment, why bother posting it?
If you use AI on a daily basis, it becomes fairly easy to spot the language.
The title is not more essential than the content. And an AI-generated title does not prove heavy AI usage in the text. But it certainly raises the odds enough for me to decide not to invest the time to read.
I have no issue with AI usage for crafting messages and even blog posts or scholarly articles. But when I visit a venue seeking original ideas, I have an issue with AI taking over because it typically produces bloated synthesis rather original ideas.
I have spent myself a lot of time pushing back on AI tools saying “AI is not intelligent, it’s stupid”, and I was referring to incoherent and/or simplistic answers, and some big blunders – a friend of mine had a ChatGPT conversation where ChatGPT calculated incorrectly an average from 3 numbers, this was really-really bad.
But then I listened to a discussion at Anthropic, and of course you can dismiss it as a marketing pitch, but the two quotes I heard there stuck to me and ring true:
1) “Scientists will not be replaced by AI, but the scientists who do not use AI, will be.”
Maybe for creative writers this is not entirely true, but for editors, journalists?… i leave this quote as a food for thought.
2) “Try again in several months.”
This was directly addressing my “it’s stupid” critique, saying that AI engines improve substantially in months, so the prompt that rendered a blunder may produce something way more useful a couple of months later.
And indeed this matches my experience, as the answers I’m getting at the moment are “less stupid than what I remember”. They are still wordy and waste my time in the sense that I have to skim through this long reply, but save my thinking time for unstructured data and out-of-the-box ideas, instead of wondering how do i do this in Excel or which keywords to put for the literature search.
All in all, I still treat AI as a student, and i say it’s a bad student, because you can’t fully rely on the answer and you always have to check the references etc. But with time passing, I do see potential in AI, maybe indeed one day it will become a good student. And in the process, while critiquing the AI output, maybe I will improve my critical thinking skills.
Wonderful Post! The line between what is actually AI and what is simply continuing automation is a very thin one. I use dictation a lot to get my thoughts out faster. Is that AI? I use spell correction. Are those unethical no they are tools. AI also a tool. It does not replace the human mind.
When people talk about AI, I think they really mean the conversational search interfaces of the large language models, and the reaction is very similar to the hysteria we had about Google only 20 years ago. I fear that we may lose sight of a whole lot of things that I consider much more important like integrity of the articles being written. It may be authors are just afraid they’ll be replaced, which I don’t think will be the case. Their work is what AI feeds on.
We at Access Innovations used to offer product called SciGen that would detect computer generated papers. We detected a lot of them in the back files of publishers. Demand was high after Nature detected 84 in a large corpus. When we ran SciGen against it, we found many more. We have since used it to detect many automatically generated papers created using the MIT software and identified them for removal from historical collections / back files. However, when ChatGPT came out, we took it off the market because we were no longer able to reliably score the likelihood that the paper was automatically generated. The syntax is simply too good to detect.
Since then we’ve introduced instead a bunch of integrity tools to provide the proper name of a gene for example. There are an average of 19 synonyms for each gene name from the NIH human genome database, and trying to gather all of those spellings and words reliably in a search statement is difficult …. makes for a large prompt. Same thing goes for medicinal plant names. There are an average of 16.7 synonyms per medicinal plant according to the Medicinal Plant Names Service (MPNS at Kew. We automated the process of detection to bring them all under the preferred name. Other things like detecting, bad cell lines use which is still rampant are also automated.
While editors fight against the apparition of AI and what it might or might not do, there are massive gaps in providing data integrity to their readers. Peer reviewers are not always likely to connect to the proper naming or to the fact that this cell line was reported as contaminated, etc. they’re reading for the substance of the paper and the research it represents, perhaps without access to sources of information on naming conventions. It would make life easier for overworked editors, reviewers, and authors if publishers added integrity checks to their data. All that incorrect and precise data is being ingested into the LLMs and the reason they are using high-quality published data. Is they believe it to be a ground truth. The liability, then for incorrectly labeled data, still resides, ultimately with the source data from the publisher
I respect AI and believe it to be a continuum of what we’ve been living through for many years. As with other things, we respect we also need to view with some concern. Perhaps the best way to deal with AI is to ensure first that we understand it and second accept it is here to stay and finally that we can come together to make it better rather than fight it.
Thanks giving us a peak behind the screen. The TSK policies sound great and I agree transparency rather than an outright ban is the way to go. I’ll share another favorite quote that I live by.
“I write because I don’t know what I think until I read what I say.” – Flannery O’Connor
Well said David! Like you, I have so far personally refused to use AI for any writing I do, but I’m of the opinion that it is impossible and inappropriate to prevent others from using it to help themselves write.
At the AACR journals we aren’t currently using Pangram to routinely screen manuscripts for use of GenAI by authors, but we do use it to screen all incoming peer review comments. The impetus for this was exactly the problem you describe with people (in our case manuscript authors) using random AI detectors to claim that AI was used to generate a peer review and thus some or all of the criticisms raised by the reviewer don’t need to be addressed. But the responsibility of the journal publisher is to evaluate the quality of the study while maintaining the confidentiality of the editorial process. If GenAI was used to generate legitimate critiques, these should not be ignored.
By screening all incoming reviews, we can spot AI-generated text and take action to communicate important critiques to the authors while also warning or reminding the peer reviewer of the need for confidentiality. Sometimes the reviewer is completely unaware that the tool or service they used to assist with their review relied on GenAI and confidentiality may have therefore been compromised. This is just one reason why in my opinion blanket bans on GenAI use in peer review have the potential to do more harm than good. Better to monitor, communicate, educate, and learn.
This technology is not going away. We need to learn how to use it to our benefit while mitigating the dangers. Perversely, one of the most damaging uses of GenAI we’ve observed so far has been to supposedly improve and/or tone down the language in wholly human-written reviews. What were once useful, clear, and concise critiques using casual scientific jargon that the informed reader could easily interpret became almost useless wishy-washy grammar-laden summaries and pseudo-commentaries.
Maybe you should encourage your readers to embark upon taking Wikipedia’s ‘AI-or-not’ test. How well do you think you can actually spot AI generated content?
We have been seeing some of this bias and accusations against human generated text. Sometimes I wonder if the accusations have overshadowed the appreciation of appropriate word use and grammar. It’s not very easy to spot whether text is generated and we need these detection tools more than ever within scholarly publishing.
Referring to participants in the disagreement as “haters” isn’t very generous, but I accept that isn’t a one-sided problem. Your second-to-last paragraph is where so many of these conversations should begin but far too often don’t address at all. In a scholarly environment, I hope we can become more critically engaged with the theory and functioning of these tools, which we must remember are also largely commercial products with marketing as obfuscating as it is inescapable. The ire directed at this one machine-generated image is disproportionate, sure, but I expect the emotion is the bottled reaction to so many colleagues and vendors who treat these probabilistic models as general reasoning solutions. See the cries of “the future,” “inevitable,” and “already here” heard in meetings, pitches, and conference talks despite knowing, as a profession, that the limitations of text, bibliometrics, and our information environment make certain claims of summarizing and analysis unlikely. We haven’t yet built the universal library, and should we ever, it’s unlikely the raw data format would be word distributions. I don’t know how we address these high emotions as a community, but as a profession, it’s our duty to push back against this tide of treating rather limited tools like LLMs as universal bridges-to-anywhere without requiring an explanation and justification of the process.
To be fair, I would firmly put myself in that category of “haters”. I avoid AI tools as much as possible and have deep concerns over their impact on education, research, creativity, etc., and have been very vocal about such concerns. I do apologize if that word seemed derogatory, but to me, language has evolved to a point where it’s more a descriptive term. Merriam Webster defines it as:
https://www.merriam-webster.com/dictionary/hater
1: a person who hates someone or something
2: informal : a person who actively and aggressively criticizes and disparages something or someone (such as a celebrity or public figure)
Language aside, your points are important ones. The level of hype around AI has certainly distorted our ability to discern its value. Too many billionaires have so much money riding on its acceptance that it’s essentially being forced down our throats and integrated into so many places where it shouldn’t be, or at least where it’s not wanted. I personally think there’s great value in AI, but more in purpose-built tools rather than general LLMs which have dominated the conversation.
Great perspective, article, and comments.
I have a complicated relationship with AI, as I suspect many people in this community do. The lines are so blurry. What if the TSK AI-generated art did not come from an LLM but rather Adobe’s generative fill function? Is that still problematic for those who shared negative feedback? So many people use Canva these days and I question how is that “better” than AI-generated art. I also suspect that many stock graphics have AI behind it that people might not consider.
It’s already been mentioned that AI integrated into our Word, Google Docs, and email applications is ubiquitous and many times I’m thankful for the clarity it provides me as the reader when I know people don’t take the time to carefully consider word choices.
I am deeply concerned about AI—the money around it, environmental impact, cognitive offloading—but the parallels to the advent of the Internet in the late 90s is striking to me.* I feel the same knot in my stomach and conflicting emotions in my brain. Remember those lofty days of so many people feeling the power around the democratization of information!? Likewise, many people I knew were deeply skeptical and refused to jump into the World Wide Web.
We are just at the cusp with all the conflicting thoughts and emotions around AI. With AI usage, I wish we could turn down the vitriol, approach with a little more pragmatism, and listen to the different viewpoints.
Thank you for writing this.
O Superman…
———-
* em dashes my own!
Your point about AI leading to boring writing reminded of this opinion piece in which the author named the process “semantic ablation”:
https://www.theregister.com/software/2026/02/16/semantic-ablation-why-ai-writing-is-boring-and-dangerous/4414930
Thank you for a post that encourages good thinking (and perhaps some writing). When an article is co-written authors sometimes identify the parts they are responsible for. But a good collaboration, in my view, is always qualitatively different than the sume of its parts. Is an AI tool a “collaborator”? I have seen it referred to as a thought partner (it is definitely not that). I am an ardent resister. And I worry about how these tools are consuming us in higher education when I see discussion of creating AI curricula or AI literacy. I am a librarian survivor of information literacy. I still believe in it but I’m still not sure we know what it is. No questions here, just writing so I can think.
At the risk of over simplification, decades ago, there were debates about using calculators and typewriters, because those technologies were dangers to independent and careful thinking. Then came computers with spelling and grammar-checkers. Then LLMs. Each time, its seems humanity moved the goal posts.
I wonder if one should also hold human help to the standard of AI. Publishers don’t normally disclose the names of developmental or copyeditors or reviewers, or which areas of the text they helped change. In other cases, when we read leader interviews, we suspect there was often a group of people who edited or wrote those responses to corporate comms perfection, without attribution. What I’m trying to say is that we humans have a history with questionable attribution, but it remains the responsibility of the author to be 100% responsible for every word written. AI is of course an amplified beast, but our thinking around it feels like early days.. In 24 months or so, general caution seems to have moved from “no-no” to “disclose it”. Now disclosure percentages are low, and seem dead on arrival.
Maybe the recorded interview with no questions shared in advance is the purer thinking, though we won’t know if AI assisted some of the thinking pre-interview. Human or AI or both combined, I wonder if I should just be interested in reading what makes me think. None of the variations are a guarantee.
I’m going to stick my neck out in support of AI-generated images, while diplomatically not pushing back on the SK decision to greatly limit them. Image generation has come on leaps and bounds in the last three years, as Ethan Mollick has shown with his otter images. While there is certainly an “AI image aesthetic” resulting from simple prompts (as there is with text and web design: https://www.newyorker.com/culture/infinite-scroll/the-ai-design-aesthetic-thats-taking-over-the-internet), this is not universally true. I’d encourage everyone to look at the work of tools like Midjourney to see the sheer range of aesthetics that can be generated.
To the point: “often perpetuates harmful social biases and injustices”, while accepting this as true for raw outputs, I’d also make the case that A: image libraries are already packed with these, and B: that AI tools give us the flexibility to create images that actively work against this biases.
Perhaps a bigger challenge is asking what is the difference between AI writing and AI images? If one is okay, and one is not, what is the mark that makes AI image generation so different and unpalatable compared to writing?
To be fair, a human generating an image doesn’t require anywhere near the resources nor create the environmental damage of AI image generation (and it allows for employment for creative individuals). As someone commented above, these things feel like they’re being forced on a public that doesn’t really want/need them, not to mention the abuse they’ve already created through generation of deepfakes and CSAM. At least for our blog’s purposes, I can’t really think of a use case for why we’d prefer AI generated images.
Perhaps a bigger challenge is asking what is the difference between AI writing and AI images? If one is okay, and one is not, what is the mark that makes AI image generation so different and unpalatable compared to writing?
A good question, but maybe a better one is to ask why it feels wrong to have AI-edited posts yet no one seems to blink if an I tool is used in photoshop to edit an image. Perhaps people in our community feel more able to identify the former than the latter?
A good place to test whether you can is here: https://www.realornotquiz.com/ – as covered in this 2025 preprint: https://arxiv.org/html/2507.18640v1
I’ve seen similar experiments with AI speech, and the results are challenging. It is very hard to now tell the difference.
David thanks for a non-boring piece about a topic I find increasingly boring; discussing AI in writing. I like writing with AI. I like making fun of AI writing. I like when AI tries to rewrite something I wrote like Joan Didion meets Bukowski. I like yelling at my AI for sounding like a McKinsey bot and telling it to create its own 2 by 2 to describe how annoying it is being. I like cutting and pasting one line out of a 1500 word robo-essay and dropping it in my final composition like I’m a 90s DJ lifting a horn line off of a record I found in the one dollar bin. All I know is I’m spending no less time on my writing then before, certainly more, and I have no less respect for my reader’s time than before. I say, quit harshin’ my mellow you luddite yahhhpps (try pangramming that one.)
it sounds like you’re having fun, but I’m not sure I really want to read any of that.
Also perhaps worth noting that although the term “Luddite” is now used to describe someone who resists progress, the original Luddites were actually fighting against things like child labor, decreased pay for workers, unsafe work conditions, and the decline in quality caused by automation. I can get behind all of those causes, so please consider generating images of me throwing wooden shoes into the gears of ChatGPT.
Touché!
It seems it will take time for the nuance and sophistication to arise in AI capability – and its use. I hold there needs to be clear disclosure of how/when and what model of AI is used so the reader can judge for themselves whether it might impact the veracity/importance/interest of the given output. So, a mathematician using AI to help prove some conjecture heretofore intractable – fantastic. If an author uses AI to generate text because writing is difficult. No thanks.
It also seems that if it saves time for the initial person/company using AI – but costs the person using the AI output *more* time – then the equation is out and the complaints will continue.
Hi David,
Nice post. I think the hyperlink in the text “we ran a guest post with an AI-generated image” directs to a wrong post as it’s not a “guest post” as you mentioned, it’s an interview.
Thank you, David.
It’s encouraging to read about such a range of views on the use of AI within the SK audience.
Around the scholarly publishing sector it appears, at least to me, that both criticisms on capabilities and the use of LLMs become a new form of political correctess. Of course, generalizing is unjust to some, but signs are enough to say we are establishing an orthodoxy. At least for a blink. I don’t suggest this post tried that. I am actually thatkful to David Crotty for sparking this conversation.
It seems reasonable to me to resist to such orthodoxy.
Meanwhile, the AI adoption might be influenced by locus on control. A group of boffins found that “employees [in knowledge management] with an internal locus of control tend to view organizational AI adoption as a challenge, which subsequently encourages knowledge sharing (motivational process). Conversely, those with an external locus of control perceive organizational AI adoption as a hindrance, ultimately leading to knowledge hiding (strain process).”
https://www.nature.com/articles/s41599-026-06829-5
Why not thinking that using AI, LLMs, or automations of all kind is neither right nor wrong? Hence, is there a point to include statements about using such technologies? Are we becoming zealots by doing so?
It’s the agency we believe we have that could matter more, not the statement.