For as long as I can remember, every scholarly publishing conference I attended had a session on English as a Second Language (ESL) researchers and how publishers can make it easier for them to publish in respected peer-reviewed academic fora. Prior to 2023, I saw hundreds of peer review reports with the instruction: “Have your manuscript reviewed by a native-speaking writer.” I myself railed against publishers who practiced what I called “linguistic discrimination” due to higher rejection rates among non-Anglo authors.

On its face, Generative AI has rendered these conference sessions and reviewer comments irrelevant overnight by making free language support available on demand.

Hello in many languages written with chalk on blackboard

AI as the Great Language Equalizer

Early surveys indicate that language support is the place researchers have embraced AI most readily. A recent study of non-native English-speaking academics found that 72% used AI to improve grammar and style, 63% for proofreading, and 45% for translation. More than half said they used it specifically to address language barriers.

This is a genuine leap forward and we should not minimize it. Researchers who previously needed to pay for costly professional language editing can now test alternative phrasings, translate a rough draft, make a response to reviewers sound more natural, and spend more time focused on the research itself. For researchers at institutions with more limited resources, that can be a real game changer.

Problem solved.

When Fluent English Becomes Cause for Suspicion

However, the quantum leap in writing and editing capabilities is messier and more nuanced than it first seems. Eliminating the language barrier does not necessarily eliminate the bias behind it. In fact, AI, as it is prone to do, simultaneously creates new problems while fixing old ones. Some of these issues are features of the AI itself while others relate to the manner in which the research community, and editors especially, respond to the use of AI-assisted writing in manuscript submissions.

Submissions by researchers from non-Anglo backgrounds are often treated as suspicious precisely because their work is too polished. We also know that AI detectors can add to this suspicion. A recent study of AI detection in multilingual writing contexts found that detectors can disproportionately misclassify second-language writing because features common in L2 academic prose, including syntactic regularity, lexical repetition, and formulaic phrasing, can resemble the patterns these systems associate with AI-generated text. Researchers writing in an additional language may find themselves caught in a new bind: AI can help them overcome linguistic barriers to publication, while the systems designed to detect AI may treat features of their writing as evidence that they actually used it.

And even when nobody runs a detector, humans develop their own informal tells. A 2025 study examined nearly 80,000 peer reviews from a major computer science conference and found continuing bias against authors from countries where English is less widely spoken. After ChatGPT arrived, some of the grammatical idiosyncrasies reviewers had previously noticed began disappearing. Reviewers simply found new clues. Interviewees described associating phrases common in LLM output with authors from non-English-speaking countries. Generic transitions. Elevated vocabulary. Endless parallel constructions. Lists in nearly every sentence. The sentence that somehow always seems to need an em dash. One too many ‘delves’.

That creates an uncomfortable new version of the old gatekeeping problem. If the English is rough, the author may be judged as less credible (especially when they could have fixed it up with AI!). If the English seems suspiciously polished or stereotypically “AI,” the author may also be judged as less credible.

But these are signals, not evidence. Once editors begin treating a particular writing style as proof of how a manuscript was produced, we risk replacing one form of language discrimination with another. Most importantly, even if we know with certainty that the author used AI to help them write the paper, that tells us nothing about the underlying data, analysis and contribution of the paper.

Fluent Does Not Mean Faithful

Suspicion is only one risk. There is a more substantive one: more fluent English does not mean more accurate English.

A recent study compared ChatGPT, Grammarly, and a professional human copyeditor on manuscripts by Ugandan authors. ChatGPT was an extraordinarily enthusiastic editor, making about three times as many corrections as the human. But only 61% of its readability edits were judged improvements, while 14% made the text worse. It also removed pieces of important information.

In addition, the human editor flagged seven passages where the meaning was unclear and asked the author for clarification. Neither LLM did. That distinction matters enormously in multilingual publishing. A good human editor does not simply make a sentence sound better. Sometimes the most valuable editorial intervention is: “I am not sure what you mean here.”

AI is much more inclined to resolve the sentence for us. The result can be what I think of as “semantic slippage. The sentence becomes clearer and more fluent, but the underlying claim shifts slightly in the process. In scientific writing, slightly can matter a lot.

Table 1: Examples of Semantic Slippage

Version Text
Original “Our findings suggest that the intervention may improve adherence, although the small sample limits generalizability.”
Problematic polished revision “Our findings demonstrate that the intervention improves adherence.”

The danger is precisely that the rewritten sentence sounds so good. A clumsy mistranslation attracts our attention. A beautifully written mistranslation may not.

Translation performance itself is also uneven. A study evaluating AI-generated hospital discharge instructions found that AI performed comparably with professional translation in some areas for Spanish, but consistently worse for Chinese, Vietnamese, and Somali. The lesson is not that AI translation is bad. In some contexts, it is remarkably good. The lesson is that knowing a text was “translated by AI” tells us very little about whether that particular translation should be trusted.

When Everyone Starts Sounding the Same

Even when the meaning survives intact, something else can be lost. AI is not merely helping researchers express their ideas, it is changing what academic writing sounds like.

An analysis in Learned Publishing found dramatic post-2022 increases in a cluster of words associated with AI-generated prose. “Delve” increased almost sixteen-fold compared with its pre-ChatGPT baseline and “underscore” more than twelve-fold.” The authors rightly stress that these words cannot be used to prove AI involvement. More interesting is what their spread says about stylistic convergence.

We may also be losing something more important than stylistic variety: our uncertainty. In an exploratory analysis of more than 1.8 million preprint abstracts, Mahmud Omar found that hedging language declined while stronger “booster” language increased substantially. He is careful to say that this does not establish AI as the cause. Still, his question is an important one: are we losing our “maybes”? Academic writing should not always sound certain. Doubt and uncertainty are not linguistic weakness, but part of the intellectual content itself.

Publishers Need Better Questions

These risks point to a practical problem for publishers: “Did you use AI?” is a nearly useless question. Using AI to fix articles and prepositions is not equivalent to using it to translate a methods section, summarize evidence, draft an argument, or reinterpret results. Recent work on documenting AI use in scholarly publishing makes exactly this distinction, arguing that low-risk language polishing and more substantive AI involvement require different levels of transparency and scrutiny (you can see my proposal for how to address this in a previous Scholarly Kitchen post).

Publishers should spend less energy trying to infer AI use from vocabulary or prose style, and more energy explaining which uses are appropriate, which need verification, and which need meaningful disclosure. We cannot encourage researchers to use AI to overcome the linguistic barriers we created and then punish them when their English improves.

Standing Out by Sounding Like Yourself

I want to conclude with one final irony, and perhaps one opportunity.

For years, researchers who use English as an additional language were told to make their writing sound more like everyone else’s. AI has finally made that dramatically easier. Now that everyone has access to the same linguistic machinery, sounding like everyone else may actually become the problem.

If AI-assisted prose becomes the default, differentiation may increasingly come from using AI more selectively. Perhaps the researcher who preserves an unusual turn of phrase, leaves in a carefully chosen “perhaps,” or simply decides not to polish every sentence into frictionless academic English will ultimately sound more credible, not less (there are even tools to insert errors into your text!).

The leveling of the language playing field is real. The next challenge is preserving the intellectual and linguistic distinctiveness that made leveling it worthwhile in the first place.

Author’s note: ChatGPT 5.6 Sol, developed by OpenAI, was used to help draft this article from the author’s notes, arguments, previous LinkedIn posts, points of emphasis, and identified sources, in order to suggest structure and phrasing, and to identify additional candidate sources. The author reviewed all cited sources, substantially revised the draft, and takes full responsibility for the final text, arguments, and conclusions.

Avi Staiman

Avi Staiman

Avi Staiman is the founder and CEO of Academic Language Experts, a company dedicated to empowering English as an Additional Language authors to elevate their research for publication and bring it to the world. Avi is a core member of CANGARU, where he represents EASE in creating legislation and policy for the responsible use of AI in research. He also is the co-host of the New Books Network 'Scholarly Communication' Podcast.

Discussion

2 Thoughts on "Damned if You Do, Damned if You Don’t: Multilingual Researchers’ Use of AI in Academic Writing"

I have some additional thoughts. Well, maybe a lot of additional thoughts.

While AI can assist with multilingual research, there are other considerations the article points out. Language is so nuanced and such a personal form of expression. You have formal literary language, standards language, technical language, regional dialects…

At what point do you know when to cut through the AI jargon when using it to translate? And certain LLMs simply don’t have enough training data for some languages. It’s like reaching the point where you’re confident enough to tell Duolingo it’s blatantly wrong and report the mistake.

And what about those whose language has different word orders than the language they’re submitting in? Or languages that don’t use direct or indirect articles? Different systems of measurement? Or languages that are written and read right to left?

All good questions Chris…..and I’m not sure I have answers for all of them. I guess what I’d love to see is ESL researchers being able to use the full power LLMs provide them, while minimizing the risks and putting QA processes into place to limit how much they pollute the literature.

However, IMO, the fact that they don’t work well/perfectly/at all in many contexts shouldn’t mean we should limit or ban use entirely. There’s a reason such a large % of ESL researchers use these tools for language (more than any other use!) and we need to listen to why that is….

Leave a Comment