AI detectors

Why AI Detectors Flag ESL Writing as AI

Picture a graduate student who spends three weeks on a literature review. She writes every sentence herself, then runs it through a detector. The verdict comes back: likely AI. She didn’t cheat. Her English is just tidy, and tidy turns out to be the problem.

This isn’t a storm in a teacup. A Stanford team documented the pattern in 2023, and people are still arguing about it. The numbers come from all over the place, from journal papers and vendor blogs to university memos. ESL Humanizer pulled the main ones into a single overview of AI detector false positive statistics, with a source next to each figure.

What Did the Stanford Study on AI Detectors Find?

Weixin Liang, James Zou and their colleagues fed 91 TOEFL essays into seven popular GPT detectors. The essays came from a Chinese education forum. For comparison, the team added 88 essays by US eighth-graders from a public Hewlett Foundation dataset.

The US essays sailed through. The detectors got nearly all of them right.

The TOEFL essays had a rough ride. On average, the detectors labeled about 61% of them as AI-written.

The worst case is even grimmer. All seven detectors agreed on 18 of the 91 essays. And 89 of the 91, or 97.8%, picked up an AI label from at least one tool. So a non-native writer who tries a handful of detectors will almost certainly hear “yep, that’s a bot” from one of them. The journal Patterns published the study in July 2023.

Why Does Simple English Look Like Machine Writing?

Many detectors lean on a measure called perplexity. In plain terms, it scores how easy it is to guess the next word. Chatbots write very guessable text. So do people who work with a smaller English vocabulary and stick to safe sentence patterns. The math can’t tell the two apart.

That same predictability is why so much bot-written copy blurs together, the way AI product descriptions all sound the same. Careful learners of English often write in a similar straight line, and they pay for it.

The researchers then tried a telling experiment. They asked ChatGPT to rewrite the TOEFL essays with fancier vocabulary. The number of essays flagged by every detector fell from 18 to just 1, and the average false positive rate dropped to roughly 12%. In other words, the detectors trusted the machine’s polish more than the students’ own plain prose.

That left the authors with an awkward conclusion. If detectors punish simple writing, non-native writers may feel pushed to use AI tools just to dodge accusations of using AI. Talk about a catch-22.

Are AI Detectors Still Biased Against ESL Writers in 2026?

Vendors say the Stanford finding is out of date.

Turnitin says it tested nearly 2,000 samples from English language learners. It reports a false positive rate of 1.4% for them, against 1.3% for native writers. Each document had to clear its 300-word minimum. The company’s own guidance puts document-level false positives under 1% for papers scoring above 20% AI, and sentence-level ones near 4%.

Pangram says it kept the 91 TOEFL essays out of its training data and saw zero false positives on them.

Keep in mind who’s talking, though. These are the companies’ own figures, and independent checks are thin on the ground. Results also swing with the detector, the model version and the length of the text.

A 2026 preprint muddies the water further. Detectors that rely on how predictable each word is labeled most pre-ChatGPT papers by non-native academic writers as AI. Methods based on likelihood ratios did much better. So asking “do detectors discriminate?” won’t get you a clean yes or no. Some do. Some don’t. It depends on what’s under the hood.

Why Are Universities Turning Off AI Detection Tools?

Several schools have decided the risk isn’t worth it.

Vanderbilt University switched off Turnitin’s AI detector in August 2023. It did the math out loud. Vanderbilt students submitted about 75,000 papers in 2022. Even at Turnitin’s own 1% false positive rate, that works out to around 750 students who could face a false accusation. The university also cited a lack of clarity about how the tool works and evidence that detectors mislabel non-native writing.

OpenAI pulled the plug on its own classifier on July 20, 2023. On its challenge set, the tool caught only 26% of AI-written text and wrongly flagged 9% of human writing. The company admitted the tool wasn’t reliable enough.

Here’s how the main findings stack up:

SourceWhat it testedWhat it found
Liang et al., Patterns (2023)7 detectors, 91 TOEFL essays~61% average false positive rate
Turnitin internal study~2,000 ELL samples1.4% vs 1.3% for native writers
OpenAI (2023)Its own classifierCaught 26% of AI text, flagged 9% of human text
Weber-Wulff et al. (2023)14 toolsNone reached 80% accuracy

Debora Weber-Wulff and her co-authors ran 14 tools against 54 documents. Not one hit 80% accuracy, and only five cleared 70%. The team called the tools neither accurate nor reliable. Accuracy sank even lower once someone edited or paraphrased the text.

Independent reviewers keep landing in the same place. The gap between AI and human writing is real, but a single score can’t measure it with any confidence.

What Can an AI Detector Score Actually Tell You?

Read a detector score as an educated guess with a wide margin of error. It doesn’t prove anything on its own.

What Teachers Should Do When a Detector Flags a Student

A flag works best as the opening line of a conversation, not the closing argument. Ask for drafts and notes. Have the student talk through the argument out loud. Compare the work with earlier writing from the same person. Someone who wrote the paper can usually explain it off the cuff.

Turnitin agrees. Its own guidance says the score shouldn’t serve as the only basis for action against a student.

How Non-Native Writers Can Protect Themselves From False Flags

Writers have the tougher job here, and it’s not a fun one.

Keep a paper trail. Save outlines, notes, and a version history with timestamps. Google Docs and Word both keep one automatically. If a flag shows up, that record is your best defense.

Don’t roughen up or dumb down your writing to look more “human.” That’s a losing game, and it shouldn’t be your burden. Plenty of AI humanizer tools promise to fix scores, but they vary wildly in quality, and some mangle your meaning along the way.

If you want help polishing your own draft, pick an editor that shows every change it makes. That way you approve each tweak, and the voice stays yours. A few habits for how to edit AI-assisted writing and keep your voice go a long way here too.

Will AI Detectors Get Better?

They might. Newer methods already look sharper than the perplexity-based tools from 2023.

Still, the jury is out. Until independent testing catches up with the marketing, a flag should spark questions, not a penalty. A student’s grade, and sometimes their visa or scholarship, can hang on that call. Nobody should lose that over a number they can’t see behind.

Related: Google Wants to Answer Everything. What Happens to the Open Web?

Disclosure: This article was contributed by a guest contributor. The views and opinions expressed are those of the contributor and do not necessarily reflect those of AIInsightsNews. Publication does not constitute an endorsement of any product, service, or company mentioned.

Tags: