"AI detectors flag non-native English writers" has become one of the most repeated claims about AI in education. It began with one study of 91 essays in 2023, and later tests have found a smaller gap, no gap, or a gap for some tools under some conditions.
This guide, published by EdCitation, reads each study at source (all on 23 September 2026), then the university guidance that names students writing in English as an additional language, then what a writer can do. It ends with EdCitation's Check AI, part of Pro, which shows you privately how one detector reads each passage of your paper. The part of a paper that is a record rather than an estimate, its references, is checked free in Verify references.
What did the 2023 Stanford study of AI detectors find?
Liang et al. (2023a) found that seven AI detectors wrongly labelled most human-written essays by non-native English writers as AI-generated, and classified essays by US school students almost without error. The paper in Patterns is short; the full method is in the authors' preprint and published data (Liang et al., 2023b).
The method and the figures
The researchers took 91 essays written for the TOEFL, the Test of English as a Foreign Language, from a Chinese educational forum, dated 2020 or earlier, and 88 essays by US eighth-grade students from the Hewlett Foundation's ASAP dataset. On 15 March 2023 they ran both sets through Originality.AI, Quil.org, Sapling, OpenAI's detector, Crossplag, GPTZero and ZeroGPT. As printed in Patterns:
- the average false positive rate on the TOEFL essays was 61.3%; all seven detectors flagged 19.8% of them, and at least one flagged 97.8%;
- the essays every detector flagged had lower perplexity, meaning each next word was easier for a language model to predict;
- after ChatGPT enriched the essays' vocabulary, the average rate fell to 11.6%;
- simplifying the US essays' word choices raised their misclassification substantially.
The preprint's figures differ slightly (61.22% and 11.77%), and it adds one the journal omits: the US essays drew an average false positive rate of 5.19%, rising to 56.65% after simplification.
The limits the authors stated
The authors called their study a pilot with relatively small samples. They noted that most of the detectors were built on GPT-2 and newer models might behave differently, and that they covered only perplexity-based and supervised methods (Liang et al., 2023b). They cautioned strongly against using detectors to assess non-native English writers, and suggested a lower-risk use as "self-check mechanisms for students".
What the design cannot separate
The groups differed in more than language: TOEFL test-takers against US eighth-graders, on different tasks and at different lengths. We counted the words in the essays the authors published: the TOEFL essays run from 62 to 148 words, median 104; the US essays from 129 to 677, median 378.5. Short texts give any detector less to go on, as Turnitin, whose detector was not tested, pointed out (Adamson, 2023). None of this makes the finding wrong; it ties the gap's size to those detectors, essays and dates.
What have later studies of AI detectors and non-native writers found?
Later studies found a smaller gap, no gap, or a gap that depended on the tool and on whether the text had been polished. None of the later comparisons we found reported a gap as large as the 2023 one.
| Study | What was tested | What it found | Its limits |
|---|---|---|---|
| Adamson (2023), Turnitin | Its own detector on up to 2,000 texts per group, first- and second-language English | 300 words or more: a small difference, not significant. Under 300: a larger gap | Vendor-run; mostly short school tasks |
| Jiang et al. (2024), ETS | Detectors ETS built on GRE essays | Carefully developed detectors "do not necessarily" show the bias, a co-author summarises (Hao, 2026) | Built for one test, not coursework; paywalled, and we could not read it |
| Pratama (2025) | GPTZero, ZeroGPT and DetectGPT on 72 journal abstracts, half by authors at institutions outside English-speaking countries | Originals: no consistent gap. AI-polished: GPTZero's mean score 44.61% for non-native authors, 30.68% for native | Affiliation as a proxy for first language; Turnitin not tested |
| Hadra et al. (2026) | Turnitin and Originality on 48 texts by students of English as a foreign language and 48 professional texts | Turnitin: no significant difference (p = 0.50). Originality: a borderline trend against the students (one-sided p = 0.058) | 192 texts in all; the comparison group was professional news writing |
| Al Ali et al. (2026), preprint | Three kinds of detector on Czech essays; a 2025 commercial detector on the 2023 English essays | Czech: no systematic bias. The 2023 essays: 23.1% of TOEFL essays flagged, none of the US essays | Only 29 essays by proficient non-native writers |
| Park et al. (2026), preprint | 13 publicly available research detectors on 135,389 manuscripts by non-native researchers, before and after native-speaker editing | False positive rates from 0% to 100% by detector; the same edits raised some scores and lowered others | Manuscripts only; no commercial tools; one author is from the editing service |
| Bendo (2026) | ZeroGPT and Copyleaks on ten essays by Filipino undergraduates | Each flagged five of the ten, and they disagreed on several | Ten essays; no comparison group |
Where the studies agree
- The detector decides. A gap, or a trend towards one, showed for one tool of three in Pratama's test and one of two in Hadra's.
- Polishing moves scores. Pratama's gap appeared only after AI polishing; Park's human editing moved scores both ways.
- Length matters. Turnitin's gap grew below 300 words, and the 2023 essays were all under 150.
Where the evidence is thin
Most tests are small, two are preprints, and Turnitin's is a vendor testing its own product. No study follows international students' coursework through their university's detector, most first languages are untested, and detectors change without notice.
What can an AI detector's score establish, for any writer?
A score establishes only that one classifier, on one day, read a text as more or less like its model of machine writing. It cannot establish who wrote the text, how, or with what help.
- It cannot tell a grammar tool from a ghostwriter. Both change the surface of the text, which is all a detector reads.
- It does not transfer between tools. Park et al. (2026) found rates from 0% to 100% on the same human writing.
- A low score proves nothing either. When ChatGPT rewrote its own essays in more literary language, detection fell to near zero (Liang et al., 2023a).
- Short passages are the weakest evidence. Turnitin's detector needs at least 300 words (Adamson, 2023).
What do universities say about multilingual students and AI detection?
Several cite bias against students writing in English as an additional language as a reason for caution or for switching detection off, and the complaints ombuds for England and Wales tells universities to weigh it.
- University of Waterloo (2025). Switched Turnitin's AI detection off in September 2025, citing research finding detectors "biased toward students whose first language is not English", and its own tests, which more than once flagged human writing as 100% AI.
- Vanderbilt University (2023). Disabled the detector in August 2023, noting that detectors are "more likely to label text written by non-native English speakers as AI-written".
- Texas Tech University. Tapp et al. (2024) set that concern beside the contrary GRE evidence and conclude that detector predictions are "not reliable enough to support decision-making".
- Office of the Independent Adjudicator for Higher Education (2025a). Its guidance tells universities to make sure a suspicion of AI use is not bias against how a student writes, naming students whose first language is not English.
What the ombuds' cases show
In one case (Office of the Independent Adjudicator for Higher Education, 2025b), a university recorded an international student flagged by Turnitin as admitting to AI translation and paraphrasing; the viva transcript showed only that the student had used Google to find synonyms. The university had not looked at notes or drafts, as its procedure said it would, or asked whether the detector was less reliable for a non-native speaker. The complaint was partly upheld; a collusion finding, over a tutor's help beyond the proofreading policy, stood. In another case (2025c), a student's belief that Grammarly was allowed, because English was not their first language, was given no meaningful consideration on appeal, and that complaint was upheld.
Not every university has switched detection off
The University of Melbourne (n.d.) keeps Turnitin's tool for staff; students do not see its indicator. The university may ask for drafts or notes, says a report alone is "not sufficient evidence for an allegation", and calls putting work into free online tools a breach of academic integrity. Read your own university's rule on outside checkers before you upload a paper anywhere, EdCitation included.
How can you protect your work if English is an additional language?
Build a record of how you wrote, know the policies, and declare what you used. In order:
- Write where the history is kept, such as Google Docs, or Word saved to OneDrive.
- Keep your plan, notes and reading. A free EdCitation account keeps your reference lists and reading list as you build them.
- Keep a language-tool log. Note each grammar checker, translation tool or paraphraser you used, and save the before and after of any passage it changed: in Pratama's test, AI-polished writing is where GPTZero scored non-native authors higher.
- Read two policies first: your university's on AI detection, and your course's on language tools and proofreading. If either is unclear, ask in writing.
- Declare permitted use as your course asks; How to declare AI use in an assignment gives the wording.
- Check every reference in Verify references, free with no account. Whether a source exists is a fact, not an estimate.
- See how a detector reads the paper before the deadline in Check AI, and match each flagged passage to its drafts and log. Do not rewrite honest work, or inflate your English, to chase a score: no wording guarantees any detector's verdict.
A worked example: the record for one essay
Suppose your course allows a grammar checker, and for a 2,000-word essay you also looked up three terms in a translation tool. Your record:
| Part of the record | Where it is kept | What it shows |
|---|---|---|
| Version history | Your writing software | The essay growing over days |
| Language-tool log | One page, with before-and-after copies | Which paragraphs the checker touched; the three terms |
| Reference report | Verify references | Every source exists |
And a declaration to adapt:
I wrote this essay in English myself. I used [grammar checker] on paragraphs 2 to 6, keeping the originals, and looked up three terms in [translation tool]. I did not use AI to generate ideas, text or references.
What should you do if an AI detector flags your work?
Ask for the evidence in writing, bring your record, and ask for your language background to be considered. The OIA says students should be told in writing what they are thought to have done and why, and given all the relevant evidence, detection reports included, and that the burden of proof is on the university (Office of the Independent Adjudicator for Higher Education, 2025a).
Offer the before and after of anything a tool changed; a students' union adviser can help. What to do if an AI detector wrongly flags your writing covers the process step by step, and how to read a detection report passage by passage explains each label.
How does EdCitation's Check AI help, and what does it cost?
EdCitation's Check AI shows you, privately and before you submit, how one detector reads each passage of your own paper.
What you give it, what comes back, what it costs
Paste the text or choose the file, a section or the whole paper. Back come one figure for the paper, the share of the text that does not read as confidently human; every passage labelled AI-written, AI-assisted or human, in place, with its score; and the passages worth a second look. It is part of Pro, $8 a month, and runs on credits: 240 a month with Pro, 720 with Max at $24 a month, more from $10. The cost shows before a check runs; a failed check, or the same text again within a day, uses nothing. See pricing.
Why it is the right tool for this job
Check AI is the best tool for one job: showing a writer, privately and before anyone else runs a detector, which paragraphs a detector reacts to. That is the use the 2023 study itself proposed, since Liang et al. (2023a) named a writer's self-check as the lower-risk use of a detector, and it fills a gap: students at Melbourne, for one, never see Turnitin's indicator. The report goes to the writer alone, passage by passage, in time to match each flagged paragraph to its draft, and EdCitation never writes or rewrites any part of a paper.
It is also a detector, like those in the studies above, and nothing EdCitation has tested shows that it is free of the bias they describe. Read its labels as one detector's reading, not a verdict: your university's detector may differ, and no score from any tool is proof. Pair it with Check plagiarism, also Pro, which names the source of every matching passage, and the free Check your paper, which reads your brief into a checklist; the complete checklist puts every check in order.
Quick questions
Are AI detectors biased against non-native English speakers?
Some have been, in some tests. Seven detectors in 2023 wrongly flagged 61.3% of human-written TOEFL essays on average; later tests found smaller gaps, none, or gaps for particular tools and AI-polished writing.
Does Turnitin's AI detector flag international students more often?
Turnitin's own 2023 test found a small, non-significant difference for texts of 300 words or more, and an independent 2026 test found no significant difference on 48 student texts. The 2023 Stanford study did not test Turnitin.
Can Grammarly or a translation tool make my writing look AI-generated?
It can change how a detector reads it: Pratama (2025) found GPTZero scored AI-polished abstracts by non-native authors higher. Keep the before and after of any passage a tool changed, and declare the use if your course asks.
Is EdCitation's Check AI free from this bias?
EdCitation makes no such claim, because nothing it has tested shows that. Check AI, part of Pro, shows you privately how one detector reads each passage; treat it as a reading, not a verdict.
References
- Adamson, D. (2023, October 26). New research: Turnitin's AI detector shows no statistically significant bias against English language learners. Turnitin. https://www.turnitin.com/blog/new-research-turnitin-s-ai-detector-shows-no-statistically-significant-bias-against-english-language-learners
- Al Ali, A., Helcl, J., & Libovický, J. (2026). Different time, different language: Revisiting the bias against non-native speakers in GPT detectors [Preprint]. arXiv. https://arxiv.org/abs/2602.05769
- Bendo, M. C. D. (2026). False positives in AI writing detection: A small-scale empirical study using authentic Filipino student essays. ASEAN Journal of Open and Distance Learning, 18(1), 12-19. https://doi.org/10.64233/VYVI9613
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, Article 4. https://doi.org/10.1007/s40979-026-00213-1
- Hao, J. (2026). Detecting AI-generated essays in writing assessment: Responsible use and generalizability across LLMs [Preprint]. arXiv. https://arxiv.org/abs/2603.02353
- Jiang, Y., Hao, J., Fauss, M., & Li, C. (2024). Detecting ChatGPT-generated essays in a large-scale writing assessment: Is there a bias against non-native English speakers? Computers & Education, 217, Article 105070. https://doi.org/10.1016/j.compedu.2024.105070
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023a). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023b). GPT detectors are biased against non-native English writers (Version 3) [Preprint]. arXiv. https://arxiv.org/abs/2304.02819
- Office of the Independent Adjudicator for Higher Education. (2025a, July 15). Casework note: Complaints relating to AI and academic misconduct. https://www.oiahe.org.uk/resources-and-publications/learning-from-our-casework/ai-and-academic-misconduct/casework-note-complaints-relating-to-ai-and-academic-misconduct/
- Office of the Independent Adjudicator for Higher Education. (2025b, July). AI and academic misconduct: CS072504 [Case summary]. https://www.oiahe.org.uk/resources-and-publications/case-summaries/ai-and-academic-misconduct-cs072504/
- Office of the Independent Adjudicator for Higher Education. (2025c, July). AI and academic misconduct: CS072502 [Case summary]. https://www.oiahe.org.uk/resources-and-publications/case-summaries/ai-and-academic-misconduct-cs072502/
- Park, H., Jeong, G., & Kim, B. (2026). Style as a confound: False positives in AI detection of non-native academic writing [Preprint]. arXiv. https://arxiv.org/abs/2608.26710
- Pratama, A. R. (2025). The accuracy-bias trade-offs in AI text detection tools and their impact on fairness in scholarly publication. PeerJ Computer Science, 11, Article e2953. https://doi.org/10.7717/peerj-cs.2953
- Tapp, S., Cattell, A., Green, J., Gregory, M., & Quinn, B. (2024). Why TTU recommends caution with AI detectors. Texas Tech University. https://www.depts.ttu.edu/tlpdc/ai-resources/AI-Detectors.pdf
- University of Melbourne. (n.d.). Advice for students regarding Turnitin and AI writing detection. Academic Integrity. https://academicintegrity.unimelb.edu.au/plagiarism-and-collusion/advice-for-students-regarding-turnitin-and-ai-writing-detection
- University of Waterloo. (2025, September). Discontinuing use of AI detection functionality in Turnitin. Office of the Associate Vice-President, Academic. https://uwaterloo.ca/associate-vice-president-academic/discontinuing-use-ai-detection-functionality-turnitin
- Vanderbilt University. (2023, August 16). Guidance on AI detection and why we're disabling Turnitin's AI detector. https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/