# How AI detectors work and how accurate they are

**How do AI detectors work, and how accurate are they?** AI detectors work as statistical classifiers. They estimate how much a text resembles what a language model would write, and they do not observe who wrote it. In independent tests, the best of 14 tools scored under 80%, machine-paraphrased AI text was caught about a quarter of the time, and human essays by non-native English writers were flagged 61% of the time on average.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/how-ai-detectors-work-and-how-accurate-they-are

An AI detector returns a number, 62% AI, or a band of highlighted sentences. It is easy to read that as a finding about who wrote the text. It is not. It is an estimate of how much the text resembles what a language model tends to write, made by a program that never saw the writer.

This guide explains how AI detectors work, then sets the independent accuracy figures beside the vendors' own. The two disagree, and the gap is the most useful thing to know. It ends with how to read a detector's output on your own paper, using EdCitation's [Check AI](https://edcitation.com/tools/check-ai), and with the one check on a paper that is not an estimate at all: whether its references exist, which [Verify references](https://edcitation.com/verify-references) looks up in the publisher's record.

## How does an AI detector decide that a text is AI-written?

An AI detector is a classifier: a statistical model trained on samples of human and machine writing, which learns the patterns that separate the two and scores a new text by how far it leans towards the machine side. The output is a probability and nothing more. The detector has no record of drafts, keystrokes or the writer; it judges only the surface of the text.

### Perplexity and burstiness, plainly

Early detectors leaned on two measures. Perplexity is how surprised a language model is by each next word: Liang et al. (2023) describe it as how "surprised" or "confused" a model is when guessing the next word. A model chooses likely words, so its own output has low perplexity, and a detector reads low perplexity as a machine signal. Burstiness is how much that perplexity varies across a document. Human writing swings between plain and surprising sentences; machine writing stays level.

Two cautions. GPTZero, which made the terms familiar, says that since autumn 2023 it no longer relies on them, having moved to a deep-learning classifier in which they are one of seven indicators (GPTZero, 2023). And predictable prose is not machine prose: formal academic writing, set phrases, and writing in a second language with a smaller vocabulary are all predictable, and all score as machine-like.

### What trained classifiers do instead

Most current detectors are trained classifiers, described by their makers only in outline. Turnitin reports a document score and a sentence score, and in May 2023 raised its minimum to 300 words (Chechitelli, 2023b); when it finds AI writing in 1% to 19% of a document it shows an asterisk instead of a number and highlights nothing, because false positives were commoner in that band (Turnitin, n.d.). A classifier is only as good as its training data: Perkins et al. (2024) note that OpenAI's classifier was not trained on writing by non-native English speakers.

## What is AI watermarking, and does it change anything?

A watermark is the one method that does not guess: the model that generates the text leaves a statistical signature as it writes. Google's SynthID for text adjusts the probabilities of candidate words during generation, in a pattern a reader cannot see and a matching test can. Dathathri et al. (2024) describe it in *Nature*, report a live experiment over nearly 20 million Gemini responses, and state that it is applied in the Gemini app.

Its limits are on Google's own developer page. The watermark is less effective on factual responses, where there is little room to vary the wording, and its confidence can be greatly reduced when the text is thoroughly rewritten or translated (Google, n.d.). The *Nature* paper adds that such watermarks are weakened by edits, including paraphrasing by another model. A watermark cannot say that a text is human, and a university's detector is not a watermark check.

## Why does paraphrasing defeat most detectors?

Because detectors measure the surface of a text, and paraphrasing changes the surface while leaving the content alone. Weber-Wulff et al. (2023) ran AI-written essays through Quillbot on its default settings: accuracy on those documents fell to 26%, against 96% on ordinary human writing, and on average 71.4% went undetected. Editing by hand, swapping synonyms and reordering clauses, brought accuracy to about 50%.

Perkins et al. (2024) tested six ways of disguising AI text, applied by the chatbots that wrote it. Adding spelling errors left the detectors with 12.9% accuracy, more varied sentence lengths 15.9%, paraphrasing 18.4%. Asking for more complex prose barely helped, a 2-point drop. Turnitin fell furthest, from 50% to 7.9%. Liang et al. (2023) found the simplest version: asking ChatGPT to rewrite its own essays in more literary language pushed detection to near zero.

### Mixed texts are the hard case

The commonest real case is a mixture, and detectors do worst there. Hadra et al. (2026) built 48 texts that were half student writing and half AI writing and ran them through Turnitin and Originality: both performed poorly, and Originality's recall on them was near zero. Turnitin's own sentence-level false positives are commoner at the transitions between human and AI writing (Chechitelli, 2023c).

## How accurate are AI detectors in independent tests?

Less accurate than their vendors say, and worst on edited, mixed and non-native writing.

| Study | Tested | Found |
| --- | --- | --- |
| Weber-Wulff et al. (2023) | 14 tools, 54 documents, March to May 2023 | No tool reached 80% accuracy; five passed 70%; Turnitin first. Accuracy fell by 20% on machine-translated human writing and to 26% on paraphrased AI text. |
| Liang et al. (2023) | 7 detectors, 91 TOEFL essays, 88 US eighth-grade essays, all human-written | US essays near-perfectly classified. TOEFL essays: false positive rate 61.3% on average, 19.8% flagged by every detector, 97.8% by at least one; 11.6% after a rewrite with richer vocabulary. |
| Elkhatat et al. (2023) | 5 tools, 15 GPT-3.5 paragraphs, 15 GPT-4, 5 human | More accurate on GPT-3.5 than GPT-4. OpenAI's classifier: 100% sensitivity on GPT-3.5, 0% specificity. GPTZero: 93% and 80%. |
| Perkins et al. (2024) | 7 detectors (the abstract says six), 805 tests, September to October 2023 | Mean accuracy 39.5% on unaltered AI text, 67% on human text, 22.2% after evasion. Copyleaks caught the most AI text and flagged 50% of human samples. |
| Hadra et al. (2026) | Turnitin and Originality, 192 texts, January to May 2025 | Accuracy 0.61 and 0.69; both poor on half-and-half texts. Turnitin 0.86 on humanities, 0.51 on science; Originality 0.96 and 0.58. |
| Van Vlasselaer et al. (2026) | 4 tools, 160 documents: human, AI, hybrid, humanised AI | All four identified the fully human texts; false positives rare. Three underestimated AI content, most of all from the newest model; one did markedly better. |

### What the two error rates mean

A false positive accuses a human. Weber-Wulff et al. (2023) counted a false accusation ratio of 2.4% on human-written English across 14 tools, 11.1% on machine-translated human writing, and 50% for one tool, GPT Zero. A false negative misses AI text, and it was the commoner error: the authors called the tools "neither accurate nor reliable", with a lean towards calling text human-written.

The two are traded against each other. In Perkins et al. (2024), Turnitin made no false accusations and missed 84% of the manipulated AI text; Copyleaks caught the most AI text and accused half the human writers. Neither setting makes the score a fact about authorship. The tests are also small, and run months apart on tools that change without notice: a 2023 figure is evidence about 2023.

## What do the vendors claim, and where is the disagreement?

The vendors publish low error rates measured on their own data; independent tests measure higher ones on theirs. Every vendor figure below was read on the vendor's own pages on 22 September 2026.

**Turnitin.** Its document-level false positive rate is under 1% for documents with 20% or more AI writing, validated, it says, on 800,000 papers written before ChatGPT; its sentence-level rate is around 4% (Chechitelli, 2023b, 2023c). It says the rate "is not zero" and that a flagged sentence is there "to initiate a conversation, not to draw a conclusion" (Chechitelli, 2023a, 2023c). Against that, Hadra et al. (2026) measured Turnitin's accuracy at 0.61 across human, AI and mixed texts, and Perkins et al. (2024) found it missed 84% of manipulated AI text.

**GPTZero.** Its home page states 99% accuracy; its technology page says it holds its false positive rate at no more than 1%, with 1.1% on TOEFL texts and 96.5% accuracy on mixed documents (GPTZero, n.d.-a, n.d.-b). Independent tests found a 50% false accusation ratio in Weber-Wulff et al. (2023), a 10% ratio and the lowest baseline accuracy of seven tools in Perkins et al. (2024), and 80% specificity in Elkhatat et al. (2023). The tool has changed since those tests.

**OpenAI.** OpenAI withdrew its classifier in July 2023 because it could not detect AI output accurately (Perkins et al., 2024). As Elkhatat et al. (2023) report from OpenAI's own tests, it caught 26% of AI text and labelled 9% of human text as AI. OpenAI's page could not be read on the day this guide was written, so those figures are second-hand.

A vendor's rate describes the test set the vendor chose; an independent rate describes a small sample on a date. What both agree on, in the vendors' own words, is that the score is not proof.

## Why have some universities switched detection off?

Because the error rate, at their scale, meant a steady flow of wrong flags. Vanderbilt University disabled Turnitin's detector in August 2023 and did the arithmetic on Turnitin's own claim: 75,000 papers submitted in 2022 at a 1% false positive rate would mean around 750 wrongly labelled (Vanderbilt University, 2023). The University of Waterloo discontinued the feature in September 2025, citing research on unreliability and bias against students whose first language is not English, and its own tests, in which human-written text was flagged as 100% AI-generated more than once (University of Waterloo, 2025). Curtin University disabled it across all campuses from 1 January 2026, with text matching staying on (Curtin University, 2025).

Many universities still run detection, and your institution's policy governs your case. If you are flagged, [What to do if an AI detector wrongly flags your writing](https://edcitation.com/newsletter/ai-detector-false-positive-what-to-do) walks through the process and the evidence that answers a score.

## How do I see how a detector reads my paper before I submit?

Run one yourself, and read the result passage by passage, because that is where the errors live. Turnitin's own figures put sentence-level false positives at around 4% and commonest where human and AI writing meet (Chechitelli, 2023c), and Hadra et al. (2026) found both commercial detectors at their worst on mixed texts. EdCitation's Check AI works at that level: every passage is labelled AI-written, AI-assisted or human, the score for each label sits next to it, and the report goes to nobody but the person who ran it. A marker's report gives a document number after the deadline; a passage-level view before it shows which paragraphs a classifier reads as machine-like, so the drafts and notes for those paragraphs can be ready. It is a Pro tool ($8 a month, with 240 credits); a check tells you its credit cost up front, and a failed one uses none. See [pricing](https://edcitation.com/pricing).

EdCitation says plainly that neither its AI check nor its plagiarism check is proof of anything. Check AI is a classifier like every tool in the table above, with the same limits. A low score does not show that you wrote the paper and a high one does not show that you did not; your university's detector may also disagree. Do not rewrite honest work to chase a score; no wording guarantees any detector's verdict. EdCitation never writes any part of a paper.

### The part of a paper that can be checked for certain

Every detector score in this guide is an estimate; whether a reference exists is not. No detector of the 14 that Weber-Wulff et al. (2023) tested reached 80% accuracy; a reference check asks a different question, with a record behind the answer. That is why EdCitation is the best tool for this job: it checks each reference against the publisher's record and never composes one, which no chatbot can promise and no classifier even attempts. [Verify references](https://edcitation.com/verify-references) is free with no account. An entry is returned as verified, as "check this" or as not found; retractions are flagged; and where the look-up could not be made, the result says so rather than calling the reference missing. Max, at $24 a month, adds Theoretics QA and the Library to what Pro offers.

| Question about a paper | What kind of answer | Where EdCitation answers it |
| --- | --- | --- |
| Does a passage read as machine-written? | An estimate, wrong in both directions in the tests above | [Check AI](https://edcitation.com/tools/check-ai), Pro |
| Does a passage match published or web text? | A match to a named source you can open and compare | [Check plagiarism](https://edcitation.com/tools/check-plagiarism), Pro |
| Does each reference exist? | A look-up in the publisher's record | [Verify references](https://edcitation.com/verify-references), free |
| Does the paper do what its brief asks? | A checklist read from the brief | [Check your paper](https://edcitation.com/check), free |

## Quick questions

### Can an AI detector tell who wrote a text?

No. It estimates how closely a text resembles language-model output. It has no access to drafts or the writer, and Turnitin's own guidance says a flagged sentence should start a conversation, not settle one.

### What are perplexity and burstiness?

Perplexity is how surprised a language model is by each next word; burstiness is how much that surprise varies across a document. Low, even perplexity reads as machine-like. GPTZero, which popularised the terms, says it stopped relying on them in autumn 2023.

### What is Turnitin's false positive rate?

Turnitin states under 1% at document level for documents with 20% or more AI writing, and around 4% per sentence. It shows an asterisk instead of a number for 1% to 19% AI writing, where false positives are commoner. These are the company's own figures.

### Are AI detectors biased against non-native English writers?

The best-known test says yes. In Liang et al. (2023), seven detectors wrongly flagged human-written TOEFL essays at an average rate of 61.3%, while essays by US school students were almost all classified correctly.

### Is there an AI detector I can run on my own paper first?

Yes. EdCitation's [Check AI](https://edcitation.com/tools/check-ai), part of Pro, gives each passage a label (AI-written, AI-assisted or human) and a score, in a report only you can see. It is an estimate like every detector in this guide; [How to check your paper for AI before submitting](https://edcitation.com/newsletter/how-to-check-your-paper-for-ai-before-submitting) explains how to read it.

## References

- Chechitelli, A. (2023a). *Understanding false positives within our AI writing detection capabilities* (16 March 2023). Turnitin. [https://www.turnitin.com/blog/understanding-false-positives-within-our-ai-writing-detection-capabilities](https://www.turnitin.com/blog/understanding-false-positives-within-our-ai-writing-detection-capabilities)
- Chechitelli, A. (2023b). *AI writing detection update from Turnitin's Chief Product Officer* (23 May 2023). Turnitin. [https://www.turnitin.com/blog/ai-writing-detection-update-from-turnitins-chief-product-officer](https://www.turnitin.com/blog/ai-writing-detection-update-from-turnitins-chief-product-officer)
- Chechitelli, A. (2023c). *Understanding the false positive rate for sentences of our AI writing detection capability* (14 June 2023). Turnitin. [https://www.turnitin.com/blog/understanding-the-false-positive-rate-for-sentences-of-our-ai-writing-detection-capability](https://www.turnitin.com/blog/understanding-the-false-positive-rate-for-sentences-of-our-ai-writing-detection-capability)
- Curtin University. (2025, September 4). *Update on Turnitin AI-detection tool*. [https://www.curtin.edu.au/news/oasis-news/update-on-turnitin-ai-detection-tool/](https://www.curtin.edu.au/news/oasis-news/update-on-turnitin-ai-detection-tool/)
- Dathathri, S., See, A., Ghaisas, S., Huang, P.-S., McAdam, R., Welbl, J., Bachani, V., Kaskasoli, A., Stanforth, R., Matejovicova, T., Hayes, J., Vyas, N., Al Merey, M., Brown-Cohen, J., Bunel, R., Balle, B., Cemgil, T., Ahmed, Z., Stacpoole, K., . . . Kohli, P. (2024). Scalable watermarking for identifying large language model outputs. *Nature, 634*, 818-823. [https://doi.org/10.1038/s41586-024-08025-4](https://doi.org/10.1038/s41586-024-08025-4)
- Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023). Evaluating the efficacy of AI content detection tools in differentiating between human and AI-generated text. *International Journal for Educational Integrity, 19*, Article 17. [https://doi.org/10.1007/s40979-023-00140-5](https://doi.org/10.1007/s40979-023-00140-5)
- Google. (n.d.). *SynthID Text*. Google AI for Developers. [https://ai.google.dev/responsible/docs/safeguards/synthid](https://ai.google.dev/responsible/docs/safeguards/synthid)
- GPTZero. (n.d.-a). *AI detector*. [https://gptzero.me/](https://gptzero.me/)
- GPTZero. (n.d.-b). *Technology*. [https://gptzero.me/technology](https://gptzero.me/technology)
- GPTZero. (2023, March 1). *What is perplexity and burstiness for AI detection?* [https://gptzero.me/news/perplexity-and-burstiness-what-is-it/](https://gptzero.me/news/perplexity-and-burstiness-what-is-it/)
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. *International Journal for Educational Integrity, 22*, Article 4. [https://doi.org/10.1007/s40979-026-00213-1](https://doi.org/10.1007/s40979-026-00213-1)
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. *Patterns, 4*(7), Article 100779. [https://doi.org/10.1016/j.patter.2023.100779](https://doi.org/10.1016/j.patter.2023.100779)
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. *International Journal of Educational Technology in Higher Education, 21*, Article 53. [https://doi.org/10.1186/s41239-024-00487-w](https://doi.org/10.1186/s41239-024-00487-w)
- Turnitin. (n.d.). *Why is the AI Writing Detection report score showing as \*%?* Turnitin Help Center. [https://helpcenter.turnitin.com/hc/en-us/articles/46245332050317-Why-is-the-AI-Writing-Detection-report-score-showing-as](https://helpcenter.turnitin.com/hc/en-us/articles/46245332050317-Why-is-the-AI-Writing-Detection-report-score-showing-as)
- University of Waterloo. (2025, September). *Discontinuing use of AI detection functionality in Turnitin*. [https://uwaterloo.ca/associate-vice-president-academic/discontinuing-use-ai-detection-functionality-turnitin](https://uwaterloo.ca/associate-vice-president-academic/discontinuing-use-ai-detection-functionality-turnitin)
- Vanderbilt University. (2023, August 16). *Guidance on AI detection and why we're disabling Turnitin's AI detector*. [https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/](https://www.vanderbilt.edu/brightspace/2023/08/16/guidance-on-ai-detection-and-why-were-disabling-turnitins-ai-detector/)
- Van Vlasselaer, M., Van Droogenbroeck, F., & Spruyt, B. (2026). Who wrote this? Evaluating the reliability of AI detection tools in higher education. *International Journal for Educational Integrity, 22*, Article 16. [https://doi.org/10.1007/s40979-026-00226-w](https://doi.org/10.1007/s40979-026-00226-w)
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. *International Journal for Educational Integrity, 19*, Article 26. [https://doi.org/10.1007/s40979-023-00146-z](https://doi.org/10.1007/s40979-023-00146-z)
