Since early 2023, researchers have run AI detectors on writing whose author they already knew and counted every verdict. Are AI detectors accurate? In those papers it depends on the text, the tool, its version and the threshold, and the figures run from perfect to near-total failure.
This roundup, published by EdCitation and checked on 2 October 2026, gives each study's figures as the paper states them. We read every paper at its abstract and its results, in the full text. How detectors score text is in How AI detectors work and how accurate they are; bias has its own guide. EdCitation's Check AI shows how one detector reads your paper, passage by passage, and the free Verify references checks the part that is a record, not an estimate: the references.
Are AI detectors accurate?
On unedited text from a known model, some are very accurate; on edited, mixed or disguised text, most are not. Walters (2023) found Copyleaks and Turnitin correct on all 126 essays tested, while Weber-Wulff et al. (2023) found none of 14 tools reached 80% accuracy. Both stand: the teams tested different texts, versions and settings. Four things decide the figure:
- The text. Plain AI output is the easy case; paraphrased, humanised and half-human texts are hard.
- The version. Detectors are retrained often, so a 2023 result describes a 2023 model.
- The threshold. Allowed more false positives, a tool catches more AI text.
- The writers. Second-language essays drew more false flags (Liang et al., 2023), and abstracts dense with academic vocabulary higher scores (Karr et al., 2026).
The studies, in one table
Peer-reviewed papers and named preprints that ran detectors on texts of known origin, by publication date.
| Study | Venue | Tested | Sample | Headline result, as published |
|---|---|---|---|---|
| Gao et al. (2023) | npj Digital Medicine | GPT-2 Output Detector and blinded reviewers, December 2022 | 50 real, 50 ChatGPT medical abstracts | Detector AUROC 0.94; reviewers caught 68% of generated abstracts, wrongly flagged 14% of real ones |
| Krishna et al. (2023) | NeurIPS 2023 | Watermarking, DetectGPT, GPTZero, OpenAI's classifier, against a paraphraser | GPT-2, OPT-13B and GPT-3.5 text | At a 1% false positive rate, paraphrasing cut DetectGPT from 70.3% to 4.6% |
| Liang et al. (2023) | Patterns | 7 detectors | 91 TOEFL and 88 US eighth-grade essays, all human | TOEFL essays: average false positive rate 61.3%; US essays classified accurately |
| Weber-Wulff et al. (2023) | International Journal for Educational Integrity | 14 tools, Turnitin included, March to April 2023 | 54 documents | All under 80% accuracy, 5 over 70%; paraphrased AI text 26% |
| Ibrahim et al. (2023) | Scientific Reports | GPTZero, OpenAI's classifier | Student and ChatGPT answers, 32 courses | Student answers flagged: 18% and 5%; after Quillbot, 95% of ChatGPT answers undetected |
| Walters (2023) | Open Information Science | 16 detectors, June to July 2023 | 126 essays: GPT-3.5, GPT-4, students | Copyleaks and Turnitin 100%, Originality.ai 98%, the rest 63% to 88% |
| Perkins et al. (2024) | International Journal of Educational Technology in Higher Education | 7 detectors, September to October 2023 | 114 samples, 805 tests | 39.5% accuracy on unaltered AI text, 67% on human text; simple edits cut accuracy by 17.4% |
| Dugan et al. (2024) | ACL 2024 | 12 detectors, GPTZero, Originality, Winston and ZeroGPT among them | Over 6 million generations, 11 models | Detectors "easily fooled" by attacks, sampling changes and unseen models |
| Sadasivan et al. (2025), preprint | arXiv (record says published in TMLR) | Four kinds of detector, against repeated paraphrasing | Passages of about 300 tokens | Watermark detection at a 1% false positive rate: 99.3% down to 9.7% |
| Hadra et al. (2026) | International Journal for Educational Integrity | Turnitin and Originality, January to May 2025 | 192 texts, including half-and-half | Accuracy 0.61 and 0.69; both poor on mixed texts |
| Karr et al. (2026), preprint | arXiv | Commercial detectors, GPTZero among them | 642 abstracts, AI-edited, then humanised | GPTZero flagged 37.6% to 48.5% of lightly AI-edited abstracts; 4% or fewer detectable after humanising |
How to read an accuracy figure
The papers report different measures, and they are not interchangeable.
- Accuracy is the share of all texts sorted correctly, so it depends on the test's mix of human and AI texts.
- False positive rate is the share of human texts called AI: the error that accuses a student.
- Detection at a fixed false positive rate is how much AI text a tool catches when allowed to flag only 1% or 5% of human writing. Held to 1%, GPTZero caught 7.1% of unedited GPT-3.5 text in Krishna et al. (2023).
- AUROC sums up every threshold, from 0.5 (a coin toss) to 1 (perfect).
Dugan et al. (2024) found detectors reached the accuracies quoted in viral reports only at similarly high false positive rates. An accuracy figure without its false positive rate says little.
What do the studies agree on?
They agree on four points, across tools, years and fields.
Disguising AI text works. Repeated paraphrasing cut a watermark detector from 99.3% to 9.7% (Sadasivan et al., 2025), and a Quillbot rewrite hid 95% of ChatGPT's course answers (Ibrahim et al., 2023). Not every trick beats every tool: in Dugan et al. (2024), paraphrasing raised Originality's accuracy, while swapping letters for look-alike characters cut it from 85.0% to 9.3%.
Mixed and lightly edited writing is the hard case. Turnitin and Originality both did poorly on half-human, half-AI texts (Hadra et al., 2026). GPTZero flagged up to 48.5% of abstracts after a light AI edit, which the authors treat as guideline-compliant help, yet caught 4% or fewer of rewrites once a humaniser had run; honest editing, they conclude, now carries more risk of sanction than evasion (Karr et al., 2026).
Tools differ more than the label suggests. Accuracy ran from 63% to 100% across the 16 tools Walters (2023) tested. In Perkins et al. (2024), Copyleaks caught the most AI text and wrongly flagged half the human samples; four of the seven detectors flagged none.
Results age quickly. OpenAI took its classifier down in July 2023 for low accuracy, as Krishna et al. (2023) note. Turnitin updated its model in October 2025 and February 2026, and Originality.ai retired its older models in September 2026, their own pages say.
Where do the studies disagree?
They disagree most on how often human writing is wrongly flagged, and on whether detection can ever be reliable.
How often human writing is flagged
At the low end, Copyleaks and Turnitin made no errors on Walters's 126 essays, though the author notes Turnitin may have been trained on those student papers; four other tools flagged 10% to 14% of the human essays. Dugan et al. (2024) found the commercial detectors' false positive rates at or below 1.7% at commonly chosen thresholds. At the high end sit Liang et al. (2023), with 61.3% on non-native English essays, and GPTZero's 18% on student answers (Ibrahim et al., 2023). The gap tracks who wrote the human text, its length, the threshold and the year.
Whether reliable detection is possible
Sadasivan et al. (2025) prove that as AI text grows more like human text, even the best possible detector weakens: when the two are close, its AUROC falls below 0.7. Krishna et al. (2023) proposed that AI companies check texts against a store of their own outputs, which caught 80% to 97% of paraphrased text while flagging 1% of human text; Sadasivan's repeated paraphrasing cut that defence from 100% to below 60%.
People are no shortcut. Blinded reviewers caught 68% of generated abstracts and wrongly flagged 14% of real ones, while the detector, at its best cut-off, caught 86% and flagged 6% (Gao et al., 2023).
What do the detector makers say about their own limits?
Turnitin, GPTZero and Originality.ai each say on their own pages, read on 2 October 2026, that a score should not decide a case alone.
- Turnitin. Its release notes say an AI writing score "should not be used as the sole basis for adverse actions against a student", record that early versions produced more false positives in a document's first and last sentences, and show that since July 2024 scores from 1% to 19% are hidden to avoid false positives (Turnitin, n.d.).
- GPTZero. Its FAQ says its results "should not be used to punish students", that it is not trained to spot AI text heavily modified after generation, and that it can flag highly procedural human text (GPTZero, n.d.).
- Originality.ai. It does not believe detection scores alone should be used for disciplinary action, since its false positive rate, "even if low", is too high to rely on (Originality.ai, n.d.).
Makers' and testers' figures differ, as How AI detectors work sets out; on that point they agree.
What do the studies not show?
They do not give the error rate of the detector your university runs today, on your writing.
- Most tested 2023 versions, since replaced.
- Most samples are small: 54 documents, 126 essays, 192 texts.
- Mixed texts are built, not observed, and real student use is messier.
- None follows real cases. What carries weight in misconduct cases is covered in The em dash is not proof of AI.
What we could not confirm: the journal page of Sadasivan et al. (2025) showed a robot check, so we read the arXiv version; a 2026 journal article arguing that most detector findings are false would not open, and is left out. Karr et al. (2026) is not yet peer reviewed, and its labels are, in its own word, proxies.
We checked this roundup's own references
On 2 October 2026 we pasted the 11 study references below into EdCitation's Verify references. Back came 11 verified, 0 doubtful, 0 not found, 0 retracted, 0 unchecked. For the nine journal and conference papers, the reason given was "The DOI resolves to this record and the title matches."; for the two arXiv preprints, "The page is at this address, and its title matches."
The detector studies disagree because they estimate. A reference check does not: the record is there or it is not.
How should you use these findings on your own paper?
Treat any score as one reading, and build the record that answers it.
- Write where the history is kept, such as Google Docs, or Word saved to OneDrive, and keep notes and drafts. What to do if an AI detector wrongly flags your writing lists the evidence that counts.
- Check every reference in Verify references, free with no account. A real, accurate list is part of the record of real reading, and anyone can check it.
- See how a detector reads each passage in Check AI, and match each flagged passage to its drafts. Do not rewrite honest work to chase a score: scores shift with the tool and version.
- Follow your course's rules on AI tools, and declare permitted use; How to declare AI use in an assignment has the wording.
- If flagged, ask which detector, version and threshold. Those moved the results in every study above.
Why use EdCitation's Check AI, and what does it cost?
Check AI is the best tool for one job: showing you, privately and before anyone else runs a detector, which passages of your own paper a detector reacts to. The studies put the errors there. Turnitin's notes found false positives commoner in opening and closing sentences, and Hadra et al. (2026) found both detectors weakest on mixed texts, so one score for a whole paper hides the passages you would need to explain.
Paste your text or choose the file. Back comes one figure for the paper, the share that does not read as confidently human, and every passage labelled AI-written, AI-assisted or human, in place, with its score. Only you see the report. Check AI is part of Pro, $8 a month, and runs on credits: 240 a month with Pro, 720 with Max at $24 a month, more from $10. The cost shows before a check runs; a failed check, or the same text again within a day, uses nothing. See pricing.
It is a detector, and every limit in this roundup applies to it. Another detector, your university's included, may read the same text differently, and no score is proof of anything. EdCitation finds, cites and checks sources, and the writing is always yours: EdCitation never writes any part of anyone's work. What Check AI shows, passage by passage explains each label.
Quick questions
Can AI detectors be trusted to prove cheating?
No. The studies found errors in both directions, and Turnitin, GPTZero and Originality.ai each say a score should not decide a case alone.
What is the false positive rate of AI detectors?
It varies by tool and writer: none for Copyleaks and Turnitin on 42 student essays (Walters, 2023), and 61.3% on average across seven detectors on non-native English essays (Liang et al., 2023).
Which AI detector is the most accurate?
No tool won every study. Copyleaks and Turnitin led one 2023 test, Originality beat Turnitin in tests run in 2025, and even the strongest failed under some conditions in Dugan et al. (2024).
Does paraphrasing beat AI detectors?
In most tests it cut detection sharply, so a low score proves nothing. Rewriting tools carry risks of their own, set out in our guide to AI humanizers.
Can I see how a detector reads my paper first?
Yes. EdCitation's Check AI, part of Pro, labels each passage AI-written, AI-assisted or human, with its score, in a report only you see: one detector's estimate, not a verdict.
References
- Dugan, L., Hwang, A., Trhlík, F., Zhu, A., Ludan, J. M., Xu, H., Ippolito, D., & Callison-Burch, C. (2024). RAID: A shared benchmark for robust evaluation of machine-generated text detectors. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 12463–12492). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.acl-long.674
- Gao, C. A., Howard, F. M., Markov, N. S., Dyer, E. C., Ramesh, S., Luo, Y., & Pearson, A. T. (2023). Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. npj Digital Medicine, 6, Article 75. https://doi.org/10.1038/s41746-023-00819-6
- GPTZero. (n.d.). Frequently asked questions. https://gptzero.me/faq
- Hadra, M., Cambridge, K., & Mesbah, M. (2026). Evaluating the accuracy and reliability of AI content detectors in academic contexts. International Journal for Educational Integrity, 22, Article 4. https://doi.org/10.1007/s40979-026-00213-1
- Ibrahim, H., Liu, F., Asim, R., Battu, B., Benabderrahmane, S., Alhafni, B., Adnan, W., Alhanai, T., AlShebli, B., Baghdadi, R., Bélanger, J. J., Beretta, E., Celik, K., Chaqfeh, M., Daqaq, M. F., El Bernoussi, Z., Fougnie, D., Garcia de Soto, B., Gandolfi, A., . . . Zaki, Y. (2023). Perception, performance, and detectability of conversational artificial intelligence across 32 university courses. Scientific Reports, 13, Article 12187. https://doi.org/10.1038/s41598-023-38964-3
- Karr, J. A., Jr., Khvatskii, G., Hua, T., & Chawla, N. V. (2026). Why AI detection fails for academic integrity [Preprint]. arXiv. https://arxiv.org/abs/2608.11256
- Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. In Advances in Neural Information Processing Systems 36 (pp. 27469–27500). Neural Information Processing Systems Foundation. https://doi.org/10.52202/075280-1195
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. Patterns, 4(7), Article 100779. https://doi.org/10.1016/j.patter.2023.100779
- Originality.ai. (n.d.). AI detector by Originality.ai has 99%+ accuracy: Originality.ai study. https://originality.ai/blog/ai-accuracy
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). Simple techniques to bypass GenAI text detectors: Implications for inclusive education. International Journal of Educational Technology in Higher Education, 21, Article 53. https://doi.org/10.1186/s41239-024-00487-w
- Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2025). Can AI-generated text be reliably detected? [Preprint]. arXiv. https://arxiv.org/abs/2303.11156
- Turnitin. (n.d.). AI writing detection model. Turnitin Guides. https://guides.turnitin.com/hc/en-us/articles/28294949544717-AI-writing-detection-model
- Walters, W. H. (2023). The effectiveness of software designed to detect AI-generated writing: A comparison of 16 AI text detectors. Open Information Science, 7(1), Article 20220158. https://doi.org/10.1515/opis-2022-0158
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26. https://doi.org/10.1007/s40979-023-00146-z