# Can teachers detect ChatGPT? What the evidence shows

**Can teachers detect ChatGPT?** Teachers can sometimes detect ChatGPT, but not reliably, and neither can detector software. Fleckenstein et al. (2024) found experienced teachers identified 37.8% of ChatGPT essays; at the University of Reading, 94% of GPT-4 exam answers went undetected. Universities therefore rely on meetings, vivas, drafts and the reference list. A reference that does not exist is the one signal that can be checked outright.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/can-teachers-detect-chatgpt

A teacher reading an essay cannot reliably tell whether ChatGPT wrote it, and the software built to help gives an estimate, not proof. That is what the published tests say, and it is why universities now lean on other things: a conversation about the work, an exam in the room, the drafts, and the reference list.

This guide gives the figures from those tests, and what universities and the UK exam boards say on their own pages, read on 22 September 2026. It is not advice on hiding AI use: where a course forbids it, using it is misconduct whatever a detector says, and where a course allows it, [cite it](https://edcitation.com/newsletter/how-to-cite-chatgpt-and-ai-tools). It ends with what a student who wrote the work can check before a marker does: the references, which EdCitation's free [Verify references](https://edcitation.com/verify-references) looks up in the publisher's record, and a detector's reading of the paper, which [Check AI](https://edcitation.com/tools/check-ai) shows to the writer alone.

## Can teachers detect ChatGPT by reading the essay?

Not reliably, on the one study that tested teachers rather than the public. Fleckenstein et al. (2024) mixed argumentative essays in English by upper-secondary students in Germany and Switzerland with essays ChatGPT wrote to the same TOEFL prompt at the same two quality levels, and asked teachers which were which.

### Trainee teachers

The 89 trainee teachers in the first study, at a German university, correctly identified 45.1% of the AI essays and 53.7% of the student essays. A test of independence found no relation between the real source and the one they named. Their average confidence was about 77%.

### Experienced teachers

The 200 practising teachers in the second study, mostly in the UK and North America, correctly identified 37.8% of the AI essays and 73.0% of the student essays. They called about a third of all essays AI and tended to call any weak essay a student's, so the AI essays they missed most were the weak ones. They did better on strong essays, and the authors note that experienced teachers even marked AI text higher when its quality was high. Both groups were overconfident, the experienced teachers at about 80%.

These were short school essays, judged with no knowledge of the writer, and a teacher who has read a student's work all year has more to go on. But the folklore that a marker can always tell has no support in the one study that tested it.

## What happened when AI answers were put into real exams?

They passed, and almost none were caught. Scarfe et al. (2024) ran a real-world test at the University of Reading in the summer of 2023. With the university's knowledge but not the markers', they set up 33 fake student accounts and submitted unedited GPT-4 answers to the take-home online exams of five undergraduate psychology modules, about 5% of each module's submissions.

94% of the AI submissions went undetected; on a stricter count, where the flag had to mention AI, 97%. The AI answers scored on average just over half a grade boundary above the real students, and only one final-year module resisted.

The limits are one department, one summer, GPT-4 as it was in 2023, and unsupervised exams, which is the point: the authors' own reading is that a pass gained this way shows no knowledge at all.

## How accurate is AI detection software?

Not accurate enough to be proof, on the two largest independent tests, and the maker of the most used detector says as much.

Weber-Wulff et al. (2023) ran 12 public tools and two commercial systems, Turnitin and PlagiarismCheck, over the same documents. Their conclusion is that the tools are "neither accurate nor reliable": every one scored below 80% accuracy and only five above 70%. Turnitin scored highest on every measure. The detail matters:

| Kind of text | How often the tools got it right |
| --- | --- |
| Written by a person | 96% |
| Written by a person in another language, then machine-translated | 20% lower |
| Written by AI, unedited | 74% |
| Written by AI, then edited by a person | 42% |
| Written by AI, then machine-paraphrased | 26% |

False positive rates on human text ran from 0% (Turnitin) to 50% (GPTZero); false negative rates from 8% (GPTZero) to 100% (Content at Scale). The main error is missing AI, not accusing people.

The error that does fall on people falls unevenly. Liang et al. (2023) ran seven detectors on 91 essays from the TOEFL English exam and 88 by US eighth-grade students. The US essays were classified almost perfectly; the TOEFL essays, all written by people, were flagged as AI at an average rate of 61.3%. When GPT-4 rewrote them with a richer vocabulary the rate fell to 11.6%: the detectors were reacting to plain language, not authorship. [What to do if an AI detector wrongly flags your writing](https://edcitation.com/newsletter/ai-detector-false-positive-what-to-do) covers that case.

Turnitin's own position, from its chief product officer, is a false positive rate under 1%, with the statements that the rate "is not zero", that Turnitin does not decide misconduct, and that the instructor must apply their own judgement (Chechitelli, 2023). Tools have changed since 2023; what has not is that a detector reports a likeness, not a record of how the text was made.

## How do universities detect AI use, if not with a detector?

By procedure: a documented concern, a meeting or viva, comparison with the student's other work, and a look at the references. Their own pages, read on 22 September 2026:

| Institution, page | What it says |
| --- | --- |
| King's College London, AI guidance, April 2025 | Did not switch on Turnitin's AI score, citing reliability and false positives. A suspected student may be invited to an investigatory meeting under the normal misconduct procedure. |
| University of Strathclyde, Guidance on using Turnitin, v2.0, August 2025 | Similarity reports cannot be used to identify, start, investigate or evidence an AI allegation. The Turnitin AI checker and any other AI detection service are not permitted. |
| University of Liverpool, process for misuse of generative AI, July 2024 | Indicators for markers, first among them false references, are "guidance only", not proof. The academic integrity officer hears the examiner, lets the student explain, and can meet both. |
| Middlesex University, academic integrity policy 2025-26 | Cases must be evidenced and documented first; "where appropriate" a viva lets the student show their understanding. Fake referencing is an offence in its own right. |
| Vanderbilt University, academic integrity and generative AI | A report to the honour council cannot rest solely on an AI detector score. The council weighs the instructions, the syllabus, the student's full body of work, and known indicators, fake citations among them. |
| Cornell University, Center for Teaching Innovation | Does not recommend detection algorithms. Evidence it names: citations to sources that do not exist, inconsistency with outlines or drafts, and a student unable to explain the work aloud. |

The common thread is a conversation: Middlesex warns that missing a viva without good reason is read as acceptance, and Vanderbilt's honour council decides whether unauthorised aid was "more likely than not" used.

The other response is to change the assessment. The University of Sydney's two-lane approach, in force from the second semester of 2025, puts some of each programme's assessment in a secured lane of in-person orals, vivas, in-class work, tests and exams (Bridgeman & Liu, 2024), on the stated ground that AI use cannot be restricted or detected in work that is not supervised face to face.

## Do schools detect ChatGPT the same way?

Broadly yes, with the teacher's knowledge of the pupil doing more of the work. The Joint Council for Qualifications (2025), the joint body of the UK exam boards, tells teachers they must know a candidate's usual standard well enough to judge whether a piece of work is within their capability, and that some direct supervision of coursework is always required. Its indicators include references that cannot be found or verified. It permits detection tools as a check, warns that their accuracy varies with the AI tool, its version and the share of AI text, and suggests a short conversation with the pupil about the work.

## What is the one signal that can actually be proved?

A reference that does not exist. Everything else here is a probability: a score, a change of style, an essay that seems too good. A reference is either in the publisher's record or it is not.

It is now the fastest-growing evidence in real cases. Munoz et al. (2026) coded 1,162 misconduct cases involving generative AI at one regional Australian university between January 2023 and December 2025. Fabricated references grew from 10.4% of evidence items in 2023 to 30.0% in 2025, the largest shift in the data, while AI detector output fell to 0.5% of items in 2025. Similarity reports were rated weak on inferential force in 100% of the items in which they appeared; a student's admission at interview was rated strong in 97.2%.

The study rates a fabricated reference as moderate evidence rather than strong, for two honest reasons: a student can invent one without AI, and a course that allows AI for editing can let one in by accident. It also exists only where there is a reference list. It is still the only signal that can be settled in minutes.

Walters and Wilder (2023) checked all 636 references in 84 short literature reviews written by ChatGPT: 55% of GPT-3.5's and 18% of GPT-4's were not real, and 43% and 24% of the real ones had substantive errors. [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references) explains the mechanism; [How to spot fabricated references in student work](https://edcitation.com/newsletter/how-to-spot-fabricated-references-in-student-work) is written for the marker.

## What can I do if I wrote the work myself?

Two checks, one for each kind of signal, and a short table of what else a university looks at.

For the thing that can be proved, EdCitation is the best tool there is, because it goes to the publisher's record for every reference and never makes one up, which is exactly what a chatbot cannot promise. The Munoz et al. (2026) figures above show why it matters: fabricated references rose to 30.0% of the evidence in real cases by 2025, while detector output fell to 0.5%. [Verify references](https://edcitation.com/verify-references) takes the paper or a pasted list and sorts every reference into verified, "check this" or not found, flagging retractions; if a look-up cannot be completed, the reference is not counted as missing. Upload the whole paper and the in-text citations are matched to the list too. It is free, with no account, and it is the check to run before a marker does. If a source turns out not to exist, [Find sources](https://edcitation.com/) searches about 300 million published works by the claim your sentence makes, so the replacement is real, and [Cite a source](https://edcitation.com/cite) builds its reference from the DOI.

For the student who wrote every word and wants to see how a detector reads it, Check AI puts one of three labels on each passage (AI-written, AI-assisted, human) with the score that produced it, and the result is seen by whoever ran the check and nobody else. It is not the detector your university uses and cannot say what that detector will report. Check AI comes with Pro, at $8 a month and 240 credits; the credits a check needs are shown before you run it, and none are taken if it fails. The $24 Max plan adds Theoretics QA and the Library. EdCitation says plainly that the result is not proof, and the tests above are why: the same texts drew false positive rates from 0% to 50% across the 14 tools in Weber-Wulff et al. (2023), and no wording can promise that honest work will not be flagged. A high score on your own writing is a reason to gather your drafts and version history before you hand in, not a reason to rewrite good sentences. EdCitation never writes any part of a paper.

### What else a university looks at, and what you can check first

| What a university looks for | Can you check it first? | How |
| --- | --- | --- |
| References that do not exist (Liverpool, Middlesex, Vanderbilt, Cornell, the JCQ) | Yes, against the record | [Verify references](https://edcitation.com/verify-references), free |
| Work unlike your outline or drafts (Cornell) | Yes, by keeping them | Write in a file that keeps versions, and keep the plan |
| Work you cannot explain aloud (Cornell, Middlesex) | Yes, by rehearsing | Explain the argument and each source to someone else |
| A detector score, where one is used | Only as an estimate | [Check AI](https://edcitation.com/tools/check-ai), Pro; not your university's tool |

## So, can teachers detect ChatGPT?

Sometimes, never with certainty. Teachers reading cold identified between 37.8% and 45.1% of ChatGPT essays in Fleckenstein et al. (2024), and markers at Reading missed 94% of GPT-4 exam answers in Scarfe et al. (2024). The software is an estimate, and the universities above say a score is not a finding. What a university can do is ask you to explain your work, compare it with your other work, assess you in the room, and check whether your sources exist. Only the last has a yes or no answer.

Not confirmed: no published test covers every current detector, and the two quoted are from 2023; the Reading paper does not state the total number of answers submitted in one place; and a 2026 analysis of UK cases sat behind a bot check and was left out.

## Quick questions

### Can teachers tell if you use ChatGPT?

Not reliably by reading alone. In Fleckenstein et al. (2024), experienced teachers identified 37.8% of ChatGPT essays and trainee teachers 45.1%, at confidence near 80%. A teacher who knows your usual work, or asks you to explain it, has more to go on.

### Do universities use AI detectors?

Some do, some have switched them off. King's College London and the University of Strathclyde do not use Turnitin's AI detection; Vanderbilt will not accept a detector score as the sole basis of a report.

### How accurate is Turnitin's AI detector?

Turnitin's own figure is a false positive rate under 1%, and it says the rate is not zero. In Weber-Wulff et al. (2023) Turnitin scored highest of 14 tools, with no false positives in that sample, but under 80% accuracy overall.

### Can a professor prove I used ChatGPT?

Not from a detector score. What a university can establish is what you say in a meeting or viva, how the work compares with your other work, and whether your references exist. The last you can check yourself first, free, with EdCitation's [Verify references](https://edcitation.com/verify-references).

### Is a viva used to check for AI?

Yes, at some universities. Middlesex University's 2025-26 policy provides for a viva where appropriate, and the University of Sydney places vivas and interactive orals in its secured assessment lane.

## References

- Bridgeman, A., & Liu, D. (2024, July 2). Frequently asked questions about the two-lane approach to assessment in the age of AI. *Teaching@Sydney*. [https://educational-innovation.sydney.edu.au/teaching@sydney/frequently-asked-questions-about-the-two-lane-approach-to-assessment-in-the-age-of-ai/](https://educational-innovation.sydney.edu.au/teaching@sydney/frequently-asked-questions-about-the-two-lane-approach-to-assessment-in-the-age-of-ai/)
- Chechitelli, A. (2023, March 16). *Understanding false positives within our AI writing detection capabilities*. Turnitin. [https://www.turnitin.com/blog/understanding-false-positives-within-our-ai-writing-detection-capabilities](https://www.turnitin.com/blog/understanding-false-positives-within-our-ai-writing-detection-capabilities)
- Cornell University Center for Teaching Innovation. (n.d.). *AI & academic integrity*. [https://teaching.cornell.edu/generative-artificial-intelligence/ai-academic-integrity](https://teaching.cornell.edu/generative-artificial-intelligence/ai-academic-integrity)
- Fleckenstein, J., Meyer, J., Jansen, T., Keller, S. D., Köller, O., & Möller, J. (2024). Do teachers spot AI? Evaluating the detectability of AI-generated texts among student essays. *Computers and Education: Artificial Intelligence, 6*, Article 100209. [https://doi.org/10.1016/j.caeai.2024.100209](https://doi.org/10.1016/j.caeai.2024.100209)
- Joint Council for Qualifications. (2025, April 30). *AI use in assessments: Your role in protecting the integrity of qualifications*. [https://www.jcq.org.uk/exams-office/malpractice/artificial-intelligence/](https://www.jcq.org.uk/exams-office/malpractice/artificial-intelligence/)
- King's College London. (2025, April). *Macro-level guidance: University-wide principles and policy*. [https://www.kcl.ac.uk/about/strategy/learning-and-teaching/ai-guidance/macro-level](https://www.kcl.ac.uk/about/strategy/learning-and-teaching/ai-guidance/macro-level)
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. (2023). GPT detectors are biased against non-native English writers. *Patterns, 4*(7), Article 100779. [https://doi.org/10.1016/j.patter.2023.100779](https://doi.org/10.1016/j.patter.2023.100779)
- Middlesex University. (2025). *Academic integrity and misconduct policy and procedures 2025-26*. [https://www.mdx.ac.uk/media/middlesex-university/about-us-pdfs/policy-pdfx27s/Policy-and-Procedures-for-Academic-Integrity-and-Misconduct-2025-2026.pdf](https://www.mdx.ac.uk/media/middlesex-university/about-us-pdfs/policy-pdfx27s/Policy-and-Procedures-for-Academic-Integrity-and-Misconduct-2025-2026.pdf)
- Munoz, A., Hinchcliff, M., Langfield, C., & Rogerson, A. (2026). How strong is the evidence in generative AI-related academic misconduct allegations? A mixed-methods analysis. *International Journal for Educational Integrity, 22*, Article 26. [https://doi.org/10.1007/s40979-026-00235-9](https://doi.org/10.1007/s40979-026-00235-9)
- Scarfe, P., Watcham, K., Clarke, A., & Roesch, E. (2024). A real-world test of artificial intelligence infiltration of a university examinations system: A "Turing Test" case study. *PLOS ONE, 19*(6), Article e0305354. [https://doi.org/10.1371/journal.pone.0305354](https://doi.org/10.1371/journal.pone.0305354)
- University of Liverpool, Centre for Innovation in Education. (2024, July 19). *Academic integrity process for the misuse of generative AI*. [https://www.liverpool.ac.uk/media/livacuk/centre-for-innovation-in-education/digital-education/generative-ai-teach-learn-assess/academic-integrity-process-for-the-misuse-of-gai.pdf](https://www.liverpool.ac.uk/media/livacuk/centre-for-innovation-in-education/digital-education/generative-ai-teach-learn-assess/academic-integrity-process-for-the-misuse-of-gai.pdf)
- University of Strathclyde. (2025, August). *Guidance on using Turnitin* (Version 2.0). [https://www.strath.ac.uk/media/ps/cs/gmap/academicaffairs/policies/Guidance_on_using_Turnitin.pdf](https://www.strath.ac.uk/media/ps/cs/gmap/academicaffairs/policies/Guidance_on_using_Turnitin.pdf)
- Vanderbilt University. (n.d.). *Academic integrity and generative AI*. [https://www.vanderbilt.edu/generative-ai/academic-integrity/](https://www.vanderbilt.edu/generative-ai/academic-integrity/)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. *International Journal for Educational Integrity, 19*, Article 26. [https://doi.org/10.1007/s40979-023-00146-z](https://doi.org/10.1007/s40979-023-00146-z)
