A reference check for a whole class is a different job from checking one essay that feels wrong. Sixty scripts with fifteen references each is nine hundred lookups, and the check must be the same for every student or it cannot be defended at an appeal. It must also keep apart three results that look alike on a marking sheet: a source that does not exist, a real source cited wrongly, and a source nobody managed to check.
This guide, published by EdCitation for instructors, markers and course leaders, sets out a method, then shows the lookups done by EdCitation's free Verify references and by References from a file, the Pro tool that reads a paper's reference list and in-text citations from the file. The studies and guidance were read on 23 September 2026.
Should I check every script or a sample?
Check every script's list when a lookup does the searching, and sample only when you search by hand. Either way, fix the rule before you open the first script, and let the stakes decide how far the hand check goes.
Why the rule comes before the marking
The University of Bath (2024) warns its staff that judging whether writing came from AI is "often subject to unconscious bias which may single out specific groups of students". Checking only the scripts that felt wrong carries the same risk; a rule set in advance does not.
| Stakes | Example | What to check | What to confirm by hand |
|---|---|---|---|
| Low | A draft bibliography | A fixed sample from every script: the first entry, the last, and any without a DOI | Only what feedback will mention |
| Medium | A module essay or report | Every reference, looked up in Verify references | Every "not found" and "check this" |
| High | A dissertation or thesis | Every reference and in-text citation, with References from a file, and the sources behind the main claims read | Every flag, logged |
The table is our recommendation, not a published standard; your department's own rule wins.
What a sample misses
A sample finds a list full of fabricated entries quickly and usually misses a single one: pick five references at random from fifteen and any given entry is in the sample one time in three. For summative work, look every entry up; with EdCitation that costs nothing but the paste.
How often do AI chatbots invent references?
Often enough to check for, though published rates vary with the model, the topic and how each study defined a fake.
Walters and Wilder (2023) checked all 636 references in 84 short literature reviews that GPT-3.5 and GPT-4 wrote on topics typical of first-year composition classes in the United States. Of the 222 works GPT-3.5 cited, 55% were fabricated; of GPT-4's 414, 18%. For book chapters, the figure was 70% for both.
Chelli et al. (2024) asked GPT-3.5, GPT-4 and Bard to find the studies for 11 published systematic reviews on shoulder rotator cuff conditions, and checked 471 references. The hallucination rate was 39.6% for GPT-3.5 (55 of 139), 28.6% for GPT-4 (34 of 119) and 91.4% for Bard (95 of 104).
The studies drew the line in different places
Walters and Wilder counted a work as real if its title and authors matched a published work or nearly did, and treated a wrong journal title as a citation error. Chelli et al. counted a reference as hallucinated when two of its title, first author and year were wrong. On either definition, a real paper with the wrong year is a real paper.
A marker's line is the one in the regulations. Manchester Metropolitan University (2025), for example, lists "creating false references" as falsification, including "'hallucinated' references", for example ones generated by artificial intelligence. No published figure was found for the share of student submissions carrying a fabricated reference, so none is given here.
What do "not found" and "could not check" mean?
"Not found" means no record turned up where the check looked; "could not check" means the check did not happen, because an index did not answer. The first is a reason to search by hand. The second is evidence of nothing, and EdCitation never shows it as "not found".
| Result | What it tells you | What the marker does |
|---|---|---|
| Verified | A published record matches the entry | Nothing more, unless a claim needs reading against it |
| Check this | A record exists but differs from the entry | Compare them field by field: a wrong year, or a different work |
| Not found | No publisher's or registry's record matches it, nor, for a book, a library catalogue's | Confirm by hand before saying anything to the student |
| Could not check | An index did not answer | Try again later or search by hand; never write it into a concern |
| Retracted | The journal has withdrawn the paper | A point for feedback, not an allegation; fine if the retraction is the subject |
Cornell University (n.d.) lists, among the evidence instructors should gather, work citing "content that does not exist or cannot be verified". Read "cannot be verified" as the marker's verdict after a proper search; a checker's "could not check" is a search that never ran.
How do I confirm a reference by hand before raising it?
Search everywhere the work would have to be if it existed, write each search down, and treat a near match as a citation error. Walters and Wilder (2023) are the model: after searching Google Scholar, PubMed, Scopus, WorldCat and more, their final check on each apparent fake was to browse the journal's volume and issue and search the publisher's own site.
- Put the full title, in quotation marks, into Crossref and Google Scholar, and PubMed for health subjects, with and without the subtitle.
- Resolve the DOI, if there is one, and check that the page it opens carries the same title and authors. Steps 1 and 2 are the lookups Verify references has already made; repeat them yourself for anything you will raise.
- Search the first author's publication list or university profile.
- Go to the journal. Browse the volume and issue given and search the publisher's site.
- Try a library catalogue such as WorldCat for a book or chapter.
- Place the kind of source. A report, thesis, conference paper or source in another language may be in no index; check the organisation's site or the repository.
- Log it: the entry, each place searched, the date and the result. The log is what the student is shown.
A match or near match on title and authors is a real work cited badly: a referencing error. Only an entry that fails every step, and should have been found, is a reference you cannot find, and the student may still know more.
What do universities tell staff to do when a reference cannot be found?
The three universities read here agree that an unfindable reference is evidence to examine, not a finding. They disagree about who acts first and on what standard.
| University | False references | What happens first | Standard |
|---|---|---|---|
| Bath (2024) | "Not evidence per se" of AI misuse, but may be an assessment offence | With evidence in the text, a formal claim; without, a viva voce, with permission | "Reasonable grounds" to raise a claim |
| Cornell (n.d.) | Content that "does not exist or cannot be verified" is evidence to gather | A conversation with the student, then documentation | "Clear and convincing": "far more likely to be true than false" |
| Manchester Metropolitan (2025) | "Creating false references", hallucinated ones included, is falsification | A report to Assessment Management "as soon as suspected" | Misconduct "likely to have occurred" |
A marker at Manchester Met reports first; Cornell's instructors may talk first. Your own institution's procedure decides, so read it before you write to the student.
How do I talk to a student about a reference I cannot find?
Ask for the source, show the search that failed, and keep the question about the reference rather than about AI.
Ask for the source, with your search attached
"I searched for the Lee (2021) article in your reference list in Crossref, Google Scholar and the journal's own archive on 23 September, and could not find it. Can you send me your copy, or tell me where you read it, within a week?" That says what you did and asks for what only the student can supply. Cornell adds that a conversation may bring out new information that clarifies the problem.
Ask about the process, not the tool
Bath places the offence in the attribution: it "would not be that a GenAI tool has been used", but that text in the argument is unattributed or falsely attributed. Its viva tips suggest asking what the student read and how they filled gaps in what they knew. Rule out the innocent explanations first:
- a reference copied, errors and all, from another paper's list;
- a source met on a lecture slide and never read;
- a year, volume or journal title miscopied;
- a real report, thesis or older book that no index holds.
What about data protection when checking student work?
Your institution's rules decide what may leave your hands, whatever the tool, so ask its data protection office before the first script goes anywhere. A paper carries a student's name, often a student number, and their writing.
Under UK law, the Information Commissioner's Office (n.d.) explains, an organisation that has a service handle personal data for it "must only use a processor that can provide 'sufficient guarantees'", needs a contract, and stays responsible for that processor's compliance. Other countries' laws differ; everywhere, your institution's approved tools and data protection office decide. Some draw tight lines: Bath told staff not to submit students' work to AI detectors, calling it "potentially illegal".
Three habits keep student data protection simple:
- Paste the reference list, not the paper. A list holds the student's sources, not their writing, and Verify references takes one with no account.
- Leave the cover page behind. No name or student number needs to travel with a list.
- Upload whole papers only where your rules allow. References from a file keeps each paper in My papers so the next check shows what changed, a copy your retention rules must cover.
For a department, the decision belongs with the institution, which is where the Institution licence sits.
What came back when we checked a test list?
One entry verified, one flagged as retracted, and the invented one not found. On 23 September 2026 we sent three references to the lookup behind Verify references: a well-known real paper without its DOI, a well-known retracted paper, and a reference we invented ourselves for this test, with made-up authors, title and journal (Crossref lists no journal of that name).
| Entry as pasted | Result | Reason given | Record returned |
|---|---|---|---|
| Kruger and Dunning (1999), Journal of Personality and Social Psychology, no DOI | Verified | "Title, year and first author match a published record." | The 1999 article, with the DOI the entry lacked |
| Wakefield et al. (1998), The Lancet, with DOI | Retracted | "The DOI resolves to this record and the title matches. This paper has since been retracted." | Now titled "RETRACTED: Ileal-lymphoid-nodular hyperplasia…", with a 2004 correction and the 2010 retraction notice |
| Marston and Adeyemi (2021), Journal of Academic Referencing Practice, invented by us | Not found | "No publisher's or registry's record matches this reference." | None |
What a marker takes from the run
The invented reference came back not found, and the record-level reason is only a lookup's result: before raising it, step 4 confirms it by hand, here in a minute, since there is no journal to browse. Give every "check this" the same hand check, since it means a record exists that differs from the entry, and never take a record the checker offers as the source the student meant.
The retracted paper is the opposite case, real and withdrawn: a matter for feedback, with a pointer to what a retraction means.
Which EdCitation tool fits a whole class?
Verify references for a free check of every list, and References from a file when you mark from the paper and will check it again. Both are the best tools for this job because each looks every reference up in the publisher's record and never writes one. Walters and Wilder (2023) note, from earlier studies, that ChatGPT often answers wrongly when asked whether a citation is correct; a lookup reports only what the record holds, which is why, in our run, a 1998 paper came back with its 2010 retraction attached.
Verify references is free with no account. Paste a list or upload a paper: each reference comes back verified, "check this" or not found, retracted papers are flagged, and "could not check" is never shown as "not found". With the whole paper uploaded, every in-text citation is matched to the list and every entry back to the text.
References from a file, part of Pro, takes the paper as a Word document or PDF and returns a verdict on every reference, with retractions flagged and published corrections linked; every citation matched to the list, with the place a source is cited but not listed, or listed and never cited; and, for a reference that cannot stand, published sources that fit, something concrete for feedback. The paper stays in My papers, so a resubmission shows what was fixed. A free account includes three checks a month, not enough for a class; Pro, at $8 a month, has no limit (see pricing).
For a department, the Institution licence covers every student, with integration into the LMS and sign-in, the university's own branding, training and a single invoice, so students can run the same check before they submit.
Further reading: every EdCitation tool for instructors, checking submissions against the assignment brief with the free Check your paper, how to spot fabricated references in student work, and the student's side, upload your paper and check every reference. EdCitation never writes any part of a paper and gives no grade.
Quick questions
Should I run every student's reference list through a checker?
For assessed work, yes: one lookup for every list is quicker than sampling by hand and fairer than checking the scripts that looked wrong. EdCitation's free Verify references takes a pasted list.
Is a reference I cannot find proof of academic misconduct?
No. It is evidence to examine: confirm it by hand, ask the student for the source, and follow your institution's procedure, which may require a report first.
What is the difference between "not found" and "could not check"?
"Not found" means no record turned up where the check looked, and calls for a search by hand. "Could not check" means an index did not answer, so no search took place; it is never evidence against a student.
Does a fabricated reference prove the student used AI?
No. It shows a source cannot be found, not how the reference got there; Bath places the offence in the false attribution, not the tool.
How much does it cost to check a whole class's references?
Nothing with Verify references, which is free with no account. References from a file, which reads each paper from the file and keeps it for the next check, is part of Pro at $8 a month, with no limit on checks.
References
- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research, 26, Article e53164. https://doi.org/10.2196/53164
- Cornell University, Center for Teaching Innovation. (n.d.). AI & academic integrity. https://teaching.cornell.edu/generative-artificial-intelligence/ai-academic-integrity
- Information Commissioner's Office. (n.d.). What responsibilities and liabilities do controllers have when using a processor? https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/contracts-and-liabilities-between-controllers-and-processors-multi/responsibilities-and-liabilities-for-controllers-using-a-processor/
- Manchester Metropolitan University. (2025). Academic misconduct policy 2025/26. https://www.mmu.ac.uk/legal/policies/misconduct-policy-25-26
- University of Bath. (2024). GenAI and academic misconduct: Guidance for colleagues [PDF]. Teaching Hub. https://teachinghub.bath.ac.uk/wp-content/uploads/2024/02/Guidance-on-GenAI-Suspected-Misuse.pdf
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5