A reference list used to be the dull part of marking. Now it is often the most informative, because a language model asked for sources will write references that look right and point at nothing, and some of them reach submitted work.
For a marker this has one great advantage over other signs of AI misuse. Whether a paper exists is a matter of record. It can be checked, the check can be shown to the student, and the result does not depend on anyone's impression of the prose. This guide, for lecturers, teaching assistants and librarians, covers what to look for, how to check quickly, and how to act fairly on what you find. The quick check needs no special access: paste a student's reference list into EdCitation's free Verify references and every entry is looked up in the publisher's record.
What do fabricated references look like?
A fabricated reference is usually built from real parts: real authors, a real journal and a plausible title that nobody wrote. The format is often flawless, so look at the facts, not the punctuation.
| Sign | What the evidence says | How to confirm |
|---|---|---|
| The DOI opens a different paper, or nothing | In a 2025 test of GPT-4o by Linardon et al. (2025), 33 fabricated references carried a DOI. 21 were real DOIs that belonged to unrelated articles, and 12 did not work | Resolve the DOI at https://doi.org/ and compare title and authors. Or give the DOI to EdCitation's Cite a source: the reference it builds carries the title of whatever paper the DOI is registered to |
| Real authors, real journal, unknown title | Walters and Wilder (2023) found that 55% of the references GPT-3.5 wrote and 18% of those GPT-4 wrote were fabricated, in 636 references across 84 short literature reviews | Search the exact title in quotation marks. Check the first author's publication list |
| Authors who do publish on the topic | Gravel et al. (2023) found 41 of 59 references fabricated, and yet 56 of the 59 named authors with previous publications on a related topic, and 71% of the fakes were attributed to a known medical journal | Treat a familiar author name as no evidence at all until the title is found |
| Volume, issue or pages that do not fit | Bhattacharyya et al. (2023) found wrong volume, page numbers and year among the most frequent errors in references written by ChatGPT-3.5, and an incorrect PubMed ID in 93% of the papers | Open that issue's table of contents on the journal's site |
| A list that is uniformly neat | Every source recent, every title a close restatement of the essay's own claims, no set readings, no reports, no books. This is our own observation, not a measured finding | Treat it as a reason to sample, nothing more |
Wrong details are not fabrication
Students miscopy years and page numbers, and so do AI tools writing about papers that do exist. In the Walters and Wilder study many of the real references carried substantive errors too, 43% of GPT-3.5's and 24% of GPT-4's. A wrong page range on a paper that exists is a referencing error, to be marked as one. A paper that does not exist is a different matter, and the two should not be recorded in the same sentence of your notes.
The word "hallucinated" is used loosely for both. Keep your own note specific: either the source could not be found where it should be, or the source exists and the details are wrong.
Why do some reference lists have more fabricated citations than others?
The rate depends on the topic, the kind of source and how narrow the question was, which is useful when you decide where to sample.
Obscure topics. Linardon et al. (2025) had GPT-4o write six literature reviews and checked all 176 references. Fabrication ran at 4 of 68 on major depressive disorder, 6%, against 17 of 60 on binge eating disorder, 28%, and 14 of 48 on body dysmorphic disorder, 29%. The less that has been written about a subject, the more the model fills in.
Narrow questions. In the same study, asking for a specialised review of digital interventions for binge eating disorder produced a 46% fabrication rate, against 17% for a general overview of the same condition. Final-year projects and dissertations ask exactly those narrow questions.
Books and chapters. In the Walters and Wilder (2023) data, 70% of GPT-4's book chapter citations were fabricated, far above the 18% for its references as a whole. A chapter citation with no DOI deserves more of your attention than a journal article with one.
Web links, in work built with a search-backed tool. Rao, Wong and Callison-Burch (2026), in a preprint that has not yet been peer reviewed, tested ten search-backed models and deep research agents and found that 3% to 13% of cited URLs had no record in the Wayback Machine and probably never existed, with 5% to 18% not resolving at all. Turning search on reduces the problem and does not remove it.
How do I spot-check a reference list in five minutes?
Sample five references and check each against the publisher's record. Five will not catch everything, but it tells you whether the list deserves a full check.
- Choose the sample. Take the two sources the argument leans on most, two that look too neat, and one at random. Weight the sample towards book chapters, obscure topics and sources without a DOI.
- Resolve the DOI. Does the page that opens carry the same title, the same authors and the same journal?
- No DOI, or it failed. Search the exact title in quotation marks in Crossref or Google Scholar, and in PubMed for health subjects. A published article's exact title nearly always turns up.
- Check one issue. For one reference, open the journal's table of contents for that volume and issue, and see whether the pages belong to that article.
- Check one claim. For the source that matters most, read the abstract against the sentence it supports. A real paper cited for something it does not say is its own problem.
- Decide. If all five hold, move on. If one fails, check the whole list before you form a view. That full check is one paste into EdCitation's Verify references, which marks each entry verified, "check this" or not found, so your own time goes on the entries it could not confirm.
Apply the same sampling to every script in the batch. A check run on some students and not others is hard to defend later.
What does "not found" not prove?
A missing record does not prove that a source is invented, because no index covers everything. Two results must be kept apart.
No record where there should be one. A recent article, said to be in a well-known indexed journal, with a DOI that does not resolve and a title that no index knows. That is strong evidence.
Could not be checked. The indexes cover this kind of source poorly, and the check belongs somewhere else.
The kinds of source the indexes cover poorly
- reports and other grey literature from governments, charities and companies
- books and book chapters, especially older ones
- theses and conference papers
- articles in small, regional or older journals without DOIs
- sources in languages other than English, or cited under a translated title
For these, the check is at the source: the organisation's website, a library catalogue such as WorldCat, a national repository. Older web references fail for an ordinary reason as well. In more than 3.5 million articles published from 1997 to 2012, Klein et al. (2014) found one in five affected by reference rot, a cited web address that had stopped working or no longer showed the cited content. A dead link in a list of older sources is weak evidence of anything.
A student may also have copied a reference in good faith from another paper's list, errors and all. How to check whether a reference is real sets out the same distinction for students.
How do I handle a suspected fabrication fairly?
Begin with a request for the source, not with an accusation, and follow your institution's procedure from the first step. A student who read the paper can usually produce a PDF, a link or a library record.
Keep two questions apart. Whether the source exists is a matter of record. How the reference came to be in the essay is not, and it needs a conversation. An invented reference does not by itself show that AI wrote the essay, and an AI detector's score does not settle it either. The public evidence on detector error is gathered in our guide for students whose writing is wrongly flagged, and it is worth reading before you rely on a score in an academic integrity case.
On process, the Good Practice Framework of the Office of the Independent Adjudicator, which reviews student complaints in England and Wales, is a useful statement of fairness wherever you teach:
- the burden of proof rests on the provider, which has to show that the student did what is alleged
- where the regulations do not say otherwise, the standard is the balance of probabilities
- the student should be given, in advance, copies of all the information the decision maker will consider, with reasonable notice of any meeting or hearing
- the student should have a fair opportunity to respond to it, and may appoint a representative
- telling the student that disciplinary action is being considered, as soon as possible after the event, is good practice
- reasons should be given for the decision and for any penalty
The framework also expects procedures to allow a case to be resolved informally and early where that is appropriate, which is often the right answer when a first-year student has copied a reference list from a source they did not read.
Keep a record of each check: the reference, where you searched, what you found and the date. It is the evidence you will share, and it is what turns "this looks fake" into something a student can answer. If a checker such as EdCitation's did the first pass, note its result for each entry, then repeat the search by hand for every entry you intend to raise. The tool's "not found" tells you where to look; your own search of the indexes that should hold the paper, written down, is what the student is shown.
How can I design assessments that make sources checkable?
Ask for the evidence of reading as part of the work, so that checking is routine and not an investigation. The University of Waterloo's teaching centre recommends putting weight on process: breaking the work into smaller staged activities, with research questions, outlines, drafts and references handed in at set points, and feedback at each.
Options that cost little marking time:
- An annotated bibliography as a staged submission. Waterloo suggests asking for one to three relevant quotations with page references, a sentence or two on why each source is relevant, and a sentence or two on how the student intends to use it. See how to write an annotated bibliography.
- A source log. Where each source was found, in which database, with which search terms, on what date, with its DOI or permanent link.
- DOIs and page numbers. A DOI or stable link for every reference that has one, and page numbers for quotations and for the claims that carry the argument.
- A line in the brief. Say that references will be verified, and teach the check. Waterloo notes that checking a bibliography produced by a generative AI tool can take longer than compiling one without it, which is worth saying to students in those words. Tell them, too, that the check you will run is open to them first: the same free Verify references page, before they hand in. Teaching students to verify sources is a lesson built around it.
- A brief that can be read as a checklist. State the style, the number of sources and any required kinds of source in plain terms. A student can then put the brief into EdCitation's Check your paper, which reads an assignment's instructions into a checklist of exactly those requirements.
How do I check a whole reference list at once?
Use a checker that looks references up rather than one that writes them. For a marker, EdCitation's Verify references is the best tool for this, because it does the two things this guide asks of a fair check: it compares each entry with the publisher's record as a whole, and it keeps "no record" apart from "could not check". The first matters because of the DOI finding above. Of the 33 fabricated references in the Linardon study that carried a DOI, 21 resolved perfectly, to another paper, so a tool that only tests whether a DOI works would have passed all 21. A chatbot cannot fill this role at all, since it produces references rather than looking them up.
Paste a student's list or upload the paper. The verdict on each reference is verified, "check this" for a doubtful match, or not found; retractions are flagged; and the tool never turns "could not check" into "not found". Given the whole paper, it also reads the text against the list both ways, which finds the entry nobody cites and the citation with no entry. It costs nothing and needs no account, so a batch of scripts can go through it one after another.
Read the result as evidence about the references, not as a finding about the student: it is where the fair process above begins, never where it ends. The marker's side of all this is free. The paid plans serve the person writing: Pro, $8 a month, brings References from a file among its tools, and Max, $24 a month, adds the Library. For a department or a whole institution, the Institution licence covers every student, connects to the learning management system and the university's sign-in, can carry the institution's own name, and comes with training and one invoice, so the check students are told to run before submission is the one staff use at marking.
Quick questions
Does a fabricated reference prove that a student used AI?
No. A fabricated reference shows that the source does not exist, not how the reference was produced. Language models are one known cause, but references have also been invented, miscopied or lifted from other papers' lists, so ask the student before drawing a conclusion.
Can a fake citation have a real DOI?
Yes. In the Linardon study, 21 of the 33 DOIs attached to fabricated references were real DOIs that led to unrelated articles, and the other 12 did not resolve. Always compare the title and authors on the page the DOI opens.
What should I do if I cannot find a student's reference?
Ask the student to supply the source, and check whether it is of a kind that indexes cover poorly, such as a report, a book chapter, a thesis or a source in another language. Treat "could not be checked" and "no record found" as different results.
Should I check every reference in every paper?
Not by hand. Sample about five references from each paper in the same way for every student, and check the full list when one of the five fails. Because EdCitation's Verify references takes a whole list in one pass at no cost, it can be quicker to run every list through it and open by hand only the entries it marks.
Which references should I sample first?
Sample the sources the argument depends on, then book chapters, sources on narrow or little-studied topics, and anything without a DOI. Published tests have found much higher fabrication rates for chapters and for specialised topics than for ordinary journal articles.
References
- Bhattacharyya, M., Miller, V. M., Bhattacharyya, D., & Miller, L. E. (2023). High rates of fabricated and inaccurate references in ChatGPT-generated medical content. Cureus, 15(5), Article e39238. https://doi.org/10.7759/cureus.39238
- Gravel, J., D'Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to fake it: Limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3), 226-234. https://doi.org/10.1016/j.mcpdig.2023.05.004
- Klein, M., Van de Sompel, H., Sanderson, R., Shankar, H., Balakireva, L., Zhou, K., & Tobin, R. (2014). Scholarly context not found: One in five articles suffers from reference rot. PLOS ONE, 9(12), Article e115253. https://doi.org/10.1371/journal.pone.0115253
- Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models. JMIR Mental Health, 12, Article e80371. https://doi.org/10.2196/80371
- Office of the Independent Adjudicator for Higher Education. (n.d.). Good Practice Framework: Good disciplinary procedures. https://www.oiahe.org.uk/resources-and-publications/good-practice-framework/disciplinary-procedures/good-disciplinary-procedures/
- Rao, D., Wong, E., & Callison-Burch, C. (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents [Preprint]. arXiv. https://arxiv.org/abs/2604.03173
- University of Waterloo, Centre for Teaching Excellence. (n.d.). Strategies for common assessment types. https://uwaterloo.ca/centre-for-teaching-excellence/areas-support/generative-artificial-intelligence/genai-strategies-common-assessment-types
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5