Fake citations in a literature review do more damage than anywhere else in a dissertation, because every later chapter stands on what it says the field has found. The invented paper is now the easy catch: a minute per entry by hand, and EdCitation's free Verify references looks up a whole list at once. The harder cases are real papers cited for things they do not say, quotations that are not in the source, a review assembled from abstracts, and a reference list that is genuine and fifteen years old.
This guide is for the marker or supervisor. It does not repeat the signs of an invented reference or the fair process, which are in how to spot fabricated references in student work, or the single-reference check, which is in how to check whether a reference is real.
What kinds of fake citation appear in a literature review?
Six kinds, and only the first two are caught by looking a reference up. The rest need the source open.
| Failure | What it looks like | In published tests | How it is caught |
|---|---|---|---|
| Invented reference | Real authors, real journal, a title nobody wrote | 55% of GPT-3.5's references and 18% of GPT-4's (Walters and Wilder, 2023); 19.9% of GPT-4o's (Linardon et al., 2025) | Resolve the DOI and compare; search the exact title; a whole list at once with EdCitation's Verify references |
| Real paper, wrong details | Year, volume, pages or DOI do not fit | 24% of GPT-4's real references (Walters and Wilder, 2023); 45.4% of GPT-4o's (Linardon et al., 2025) | Compare with the publisher's record. A referencing error, not fabrication |
| Citation-claim mismatch | The paper does not say what it is cited for | 25.4% in medical journals (Jergas and Baethge, 2015); 25% in five leading science journals (Smith and Cumberledge, 2020) | Read the cited passage against the sentence. Human reading only |
| Quotation not in the source | Quoted words the source does not contain, or was itself quoting | Not counted separately in any study read here | Search the source's full text for a phrase of the quotation |
| Review written from abstracts | Only claims the abstract makes; no page numbers; no limits or subgroups | Not measured directly | Ask for a claim only the full text supports, with its page |
| Stale list | Real references, the newest years old | No published measure found | Sort by year; compare the newest with the field's recent work, found with Find sources limited to recent years |
Why the invented paper is the easy catch
Whether a paper exists is a matter of record. The journal name is no guide: Walters and Wilder (2023) found that only 2% of the fabricated article citations named a journal that did not exist. The title is, because a published paper's exact title in quotation marks turns up in Crossref, Google Scholar or PubMed. A resolving DOI is not enough either: when GPT-4o attached a DOI to a fabricated reference, Linardon et al. (2025) found that 21 of the 33 opened unrelated articles and 12 did not work. Compare the title and authors on the page that opens.
Why the mismatch is the hard one
Citing a real paper for a claim it does not make is the ordinary condition of published research, and only reading catches it. Jergas and Baethge (2015) pooled 28 studies of quotation accuracy in medical journals and estimated that 25.4% of citations were quotation errors: 11.9% major, meaning the cited paper was not in accordance with the claim at all, and 11.5% minor. Smith and Cumberledge (2020) checked 248 citations from Science, Nature, Nature Communications, Science Advances and PNAS: 75% fully substantiated, 3.6% partially, 8.5% unsubstantiated, and 12.9% impossible to substantiate, where the sentence made no checkable claim. Pavlovic et al. (2021) traced 4,912 citations to 13 highly cited papers, and the most common problem, at 38.4%, was the citation of a finding the paper did not contain. The studies disagree about the number more than the problem: Mogull (2017) recalculated the medical figure at 14.5%, with a range of 4.1% to 34.2% across specialties. Take one in four as the working figure for work that editors have reviewed.
Which references should I check first, and how many?
Check the references the argument rests on first, then the kinds the indexes cover worst, and check the same number in every script.
- The key-claim references. Mark the three to five sentences the argument stands on. Check each reference for existence, and for these alone read the cited passage.
- Every entry without a DOI. Walters and Wilder (2023) found 70% of the book chapters cited by both models fabricated, against 18% of GPT-4's references overall. A chapter, report or thesis in a generated list deserves the time.
- The newest entries. An article said to be in a known journal in the last two years, with no record in Crossref or PubMed, is the strongest "not found" a marker can get, and the newest date shows when the search stopped.
- Any journal you do not recognise. Rare among fakes, so it stands out. Walters and Wilder (2023) found at least two of GPT-3.5's fabricated citations placed in journals whose publishers have been identified as predatory.
- One at random, so the sample is not only the entries that look odd.
How many is enough
Smith and Cumberledge (2020) report that two reviewers took months to check 250 citations, so a full claim check is not a marking task. The existence check is cheap: the whole list, every time, pasted into Verify references, which leaves your reading time for the claims. The claim check is dear: five cited passages in a 40-reference undergraduate review, ten in a dissertation, every key-claim sentence in a thesis chapter you supervise. If one fails, widen the sample. Use the same numbers for every script.
How do I check one claim against the abstract and then the full text?
Open the abstract first, because it settles most claims in a minute, and open the full text when it does not. The sentence below is a teaching example; the source is real.
The student writes: "Around half of the references produced by ChatGPT are fabricated (Walters & Wilder, 2023)."
- Confirm the reference exists. The DOI resolves to Scientific Reports, volume 13, article 14045, by those authors, with that title.
- Read the abstract against the sentence. It gives two figures, 55% for GPT-3.5 and 18% for GPT-4. "Around half" is true of one model and not of the other, and the difference is the paper's main finding. Verdict so far: partially substantiated, in Smith and Cumberledge's terms.
- Read the full text where the number lives. Table 3 gives the denominators, 222 and 414 references, and adds: 73% of GPT-3.5's journal articles were fabricated, 70% of the chapters from both models, 23% and 8% of the books, in short papers of the kind set in first-year composition courses. A sentence the source supports: in an April 2023 test, 55% of the references GPT-3.5 wrote for short literature reviews and 18% of GPT-4's did not exist (Walters & Wilder, 2023).
- Record it. The sentence, the reference, what the abstract and the full text say, the verdict and the date. That note is what you show the student.
The quotation check
Suppose the same student writes that Walters and Wilder (2023) conclude that ChatGPT is "confidently wrong". The phrase is in the full text, in the Discussion, and the paper attributes it to an earlier study by Gravel and colleagues. The words exist; the attribution is wrong; a marker who stops at the search box would pass it. A quotation check asks whether the phrase is in the source and whether it is the source's own. Ask for the page number; Smith and Cumberledge (2020) end by asking journals to require page numbers for the same reason.
How do I tell a review written from abstracts, or from a stale list?
Ask for something only the full text contains, and read the dates.
The review written from abstracts
A review written from abstracts cites correctly and reads thinly: every figure is one the abstract gives, there are no page numbers, and limitations and negative findings never appear. The check is one question per key paper: name a result, a sample size or a caveat only the body contains, and ask where it is. Greenberg (2009) found the published form of this in one field's literature: 17 citations, from seven papers, to 12 meeting abstracts cited as full papers. A model prompted for a review does a version of the same thing. Chelli et al. (2024) gave ChatGPT and Bard the inclusion criteria of 11 published systematic reviews: 28.6% of GPT-4's references and 91.4% of Bard's were hallucinated, and only 16 of GPT-4's 119, 13.4%, were papers the human reviewers had included.
The stale list
A stale list is real and out of date, and no checker flags it, because every entry in it is genuine. Sort the references by year, look at the newest three, and compare them with what the field published in the last two or three years. The quickest way to see that recent work is a topic search in EdCitation's Find sources with the years filter set to the last three: if it returns a run of relevant papers newer than anything in the list, the search behind the review probably stopped early. None of the studies read for this guide measured how old the references in a generated review are, so there is no figure to give, only the check. Walters and Wilder (2023) note that for older works ChatGPT often gave the date a paper was posted online rather than its publication date, so a year in a generated list can be the year of a repost. Chelli et al. (2024) found the models' lists leaned towards American first authors: 44% for GPT-3.5 and 33% for GPT-4, against 16.5% in the human-written reviews.
What does a failed check prove?
A failed existence check proves that the record has no such paper; a failed claim check proves that the source does not say what the sentence says. Neither proves how the citation got there.
Keep "not found" apart from "could not check", because reports, theses and older chapters are missing from the indexes and can still be real. A mismatch is weaker evidence still: it runs at about one in four in edited journals, and Pavlovic et al. (2021) found that one fifth of the inaccurate citations they traced came from chains, a citing author copying a claim and its reference from an earlier paper that had already got it wrong. A student who copies a review's citation inherits its error in good faith. Mark it as a referencing fault, and ask.
How do I check every reference in a literature review at once?
Hand the list to something that looks up and never composes. For the existence and retraction check of a whole list, EdCitation's Verify references is the best tool, since it takes every entry to the publisher's record (Crossref, DataCite, PubMed, Open Library, and Retraction Watch for withdrawals), and no chatbot can do that: Walters and Wilder (2023) report that ChatGPT often answers inaccurately when asked to verify a work it has cited, and in the Chelli study Bard's own list was 91.4% hallucinated. Paste the list, or upload the review itself. Each entry is returned as verified, doubtful (labelled "check this") or not found, with retracted papers marked; when an index gives no answer, the entry says "could not check" and never "not found". It is free and needs no account.
Uploading the whole review adds a second check at no cost: each citation in the review is looked for in the list, and each entry in the list is looked for in the text. That matters here because among the irregularities Walters and Wilder (2023) recorded in generated papers were in-text citations with no entry in the list. References from a file, part of Pro at $8 a month, runs the same reference and citation check on a file for the person writing; Max ($24 a month) brings Theoretics QA and the Library. A department that wants every candidate and supervisor on the same check can take the Institution licence, which covers every student and integrates with the LMS and sign-in.
The claim-mismatch check is human reading, and no tool on this site does it for you: whether a paper says what the student says is answered by opening it.
Quick questions
Can a citation checker detect a real paper cited for the wrong claim?
No. A checker looks a reference up in the publisher's record and says whether the paper exists, whether its details are right and whether it has been retracted. Whether it supports the sentence is found only by reading the cited passage.
How many references should I check in a dissertation literature review?
Run the whole list through an existence check, which EdCitation's Verify references does free in one pass, then read the cited passage for about ten references: those behind the key claims, any without a DOI, the newest, and one at random. Widen the sample if any fails, and use the same numbers for every candidate.
What is a citation-claim mismatch?
A citation-claim mismatch is a real source cited for something it does not say: a finding it does not contain, a result read the wrong way, or a caveat left out. Studies of published medical and science journals put the rate at about one citation in four.
Does a citation-claim mismatch prove the student used AI?
No. Mismatches run at about a quarter of citations in edited journals, and a fifth of them are inherited from an earlier paper's error. Mark it as a referencing fault and raise it with the student.
References
- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J. L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research, 26, Article e53164. https://doi.org/10.2196/53164
- Greenberg, S. A. (2009). How citation distortions create unfounded authority: Analysis of a citation network. BMJ, 339, Article b2680. https://doi.org/10.1136/bmj.b2680
- Jergas, H., & Baethge, C. (2015). Quotation accuracy in medical journal articles: A systematic review and meta-analysis. PeerJ, 3, Article e1364. https://doi.org/10.7717/peerj.1364
- Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: Experimental study. JMIR Mental Health, 12, Article e80371. https://doi.org/10.2196/80371
- Mogull, S. A. (2017). Accuracy of cited "facts" in medical research articles: A review of study methodology and recalculation of quotation error rate. PLOS ONE, 12(9), Article e0184727. https://doi.org/10.1371/journal.pone.0184727
- Pavlovic, V., Weissgerber, T., Stanisavljevic, D., Pekmezovic, T., Milicevic, O., Lazovic, J. M., Cirkovic, A., Savic, M., Rajovic, N., Piperac, P., Djuric, N., Madzarevic, P., Dimitrijevic, A., Randjelovic, S., Nestorovic, E., Akinyombo, R., Pavlovic, A., Ghamrawi, R., Garovic, V., & Milic, N. (2021). How accurate are citations of frequently cited papers in biomedical literature? Clinical Science, 135(5), 671-681. https://doi.org/10.1042/CS20201573
- Smith, N., Jr., & Cumberledge, A. (2020). Quotation errors in general science journals. Proceedings of the Royal Society A, 476(2242), Article 20200538. https://doi.org/10.1098/rspa.2020.0538
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5