Ask an AI tool for sources and it will hand you a tidy list: authors, year, title, journal, volume, pages, often a DOI. Some of those papers exist. Some do not, and nothing on the page tells you which.
This is not a rare glitch. It has been measured in peer-reviewed studies across several versions of ChatGPT, and the rate has fallen without reaching zero. An invented reference is also easy to catch once you know where to look, and the checks below take a minute each. For a list of thirty, EdCitation's free Verify references runs the lookups on every entry together.
Why does ChatGPT make up references?
ChatGPT makes up references because a language model writes the most likely next words and does not look anything up. Unless it is connected to a search tool, there is no database of papers behind the answer, only patterns learned from a very large amount of text.
A reference is the easiest shape in academic writing to imitate
A reference is one of the most regular patterns there is: surnames and initials, a year in brackets, a title in the vocabulary of the field, a journal, a volume and a page range. A model that has read millions of reference lists can produce that shape perfectly. Whether a paper with that exact title was ever published is a separate fact, and the model has no step at which it checks.
So the parts are usually real while the whole is not. Gravel et al. (2023) asked ChatGPT 20 medical questions and checked the 59 references it produced: 41 of them, 69%, were fabricated, and yet 56 of the 59 named authors with real publications on a related topic. Of the fabricated references, 71% were attributed to a known medical journal, with a year, a volume and a page range that fitted that journal but belonged to an unrelated article. The authors describe the fakes as looking as though they had been assembled by merging existing references, which is a fair description of what the model is doing.
Why the model does not say "I am not sure"
There is also a reason the tool rarely admits doubt. Kalai et al. (2025) argue that the way models are trained and scored rewards a confident guess over an admission of uncertainty: models are "optimized to be good test-takers", and on a scored test a guess beats a blank. Nothing in that process gives the model a reason to stop and say that it cannot find the paper it is about to cite.
How often are AI-generated references fake?
In published tests, between about one in five and about seven in ten of the references written by ChatGPT did not exist, depending on the model version, the field and the topic. Four studies that checked every reference by hand give the clearest picture.
| Study | What was tested | References checked | Fabricated | Real, but with errors |
|---|---|---|---|---|
| Gravel et al. (2023), Mayo Clinic Proceedings: Digital Health | ChatGPT answering 20 medical questions | 59 | 69% (41 of 59) | 8 of the 11 real journal articles were miscited |
| Bhattacharyya et al. (2023), Cureus | ChatGPT-3.5 writing 30 short medical papers | 115 | 47% | 46% of all references |
| Walters and Wilder (2023), Scientific Reports | GPT-3.5 and GPT-4 writing 84 literature reviews on 42 topics | 636 | 55% (GPT-3.5), 18% (GPT-4) | 43% (GPT-3.5) and 24% (GPT-4) of the real ones |
| Linardon et al. (2025), JMIR Mental Health | GPT-4o writing 6 literature reviews, June 2025 | 176 | 19.9% | 45.4% of the real ones |
What makes the rate go up or down
Three things move it, and none of them makes the problem go away.
The model version. Newer models do better. In the Bhattacharyya study only 7% of the references were both real and accurate. Two years later, Linardon et al. (2025) found that 19.9% of GPT-4o's references were fabricated and that 45.4% of the real ones carried errors, so nearly two thirds were still fabricated or wrong in some detail.
The topic. In that same study, 4 of 68 references on major depressive disorder were fabricated, 6%, against 17 of 60 for binge eating disorder, 28%, and 14 of 48 for body dysmorphic disorder, 29%. The less that has been written on a subject, the more the model fills in. Student topics are usually narrow, which puts them at the wrong end of that range, and is the strongest reason to find sources in a search of published work, such as EdCitation's Find sources, rather than ask a model to supply them.
The kind of source, and how narrow the question is. Walters and Wilder (2023) report particular trouble with book chapters: GPT-4 fabricated 70% of the book chapters it cited, against 18% of its references overall. Linardon et al. (2025) found the same effect from a narrower prompt: a prompt for a specialised review of digital interventions for binge eating disorder drew a 46% fabrication rate, against 17% when the prompt asked for a general overview of the disorder.
Do AI tools with web search still invent references?
Yes, though less often. A tool that searches the web can point at pages that exist, but it still writes its answer with a language model, and the citation can go wrong between the search and the sentence.
Rao, Wong and Callison-Burch (2026) tested ten search-backed models and deep research agents on more than 53,000 citation URLs, and three models on a further 168,000 URLs across 32 academic fields, in a preprint not yet peer reviewed. Between 3% and 13% of the cited URLs were hallucinated, meaning they had no record in the Wayback Machine and probably never existed. Between 5% and 18% did not resolve at all. Deep research agents cited more sources per question than ordinary search-backed chatbots, and hallucinated URLs at higher rates, so a longer bibliography is not a safer one.
A link that opens is not the end of the check either. The page has to say what your sentence claims. Open it and read the relevant part.
What does a DOI prove, and what does it not prove?
A DOI proves only that somebody registered that string for something, not that it belongs to the reference printed next to it. The DOI Foundation describes a DOI as a persistent identifier that redirects you to the object it points to, and adds that persistence comes from "organizations, not technology". None of that is a statement about the reference in your draft.
The numbers make the point. Of the 33 fabricated references that carried a DOI in the Linardon et al. (2025) study, 21 had a valid DOI that led to an unrelated article, and 12 did not work at all. Bhattacharyya et al. (2023) found the same for the other common identifier: an incorrect PubMed ID was the most frequent error of all, appearing in 93% of the papers ChatGPT wrote.
The exception runs the other way too. A link that fails does not prove that a source was invented. Klein et al. (2014) analysed over 3.5 million articles from 1997 to 2012 and found reference rot in one in five: a web reference that has stopped resolving, or no longer shows what it showed on the day it was cited. In an older reference list a dead URL is ordinary; a dead DOI on a paper supposedly published last year is not.
What happens when invented references are handed in?
The best-documented case is a legal one. In Mata v. Avianca, two New York lawyers and their firm were ordered to pay a $5,000 penalty for filing a brief built on court decisions that ChatGPT had produced and that did not exist.
The brief cited opinions complete with quotations and citations. When the court and the other side could not find them, one lawyer asked ChatGPT whether a case was real, and put the screenshots in the record: the tool answered that the cases were real and could be found through Westlaw, LexisNexis and the Federal Reporter. On 22 June 2023, Judge P. Kevin Castel of the Southern District of New York imposed the penalty and ordered the lawyers to write to their client and to each of the six judges who had been falsely named as the author of a fake opinion.
The judge was careful about where the fault lay. The order says there is "nothing inherently improper" about using a reliable artificial intelligence tool for assistance; the failure was that nobody checked, and that the lawyers stood by the fake opinions after the court questioned them. Note in particular what did not work: asking the chatbot to confirm its own output.
In coursework the principle is the same. A reference says that a source exists and that you used it. An invented one can be treated as fabrication under an academic integrity policy, whoever or whatever typed it. Universities differ on how they class it, so read your own institution's wording rather than assuming.
How do I catch an invented reference?
Check each reference against the publisher's record, not against the AI tool. For one reference, this is the order that finds a fake fastest.
- Resolve the DOI. Put it after
https://doi.org/in your browser. If it fails, or lands on a paper with a different title or different authors, the reference is wrong. - Search the exact title. Enclose the full title in quotation marks and search Google Scholar, Crossref or PubMed. A published paper's exact title nearly always turns up. Near misses with other words do not count.
- Check the issue. Open the journal's page for that volume and look at the page range. Wrong volume, pages and year were among the most frequent errors in the Bhattacharyya study, even in references to real papers.
- Check the authors. Look at the first author's publication list. A real researcher paired with a title they never wrote is the classic mix, and it is what 56 of the 59 references in the Gravel study looked like.
- Open the source and read it. Confirm that it says what your sentence says; a real paper that does not support your claim is still a problem.
- Separate "not found" from "could not check". Reports, theses and older books are missing from many indexes and can still be real.
Steps 1 and 2 are the lookups EdCitation's Verify references makes for a whole list. Step 5, reading the source, is always yours, and so is the judgement in step 6 on a report or thesis the indexes do not hold. Never ask the AI tool whether its own reference is real.
What "not found" does not prove
A missing record is evidence, not proof, and how strong it is depends on where you looked. For a paper said to have appeared recently in a well-known journal, nothing in Crossref, PubMed and Google Scholar together is strong evidence that it does not exist. For a government report, a thesis, a conference paper, an older book chapter or a source published in another language, a blank result usually means the indexes do not cover it. Those are checked at the source: the organisation's own website, a library catalogue, the repository. How to check whether a reference is real works through both cases in detail.
How do I check a whole reference list at once?
Put the list through a checker that never writes a reference, only looks each one up. EdCitation's Verify references is the best tool for this, for the reason this article opened with: a chatbot produces likely text, and when one was asked in Mata v. Avianca whether its cases were real, it said yes. Verify references can report only what the publisher's record holds, or that an index did not answer. Upload your paper or paste your reference list and every entry receives a verdict; with the whole paper uploaded, each in-text citation is also matched to the list, and each entry back to the text. It is free and needs no account. The paid plans are for other jobs: Pro at $8 a month brings References from a file and Mechanics QA, Max at $24 a month the Library and Theoretics QA.
| What the check says | What it tells you | What to do |
|---|---|---|
| Verified | The record has this work | Read it for the claim you cite it for |
| Doubtful, shown as "check this" | Something in the entry needs a second look | Compare the entry with the record and correct it |
| Not found | No record where the check looked | For a recent journal article, treat it as invented; for a report, thesis or old book, check at the source |
| Could not check | An index did not answer | Try again, or check by hand; it is never reported as "not found" |
| Retracted (a flag on the entry) | The journal has withdrawn the paper | Replace it, unless the retraction is your subject |
If a reference fails, do not delete it and hope. The sentence it supported now has nothing under it. Search for the claim itself in Find sources, read what comes back, and let Cite a source build the new reference from its DOI in the style your course uses. If you used an AI tool in a way your course allows, say so and cite the tool as your course asks.
Quick questions
Does ChatGPT still make up references in 2026?
Yes, less often than in 2023. The most recent peer-reviewed test cited here, run on GPT-4o in June 2025, found 19.9% of references fabricated, and a 2026 preprint found that 3% to 13% of the URLs cited by search-backed AI tools had probably never existed.
Why does a fake reference have a real DOI?
A language model can reproduce a DOI it has seen and attach it to a reference it has invented. In the Linardon study, 21 of the 33 DOIs on fabricated references were real DOIs that belonged to unrelated articles, so check that the page the DOI opens carries the same title and authors.
Can I ask the AI tool whether its reference is real?
No. The tool answers that question the same way it wrote the reference, by producing likely text. In Mata v. Avianca, ChatGPT told a lawyer that non-existent cases were real and could be found in the standard legal databases.
Which subjects get the most fabricated references?
The narrower and less written-about the topic, the higher the rate. In one 2025 test, fabrication ran at 6% on major depressive disorder and at 28% and 29% on two less studied conditions, and book chapters were fabricated far more often than journal articles.
Is it misconduct to hand in a reference that an AI tool invented?
It can be. A reference list states that each source exists and that you used it, and universities deal with invented sources under their academic integrity policies. Read your own institution's policy, and check every reference before you submit; EdCitation's free Verify references takes the whole list at once.
References
- Bhattacharyya, M., Miller, V. M., Bhattacharyya, D., & Miller, L. E. (2023). High rates of fabricated and inaccurate references in ChatGPT-generated medical content. Cureus, 15(5), Article e39238. https://doi.org/10.7759/cureus.39238
- DOI Foundation. (n.d.). What is a DOI? https://www.doi.org/the-identifier/what-is-a-doi/
- Gravel, J., D'Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to fake it: Limited responses and fabricated references provided by ChatGPT for medical questions. Mayo Clinic Proceedings: Digital Health, 1(3), 226-234. https://doi.org/10.1016/j.mcpdig.2023.05.004
- Kalai, A. T., Nachum, O., Vempala, S. S., & Zhang, E. (2025). Why language models hallucinate [Preprint]. arXiv. https://arxiv.org/abs/2509.04664
- Klein, M., Van de Sompel, H., Sanderson, R., Shankar, H., Balakireva, L., Zhou, K., & Tobin, R. (2014). Scholarly context not found: One in five articles suffers from reference rot. PLOS ONE, 9(12), Article e115253. https://doi.org/10.1371/journal.pone.0115253
- Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models. JMIR Mental Health, 12, Article e80371. https://doi.org/10.2196/80371
- Mata v. Avianca, Inc., No. 22-cv-1461 (PKC) (S.D.N.Y. June 22, 2023). Opinion and order on sanctions. https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf
- Rao, D., Wong, E., & Callison-Burch, C. (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents [Preprint]. arXiv. https://arxiv.org/abs/2604.03173
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5