# How often does ChatGPT make up references? What studies found

**How often do ChatGPT and other AI chatbots make up references, according to published studies?** Published studies found AI chatbots inventing anywhere from none to more than half of the references they wrote, depending on the model, the field and the test. GPT-3.5 fabricated 55% in April 2023, GPT-4o 19.9% in June 2025, and GPT-5.3, with web search off, 27% in March 2026. Every study that checked the details also found errors in real references.

Published 2026-09-26 by EdCitation. https://edcitation.com/newsletter/what-studies-say-about-ai-fabricated-references

Since early 2023, researchers have asked AI chatbots for references and looked every one up. The answer to "how often does ChatGPT make up citations?" depends on which model, which month and which field, and the published figures run from none at all to more than nine in ten.

This roundup, published by EdCitation, gives each figure as its paper states it, with the model and the date it was tested, because a result about a 2023 model says little about the one you used last night. We add studies as they appear; this version was checked in September 2026. Why models invent references at all is in [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references), and the check for your own list is EdCitation's free [Verify references](https://edcitation.com/verify-references), which looks each entry up and never writes one.

## How often does ChatGPT make up references?

In the studies we could read, ChatGPT fabricated between 18% and 69% of its references in 2023, 19.9% in June 2025 (GPT-4o) and 27% in March 2026 (GPT-5.3, web search off). One small January 2026 test found none of its 31 references invented. The rate has fallen with each generation of models, and no study that checked every detail found a model that got every reference right.

Two numbers matter: the share of references that do not exist, and the share of real ones with a wrong author, year, volume or DOI. Either can cost marks.

## The studies, in one table

Peer-reviewed studies, and preprints marked as such, that asked a chatbot for references and checked each one, in order of test date.

| Study | Model and when tested | What was asked | References | Did not exist | Real, with errors |
| --- | --- | --- | --- | --- | --- |
| [Gravel et al. (2023)](https://doi.org/10.1016/j.mcpdig.2023.05.004) | ChatGPT 3.5, Feb 2023 | 20 medical questions | 59 | 69% (41) | 8 of 11 real articles |
| [McGowan et al. (2023)](https://doi.org/10.1016/j.psychres.2023.115334) | ChatGPT, Mar 2023 | Psychiatry literature search | 35 | Only 2 were real; 21 were pastiches | 12 resembled real papers |
| [Walters and Wilder (2023)](https://doi.org/10.1038/s41598-023-41032-5) | GPT-3.5 and GPT-4, early Apr 2023 | 84 short literature reviews | 636 | 55% and 18% | 43% and 24% of real ones |
| [Bhattacharyya et al. (2023)](https://doi.org/10.7759/cureus.39238) | ChatGPT-3.5, 12 Apr 2023 | 30 short medical papers | 115 | 47% | 46% of all |
| [Buchanan et al. (2024)](https://doi.org/10.1177/05694345231218454) | GPT-3.5 and GPT-4, 2023 | Every *Journal of Economic Literature* topic | Not in abstract | "More than 30%" (GPT-3.5), slightly less (GPT-4) | Not in abstract |
| [Chelli et al. (2024)](https://doi.org/10.2196/53164) | GPT-3.5, GPT-4, Bard, Jul 2023 | Repeat 11 systematic reviews | 471 | 39.6%, 28.6%, 91.4% | Counted with the fabricated |
| [Mugaanyi et al. (2024)](https://doi.org/10.2196/52935) | GPT-3.5, Jul-Aug 2023 | Introductions, 10 topics | 102 | 27.3% (sciences), 23.4% (humanities) not confirmed | DOI right in 32.7% and 8.5% |
| [Pastucha et al. (2026)](https://doi.org/10.12659/MSM.950916) | ChatGPT 4.0, Gemini 1.5 Pro (Jul 2024); 4.1, 2.5 Pro, search on and off (Jun 2025) | References on 25 ENT guidelines, 3 days | 1,947 | 1% to 22% (ChatGPT), 0% to 26% (Gemini) | Best accuracy score 51% |
| [Aydin et al. (2026)](https://doi.org/10.32708/uutfd.1870116) | Free ChatGPT, Gemini, Claude, Copilot, Nov 2024 and Jan 2026 | 5 medical-education introductions | Not totalled | ChatGPT: 9 of 14 (2024), 0 of 31 (2026), our count | Not checked |
| [Cabezas-Clavijo and Sidorenko-Bautista (2026)](https://doi.org/10.1515/jdis-2025-0326) | 8 free chatbots, 7-9 Feb 2025 | 10 APA references for a final-year project, 5 fields | 400 | 39.8% wrong or fabricated | 33.8% partly correct |
| [Linardon et al. (2025)](https://doi.org/10.2196/80371) | GPT-4o, Jun 2025 | 6 literature reviews | 176 | 19.9% | 45.4% of real ones |
| [Naser (2026)](https://arxiv.org/abs/2603.03299), preprint | 10 models by API, date not stated | References in 4 fields | 69,557 | 11.4% to 56.8% unverified | Not reported |
| [Seifi and Seyfi (2026)](https://doi.org/10.1097/CCE.0000000000001474) | GPT-5.3, DeepSeek-V3, Grok-4, search off, 10 Mar 2026 | 10 references on each of 10 topics | 300 | 27%, 8%, 50% | Any error: 69%, 23%, 73% |

### How to read the table

Each paper's own definition holds, and they differ. Chelli et al. (2024) counted a reference as hallucinated when two of its title, first author and year were wrong, so a garbled real paper is counted with the fakes. Cabezas-Clavijo and Sidorenko-Bautista (2026) report "wrong or fabricated" as one figure. Naser (2026) counts a reference as unverified when an automatic search of three databases found no match, which can include real works the databases missed. Aydin et al. (2026) checked only existence and report counts prompt by prompt; the ChatGPT totals are ours, added from their table. Seifi and Seyfi (2026) count any inaccuracy, fabricated references included.

## What do the studies agree on?

**Newer models invent fewer.** Walters and Wilder (2023) found 55% of GPT-3.5's references fabricated against 18% of GPT-4's, tested the same week. Chelli et al. (2024) saw the same order three months later. Linardon et al. (2025) put GPT-4o at 19.9%.

**Real references are often wrong.** Bhattacharyya et al. (2023) found only 7% of ChatGPT-3.5's references both real and accurate. Two years later, 45.4% of GPT-4o's real references carried an error, most often the DOI (Linardon et al., 2025).

**The field and the source type matter.** In the preprint version of their study, Cabezas-Clavijo and Sidorenko-Bautista (2025) found 52.5% of references wrong or fabricated in engineering against 26.3% in the humanities, and 78% of journal-article references wrong or fabricated against 12.9% of books. Mugaanyi et al. (2024) found DOIs right in 32.7% of natural-science references and 8.5% of humanities ones.

**No list is safe unchecked.** Even Aydin et al. (2026), whose results were the most favourable, tell researchers to keep using verification protocols.

## Where do the studies disagree?

They disagree most on the newest models, and the reasons lie in the methods.

### Two 2026 tests, opposite results

Aydin et al. (2026) found every one of the 31 references from the free ChatGPT real in January 2026, against 5 of 14 in November 2024. Seifi and Seyfi (2026), two months later, found GPT-5.3 fabricated 27% of 100 references with web search off, and got some detail wrong in 69%. The first asked for introductions on broad medical-education topics and ran each prompt once; the second asked for ten references on each of ten narrow neurocritical care topics and blocked retrieval. Both can be right: a broad, well-studied topic is the easy case, a specialist one with search off the hard one.

### The same chatbot, different results

Grok and DeepSeek fabricated none of their 50 references each in February 2025 (Cabezas-Clavijo & Sidorenko-Bautista, 2025). In March 2026, Grok-4 fabricated 50% and DeepSeek-V3 8% in neurocritical care (Seifi & Seyfi, 2026). Naser (2026) adds that a newer model is not always better: the preprint reports that generational gains were "not guaranteed".

### Web search on or off

Where a study tested both, search helped without solving the problem. Pastucha et al. (2026) found the two web-search versions more accurate than the four without, yet the best, ChatGPT-4.1 with browsing in June 2025, scored 51% for accuracy. [ChatGPT's search, explained](https://edcitation.com/newsletter/chatgpt-fake-citations-how-to-find-and-fix) covers what changes when it is on.

## What do the studies not show?

They do not give the rate for the model and prompt you used. Know these limits before quoting a figure.

- **Most test retired models.** GPT-3.5, GPT-4, Bard and GPT-4o dominate the table.
- **Most are medical.** Eight of the thirteen studies ask medical or health questions.
- **Many are small.** One run per prompt is common, and Aydin et al. (2026) call their results a snapshot of a non-deterministic system.
- **Definitions differ**, so two percentages side by side may not measure the same failure.

What we could not confirm: how many references Buchanan et al. (2024) checked; the version of ChatGPT and the June 2023 results in McGowan et al. (2023), whose full text we could not open; and the test dates in Naser (2026). A 2026 anatomy study reporting "hallucination rates" for ChatGPT 5.2, Gemini 3 Pro and DeepSeek V3.2 is left out of the table because we could not read how it defines one.

## Do fabricated references reach published papers?

Yes, rarely per reference and visibly per paper. Russinovich et al. (2026), in a preprint, checked accepted papers at four computer science and security conferences and found likely hallucinated references in usually under 1% of entries, yet roughly one in twenty 2025 NeurIPS and USENIX Security papers carried at least two. What happens to a student who hands one in is covered in [the leader guide](https://edcitation.com/newsletter/why-ai-tools-invent-references).

## We ran the studies' own references through EdCitation

We put eight of the reference entries from this page into EdCitation's [Verify references](https://edcitation.com/verify-references), with one more we invented for the test: "Hartwell, D. R., & Mensah, K. (2024). Citation fabrication by large language models across twelve disciplines: A systematic review and meta-analysis. *Research Integrity and Peer Review, 9*, Article 14." No such paper exists; we wrote it to look like the others.

| Entry | Verdict | Reason given, as returned |
| --- | --- | --- |
| Walters and Wilder (2023), and six more journal articles with DOIs | verified | "The DOI resolves to this record and the title matches." |
| Naser (2026), arXiv preprint | verified | "The page is at this address, and its title matches." |
| The invented entry | not found | "No publisher's record or library catalogue entry matches this reference." |

The counts: 8 verified, 0 doubtful, 1 not found, 0 retracted, 0 unchecked. That lookup is the step missing from every study above, where the model writes the reference and nothing checks it.

### What Cite a source returned for four of the DOIs

We gave [Cite a source](https://edcitation.com/cite) four of the studies' DOIs. Each came from Crossref with no retraction notice. One thing was already right, and one needed correcting by hand:

- **Sentence case in APA 7.** For Seifi and Seyfi (2026), the publisher deposited the title in title case, and APA came back as "Hallucination rate of peer-reviewed citations generated by large language models in neurocritical care", the sentence case APA 7 wants. Linardon et al. (2025) and Pastucha et al. (2026) came back the same way, and McGowan et al. (2023) kept "ChatGPT and Bard" capitalised.
- **A record without its volume.** For Cabezas-Clavijo and Sidorenko-Bautista (2026) the publisher's record carries no volume, issue or pages, so APA ended at the journal name and Harvard added "[Preprint]". The journal's own page gives volume 11, issue 2, pages 102-116, which our reference list uses.

That is why EdCitation is the best tool for this job and a chatbot is the wrong one: EdCitation builds a reference only from a record that exists, and where the record is thin, what it returns shows the gap rather than filling it. Verify references and Cite a source are free with no account. [Pro](https://edcitation.com/pricing), $8 a month, adds [References from a file](https://edcitation.com/tools/references-from-a-file), which checks every reference in an uploaded paper and matches each in-text citation to the list; Max, $24 a month, adds Theoretics QA and the Library.

## How do I use these findings on my own reference list?

Treat any reference a chatbot gave you as unchecked until a record says otherwise.

1. **Paste the whole list into [Verify references](https://edcitation.com/verify-references).** Each entry comes back verified, "check this" or not found, with a reason, and "could not check" is never shown as "not found".
2. **Fix what is real but wrong.** In recent tests this group is as large as the invented one, or larger. Rebuild the entry from its DOI in [Cite a source](https://edcitation.com/cite), then correct missing details as above.
3. **Replace what does not exist.** Search the claim in [Find sources](https://edcitation.com/), read what you find, and cite that; [My reference cannot be found](https://edcitation.com/newsletter/my-reference-cannot-be-found-what-now) explains what to try before deciding a source is invented.
4. **Open and read every source you keep**, since no lookup can tell you whether it says what your sentence claims. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) covers that step.

## Quick questions

### What is the most recent study on ChatGPT fake references?

In this roundup, Seifi and Seyfi (2026), who tested GPT-5.3 on 10 March 2026 with web search off and found 27% of its neurocritical care references fabricated. Aydin et al. (2026) found all 31 of ChatGPT's references real in January 2026, on broader topics.

### Which AI chatbot makes up the fewest references?

No chatbot wins every study. Grok and DeepSeek fabricated none in one February 2025 test, yet Grok-4 fabricated half its references in a March 2026 one, on a different field and prompt.

### Is the 55% figure for ChatGPT still true?

No. The 55% is GPT-3.5 in April 2023 (Walters & Wilder, 2023). Later models did better, and later studies should be quoted for current tools, each with its model and test date.

### Does web search stop AI from inventing references?

It helps without solving it. In one 2025 test the search-enabled versions of ChatGPT and Gemini were the most accurate, and the best still scored 51%. EdCitation's free Verify references checks each entry against the publisher's record either way.

## References

- Aydin, M. O., Vatansever, A., & Erer Kafa, S. (2026). From hallucination to precision: A longitudinal analysis of reference accuracy and plagiarism in AI-generated medical literature (2024-2026). *Journal of Uludağ University Medical Faculty, 52*, Article 1870116. [https://doi.org/10.32708/uutfd.1870116](https://doi.org/10.32708/uutfd.1870116)
- Bhattacharyya, M., Miller, V. M., Bhattacharyya, D., & Miller, L. E. (2023). High rates of fabricated and inaccurate references in ChatGPT-generated medical content. *Cureus, 15*(5), Article e39238. [https://doi.org/10.7759/cureus.39238](https://doi.org/10.7759/cureus.39238)
- Buchanan, J., Hill, S., & Shapoval, O. (2024). ChatGPT hallucinates non-existent citations: Evidence from economics. *The American Economist, 69*(1), 80-87. [https://doi.org/10.1177/05694345231218454](https://doi.org/10.1177/05694345231218454)
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2025). *Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate* [Preprint]. arXiv. [https://arxiv.org/abs/2505.18059](https://arxiv.org/abs/2505.18059)
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2026). Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are entirely accurate. *Journal of Data and Information Science, 11*(2), 102-116. [https://doi.org/10.1515/jdis-2025-0326](https://doi.org/10.1515/jdis-2025-0326)
- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. *Journal of Medical Internet Research, 26*, Article e53164. [https://doi.org/10.2196/53164](https://doi.org/10.2196/53164)
- Gravel, J., D'Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to fake it: Limited responses and fabricated references provided by ChatGPT for medical questions. *Mayo Clinic Proceedings: Digital Health, 1*(3), 226-234. [https://doi.org/10.1016/j.mcpdig.2023.05.004](https://doi.org/10.1016/j.mcpdig.2023.05.004)
- Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: Experimental study. *JMIR Mental Health, 12*, Article e80371. [https://doi.org/10.2196/80371](https://doi.org/10.2196/80371)
- McGowan, A., Gui, Y., Dobbs, M., Shuster, S., Cotter, M., Selloni, A., Goodman, M., Srivastava, A., Cecchi, G. A., & Corcoran, C. M. (2023). ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. *Psychiatry Research, 326*, Article 115334. [https://doi.org/10.1016/j.psychres.2023.115334](https://doi.org/10.1016/j.psychres.2023.115334)
- Mugaanyi, J., Cai, L., Cheng, S., Lu, C., & Huang, J. (2024). Evaluation of large language model performance and reliability for citations and references in scholarly writing: Cross-disciplinary study. *Journal of Medical Internet Research, 26*, Article e52935. [https://doi.org/10.2196/52935](https://doi.org/10.2196/52935)
- Naser, M. Z. (2026). *How LLMs cite and why it matters: A cross-model audit of reference fabrication in AI-assisted academic writing and methods to detect phantom citations* [Preprint]. arXiv. [https://arxiv.org/abs/2603.03299](https://arxiv.org/abs/2603.03299)
- Pastucha, M., Skarżyński, H., Kochanek, K., & Jedrzejczak, W. W. (2026). Reference accuracy in large language model chatbots: A metric for inherent misinformation? *Medical Science Monitor, 32*, Article e950916. [https://doi.org/10.12659/MSM.950916](https://doi.org/10.12659/MSM.950916)
- Russinovich, M., Kumar, R. S. S., & Salem, A. (2026). *Phantom references: Hallucinated citations that survive peer review at top-tier conferences* (Version 2) [Preprint]. arXiv. [https://doi.org/10.48550/arXiv.2607.00738](https://doi.org/10.48550/arXiv.2607.00738)
- Seifi, A., & Seyfi, A. (2026). Hallucination rate of peer-reviewed citations generated by large language models in neurocritical care. *Critical Care Explorations, 8*(9), Article e1474. [https://doi.org/10.1097/CCE.0000000000001474](https://doi.org/10.1097/CCE.0000000000001474)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
