# ChatGPT fake citations and how to find and fix them

**How do I find and fix fake citations from ChatGPT?** To find fake citations from ChatGPT, check every reference against the publisher's record: EdCitation's free Verify references looks each one up and flags what does not exist, or follow each DOI and search each exact title yourself. In the most recent peer-reviewed test, 19.9% of the references GPT-4o wrote did not exist. Never ask ChatGPT to confirm its own reference.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/chatgpt-fake-citations-how-to-find-and-fix

Fake citations from ChatGPT look like every other citation: real names, a plausible title, a journal that exists, a year, often a DOI. The only way to tell is to look each one up, and the way to have fewer to look up is to understand how ChatGPT produces them.

This guide is about ChatGPT alone: what OpenAI says, what the tests found for each version, the five ways its references go wrong, and how to ask for ones you can check. The checking procedure is the same for every AI tool and lives in [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references) and [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real). If a ChatGPT list is already in your draft, EdCitation's free [Verify references](https://edcitation.com/verify-references) looks up every entry in the publisher's record in one pass, the step ChatGPT itself never takes.

## What does OpenAI say about sources and citations in ChatGPT?

OpenAI says that ChatGPT can cite web sources when it searches, and that those citations "can be incomplete, outdated, or incorrect". Both statements are on the help centre page for web search, checked on 22 September 2026 (OpenAI, 2026a).

That page says web search is available on the Free, Go, Plus, Pro, Business, Enterprise and Edu plans, and to people who are not signed in. A response that used search "may include citations"; selecting one opens its source, and a Sources button lists the cited pages. Then comes the caution: open a cited source to check that it supports the answer.

### What OpenAI says about hallucination

OpenAI's research post of 5 September 2025 says plainly that "ChatGPT also hallucinates", and explains why: training and evaluation reward a guess over an admission of uncertainty (OpenAI, 2025a). The post gives the trade in numbers. On the SimpleQA test, gpt-5-thinking-mini declined to answer 52% of questions and was wrong on 26%; the older o4-mini declined on 1% and was wrong on 75%. A reference list is a run of exactly this kind of question, and a model tuned to guess fills it.

### The knowledge cutoff

Without search, ChatGPT writes from training data, and training stops at a date. For GPT-5.6 Sol, the reasoning model OpenAI began rolling out to paid ChatGPT plans on 9 July 2026, the model page gives a knowledge cutoff of 16 February 2026 (OpenAI, 2026c). Nothing after that date can be in a reference written without search, and nothing before it is looked up.

### What OpenAI says about deep research

Deep research is ChatGPT's long-form mode: it searches, reads and writes a report. The help centre says every report includes "citations or source links so you can verify the information" (OpenAI, 2026b). The launch post was more candid: deep research "can sometimes hallucinate facts in responses or make incorrect inferences", at a lower rate than other ChatGPT models by OpenAI's internal measure, and was "often failing to convey uncertainty accurately" (OpenAI, 2025b). OpenAI is telling you to check.

## How often does ChatGPT make up references?

In peer-reviewed tests run between February 2023 and June 2025, between 18% and 69% of the references ChatGPT wrote did not exist, depending on the model version and the task, and the rate fell with each new version. The [leader article](https://edcitation.com/newsletter/why-ai-tools-invent-references) tabulates the four hand-checked studies in full; the rows below show the trend by version.

| Model in ChatGPT | When tested | Study | References that did not exist |
| --- | --- | --- | --- |
| GPT-3.5 | April 2023 | Walters and Wilder (2023), 84 literature reviews | 55% |
| GPT-4 | April 2023 | Walters and Wilder (2023) | 18% |
| GPT-3.5 | July 2023 | Chelli et al. (2024), 11 systematic reviews | 39.6% (55 of 139) |
| GPT-4 | July 2023 | Chelli et al. (2024) | 28.6% (34 of 119) |
| GPT-4o | June 2025 | Linardon et al. (2025), 6 literature reviews | 19.9% (35 of 176) |

The rows are not exactly comparable: Chelli et al. (2024) counted a reference as hallucinated when two of the title, first author and year were wrong. And none of these studies reports switching on web browsing. ChatGPT search did not launch until 31 October 2024 (OpenAI, 2024), so the 2023 tests measure the model writing from training data alone.

The fall is real, and it is not a cure: Linardon et al. (2025) also found that 64 of GPT-4o's 141 real references, 45.4%, carried an error, most often in the DOI. And none of these studies tested the models ChatGPT runs now. OpenAI retired GPT-4o and GPT-4.1 from ChatGPT on 13 February 2026 (OpenAI, 2026d), and we found no peer-reviewed count of fabricated references for the GPT-5.6 models.

### With web search on

Search changes the failure, not the need to check. The Tow Center for Digital Journalism gave eight AI search tools excerpts from 200 news articles and asked each to name the article, publisher, date and URL. ChatGPT Search got 134 of the 200 wrong, signalled doubt in only 15 answers, and never declined to answer (Jaźwińska & Chandrasekar, 2025).

Rao, Wong and Callison-Burch (2026), in a preprint not yet peer reviewed, checked more than 53,000 URLs cited by ten search-backed models and agents. For OpenAI's gpt-4o-search-preview, 8.8% of the cited URLs had probably never existed. For OpenAI's deep research agent the invented share was lower, 3.5%, but 10.1% of its URLs did not resolve at all.

## How does ChatGPT get a reference wrong?

ChatGPT gets a reference wrong in five recognisable ways, each measured in a study or recorded by a court.

**Real authors, a title nobody wrote.** Gravel et al. (2023) asked ChatGPT 20 medical questions in February 2023 and checked its 59 references: 41 (69%) were fabricated, yet 56 of the 59 (95%) named authors who had really published on a related subject, and 29 of the 41 fakes (71%) sat in a known journal with a volume and page range that belonged to a different article. Only the paper counts.

**A borrowed identifier.** Of the 33 fabricated references with a DOI in the Linardon et al. (2025) test, 21 had a DOI that worked and led to an unrelated article. Chelli et al. (2024) found the DOI correct on only 16% of real GPT-3.5 references and 20% of real GPT-4 ones, and Bhattacharyya et al. (2023) found the PubMed ID wrong in 93% of the papers ChatGPT-3.5 wrote. In *Johnson v. Dunn*, one citation ChatGPT produced carried a Westlaw number that opened a maritime injury case which "does not discuss discovery".

**A real source cited for something it does not say.** In the same case, another citation named a case that exists, but "no case with that combination of style and proposition exists". The Tow Center found the search tools shared "a common tendency to cite the wrong article", often a syndicated copy rather than the original.

**Search on, sentence still wrong.** A search-backed answer can link to a page that exists and still misstate it, which is why OpenAI's own page tells you to open the source and confirm it supports the answer.

**"Are these real?" answered with confidence.** When Gravel et al. (2023) challenged ChatGPT about a reference, it replied that the reference was "available in Pubmed" and supplied a link to an unrelated record. The lawyers in *Mata v. Avianca* got the same reassurance. It answers that question the way it wrote the reference, by producing likely text.

Two of the five are caught by looking the reference up, and two only by reading the source:

| ChatGPT's failure | What catches it | Where EdCitation fits |
| --- | --- | --- |
| Real authors, a title nobody wrote | The exact title has no record | [Verify references](https://edcitation.com/verify-references) looks up every entry |
| A borrowed DOI or PubMed ID | The identifier opens a different paper | The same lookup, entry by entry |
| A real page cited for a claim it does not make, search on or off | Reading the passage | No lookup can; open the source |
| "Are these real?" answered yes | Any check made outside ChatGPT | A checker that never writes a reference |

### What a court did about ChatGPT citations in 2025

The stricter ruling came two years after *Mata v. Avianca*, which the leader article covers. In *Johnson v. Dunn*, three lawyers at the firm Butler Snow filed two motions in a federal court in Alabama containing five citations that a partner had obtained from ChatGPT "without verifying their accuracy"; he had checked none of them in Westlaw or PACER. On 23 July 2025 the court wrote that "Fabricating legal authority is serious misconduct that demands a serious sanction", publicly reprimanded all three, ordered the ruling published, disqualified them from the case and referred them to the state bar.

## How do I ask ChatGPT for references I can check?

Ask for identifiers you can follow, restrict ChatGPT to what it retrieved, and give it permission to say "not found". None of this makes the list safe; it makes the list checkable. Turn web search on first.

1. **Ask for the DOI and URL of every source.** "Give me five peer-reviewed sources on [topic]. For each, give the full reference, the DOI, and the URL of the page you retrieved it from."
2. **Limit it to what it retrieved.** "Cite only sources you retrieved in this conversation with web search. Do not cite anything from memory." The Linardon et al. (2025) prompt asked for peer-reviewed sources and still received 19.9% fakes, so this reduces the problem rather than removing it.
3. **Ask for the passage.** "For each source, quote the sentence or two that supports the claim, and say where in the source it appears." You then have something to look for when you open the page.
4. **Permit "not found".** "If you cannot retrieve a source for a claim, write 'not found' rather than guessing." In the Tow Center test ChatGPT Search never declined to answer, so expect this to work only sometimes.
5. **Check every reference anyway**, including those with a working link. Pasting the list into [Verify references](https://edcitation.com/verify-references) does the lookups; reading the passage is still yours.

No prompt changes how the sentence is made: even with search on, a language model writes the citation. For the finding itself there is a way round ChatGPT. Give EdCitation's [Find sources](https://edcitation.com/) the claim your sentence makes, or the topic, and it searches about 300 million published works for it, so what comes back is a record to read, not a reference to catch out.

## How do I turn on web search in ChatGPT?

Select View all tools in the message box and then Search, or type / and choose Search; ChatGPT may also search on its own when it decides a question needs current information. The steps are OpenAI's, checked on 22 September 2026 (OpenAI, 2026a).

- **To search a question you have already asked**, use the refresh control under the answer and choose "Search the web", where it is offered.
- **To read the sources**, select a citation, or the Sources button.
- **To use deep research**, type /Deepresearch, or open the tools menu (+) and select Deep research. Under Sites and Manage sites you can restrict the research to sites you name, or prioritise them and still search the wider web. Usage varies by plan, and a counter in the product shows what you have left (OpenAI, 2026b).

Search turns the citations into links you can open. It does not stop the model getting the title, authors or year wrong beside a link that works.

## How do I find and fix fake citations in a ChatGPT reference list?

Check each reference against the publisher's record, then replace the ones that fail with sources you have read. Resolve the DOI and see whether the page it opens carries the same title and authors; failing a DOI, put the exact title in quotation marks into Crossref, PubMed or Google Scholar; then read the part you are citing. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) goes through each step, including what "not found" does and does not prove.

For a whole list, EdCitation's [Verify references](https://edcitation.com/verify-references) is the best tool for the job, for a reason ChatGPT cannot match: it looks each entry up in the publisher's record and never writes one. When Gravel et al. (2023) challenged ChatGPT about a fabricated reference, it said the paper was in PubMed; a lookup finds the record or reports that it did not. Paste the list or upload the paper and every entry is marked verified, doubtful (shown as "check this") or not found, a withdrawn paper carries a retraction flag, and an index that did not reply leaves the entry at "could not check", never at "not found". That check is free, with no account. The paid plans cover the rest of the paper: [Pro](https://edcitation.com/pricing), $8 a month, adds [References from a file](https://edcitation.com/tools/references-from-a-file) and Mechanics QA, and Max, $24 a month, adds Theoretics QA and the Library.

To fix a reference that fails, do not just delete it; the claim it was holding up now has nothing under it. Search for the claim itself with [Find sources](https://edcitation.com/), read what you find, and let [Cite a source](https://edcitation.com/cite) set the new reference from its DOI in APA 7, MLA 9 or whichever of its six styles your course uses, screened for retraction. If your course allows ChatGPT, say so and [cite it as the course asks](https://edcitation.com/newsletter/how-to-cite-chatgpt-and-ai-tools).

## Quick questions

### Does ChatGPT still make up citations in 2026?

Yes, though less often than in 2023. Linardon et al. (2025), the latest peer-reviewed count, found 35 of GPT-4o's 176 references fabricated in June 2025, and a 2026 preprint found that 8.8% of the URLs cited by OpenAI's gpt-4o-search-preview model had probably never existed.

### Does turning on web search in ChatGPT stop fake citations?

No. Search lets ChatGPT link to pages that exist, and OpenAI's help page still warns that citations can be "incomplete, outdated, or incorrect". In the Tow Center test, ChatGPT Search answered 134 of 200 source-identification questions wrongly.

### Can I ask ChatGPT to check whether its citations are real?

No. ChatGPT answers that question the way it wrote the reference, by producing likely text: in one published test it confirmed that a fabricated reference was in PubMed and linked to an unrelated record, and in *Mata v. Avianca* the lawyers who relied on that reassurance were sanctioned. Ask something that cannot reassure you, such as EdCitation's Verify references, which looks each reference up and reports what the record holds.

### Is ChatGPT deep research safe for a reference list?

Safer, not safe. OpenAI's launch post said deep research can still "hallucinate facts", and a 2026 preprint found that 3.5% of the URLs cited by OpenAI's deep research agent had probably never existed and 10.1% did not resolve. Check every reference in the report as you would any other.

## References

- Bhattacharyya, M., Miller, V. M., Bhattacharyya, D., & Miller, L. E. (2023). High rates of fabricated and inaccurate references in ChatGPT-generated medical content. *Cureus, 15*(5), Article e39238. [https://doi.org/10.7759/cureus.39238](https://doi.org/10.7759/cureus.39238)
- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. *Journal of Medical Internet Research, 26*, Article e53164. [https://doi.org/10.2196/53164](https://doi.org/10.2196/53164)
- Gravel, J., D'Amours-Gravel, M., & Osmanlliu, E. (2023). Learning to fake it: Limited responses and fabricated references provided by ChatGPT for medical questions. *Mayo Clinic Proceedings: Digital Health, 1*(3), 226-234. [https://doi.org/10.1016/j.mcpdig.2023.05.004](https://doi.org/10.1016/j.mcpdig.2023.05.004)
- Jaźwińska, K., & Chandrasekar, A. (2025). AI search has a citation problem. *Columbia Journalism Review*. [https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- *Johnson v. Dunn*, No. 2:21-cv-1701-AMM (N.D. Ala. July 23, 2025). Sanctions order. [https://www.courthousenews.com/wp-content/uploads/2025/07/johnson-vs-dunn-attorney-sanctions-order.pdf](https://www.courthousenews.com/wp-content/uploads/2025/07/johnson-vs-dunn-attorney-sanctions-order.pdf)
- Linardon, J., Jarman, H. K., McClure, Z., Anderson, C., Liu, C., & Messer, M. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models. *JMIR Mental Health, 12*, Article e80371. [https://doi.org/10.2196/80371](https://doi.org/10.2196/80371)
- *Mata v. Avianca, Inc.*, No. 22-cv-1461 (PKC) (S.D.N.Y. June 22, 2023). Opinion and order on sanctions. [https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf](https://storage.courtlistener.com/recap/gov.uscourts.nysd.575368/gov.uscourts.nysd.575368.54.0_3.pdf)
- OpenAI. (2024). *Introducing ChatGPT search*. [https://openai.com/index/introducing-chatgpt-search/](https://openai.com/index/introducing-chatgpt-search/)
- OpenAI. (2025a). *Why language models hallucinate*. [https://openai.com/index/why-language-models-hallucinate/](https://openai.com/index/why-language-models-hallucinate/)
- OpenAI. (2025b). *Introducing deep research*. [https://openai.com/index/introducing-deep-research/](https://openai.com/index/introducing-deep-research/)
- OpenAI. (2026a). *Searching the web with ChatGPT*. OpenAI Help Center. [https://help.openai.com/en/articles/9237897-chatgpt-search](https://help.openai.com/en/articles/9237897-chatgpt-search)
- OpenAI. (2026b). *Deep research in ChatGPT*. OpenAI Help Center. [https://help.openai.com/en/articles/10500283-deep-research-faq](https://help.openai.com/en/articles/10500283-deep-research-faq)
- OpenAI. (2026c). *GPT-5.6 Sol*. OpenAI API documentation. [https://developers.openai.com/api/docs/models/gpt-5.6-sol](https://developers.openai.com/api/docs/models/gpt-5.6-sol)
- OpenAI. (2026d). *Model release notes*. OpenAI Help Center. [https://help.openai.com/en/articles/9624314-model-release-notes](https://help.openai.com/en/articles/9624314-model-release-notes)
- Rao, D., Wong, E., & Callison-Burch, C. (2026). *Detecting and correcting reference hallucinations in commercial LLMs and deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2604.03173](https://arxiv.org/abs/2604.03173)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
