# Claude fake citations and how to find and fix them

**Does Claude make up citations, and how do I check them?** Yes. Claude can make up citations when it answers from memory, and with web search on it can attach a real link to a claim the page does not make. In a 2025 test of free chatbots, 64% of the 50 references Claude produced were fabricated or wrong. Check each against the publisher's record: EdCitation's free Verify references does that for a whole list.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/claude-fake-citations-how-to-find-and-fix

Ask Claude for sources and it answers in the shape you asked for: authors, year, title, journal, sometimes a DOI. Whether those papers exist depends on how you asked and on what Claude was allowed to look up.

Whether Claude makes up citations has been measured. One comparison tested the free version of Claude before it had web search and found 64% of its references fabricated or wrong. Two 2026 preprints tested Claude with search on and found a different failure: the link opens, but the page does not say what Claude said it does.

## Does Claude make up citations?

Yes. When Claude answers from what it learned in training, it writes the most likely next words, and a reference is a very regular kind of text. Nothing in that step looks a paper up; that is a separate job, and the one EdCitation's free [Verify references](https://edcitation.com/verify-references) does for a whole list. [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references) explains the mechanism and tabulates the ChatGPT studies; Walters and Wilder (2023) checked 636 citations and found 55% of those from GPT-3.5 and 18% of those from GPT-4 fabricated.

### What Anthropic says about accuracy

Anthropic's help centre says it plainly (checked on 22 September 2026). Its page on incorrect responses says Claude "can occasionally produce responses that are incorrect or misleading" and warns that Claude "can display quotes that may look authoritative or sound convincing, but are not grounded in fact". Users "should not rely on Claude as a singular source of truth", and after a web search they should examine the cited original sources.

Anthropic's developer guide on reducing hallucinations recommends letting Claude say it does not know, asking for word-for-word quotes, having it cite a source for each claim and retract any it cannot support, and restricting it to the documents you gave it. It then says these techniques "don't eliminate them entirely".

### The knowledge cutoff, and why it matters

Every Claude model has a date after which it knows nothing unless it searches. Anthropic's models overview (checked on 22 September 2026) gives each model a "reliable knowledge cutoff": June 2026 for Claude Fable 5.1, May 2026 for Claude Opus 5, January 2026 for Claude Sonnet 5 and February 2025 for Claude Haiku 4.5. The help centre adds that the models "may not be aware of events or information that occurred after their respective cutoff dates". A paper published after the cutoff cannot be known without a search, and earlier papers are known, if at all, as patterns in text rather than as records.

## What do the studies say about Claude's references?

The evidence on Claude is thinner than on ChatGPT, and it splits by whether search was on. None of the medical studies in the leader article tested Claude.

### Claude without search: the eight-chatbot test

Cabezas-Clavijo and Sidorenko-Bautista (2025) asked eight chatbots, in their free versions, for academic references for a final-year project in five disciplines, on 7 to 9 February 2025. The Claude tested was Claude 3.5 Sonnet, and Anthropic's own announcement dates web search in Claude to March 2025 for paid plans and May 2025 for all plans, so it was answering from memory. Each chatbot produced 50 references, checked by searching the title in quotation marks in Google and Google Scholar and scored on five elements.

Across all 400 references, 26.5% were fully correct, 33.8% real but with errors, and 39.8% wrong or fabricated. Claude's figure was 64%, the third highest of the eight, with 3.0 missing or wrong elements per reference on average. Its references were recent, dated 2019 to 2023 with an average age of 3.2 years. And 66% were journal articles, the format most often invented: across the sample, 78% of journal references were wrong or fabricated against 12.9% of books. The study has since appeared in the *Journal of Data and Information Science*. We found no peer-reviewed repeat on a current Claude model.

### Claude with web search: two 2026 preprints

Rao, Wong and Callison-Burch (2026), in a preprint, checked every URL in the pre-collected outputs of ten search-backed systems. A URL was hallucinated if it did not resolve and had never been archived by the Wayback Machine, and stale if it did not resolve but had been. Claude 3.5 Sonnet with search: 641 URLs, 7.8% non-resolving, 3.0% hallucinated. Claude 3.7 Sonnet with search: 1,735 URLs, 8.5% non-resolving, 3.2% hallucinated. Those were the lowest hallucinated rates of the ten, on smaller samples than most.

The same preprint ran Claude Sonnet 4.5 through Anthropic's web search API on 2,177 expert questions in 32 fields: 61,407 URLs, 28.3 per question, 9.38% non-resolving, from 4.0% in mathematics to 17.4% in healthcare and medicine. A "claude-research" entry was excluded for producing a single URL, so the preprint says nothing about Claude's Research mode.

Onweller et al. (2026), a preprint from a PricewaterhouseCoopers team, ran 14 models as deep research agents with web search on 130 queries and scored whether each cited link worked, whether the page was relevant, and whether it supported the claim attributed to it, the last judged by a model checked against manual reviews. For the five Claude models, links worked 97.2% to 99.2% of the time and pages were relevant 83.9% to 95.7% of the time, but the claim was supported only 51.8% (Claude Sonnet 4.5) to 76.8% (Claude Opus 4.5) of the time. Letting Claude Opus 4.6 search more left the links good and cut the fact-check score from 80.0% at two tool calls to 57.9% at 150.

### What did not test Claude

The Tow Center's March 2025 test of AI search engines (Jaźwińska and Chandrasekar, 2025), quoted for finding more than 60% of answers wrong, covered eight tools from OpenAI, Perplexity, DeepSeek, Microsoft, xAI and Google, and not Claude. Its figures do not apply to Claude, and neither do the ChatGPT figures in the leader article.

## How does Claude go wrong with references?

In those tests Claude's failures took four forms, each caught by a different check.

**A real journal, a recent year, a paper that does not exist.** The no-search failure: Claude leaned towards journal articles and recent dates, the references most often fabricated. Follow the DOI or search the exact title.

**A real paper with wrong details.** Claude's references averaged three wrong or missing elements out of five. Compare every element with the publisher's page.

**A live link that does not say it.** The with-search failure: in Onweller et al. (2026) almost every Claude link opened, and between a quarter and a half of the claims hung on them were unsupported. Open the page and find the sentence.

**A dead link.** In Rao, Wong and Callison-Burch (2026) most of Claude's non-resolving links were stale pages rather than invented ones, and the rate was highest in medicine. Search the title before condemning the reference.

The fifth failure is the user's: asking Claude whether its own list is real. No study we read tested that on Claude, and nothing in Anthropic's pages suggests the answer is worth more than any other generated sentence. When Rao, Wong and Callison-Burch got Claude to correct its own links, a tool fetched each URL, and non-resolving links fell from 4.9% to 0.8%. The tool did the checking.

## How do I turn on web search in Claude?

Open the + menu at the bottom of the chat window and choose Web search; a tick shows it is on (Anthropic's help centre, checked on 22 September 2026). Anthropic's announcement says web search is on all Claude plans.

Once on, Claude decides when to search: the help page says it "invokes a search tool" when a topic benefits from current information, and that writing "Search the web" or "Use web search" in the prompt forces one. Answers then carry "direct citations to sources", source links and "relevant quotes when appropriate", and the page's advice is to "cross-reference cited sources". With search on, Claude can also fetch a page whose URL you give it, a good way to make it read the DOI page of a reference it has just offered.

### What Research adds, and what it does not

Research, Anthropic's deep research mode, is on the Pro, Max, Team and Enterprise plans, is chosen from the same + menu, and needs web search on. The help centre describes it as "conducting multiple searches that build on each other" and returning "easy-to-check citations". Easier to check is not more likely to be right: in Onweller et al. (2026), more searching lowered the share of claims the cited page supported.

### The API's Citations feature

Developers can turn on Citations in Anthropic's API. Its documentation (checked on 22 September 2026) says it grounds Claude's answer in documents supplied with the request, as PDF, plain text or custom content, chunks them into sentences, and returns the cited text with a location: character positions for text, page numbers for PDFs. The API extracts the cited text itself, so citations are "guaranteed to contain valid pointers to the provided documents". Citations does not look anything up. It points into a file you supplied, and does not confirm that a paper exists or that a reference to it is correctly formed.

### Projects and uploaded files

Projects are on every plan, with up to five on Free, and anything uploaded to a project is used as context across its chats. When Claude quotes from a paper you uploaded, the paper is real because you supplied it, and the quotation can be checked against the file. A reference list Claude then writes for those papers is still generated text, and the year, volume or pages can be wrong. Build those references from each paper's DOI with EdCitation's [Cite a source](https://edcitation.com/cite) instead: it takes them from the record, in APA 7, MLA 9, Chicago, Harvard, IEEE or Vancouver.

## How do I ask Claude for references I can check?

Ask for the DOI and URL of every source, keep Claude to what it retrieved, and let it answer "not found". That is Anthropic's developer guidance, turned into a chat.

1. **Turn on web search**, and Research if the job needs more than a few searches.
2. **Ask for identifiers and provenance.** "Find peer-reviewed sources on [topic]. For each, give the authors, year, title, journal, DOI and the URL you retrieved it from. Cite only sources you found with web search in this conversation."
3. **Allow "not found".** "If no source you found supports a claim, write 'not found' rather than suggesting one from memory."
4. **Ask for the passage.** "For each source, quote the sentence that supports the claim."
5. **Restrict Claude to your own documents when you have them.** "Use only the files I uploaded. Do not add sources from memory."
6. **Check every reference outside Claude** before it goes into your list: paste Claude's list into EdCitation's [Verify references](https://edcitation.com/verify-references), then open the sources your argument leans on.

| What you ask for | What it catches | What it misses |
| --- | --- | --- |
| A DOI and URL for each source | A reference with no record fails at the first click | A real DOI attached to the wrong title |
| "Cite only what you retrieved" | References written from memory | A live page cited for a claim it does not make |
| A quoted sentence | A source that does not say it | A quotation altered on the way |
| Then, the whole list through [Verify references](https://edcitation.com/verify-references) | Papers with no record, doubtful details, retractions | Whether the page supports Claude's sentence |

None of these removes the problem, which is why the last row is a lookup and not another prompt.

## How do I check the references Claude gave me?

Against the publisher's record, not against Claude. Follow the DOI and check that the page names the same title and authors; where there is no DOI, look the exact title up, in quotation marks, in Crossref, PubMed or Google Scholar; then open the source and find the sentence you are citing it for. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) takes each step in turn, and explains what a blank result does and does not prove.

EdCitation's [Verify references](https://edcitation.com/verify-references) is the best tool for checking a whole list, and it is the kind of tool Claude needed in the one experiment where its links got better: once Rao, Wong and Callison-Burch (2026) gave Claude a tool that fetched each URL, its non-resolving links fell from 4.9% to 0.8%. Verify references never writes a reference; it looks each one up in the publisher's record, which no chatbot can say of itself. Upload the paper or paste the list; each entry returns as verified, doubtful ("check this") or not found, withdrawn papers are marked as retracted, and an entry no index answered for reads "could not check" instead of "not found". With the whole paper uploaded, it also pairs each in-text citation with its entry in the list, and each entry with the place it is cited. All of that is free, without an account; [Pro](https://edcitation.com/pricing) at $8 a month and Max at $24 a month add the paid tools, from [References from a file](https://edcitation.com/tools/references-from-a-file) to Theoretics QA and the Library.

Where a reference fails, the claim above it has lost its support. Search for that claim in [Find sources](https://edcitation.com/), open what comes back, and have Cite a source set the replacement from its DOI.

## Quick questions

### Does Claude make up citations when web search is on?

Less often, and differently. In a 2026 preprint, 3.0% to 3.2% of the URLs cited by Claude 3.5 and 3.7 Sonnet with search had probably never existed. The larger risk with search on is a real page cited for a claim it does not make.

### Can I ask Claude to check its own references?

Not usefully. Claude answers that question the way it wrote the reference, by generating likely text, and Anthropic's help centre says not to treat Claude as a single source of truth. The one experiment in which Claude corrected its own links gave it a tool that fetched each URL.

### Does Claude's Citations feature verify references?

No. It points to passages in documents you supplied, and Anthropic's documentation guarantees only that those pointers are valid within those documents. It does not check that a paper exists; a lookup in the publisher's record does, which is the job of EdCitation's Verify references.

### Is Claude better or worse than ChatGPT at references?

It depends on the test. In the February 2025 test of free chatbots without search, 64% of Claude's references were fabricated or wrong against 38% of ChatGPT's. In the 2026 test of search-backed systems, the two Claude models had the lowest rates of invented URLs.

### Why does Claude's knowledge cutoff matter for references?

A model cannot know a paper published after its cutoff, and knows earlier papers only as patterns in text. Anthropic gives each model a reliable knowledge cutoff, from February 2025 for Claude Haiku 4.5 to June 2026 for Claude Fable 5.1. Recent literature needs web search on.

## References

- Anthropic. (n.d.-a). *Citations*. Claude Platform Docs. [https://platform.claude.com/docs/en/build-with-claude/citations](https://platform.claude.com/docs/en/build-with-claude/citations)
- Anthropic. (n.d.-b). *Claude is providing incorrect or misleading responses. What's going on?* Claude Help Center. [https://support.claude.com/en/articles/8525154-claude-is-providing-incorrect-or-misleading-responses-what-s-going-on](https://support.claude.com/en/articles/8525154-claude-is-providing-incorrect-or-misleading-responses-what-s-going-on)
- Anthropic. (n.d.-c). *Enable and use web search*. Claude Help Center. [https://support.claude.com/en/articles/10684626-enable-and-use-web-search](https://support.claude.com/en/articles/10684626-enable-and-use-web-search)
- Anthropic. (n.d.-d). *How can I create and manage projects?* Claude Help Center. [https://support.claude.com/en/articles/9519177-how-can-i-create-and-manage-projects](https://support.claude.com/en/articles/9519177-how-can-i-create-and-manage-projects)
- Anthropic. (n.d.-e). *How up-to-date is Claude's training data?* Claude Help Center. [https://support.claude.com/en/articles/8114494-how-up-to-date-is-claude-s-training-data](https://support.claude.com/en/articles/8114494-how-up-to-date-is-claude-s-training-data)
- Anthropic. (n.d.-f). *Models overview*. Claude Platform Docs. [https://platform.claude.com/docs/en/models/overview](https://platform.claude.com/docs/en/models/overview)
- Anthropic. (n.d.-g). *Reduce hallucinations*. Claude Platform Docs. [https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations](https://platform.claude.com/docs/en/test-and-evaluate/strengthen-guardrails/reduce-hallucinations)
- Anthropic. (n.d.-h). *Use research on Claude*. Claude Help Center. [https://support.claude.com/en/articles/11088861-use-research-on-claude](https://support.claude.com/en/articles/11088861-use-research-on-claude)
- Anthropic. (2025, March 20). *Claude can now search the web*. Claude Blog. [https://claude.com/blog/web-search](https://claude.com/blog/web-search)
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2025). *Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate* [Preprint]. arXiv. [https://arxiv.org/abs/2505.18059](https://arxiv.org/abs/2505.18059)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. *Columbia Journalism Review*. [https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- Onweller, H., Lumer, E., Huber, A., Ramchandani, P., Subbiah, V. K., & Feld, C. (2026). *Cited but not verified: Parsing and evaluating source attribution in LLM deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2605.06635](https://arxiv.org/abs/2605.06635)
- Rao, D., Wong, E., & Callison-Burch, C. (2026). *Detecting and correcting reference hallucinations in commercial LLMs and deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2604.03173](https://arxiv.org/abs/2604.03173)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
