# Does DeepSeek make up references, and how to check them

**Does DeepSeek make up references?** Yes. DeepSeek makes up references when it answers from memory, less often than some tools and more often than others, depending on the version and the topic: published tests found anything from no fabricated references to not one reference fully correct. Search mode cites real pages but can misattribute them. EdCitation's free Verify references checks every reference against the publisher's record.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/deepseek-fake-citations-how-to-check

Ask DeepSeek for sources and it gives you something that looks like a reference list. Whether the papers in it exist depends on which DeepSeek you asked, whether Search was on, and how narrow your topic was. DeepSeek does make up references. In some tests it also made up fewer than most other chatbots tested.

That is what makes DeepSeek's citations awkward to describe. One 2025 study found none in 50 references; another found not one reference fully correct. The gap is mostly the version, the mode and the topic. The mechanism itself, a model writing likely text with no step at which it looks anything up, is explained in [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references). Whichever DeepSeek you used, the check afterwards is the same, and EdCitation's free [Verify references](https://edcitation.com/verify-references) runs it on the whole list against the publisher's record.

## What does DeepSeek say about sources and citations?

DeepSeek says its output may be wrong and should be checked. Its terms of use, last updated 27 March 2026, state that "The Outputs may contain errors or omissions and are for your reference only", and that inaccuracy in AI-generated content "cannot be entirely avoided" (DeepSeek, 2026a). Its disclosure of how the models work says "we cannot guarantee that the model will not produce hallucinations", and describes labels in the app reminding users that content is AI-generated and may be inaccurate (DeepSeek, n.d.).

None of the DeepSeek pages read for this article states a knowledge cutoff for a current model. The models and pricing page, the model cards for DeepSeek-V4-Pro and DeepSeek-V4.1-Flash and the release notes back to January 2025 give parameters, context length and benchmarks, but no cutoff date (checked on 22 September 2026). When DeepSeek answers from memory, you do not know how recent that memory is.

### What the Search toggle does

DeepSeek's home page shows its chat box with two buttons beneath it, DeepThink and Search (checked on 22 September 2026). Search is off unless you turn it on. DeepSeek publishes no user guide for it, so the clearest description of what it does is the prompt template released with DeepSeek-R1 for the web search feature (DeepSeek-AI, 2025). The template pastes the search results into the prompt, tells the model that not all of them are relevant and that it must "evaluate and filter" them, and asks it to cite in the body of the answer as [citation:X], where X is the number of a retrieved result.

So with Search on, a numbered citation is a retrieved page that can be opened. The sentence above it is still written by the model, so the page can be real and the attribution wrong. The terms of use add that output "may include information from third-party websites or other external sources" for whose validity DeepSeek takes no responsibility (DeepSeek, 2026a).

### DeepThink, Expert Mode, and whether reasoning changes anything

Reasoning adds no lookup step. DeepThink is the chat app's reasoning button, and the V4 preview note of 24 April 2026 invites users to try V4 at chat.deepseek.com "via Expert Mode / Instant Mode" (DeepSeek, 2026b). The *Nature* paper on DeepSeek-R1 lists among its limitations that the model "cannot make use of tools, such as search engines and calculators" (Guo et al., 2025). A reasoning trace on its own is the model talking to itself.

### Open weights: a DeepSeek answer in another app may have no search at all

DeepSeek publishes its weights. DeepSeek-R1, DeepSeek-V4-Pro and DeepSeek-V4.1-Flash are on Hugging Face under the MIT licence, which permits commercial use and derivative works (DeepSeek-AI, 2026), so many apps and services run a DeepSeek model without being DeepSeek. The Search toggle belongs to DeepSeek's own app. A DeepSeek model inside another product has whatever retrieval that product added, which may be none, so ask whether "DeepSeek" means the app, the API or a copy of the weights.

## What has been measured about DeepSeek's references?

Seven tests published between March 2025 and August 2026 checked DeepSeek's references by hand, with results from no fabricated references at all to no reference fully correct. The spread is the finding.

| Study | DeepSeek tested | Task | Result |
| --- | --- | --- | --- |
| Jaźwińska and Chandrasekar (2025), Tow Center | DeepSeek Search | Name the source of 200 news excerpts | Source misattributed 115 of 200 times |
| Spennemann (2025) | R1, February 2025 | 50 references on each of 4 archaeology topics | 85% correct, 7% confabulated |
| Cabezas-Clavijo and Sidorenko-Bautista (2026) | V3, free app, 7 to 9 February 2025 | 50 references across 5 disciplines | 48% fully correct, none fabricated |
| Uldin et al. (2025) | R1 | 10 questions on recent radiology research | 5 of 10 answers significantly inaccurate; references fictitious |
| Gumilar et al. (2025), preprint | R1, chat app | Recent references on 7 biomedical topics | 91.43% of bibliographic fields wrong; none fully correct |
| Seifi and Seyfi (2026) | V3, web interface, retrieval off, March 2026 | 10 references on each of 10 neurocritical care topics | 23% with an error, 8% fabricated |
| Ülkir and Paslı (2026) | V3.2 | 120 anatomy questions | 47.5% hallucinated; 41.0% fully supported |

### Reading the seven results

Topic explains most of the spread. The good results came from broad disciplinary lists, where a model can name the standard works: 84% of the references DeepSeek gave Cabezas-Clavijo and Sidorenko-Bautista (2026) were books, and 45% of its real references were also produced by ChatGPT. Spennemann (2025) found DeepSeek R1's 7% confabulation rate statistically indistinguishable from ChatGPT-4o's 10%, and Seifi and Seyfi (2026), checking every reference in PubMed, Crossref, Google Scholar and the DOI resolver, found DeepSeek-V3 the lowest of three models at 8% fabricated, against 27% for GPT-5.3 and 50% for Grok-4.

The bad results came from recent, narrow biomedical literature, where every model does worst. Gumilar et al. (2025), a preprint not yet peer reviewed, scored each DeepSeek-R1 reference on five bibliographic fields and found none fully correct. Ülkir and Paslı (2026) put DeepSeek V3.2 at 47.5% hallucinated against 23.2% for ChatGPT 5.2 and 45.8% for Gemini 3 Pro. Version matters too: the two worst results are both R1, the reasoning model.

The evidence on DeepSeek is thinner than on ChatGPT: seven tests, one a preprint, most with 50 to 100 DeepSeek references, and none comparing the same model with Search on and off. Rao, Wong and Callison-Burch (2026), the largest test of citation URLs from search-backed tools, included no DeepSeek model. The 55% and 18% fabrication rates in Walters and Wilder (2023) are ChatGPT's figures, not DeepSeek's.

## How does DeepSeek go wrong with references?

In the ways every language model does, plus a few the tests make visible.

- **A fake that looks complete.** In Seifi and Seyfi (2026) the other two models sometimes left identifiers off their fabricated references; DeepSeek-V3 attached an identifier to every one of its fabrications.
- **A DOI that opens a different paper.** A DOI is a string the model reproduces from anything it has read. In Gumilar et al. (2025), 97.14% of DeepSeek-R1's DOIs were wrong; in Spennemann (2025), the same month, DeepSeek gave no DOIs at all.
- **The standard list, not the literature.** Cabezas-Clavijo and Sidorenko-Bautista (2026) found heavy overlap between the real references of DeepSeek, ChatGPT, Gemini and Grok, and Spennemann (2025) found his models' genuine sources also on Wikipedia pages. A DeepSeek list maps the famous works, not last year's papers.
- **References only when pushed.** In Uldin et al. (2025), DeepSeek-R1 answered without references and produced fictitious ones when asked. A reference requested after the answer is written to fit the answer.
- **With Search on, the right page under the wrong sentence.** Jaźwińska and Chandrasekar (2025) gave eight search tools excerpts from 200 news articles and asked for headline, publisher, date and URL. Collectively the tools answered more than 60% of queries wrongly, and DeepSeek Search misattributed the source 115 times out of 200.
- **Mandarin-language sources your indexes do not cover.** DeepSeek-R1 is "optimized for Chinese and English" (Guo et al., 2025), and its lists can include Chinese journals and reports that are real and absent from Crossref, PubMed and Google Scholar. That is "could not check", not "not found": look on the journal's own site or in a catalogue covering Mandarin-language publishing.
- **Being asked to verify its own list.** DeepSeek answers with the process that wrote the list. With Search on it may find a page that mentions the reference.

## How do I ask DeepSeek for references I can check?

Ask for identifiers, retrieved sources and quotations, and let it say "not found". These prompts do not stop fabrication; they make it faster to catch.

1. **Turn Search on and hold it to the results.** "Cite only sources from this search. For each, give the title, authors, year, journal, DOI and the URL you retrieved. Do not add sources from memory."
2. **Ask for the passage.** "For each citation, quote the sentence from the page that supports the claim, with its citation number."
3. **Give it a way out.** "If no retrieved source says this, write 'not found' instead of suggesting one."
4. **Ask for DOIs alone.** "List each reference as its DOI only, one per line." A column of DOIs is quick to resolve at doi.org, and EdCitation's [Cite a source](https://edcitation.com/cite) builds a reference from each DOI out of the record itself, so a DOI that belongs to another paper comes back with that paper's title.
5. **Use memory for orientation only.** With Search off, "Name the five best-known books and papers on this topic" gets the standard works, to be checked one by one.

Asking for a DOI does not make a reference more likely to exist: DeepSeek-V3 supplied one for every fabricated reference in Seifi and Seyfi (2026). The DOI is for you to check, not to trust.

## How do I switch on Search in DeepSeek, and what changes?

Select Search under the message box before you send the question. The app needs an account and was not signed into for this article; the buttons are as DeepSeek's home page shows them (checked on 22 September 2026).

1. Sign in at chat.deepseek.com or open the DeepSeek app.
2. Turn Search on under the message box. DeepThink or Expert Mode can be added, but it is not a second search.
3. Ask in the form above: retrieved sources only, DOI and URL for each, "not found" allowed.
4. Open every citation and confirm that the page is the source, not a news story about it, and says what your sentence says. Where it is a story about a study, search for the study itself in EdCitation's [Find sources](https://edcitation.com/).

What changes is that a bad citation becomes a bad reading of a real page rather than an invention. What does not change is that the reference list is still written by the model, that a retrieved page may be an abstract, a preprint or a mention rather than the paper, and that Search is a web search, not a lookup in Crossref or PubMed.

## How do I check what DeepSeek gave me?

Check each reference against the publisher's record, never against DeepSeek. Open the DOI and hold the title and authors on that page against the reference; for an entry with no DOI, search Crossref, PubMed or Google Scholar for its exact title in quotation marks; then read the passage you are citing. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) sets out every step.

What DeepSeek gave you decides the check:

| What DeepSeek gave you | The check | In EdCitation |
| --- | --- | --- |
| A list written with Search off | Every entry against the record | [Verify references](https://edcitation.com/verify-references), the list in one pass |
| A numbered citation from Search | Is the page the paper, and does it say this? | [Find sources](https://edcitation.com/) for the paper behind a news story |
| A book, as 84% of one test's references were | The ISBN or a library catalogue | [Cite a source](https://edcitation.com/cite) builds the reference from the ISBN |
| A Mandarin-language journal or report | The journal's own site, or a catalogue that covers it | Treat an empty lookup as unchecked, not as fake |

EdCitation's Verify references is the best tool for a DeepSeek list because it does what DeepSeek's Search does not: it never writes a reference, and it looks every entry up in the publisher's record, Crossref and PubMed among them, rather than on the open web. That matters with DeepSeek in particular. DeepSeek-V3 attached an identifier to every reference it fabricated in Seifi and Seyfi (2026), so a list that looks complete can still be hollow. Paste the list, or upload the paper, and each entry is returned verified, doubtful (flagged "check this") or not found; retracted papers carry a flag; and a lookup that an index never answered is labelled "could not check", not "not found". The check is free and needs no account. The paid plans, Pro at $8 a month and Max at $24, add other tools rather than a better version of this one; the [pricing page](https://edcitation.com/pricing) sets them side by side.

When an entry fails, the sentence it supported still needs a source: search for the claim itself, read what you find, and cite that instead.

## Quick questions

### Does DeepSeek make up fewer references than ChatGPT?

Sometimes. In February 2025 tests DeepSeek fabricated none or 7% of its references, and in March 2026 DeepSeek-V3 fabricated 8% against 27% for GPT-5.3. In a 2026 anatomy test DeepSeek V3.2 hallucinated 47.5% against 23.2% for ChatGPT 5.2. No version of DeepSeek has been shown not to invent references.

### Does DeepThink or Expert Mode stop DeepSeek inventing references?

No. Reasoning mode adds a visible chain of thought from the same model and no lookup. The DeepSeek-R1 paper says the model cannot use search engines, and the two worst published results were both produced by R1.

### Does turning on Search in DeepSeek fix its citations?

It changes them. Citations become numbered links to retrieved pages, so each can be opened. In the Tow Center test DeepSeek Search still misattributed the source of 115 of 200 excerpts, and a real page can still fail to say what the sentence above it claims.

### Can I ask DeepSeek to check whether its own references are real?

No. It answers with the process that wrote the list. Resolve the DOI, search the exact title in Crossref, PubMed or Google Scholar, or paste the list into EdCitation's Verify references, which looks each one up and writes nothing.

## References

- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2026). Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are entirely accurate. *Journal of Data and Information Science, 11*(2), 102-116. [https://doi.org/10.1515/jdis-2025-0326](https://doi.org/10.1515/jdis-2025-0326)
- DeepSeek. (n.d.). *Model mechanism and training methods of DeepSeek*. [https://cdn.deepseek.com/policies/en-US/model-algorithm-disclosure.html](https://cdn.deepseek.com/policies/en-US/model-algorithm-disclosure.html)
- DeepSeek. (2026a). *DeepSeek terms of use*. [https://cdn.deepseek.com/policies/en-US/deepseek-terms-of-use.html](https://cdn.deepseek.com/policies/en-US/deepseek-terms-of-use.html)
- DeepSeek. (2026b). *DeepSeek-V4 preview release*. DeepSeek API Docs. [https://api-docs.deepseek.com/news/news260424](https://api-docs.deepseek.com/news/news260424)
- DeepSeek-AI. (2025). *DeepSeek-R1* [README and model card]. GitHub. [https://github.com/deepseek-ai/DeepSeek-R1](https://github.com/deepseek-ai/DeepSeek-R1)
- DeepSeek-AI. (2026). *DeepSeek-V4-Pro* [Model card]. Hugging Face. [https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro)
- Gumilar, K. E., Ibrahim, I. H., Omran, K. E., Rajanagara, A. S., Yu, Z., Hsu, Y., Lin, C., Hung, T., Kousar, I., Tahir, N., Nhung, N. T. T., Ling, Y., Dachlan, E. G., Yang, J., Liao, L., & Tan, M. (2025). *Accuracy and hallucination of DeepSeek and ChatGPT in scientific figure interpretation and reference retrieval* [Preprint]. Research Square. [https://doi.org/10.21203/rs.3.rs-6676676/v1](https://doi.org/10.21203/rs.3.rs-6676676/v1)
- Guo, D., Yang, D., Zhang, H., Song, J., Wang, P., Zhu, Q., Xu, R., Zhang, R., Ma, S., Bi, X., Zhang, X., Yu, X., Wu, Y., Wu, Z. F., Gou, Z., Shao, Z., Li, Z., Gao, Z., Liu, A., … Zhang, Z. (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. *Nature, 645*, 633-638. [https://doi.org/10.1038/s41586-025-09422-z](https://doi.org/10.1038/s41586-025-09422-z)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. *Columbia Journalism Review*. [https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- Rao, D., Wong, E., & Callison-Burch, C. (2026). *Detecting and correcting reference hallucinations in commercial LLMs and deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2604.03173](https://arxiv.org/abs/2604.03173)
- Seifi, A., & Seyfi, A. (2026). Hallucination rate of peer-reviewed citations generated by large language models in neurocritical care. *Critical Care Explorations, 8*(9), Article e1474. [https://doi.org/10.1097/CCE.0000000000001474](https://doi.org/10.1097/CCE.0000000000001474)
- Spennemann, D. H. R. (2025). The origins and veracity of references 'cited' by generative artificial intelligence applications: Implications for the quality of responses. *Publications, 13*(1), Article 12. [https://doi.org/10.3390/publications13010012](https://doi.org/10.3390/publications13010012)
- Uldin, H., Saran, S., Gandikota, G., Iyengar, K. P., Vaishya, R., Parmar, Y., Rasul, F., & Botchu, R. (2025). A comparison of performance of DeepSeek-R1 model-generated responses to musculoskeletal radiology queries against ChatGPT-4 and ChatGPT-4o: A feasibility study. *Clinical Imaging, 123*, Article 110506. [https://doi.org/10.1016/j.clinimag.2025.110506](https://doi.org/10.1016/j.clinimag.2025.110506)
- Ülkir, M., & Paslı, B. (2026). Reference hallucination, citation reliability, and readability of large language models in anatomy-related question answering. *Clinical Anatomy*. Advance online publication. [https://doi.org/10.1002/ca.70187](https://doi.org/10.1002/ca.70187)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
