# Gemini fake citations and how to find and fix them

**Does Gemini make up sources, and how do I check them?** Yes. Google Gemini makes up sources, and Google's own help page says Gemini can hallucinate and misrepresent how it cites. In 2023 Bard fabricated 91% of its references, and in a 2026 preprint 13% of the URLs cited by Gemini Deep Research had probably never existed. EdCitation's free Verify references checks every Gemini reference against the publisher's record.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/gemini-fake-citations-how-to-find-and-fix

Ask Google Gemini for sources and it answers like a search engine: a list of papers, often with links, sometimes with a Sources button under the reply. That is what makes Gemini fake citations harder to spot than ChatGPT's. A link looks like proof. It is not.

Google's own help page says Gemini "can hallucinate and present inaccurate information as factual", and gives as an example that it can misrepresent how it carries out functions "like citing sources". The studies say the same in numbers.

## Does Gemini make up sources?

Yes. Gemini makes up sources, and in a way that is easy to miss, because the answer often carries a real link next to a wrong reference. The rate depends on the version, on whether search grounding was on, and on how narrow the question was.

The mechanism is the one behind every language model, and it has been measured: Walters and Wilder (2023) found that 55% of GPT-3.5's references and 18% of GPT-4's did not exist. [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references) explains it. The test of a Gemini reference is therefore not whether its link opens but whether the publisher's record holds that paper, and that is the test EdCitation's free [Verify references](https://edcitation.com/verify-references) runs on every entry in a list.

## What have the studies found about Gemini and Bard?

Every published test of Gemini or Bard that we read found fabricated or wrong references, from nearly all of them in Bard in 2023 to a minority in grounded Gemini in 2026. The methods differ so much that the figures are not one trend line.

| Study | What was tested | Result for Gemini or Bard |
| --- | --- | --- |
| McGowan et al. (2023), *Psychiatry Research* | Bard 2.0, June 2023, two psychiatry searches | 15 linked citations, none accurate; 5 of 10 links on the broad topic went to unrelated papers |
| Chelli et al. (2024), *Journal of Medical Internet Research* | Bard (PaLM 2), July 2023, 11 systematic reviews | 95 of 104 references hallucinated (91.4%) |
| Omar et al. (2025), *Computers in Biology and Medicine* | Gemini Ultra against GPT-4, referenced medical introductions | Gemini 77.2% correctness and 68.0% accuracy, GPT-4 54.0% and 49.2%; both "produced fabricated evidence" |
| Jaźwińska and Chandrasekar (2025), Tow Center | Gemini among eight AI search tools, 200 news articles | One completely correct answer; over half of responses cited fabricated or broken URLs |
| Venkit et al. (2025), preprint | Gemini Deep Research, August 2025 | 50.3% citation accuracy; 53.6% of statements unsupported by its own listed sources |
| Civelekler and Çıtırık (2026), *Czech and Slovak Ophthalmology* | Gemini ("Gemini Ultra 2.5" in the paper), 35 clinical paragraphs | 5 of 35 references accurate (12.9%), the lowest of four tools |
| Özbek and Bağcıer (2026), *Indian Journal of Orthopaedics* | ChatGPT, Gemini and Perplexity, 3,150 references | Hallucination score 4.01, between ChatGPT (1.81) and Perplexity (6.51); higher is worse |
| Rao et al. (2026), preprint | Gemini 2.5 with search, and 2.5 Pro Deep Research, among ten systems | 4.8% and 4.6% of cited URLs hallucinated with search on; 13.3% for Deep Research, the highest of ten |

### Bard in 2023: the worst figures in the literature

Chelli et al. (2024) counted a reference as hallucinated if two of its title, first author and year were wrong. Bard produced 104 references and 95 were hallucinated, 91.4%, against 39.6% for GPT-3.5 and 28.6% for GPT-4, and the authors describe a "try-and-repeat approach": several versions of one invented paper with slightly different titles and journals. McGowan et al. (2023) found the pattern that still matters most: on the specific topic, all five of Bard's links opened real papers with the same title, and all five citations named the wrong authors.

### Gemini with search on: better, not fixed

Rao et al. (2026), in a preprint not yet peer reviewed, checked whether each cited URL resolved and had ever existed. With search on, Gemini 2.5 Pro and Flash hallucinated 4.8% and 4.6% of their cited URLs; Gemini 2.5 Pro Deep Research cited 113.1 URLs per query, more than any other system, and 13.3% of them were hallucinated, the highest rate of the ten. A longer Gemini bibliography is not a safer one.

The evidence on the current app models is thin: we found no peer-reviewed test that names Gemini 3 or later, and the two 2026 journal studies do not state the model version precisely. Treat the figures as a range, not a verdict on the version in front of you.

## What does Google say about Gemini's sources and accuracy?

Google says plainly that Gemini can hallucinate, and adds that it can misrepresent how it cites. Checked on 22 September 2026, the page "Learn about responses from Gemini Apps" gives as an example of hallucination that Gemini can misrepresent how it carries out functions "like citing sources, or providing fresh information". So asking Gemini whether it looked a paper up is not a check.

Other things Google's pages say, checked the same day:

- **Sources are optional.** A Sources button appears under a response, or in line, when sources are available, and "Not all responses include related links or sources". No button means Gemini gave no links for that answer.
- **The model has a cutoff.** The model card for Gemini 3.6 Flash, the free plan's model (Google DeepMind, 2026), gives a knowledge cutoff of March 2026 for some domains and January 2025 for others, and says the model "may exhibit some of the general limitations of foundation models, such as hallucinations".

## What does Gemini's double-check button actually check?

Gemini's double-check feature compares each statement in a response with Google Search results and colours it by whether similar text was found. It does not check whether a cited paper exists.

Google's help page as archived on 28 January 2026 described it like this. Under a response, open More and choose Double-check response. Green means "Google Search found content that's likely similar to the statement", and the link shown "is not necessarily what the Gemini app used to generate its response". Orange means Google Search found content that is likely different, or found nothing relevant. The page also said that even with sources shown Gemini "can still get things wrong".

### Where the feature stands on 22 September 2026

The current version of that page, checked on 22 September 2026, is titled "View related sources from Gemini Apps" and no longer describes double-check. In Google's community forum a user asked on 20 May 2026 whether it had been phased out; a volunteer product expert replied that it still sits under the More menu, but community answers are not Google's documentation, and we could not confirm the option in the app. If yours shows it, remember that the highlight is about the claim, not the citation: a sentence can be green because many pages agree with the claim while the paper cited for it does not exist.

Set beside a reference checker, the difference is plain:

| | Gemini's double-check | EdCitation's Verify references |
| --- | --- | --- |
| What it compares | Each statement with Google Search results | Each reference with the publisher's record |
| Whether the cited paper exists | Not checked | Verified, doubtful ("check this") or not found |
| Whether the source supports the sentence | Compared with web text, not the cited source | Not checked: open the source and read it |
| Where it runs | Inside Gemini, where the option still appears | [On the site](https://edcitation.com/verify-references), free, no account |

## How does Gemini typically go wrong with references?

Gemini's characteristic failures involve links and real papers rather than obviously invented ones.

- **A real title, the wrong authors.** In McGowan et al. (2023), all five of Bard's links on the specific topic opened a real paper with the same title, and all five citations named the wrong authors.
- **A link to an unrelated paper.** Five of Bard's ten citations on the broad topic linked to articles on something else, and Civelekler and Çıtırık (2026) name "DOI mismatches and the generation of irrelevant or unverifiable references" as the main error types in their test.
- **A URL that never existed.** In the Tow Center test, more than half of Gemini's responses cited fabricated or broken URLs (Jaźwińska and Chandrasekar, 2025). Rao et al. (2026) found URLs that had never existed in 4.6% to 13.3% of Gemini's citations depending on the mode.
- **A real source cited for a claim it does not make.** Venkit et al. (2025) found that 53.6% of the relevant statements in Gemini Deep Research reports were not supported by any of the report's own listed sources, and that only 50.3% of its citations accurately supported the statement they were attached to.
- **Grounding on, sentence still wrong.** Grounded or not, the sentence is written by the model, and Google's API documentation says the model itself "determines if a Google Search can improve the answer", so a given reference may never have been searched for.

## How do I ask Gemini for references that can be checked?

Ask for the identifiers a checker needs, and ask Gemini to say when it has not found something. No prompt removes the problem, but a good one makes every reference checkable in a minute.

### Prompts that make checking possible

1. **Ask for a DOI and a URL for every reference.** "For each source, give the DOI and the URL of the publisher's page, or write 'no DOI found' rather than guessing."
2. **Ask it to cite only what it retrieved.** "Cite only sources you have actually found in this session. Do not cite from memory."
3. **Ask for the words.** "Quote, for each source, the sentence that backs the claim."
4. **Ask for "not found" rather than a guess.** "If you cannot find a source for a claim, say 'no source found' and leave the claim unsupported."
5. **Take the list out of Gemini to check it.** Paste it into EdCitation's [Verify references](https://edcitation.com/verify-references), which looks each entry up in the record, then open the sources your argument rests on.

Never ask Gemini whether its own reference is real. Google's help page says the tool can misrepresent how it cites sources, and the answer is produced the same way the reference was.

### Grounding, Deep Research and the API: what each changes

- **In the Gemini app**, search grounding is not a switch you set. The Sources button tells you whether links were provided for that answer, and no button means none.
- **Deep Research** is the mode to use when you want links. Checked on 22 September 2026, Google's help page says Google Search is included as a source by default, and that Google AI Pro and Ultra plans get higher limits on the number of reports. On the UK plans page the same day, Deep Research is on the free plan, with Plus (£4.99 a month), Pro (£18.99) and Ultra (from £79.99) raising usage limits. It is also the mode with the highest hallucinated-URL rate in Rao et al. (2026).
- **In the Gemini API**, grounding is explicit. The documentation says the google_search tool "connects the Gemini model to real-time web content" and returns citations as annotations giving the URL, the title and the span of text each supports. That is a list of URLs to check, not a list that has been checked.

## How do I check Gemini's references?

Check each reference against the publisher's record, not against Gemini. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) has the full procedure; the short version is to resolve the DOI at doi.org and compare the title and authors, search Google Scholar, Crossref or PubMed for the exact title in quotation marks, open the paper and read the part you are citing, and treat a source the indexes do not cover as "could not be checked", never as "not found".

EdCitation's [Verify references](https://edcitation.com/verify-references) is the best tool for checking a Gemini list. It never writes a reference and it checks each one against the publisher's record, which is exactly what Gemini cannot do: Google's own help page says Gemini can misrepresent how it cites, and in McGowan et al. (2023) Bard's links opened real papers while every citation named the wrong authors, a mismatch that only comparing the reference with the record exposes. Paste the list or upload the paper. Every entry is reported as verified, doubtful ("check this") or not found, any retracted paper is flagged as such, and an index that failed to answer produces "could not check" rather than a false "not found". The check costs nothing and needs no account; EdCitation's Pro plan ($8 a month) and Max plan ($24 a month) are for the paid tools listed on the [pricing page](https://edcitation.com/pricing), not for this one.

When a Gemini reference fails, the claim has no support. Search for it in EdCitation's [Find sources](https://edcitation.com/), which covers about 300 million published works, read what you find, and let [Cite a source](https://edcitation.com/cite) build the reference from its DOI.

## Quick questions

### Does Gemini make up references less often than ChatGPT?

The direct comparisons point both ways: Chelli et al. (2024) found Bard far worse than GPT-4, Omar et al. (2025) found Gemini Ultra better than GPT-4, and Özbek and Bağcıer (2026) found ChatGPT more reliable than Gemini. All three found fabricated references from both.

### Does Gemini Deep Research invent citations?

Yes. In the Rao et al. (2026) preprint, 13.3% of the URLs cited by Gemini 2.5 Pro Deep Research had probably never existed, the highest rate of the ten systems tested. Google's overview page for the feature says "Check responses", and its developer documentation recommends "reviewing the citations provided in the response to verify the sources".

### Does a green double-check highlight mean the source is real?

No. Google's description of the feature said green means Google Search found content likely similar to the statement, and that the link offered is not necessarily what Gemini used. It checks the claim against the web, not whether the paper cited for it exists; that needs a lookup in the publisher's record, the job of EdCitation's Verify references.

### If Gemini gives a link, is the reference real?

Not necessarily. In McGowan et al. (2023), Bard's links opened real papers with the same title while the citations named the wrong authors. Open the link and compare the authors, the year and the title with the reference.

## References

- Chelli, M., Descamps, J., Lavoué, V., Trojani, C., Azar, M., Deckert, M., Raynier, J.-L., Clowez, G., Boileau, P., & Ruetsch-Chelli, C. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: Comparative analysis. *Journal of Medical Internet Research, 26*, Article e53164. [https://doi.org/10.2196/53164](https://doi.org/10.2196/53164)
- Civelekler, M., & Çıtırık, M. (2026). Evaluation of AI citation accuracy in anterior segment research. *Czech and Slovak Ophthalmology, 82*. Advance online publication. [https://doi.org/10.31348/2026/21](https://doi.org/10.31348/2026/21)
- Google. (n.d.-a). *Gemini Deep Research agent*. Gemini API. [https://ai.google.dev/gemini-api/docs/interactions/deep-research](https://ai.google.dev/gemini-api/docs/interactions/deep-research)
- Google. (n.d.-b). *Grounding with Google Search*. Gemini API. [https://ai.google.dev/gemini-api/docs/grounding](https://ai.google.dev/gemini-api/docs/grounding)
- Google. (n.d.-c). *Learn about responses from Gemini Apps*. Gemini Apps Help. [https://support.google.com/gemini/answer/16279220](https://support.google.com/gemini/answer/16279220)
- Google. (n.d.-d). *Use Deep Research in Gemini Apps*. Gemini Apps Help. [https://support.google.com/gemini/answer/15719111](https://support.google.com/gemini/answer/15719111)
- Google. (n.d.-e). *View related sources from Gemini Apps*. Gemini Apps Help. [https://support.google.com/gemini/answer/14143489](https://support.google.com/gemini/answer/14143489)
- Google. (2026, January 28). *View related sources & double-check responses from Gemini Apps* [Archived help page]. Archived January 28, 2026, at [https://web.archive.org/web/20260128165812/https://support.google.com/gemini/answer/14143489?hl=en](https://web.archive.org/web/20260128165812/https://support.google.com/gemini/answer/14143489?hl=en)
- Google DeepMind. (2026). *Gemini 3.6 Flash model card*. [https://deepmind.google/models/model-cards/gemini-3-6-flash/](https://deepmind.google/models/model-cards/gemini-3-6-flash/)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. *Columbia Journalism Review*. [https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- McGowan, A., Gui, Y., Dobbs, M., Shuster, S., Cotter, M., Selloni, A., Goodman, M., Srivastava, A., Cecchi, G. A., & Corcoran, C. M. (2023). ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. *Psychiatry Research, 326*, Article 115334. [https://doi.org/10.1016/j.psychres.2023.115334](https://doi.org/10.1016/j.psychres.2023.115334)
- Omar, M., Nassar, S., Hijazi, K., Glicksberg, B. S., Nadkarni, G. N., & Klang, E. (2025). Generating credible referenced medical research: A comparative study of OpenAI's GPT-4 and Google's Gemini. *Computers in Biology and Medicine, 185*, Article 109545. [https://doi.org/10.1016/j.compbiomed.2024.109545](https://doi.org/10.1016/j.compbiomed.2024.109545)
- Özbek, İ. C., & Bağcıer, F. (2026). Reference hallucination in AI-assisted academic writing: A comparative analysis of ChatGPT, Gemini, and Perplexity in rotator cuff literature. *Indian Journal of Orthopaedics, 60*, 1949-1956. [https://doi.org/10.1007/s43465-026-01807-0](https://doi.org/10.1007/s43465-026-01807-0)
- Rao, D., Wong, E., & Callison-Burch, C. (2026). *Detecting and correcting reference hallucinations in commercial LLMs and deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2604.03173](https://arxiv.org/abs/2604.03173)
- Venkit, P. N., Laban, P., Zhou, Y., Huang, K.-H., Mao, Y., & Wu, C.-S. (2025). *DeepTRACE: Auditing deep research AI systems for tracking reliability across citations and evidence* [Preprint]. arXiv. [https://arxiv.org/abs/2509.04499](https://arxiv.org/abs/2509.04499)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
