# Do Meta AI, Mistral and other chatbots invent references?

**Do Meta AI, Mistral and other chatbots invent references?** Yes. Meta AI and Mistral's Le Chat can invent references, because both write answers with a language model and neither searches for every answer. In a February 2025 test, about a third of the 50 references Le Chat gave were wrong or fabricated, and Meta AI's reference lists scored a median 29.7% for completeness. EdCitation's free Verify references checks every one against the publisher's record.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/meta-ai-and-other-chatbots-fake-citations

Ask Meta AI in WhatsApp for five sources and it will give you five. Ask Mistral's Le Chat, renamed Vibe in May 2026, and it will do the same. Both can produce a fake reference that looks right and points at nothing, for the reason ChatGPT does: a language model writes the answer, and a model that has not looked anything up is guessing at the title. The mechanism is explained in [why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references).

The two differ in when they look things up, how they show what they found, and what their makers say about trusting the result. This article covers both, then the chatbots built on the same open-weights models, with a one-minute test for whether the one in front of you searches at all. The evidence here is thinner than on ChatGPT; where none exists, this article says so.

## Does Meta AI in WhatsApp search the web before it answers?

Sometimes, and Meta AI decides when. Meta's help centre says Meta AI "may use sources and links from the internet to inform its responses", shown by selecting Sources under the response (checked on 22 September 2026). Meta's AI terms, dated 13 May 2026, say Meta may share information with "select partners such as search engines". No page we read says every answer is searched; the only web search switch Meta describes is for incognito chats in the Meta AI app, where search "begins toggled on", and we found none described for WhatsApp. So an answer without a Sources link has not been searched, and any reference in it is the model's guess. Before a guess goes into your work, EdCitation's free [Verify references](https://edcitation.com/verify-references) can look it up in the publisher's record.

### What Meta says about accuracy

Meta's help centre says an AI's response "may be inaccurate or inappropriate" and should not be used for important decisions. The AI terms make no warranty about accuracy, say outputs "may not be accurate, complete or current information", and tell you not to rely on them for decisions about medicine, finance, law or pharmaceuticals. A reference list is a set of claims about what exists; Meta's own terms tell you to check each one.

### Which model is behind Meta AI now

When Meta launched the Meta AI app on 29 April 2025 it said the app was "built with Llama 4", the open-weights model whose card is quoted below. On 24 July 2026 Meta said its closed Muse Spark 1.1 model was powering Meta AI in the app and at meta.ai, with WhatsApp to follow "in the coming weeks". So a WhatsApp chat may be answered by a different model from the app, and Meta's pages do not say which. A chat inside a messaging app is not a research tool: it keeps no reference list, and the only trace of a search is the Sources link when one appears. A free EdCitation account keeps reference lists of what you cite, and [Cite a source](https://edcitation.com/cite) builds each entry from the record rather than having a model type it.

## Does Mistral's Le Chat show sources for its references?

Only when it has searched, and the product is now called Vibe. Mistral announced on 28 May 2026 that "Le Chat is now Vibe"; chat.mistral.ai still opens, and conversations, settings and plans carry over (checked on 22 September 2026). Vibe has three modes: Chat, Work and Code.

### Web search, Deep Research and the plans

Mistral's pricing page lists a Free plan with "limited messages and web searches", Pro at $14.99 a month ($5.99 for verified students) with "more messages and web searches", Team at $24.99 per user a month, and Enterprise on request. Deep Research is now available only as a Skill in Work mode; starting it in Chat redirects you to Work. The Skill "plans the search, runs multiple web queries, synthesizes the sources, and returns a structured brief with citations", and Work mode also lists "Web search and Open URL" among its tools. Mistral's pages do not describe how sources appear in an ordinary Chat-mode answer, or which plan limits apply to Deep Research. Its product page names Mistral Medium 3.5 as its latest model.

### What Mistral says about accuracy

Mistral's help centre says Vibe can occasionally give "incorrect answers or facts" and that its models have "a limited understanding of the world and events". Its consumer terms, dated 5 August 2026, say outputs "may occasionally be inaccurate", that its products are "not authoritative or infallible sources of information", and that you should verify an output before relying on it.

## What have studies found about Meta AI, Llama and Mistral references?

Four studies between 2024 and 2026 tested these products or the models behind them, and each found invented or wrong references.

| Study | What was tested | References checked | What was found |
| --- | --- | --- | --- |
| Agrawal et al. (2024), *Findings of EACL 2024* | Llama 2 Chat at 7B, 13B and 70B, with three OpenAI models, asked for five paper titles on each of 200 computing topics | 1,000 titles per model | 68.3% (7B), 76.7% (13B) and 66.2% (70B) of the Llama 2 titles returned nothing when searched in quotation marks; GPT-4 was at 46.8% |
| Cabezas-Clavijo and Sidorenko-Bautista (2025), preprint, published in *JDIS* in 2026 | Le Chat's free version on Mistral Large, 7 to 9 February 2025, with seven other chatbots; ten APA references in each of five disciplines | 50 per chatbot, 400 in all | About 22% of Le Chat's references fully correct, about 46% partly correct and about 32% wrong or fabricated, read from the paper's Figure 1 |
| Tiller et al. (2026), *BMJ Open* | Meta AI with Gemini, DeepSeek, ChatGPT and Grok, 50 medical questions, February 2025 | Median 10 references per answer | Meta AI's reference lists scored a median 29.7% for completeness; no chatbot gave a fully complete and accurate list for any question |
| Naser (2026), preprint | Llama 4 Scout, Llama 4 Maverick and Mistral Small 3 with seven other models, through their APIs without search | 8,952, 8,105 and 7,976 citations | 49.4%, 33.5% and 48.3% of their citations could not be verified in Crossref, OpenAlex or Semantic Scholar |

### What the Meta AI and Le Chat tests showed

Tiller et al. (2026) scored a reference zero if it was not a published journal article, if its link was broken, or if the link opened an article "different from the one stated"; Meta AI's median of 29.7% sat with ChatGPT's 24.0%, and the two "did not differ from one another". The questions were medical, and the paper does not say which Meta AI interface was used. Cabezas-Clavijo and Sidorenko-Bautista (2025) found Le Chat neither among the worst, where Copilot fabricated 100%, Perplexity 72% and Claude 64%, nor among the two that fabricated none, Grok and DeepSeek. Its distinctive failure was age: its references averaged 45.3 years old, against 14.7 across all eight chatbots. The Naser (2026) figures are pure-model results with no search.

### What has not been tested

The Tow Center's March 2025 test (Jaźwińska and Chandrasekar, 2025) covered ChatGPT Search, Perplexity, Perplexity Pro, DeepSeek Search, Copilot, Grok-2, Grok-3 and Gemini, not Meta AI or Le Chat. Rao, Wong and Callison-Burch (2026) measured hallucinated URLs from ten search-backed systems, all from OpenAI, Google and Anthropic. The ChatGPT rates in the leader article, from Walters and Wilder (2023) onwards, are ChatGPT rates. We found no published test of Meta AI inside WhatsApp, or of Vibe's Deep Research.

## How do Meta AI and Le Chat typically go wrong with references?

In three ways the studies document, and one that follows from how the products work.

| Failure | The evidence | What catches it | In EdCitation |
| --- | --- | --- | --- |
| A real work with wrong details | Nearly half of Le Chat's references, wrong in the year or in where the work could be found | Comparing every field with the record | [Verify references](https://edcitation.com/verify-references) |
| A real work that is decades old | Le Chat's references averaged 45.3 years: not fake, but the wrong answer about current literature | A search for recent work on the same claim | [Find sources](https://edcitation.com/), with its years filter |
| A link that opens on something else | Tiller et al. (2026) scored such a reference zero | Opening the link and comparing the title | Rebuild the entry from the right DOI with Cite a source |
| A Sources link that does not cover the reference | Meta AI's link shows pages it retrieved, not that the reference came from one | Finding the paper on the page | Verify references, then read the page |

## How do I ask Meta AI or Le Chat for references I can check?

Ask for identifiers, ask it to cite only what it retrieved, and ask it to say "not found" instead of guessing. No prompt removes the problem; these make the output checkable.

1. **Get search on if you can.** In Vibe, use Work mode or type `/deep-research`. In the Meta AI app, an incognito chat has web search on by default. In WhatsApp, watch for the Sources link.
2. **Ask for what a checker needs.** "Give me five peer-reviewed sources on [topic]. For each, give the DOI and the URL of the publisher's page. Include only sources you retrieved in this conversation. If you cannot retrieve one, write 'not found' rather than guessing."
3. **Ask it to quote.** "For each source, quote the sentence that supports [claim] and say where on the page it appears." A model that cannot open the page cannot quote it.
4. **Do not ask it whether its references are real.** It answers that the way it wrote them. Copy the list out of the chat and into EdCitation's Verify references, which answers the question from the record.

Naser (2026) offers a hint, not a check: a work cited by three or more different models was real 95.6% of the time.

## What about a chatbot in a note-taking app or a university assistant?

Treat it as the model it runs on, with search off, until you can see otherwise. Meta's Llama models and many of Mistral's are open-weights, so anyone can build them into a note-taking app, a writing assistant or a campus chatbot, and the model card's warnings do not travel with the product. The Llama 4 model card says the model "may in some instances produce inaccurate or other objectionable responses" and that testing "has not covered, nor could it cover, all scenarios"; Mistral's consumer terms say outputs "may occasionally be inaccurate".

A product that wraps Llama 4 with a search step and shows links is a different tool from one that wraps Llama 4 alone, and nothing in the chat window tells you which you have, so test it.

### How to tell in one minute whether the chatbot searches

1. Ask "What is today's date, and what is one news story from this week?" A model without search gives a date from its training data, or says it cannot browse.
2. Ask for one source on your topic and look for a link. No link means no search. A link means open it and check that the page says what the chatbot said.
3. Ask "Did you search the web for that, or answer from memory?" Treat "from memory" as the answer, and "I searched" as unproven until a link opens.

A university's own assistant may be connected to the library catalogue rather than the open web; ask. A link into the catalogue can be checked; a bare reference cannot.

## How do I check the references these tools give me?

Check each one against the publisher's record, never against the chatbot: follow the DOI, compare the title and authors on the page that opens, and where a reference has no DOI, try its exact title, in quotation marks, in Crossref, PubMed or Google Scholar. [How to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) works through each step.

EdCitation's [Verify references](https://edcitation.com/verify-references) is the best tool for a list from Meta AI, Vibe or a campus chatbot, whichever model is underneath, because it never writes a reference and looks each one up in the publisher's record instead. That is the step these products skip: in Tiller et al. (2026), no chatbot, Meta AI included, gave a fully complete and accurate reference list for any of the 50 questions. Upload the paper or paste the list. Each entry is shown as verified, doubtful ("check this") or not found; retracted papers are flagged; and if an index gives no answer the entry reads "could not check", a different verdict from "not found". It costs nothing and asks for no account. [References from a file](https://edcitation.com/tools/references-from-a-file) and the other paid tools come with Pro, $8 a month, or Max, $24 a month; the [pricing page](https://edcitation.com/pricing) lists which plan has which.

When a reference fails, search Find sources for the claim it was holding up, and build the replacement from its DOI with [Cite a source](https://edcitation.com/cite).

## Quick questions

### Does Meta AI in WhatsApp make up sources?

It can. Meta AI writes with a language model and searches the web only some of the time, and Meta's own terms say outputs may not be accurate, complete or current. In a February 2025 audit of medical questions, its reference lists scored a median 29.7% for completeness.

### Is Mistral's Le Chat still available?

Yes, as Vibe since May 2026; chat.mistral.ai still opens and conversations carry over. Web search is included on the free plan with limits, and Deep Research now runs only in Work mode.

### Do Meta AI and Le Chat invent fewer references than ChatGPT?

The evidence does not say. In the one test with both, Le Chat fabricated fewer references than Copilot, Perplexity and Claude and more than Grok and DeepSeek, on 50 references each. In a separate medical audit, Meta AI's and ChatGPT's reference scores did not differ.

### Can I trust a reference that comes with a Sources link?

Not without opening it. The link shows pages were retrieved, not that the reference came from them. Find the paper on the page and compare title, authors and year, or look the reference up in EdCitation's Verify references.

## References

- Agrawal, A., Suzgun, M., Mackey, L., & Kalai, A. T. (2024). Do language models know when they're hallucinating references? In *Findings of the Association for Computational Linguistics: EACL 2024* (pp. 912-928). [https://doi.org/10.18653/v1/2024.findings-eacl.62](https://doi.org/10.18653/v1/2024.findings-eacl.62)
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2025). *Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate* [Preprint]. arXiv. [https://arxiv.org/abs/2505.18059](https://arxiv.org/abs/2505.18059)
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. *Columbia Journalism Review*. [https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php)
- Meta. (n.d.-a). *Report a source or link used in a response by Meta AI*. Meta Help Center. [https://www.meta.com/help/artificial-intelligence/578066098711082/](https://www.meta.com/help/artificial-intelligence/578066098711082/)
- Meta. (n.d.-b). *Start a chat with Meta AI*. Meta Help Center. [https://www.meta.com/help/artificial-intelligence/943942350800511/](https://www.meta.com/help/artificial-intelligence/943942350800511/)
- Meta. (n.d.-c). *Toggle web search for incognito chats with Meta AI on or off*. Meta Help Center. [https://www.meta.com/help/artificial-intelligence/1510309990445305/](https://www.meta.com/help/artificial-intelligence/1510309990445305/)
- Meta. (2025a, April 29). *Introducing the Meta AI app: A new way to access your AI assistant*. Meta Newsroom. [https://about.fb.com/news/2025/04/introducing-meta-ai-app-new-way-access-ai-assistant/](https://about.fb.com/news/2025/04/introducing-meta-ai-app-new-way-access-ai-assistant/)
- Meta. (2025b, April 5). *Llama 4 model card: Llama-4-Scout-17B-16E-Instruct*. Hugging Face. [https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct](https://huggingface.co/meta-llama/Llama-4-Scout-17B-16E-Instruct)
- Meta. (2026a, July 24). *Meta AI doesn't just think, it acts*. Meta Newsroom. [https://about.fb.com/news/2026/07/meta-ai-muse-spark-doesnt-just-think-it-acts/](https://about.fb.com/news/2026/07/meta-ai-muse-spark-doesnt-just-think-it-acts/)
- Meta. (2026b, May 13). *Meta AI terms of service*. [https://www.facebook.com/legal/ai-terms](https://www.facebook.com/legal/ai-terms)
- Mistral AI. (n.d.-a). *Can I trust the models' output?* Mistral Help Center. [https://help.mistral.ai/en/articles/704907-can-i-trust-the-models-output](https://help.mistral.ai/en/articles/704907-can-i-trust-the-models-output)
- Mistral AI. (n.d.-b). *Choose Chat, Work, or Code*. Mistral Docs. [https://docs.mistral.ai/vibe/choose-chat-work-code](https://docs.mistral.ai/vibe/choose-chat-work-code)
- Mistral AI. (n.d.-c). *Le Chat is now Vibe*. Mistral Help Center. [https://help.mistral.ai/en/articles/682992-le-chat-is-now-vibe](https://help.mistral.ai/en/articles/682992-le-chat-is-now-vibe)
- Mistral AI. (n.d.-d). *Pricing*. [https://mistral.ai/pricing](https://mistral.ai/pricing)
- Mistral AI. (n.d.-e). *Skills*. Mistral Docs. [https://docs.mistral.ai/vibe/work/skills](https://docs.mistral.ai/vibe/work/skills)
- Mistral AI. (2026a, August 5). *Consumer terms of service*. [https://legal.mistral.ai/terms/row-consumer-terms](https://legal.mistral.ai/terms/row-consumer-terms)
- Mistral AI. (2026b, May 28). *Vibe gets to work*. [https://mistral.ai/news/vibe-agent/](https://mistral.ai/news/vibe-agent/)
- Naser, M. Z. (2026). *How LLMs cite and why it matters: A cross-model audit of reference fabrication in AI-assisted academic writing and methods to detect phantom citations* [Preprint]. arXiv. [https://arxiv.org/abs/2603.03299](https://arxiv.org/abs/2603.03299)
- Rao, D., Wong, E., & Callison-Burch, C. (2026). *Detecting and correcting reference hallucinations in commercial LLMs and deep research agents* [Preprint]. arXiv. [https://arxiv.org/abs/2604.03173](https://arxiv.org/abs/2604.03173)
- Tiller, N. B., Marcon, A. R., Zenone, M., Kidd, K. E., Jeukendrup, A. E., Master, Z., & Caulfield, T. (2026). Generative artificial intelligence-driven chatbots and medical misinformation: An accuracy, referencing and readability audit. *BMJ Open, 16*(4), Article e112695. [https://doi.org/10.1136/bmjopen-2025-112695](https://doi.org/10.1136/bmjopen-2025-112695)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
