Microsoft Copilot answers with numbered footnotes that open real web pages, which makes its citations look safer than a chatbot's bare list of references. A footnote says only that Bing found a page, not that the page supports the sentence, that it is the paper rather than a summary of it, or that Copilot searched at all.
This guide covers Microsoft Copilot's fake citations and how to catch them: what Microsoft says its citations are, what tests found, how it goes wrong with academic references, and how to ask for references you can verify. The mechanism is in Why AI tools invent references and the manual checks in How to check whether a reference is real. Where Copilot's footnotes lead to Bing results, EdCitation's free Verify references takes each reference to the publisher's record instead. Product details were checked on Microsoft's pages on 22 September 2026.
What does Microsoft say Copilot's citations are?
Microsoft says a Copilot citation is a link to the source Copilot used, and that you should open it. Its support page on what Copilot uses lists three kinds of source: the web, meaning "current, publicly available information on the web"; your work data, inside an organisation; and whatever you attach to the prompt. It says this grounding gives "citations you can check", and tells you to "review the sources and confirm critical details before you share or act on them" (Microsoft, n.d.-e). The supplemental terms of use, effective 18 August 2026, are blunter: "Copilot can make mistakes, and it may not work as intended", so "always verify the accuracy of information presented by Copilot before you rely on it" (Microsoft, 2026b).
Which Copilot you are using matters
"Copilot" is several products that ground answers differently. Microsoft's overview page, updated on 17 September 2026, says Microsoft 365 Copilot is now named simply Microsoft Copilot, and separates three experiences (Microsoft, 2026c):
| Product | Who has it | Where its citations come from |
|---|---|---|
| The Copilot app | Anyone | Bing web search, when Copilot searches or you pick Search mode |
| Copilot Chat (Basic) | Eligible Microsoft 365 work and school accounts | Web data; work files only if you upload or open them |
| Copilot with the add-on licence | Organisations paying for it | Web data plus the organisation's files, mail and chats |
The consumer app's models are OpenAI's: Think Deeper is "powered by the latest reasoning models from OpenAI" and Smart mode by GPT-5 (Microsoft, n.d.-b). A study of ChatGPT is still not a study of Copilot: the search step and the interface differ, and that is where citations are won or lost.
How the search behind a citation works
Copilot does not send your prompt to Bing. Microsoft's Learn article on web search, dated 18 August 2026, says Copilot picks terms from the prompt, generates a query of a few words, and sends that. The queries are shown in the citation section of the response, but only in Copilot Chat, not in the Copilot pane inside Word or PowerPoint, and only for 24 hours (Microsoft, 2026a). A citation is the result of a search on terms you did not choose; seeing them is the first check.
What have tests of Bing Chat and Copilot found?
When Copilot searches, its links point at something real but often not the right thing; when it does not search, it invents. The Copilot-specific evidence is thin and the versions tested are old.
| Study | Test | Bing Chat or Copilot result |
|---|---|---|
| Liu et al. (2023) | Four generative search engines, 1,450 queries, early 2023 | 89.5% of citations supported their sentence; 58.7% of sentences fully supported |
| Aljamaan et al. (2024) | Six chatbots, 10 medical prompts, 10 references each | Median hallucination score 11, the highest, level with ChatGPT 3.5 |
| Siyad et al. (2025) | Free Copilot, ten dental references per topic, March 2024 | 89.4% of 47 titles existed; 10.6% of references had a correct link |
| Jaźwińska and Chandrasekar (2025) | Eight AI search tools, 200 news excerpts | No answer to 104; 16 fully correct, 66 wholly wrong (counted from the report's chart) |
| Cabezas-Clavijo and Sidorenko-Bautista (2025) | Eight free chatbots, 10 APA references in each of five fields, February 2025 | All 50 references fabricated |
With search on: Bing Chat in 2023 and Copilot in 2025
Liu et al. (2023) had 34 annotators check four generative search engines in early 2023. Bing Chat had the best citation precision of the four, 89.5%, but a recall of 58.7%, so about two sentences in five were not fully backed by any citation. It paraphrased its sources most closely, and on open essay questions it copied statements from pages on the wrong topic: the citation checked out and the answer did not.
In 2025 the Tow Center asked eight AI search tools to identify the source of 200 news excerpts (Jaźwińska & Chandrasekar, 2025). Copilot gave no answer to 104 of the 200, more than it answered, although no publisher in the set had blocked it; of the 96 it answered, 16 were completely correct, 14 partly wrong and 66 completely wrong. The counts are read from the report's chart.
Asked for academic references
Cabezas-Clavijo and Sidorenko-Bautista (2025) asked the free versions of eight chatbots, on 7 to 9 February 2025, for ten academic references in APA 7 in each of five fields, and searched every title. Across all eight, 39.8% of the 400 references were wrong or fabricated. Copilot's 50 were all fabricated, with 4.2 errors per reference out of five, and it gave the same ten in every field, changing only the title and the journal: the same Green and Brown, 2013, volume 27, issue 4, pages 201 to 215, appeared under a mechanical engineering title, an organic chemistry title and an art history title.
Siyad et al. (2025), asking the free Copilot in March 2024 for ten Vancouver references with links on cone-beam CT topics, got 47 unique references of which 89.4% had a title that existed, but 63.8% came with no link and only 10.6% had a correct one. Aljamaan et al. (2024), testing the Bing chatbot in 2023, found its median reference hallucination score was 11, the highest recorded and level with ChatGPT 3.5; relevance to the prompt failed most often, in 61.6% of references. Rao et al. (2026), who checked more than 53,000 URLs cited by search-backed models, included no Microsoft product, and no published test of Deep Research or Researcher was found.
Why does Microsoft Copilot give fake citations?
Copilot gives fake citations for the reason every chatbot does, plus one of its own: it does not always search, and nothing on screen says which answers were searched. A language model writes the most likely next words, and a reference is the easiest academic shape to imitate: Walters and Wilder (2023) found 55% of GPT-3.5's references and 18% of GPT-4's were fabricated, and Copilot's modes run on the same family of models.
By Microsoft's own description Copilot searches only when it judges that web information would improve the answer, searches on a few words it chose, and writes the sentence afterwards (Microsoft, 2026a). Each gap has a failure to go with it.
The five ways a Copilot citation goes wrong
- The page exists but does not support the sentence. One Bing Chat citation in ten failed this test in Liu et al. (2023). Read the passage, not the URL.
- The citation is a secondary page, not the paper. Bing ranks pages, so Copilot cites what ranks, which is often a write-up of the paper. Cite the paper.
- A real paper cited for a claim it does not make. Relevance to the question was the component that failed most often in Aljamaan et al. (2024).
- References written without a search. Ask for "ten references in APA" and Copilot may write ten, which is how Cabezas-Clavijo and Sidorenko-Bautista (2025) got fifty fakes. A list like that has to pass a lookup in the record, such as Verify references, before any of it is used.
- Asked to verify its own list, it answers from the same model. A chatbot confirms a fake reference the same way it wrote it.
How do I make Copilot search, and open its citations?
Choose the mode that searches, then use the Sources button, because an answer without visible sources may not have searched at all.
- Pick a searching mode. Microsoft's conversation modes page lists Quick response, Think Deeper, Study and Learn, Smart and Search, and it is Search that it describes as bringing "up-to-date answers from the web, with citations" (Microsoft, n.d.-b). In Copilot Chat for work, check the Web search toggle under Settings, Personalization, Advanced; if an administrator has turned web search off, the toggle is dimmed (Microsoft, 2026a).
- Find the Sources button under the answer. It lists every source used and, in Copilot Chat, the exact query sent to Bing (Microsoft, n.d.-a; Microsoft, 2026a). If the query misses the point of your question, so will the citations.
- Open each inline citation. Hover over the number and select the source to open it in a side pane, or open it from the Sources list (Microsoft, n.d.-a).
- Read the passage, not the page title. If the sentence Copilot relied on is not there, the citation has failed; look for a published source that does say it in EdCitation's Find sources.
- Note where citations do not appear. Not in the Copilot pane inside Word or PowerPoint, and the queries leave the thread after 24 hours (Microsoft, 2026a).
Deep Research, Think Deeper and Researcher
Microsoft is retiring Deep Research in the consumer app from 18 August 2026 and points Microsoft 365 Premium subscribers to Researcher (Microsoft, n.d.-c). Researcher, for Microsoft 365 Premium and Pro subscribers and organisations with the add-on licence, produces a "structured, source-cited report" from work content and the web (Microsoft, n.d.-d). Search mode is available to every user; a paid plan buys priority access, not a different kind of citation.
How do I ask Copilot for references I can check?
Ask for identifiers that can be looked up, tell Copilot to cite only pages it opened, and give it permission to say "not found". Prompts like these make the output checkable, not correct.
- "Search the web. Give me up to five peer-reviewed papers on [topic]. For each, give the DOI, the URL of the page you opened, and one sentence quoted from that page that supports the claim. If you did not open a page for a paper, write 'not found'."
- "Cite only sources that appear in your Sources list. Do not add references from memory."
Then resolve each DOI at doi.org and compare the title and authors; a DOI that resolves is not yet proof. No prompt removes the problem. The prompt in Cabezas-Clavijo and Sidorenko-Bautista (2025) was specific and got fifty fakes; Siyad et al. (2025) asked for links and got them for a third of the references; a quotation can be invented as easily as a title.
Prompts that make it worse
A fixed number of references pushes the model to fill the list; a narrow topic pushes it towards invention, because the less there is to retrieve, the more it writes; and "confirm these are real" produces reassurance, not a check. If the number comes from your assignment, EdCitation's Check your paper reads the brief into a checklist, number of sources included, and the sources themselves are better found among published works than written to order.
How do I check the references Copilot gave me?
Check each reference against the publisher's record, never against Copilot: resolve the DOI, put the exact title in quotation marks into a scholarly index, compare the authors and the issue, and read the part you rely on. How to check whether a reference is real covers the full procedure, and what "not found" does and does not mean.
EdCitation's Verify references is the best tool for a Copilot list, because it goes where Copilot's search does not: to the publisher's record rather than to Bing, and it never writes a reference of its own. The fifty references Copilot wrote for Cabezas-Clavijo and Sidorenko-Bautista (2025) were all fabricated, down to one Green and Brown paper reused across five fields, and a fabricated reference has no record for a lookup to find. Paste the list or upload the paper. Each entry gets one of three verdicts, verified, doubtful ("check this") or not found; a retracted paper carries a flag; and "could not check", which appears when an index does not answer, is kept apart from "not found". The check is free and asks for no account.
What Copilot gave you decides what to do next:
| What Copilot gave you | What to do | In EdCitation |
|---|---|---|
| A footnote to a web page | Read the passage behind it | Find sources for the paper, if the page is a write-up |
| A list of references in APA 7 | Look every one up, then rebuild the survivors | Verify references, then Cite a source in APA 7 from the DOI |
| References already in a Word document | Check them where they sit, with the in-text citations | References from a file, on Pro at $8 a month |
| A citation to a SharePoint file or an email | Nothing to verify: it is not literature | Not a job for a reference checker |
At work, a Copilot citation to a SharePoint document or an email says only that the file exists and you may see it (Microsoft, 2026c). Pro also brings Mechanics QA, which checks the Word file's formatting against the style; Max, $24 a month, adds Theoretics QA and the Library.
Quick questions
Does Microsoft Copilot make up references?
Yes, when it writes references without searching. In a February 2025 test, all 50 references the free Copilot gave for coursework were fabricated, with the same authors, year, volume and pages reused across five fields.
Are Copilot's citations real links?
Usually, when it has searched: its footnotes open pages that Bing returned. In the 2023 test, 89.5% of Bing Chat's citations supported their sentence, but two sentences in five had no full support, so open every one and read the passage.
Can I ask Copilot to check whether its references are real?
No. It answers from the same model that wrote them. Resolve the DOI at doi.org, or paste the list into EdCitation's Verify references, which looks each reference up in the publisher's record rather than in Bing.
Is Copilot better than ChatGPT at references?
It turns on whether Copilot searched. In a March 2024 dental test, 89.4% of Copilot's titles existed against 30.9% of ChatGPT 3.5's; in a February 2025 test Copilot fabricated all 50 of its references and ChatGPT fewer. Neither replaces checking.
References
- Aljamaan, F., Temsah, M.-H., Altamimi, I., Al-Eyadhy, A., Jamal, A., Alhasan, K., Mesallam, T. A., Farahat, M., & Malki, K. H. (2024). Reference hallucination score for medical artificial intelligence chatbots: Development and usability study. JMIR Medical Informatics, 12, Article e54345. https://doi.org/10.2196/54345
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2025). Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate [Preprint]. arXiv. https://arxiv.org/abs/2505.18059
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. Columbia Journalism Review. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
- Liu, N. F., Zhang, T., & Liang, P. (2023). Evaluating verifiability in generative search engines [Preprint]. arXiv. https://arxiv.org/abs/2304.09848
- Microsoft. (n.d.-a). Control and review sources of Microsoft Copilot Chat's responses. Microsoft Support. Retrieved September 22, 2026, from https://support.microsoft.com/en-us/microsoft-365-copilot/control-review-sources-copilot-chat
- Microsoft. (n.d.-b). Conversation modes: Quick, Think Deeper, Deep Research. Microsoft Support. Retrieved September 22, 2026, from https://support.microsoft.com/en-us/microsoft-copilot/conversation-modes-in-microsoft-copilot
- Microsoft. (n.d.-c). Deep Research in Microsoft Copilot. Microsoft Support. Retrieved September 22, 2026, from https://support.microsoft.com/en-us/microsoft-copilot/deep-research-in-microsoft-copilot
- Microsoft. (n.d.-d). Get started with Researcher in Microsoft 365 Copilot. Microsoft Support. Retrieved September 22, 2026, from https://support.microsoft.com/en-us/microsoft-365-copilot/get-started-with-researcher-in-microsoft-365-copilot
- Microsoft. (n.d.-e). What information does Copilot use to answer my prompt? Microsoft Support. Retrieved September 22, 2026, from https://support.microsoft.com/en-us/microsoft-365-copilot/what-information-does-copilot-use-to-answer-my-prompt
- Microsoft. (2026a). Data, privacy, and security for web search in Microsoft Copilot and Microsoft Copilot Chat. Microsoft Learn. https://learn.microsoft.com/en-us/microsoft-365/copilot/manage-public-web-access
- Microsoft. (2026b). Microsoft Copilot supplemental terms of use. https://www.microsoft.com/en-us/microsoft-copilot/for-individuals/termsofuse
- Microsoft. (2026c). What is Microsoft Copilot? Microsoft Learn. https://learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-overview
- Rao, D., Wong, E., & Callison-Burch, C. (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents [Preprint]. arXiv. https://arxiv.org/abs/2604.03173
- Siyad, A. M., Akhila, A. S., Ramanarayanan, S., Mustafa, S. M., James, J. M., & Babu, P. (2025). Assessing the validity and accuracy of artificial intelligence technologies for identifying relevant literature in dentistry. Journal of Nature and Science of Medicine, 8(2), 135-138. https://doi.org/10.4103/jnsm.jnsm_170_24
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5