Ask Grok for sources and it answers with a list, sometimes with links, sometimes with a post on X. Some of the papers exist. Some do not. Some are real but cited for something they never say, and a link that opens is not proof of any of it.
Grok differs from the other chatbots in one respect: xAI built it to search X as well as the web. That helps with news and hurts with references, because a post on X is not a scholarly source. This guide is about Grok alone; the mechanism, and the studies on ChatGPT, are in Why AI tools invent references. For the scholarly references in whatever Grok hands you, EdCitation's free Verify references checks each against the publisher's record; for the posts, the guide shows how to reach the source behind them.
Does Grok make up sources?
Yes. Grok makes up sources when it answers from memory, and when it searches it can still attach a dead link, an unrelated page or a post on X to a claim. Both have been measured, and both fit what xAI says about the model.
What xAI says Grok knows without search
xAI's developer documentation, checked on 22 September 2026, gives Grok 4.7 a knowledge cutoff of May 2026 and says Grok "has no knowledge of current events or data beyond what was present in its training data" unless the web search or X search tools are enabled. The help centre on X says the model uses "next-token prediction model weights": it predicts the next likely word.
The same help page warns that Grok "may confidently provide factually incorrect information, missummarize, or miss some context" and asks users to "independently verify any information you receive". The consumer terms, last updated on 11 September 2026, say outputs may contain "hallucinations" and that the user "should not rely on Output as the truth". (The site's footer read SpaceXAI LLC when checked; this guide uses xAI.)
Why the fake ones look real
The parts are usually real while the whole is not. In Walters and Wilder (2023), GPT-3.5 invented 55% of its references and GPT-4 18%, and the inventions borrowed real authors and real journals. Nothing in Grok's design exempts it from that pattern; the 2026 test below shows it.
How does Grok search the web and X, and how does it show sources?
Grok decides for itself whether to search, and can draw on public posts on X as well as web pages. The help centre on X says Grok can "decide whether or not to search X public posts and conduct a real-time web search on the Internet". xAI's product page, checked on 22 September 2026, says Grok "Searches the web and X live", promises "Live citations from primary sources across the web", and tells users to "switch modes for search, reasoning, voice, or imagery" in the composer.
DeepSearch, Think and the sources panel
The names have changed since launch. xAI's Grok 3 announcement of 19 February 2025 introduced DeepSearch, "our first agent", and a "Think" button that shows the model's reasoning. Grok's release notes mention a DeepSearch button on 2 December 2025, a "Think Harder" pill on 15 August 2026 and, on 22 August 2026, a "sources panel". Whatever the button is called, with a search mode on Grok reads pages and lists them; with it off, it answers from training data.
What a Grok citation is, according to xAI's own docs
xAI's developer documentation says the citations field on a response lists every URL "the agent encountered during its search process", and "Not every URL in this list will necessarily be directly referenced in the final answer". It adds that "Enabling inline citations does not guarantee that the model will cite sources on every answer". So a source list can include pages the model looked at and discarded, an answer can arrive with no citations, and because the X search tool produces citations in the same form, a post can sit beside a journal article.
Posts on X are not scholarly sources
Nothing on the xAI pages read for this guide calls a post on X a scholarly source. The help centre describes the value of X posts as "up-to-date information and insights", which is what they are: a lead, a date, a name to search. If Grok cites a post for a factual claim, find the paper, report or dataset the post is about and cite that.
Which plans include which features
Checked on 22 September 2026, xAI's pricing page and the X help centre list the following, at US web prices.
| Plan | Price on the page | What the page says |
|---|---|---|
| Grok free | $0 | "Real-time web and X search", "within generous limits" |
| SuperGrok | $30 a month | "Grok 4.6 model", "Expert", "Higher rate limits across all features" |
| SuperGrok Plus | $100 a month | "Significantly higher usage", priority access |
| X Premium | $8 a month or $84 a year | "increased usage limits on Grok" |
| X Premium+ | $40 a month or $395 a year | "SuperGrok access" |
SuperGrok Lite, SuperGrok Heavy, Business and Enterprise tiers are also listed. A subscription buys limits and model access, not a different attitude to references.
What has been measured about Grok's references?
Three tests have measured Grok's references directly, and they disagree, mainly on whether search was on and how narrow the topic was. The evidence on Grok is thinner than on ChatGPT: three studies, only one of them peer reviewed, and none peer reviewed with search on.
| Study | What was tested | What it found about Grok |
|---|---|---|
| Jaźwińska and Chandrasekar (2025), Tow Center, Columbia Journalism Review | 8 AI search tools, 200 news queries each, search on | Grok 3: 94% incorrect; 154 of 200 citations led to error pages. Grok-2 often linked to a homepage, not the article |
| Cabezas-Clavijo and Sidorenko-Bautista (2025), preprint | 8 free chatbots, 50 references each, 7 to 9 February 2025 | Grok: 60% fully correct, none fabricated, 0.4 errors per reference; average age 21.4 years; 80% books |
| Seifi and Seyfi (2026), Critical Care Explorations | Grok-4, GPT-5.3, DeepSeek-V3; retrieval disabled; 100 references each; 10 March 2026 | Grok-4: 73% with an inaccuracy, 50% completely fabricated. DeepSeek-V3: 23% and 8%. GPT-5.3: 69% and 27% |
The search-on test: links that lead nowhere
Jaźwińska and Chandrasekar (2025) tested eight generative search tools on 1,600 queries. Grok 3 was worst at 94% incorrect, and "More than half of responses from Gemini and Grok 3 cited fabricated or broken URLs". The premium tools, Grok 3 and Perplexity Pro, had higher error rates than the free versions because they gave "definitive, but wrong, answers rather than declining to answer the question directly".
The memory test: half of them invented
Seifi and Seyfi (2026) queried Grok-4 with retrieval disabled and had two experts, blinded to the model, check every reference against PubMed, DOI, Google Scholar and Crossref. Of Grok-4's 100 references, 73% contained an inaccuracy and 50% were "completely fabricated", the highest rate of the three models; Grok-4 was 3.17 times more likely to hallucinate than DeepSeek-V3.
The test where Grok invented nothing, and why
Cabezas-Clavijo and Sidorenko-Bautista (2025) asked eight free chatbots for ten references for a final degree project in each of five disciplines, and Grok produced the best list: 60% fully correct, none fabricated. Its references were old, 21.4 years on average, and 80% were books, the well-known works of each field. The paper lists the model as Grok-3, though the test dates fall before xAI announced Grok 3 on 19 February 2025, so the version is not certain. Ask for classic textbooks and Grok does well; ask for recent papers on a narrow question and you are back in the 2026 conditions.
Rao, Wong and Callison-Burch (2026) found 3% to 13% hallucinated URLs across ten search-backed systems, but their preprint does not name Grok.
What have courts recorded about Grok citations?
Two published orders name Grok, and both show the failure a marker sees: nobody read the sources.
In Billups v. Louisville Municipal School District, the federal court for the Northern District of Mississippi found four problematic citations in one memorandum, one of them to a nonexistent case. The sanctions order of 19 December 2025 records that the drafting attorney "admitted that she used 'Grok', an external AI tool, to assist in drafting and research without verifying the accuracy of the output". The court disqualified all three attorneys from the case, sent the order to the Mississippi Bar, and ordered the firm to audit its filings.
In Noland v. Land of the Free, L.P., the California Court of Appeal found that 21 of the 23 case quotations in an opening brief were fabrications. Counsel said he had "enhanced" his drafts with ChatGPT, run them through other AI platforms "to check for errors", and not read the result; the court records that he blamed "generative AI sources such as ChatGPT, Claude, Gemini, and Grok". The opinion of 12 September 2025 imposed a $10,000 sanction. One chatbot's output had been checked by asking other chatbots, which is not checking.
How does Grok typically go wrong with a reference?
The patterns the sources above document, each with its check:
- A link to an error page. 154 of 200 in the Tow Center test. Open every link.
- A link to the homepage, not the article. Grok-2's habit. The page must show the title and authors.
- A post on X in the place of a source. Trace the post to the paper it is about; if the post gives no link, search its claim in EdCitation's Find sources.
- Real authors, invented title, when search was off. Half of Grok-4's references in the 2026 test. Search the exact title in quotation marks.
- A real paper with wrong details. 73% of Grok-4's references. Compare every element, then read the passage.
- A source list padded with pages it looked at. Keep only the ones the sentence rests on.
- Confidence where "not found" was the honest answer. No link and no DOI means unverified.
How do I ask Grok for references I can check?
Ask Grok to search, to cite only what it retrieved, to give a DOI and a URL for each source, to quote the passage, to leave X out, and to say "not found" rather than guess. These prompts make its references checkable; none makes them true.
Prompts that work
- "Search before answering. Cite only sources you retrieved in this conversation. For each one give the DOI, the URL and the sentence you are relying on, quoted exactly."
- "Do not use posts on X as sources. Journal articles, books and official reports only."
- "If you cannot find a source for a claim, write 'not found' next to it. Do not supply a plausible reference."
What to switch on
- In the app. Choose a search or reasoning mode in the composer rather than leaving it to the model, and open the sources panel. An answer with no sources came from memory.
- In the API. Enable the
web_searchtool and use itsallowed_domainsparameter, which takes up to five domains, to confine a search todoi.org,pubmed.ncbi.nlm.nih.govand a publisher or two. Leavex_searchoff. - Not a switch. No plan, mode or prompt removes the problem. xAI's terms put the responsibility for "evaluating the Output for accuracy" on the user.
How do I check what Grok gave me?
Check each reference against the publisher's record, never against Grok. The full method is in How to check whether a reference is real; in short:
- Follow the DOI at doi.org and compare the title and authors on the page with the reference.
- Without a DOI, put the exact title in quotation marks and search Crossref, PubMed or Google Scholar.
- Open the source and find the sentence Grok quoted.
- Keep "could not check" (a report, a thesis, a post) apart from "not found" (a recent article in no index).
- For more than a handful of references, let EdCitation's Verify references do steps 1 and 2 for the whole list, and spend your time on step 3.
Grok's citations are of three kinds, and only one can carry a factual claim:
| What Grok cited | What it is | What to do |
|---|---|---|
| A journal article or book | A scholarly source, if it exists | Verify the list; Cite a source builds the entry from the DOI or ISBN |
| A post on X | A lead, not a source | Search the claim in Find sources and cite the paper |
| A URL from the citations field the answer never used | A page the search passed over | Leave it out |
EdCitation's Verify references is the best tool for the first kind, because it answers the question Grok cannot: whether the reference is in the publisher's record. It never writes one, which is the difference that mattered in Noland: there one chatbot's output was "checked" by other chatbots, and 21 of 23 quotations were still fabricated. Upload the paper or paste the list; every entry is marked verified, doubtful ("check this") or not found, retractions are flagged, and when an index does not respond the entry says "could not check", which is never presented as "not found". That is free, with no account. Of the paid plans, Pro at $8 a month adds tools such as References from a file, and Max at $24 a month adds the Library, which suits the lists Grok did best on. It searches a catalogue of textbooks and scholarly books, shows where to buy a copy or find one in a library near you, and opens more than 5 million books to read or borrow.
Quick questions
Does Grok make up references?
Yes. With retrieval disabled, Seifi and Seyfi (2026) found 50 of 100 Grok-4 references completely fabricated and 73 with some inaccuracy. With search on, the Tow Center test found 154 of 200 Grok 3 citations led to error pages.
Does DeepSearch or a search mode stop Grok inventing sources?
No. Search changes the failure from an invented paper to a wrong or dead link, and xAI's own documentation says a citation list can include pages the model looked at but did not use.
Can I cite a post on X that Grok found?
Only as a post, and only if the post itself is what you are discussing. For a factual claim, find the paper, report or dataset behind the post and cite that.
Can I ask Grok to check its own references?
No. It answers that question the same way it wrote the list, and in Noland, where counsel ran AI-drafted briefs through other AI tools to check them, 21 of 23 quotations were still fabricated. Use a lookup instead, such as EdCitation's Verify references, which checks each entry against the publisher's record.
Is Grok worse than ChatGPT at references?
The tests disagree. In the February 2025 free-tier comparison Grok did best of eight chatbots; in the March 2026 retrieval-off test Grok-4 fabricated 50% against 27% for GPT-5.3.
References
- Billups v. Louisville Municipal School District, No. 1:24-cv-74-SA-RP (N.D. Miss. Dec. 19, 2025). Sanctions order. https://storage.courtlistener.com/recap/gov.uscourts.msnd.49169/gov.uscourts.msnd.49169.79.0.pdf
- Cabezas-Clavijo, Á., & Sidorenko-Bautista, P. (2025). Assessing the performance of 8 AI chatbots in bibliographic reference retrieval: Grok and DeepSeek outperform ChatGPT, but none are fully accurate [Preprint]. arXiv. https://arxiv.org/abs/2505.18059
- Grok. (2026, September 12). Grok release notes. https://grok.com/release-notes
- Jaźwińska, K., & Chandrasekar, A. (2025, March 6). AI search has a citation problem. Columbia Journalism Review, Tow Center for Digital Journalism. https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php
- Noland v. Land of the Free, L.P., 114 Cal. App. 5th 426 (Cal. Ct. App. Sept. 12, 2025). https://www.law.berkeley.edu/wp-content/uploads/archive/2025/12/Noland-v-Land-of-the-Free-LP.pdf
- Rao, D., Wong, E., & Callison-Burch, C. (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents [Preprint]. arXiv. https://arxiv.org/abs/2604.03173
- Seifi, A., & Seyfi, A. (2026). Hallucination rate of peer-reviewed citations generated by large language models in neurocritical care. Critical Care Explorations, 8(9), Article e1474. https://doi.org/10.1097/CCE.0000000000001474
- SpaceXAI. (n.d.-a). Citations. SpaceXAI Docs. https://docs.x.ai/developers/tools/citations
- SpaceXAI. (n.d.-b). Grok. https://x.ai/grok
- SpaceXAI. (n.d.-c). Grok models and pricing. SpaceXAI Docs. https://docs.x.ai/developers/models
- SpaceXAI. (n.d.-d). Pricing: Compare Grok plans. https://x.ai/pricing
- SpaceXAI. (n.d.-e). Web search. SpaceXAI Docs. https://docs.x.ai/developers/tools/web-search
- SpaceXAI. (2026, September 11). Terms of service: Consumer. https://x.ai/legal/terms-of-service
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5
- X. (n.d.-a). About Grok, your AI assistant on X. X Help Center. https://help.x.com/en/using-x/about-grok
- X. (n.d.-b). About X Premium. X Help Center. https://help.x.com/en/using-x/x-premium
- xAI. (2025, February 19). Grok 3 Beta: The age of reasoning agents. https://x.ai/news/grok-3