# Are Elicit, Consensus and Scite citations reliable?

**Are the citations from AI research assistants like Elicit, Consensus and Scite reliable?** Elicit, Consensus and Scite cite real papers, because they search indexes rather than writing from memory. What goes wrong is the summary, the extracted claim, the meter and the label: Elicit found 39.5% of the studies in four reviews, and Scite filed 96 of 98 citations as mentioning. Read each paper, then check the list with EdCitation's free Verify references, which looks each entry up.

Published 2026-09-22 by EdCitation. https://edcitation.com/newsletter/ai-research-assistants-elicit-consensus-scite-citations

AI research assistants such as Elicit, Consensus and Scite work differently from a chatbot. Each searches an index of published papers first and writes second, so the papers they cite nearly always exist. That removes the invented reference, the best-known failure of ChatGPT, and moves the risk somewhere quieter: into the one-line summary, the number pulled from an abstract, the answer stitched together across twenty papers, the label on a citation, and the papers the index never held. EdCitation publishes this comparison.

Every fact about the three products was read on that product's own site on 22 September 2026; prices and features change. SciSpace works the same way and is widely used; its help and data-source pages sat behind a verification step when we checked, so this guide says nothing about it rather than repeat its marketing.

## Why do Elicit, Consensus and Scite cite real papers when ChatGPT does not?

They cite real papers because each paper is retrieved from an index before any text is written. A language model on its own predicts the next words: Walters and Wilder (2023) found that 55% of the references written by GPT-3.5 and 18% of those written by GPT-4 did not exist. [Why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references) sets out the mechanism. Preisler (2024) tested Elicit and Scite with two other assistants and found no fabricated references at all; the errors were in bibliographic details, which is the failure that reaches a marker.

### The failure that survives a real index

A real index removes the invented paper. It does nothing about the wrong claim, because the sentence you read is still written by a language model. Peters and Chin-Yee (2025) tested ten language models on 4,900 summaries of scientific texts: three of them, DeepSeek, ChatGPT-4o and LLaMA 3.3 70B, produced conclusions broader than the original in 26% to 73% of cases, model summaries were nearly five times as likely as human-written ones to overgeneralise, and asking the model to be accurate made it worse. Every product below writes that kind of summary, and none of them checks it against the paper for you.

## Where does Elicit go wrong?

Elicit misses studies, occasionally reports things a paper does not contain, and reads only the abstract of a paper it cannot open. Three peer-reviewed tests show each of those.

- Lau and Golder (2025) ran Elicit Pro in Review mode in February and March 2025 against four published reviews: it found 39.5% of the studies those reviews had included, against 94.5% for the original database searches. Its precision was far higher, 41.8% against 7.55%, and it turned up 11 eligible studies the originals had missed, but a search that finds four studies in ten is not a literature search on its own.
- Lagisz et al. (2026), on a paid plan across seven reviews, found Elicit reporting conflict-of-interest and author-contribution statements in papers that had none, and its written reasoning for the same paper matched only 30% of the time between two accounts.
- Hilkenmeier et al. (2025), across 43 studies and 602 data points, put Elicit's extraction at 81.4% accurate against 86.7% for a human reviewer, not a significant difference; where the two agreed, the value was correct every time.

So what Elicit extracts is usually right, and what it leaves out, or occasionally supplies from nowhere, is the part to check.

### What Elicit's own pages say

Elicit says it searches more than 138 million papers from Semantic Scholar, PubMed and OpenAlex, does not cover books, dissertations or non-academic publications, reads the title and abstract rather than the full text of a paywalled paper, and has gaps in Mandarin-language publishing (checked on 22 September 2026). Its help pages say it "can miss the nuance of a paper or misunderstand what a number refers to" and does not judge whether one paper is more trustworthy than another. Its own evaluation, posted on 6 May 2026 and built from 994 open-access Cochrane reviews, reports 95.0% recall for search and 95.6% correct extraction; that is the vendor's test on the vendor's choice of reviews, and the independent figures above are lower. Basic is free; Pro is $49 per user a month billed as $588 a year; Scale is $169 a month billed as $2,028 a year.

## What does the Consensus Meter get wrong?

The Consensus Meter is a count of labels on twenty papers, not a finding. For a yes-or-no question, Consensus takes the top 20 papers its search returns, has a model label each one Yes, No, Possibly or Mixed, and shows the split; it needs at least five papers to appear (checked on 22 September 2026).

Consensus's own help page lists the failures. The Meter is "not a perfect reflection of all the science on a topic", because it uses the 5 to 20 most relevant results for one search. The model "will occasionally incorrectly classify results". It can miss the detail in the question: Consensus's own example is a question about adults answered Yes by a paper about children. Tay (2025) adds that a study of 50 people counts on the Meter the same as one of 5,000. We found no peer-reviewed test of the Meter's labels, so its error rate is unknown.

### What Consensus searches, and what it admits

Consensus says its database holds more than 220 million papers from Semantic Scholar, OpenAlex, its own crawl and full-text agreements with seven publishers including Wiley, Sage and the APA, updated weekly; without full text it works from the abstract. Retracted papers carry a badge and, Consensus says, are excluded from its AI analyses. Its Responsible AI page separates three kinds of hallucination, fake sources, wrong facts from memory and misread sources, and says that because search comes first "only the third type is possible", and that a model "can misinterpret a paper and summarize it incorrectly". The third type is the one a marker catches. The free plan gives 10 Pro messages and up to 3 Deep reviews a month; Pro is $12 a month billed as $144 a year and Deep is $45 a month billed as $540 a year.

## How often is a Scite Smart Citation label wrong?

Often enough that a count of contrasting citations means little. A Smart Citation is the sentence in which a later paper cites a work, with a label from a deep learning model: supporting, contrasting or mentioning. Scite's help centre says the label describes rhetorical function, not sentiment: a supporting citation offers its own evidence for the claim, and a sentence beginning "consistent with these findings" without new evidence is filed as mentioning (checked on 22 September 2026). Scite's home page puts its coverage at more than 318 million full-text articles and 1.6 billion citations.

Nicholson et al. (2021), the paper Scite's help pages point to, reports early F-scores of 96.3% for mentioning, 55.3% for supporting and 20.5% for contrasting, and says the production model is tuned so that every class has a precision above 80%. The cost is that doubtful cases are filed as mentioning. Bakker et al. (2023) saw exactly that: of 98 citations they classified by hand as 42 supporting, 39 mentioning and 17 contrasting, Scite labelled 2 supporting, 96 mentioning and none contrasting. A supporting label is probably right; zero contrasting citations proves nothing, and Scite's own pages add that a retracted paper can show none because nobody published a contrast or Scite has not ingested it.

Scite's report pages show corrections and retractions from Crossref and PubMed, and its paid Reference Check takes a manuscript PDF and lists the references with editorial notices. The free Connect plan has no Assistant or search; Basic is $20 a month billed yearly and Pro $50, each with a 7-day trial.

## What can still go wrong when the papers are real?

| What goes wrong | Where it shows | What to do |
| --- | --- | --- |
| The summary says more than the paper | Every product's summary and synthesised answer | Read the passage the claim rests on |
| A claim from the abstract that the full text qualifies | Elicit on paywalled papers; Consensus without full text | Open the full text before citing a number |
| A meter or a count read as a verdict | The Consensus Meter; Scite's counts | Treat it as a first sort, then read the studies |
| A retracted paper in the results | Consensus badges and excludes retractions; Scite shows notices; Elicit's pages say nothing | Check the publisher's page yourself |
| A source the index never held | Elicit excludes books and dissertations; Consensus says it lacks some research | Search the library catalogue too |
| An exported reference with a wrong detail | All three export from index metadata | Take the DOI and build the reference from the record |

The last row is the one a marker sees: Preisler (2024) found no invented references but did find wrong and missing bibliographic details.

## Where is EdCitation the better tool?

For checking that every reference in a finished paper is real, EdCitation is the best tool on this page, and the only one that does it free with no account. [Verify references](https://edcitation.com/verify-references) takes the paper or the pasted list and looks every entry up in the publisher's record, through Crossref, DataCite, PubMed, Open Library and the Retraction Watch database, and never writes a reference. Each entry comes back verified, doubtful or not found, retracted papers are flagged, and an index that did not answer is shown as "could not check", never as "not found". None of the three research assistants does this: nothing on Elicit's or Consensus's pages starts from a finished list, and Scite's Reference Check does, from a PDF, on a paid plan.

For the reference itself, [Cite a source](https://edcitation.com/cite) is the better way to build it than exporting from any of the three: paste the DOI and the reference is set from the publisher's record in APA 7, MLA 9, Chicago 18, Harvard, IEEE or Vancouver, screened for retraction first, so the wrong and missing details Preisler (2024) found in the assistants' references do not arise. For a claim the index never held, [Find sources](https://edcitation.com/) searches about 300 million published works by topic or by the claim itself, free.

What the three do that EdCitation does not, in a sentence: search by research question and return a synthesised answer, extract data from papers into a table, and classify how later papers cite a work. EdCitation never writes any part of a paper and never summarises one for you; it finds, cites, verifies, and checks a paper against its own assignment.

## How do I use an AI research assistant so the result can be checked?

1. Ask the question, then open every paper you intend to cite.
2. Read the passage, not the snapshot. Find the sentence in the paper that says what the summary says.
3. Take the DOI, not the formatted reference, and build the reference from it with [Cite a source](https://edcitation.com/cite).
4. Treat the Consensus Meter and Scite's counts as a first sort, and read the studies behind the positions that matter.
5. Search beyond the index for books, reports, theses and non-English work; [how to check whether a reference is real](https://edcitation.com/newsletter/how-to-check-a-reference-is-real) covers those.
6. Before you submit, put the whole reference list through [Verify references](https://edcitation.com/verify-references).

## Quick questions

### Does Elicit make up references?

Not in the published tests. Elicit retrieves papers from an index of more than 138 million records before writing anything, and Preisler (2024) found no fabricated references from it. Its documented failures are the studies it misses and the details it gets wrong, so read the paper before you cite an extracted value.

### Is the Consensus Meter a meta-analysis?

No. It is an AI classification of the top 20 papers for one search into Yes, No, Possibly and Mixed; Consensus says the model can misclassify and can miss the detail in the question, and no peer-reviewed test of it exists. Use it to see where the disagreement lies, then read the studies.

### What does a Scite "supporting" citation prove?

That a model judged a later paper to have offered its own evidence for the claim. Scite tunes the classifier for precision above 80%, so a supporting label is usually right, but in one hand-checked test it found 2 of 42 supporting citations and none of 17 contrasting ones.

### Which tool checks that my reference list is real?

EdCitation's Verify references, free and with no account: every entry is looked up in the publisher's record, never written, and comes back verified, doubtful or not found with retractions flagged. Of the three assistants, only Scite offers a comparable check, on a paid plan and from a PDF.

## References

- Bakker, C., Theis-Mahon, N., & Brown, S. J. (2023). Evaluating the accuracy of scite, a smart citation index. *Hypothesis: Research Journal for Health Information Professionals, 35*(2). [https://doi.org/10.18060/26528](https://doi.org/10.18060/26528)
- Consensus. (n.d.-a). *Consensus research database*. The Consensus Help Center. [https://help.consensus.app/en/articles/10055108-consensus-research-database](https://help.consensus.app/en/articles/10055108-consensus-research-database)
- Consensus. (n.d.-b). *Responsible AI & limitations*. The Consensus Help Center. [https://help.consensus.app/en/articles/10046838-responsible-ai-limitations](https://help.consensus.app/en/articles/10046838-responsible-ai-limitations)
- Consensus. (2026, April 22). *The Consensus Meter*. The Consensus Help Center. [https://help.consensus.app/en/articles/10069920-the-consensus-meter](https://help.consensus.app/en/articles/10069920-the-consensus-meter)
- Elicit. (n.d.). *Elicit's source for papers*. Elicit Help Center. [https://support.elicit.com/en/articles/14758040-elicit-s-source-for-papers](https://support.elicit.com/en/articles/14758040-elicit-s-source-for-papers)
- Elicit. (2026, May 6). *Evaluating Elicit's systematic literature review capabilities*. [https://elicit.com/blog/evaluating-elicit-slr](https://elicit.com/blog/evaluating-elicit-slr)
- Hilkenmeier, F., Pelzer, M., Stierle, C., & Fink-Lamotte, J. (2025). Evaluating the AI tool "Elicit" as a semi-automated second reviewer for data extraction in systematic reviews: A proof-of-concept. *Social Science Computer Review*. Advance online publication. [https://doi.org/10.1177/08944393251404052](https://doi.org/10.1177/08944393251404052)
- Lagisz, M., Mizuno, A., Morrison, K., Pollo, P., Ricolfi, L., Yang, Y., & Nakagawa, S. (2026). Using Elicit AI research assistant for data extraction in systematic reviews: A feasibility study across environmental and life sciences. *Research Synthesis Methods*. Advance online publication. [https://doi.org/10.1017/rsm.2026.10080](https://doi.org/10.1017/rsm.2026.10080)
- Lau, O., & Golder, S. (2025). Comparison of Elicit AI and traditional literature searching in evidence syntheses using four case studies. *Cochrane Evidence Synthesis and Methods, 3*(6), Article e70050. [https://doi.org/10.1002/cesm.70050](https://doi.org/10.1002/cesm.70050)
- Nicholson, J. M., Mordaunt, M., Lopez, P., Uppala, A., Rosati, D., Rodrigues, N. P., Grabitz, P., & Rife, S. C. (2021). scite: A smart citation index that displays the context of citations and classifies their intent using deep learning. *Quantitative Science Studies, 2*(3), 882-898. [https://doi.org/10.1162/qss_a_00146](https://doi.org/10.1162/qss_a_00146)
- Peters, U., & Chin-Yee, B. (2025). Generalization bias in large language model summarization of scientific research. *Royal Society Open Science, 12*(4), Article 241776. [https://doi.org/10.1098/rsos.241776](https://doi.org/10.1098/rsos.241776)
- Preisler, A. (2024). *Correctness and quality of references generated by AI-based research assistant tools: The case of Scopus AI, Elicit, SciSpace and Scite in the field of business administration* [Master's thesis, University of Graz]. ResearchGate. [https://doi.org/10.13140/RG.2.2.15104.34567](https://doi.org/10.13140/RG.2.2.15104.34567)
- Research Solutions. (n.d.-a). *How are citations classified?* Research Solutions Help & Support Center. [https://help.researchsolutions.com/hc/en-us/articles/31949617584148-How-are-citations-classified](https://help.researchsolutions.com/hc/en-us/articles/31949617584148-How-are-citations-classified)
- Research Solutions. (n.d.-b). *How to use the Scite Reference Check*. Research Solutions Help & Support Center. [https://help.researchsolutions.com/hc/en-us/articles/31949485592596-How-to-use-the-Scite-Reference-Check](https://help.researchsolutions.com/hc/en-us/articles/31949485592596-How-to-use-the-Scite-Reference-Check)
- Scite. (n.d.). *AI for research*. [https://scite.ai/](https://scite.ai/)
- Tay, A. (2025, November 15). *A 2025 deep dive of Consensus: Promises and pitfalls in AI-powered academic search* [Substack newsletter]. [https://aarontay.substack.com/p/a-2025-deep-dive-of-consensus-promises](https://aarontay.substack.com/p/a-2025-deep-dive-of-consensus-promises)
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. *Scientific Reports, 13*, Article 14045. [https://doi.org/10.1038/s41598-023-41032-5](https://doi.org/10.1038/s41598-023-41032-5)
