A similarity score arrives with no explanation: 23%, a colour, a list of sources, and a marker who may or may not have read any of it.
The number is easier to understand once you know how it was made. A plagiarism checker is a text-matching engine. It looks for sequences of words that also appear in its database and reports how much of your paper they cover. That is all it measures, and every company that sells one says so on its own pages.
This guide explains the matching, the databases, what the percentage counts and misses, and why no score is a finding of plagiarism. EdCitation publishes it and sells a plagiarism check of its own, Check plagiarism; every product fact here was read on the company's own site on 22 September 2026. The usual fix for a match is credit rather than new wording, and EdCitation's free Cite a source builds that credit from the publisher's record.
How does a plagiarism checker find matching text?
It cuts the paper into short phrases, gives each one a fingerprint, and looks for the same fingerprints in an index of other texts. Turnitin describes its own process on its site: words are grouped into phrases, common words such as "and" and "the" are removed, each phrase is stored under a fingerprint ID, a document can hold up to 80,000 of them, and each is compared with "7 trillion possible phrase matches" in the content databases (Hanson, 2021). Turnitin (2016) contrasts this with matching fixed strings of three or four words, which finds exact copying but not paraphrase.
Nothing in that process reads for meaning. The engine does not know what a sentence says, who wrote it, or whether it sits inside quotation marks; it knows that a phrase in your file also exists in a file it has indexed. Weber-Wulff (2019) calls the products "black box" algorithms that produce a score. Researchers therefore say text-matching software, and the vendors agree: Foltýnek et al. (2020) open their test of 15 systems by stating that software "cannot determine plagiarism", and Turnitin's own FAQ says its Similarity product "does not determine whether plagiarism has occurred" (Turnitin, n.d.-a).
The steps, in order
- The text is extracted from the file. Turnitin's report now carries a flags panel for replaced or hidden characters, the known ways of disturbing this step (Turnitin, n.d.-a).
- The text is cut into short phrases with common words removed, and each phrase gets a fingerprint ID.
- Every fingerprint is looked up in the index.
- Runs of matched phrases are grouped into passages, each tied to one primary source (Hanson, 2021).
- Matched words are divided by total words. That is the similarity score.
What is in the database a checker compares against?
Three kinds of text: crawled web pages, licensed publications, and papers previously submitted to that checker. The figures are the companies' own, read on 22 September 2026.
| Checker | What its own pages say it compares against |
|---|---|
| Turnitin | More than 1.2 billion student submissions, 70 billion current and archived web pages, and 180 million journal articles (Hanson, 2021); the product page now says "20+ years of internet content" and publisher content in 170 languages (Turnitin, n.d.-a) |
| Crossref Similarity Check | Turnitin's iThenticate at a reduced rate for Crossref members who deposit full-text links, so their own articles join the index; more than 78 million full-text scholarly items plus the web (Crossref, n.d.-a) |
| Grammarly | More than 16 billion web pages, plus academic papers in ProQuest's databases (Grammarly, n.d.-a) |
| EdCitation Check plagiarism | Published and web text, with the source of each match shown; not papers other students have submitted |
A student paper is compared with other student papers, so a friend's essay from last year sits in the same index as a journal. Turnitin's student guide describes the case where a copy of your paper is submitted first and scores low, and your original then scores 100% (Turnitin, n.d.-b).
What no database holds
Anything the checker has not indexed, which is more than students expect. Foltýnek et al. (2020) built their test from openly available sources only, and even so a Slovak paper in an open-access journal was found by none of the 15 systems. Print-only books, paywalled papers the vendor has not licensed, and work submitted to a different checker are all invisible. A tool that finds no source has not shown the text is original: "the text can still be plagiarized" (Foltýnek et al., 2020).
What does the similarity score actually count?
The share of words in the paper that fall inside a matched passage. Turnitin's student guide defines the score as the percentage of matched text, and Crossref's documentation for iThenticate gives the arithmetic: matching words divided by the document's total words (Crossref, n.d.-b). Turnitin colours the result: blue for no matching text, green from one word to 24%, yellow to 49%, orange to 74%, red to 100% (Turnitin, n.d.-b). The bands carry no verdict: the guide says there is no fixed number, and every school, instructor or assignment can set its own expectation.
Why a quotation, a reference list and a standard phrase all match
Because each is a sequence of words that also exists somewhere else, which is the only thing the engine tests. A quotation is a match by definition, and Turnitin's student guide says that quotations and references used correctly are still highlighted (Turnitin, n.d.-b). A reference list matches because two papers citing the same source in the same style produce identical entries; Foltýnek et al. (2020) found that in their translated test documents the reference list was the only text most systems matched. Standard phrases match for the same reason. Elsevier's guide for its editors warns that a high percentage "does not necessarily indicate plagiarized text" and names legitimate citation and commonly used phrasing as ordinary causes (Elsevier, n.d.).
Checkers therefore offer exclusions: a Turnitin instructor can drop quotations, bibliographies and small matches from the score (Hanson, 2021), so a student's percentage and a marker's can differ.
Why a low score is not clean and a high score is not guilt
A score measures overlap with one index on one day. Turnitin's student guide lists the cases: a long paper can show 0% after rounding while the report still holds matches; a final draft can show 100% because an earlier draft went into the same repository (Turnitin, n.d.-b). Crossref tells editors not to set a score "over which you automatically reject manuscripts" (Crossref, n.d.-b). If a number cannot decide a journal submission, it cannot decide an essay.
Why do paraphrase and translation slip past a checker?
Because the engine matches sequences of words, and a paraphrase or a translation keeps the idea while changing the sequence. Turnitin (2016) says its fingerprinting can find text that is "poorly paraphrased". The published tests measure how far that reaches.
Foltýnek et al. (2020) is the largest independent test: documents in eight languages submitted to 15 systems, each source copied three ways, pasted, disguised by manual synonym replacement, and manually paraphrased, plus a translation of the English Wikipedia article on plagiarism detection, half by Google Translate and half by hand. Each result was scored from 0, a sentence or less found, to 5, all or almost all found.
| What the test document did | What the 15 systems found |
|---|---|
| Copy and paste | Acceptable results from every system except DPV, intihal.net and Dupli Checker |
| Manual synonym replacement | A considerable drop in every system except Urkund, PlagiarismCheck.org and Turnitin |
| Manual paraphrase | No system satisfactory; PlagiarismCheck.org, Urkund, PlagScan and Turnitin somewhat better than the rest |
| Translation | No system except Akademia, whose database is small; the others matched mainly the reference lists |
| Original, unpublished document | Scored in reverse for false positives; systems "at times also identify non-plagiarized material as problematic" |
No system reached the "useful" band of 3.75 to 5. Five sat in the "partially useful" band from 2.5 to 3.75, Turnitin and Urkund among them with PlagAware, PlagScan and StrikePlagiarism.com; seven were "marginally useful" and three were judged unsuited to academic institutions (Foltýnek et al., 2020). That is the gap between the marketing and the test.
What paraphrasing tools and machine translation do to the score
They can take it to zero. Prentice and Kinden (2018) traced the strange, unidiomatic language in first-year health science essays to free online paraphrasing tools, which had swapped standard medical terms for synonyms. Some of the suspicious essays produced Turnitin matches; others "resulted in an index of 0%". The give-away was vocabulary: 73 substitutes for 21 standard medical terms.
Translation is worse. Dilber and Yoşumaz (2026) translated ten English manuscripts into Spanish, Portuguese and French and ran them through Turnitin, iThenticate and Grammarly. Their abstract reports that the tools were largely ineffective on the translated texts, that plagiarism by translation can become "virtually undetectable by software", and that Grammarly showed limited detection even on the English originals. The full text is paywalled, so the figure for each tool could not be read for this article.
A paraphrasing tool lowers the score and leaves the copying in place. The rules for a paraphrase that is actually yours are in When to paraphrase, quote or summarise. EdCitation offers no paraphraser and never rewrites a sentence, because the repair is credit: when an idea is in your notes but its origin is not, Find sources searches about 300 million published works for the claim itself, so the original can be found and cited.
Why is a similarity score not a finding of plagiarism?
Because plagiarism is about attribution and intent, and a text match is evidence of neither. Turnitin (2016) states that no software can determine intent. Crossref's documentation puts it in one line: iThenticate "does not check for plagiarism; it checks for similarity" (Crossref, n.d.-b).
The errors run both ways. Weber-Wulff (2016), summarising the tests she ran between 2004 and 2013, reports that systems can flag correctly referenced material as unoriginal and can report no problems at all for heavily plagiarised texts. Her comment in Nature argues that because the systems do catch some plagiarism, users have come to believe they document all of it, and asks academics and editors to read more carefully instead (Weber-Wulff, 2019). Foltýnek et al. (2020) add that the reports serve as evidence in disciplinary cases, where a human reading is always needed.
How to read a similarity report, match by match
The percentage is the least useful line on it. Read the matches in this order:
- Open each match and look at what is highlighted, not how much.
- A quotation: check the quotation marks are in place and the citation carries a page number. If its entry is missing from the list, EdCitation's Cite a source builds it from a DOI, title or ISBN.
- A reference list entry, a short standard phrase, a method description: expected. Move on, but note that a matched entry proves nothing about whether its source is real; Verify references answers that separately.
- A run of sentences outside quotation marks matched to one source, with your synonyms swapped in: patchwriting. Rewrite it from understanding or quote it properly, and cite it either way.
The similarity report explained this way is a list of places to check, and its meaning is set by what you find at each one.
How can I see the matches before a marker does?
Check the finished draft yourself and go through the matches with that list in hand. Check plagiarism is EdCitation's plagiarism check, and it is the best tool for this job because it hands the writer, before anyone else reads the paper, the part of a report this guide says matters: each matching passage, marked in place in your own text, with the source it matches. A forgotten quotation mark or a paraphrase that stayed too close can then be found by you rather than by the person grading it. It cannot predict the percentage a marker's checker will give, since that index holds other students' papers and the marker sets the exclusions. The report stays private to you. It runs on credits (Pro, $8 a month, comes with 240; Max, $24, with 720), with the cost shown before a check runs and nothing used if it fails. See pricing.
Turnitin is sold to institutions, and its student guide says the score is not shown when the instructor has turned off student access to the report (Turnitin, n.d.-b). Grammarly's plans page lists plagiarism detection under Pro at $12 a month and not under Free (Grammarly, n.d.-b), and the one published test that included it found its detection limited even on English originals (Dilber & Yoşumaz, 2026). What Turnitin does that EdCitation does not is compare a paper with the papers other students handed in.
Two limits, stated plainly. A match is not proof of plagiarism and a clean report is not proof of originality, and EdCitation says so. EdCitation never writes or rewrites any part of a paper: the judgement about each match stays with the writer. The repairs are free and need no account: Cite a source looks a matched work up in the publisher's record and builds its reference, where a chatbot would generate one, and Verify references checks the finished list, entry by entry, against the record.
Quick questions
What percentage on a plagiarism checker is acceptable?
There is no universal figure. Turnitin's student guide says there is no fixed number and that every school, instructor or assignment may set its own expectation. Ask the marker, and read the matches rather than the percentage.
Does the similarity score include quotations and references?
Yes, unless they are excluded. Turnitin's student guide says correctly used quotations and references are still highlighted, and an instructor can exclude quotations, bibliographies and small matches, so the number a marker sees may differ from yours.
Can a plagiarism checker detect paraphrasing?
Only when the paraphrase stays close to the source. Foltýnek et al. (2020) tested 15 systems and found that almost none could satisfactorily identify manual paraphrase, though Turnitin, Urkund, PlagScan and PlagiarismCheck.org did better than the rest.
Can a plagiarism checker detect translated text?
Almost never. In the Foltýnek et al. (2020) test no system except Akademia found the translated documents, and Dilber and Yoşumaz (2026) reached the same conclusion for Turnitin, iThenticate and Grammarly.
Does a plagiarism checker show whether my references are real?
No. A checker matches words, so an entry copied from someone else's list matches whether or not its source exists. EdCitation's Verify references looks each entry up in the publisher's record, free with no account, and never reports one it could not check as not found.
References
- Crossref. (n.d.-a). Similarity Check. Retrieved September 22, 2026, from https://www.crossref.org/services/similarity-check/
- Crossref. (n.d.-b). Understanding your Similarity Report. Retrieved September 22, 2026, from https://www.crossref.org/documentation/similarity-check/similarity-report-understand/
- Dilber, C., & Yoşumaz, İ. (2026). The impact of language translation on plagiarism rates: Evidence from Turnitin, iThenticate, and Grammarly. Journal of Academic Ethics, 24(1), Article 4. https://doi.org/10.1007/s10805-025-09681-5
- Elsevier. (n.d.). Editor guide to CrossCheck plagiarism reports. Retrieved September 22, 2026, from https://www.elsevier.support/publishing/answer/editor-guide-to-crosscheck-plagiarism-reports
- Foltýnek, T., Dlabolová, D., Anohina-Naumeca, A., Razı, S., Kravjar, J., Kamzola, L., Guerrero-Dib, J., Çelik, Ö., & Weber-Wulff, D. (2020). Testing of support tools for plagiarism detection. International Journal of Educational Technology in Higher Education, 17, Article 46. https://doi.org/10.1186/s41239-020-00192-4
- Grammarly. (n.d.-a). Plagiarism checker. Retrieved September 22, 2026, from https://www.grammarly.com/plagiarism-checker
- Grammarly. (n.d.-b). Grammarly prices and plans. Retrieved September 22, 2026, from https://www.grammarly.com/plans
- Hanson, G. (2021, November 15). How to bring deeper meaning to the Similarity Score. Turnitin. https://www.turnitin.com/blog/how-to-bring-deeper-meaning-to-the-similarity-score
- Prentice, F. M., & Kinden, C. E. (2018). Paraphrasing tools, language translation tools and plagiarism: An exploratory study. International Journal for Educational Integrity, 14, Article 11. https://doi.org/10.1007/s40979-018-0036-7
- Turnitin. (n.d.-a). Turnitin Similarity. Retrieved September 22, 2026, from https://www.turnitin.com/products/similarity/
- Turnitin. (n.d.-b). Understanding the similarity score for students. Turnitin Guides. Retrieved September 22, 2026, from https://guides.turnitin.com/hc/en-us/articles/23713493434253-Understanding-the-similarity-score-for-students
- Turnitin. (2016, July 21). The detection is in the details. https://www.turnitin.com/blog/the-detection-is-in-the-details
- Weber-Wulff, D. (2016). Plagiarism detection software: Promises, pitfalls, and practices. In T. Bretag (Ed.), Handbook of academic integrity (pp. 625–638). Springer. https://doi.org/10.1007/978-981-287-098-8_19
- Weber-Wulff, D. (2019). Plagiarism detectors are a crutch, and a problem. Nature, 567(7749), 435. https://doi.org/10.1038/d41586-019-00893-5