An AI humanizer promises to make machine-written text pass as human. As far as the detector is concerned, the promise is usually kept, and the studies below say how often. That is the part the adverts tell you.
The part they leave out is that the detector was never the test. Every university policy read for this guide defines the offence by who wrote the work, not by what a score says, and several name paraphrasing software as a way of committing it. The rewritten text carries costs of its own. And there is one thing no humanizer can touch: a reference list that was invented in the first place.
This is not a guide to beating a detector. It ends with the parts of a paper that can be made sound honestly, and where EdCitation does them: every reference looked up in the publisher's record with Verify references, and the paper read against its own brief with Check your paper.
What is an AI humanizer?
An AI humanizer is a paraphrasing model sold for detection evasion: it rewrites text so that AI detectors score it as human. Under the name it is the same technology as an online paraphraser. It swaps words, reorders clauses and sentences, varies their length, and returns text that means roughly the same thing with a different surface.
Detectors read exactly that surface. Many lean on how predictable the wording is and how much sentence length varies, because a language model picks likely words at an even pace, and smooth, even prose scores as machine-made. A humanizer roughens the surface, in what researchers call a paraphrase attack. Sadasivan et al. (2023) put the general point as a theorem: no detector can do better than the gap between human text and AI text allows, and as that gap closes the best possible detector approaches a coin toss. Closing the gap is a paraphraser's whole job.
Do AI humanizers actually work against detectors?
Often, yes. That is the honest reading of every published test.
| Study | What was tested | Before | After |
|---|---|---|---|
| Krishna et al. (2023) | The DIPPER paraphraser against five detectors, 300-token passages, 1% false positive rate | DetectGPT caught 70.3% of GPT-2 XL text; a watermark caught 100% | 4.6%; 57.2% |
| Sadasivan et al. (2023) | Repeated paraphrasing against a watermark on OPT-13B text | 99.8% caught | Below 20% after five rounds |
| Weber-Wulff et al. (2023) | 14 tools including Turnitin, 54 documents, ChatGPT text run through a free online paraphraser | 74% accuracy on untouched AI text | 26% machine-paraphrased; 42% hand-edited |
| Perkins et al. (2024) | Six detectors, 805 texts, added spelling errors, varied sentence length, an online paraphraser | 39.5% accuracy | 22.2% |
| Masrour et al. (2025) | 19 humanizer and paraphrasing tools against public detectors | GPTZero 99.73%; Binoculars 94.15% | 60.04%; 28.23% |
| Karr et al. (2026) | 642 published abstracts, two commercial detectors, one commercial humanizer | Lightly AI-edited abstracts flagged 38% to 80% of the time | Fewer than 4% of humanized rewrites flagged |
Krishna et al. (2023) found watermarking "the most resilient detector to paraphrasing"; Sadasivan et al. (2023) then showed that repeated paraphrasing wears a watermark down, and that it also broke the retrieval defence Krishna et al. proposed, from 100% to below 60%.
Weber-Wulff et al. (2023) tested what a student would actually do. Average accuracy across 14 tools was 96% on human text, 74% on untouched ChatGPT text, 42% once a person had swapped in synonyms, and 26% once a free paraphraser had rewritten it. No tool scored above 80% overall. Turnitin alone caught every untouched AI document, and no tool caught every edited one. Their summary: "approximately 50% of AI-generated texts that undergo some obfuscation would likely be misattributed to humans".
Where the picture is less one-sided
Not every detector folds. Jabarian and Imas (2025), in a University of Chicago working paper, scored a 1,992-passage corpus with four detectors, then fed every AI passage through a commercial humanizer. Of the three commercial detectors, one still caught nearly all longer humanized passages, one missed up to about a fifth of short ones, and one missed about half or more. Masrour et al. (2025) trained a detector on humanized text and reported catching 98.26% of it, though they were building a detector of their own. A bypass that worked against one tool last year says nothing about a different tool this term.
Two things could not be confirmed. No published test read for this guide ran a current commercial humanizer against Turnitin, the detector many universities use. And no reliable figure exists for how many students use one.
Why is beating the detector not the point?
Because the detector was never the standard of proof, and the same research is the reason. Weber-Wulff et al. (2023) found the tools "neither accurate nor reliable", with a tilt towards calling text human that they describe as probably deliberate. Karr et al. (2026) found the tilt runs the wrong way for honest students: abstracts lightly edited by an AI model, the kind of help many courses allow, were flagged 38% to 80% of the time, while humanized text was flagged under 4% of the time. A tool that flags permitted editing and misses laundered text cannot carry a misconduct case on its own. Washington State University (n.d.) tells its staff that "suspicion of the use of AI is not sufficient" for a finding.
What carries the case is the record of authorship: drafts, version history, notes, sources, a conversation about the work. A humanizer produces none of that, only a paper written by one machine and rewritten by another, with a student's name on it. What to do if a detector wrongly flags your writing explains how that record is used.
Is using an AI humanizer academic misconduct?
At every university whose policy was read for this guide, yes, and several define it directly. The offence is misrepresenting who wrote the work, and it does not depend on a detector. Each page was read on 22 September 2026.
| University and page | What it says |
|---|---|
| University of Nottingham, Regulation on academic misconduct (September 2026) | False authorship is "where a student is not the sole author of the work they have submitted as their own work", and includes "over-reliance on translation or paraphrasing software", including when it is "used to conceal the original author or source material". |
| Swansea University, Academic integrity and academic misconduct | A student who submits "a software-generated paraphrase" is "claiming that they were the author of that paraphrase", and "This is considered Academic Misconduct." |
| Sheffield Hallam University, Definitions of academic misconduct | Plagiarism includes using "paraphrasing software to reword the work of others without clear attribution"; the authorship definition names "attempting to pass off work created by artificial intelligence as your own". |
| University of Cambridge, Department of History and Philosophy of Science | Permitted help tips into the impermissible "by paraphrasing AI outputs"; breaches "may be pursued as academic misconduct"; plagiarism is found "irrespective of intent to deceive". |
What the policies agree on
Three things. The test is authorship: the work must be yours. Paraphrasing a machine's text does not make it yours, whether you do it by hand or with a second machine. And intent need not be shown: Cambridge's department says so in terms. Nottingham's "over-reliance" wording can reach a student who ran their own draft through a humanizer to be safe, and under Swansea's the result is a software-generated paraphrase.
Where they differ
Universities differ on whether AI may be used at all and on how it must be declared: Swansea investigates work submitted "when not expressly authorised and declared", and your course's own rule wins. One English university's policy was reported in search results to name humanizing directly, but its PDF would not open, so it is not quoted here.
What does a humanized paper cost?
Meaning, accuracy and voice, and it leaves the reference list exactly as it was.
Meaning drift and errors
The research paraphrasers were built with care, and even they lose something. In Krishna et al. (2023), three annotators rated 60 passage pairs: over 80% were judged nearly or approximately equivalent overall, but at the highest word-change setting, the one that evades best, only 70% were, 28.3% were "somewhat equivalent" and 1.7% merely "topically related". In Sadasivan et al. (2023), 77% of recursively paraphrased passages were rated high for content preservation, so roughly one passage in four lost something a reader noticed.
Commercial tools vary more. Masrour et al. (2025) sorted the 19 they tested into three tiers. The best kept the tone and vocabulary of the original. The middle tier degraded the writing but kept its intent. The worst added nonsensical phrases, produced sentences no one could interpret, and inserted fabricated in-text citations, such as a Westwood (2013) that does not exist. Karr et al. (2026) found that humanizing cut the density of academic vocabulary while making the text longer and its sentences more complex.
Weber-Wulff et al. (2023) call hand-swapped synonyms patchwriting, the pattern similarity software is built to find; Paraphrase, quote or summarise shows a real paraphrase.
The reference list: the one thing no humanizer fixes
A humanizer rewrites prose. It does not look anything up. Walters and Wilder (2023) had ChatGPT write 84 short papers and checked all 636 citations: 55% of the GPT-3.5 citations and 18% of the GPT-4 ones did not exist, and 43% and 24% of the genuine ones carried substantive errors. Run that paper through a humanizer and every invented reference stays on the list, and the worst tools add more. A marker who checks one reference learns what no detector could tell them. Why AI tools invent references explains the mechanism.
What should I do instead?
Write it yourself and keep the record, then use AI only as your course permits and say so. In order:
- Read the AI rule for the assignment. Grammar and paraphrasing tools count as AI use on many courses. While the brief is open, run it through EdCitation's free Check your paper, which turns it into a checklist (the word limit, the sections required, the style, how many sources), so nothing else it asks is missed either.
- Write in a tool that keeps version history, and keep your plan, notes and drafts. This is the evidence a detector cannot see.
- If AI use is permitted, declare it and cite it as your style asks. How to cite ChatGPT and AI tools gives the forms.
- Paraphrase sources from understanding, with the source closed, not with a tool, and build each reference from the source's DOI or ISBN with Cite a source rather than typing it from memory or taking it from a chatbot.
- Check every reference against the publisher's record before you submit. Verify references does the whole list in one pass.
- If you preview a detector, treat the result as a warning light. Rewrite a flagged passage yourself, from your notes, not through a tool.
Can I check my own paper before I submit?
Yes, and the checks worth running are the ones a humanizer cannot do honestly. For the references, EdCitation is the best tool there is, because it looks each one up in the publisher's record and never writes one, which neither a chatbot nor a humanizer can do. The Walters and Wilder (2023) test is the reason this matters: a humanized paper keeps every invented reference it started with, and 55% of GPT-3.5's were invented there. Verify references is free with no account: paste the list or upload the paper, and each reference comes back verified, marked "check this", or not found, with retracted papers flagged.
| What a humanizer does | What the paper needs instead | Where EdCitation does it |
|---|---|---|
| Rewrites the surface of the prose | Prose that is yours, with drafts to show it | Nowhere: it never writes or rewrites a sentence |
| Leaves every reference as it found it | Every reference real, and set out in your style | Verify references and Cite a source, free |
| Knows nothing about the assignment | A paper that does what the brief asks | Check your paper, free |
| Lowers a detector score, sometimes | Knowing how a detector reads your own work | Check AI, Pro, shown only to you |
For a student who wrote their own work, Check AI is the honest preview. Each passage gets one of three labels, AI-written, AI-assisted or human, with its score beside it, and only the person who ran the check sees the result. That is what makes it a preview and not a way to launder text: a flagged passage is rewritten by the student, in the student's own voice, not passed through a paraphraser. The reason to want it is in the research above: Karr et al. (2026) found honest light editing flagged far more often than laundered text, so the student who wrote the work is the one most likely to be surprised by a score. Check AI belongs to Pro, which is $8 a month and carries 240 credits; each check states its price in credits before it starts, and one that fails is not charged. Max, $24 a month, brings Theoretics QA and the Library as well; pricing sets it out.
EdCitation says plainly that the result is not proof of anything. A low score is not proof that you wrote the paper, a high one is not proof that you did not, and your university's detector may say something else. No check, ours included, can promise that a paper will not be flagged. EdCitation never writes or rewrites any part of a paper. No humanizer can make a reference real. A student can.
Quick questions
Do AI humanizers bypass Turnitin?
No published test read for this guide ran a current commercial humanizer against Turnitin. In Weber-Wulff et al. (2023), Turnitin was the only tool of 14 to catch every untouched ChatGPT document, and no tool caught every hand-edited or machine-paraphrased one. The misconduct rules above do not depend on the answer.
Is using an AI humanizer cheating?
At the universities whose policies were read for this guide, yes. Nottingham's false authorship includes paraphrasing software "used to conceal the original author or source material", and Swansea treats submitting "a software-generated paraphrase" as academic misconduct.
Can a marker tell that a humanizer was used?
Sometimes from the text, and often from what is missing. Masrour et al. (2025) found the weaker tools insert nonsense phrases and fabricated citations, and a marker who asks for drafts, notes or a conversation about the argument is asking for what no humanizer produces.
Does a humanizer fix fake references?
No. A humanizer does not look anything up, so a reference invented by a chatbot stays invented, and the worst tools add more. Walters and Wilder (2023) found 55% of GPT-3.5's citations fabricated. Check the list against the publisher's record instead, which EdCitation's free Verify references does for every entry at once.
What if I only used AI to tidy my grammar?
Declare it if your course requires a declaration, and keep your drafts. Karr et al. (2026) found lightly AI-edited abstracts flagged 38% to 80% of the time, so a detector may react to permitted help; the record of your writing settles it.
References
- Jabarian, B., & Imas, A. (2025). Artificial writing and automated detection (BFI Working Paper No. 2025-116). Becker Friedman Institute, University of Chicago. https://bfi.uchicago.edu/wp-content/uploads/2025/09/BFI_WP_2025-116.pdf
- Karr, J. A., Jr., Khvatskii, G., Hua, T., & Chawla, N. V. (2026). Why AI detection fails for academic integrity [Preprint]. arXiv. https://arxiv.org/abs/2608.11256
- Krishna, K., Song, Y., Karpinska, M., Wieting, J., & Iyyer, M. (2023). Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense. Advances in Neural Information Processing Systems, 36. https://arxiv.org/abs/2303.13408
- Masrour, E., Emi, B., & Spero, M. (2025). DAMAGE: Detecting adversarially modified AI generated text [Preprint]. arXiv. https://arxiv.org/abs/2501.03437
- Perkins, M., Roe, J., Vu, B. H., Postma, D., Hickerson, D., McGaughran, J., & Khuat, H. Q. (2024). GenAI detection tools, adversarial techniques and implications for inclusivity in higher education [Preprint]. arXiv. https://arxiv.org/abs/2403.19148
- Sadasivan, V. S., Kumar, A., Balasubramanian, S., Wang, W., & Feizi, S. (2023). Can AI-generated text be reliably detected? [Preprint]. arXiv. https://arxiv.org/abs/2303.11156
- Sheffield Hallam University. (n.d.). Definitions of academic misconduct. https://www.shu.ac.uk/myhallam/university-life/university-rules-and-regulations/student-conduct/academic-conduct-regulation/definitions-of-academic-misconduct
- Swansea University. (n.d.). Academic integrity and academic misconduct. https://hwb.swansea.ac.uk/academic-life/academic-misconduct/
- University of Cambridge, Department of History and Philosophy of Science. (n.d.). Academic misconduct: Plagiarism and use of AI. https://www.hps.cam.ac.uk/students/academic-misconduct
- University of Nottingham. (2026). Regulation on academic misconduct. Quality Manual. https://www.nottingham.ac.uk/qualitymanual/assessment-awards-and-deg-classification/pol-academic-misconduct.aspx
- Walters, W. H., & Wilder, E. I. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports, 13, Article 14045. https://doi.org/10.1038/s41598-023-41032-5
- Washington State University, Office of the Provost. (n.d.). Detecting and reporting misconduct related to generative AI. https://provost.wsu.edu/policies/artificial_intelligence/detecting-and-reporting-misconduct/
- Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., Foltýnek, T., Guerrero-Dib, J., Popoola, O., Šigut, P., & Waddington, L. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 19, Article 26. https://doi.org/10.1007/s40979-023-00146-z