The words AI overuses are real and measurable, but only in bulk. Studies published since 2024 have counted words across millions of abstracts, papers and peer reviews, and found that "delves", "underscores", "showcasing", "intricate", "meticulous" and "commendable" became far more common after ChatGPT was released. What they measured is word frequency across a whole body of text. What none of them measured, and what their authors say their methods cannot do, is tell whether any one essay or paper was written with AI.
EdCitation publishes this guide. Every figure below comes from the paper itself, read on 24 September 2026, with its method and limits; preprints are marked. Where the guide gives a view rather than a finding, it says so. Each paper can be found in EdCitation's free Find sources, a search of about 300 million published works that takes a topic, or the claim a sentence has to rest on.
Our piece on the em dash makes one point in full, and it is not argued again here: a writer who has always used such words is not shown by them to be using AI. This guide is about the studies themselves.
Which words does AI overuse, according to the studies?
A family of style words, not a single tell: verbs such as "delves", "underscores" and "showcasing", adjectives such as "intricate", "meticulous", "commendable" and "pivotal", and linking words such as "additionally" and "notably". Nouns that name a topic barely feature. Nine studies, what each counted, and the limit each states:
| Study | What was counted | Headline finding | What it cannot show |
|---|---|---|---|
| Gray (2024), preprint | Full text of articles in the Dimensions database, to 2023 | 60,000 to 85,000 articles in excess in 2023 | Which papers used a model |
| Liang et al. (2024) | Peer reviews for four AI conferences | Between 6.5% and 16.9% of review text substantially modified by a model | That any review was written from scratch |
| Geng and Trotta (2024), preprint | One million arXiv abstracts to January 2024 | About 35% in computer science, relative to one simple revision prompt | How many people used a model |
| Kobak et al. (2025) | 15.1 million PubMed abstracts, 2010 to 2024 | A floor of 13.5% of 2024 abstracts processed with a model | Which abstracts those were |
| Liang, Zhang, Wu, et al. (2025) | 1,121,912 papers and preprints to September 2024 | Up to 22% in computer science, up to 9% in mathematics | Any single paper |
| Liang, Zhang, Codreanu, et al. (2025) | Complaints, press releases, job posts, UN releases | Just under 10% to 24% of text by late 2024 | Heavily edited text, so each figure is a floor |
| Juzek and Ward (2025) | 26.7 million PubMed abstracts to May 2024 | 21 "focal words" whose rise is likely due to model use | Why models favour them |
| Geng and Trotta (2025) | 1,294,653 arXiv abstracts, 2018 to 2024 | "delve" fell from about April 2024 | Cause: it reports correlations |
| Gray (2025), preprint | Full text of Dimensions articles, to 2024 | More than 10% of 2024 papers | That any given paper was assisted |
How did researchers find the words AI overuses?
By comparing how often words appeared before and after ChatGPT's release on 30 November 2022, in collections too large for chance to explain the change. Three methods recur, and each shapes what its figure means.
Excess words against a projected baseline
Kobak et al. (2025) borrowed the idea of excess deaths from studies of the pandemic. For each word common enough to measure in 15.1 million English-language PubMed abstracts, they took the share of abstracts containing it each year, projected a 2024 value from 2021 and 2022, and compared that projection with what 2024 showed. "Delves" turned up 28 times as often as projected; "potential" appeared in 5.2 percentage points more abstracts than expected. No list of suspect words went in first: the counting found them.
Marker words and control words
Gray (2024) took twelve adjectives and twelve adverbs that Liang et al. (2024) had found overrepresented in model-written reviews, and counted the articles containing each, beside control words such as "blue", "red" and "later" that no model should favour. Most controls moved less than 5% in a year. In 2023, "intricate" rose 117% and "meticulously" 137%; "outwith", a Scottish English word tested separately, rose 185%.
Mixture models trained on known AI text
Liang et al. (2024) had a model write peer reviews of real papers, then estimated what share of real reviews best fitted a mix of human and model vocabulary. On older reviews mixed with known amounts of model text, the estimate stayed within 2.4 percentage points of the truth, and reviews a model had only proofread barely moved it. The figures describe text substantially changed by a model, not spell-checked text.
What did the studies find in abstracts, papers and reviews?
That the change arrived within about a year of ChatGPT and differed widely by field, country and venue.
Biomedical abstracts
Kobak et al. (2025) counted 454 excess words in 2024, against 190 in 2021 at the height of the pandemic's effect. The pandemic's excess words were almost all content words such as "lockdown"; of 2024's 379 excess style words, 66% were verbs and 14% adjectives. Two separate sets of marker words gave 13.6% and 13.4%, averaged to a floor of 13.5% of 2024 abstracts, at least 200,000 papers a year. The floor was about 5% for the United Kingdom and Australia, about 20% for China, South Korea and Taiwan, and 7% for Nature, Science and Cell together.
The authors add a caution before anyone reads that as a judgement on a country's writers: native English speakers may use models as often and simply be better at deleting the telltale words, which a word count cannot see. The related problem with detectors is in our guide on non-native English writers.
Peer reviews
Liang et al. (2024) estimated that the share of review sentences substantially modified by a model rose from 1.6% to 10.6% at ICLR, from 1.9% to 9.1% at NeurIPS and from 2.4% to 6.5% at CoRL, and stood at about 16.9% at EMNLP 2023. Reviews for 15 Nature portfolio journals showed no significant rise. In ICLR 2024 reviews, "meticulous" was 34.7 times as likely to occur in a sentence as before.
Papers across the sciences
Gray (2025) estimated that more than 10% of papers published in 2024 show signs of model involvement. "Additionally", in 13.8% of papers in 2016, never grew more than 9% a year up to 2022, then grew 20% and 39% in the next two. Geng and Trotta (2024) found "is" and "are" down by more than 10% in 2023 arXiv abstracts, a shift invisible in any one text.
What did the studies not measure?
Whether any single text was written by AI. Every study here works on a body of text, and none studied student essays: the collections are abstracts, papers, peer reviews, complaints, press releases and job posts.
A floor for a collection, not a count of writers
Kobak et al. (2025) state that their analysis works on the whole corpus and cannot identify individual abstracts. Their 13.5% is a floor, because an abstract that went through a model without picking up a marker word is not counted. Liang, Zhang, Codreanu, et al. (2025) call their figures a lower bound, since heavily edited model text escapes the method. Liang et al. (2024) say their estimate is not evidence of reviews written from scratch: a reviewer who had a model expand their own bullet points would leave the same signal.
Why a word cannot point at one writer
Gray (2024) says the markers do not show that any particular paper was written with a model, and that their absence does not show one was not used. Gray (2025) puts it most simply: "sometimes, humans simply do write like that." Kobak et al. (2025) add that a rise caused by a model and a rise caused by people adopting the model's favourite words look the same in their counts. Each rate is a share of documents containing a word at least once, so a writer who has written "crucial" for twenty years adds to the same count as a model does.
| Question | Can the vocabulary studies answer it? | What does answer it |
|---|---|---|
| How much more often did "delves" appear in 2024 abstracts? | Yes: 28 times the projection (Kobak et al., 2025) | The count itself |
| Roughly what share of a collection was shaped by a model? | Yes, as an estimate or a floor | The corpus studies above |
| Was this essay written with AI? | No, and the authors say so | Nothing in these papers; the writer's drafts and notes are the record |
| Has this student always written "delve"? | No | Their writing from before 2022 |
| Does each reference in the essay exist? | No | EdCitation's free Verify references, from the publisher's record |
Why does AI overuse words like "delve"?
Nobody outside the companies that build the models knows for certain. Juzek and Ward (2025) searched 26.7 million PubMed abstracts for words that rose sharply between 2020 and 2024 with no scientific or news explanation, then kept those that ChatGPT-3.5 also overused when it wrote 9,953 abstracts from summaries of real ones. That left 21 focal words.
What the tests ruled out, and what they did not
The focal words were far rarer in arXiv abstracts, news text and Wikipedia than in the model's abstracts, which counts against simple copying of training data, and a corpus of the world's Englishes showed no variety that favours them. Comparing Meta's Llama 2 before and after its chat tuning with human feedback, the tuned model was markedly less surprised by AI-written abstracts, consistent with human raters' preferences playing a part. The direct test was inconclusive: 201 participants recruited in India showed no significant overall preference between abstracts with and without focal words, though they tended to reject versions opening with "delves".
Are the words AI overuses changing?
Yes, quickly, which is one more reason a word list settles nothing about a person. Geng and Trotta (2025) tracked 1,294,653 arXiv abstracts from 2018 to 2024. "Delve" and "intricate" fell from about April 2024, soon after studies had named them, while "significant" and "additionally", which drew less attention, kept rising. The authors read this as writers still using models while editing out the notorious words.
Gray (2025) therefore built the 2024 estimate from markers that had not fallen, with "red", "blue" and "yellow" as controls: a surplus of 0.6% of papers for the colours, and just over 12% for any two of the eight strongest markers. In our view, the fall of "delve" bears out a warning Gray (2024) gave: a paper without the markers is not thereby free of model use.
How do you read a claim about AI words?
Go to the paper, not the headline, and check these things before repeating a figure:
- Find the study itself. EdCitation's Find sources found it for this guide: on 24 September 2026, "delve excess vocabulary PubMed abstracts LLM", sorted by most cited, listed Kobak et al. (2025) first, open access, with 144 citations. A broader phrasing brought back work on vocabulary learning, so search with the study's own distinctive words.
- Note what was counted. A figure for abstracts says nothing about reviews, press releases or essays.
- Read the figure's type. The 13.5% is a floor for abstracts; the 16.9% is an estimate for review sentences.
- Find the baseline. Every rise is measured against a projection or an earlier average, and control words show ordinary drift.
- Read the limits section, and quote it beside the number.
- Cite the version you read with Cite a source, and check the finished list with Verify references.
How can EdCitation help you work with these studies?
EdCitation is the best tool for turning these studies into references you can trust: every entry it builds is read from the publisher's record, and it never composes a reference, which no chatbot can promise. Cite a source is free with no account. Given the DOI of the Patterns paper (Liang, Zhang, Codreanu, et al., 2025) on 24 September 2026, it returned all six authors, the volume, issue and article number, and no retraction notice. Geng and Trotta's ACL paper came back with its pages and publisher, and its title, deposited in title case, set in APA's sentence case with the initialism LLM kept. Searched by its title, Juzek and Ward's paper now comes back with its 2024 arXiv preprint offered first; the COLING version cited here has no DOI, so that entry was built from the ACL Anthology page.
Verify references, free as well, takes the whole list at once. Its verdict on each entry is verified, "check this" or not found; it points out any paper its journal has withdrawn, and says plainly when a lookup failed rather than reporting the source as missing. Run on this guide's nine references, it verified all nine, checking the two without a DOI at their own pages.
Check AI looks at your own draft, not at the studies. Pro, at $8 a month, includes it with 240 credits: it sorts a draft's passages into AI-written, AI-assisted and human, with the score for each, and only you see the report. Its cost is shown first, and a failed check costs nothing. Treat the result as an estimate: a different detector, such as the one your university uses, may score the same passages otherwise, and a score on its own proves nothing. Reading a Check AI report walks through it, and Pricing shows the $24 Max plan, which brings Theoretics QA and the Library as well.
Quick questions
Does writing "delve" mean AI wrote my essay?
No. The studies that found "delves" rising measured it across millions of abstracts, and Kobak et al. (2025) state that their corpus-level analysis cannot pick out single abstracts; an essay is further still from what they measured.
What share of research papers is written with AI?
No study counts that directly. Estimates of text shaped by a model run from over 1% of 2023 articles (Gray, 2024) to a floor of 13.5% for 2024 PubMed abstracts and up to 22% in computer science, each for its own collection and method.
Should I take words like "crucial" out of my writing?
Choose words for your reader. Removing them proves nothing about authorship, and writers have been editing the best-known ones out since 2024 (Geng & Trotta, 2025). Keep drafts and notes, which show how the work was written.
How can I find and cite the studies on AI words?
Search their titles or distinctive words in EdCitation's Find sources, build each entry from its DOI in Cite a source, and check the list in Verify references, all free with no account.
References
- Geng, M., & Trotta, R. (2024). Is ChatGPT transforming academics' writing style? [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2404.08627
- Geng, M., & Trotta, R. (2025). Human-LLM coevolution: Evidence from academic writing. In Findings of the Association for Computational Linguistics: ACL 2025 (pp. 12689–12696). Association for Computational Linguistics. https://doi.org/10.18653/v1/2025.findings-acl.657
- Gray, A. (2024). ChatGPT "contamination": Estimating the prevalence of LLMs in the scholarly literature [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2403.16887
- Gray, A. (2025). Estimating the prevalence of LLM-assisted text in scholarly writing [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2512.01560
- Juzek, T. S., & Ward, Z. B. (2025). Why does ChatGPT "delve" so much? Exploring the sources of lexical overrepresentation in large language models. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 6397–6411). Association for Computational Linguistics. https://aclanthology.org/2025.coling-main.426/
- Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances, 11(27), Article eadt3813. https://doi.org/10.1126/sciadv.adt3813
- Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D. A., & Zou, J. Y. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. In Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 29575–29620). PMLR. https://proceedings.mlr.press/v235/liang24b.html
- Liang, W., Zhang, Y., Codreanu, M., Wang, J., Cao, H., & Zou, J. (2025). The widespread adoption of large language model-assisted writing across society. Patterns, 6(12), Article 101366. https://doi.org/10.1016/j.patter.2025.101366
- Liang, W., Zhang, Y., Wu, Z., Lepp, H., Ji, W., Zhao, X., Cao, H., Liu, S., He, S., Huang, Z., Yang, D., Potts, C., Manning, C. D., & Zou, J. (2025). Quantifying large language model usage in scientific papers. Nature Human Behaviour, 9(12), 2599–2609. https://doi.org/10.1038/s41562-025-02273-8