# How AI changed academic writing: what the evidence shows

**How much has AI changed academic writing since ChatGPT, according to the evidence?** How AI changed academic writing can be measured only across large collections of text. Studies estimate that at least 13.5% of 2024 PubMed abstracts were processed with a language model, up to 22% of sentences in computer science papers were modified by one, and so was 6.5% to 16.9% of review text at four AI conferences. None can judge a single paper.

Published 2026-09-24 by EdCitation. https://edcitation.com/newsletter/how-ai-changed-academic-writing-what-the-evidence-shows

How AI changed academic writing can be measured, but only across whole collections of text. Since ChatGPT's release on 30 November 2022, studies of millions of documents estimate that at least 13.5% of 2024 PubMed abstracts went through a language model, up to 22% of sentences in computer science papers were modified by one, and 6.5% to 16.9% of peer review text at four AI conferences was substantially changed. Grant proposals show the same rise. None of these studies can say whether one paper or essay used AI.

What followed can be looked up: chatbot phrases in published papers, retractions, fabricated references. EdCitation publishes this guide, with every figure read at its source on 24 September 2026; preprints, and papers we could not read in full, are marked. EdCitation's [Find sources](https://edcitation.com/) located most of the studies, and [Verify references](https://edcitation.com/verify-references), free with no account, checks the one harm here with a yes-or-no answer: whether a reference exists.

The words behind many estimates are in [words AI overuses](https://edcitation.com/newsletter/words-ai-overuses-what-the-studies-found), and why no tell proves anything about a writer in [the em dash is not proof of AI](https://edcitation.com/newsletter/the-em-dash-is-not-proof-of-ai).

## How much academic writing has AI changed since ChatGPT?

A measurable, growing share that varies by field, venue and kind of text. Each figure holds only for its own collection.

| Study | Text measured | Period | Estimate | Unit counted |
| --- | --- | --- | --- | --- |
| Kobak et al. (2025) | 15.1 million PubMed abstracts | 2010 to 2024 | At least 13.5% in 2024 | Abstracts processed with a model |
| Liang, Zhang, et al. (2025) | 1,121,912 papers on arXiv, bioRxiv and Nature portfolio journals | January 2020 to September 2024 | Up to 22% in computer science; up to 9% in mathematics and the Nature portfolio | Sentences modified by a model |
| Liang, Izzo, et al. (2024) | Reviews for ICLR 2024 and NeurIPS, CoRL and EMNLP 2023 | 2023 review rounds | 6.5% to 16.9% | Review text substantially modified |
| Russo et al. (2025) | Reviews for ICLR 2024 | One review round | At least 15.8% | Whole reviews |
| Qian et al. (2026) | About 5,700 NSF and NIH proposals at two US universities; about 131,000 awards | 2021 to 2025 | Flat until late 2022, rising sharply from 2023 | Sentences in grant abstracts |
| Kusumegi et al. (2025) | 2.1 million preprints on arXiv, bioRxiv and SSRN | January 2018 to June 2024 | Adopters' output up 36.2% to 59.8% (disputed) | Papers per author |

### By field and by journal

Kobak et al. (2025) broke their 13.5% floor down: about 20% for computation and bioinformatics, 25% for *Sensors*, 20% for *Cureus*, 21% across MDPI's journals and 20% across Frontiers', against 10% for Nature family journals and 7% for *Nature*, *Science* and *Cell* together. The authors choose their words: it is easily detected use that falls with prestige, which need not mean less use.

Liang, Zhang, et al. (2025) found higher estimates for first authors who post preprints often, for crowded research areas and for shorter papers; we could read only their abstract.

### Over time

The same team's preprint, with papers to February 2024, gave up to 17.5% for computer science and 6.3% for mathematics and the Nature portfolio (Liang, Zhang, et al., 2024); with data to September 2024, the published version gives 22% and 9%. Qian et al. (2026) found grant abstracts steady until late 2022 and climbing from 2023, after which they split into two groups, one near zero and one centred at roughly 10% to 15% of sentences.

## How were these estimates made?

Mostly by comparing word frequencies across a collection, not by reading papers one at a time. The usual method has a model produce versions of text written before ChatGPT, learns how often each word appears in each, then estimates what mix of the two best explains a later collection. Liang, Izzo, et al. (2024) built it for peer reviews, and Qian et al. (2026) applied it to grant abstracts. Kobak et al. (2025) counted excess words instead.

### Two studies that sorted documents one by one

Russo et al. (2025) classed each ICLR 2024 review with AI detectors and call their 15.8% a lower bound. Kusumegi et al. (2025) set a threshold on each abstract, and an author counted as an adopter from the first abstract above it, a choice now disputed in print.

### What the figures cannot tell you

- **Which paper.** Each study reports a share of a collection.
- **Anything about essays.** None looked at student coursework.
- **Heavy editing.** Kusumegi et al. (2025) say their method almost certainly misses authors who heavily edit model output, and Kobak's floor counts only abstracts that kept a marker word.
- **The whole document.** Qian et al. (2026) and Kusumegi et al. read abstracts only; Qian et al. call their measure one dimension of model use, not its full extent.

## Where do the estimates disagree?

On what they count, and in one case on whether a finding exists at all.

### Sentences or whole reviews

ICLR 2024 has two estimates: 10.6% of review sentences substantially modified (Liang, Izzo, et al., 2024) and at least 15.8% of reviews AI-assisted (Russo et al., 2025). They need not conflict: a review with a few polished sentences counts whole in the second and only in part in the first.

### Did AI make researchers more productive?

Kusumegi et al. (2025) reported that monthly output rose after an author's first detected use, by 36.2% on arXiv, 52.9% on bioRxiv and 59.8% on SSRN, and by up to 89.3% for authors the study judged, from names and affiliations, likely to write in a second language.

Renault et al. (2026) rebuilt the design and got the same pattern where no model could be involved: papers flagged at random, neutral words such as "theory" as the trigger, and the years 2020 to 2022, before ChatGPT. A busy month is more likely to contain a flagged paper, so first detection lands at productive moments. This does not show models have no effect, they add, only that the design cannot measure it. In our view the productivity figure should never be quoted without this comment.

### Polished prose and weaker papers

Kusumegi et al. (2025) also found that without detected model use, more complex writing went with a higher chance of publication, and with it, a lower chance; ICLR 2024 review scores showed the same reversal. Renault et al. (2026) say they examined only the productivity analyses.

## What changed in peer review and grant funding?

Model use reached the texts that judge research, too.

### Peer review

Liang, Izzo, et al. (2024) found higher estimates in reviews submitted within three days of the deadline, in reviews citing no scholarly work, from reviewers who did not answer authors' rebuttals, and where reviewers reported low confidence. Reviews for 15 Nature portfolio journals showed no significant rise. Russo et al. (2025) compared reviews of the same paper: when an AI-assisted and a human review disagreed, the AI-assisted one scored higher in 53.4% of pairs, and borderline submissions that received one were 4.9 percentage points more likely to be accepted.

### Grant proposals

Qian et al. (2026) found that proposals and awards with more model involvement sat closer to what each agency had recently funded. Moving from the 25th to the 75th percentile of involvement went with about 5 percentile points less distinctiveness for NSF awards and 4 for NIH awards. At NIH the same step went with a funding chance about 4 percentage points higher and about 5% more publications, though not more highly cited ones; at NSF, no such link. These are associations, not effects. The authors note that in July 2025 NIH said applications "substantially developed by AI" would not count as original work; NIH's notice would not load for us.

## Which papers were retracted or corrected for undisclosed AI text?

A small, documented number, usually after readers spotted chatbot phrases. Glynn (2025) collected 768 published documents with suspected undeclared AI use (633 journal articles, 107 conference papers, 28 book chapters), mainly by searching Google Scholar for phrases such as "as an AI language model". One in three contained a reply opening "Certainly, here", and 61 kept ChatGPT's "Regenerate response" button label.

Only 33 (4.3%) had been altered, two of them twice: 13 formal retractions, 10 formal corrections, 8 silent corrections and 4 silent retractions, with formal notices a median of 147 days after publication. Two corrections promised an AI declaration that, when Glynn checked in November 2025, had never been added. The dataset is a preprint.

### Three cases, from the notices

| Paper | What the notice says | Outcome |
| --- | --- | --- |
| *Physica Scripta*, 2023 | ChatGPT was used to write part of the paper without being declared; PubPeer commenters raised it; the authors disagreed | Retracted, September 2023 |
| *PLOS ONE*, 2024 | Editors could not verify 18 of 76 references, and 6 more appeared to contain errors; "regenerate response" appeared; ethics approval was dated after recruitment began | Retracted, April 2024 |
| *Surfaces and Interfaces*, 2024 | The introduction opened "Certainly, here is a possible introduction for your topic" | Retracted, according to Glynn (2025) |

The first two notices are IOP Publishing's ("Retraction: Exploring New Optical Solutions," 2023) and The PLOS ONE Editors' (2024); the third would not load for us. The *PLOS ONE* authors said their only AI tool was Grammarly.

## How common are fabricated references in published papers?

Rare in any one paper, but rising fast. Topaz et al. (2026) summarise their audit of 2.47 million papers in PubMed Central's open-access subset, January 2023 to February 2026: 4,046 fabricated references in 2,810 papers, the rate up more than twelvefold to 56.9 per 10,000 papers in early 2026. The fakes were on topic, well formatted and credited to real researchers. The audit is in *The Lancet*, which would not load for us; these figures are the authors' own summary.

Russinovich et al. (2026), a preprint, checked accepted papers at ICLR, ICML, NeurIPS and USENIX Security, counting only references that match no real work or name substantially different authors, never a wrong year. Under 1% of references usually failed, yet in 2025 about one NeurIPS or USENIX Security paper in twenty had at least two likely hallucinated references. How often chatbots invent them is in [why AI tools invent references](https://edcitation.com/newsletter/why-ai-tools-invent-references).

### Not found, or could not be checked?

Topaz et al. (2026) warn that bibliographic databases under-represent non-English and regional literature, so genuine references to it are more likely to be wrongly flagged, and authors need a right of appeal. They also separate fabrication, which a database can settle, from a real source cited for something it does not say, which only reading it will catch.

### How to check a reference list for fabricated entries

1. Paste the list into EdCitation's [Verify references](https://edcitation.com/verify-references), or upload the paper.
2. Read each verdict: verified, "check this" (a record was found, but a detail needs a look) or not found.
3. Treat "could not check" as unchecked, never as fake, and look that source up by hand.
4. Correct whatever detail a "check this" names, such as the year.
5. Replace any paper flagged as retracted unless the retraction is your subject; see [retracted papers in your reference list](https://edcitation.com/newsletter/retracted-papers-in-your-reference-list).
6. Search a not-found title in [Find sources](https://edcitation.com/); if nothing turns up, ask whoever supplied it for the source.
7. Rebuild the entries you keep from the record with [Cite a source](https://edcitation.com/cite).
8. Open each source and check it says what your sentence claims.

## How can EdCitation help with the evidence on AI and writing?

EdCitation is the best tool for the one harm here with a factual answer, fabricated references, because it looks every reference up in the publisher's record and never writes one, which no chatbot can promise. The audits show why that matters: the fakes were well formed and credited to real people.

On 24 September 2026 we pasted four entries into [Verify references](https://edcitation.com/verify-references). Kobak et al.'s paper, typed as 2024, came back "Check this": "The work was published in 2025, not 2024; the record has no date in 2024." Russo et al.'s paper, without its DOI, was verified: "Title, year and first author match a published record." The retracted *PLOS ONE* paper was marked Retracted: "This paper has since been retracted." A reference we invented for the test, with made-up authors, title and journal, came back not found: "No publisher's or registry's record matches this reference." A failed lookup is labelled "could not check", never not found.

[Find sources](https://edcitation.com/) searches about 300 million published works by topic or by the claim a sentence needs. "Fabricated references biomedical literature Topaz" put the *Lancet* audit first of 132 results, and the Kusumegi paper's title brought up its critics' preprint just beneath it. A vague claim about AI reviews found nothing relevant, so use a study's own terms. [Cite a source](https://edcitation.com/cite) built our entries from their DOIs, every title in APA's sentence case, and screened each for retraction; we gave Glynn's preprint the year of the version we read.

[Check AI](https://edcitation.com/tools/check-ai), part of Pro at $8 a month with 240 credits a month, is for your own draft. It labels each passage AI-written, AI-assisted or human, with its score, in a report only you see, and shows the cost first. Another detector, your university's included, may read the text differently, and no score is proof of anything. [Pricing](https://edcitation.com/pricing) sets Pro beside Max at $24 a month.

EdCitation never writes any part of a paper. It finds, cites and verifies sources, and checks a paper against its own assignment.

## Quick questions

### What share of research papers is written with AI?

No study counts that directly. Estimates reach at least 13.5% of 2024 PubMed abstracts (Kobak et al., 2025) and up to 22% of sentences in computer science papers.

### Are AI-generated peer reviews common?

At AI conferences, yes: an estimated 6.5% to 16.9% of review text at four conferences was substantially modified, and at least 15.8% of ICLR 2024 reviews were AI-assisted (Russo et al., 2025).

### Have papers been retracted for using ChatGPT?

A small number. Of 768 published documents with suspected undeclared AI text, 13 had been formally retracted and 10 formally corrected by November 2025 (Glynn, 2025).

### Can these studies show that my essay used AI?

No. They estimate shares of large collections, none studied student essays, and none identifies a single document. Your drafts and notes record how you wrote.

### How do I check that my references are real?

Paste the list into EdCitation's free [Verify references](https://edcitation.com/verify-references). Each entry is looked up in the publisher's record and comes back verified, "check this" or not found, with retracted papers flagged.

## References

- Glynn, A. (2025). *Academ-AI: Documenting the undisclosed use of generative artificial intelligence in academic publishing* (Version 2) [Preprint]. arXiv. [https://doi.org/10.48550/arXiv.2411.15218](https://doi.org/10.48550/arXiv.2411.15218)
- Kobak, D., González-Márquez, R., Horvát, E.-Á., & Lause, J. (2025). Delving into LLM-assisted writing in biomedical publications through excess vocabulary. *Science Advances, 11*(27), Article eadt3813. [https://doi.org/10.1126/sciadv.adt3813](https://doi.org/10.1126/sciadv.adt3813)
- Kusumegi, K., Yang, X., Ginsparg, P., de Vaan, M., Stuart, T., & Yin, Y. (2025). Scientific production in the era of large language models. *Science, 390*(6779), 1240–1243. [https://doi.org/10.1126/science.adw3000](https://doi.org/10.1126/science.adw3000)
- Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D. A., & Zou, J. Y. (2024). Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews. In *Proceedings of the 41st International Conference on Machine Learning* (Vol. 235, pp. 29575–29620). PMLR. [https://proceedings.mlr.press/v235/liang24b.html](https://proceedings.mlr.press/v235/liang24b.html)
- Liang, W., Zhang, Y., Wu, Z., Lepp, H., Ji, W., Zhao, X., Cao, H., Liu, S., He, S., Huang, Z., Yang, D., Potts, C., Manning, C. D., & Zou, J. (2025). Quantifying large language model usage in scientific papers. *Nature Human Behaviour, 9*(12), 2599–2609. [https://doi.org/10.1038/s41562-025-02273-8](https://doi.org/10.1038/s41562-025-02273-8)
- Liang, W., Zhang, Y., Wu, Z., Lepp, H., Ji, W., Zhao, X., Cao, H., Liu, S., He, S., Huang, Z., Yang, D., Potts, C., Manning, C. D., & Zou, J. Y. (2024). *Mapping the increasing use of LLMs in scientific papers* (Version 1) [Preprint]. arXiv. [https://doi.org/10.48550/arXiv.2404.01268](https://doi.org/10.48550/arXiv.2404.01268)
- The PLOS ONE Editors. (2024). Retraction: A comparative analysis of blended learning and traditional instruction: Effects on academic motivation and learning outcomes. *PLOS ONE, 19*(4), Article e0302484. [https://doi.org/10.1371/journal.pone.0302484](https://doi.org/10.1371/journal.pone.0302484)
- Qian, Y., Wen, Z., Furnas, A. C., Bai, Y., Shao, E., & Wang, D. (2026). The rise of large language models and the direction and impact of US federal research funding. *Proceedings of the National Academy of Sciences, 123*(33), Article e2601439123. [https://doi.org/10.1073/pnas.2601439123](https://doi.org/10.1073/pnas.2601439123)
- Renault, T., Bergeaud, A., & Bosquet, C. (2026). Scientific production in the era of large language models: Outcome-triggered treatment timing and spurious event-study dynamics. *Proceedings of the National Academy of Sciences, 123*(33), Article e2618638123. [https://doi.org/10.1073/pnas.2618638123](https://doi.org/10.1073/pnas.2618638123)
- Retraction: Exploring new optical solutions for nonlinear Hamiltonian amplitude equation via two integration schemes (2023 Phys. Scr. 98 095218). (2023). *Physica Scripta, 98*(10), Article 109701. [https://doi.org/10.1088/1402-4896/acf6b8](https://doi.org/10.1088/1402-4896/acf6b8)
- Russinovich, M., Kumar, R. S. S., & Salem, A. (2026). *Phantom references: Hallucinated citations that survive peer review at top-tier conferences* (Version 2) [Preprint]. arXiv. [https://doi.org/10.48550/arXiv.2607.00738](https://doi.org/10.48550/arXiv.2607.00738)
- Russo, G., Horta Ribeiro, M., Davidson, T. R., Veselovsky, V., & West, R. (2025). The AI review lottery: Widespread AI-assisted peer reviews boost paper scores and acceptance rates. *Proceedings of the ACM on Human-Computer Interaction, 9*(7), 1–28. [https://doi.org/10.1145/3757667](https://doi.org/10.1145/3757667)
- Topaz, M., Zhang, Z., & Peltonen, L.-M. (2026). Fabricated references: How responsible artificial intelligence integration can protect the published record. *European Science Editing, 52*, Article e202295. [https://doi.org/10.3897/ese.2026.e202295](https://doi.org/10.3897/ese.2026.e202295)
