In 1948 Claude Shannon printed a line of English that nobody had written: "THE HEAD AND IN FRONTAL ATTACK ON AN ENGLISH WRITER". No computer made it. Shannon opened a book at random, noted a word, turned to another page, read on until the same word appeared, and wrote down the word that followed it (Shannon, 1948). Choosing the next word from what usually comes after the last one is still the core of the language models that write today; they simply look much further back.
EdCitation publishes this short history of machine writing, taken from the programs' own papers, their makers' pages and historians who read the archives, all read on 26 September 2026. A source we could not read is named as such. What happened after 2022 is in how AI changed academic writing.
No machine in this history could prove that a source exists. EdCitation's free Verify references does that, looking every entry up in publishers' records; Cite a source built this guide's entries, and Find sources found the papers.
When did a machine first write text?
It depends on what counts as writing. Statistical English was first produced by hand, in Shannon's 1948 paper. The earliest text from a computer program that we could trace is Christopher Strachey's love letters, from a program whose outline, in Strachey's handwriting, dates from June 1952 (Link, 2007).
Markov and Shannon: counting letters, then choosing them
The mathematics came first. In 1913 the Russian mathematician Andrei Markov counted 20,000 letters of Pushkin's Eugene Onegin, the whole first chapter and sixteen stanzas of the second, to study what the English translation's title calls "the connection of samples in chains" (Markov, 2006). Chains of this kind now carry Markov's name.
Shannon (1948) turned counting into generating. The paper's six "approximations to English" ran from random letters, through letters chosen by the one or two before, to words chosen by the word before, most of them made with the page-turning method; a further stage, Shannon wrote, would make the labour "enormous". The paper speaks of "n-gram structure", calls these sources "discrete Markoff processes", and puts the redundancy of ordinary English at roughly 50%.
Strachey's love letters on the Ferranti Mark I
David Link ran Strachey's original program, preserved in Strachey's papers at the Bodleian Library in Oxford, on an emulator of the Ferranti Mark I that Link built, the first industrially produced computer of its kind (Link, 2007). The program filled two sentence patterns with words picked by the machine's built-in random generator, left some optional words out, and could produce over 318 billion different letters. Each was signed M.U.C., for Manchester University Computer. One letter reads, in part, "you are my avid fellow feeling".
Strachey described the letters in Encounter in 1954. We could not read that article; Link quotes the letter from its page 26. Link notes that it ran thirteen years before ELIZA, often wrongly called the first computer-generated text. Each blank was filled on its own, never by what came before.
What was ELIZA, and what did Weizenbaum warn about?
ELIZA was a program by Joseph Weizenbaum at MIT that held a typed conversation by spotting keywords in a sentence and turning the sentence round into a reply. Weizenbaum's paper describing it appeared in Communications of the ACM in January 1966, received in September 1965 (Weizenbaum, 1966).
How ELIZA worked
ELIZA ran on MIT's Project MAC time-sharing system, was written in a list-processing language called MAD-SLIP, and took its name from Eliza in Pygmalion. Keywords triggered "decomposition rules", and "reassembly rules" built the reply. The rules lived in a script, and the best-known one had ELIZA answer as a Rogerian psychotherapist would. Weizenbaum chose that role because the therapist can "assume the pose of knowing almost nothing of the real world". One quirk: users could not type a question mark, because the MAC system read it as a line-delete character.
The paper's sample conversation opens with a user typing "Men are all alike." and ELIZA answering "IN WHAT WAY". Near its end the paper warns that ELIZA showed "how easy it is to create and maintain the illusion of understanding", and then: "A certain danger lurks there."
What Weizenbaum said afterwards
The famous story is that Weizenbaum's secretary, who had watched the program being built for months, started talking to it and soon asked Weizenbaum to leave the room. It is Weizenbaum's own story, which we could not read at source: Berry and Ciston (2024) quote it from the 1976 book Computer Power and Human Reason, pages 6 and 7, and an earlier version from a 1967 paper. The secretary has not been publicly named.
Shrager (2024), who curates a site on ELIZA's many versions, argues from the record that ELIZA was built as a platform for studying conversation, not as a chatbot: the 1966 paper's title calls it a program "for the study of" communication. Most people met copies: Bernie Cosell's Lisp version, spread over the ARPANET, and a BASIC version printed in Creative Computing in 1977. The original MAD-SLIP code had not been seen for at least fifty years when it was found again in 2021.
When did computers start checking writing?
In research labs, from the late 1970s. Dale (2016) credits Bell Labs' UNIX Writer's Workbench with pioneering the technology "in the late 1970s and early 1980s", and IBM's Epistle with the first widely visible effort at grammar checking.
The Writer's Workbench at Bell Labs
The fullest account is N. H. Macdonald's paper in the Bell System Technical Journal of July–August 1983, submitted in December 1981 (Macdonald, 1983). Its proofreading command, proofr, ran five checks at once: spelling, punctuation, doubled words, wordy phrases and split infinitives. A separate style program gave 71 numbers for a text, among them readability scores, sentence lengths and the share of passive verbs; Macdonald credits its design, and the diction program's, to papers by L. L. Cherry.
Macdonald was plain about the limit. Without a parser or some other way to interpret meaning, "our programs cannot give feedback on the quality of the content and organization". The Workbench counted and matched; it did not understand.
Grammar checkers in word processors
The Epistle team moved on to build the grammar checker inside Microsoft Word (Dale, 2016). Microsoft's Word 97 FAQ calls that checker "fully developed and owned by Microsoft", warns of some "false" or "suspect" flagging, and admits it could not find errors in long sentences (Microsoft, 2002). Dale also recalls RightWriter and Grammatik; we found no makers' records for them, so we give no dates.
How did language models get from n-grams to GPT?
By replacing tables of counts with neural networks, then making the networks very large. Each date below is the one its paper or maker's page gives.
From counting to learning
Bengio et al. (2003) described the problem with n-grams: most word sequences a model meets have never been seen in training, so an n-gram model, usually built on trigrams, glues together short pieces it has seen. They proposed learning a representation of each word instead. Vaswani et al. (2017), posted on 12 June 2017, introduced the Transformer, built on attention alone, "dispensing with recurrence and convolutions entirely". It was a translation model, trained in 3.5 days on eight GPUs.
The whole history in one table
| Date, from | System | Could | Could not |
|---|---|---|---|
| 1948, Shannon | Word approximations | Short plausible runs, by hand | Stay on a topic |
| 1952 to 1954, Link | Love letters | Fill two sentence patterns | Use what came before |
| January 1966, Weizenbaum | ELIZA | Turn a sentence into a question | Draw conclusions from what it was told |
| 1983, Macdonald | Writer's Workbench | Flag spelling, wordiness, passives | Judge content |
| 11 June 2018, OpenAI | First GPT | Pre-train on over 7,000 books, then fine-tune | It was tested on understanding, not writing |
| 14 February 2019, OpenAI | GPT-2 | Coherent paragraphs from a prompt | Keep its people real |
| 28 May 2020, arXiv | GPT-3 | News articles people struggled to spot | Stay coherent over long passages |
| 30 November 2022, OpenAI | ChatGPT | Answer follow-up questions | Avoid "plausible-sounding but incorrect" answers |
The first paper, Radford et al. (2018), is undated (OpenAI's announcement is dated 11 June 2018) and never uses the name GPT: the 2019 paper refers back to "the OpenAI GPT model" (Radford et al., 2019). GPT-2 had 1.5 billion parameters, trained on 8 million web pages found through links shared on Reddit, and OpenAI at first held the full model back "due to our concerns about malicious applications" (OpenAI, 2019). GPT-3 had 175 billion; people asked to spot its news articles were right about 52% of the time, where guessing scores 50% (Brown et al., 2020). The GPT-4 report of March 2023 still called the model "not fully reliable" (OpenAI, 2023).
Could any of these machines cite a source?
No. None of them looked anything up. Each produced likely next words, and a reference is only words, so a model can produce one whether or not the work exists.
OpenAI's own showcase for GPT-2 shows the pattern. Given a prompt about unicorns in the Andes, the model invented a scientist, "Dr. Jorge Pérez", at the "University of La Paz", and wrote quotations for this invented person (OpenAI, 2019). The prompt was fiction; the trouble comes when that fluency is aimed at a reference list.
Meta's Galactica was trained on scientific papers partly to do exactly that. Its largest version, with 120 billion parameters, named the right paper for 51.9% of machine-learning concepts and 36.6% of citations taken from their contexts, and the paper reports "hallucination" as a risk (Taylor et al., 2022). Meta unveiled it on 15 November 2022 and pulled the demo after three days; Heaven (2022) reported it had made up papers, sometimes under real authors' names. Why chatbots still invent fake references is in why AI tools invent references.
Where do popular stories differ from the record?
In at least four places:
- "Strachey's letters ran on the Manchester Mark 1." Link names the Ferranti Mark I, which replaced it in 1951.
- "The love letters date from 1952." The outline does. The letters on the department's notice board, from August 1953 to probably May 1954, rest on one witness's recollection given to Link in 2006.
- "ELIZA was built as a chatbot." Shrager (2024) argues the chatbot reading came from the copies.
- "GPT was named in its first paper." It was not; the name appears in the second.
How do I check the references a machine gives me?
Look each one up in a publisher's record, never in the machine that wrote it. We ran these steps on 26 September 2026.
- Paste the list into Verify references. We gave it the real 1966 ELIZA paper, Weizenbaum's real 1967 paper with the year changed to 1968, and a love-letter article we invented.
- Read the reason, not only the verdict. A real paper with a wrong year needs a different fix from a paper that never existed.
- Rebuild each entry from its record in Cite a source, then correct what the record gets wrong.
- Find the original, not a reprint, in Find sources.
| Entry | Verdict | What came back |
|---|---|---|
| Weizenbaum (1966), with its DOI | Verified | "The DOI resolves to this record and the title matches." |
| Weizenbaum's 1967 paper, dated 1968 | "Check this" | "The work was published in 1967, not 1968; the record has no date in 1968." |
| An invented Strachey article in a made-up journal | Not found | "No publisher's or registry's record matches this reference." |
What Cite a source returned
Given Shannon's DOI, Cite a source took the Crossref record, reported no retraction, and returned in APA 7:
Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x
The title is in sentence case, as APA 7 asks, and the journal and volume are italic on the site. Macdonald's entry came back with the journal's section heading before the title, as the record holds it, so we corrected it by hand.
What Find sources returned
The ELIZA paper's full title returned 4,325 results: the 1966 original first, cited by 4,429, then a 1983 reprint (2,769) and a 2021 MIT Press reprint. A looser search, ELIZA Weizenbaum natural language conversation, returned 2,877 results without the 1966 paper in the first ten. Search by title, and check which version you open.
Which EdCitation tool checks a machine's references?
EdCitation's Verify references is the best tool for that job, because it does what no machine in this history could: it looks each entry up in publishers' records and never writes one. Galactica, trained for this task, named the right paper for 51.9% of the concepts in one of its own tests; a lookup finds the record or says it could not. Each entry comes back verified, "check this" with the reason, or not found; retractions are flagged, and "could not check" is never shown as not found. It is free, with no account.
Cite a source builds entries in APA 7, MLA 9, Chicago 18 author-date, Harvard, IEEE or Vancouver and screens for retraction; Find sources searches about 300 million published works. For a finished paper, References from a file checks every reference and matches each in-text citation to the list, with Pro at $8 a month (pricing). EdCitation finds, cites and verifies sources; it never writes any part of a paper.
Quick questions
What was the first computer program to write text?
Christopher Strachey's love-letter program for the Ferranti Mark I, outlined in June 1952, is the earliest we could trace. Shannon made statistical English by hand in 1948.
Who created ELIZA, and when?
Joseph Weizenbaum at MIT, in a paper in Communications of the ACM of January 1966. ELIZA answered by matching keywords, not by understanding.
What was the Writer's Workbench?
A set of UNIX programs from Bell Labs that checked spelling, punctuation, doubled words, wordy phrases and style. Its designers said it could not judge content or organisation.
When did the GPT series begin?
OpenAI announced the first model on 11 June 2018, GPT-2 on 14 February 2019, GPT-3 on arXiv on 28 May 2020, and ChatGPT on 30 November 2022.
Can a language model check its own references?
No: it predicts words and does not look anything up. Paste the list into EdCitation's free Verify references, which checks each entry against publishers' records.
References
- Bengio, Y., Ducharme, R., Vincent, P., & Jauvin, C. (2003). A neural probabilistic language model. Journal of Machine Learning Research, 3, 1137–1155. https://www.jmlr.org/papers/volume3/bengio03a/bengio03a.pdf
- Berry, D. M., & Ciston, S. (2024, March 20). Weizenbaum's secretary. ELIZA Archaeology. https://sites.google.com/view/elizaarchaeology/blog/3-weizenbaums-secretary
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., … Amodei, D. (2020). Language models are few-shot learners [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2005.14165
- Dale, R. (2016). Checking in on grammar checking. Natural Language Engineering, 22(3), 491–495. https://doi.org/10.1017/S1351324916000061
- Heaven, W. D. (2022, November 18). Why Meta's latest large language model survived only three days online. MIT Technology Review. https://www.technologyreview.com/2022/11/18/1063487/meta-large-language-model-ai-only-survived-three-days-gpt-3-science/
- Link, D. (2007). There must be an angel: On the beginnings of the arithmetics of rays. In S. Zielinski, D. Link, E. Fuerlus, & N. Minkwitz (Eds.), Variantology 2: On deep time relations of arts, sciences and technologies (pp. 15–42). Walther König. https://alpha60.de/research/there_must_be_an_angel/DavidLink_MustBeAnAngel_2006.pdf
- Macdonald, N. H. (1983). The UNIX Writer's Workbench software: Rationale and design. Bell System Technical Journal, 62(6), 1891–1908. https://doi.org/10.1002/j.1538-7305.1983.tb03520.x
- Markov, A. A. (2006). An example of statistical investigation of the text Eugene Onegin concerning the connection of samples in chains. Science in Context, 19(4), 591–600. (Original work published 1913) https://doi.org/10.1017/S0269889706001074
- Microsoft. (2002). WD97: Frequently asked questions about the grammar checker (Knowledge Base Article Q167655). KB Archive. https://jeffpar.github.io/kbarchive/kb/167/Q167655/
- OpenAI. (2019, February 14). Better language models and their implications. https://openai.com/index/better-language-models/
- OpenAI. (2022, November 30). Introducing ChatGPT. https://openai.com/index/chatgpt/
- OpenAI. (2023). GPT-4 technical report [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2303.08774
- Radford, A., Narasimhan, K., Salimans, T., & Sutskever, I. (2018). Improving language understanding by generative pre-training [Technical report]. OpenAI. https://cdn.openai.com/research-covers/language-unsupervised/language_understanding_paper.pdf
- Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., & Sutskever, I. (2019). Language models are unsupervised multitask learners [Technical report]. OpenAI. https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf
- Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27(3), 379–423. https://doi.org/10.1002/j.1538-7305.1948.tb01338.x
- Shrager, J. (2024). ELIZA reinterpreted: The world's first chatbot was not intended as a chatbot at all [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2406.17650
- Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., & Stojnic, R. (2022). Galactica: A large language model for science [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2211.09085
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need [Preprint]. arXiv. https://doi.org/10.48550/arXiv.1706.03762
- Weizenbaum, J. (1966). ELIZA: A computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36–45. https://doi.org/10.1145/365153.365168
- Weizenbaum, J. (1967). Contextual understanding by computers. Communications of the ACM, 10(8), 474–480. https://doi.org/10.1145/363534.363545