# AI transcription of research interviews: ethics and GDPR

**Can I use AI to transcribe research interviews, and what does my ethics committee need to know?** You can usually use AI to transcribe research interviews once your ethics committee has approved the tool, before any participant audio is uploaded. Universities prefer tools they license, such as Microsoft Teams, or transcription run on their own computers, and warn against unapproved cloud services. Participants must be told, and every AI transcript must be checked against the audio.

Published 2026-09-26 by EdCitation. https://edcitation.com/newsletter/ai-transcription-tools-and-research-ethics

You can usually use AI to transcribe research interviews, provided the tool is named in your ethics application and approved before a single recording is uploaded. The universities we read prefer tools they already license, or transcription that runs on their own machines, and they are wary of free cloud services whose servers sit abroad. Whatever you use, the transcript is yours to check against the audio.

This guide is published by EdCitation. We read every rule below on the university's or regulator's own page on 27 September 2026, and we name transcription products only where those pages name them. We make no claim that any product is compliant or accurate. This is general information, not legal advice: your ethics committee, data protection officer and supervisor decide for your project.

EdCitation does not transcribe anything, and it never writes any part of a paper. What it does is the reference work around the method: [Find sources](https://edcitation.com/) for the methods literature, [Cite a source](https://edcitation.com/cite) for each entry, and [Verify references](https://edcitation.com/verify-references) to match every entry to a real record before the ethics application or the thesis goes in. We ran all three on this topic and quote the results near the end.

## Can I use AI to transcribe research interviews?

Yes, at every university we read, but only with approval. The University of Oxford (2025) counts using generative AI to transcribe interviews "to save time and effort" as substantive use, and says uploading personal data to such a tool is processing that, in research, needs ethics approval. The University of Melbourne (n.d.) says AI use with participant data must be described in the ethics application first, because retrospective approvals are "not normally permissible".

Georgia State University (2023) puts it in one line: AI transcription goes in the IRB application "as a data sharing practice". Sending audio to a transcription company shares participants' voices with a third party, so expect questions on who, where and for how long.

## What do universities and regulators say about AI transcription?

They agree on the principle and differ on the tools. Every page we read wants the tool approved, participants told and the output checked; which products are approved depends on each university's contracts and security reviews.

| University or body | Country | What it allows or requires | Source |
| --- | --- | --- | --- |
| University of Oxford | UK | Transcription is substantive GenAI use; personal data only where confidentiality is guaranteed; a locally deployed model or the licensed ChatGPT Edu; declare the tool | University of Oxford (2025) |
| University College London | UK | "Strongly recommend" avoiding services controlled from outside the EU, asked about Otter.ai; use services under a UCL enterprise agreement | University College London (n.d.) |
| Loughborough University | UK | Use transcription in Microsoft Word and Teams first; names otter.ai and MeetGeek as risky; review transcripts carefully | Loughborough University (2026) |
| Georgia State University | US | Describe AI transcription in the IRB application as data sharing; a limited number of otter.ai licences through the university | Georgia State University (2023) |
| New York University Libraries | US | Consent from all parties despite New York's one-party law; follow the IRB; check outputs; provides Zoom, NYU Stream and Word online | New York University Libraries (2026) |
| University of Calgary | Canada | Rev.com, Transcription Heroes, aTrain, Buzz and MacWhisper approved for de-identified Level 2 data; Zoom to Level 3; Teams to Level 4 | University of Calgary Library (2026) |
| University of Melbourne | Australia | Ethics approval before use; information sheets explain the tool; check every transcript | University of Melbourne (n.d.) |
| Utrecht University | Netherlands | Amberscript "approved for the use of sensitive data"; ask before using another tool | Utrecht University (n.d.) |
| University of Groningen | Netherlands | Whisper run on the university's own computing cluster, no external cloud | University of Groningen (2025) |
| Singapore Management University | Singapore | Whisper on a library terminal; delete files afterwards | Ratmelia (2025) |
| Information Commissioner's Office | UK | Consent to take part is not the lawful basis; AI can trigger a DPIA | Information Commissioner's Office (n.d.-a, n.d.-b) |
| Office of the Australian Information Commissioner | Australia | Privacy law covers AI input and output; keep personal information out of public generative AI | Office of the Australian Information Commissioner (2025) |
| The reference list in your application | Anywhere | Every entry matched to a real record | EdCitation's [Verify references](https://edcitation.com/verify-references), free |

### Where the guidance disagrees

The same product can be approved at one university and discouraged at the next. Georgia State hands out otter.ai licences; UCL advises against services like it controlled from outside the EU; Loughborough lists it among the risky tools. Calgary approves several services only for de-identified data, and Teams for more sensitive material than Zoom. What decides is your university's contract, your class of data and your committee, not the product's name.

## Where is the audio processed and stored?

It depends on the kind of tool, the first thing a committee asks about. There are three routes.

- **A service your university licenses.** UCL (n.d.) explains that its Microsoft agreement keeps data in Microsoft data centres meeting UK and EU standards, and that a service without an enterprise agreement may store data in several countries, including the US.
- **A free or personal cloud account.** Loughborough University (2026) warns that some services act as data controllers, deciding what else to do with an upload, that many train their models on uploaded audio, and that hosting outside the UK may move data somewhere without adequate protection.
- **A tool that runs on a machine you control.** Nothing leaves the building unless you move it. For any service, NYU's guide asks where its servers are and whether it trains on your recordings (New York University Libraries, 2026).

### Local or offline transcription

Several universities now offer transcription that stays on their premises. The University of Groningen (2025) runs Whisper on its own computing cluster with "no communication to external cloud services". Singapore Management University's library offers Whisper on a bookable terminal and asks users to delete their files afterwards (Ratmelia, 2025). Calgary lists aTrain, Buzz and MacWhisper as running on your own machine.

Oxford adds that cloud models change version and cannot be archived, which hurts reproducibility. Da Silva (2021) describes a secure method that more than halved interview transcription time in doctoral research. Local does not mean accurate, though: it settles where the data goes, not what the transcript says.

## What should the consent form and information sheet say?

Participants should learn, before they agree, that AI will transcribe their words, where, and what becomes of the files. The University of Melbourne (n.d.) lists what to explain: the tool's role, how data are handled, whether an outside provider may keep or use them, and the risks. Two regulators add points students miss:

- **UK:** the ICO says that consent to take part in research is distinct from consent as the UK GDPR lawful basis, and that research processing usually rests on public task or legitimate interests instead (Information Commissioner's Office, n.d.-a). Your information sheet still has to say what happens; see [lawful basis for processing](https://edcitation.com/glossary/research-ethics#lawful-basis-for-processing).
- **Australia:** the OAIC says privacy obligations apply to personal information put into an AI system and to AI output containing it, and recommends, as best practice, not entering personal or sensitive information into publicly available generative AI (Office of the Australian Information Commissioner, 2025).

Calgary puts the onus on the researcher to keep data de-identified, which may mean asking participants not to say names or workplaces on the recording.

### Do I need a DPIA?

Possibly. The ICO lists innovative technology, including AI, as requiring a [data protection impact assessment](https://edcitation.com/glossary/research-ethics#data-protection-impact-assessment) when combined with another high-risk criterion, such as special category data (Information Commissioner's Office, n.d.-b). Ask your data protection office early.

### Sample wording for the information sheet

This wording is ours, for your committee to amend; keep only what is true.

> Your interview will be audio recorded. The recording will be turned into text by [tool], which [runs on University computers / is provided under a University contract and stores data in (location)]. [The provider does not use recordings to train its systems.] The researcher will check the text against the recording, replace names with pseudonyms, and delete the recording by [date or event]. If you would rather not be transcribed this way, tell the researcher before we start.

## How do you check an AI transcript against the audio?

By listening to the whole recording while reading the transcript, not by skimming the text. Koenecke et al. (2024) found that roughly 1% of Whisper transcriptions in their tests, run in 2023, contained whole phrases or sentences absent from the audio, and 38% of those inventions included explicit harms. Hallucinations were more frequent for speakers with longer pauses, as in aphasia.

Calgary asks researchers to listen to the original audio while reading, so the transcript has "nothing more or less". Melbourne warns that a tool's failure is no defence against misquoting a participant, and Loughborough flags names and technical terms. Bokhove and Downey (2018) judged automated transcripts "good enough" for a first version, with most mismatches easy to fix in review.

### A procedure from planning to write-up

This order is ours, built on the pages above.

1. Ask the ethics or data protection office which transcription tools are approved, and for which classes of data.
2. Classify your data: identifiable, special category, or de-identified before recording.
3. Choose a licensed service, a local tool or a contracted human transcriber, never a personal account.
4. In the ethics application, name the tool, where it processes data, retention and deletion; complete a DPIA if required.
5. Put the same facts in the information sheet and consent form.
6. Transcribe, then listen to every recording in full against its transcript. Mark silences, overlapping speech, names and figures.
7. Pseudonymise, store as approved, and delete recordings when you said you would.
8. Report the method in the methods chapter, and cite the transcription literature: [Cite a source](https://edcitation.com/cite) builds each entry from its record, and [Verify references](https://edcitation.com/verify-references) checks the finished list.

Quoting participants in the thesis is covered in [how to cite an interview transcript](https://edcitation.com/newsletter/how-to-cite-an-interview-transcript).

## How do you declare AI transcription in the ethics application and methods chapter?

Name the tool, its version or licence, where it ran, and who checked the output and how. Oxford asks for the name, version, date and use of a discrete generative AI tool, as opposed to a function inside software you already use, and wants limitations and validation discussed in the methods (University of Oxford, 2025). Melbourne requires such use to be disclosed, theses included. Our guide to [research ethics approval for a dissertation](https://edcitation.com/newsletter/research-ethics-approval-for-a-dissertation) covers the application as a whole, and [how to write the methodology chapter](https://edcitation.com/newsletter/methodology-chapter-how-to-write-it) covers where this paragraph sits.

A methods sentence can be short. This one is ours: "Interviews were recorded and transcribed with Microsoft Teams under the University's licence, as approved by the ethics committee (reference). I checked each transcript in full against its recording, corrected errors and replaced names with pseudonyms before analysis." For recording meetings rather than interviews, see [AI note-takers in meetings](https://edcitation.com/newsletter/ai-note-takers-in-meetings) and [recording a meeting: the law on consent](https://edcitation.com/newsletter/recording-a-meeting-the-law-on-consent).

## Where does EdCitation help with a transcription study?

In the literature, not the audio. For the references in an ethics application or methods chapter, EdCitation is, in our view, the best tool, because it looks each source up in the publisher's record and never writes one. A chatbot composes references from memory: the same fault as a transcription tool that invents a sentence.

### Verify references on a transcription reading list

We pasted six real entries into [Verify references](https://edcitation.com/verify-references): Bokhove and Downey, Da Silva, Koenecke and colleagues, McMullin's 2023 *Voluntas* paper on transcription, Samuel and Wassenaar's editorial on consent and AI transcription, and Oxford's policy page. Five came back verified. For Bokhove and Downey: "The DOI resolves to this record and the title matches." For the editorial, entered without its DOI and dated by its 2025 issue: "Title, year and first author match a published record."

Oxford's page came back unchecked, not "not found": "The site at this address turns automated readers away, so the page could not be read here. Open it in a browser and check it is the page meant." Correctly so: the page exists, and we read it in a browser.

### Cite a source on two DOIs

Given 10.1177/2059799120987766, [Cite a source](https://edcitation.com/cite) returned in APA 7 "Da Silva, J. (2021). Producing ‘good enough’ automated transcripts securely: Extending Bokhove and Downey (2018) to address security concerns. Methodological Innovations, 14(1), Article 2059799120987766." and the DOI, with journal and volume in italics. Given 10.1145/3630106.3658996 it returned, in a rerun on 27 September 2026, "Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H., & Sloane, M. (2024). Careless whisper: Speech-to-text hallucination harms. In The 2024 ACM Conference on Fairness, Accountability, and Transparency (pp. 1672–1681). ACM." The record's title case came back in sentence case, as APA 7 wants; before pasting, give Whisper, the transcription model's name, its capital back. For Samuel and Wassenaar, online in 2024 and in a 2025 issue, Cite a source gave 2025, the issue's year, as APA asks.

### Find sources on the topic

[Find sources](https://edcitation.com/) returned 872,744 works for "automated transcription qualitative interviews privacy". The first was "Automated Transcription of Interviews in Qualitative Research Using Artificial Intelligence A Simple Guide" by Jacob Rosenberg, 2024, in the *Journal of Surgery Research and Practice*. We have not read it; open any result before citing it. Several book chapters lower down listed no author: check those at the publisher.

All three tools are free with no account. Pro ($8 a month) adds [References from a file](https://edcitation.com/tools/references-from-a-file), which checks every reference in an uploaded thesis and matches each in-text citation to the list, and [Mechanics QA](https://edcitation.com/tools/mechanics-qa); see [pricing](https://edcitation.com/pricing).

## Quick questions

### Is Otter.ai allowed for research interviews?

It depends on your university. Georgia State offers otter.ai licences with IRB disclosure; UCL strongly recommends avoiding services like it controlled from outside the EU. Ask your ethics office before uploading anything.

### Is Whisper transcription safe for research data?

Where it runs matters. Groningen and Singapore Management University run Whisper on their own machines so audio stays on site, but Koenecke et al. (2024) found it sometimes invents whole sentences, so every transcript still needs checking against the audio.

### Do I have to tell participants that AI will transcribe their interview?

Yes, at every university we read. The information sheet names the tool, where data goes, whether the provider keeps it and when recordings are deleted.

### Is consent the lawful basis for AI transcription under UK GDPR?

Usually not. The ICO says consent to take part is an ethical standard distinct from the UK GDPR lawful basis, which in research is usually public task or legitimate interests.

### Can EdCitation transcribe my interviews?

No. EdCitation finds, cites and checks sources. Its free [Verify references](https://edcitation.com/verify-references) matches the methods literature in your application or thesis to real records.

## References

- Bokhove, C., & Downey, C. (2018). Automated generation of 'good enough' transcripts as a first step to transcription of audio-recorded data. *Methodological Innovations, 11*(2), Article 2059799118790743. [https://doi.org/10.1177/2059799118790743](https://doi.org/10.1177/2059799118790743)
- Da Silva, J. (2021). Producing 'good enough' automated transcripts securely: Extending Bokhove and Downey (2018) to address security concerns. *Methodological Innovations, 14*(1), Article 2059799120987766. [https://doi.org/10.1177/2059799120987766](https://doi.org/10.1177/2059799120987766)
- Georgia State University. (2023, November 29). *IRB protocols for AI-generated transcripts*. University Research Services & Administration. [https://ursa.research.gsu.edu/2023/11/29/irb-protocols-for-ai-generated-transcripts/](https://ursa.research.gsu.edu/2023/11/29/irb-protocols-for-ai-generated-transcripts/)
- Information Commissioner's Office. (n.d.-a). *Principles and grounds for processing*. Retrieved September 27, 2026, from [https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/the-research-provisions/principles-and-grounds-for-processing/](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/the-research-provisions/principles-and-grounds-for-processing/)
- Information Commissioner's Office. (n.d.-b). *When do we need to do a DPIA?* Retrieved September 27, 2026, from [https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/data-protection-impact-assessments-dpias/when-do-we-need-to-do-a-dpia/](https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/accountability-and-governance/data-protection-impact-assessments-dpias/when-do-we-need-to-do-a-dpia/)
- Koenecke, A., Choi, A. S. G., Mei, K. X., Schellmann, H., & Sloane, M. (2024). Careless Whisper: Speech-to-text hallucination harms. In *The 2024 ACM Conference on Fairness, Accountability, and Transparency* (pp. 1672–1681). ACM. [https://doi.org/10.1145/3630106.3658996](https://doi.org/10.1145/3630106.3658996)
- Loughborough University. (2026, February 6). *AI transcription tools: A time-saver or security risk?* [https://www.lboro.ac.uk/internal/news/2026/february/ai-transcription-tools/](https://www.lboro.ac.uk/internal/news/2026/february/ai-transcription-tools/)
- New York University Libraries. (2026, September 1). *Evaluating generative AI tools for academic research: For transcription*. [https://guides.nyu.edu/ai-tools/transcription](https://guides.nyu.edu/ai-tools/transcription)
- Office of the Australian Information Commissioner. (2025, January 17). *Guidance on privacy and the use of commercially available AI products*. [https://www.oaic.gov.au/privacy/privacy-guidance-for-organisations-and-government-agencies/guidance-on-privacy-and-the-use-of-commercially-available-ai-products](https://www.oaic.gov.au/privacy/privacy-guidance-for-organisations-and-government-agencies/guidance-on-privacy-and-the-use-of-commercially-available-ai-products)
- Ratmelia, B. (2025, August 15). *On-campus Whisper AI: Safe and IRB-compliant transcription for research projects*. Singapore Management University Libraries. [https://library.smu.edu.sg/topics-insights/campus-whisper-ai-safe-and-irb-compliant-transcription-research-projects](https://library.smu.edu.sg/topics-insights/campus-whisper-ai-safe-and-irb-compliant-transcription-research-projects)
- University College London. (n.d.). *Frequently asked questions about cloud services*. Data Protection. Retrieved September 27, 2026, from [https://www.ucl.ac.uk/data-protection/guidance-staff-students-and-researchers/research/frequently-asked-questions-about-cloud-services](https://www.ucl.ac.uk/data-protection/guidance-staff-students-and-researchers/research/frequently-asked-questions-about-cloud-services)
- University of Calgary Library. (2026, September 16). *NVivo: Transcription*. [https://libguides.ucalgary.ca/c.php?g=709222&p=5052798](https://libguides.ucalgary.ca/c.php?g=709222&p=5052798)
- University of Groningen. (2025, September 22). *Free secure AI transcription tool now available for everyone at the UG*. Digital Competence Centre. [https://www.rug.nl/digital-competence-centre/contact/news/2025/free-secure-ai-transcription-tool-available-for-everyone-at-the-ug?lang=en](https://www.rug.nl/digital-competence-centre/contact/news/2025/free-secure-ai-transcription-tool-available-for-everyone-at-the-ug?lang=en)
- University of Melbourne. (n.d.). *Responsible use of AI in research*. Retrieved September 27, 2026, from [https://research.unimelb.edu.au/strengths/responsible-research/trusted-research/responsible-use-of-AI-in-research](https://research.unimelb.edu.au/strengths/responsible-research/trusted-research/responsible-use-of-AI-in-research)
- University of Oxford. (2025). *Policy for using generative AI in research* (Version 1.0). [https://www.ox.ac.uk/research/support/governance-and-committees/research-policies/policy-for-using-generative-ai-in](https://www.ox.ac.uk/research/support/governance-and-committees/research-policies/policy-for-using-generative-ai-in)
- Utrecht University. (n.d.). *Transcription service Amberscript*. Research Data Management Support. Retrieved September 27, 2026, from [https://www.uu.nl/en/research/research-data-management/tools/transcription-of-audio-data/transcription-service-amberscript](https://www.uu.nl/en/research/research-data-management/tools/transcription-of-audio-data/transcription-service-amberscript)
