The short answer
You can use AI to transcribe research interviews under GDPR. Most of the work is the same as for any research data: a lawful basis, clear information to participants, a processor agreement with the service, and good security. The choice of service matters most for three questions: where the audio is processed, whether the provider trains on it, and how long it keeps it.
- Tell participants in the information sheet and consent form that recordings will be transcribed by an automated service, and what kind.
- Pick a service that processes audio in the EU or EEA and does not train its models on your recordings.
- Make sure a data processing agreement (DPA) covers the service. At universities this is usually signed centrally.
- Pseudonymise transcripts before analysis, and delete recordings once they are no longer needed.
- If the interviews touch health, ethnicity, religion, politics, sexuality or other special categories, check that your ethics approval covers the tool.
Why GDPR applies to interview recordings
A voice recording of an identifiable person is personal data. So is a transcript that contains names, workplaces, places or stories that could identify someone. For you as the researcher, pseudonymised data is still personal data under GDPR (Recital 26), because you hold the key that can re-identify it. Only properly anonymised data falls outside the regulation, and interview data is rarely fully anonymous.
When you upload a recording to a transcription service, that service processes personal data on your behalf. Under GDPR (Article 4(8) and Article 28) it is a processor, and you (or your university) are the controller. The controller decides why and how data is processed and stays responsible for it, which is why your choice of tool is part of your data management, not only a convenience.
Consent, lawful basis and ethics approval
Two different kinds of consent often get mixed up. Research ethics asks for informed consent to take part in the study. GDPR asks for a lawful basis to process personal data, and that basis is often not consent. Many public universities rely on "public interest" (Article 6(1)(e)), and for special categories such as health data on the research exemption in Article 9(2)(j), which depends on national law, with the safeguards of Article 89. Your university’s data protection officer will tell you which basis your institution uses.
Either way, participants must be told clearly how their recording will be handled. In practice your information sheet should say that the interview will be recorded, that it will be transcribed by an automated service, where that service processes data, how long recordings are kept, and how to withdraw.
- Example wording: "The interview will be audio recorded. The recording will be transcribed by an automated transcription service that processes data within the European Union and does not use it to train AI models. The recording will be deleted once the transcript has been checked, at the latest by [date]."
- If you plan to use AI for analysis as well, such as suggested codes or summaries, say so too.
- If your ethics application named a specific tool or said "manual transcription", update it before switching.
How to choose a GDPR-friendly transcription service
Accuracy is easy to test yourself on ten minutes of audio. The data questions are harder to see, so ask them directly and get the answers in writing.
| Question | What a good answer looks like | Why it matters |
|---|---|---|
| Where is the audio processed? | In the EU or EEA, with named subprocessors | Transfers outside the EEA need an adequacy decision or other safeguards under GDPR Chapter V |
| Is my data used to train models? | No, by default and in the contract | Training can make recordings part of a product you do not control |
| Is there a data processing agreement? | Yes, covering the transcription service | Required under Article 28 when a processor handles personal data |
| How long are files kept? | Only as long as you need them, deletable by you | Storage limitation (Article 5) and your own retention plan |
| Can it separate speakers? | Yes, with labels you can rename | Needed to pseudonymise and to analyse who said what |
| Can it redact personal details? | Optional masking of names and identifiers | Helps with pseudonymisation and data minimisation |
Test it on one of your own interviews
Upload 30 minutes of audio and compare the transcript with your current method. Processed in the EU, never used for training.
A GDPR-friendly workflow, step by step
This is the workflow we recommend, whichever tool you use.
- 1. Before recruiting: confirm the tool with your data protection officer or ethics committee, and write it into your information sheet and data management plan.
- 2. Recording: use a dedicated recorder or a phone in flight mode, not a cloud meeting tool you have not checked. Close microphones give fewer errors.
- 3. Transfer: move files to secure storage the same day. Do not leave recordings on personal phones or in email.
- 4. Transcription: upload to an EU-processed service. Turn on speaker labels and, if available, redaction of names and identifiers.
- 5. Check: listen through every passage you will quote and spot-check the rest. Names, numbers and dialect words are the usual errors.
- 6. Pseudonymise: replace names, places and workplaces with codes or roles (Participant 3, [hospital], [town]). Keep the key separately.
- 7. Delete: remove the original recording when your plan says so, often once the transcript is verified.
- 8. Report: in your methods section, name the transcription tool, how transcripts were checked, and how data was pseudonymised.
How to do it in Kahubi
Kahubi is a research workspace run by Avidemic AB in Sweden. Interview audio and video are transcribed by European providers (in France and Germany), AI processing runs with European providers, and your recordings and transcripts are never used to train models.
Upload recordings (up to 1 GB per file, audio or video) to an interview study. Kahubi transcribes them with speaker labels in 99+ languages and can optionally mask names and other personal details. You then rename speakers, check the transcript against the audio, and start analysis in the same project: AI-suggested codes you accept or rewrite, themes built from the codes you confirmed, and a findings draft where every quote links back to the transcript and timestamp.
Common mistakes to avoid
- Using a free consumer app whose terms allow training on your audio.
- Recording in a video meeting tool with automatic cloud transcription switched on without telling participants.
- Keeping raw recordings "just in case" with no deletion date.
- Sharing unpseudonymised transcripts with co-authors by email.
- Pasting full transcripts into a general chatbot that your institution has not approved.
Sources
- Regulation (EU) 2016/679 (General Data Protection Regulation). EUR-Lex.
- European Commission (2026). Living guidelines on the responsible use of generative AI in research, third version.
- European Data Protection Board (2025). Guidelines 01/2025 on pseudonymisation.
- Finnish Social Science Data Archive. Anonymisation and personal data.
- Otter.ai. Privacy policy.
Last updated 2026-10-09.