How to transcribe research interviews with AI, GDPR compliant

AI transcription saves days of work per study, but an interview recording is personal data about someone who trusted you. This guide covers what GDPR asks of you, how to choose a service, and a workflow from consent form to coded transcript.

Updated 2026-10-09

The short answer

You can use AI to transcribe research interviews under GDPR. Most of the work is the same as for any research data: a lawful basis, clear information to participants, a processor agreement with the service, and good security. The choice of service matters most for three questions: where the audio is processed, whether the provider trains on it, and how long it keeps it.

  • Tell participants in the information sheet and consent form that recordings will be transcribed by an automated service, and what kind.
  • Pick a service that processes audio in the EU or EEA and does not train its models on your recordings.
  • Make sure a data processing agreement (DPA) covers the service. At universities this is usually signed centrally.
  • Pseudonymise transcripts before analysis, and delete recordings once they are no longer needed.
  • If the interviews touch health, ethnicity, religion, politics, sexuality or other special categories, check that your ethics approval covers the tool.

Why GDPR applies to interview recordings

A voice recording of an identifiable person is personal data. So is a transcript that contains names, workplaces, places or stories that could identify someone. For you as the researcher, pseudonymised data is still personal data under GDPR (Recital 26), because you hold the key that can re-identify it. Only properly anonymised data falls outside the regulation, and interview data is rarely fully anonymous.

When you upload a recording to a transcription service, that service processes personal data on your behalf. Under GDPR (Article 4(8) and Article 28) it is a processor, and you (or your university) are the controller. The controller decides why and how data is processed and stays responsible for it, which is why your choice of tool is part of your data management, not only a convenience.

Consent, lawful basis and ethics approval

Two different kinds of consent often get mixed up. Research ethics asks for informed consent to take part in the study. GDPR asks for a lawful basis to process personal data, and that basis is often not consent. Many public universities rely on "public interest" (Article 6(1)(e)), and for special categories such as health data on the research exemption in Article 9(2)(j), which depends on national law, with the safeguards of Article 89. Your university’s data protection officer will tell you which basis your institution uses.

Either way, participants must be told clearly how their recording will be handled. In practice your information sheet should say that the interview will be recorded, that it will be transcribed by an automated service, where that service processes data, how long recordings are kept, and how to withdraw.

  • Example wording: "The interview will be audio recorded. The recording will be transcribed by an automated transcription service that processes data within the European Union and does not use it to train AI models. The recording will be deleted once the transcript has been checked, at the latest by [date]."
  • If you plan to use AI for analysis as well, such as suggested codes or summaries, say so too.
  • If your ethics application named a specific tool or said "manual transcription", update it before switching.

How to choose a GDPR-friendly transcription service

Accuracy is easy to test yourself on ten minutes of audio. The data questions are harder to see, so ask them directly and get the answers in writing.

QuestionWhat a good answer looks likeWhy it matters
Where is the audio processed?In the EU or EEA, with named subprocessorsTransfers outside the EEA need an adequacy decision or other safeguards under GDPR Chapter V
Is my data used to train models?No, by default and in the contractTraining can make recordings part of a product you do not control
Is there a data processing agreement?Yes, covering the transcription serviceRequired under Article 28 when a processor handles personal data
How long are files kept?Only as long as you need them, deletable by youStorage limitation (Article 5) and your own retention plan
Can it separate speakers?Yes, with labels you can renameNeeded to pseudonymise and to analyse who said what
Can it redact personal details?Optional masking of names and identifiersHelps with pseudonymisation and data minimisation

Test it on one of your own interviews

Upload 30 minutes of audio and compare the transcript with your current method. Processed in the EU, never used for training.

Transcribe 30 minutes free

A GDPR-friendly workflow, step by step

This is the workflow we recommend, whichever tool you use.

  • 1. Before recruiting: confirm the tool with your data protection officer or ethics committee, and write it into your information sheet and data management plan.
  • 2. Recording: use a dedicated recorder or a phone in flight mode, not a cloud meeting tool you have not checked. Close microphones give fewer errors.
  • 3. Transfer: move files to secure storage the same day. Do not leave recordings on personal phones or in email.
  • 4. Transcription: upload to an EU-processed service. Turn on speaker labels and, if available, redaction of names and identifiers.
  • 5. Check: listen through every passage you will quote and spot-check the rest. Names, numbers and dialect words are the usual errors.
  • 6. Pseudonymise: replace names, places and workplaces with codes or roles (Participant 3, [hospital], [town]). Keep the key separately.
  • 7. Delete: remove the original recording when your plan says so, often once the transcript is verified.
  • 8. Report: in your methods section, name the transcription tool, how transcripts were checked, and how data was pseudonymised.

How to do it in Kahubi

Kahubi is a research workspace run by Avidemic AB in Sweden. Interview audio and video are transcribed by European providers (in France and Germany), AI processing runs with European providers, and your recordings and transcripts are never used to train models.

Upload recordings (up to 1 GB per file, audio or video) to an interview study. Kahubi transcribes them with speaker labels in 99+ languages and can optionally mask names and other personal details. You then rename speakers, check the transcript against the audio, and start analysis in the same project: AI-suggested codes you accept or rewrite, themes built from the codes you confirmed, and a findings draft where every quote links back to the transcript and timestamp.

Common mistakes to avoid

  • Using a free consumer app whose terms allow training on your audio.
  • Recording in a video meeting tool with automatic cloud transcription switched on without telling participants.
  • Keeping raw recordings "just in case" with no deletion date.
  • Sharing unpseudonymised transcripts with co-authors by email.
  • Pasting full transcripts into a general chatbot that your institution has not approved.

Sources

  1. Regulation (EU) 2016/679 (General Data Protection Regulation). EUR-Lex.
  2. European Commission (2026). Living guidelines on the responsible use of generative AI in research, third version.
  3. European Data Protection Board (2025). Guidelines 01/2025 on pseudonymisation.
  4. Finnish Social Science Data Archive. Anonymisation and personal data.
  5. Otter.ai. Privacy policy.

Last updated 2026-10-09.

Frequently asked questions

Is it legal to use AI to transcribe research interviews in the EU?
Yes, if you follow GDPR: a lawful basis for processing, clear information to participants, a data processing agreement with the service, appropriate security, and limits on how long you keep data. Choosing a service that processes data in the EU and does not train on it makes this much simpler.
Do I need participants’ consent to transcribe their interview with AI?
Participants must consent to take part and be informed about how their recording is handled, including automated transcription. The GDPR lawful basis may be public interest rather than consent, depending on your institution. Ask your data protection officer which applies.
Is Otter.ai GDPR compliant for research?
Otter processes data in the United States and its privacy policy lists training its AI on de-identified audio recordings and transcripts. Many European universities do not approve it for research interviews. Check your institution’s list of approved tools.
Which transcription tools process data in the EU?
Examples include Kahubi (France and Germany), Amberscript (data stored in Germany) and university services such as Sunet Scribe in Sweden. Always confirm the current processing locations and subprocessors on the provider’s own pages.
Should I anonymise or pseudonymise interview transcripts?
Pseudonymise during analysis: replace names and identifying details with codes and keep the key separately. Full anonymisation is hard for interview data because stories themselves can identify people. Pseudonymised data is still personal data under GDPR, so it still needs protection.
How should I describe AI transcription in my methods section?
Name the tool, say where data was processed, describe how transcripts were checked against the audio (for example every quoted passage plus a 10% spot-check), and how data was pseudonymised. This matches the EU living guidelines’ call to be transparent about substantial AI use.

Related

Transcribe interviews in the EU

Speaker labels, 99+ languages, optional redaction, never used for training. Then code themes in the same project. 30 minutes free.