How to anonymise interview transcripts

Removing names is not enough. People can be recognised from their job, their town or a single unusual story. This guide explains the difference between pseudonymisation and anonymisation, the techniques data archives recommend, and what AI redaction can and cannot do.

Updated 2026-10-09

The short answer

During analysis, pseudonymise: replace names and identifying details with consistent codes or categories, and keep the key separately and securely. Before sharing or archiving, go further and remove or generalise anything that could identify someone in combination with other information. Mark every change in brackets and keep a log.

True anonymisation of interview data is hard, so plan for it from the start and be honest with participants about what you can promise.

Pseudonymisation is not anonymisation

Under GDPR, pseudonymised data is data that can no longer be linked to a person without extra information that is kept separately, such as your list of real names and codes. It is still personal data, so GDPR still applies. Anonymous data, where nobody can reasonably identify the person by any means likely to be used, falls outside GDPR.

The European Data Protection Board adopted draft guidelines on pseudonymisation for public consultation in January 2025, and in July 2026 published draft guidelines on anonymisation for consultation until 30 October 2026. The Court of Justice of the EU also ruled in September 2025 (EDPS v SRB, under rules for EU bodies that mirror GDPR) that pseudonymised data can be personal data for the organisation holding the key while not being personal data for a recipient who has no realistic way to re-identify anyone.

Know your identifiers

The Finnish Social Science Data Archive (FSD) groups identifiers into three levels. It is a useful checklist for any transcript.

LevelExamplesWhat to do
Direct identifiersName, personal ID number, voice, face, emailAlways remove or replace
Strong indirect identifiersAddress, phone number, unusual job title, rare disease, a position held by only one personUsually remove or generalise
Indirect identifiersAge, gender, occupation, region, dates, workplaceGeneralise when they could combine to identify someone

Techniques that work

These techniques come from the guidance of European and UK data archives.

  • Plan ahead. Tell participants how their data will be anonymised, and think about which details your analysis really needs.
  • Replace, do not blank out. "[colleague]" or "[small town in northern Sweden]" keeps the meaning; "XXXX" loses it.
  • Use consistent pseudonyms across transcripts, recordings and field notes.
  • Generalise: exact age becomes an age band, a specific hospital becomes "[regional hospital]".
  • Remove details that the analysis does not need.
  • Mark every change in square brackets so readers know the text was edited.
  • Keep an anonymisation log of original and replacement text, stored separately from the data.
  • Do not over-anonymise. Removing too much destroys the value of the data.
  • Remove metadata from files, such as author names and timestamps.

Test it: the motivated intruder

A practical test from the UK Information Commissioner’s Office is to imagine a "motivated intruder": a reasonably competent person with internet access who wants to identify a participant and can ask around. Could they work out who this is? Pay special attention to small communities, workplaces and rare experiences, where one detail can be enough.

This matters more now that AI exists. Staab and colleagues (2024) showed that language models can infer personal attributes such as location, income and sex from ordinary text with up to 85% accuracy. A 2026 study showed that an AI agent with web search could link 6 of 24 scientist interviews in a public interview dataset to specific published papers, in some cases identifying the interviewee.

What automated and AI redaction can do

Automated tools can speed up the work but cannot be trusted on their own. A 2024 study found that GPT-4 found almost all personal identifiers in forum posts but also flagged many things that were not identifiers. In 2025 FSD tested Microsoft Copilot tools on interview transcripts and found that some versions missed names and places, results varied between runs, and human review was still needed.

The best workflow combines both: let a tool flag candidate identifiers, then read every transcript yourself and decide.

Pseudonymising transcripts in Kahubi

When you transcribe interviews in Kahubi you can switch on redaction, which masks names and other personal details with consistent markers. You rename speakers to pseudonyms, and the transcript stays linked to the recording so you can check what was masked. Transcription and AI processing run with European providers and your data is never used for training. Treat redaction as a first pass: read the transcript and generalise context-based details yourself.

Sources

  1. Regulation (EU) 2016/679 (GDPR), Article 4(5) and Recital 26. EUR-Lex.
  2. European Data Protection Board (2025). Guidelines 01/2025 on pseudonymisation.
  3. European Data Protection Board (2026). Guidelines 02/2026 on anonymisation (public consultation).
  4. Court of Justice of the EU (2025). EDPS v Single Resolution Board, Case C-413/23 P. Press release No 107/25.
  5. Finnish Social Science Data Archive. Anonymisation and personal data.
  6. Finnish Social Science Data Archive (2025). AI for verifying text file anonymisation.
  7. UK Data Service. Anonymisation for text data.
  8. Information Commissioner’s Office. Anonymisation guidance: how do we ensure anonymisation is effective? (motivated intruder test).
  9. Singhal S, Zambrano AF, Pankiewicz M, Liu X, Porter C, Baker RS (2024). De-identifying student personally identifying information with GPT-4. Educational Data Mining 2024.
  10. Staab R, Vero M, Balunović M, Vechev M (2024). Beyond memorization: violating privacy via inference with large language models.
  11. Li T (2026). Agentic LLMs as powerful deanonymizers: re-identification of participants in the Anthropic Interviewer dataset. arXiv:2601.05918.

Last updated 2026-10-09.

Frequently asked questions

What is the difference between anonymisation and pseudonymisation?
Pseudonymisation replaces identifiers with codes while a key still exists, so the data remains personal data under GDPR. Anonymisation removes the possibility of identifying anyone by any means reasonably likely to be used, so GDPR no longer applies.
Is removing names enough to anonymise a transcript?
No. Jobs, places, dates, rare conditions and unusual stories can identify people, especially in small communities. Generalise or remove these too, and apply the motivated intruder test.
Can AI anonymise transcripts for me?
AI can flag likely identifiers quickly, but studies show it misses some and over-flags others, and results vary between runs. Use it as a first pass and review every transcript yourself.
When should I anonymise interview data?
Pseudonymise as soon as transcripts are checked, before analysis and sharing. Anonymise more thoroughly before archiving or publishing data. For long-term studies, plan when to anonymise so that later waves can still be linked.
Should I keep the key that links pseudonyms to names?
Only if you need it, for example for follow-up interviews or withdrawal requests. Store it separately and securely, and delete it when your data management plan says so.

Related

Transcribe and pseudonymise in the EU

Optional redaction of names and personal details, speaker pseudonyms, and every passage linked to the recording. 30 minutes free.