How to analyze interview data

Transcribe, code, theme, report — the qualitative analysis pipeline explained, with the discipline that separates real thematic analysis from vibes with quotes.

1. Transcribe with intention

Decide the transcription level before you start: verbatim (every um and repair — needed for conversation analysis), intelligent verbatim (cleaned but complete — the default for thematic analysis), or summary notes (rarely defensible for research). Modern speech-to-text handles the first pass; your job shifts to correcting domain terms and attributing speakers.

Anonymize as you transcribe, not after: replace names with roles ([Clinic manager]) while context is fresh. If your data contain personal or health information, check where your transcription service processes audio — EU processing is a common ethics-board requirement.

2. Pick an analysis framework and commit

Thematic analysis (Braun & Clarke’s six phases) is the workhorse: familiarization, initial codes, candidate themes, review, definition, write-up. Grounded theory suits theory-building from scratch with its constant-comparison and theoretical-sampling machinery. Narrative analysis preserves individual stories rather than fragmenting them into codes. IPA goes deep on few participants’ lived experience.

The common failure is drift: starting "thematic analysis" and quietly doing content counting, or claiming grounded theory after a single coding pass. Name the framework in your methods section, cite it, and follow its actual steps — reviewers who know the framework will check.

3. Code systematically, not impressionistically

A code is a short label for a meaningful segment: "distrust of head office", "workaround pride". First-pass coding should stay close to the data (semantic codes); interpretation comes when you group codes into themes. Keep a codebook — code name, definition, an example quote — and update it as codes merge and split.

For team coding, measure agreement early: double-code two transcripts, compare, discuss disagreements, refine the codebook, then split the rest. Report the process; some journals also want an agreement statistic.

4. Build themes, then stress-test them

A theme is not a topic ("communication") but a claim about the data ("informal channels compensate for broken formal reporting"). Candidate themes come from clustering codes; the review phase is where rigor lives — for each theme, re-read every coded extract and ask whether it genuinely supports the claim, and actively hunt for disconfirming cases.

Aim for three to six themes in a paper. More than that usually means some are topics in disguise.

5. Report against a standard

COREQ (32 items, interviews and focus groups) and SRQR are the checklists journals ask for. They want to see who interviewed, the guide, saturation reasoning, coding process, and quotes identified by participant. Write the methods section as if a skeptical reviewer will reconstruct your audit trail — because a good one will.

Where Kahubi fits

Kahubi’s interview study flow transcribes in 99+ languages with EU processing, supports speaker labels and PHI redaction, then runs framework-guided analysis: AI-suggested codes you accept or rewrite, segments linked back to the transcript, themes built from your confirmed codes, and a QDA report draft with the audit trail intact.

Last updated 2026-07-09.

Frequently asked questions

How many interviews do I need?
Saturation-based guidance lands at 9–17 interviews for homogeneous samples in most published tests; IPA studies run smaller (4–10). What convinces reviewers is not the number but showing that late interviews stopped producing new codes.
Can AI do qualitative coding?
AI can propose codes and locate candidate segments fast, which removes the mechanical burden. The interpretive act — deciding what counts, merging codes, naming themes — must stay with the researcher, and your methods section should say exactly where the line was.
What is the difference between codes and themes?
Codes label individual data segments; themes are patterns across many codes that make a claim about the dataset. Typically dozens or hundreds of codes collapse into 3–6 themes.

Related

From audio to audited themes

Transcription, coding and thematic analysis with the researcher in charge. Free plan included.