AI-assisted thematic analysis: is it valid, and how to do it well

AI can suggest codes in minutes, and studies show it often finds the same broad themes as human analysts. It also misses subtext and context, and leading qualitative methodologists have rejected it for reflexive work. Here is the evidence on both sides and a practical way to decide.

Updated 2026-10-09

The short answer

It depends on your approach. For codebook and framework approaches, where codes are applied consistently to a lot of data, AI suggestions checked by a human researcher can save time and are increasingly accepted when reported openly. For reflexive thematic analysis, where the researcher’s own interpretation is the method, many experienced qualitative researchers, including Virginia Braun and Victoria Clarke, argue that generative AI does not fit.

Either way, AI is a tool, not an analyst. The interpretation, and the responsibility for it, stays with you.

What studies comparing AI and human coding found

A growing number of studies have compared AI-generated codes and themes with those of human analysts. The pattern is consistent.

  • Prescott and colleagues (2024, JMIR AI) found that ChatGPT and Bard each identified 5 of 7 themes that human researchers had found in short text messages, in about 20 minutes compared with roughly 9.5 hours for the human team. Agreement on individual codes was much lower, and the AI took messages literally and missed subtle, interpretive meanings.
  • Bijker and colleagues (2024, Journal of Medical Internet Research) found good agreement between ChatGPT and human coders when it generated codes inductively from 537 forum posts, and weaker, more variable agreement when it had to apply an existing framework.
  • De Paoli (2024, Social Science Computer Review) took GPT-3.5 through the six phases of thematic analysis on two published interview datasets and found it could recover many of the main themes of the original analyses, while noting important limits.
  • Morgan (2023, International Journal of Qualitative Methods) found ChatGPT stronger on concrete, descriptive themes than on subtle, interpretive ones.
  • Dai and colleagues (2023) showed that an LLM-in-the-loop design, where the model suggests and humans decide, reached coding quality similar to human coders on survey data with less time and effort.

The case against AI in reflexive thematic analysis

In December 2025, Jowsey, Braun, Clarke, Lupton and Fine published an open letter in Qualitative Inquiry, signed by 419 experienced qualitative researchers from 32 countries, rejecting generative AI for reflexive qualitative research. Their argument is that reflexive analysis depends on a positioned, reflexive human researcher making meaning, and that AI use is not methodologically congruent with that. They also raise social and environmental justice concerns.

In March 2026, Friese, Nguyen-Trung, Powell and Morgan replied in the same journal, arguing that a blanket rejection closes off careful, critical use and turns the debate into pro versus anti. Both papers are worth reading before you decide.

Where AI genuinely helps qualitative researchers

Even researchers who do not want AI near their interpretation often use it for the mechanical parts of a project.

  • Transcription, which takes several hours per hour of audio when done by hand.
  • Pseudonymisation support, flagging names and places to replace.
  • A first pass of descriptive codes in large datasets, for codebook or framework analysis.
  • Checking coverage: asking which transcripts mention a code you defined, so nothing is missed.
  • Finding quotes: retrieving the passages behind a theme you have already developed.
  • Acting as a sounding board that challenges your reading.

How to use AI in qualitative analysis responsibly

  • Choose your approach first, and check whether AI fits its assumptions. Do not use AI just because it is available.
  • Read your data yourself. Familiarisation cannot be delegated.
  • Treat every AI code as a suggestion. Accept, rewrite or reject it, and keep a record of which you did.
  • Watch for what AI misses: irony, slang, silence, culture and context.
  • Use a tool that processes data in the EU and does not train on it, and check that your ethics approval covers it.
  • Report what you did. The COREQ+LLM checklist, published as a preprint in October 2026, extends COREQ with items on LLM use in qualitative research.

How Kahubi handles qualitative analysis

Kahubi is built for the human-in-the-loop model. Interviews are transcribed in the EU, and the AI suggests codes that you accept, rewrite or reject. Themes are built only from codes you confirmed, and every code links back to the passage and timestamp. The record of what was suggested and what you decided makes your methods section easy to write. If your approach rules out AI coding, you can still use Kahubi for transcription and pseudonymisation and do the coding yourself.

Sources

  1. Jowsey T, Braun V, Clarke V, Lupton D, Fine M (2025). We reject the use of generative artificial intelligence for reflexive qualitative research. Qualitative Inquiry.
  2. Friese S, Nguyen-Trung K, Powell S, Morgan DL (2026). Beyond binary positions: making space for critical and reflexive GenAI integration in qualitative research. Qualitative Inquiry.
  3. Prescott MR, et al. (2024). Comparing the efficacy and efficiency of human and generative AI: qualitative thematic analyses. JMIR AI 3:e54482.
  4. Bijker R, Merkouris SS, Dowling NA, Rodda SN (2024). ChatGPT for automated qualitative research: content analysis. Journal of Medical Internet Research 26:e59050.
  5. De Paoli S (2024). Performing an inductive thematic analysis of semi-structured interviews with a large language model. Social Science Computer Review 42(4):997-1019.
  6. Morgan DL (2023). Exploring the use of artificial intelligence for qualitative data analysis: the case of ChatGPT. International Journal of Qualitative Methods 22.
  7. Dai SC, Xiong A, Ku LW (2023). LLM-in-the-loop: leveraging large language model for thematic analysis. Findings of EMNLP 2023.
  8. Fehring L, et al. (2026). Reporting of qualitative research using large language models (COREQ+LLM): a Delphi-based extension of the COREQ reporting guideline. medRxiv preprint.
  9. Christou PA (2023). How to use artificial intelligence (AI) as a resource, methodological and analysis tool in qualitative research? The Qualitative Report 28(7).

Last updated 2026-10-09.

Frequently asked questions

Can I use ChatGPT for thematic analysis?
For codebook or framework approaches, AI suggestions reviewed by a human can save time and are increasingly accepted when reported. For reflexive thematic analysis, many methodologists including Braun and Clarke argue against it. Do not paste participant data into a consumer chatbot your institution has not approved.
What do Braun and Clarke say about AI?
In December 2025 they co-authored an open letter in Qualitative Inquiry, signed by 419 qualitative researchers, rejecting generative AI for reflexive qualitative research because reflexive analysis relies on a positioned human researcher.
How accurate is AI at coding qualitative data?
Studies find that AI often identifies the same broad, descriptive themes as human analysts, much faster, but agrees less on individual codes and misses subtext, tone and cultural meaning. Human review of every code is essential.
How do I report AI use in qualitative research?
Name the tool and version, describe exactly what it did (for example suggested initial codes) and how researchers reviewed it, and reflect on its influence. COREQ+LLM, a 2026 preprint, offers a 31-item checklist for this.

Related

Qualitative analysis with you in charge

EU transcription, AI-suggested codes you confirm, and every quote linked to the recording. 30 minutes of transcription free.