Can ChatGPT Really Do Thematic Analysis? Here's What the Research Found

You paste your interview transcripts into ChatGPT, ask for themes, and get a tidy list back in ten seconds. It reads well. It sounds confident. The question is whether it holds up in front of a dissertation committee or a peer reviewer, and on that question, the published research is now specific enough to give a real answer.

Researcher typing interview transcripts into ChatGPT on a laptop for AI-assisted qualitative analysis
Photo by Matheus Bertelli on Pexels.

The short answer

ChatGPT can generate a useful first-pass codebook, but on its own it tends to produce topic summaries rather than genuine themes, and every controlled comparison against human coders published so far has found gaps in interpretive depth, consistency, or both (Lee et al., 2024; Wachinger et al., 2025; Naeem, Smith and Thomas, 2025). It is a real assistant. It is not, at this point, a substitute for the researcher.

What “doing thematic analysis” actually requires

Thematic analysis, as defined by Braun and Clarke (2006), is not pattern-matching against a transcript. It's an active, interpretive process where the researcher builds themes around a central organising concept, shaped by their own theoretical position and engagement with the data (Braun and Clarke, 2019). That distinction, between describing what people said and interpreting what it means, is exactly where the research on ChatGPT keeps landing.

If you haven't already, our guide on the difference between a code and a theme covers this distinction in full. It matters here because it's the single most common failure mode reported across every study of ChatGPT-assisted analysis.

What the studies actually found

ChatGPT produces topics, not themes, by default. A 2024 viewpoint published in the Journal of Medical Internet Research had ChatGPT-3.5 and ChatGPT-4.0 code a real interview transcript and generate themes from a set of 81 codes (Lee et al., 2024). Both models identified five themes, matching the human analyst's theme count, but the actual themes differed: ChatGPT surfaced a “diet and nutrition management” theme the human analyst never coded as a theme or sub-theme at all, and repeating the same prompt produced different themes on different runs. The researchers describe ChatGPT as capable of “enhancing the efficiency” of analysis, while stressing that a human researcher has to review and consolidate everything it produces.

Comparative studies against human coders show real gaps, not just style differences. Wachinger et al. (2025), comparing ChatGPT directly against a human researcher on the same qualitative dataset, found ChatGPT competent at descriptive, semantic-level coding but noted it would “convincingly argue connections” to a theoretical framework even when the fit was poor. Prescott et al. (2024), in a separate comparison published in JMIR AI, found ChatGPT's themes matched human-generated themes 71% of the time under inductive coding but only 50–58% under deductive coding, where a pre-existing framework has to be applied faithfully rather than pattern-matched loosely.

Hallucination is a named, recurring problem, not an edge case. Multiple independent studies report the same failure: ChatGPT assigning a code to a quote that doesn't support it, inventing a connection the transcript doesn't contain, or generating a theme that sounds plausible but isn't grounded in the data (Lee et al., 2024; De Paoli, 2023). This is the exact concern Cook et al. (2025) raise in Academic Medicine when they argue that a human must remain “in charge of the loop,” not simply present as a rubber stamp on AI output.

Context gets lost across a long analysis. ChatGPT's token limits mean long transcripts often have to be split into chunks, and the model doesn't reliably retain earlier codes or context once a conversation moves on (Lee et al., 2024). Practically, this means asking it to stay consistent across a 20-transcript dataset in one continuous chat session is asking it to do something the architecture isn't built to do.

Why this happens: the topics vs. themes problem

This is worth dwelling on because it's the single most repeated finding across the literature, and it maps directly onto a distinction Braun and Clarke themselves have spent years warning researchers not to skip. A topic summary groups content by subject: “participants discussed remote work.” A theme makes an interpretive claim about what a pattern of content means: “remote work blurred the boundary between professional competence and personal availability.” ChatGPT's defaults lean heavily toward the first, because generating a topic requires only recognising that a subject recurred, while generating a genuine theme requires the kind of situated, reflexive interpretation Braun and Clarke (2019) argue is inherently the researcher's job, not a pattern a model can extract from text alone.

So is it worth using at all?

Yes, with real conditions attached. Every study cited above reaches the same practical conclusion: ChatGPT is genuinely useful as an assistant for the mechanical, early-stage work, generating an initial pool of codes, summarising a transcript, or acting as a second opinion to check against your own reading, but it needs a human reviewing, correcting, and taking final responsibility for every output (Lee et al., 2024; Naeem, Smith and Thomas, 2025). If you want the fuller methodological picture, including how to disclose this kind of AI use properly, see our guides on AI thematic analysis and academic validity and how to report AI-assisted thematic analysis in your methods section.

The practical problem is that a general-purpose chat tool was never built for this specific workflow. It has no persistent codebook, no way to link a generated theme back to the exact transcript passage it came from, and no memory of what you accepted or rejected in your last session. Every one of those is a rebuild-from-scratch problem the next time you open a new chat.

Try a structured workflow instead of a blank chat window

If the appeal of ChatGPT was speed, thematicanalysis.ai/analyze is built around the same idea with the missing structure attached: upload your transcripts and get an initial codebook and candidate theme clusters generated in minutes, with every suggestion linked back to its source passage so you can verify it rather than take it on faith. Try it on your own transcripts and compare the output directly against a plain ChatGPT prompt before deciding anything.

ChatGPT vs. a purpose-built thematic analysis workflow

CapabilityPlain ChatGPTthematicanalysis.ai
Codebook persistenceResets with every new chatSaved and editable across your whole project
Source traceabilityNo built-in link back to transcriptEvery code and theme links to its source extract
Consistency across a datasetDegrades as context window fillsMaintained across the full project
Grounded in Braun & Clarke's six phasesOnly if you prompt for it explicitly, and inconsistentlyStructured around the six-phase framework by default
Audit trail for methods reportingYou'd have to build one manuallyExportable alongside your analysis

The bottom line

The research is consistent: ChatGPT can speed up the early, mechanical stages of thematic analysis, but it defaults to descriptive topic summaries rather than interpretive themes, its output needs verification against the source data every time, and it isn't built to hold a codebook steady across a real research project. Used as a supervised assistant with a human firmly in charge of the interpretation, it's a reasonable place to start. Used as a replacement for the analysis itself, the published evidence doesn't support it yet.

References

  • Braun, V. and Clarke, V. (2006) 'Using thematic analysis in psychology', Qualitative Research in Psychology, 3(2), pp. 77–101. Available at: https://doi.org/10.1191/1478088706qp063oa
  • Braun, V. and Clarke, V. (2019) 'Reflecting on reflexive thematic analysis', Qualitative Research in Sport, Exercise and Health, 11(4), pp. 589–597. Available at: https://doi.org/10.1080/2159676X.2019.1628806
  • Cook, D.A., Ginsburg, S., Sawatsky, A.P., Kuper, A. and D'Angelo, J.D. (2025) 'Artificial intelligence to support qualitative data analysis: promises, approaches, pitfalls', Academic Medicine, 100(10), pp. 1134–1149. Available at: https://pubmed.ncbi.nlm.nih.gov/40560241/
  • De Paoli, S. (2023) 'Performing an inductive thematic analysis of semi-structured interviews with a large language model: an exploration and provocation on the limits of the approach', Social Science Computer Review. Available at: https://doi.org/10.1177/08944393231220483
  • Lee, V.V., van der Lubbe, S.C.C., Goh, L.H. and Valderas, J.M. (2024) 'Harnessing ChatGPT for thematic analysis: are we ready?', Journal of Medical Internet Research, 26, e54974. Available at: https://doi.org/10.2196/54974
  • Naeem, M., Smith, T. and Thomas, L. (2025) 'Thematic analysis and artificial intelligence: a step-by-step process for using ChatGPT in thematic analysis', International Journal of Qualitative Methods, 24. Available at: https://doi.org/10.1177/16094069251333886
  • Prescott, M.R., Yeager, S., Ham, L., Rivera Saldana, C.D., Serrano, V., Narez, J., Paltin, D., Delgado, J., Moore, D.J. and Montoya, J. (2024) 'Comparing the efficacy and efficiency of human and generative AI: qualitative thematic analyses', JMIR AI, 3, e54482. Available at: https://doi.org/10.2196/54482
  • Wachinger, J., Bärnighausen, K., Schäfer, L.N., Scott, K. and McMahon, S.A. (2025) 'Prompts, pearls, imperfections: comparing ChatGPT and a human researcher in qualitative data analysis', Qualitative Health Research, 35(9), pp. 951–966. Available at: https://doi.org/10.1177/10497323241244669