AI and qualitative methods

AI for Qualitative Coding: Using It on Interviews and Surveys Without Losing Interpretive Authority

A recent thread on r/QualitativeResearch argued that handing coding to software and chatbots leads to thin analysis. The researchers in it would rather code by hand, because working through every line is how they get close to the data, and because they don't want a model deciding what their participants meant. They have a point. So do the researchers on the other side, who use AI every week and still consider the analysis their own. This guide explains how AI coding tools work on interview transcripts and survey responses, what the published evidence says, the strongest version of the case against them, and how to use them, if you choose to, without handing over your interpretive authority.

A student researcher pausing to think with a highlighter over her notes and laptop while reviewing qualitative data
Photo by Zen Chung on Pexels.

Can you use AI to code qualitative data?

Yes, as an assistant, with limits. Studies comparing AI and human coding find that AI produces a reasonable first pass of descriptive codes quickly but misses meaning that depends on context, emotion or culture (Morgan, 2023; Hamilton et al., 2023). Hybrid workflows, where a researcher reviews and reworks the AI's output, have produced better codebooks than either humans or AI working alone (Barany et al., 2024).

It is not settled for every method. In 2025, 419 qualitative researchers, including Braun and Clarke, rejected generative AI for reflexive approaches at every stage, including initial coding (Jowsey et al., 2025). Whatever tool you use, interpretation stays with the researcher: you read the data, check every code against its quote, and decide what it means.

What AI coding tools actually do

“AI coding” covers three quite different things, and a lot of the online argument comes from people meaning different ones.

  • General chatbots such as ChatGPT or Claude. You paste text and ask for codes or themes. Flexible, but the output depends heavily on your prompt, long transcripts get cut off, and nothing links a code back to the passage it came from unless you ask. Our guide to whether ChatGPT can do thematic analysis covers this route.
  • AI features inside CAQDAS software such as NVivo, ATLAS.ti and MAXQDA's AI Assist. These suggest codes or summaries inside a project you are already coding by hand.
  • Purpose-built analysis tools such as thematicanalysis.ai that take your transcripts or survey export, produce codes with definitions and supporting quotes, and group them following a named method.

All three work the same way underneath. A large language model reads the text and predicts which labels best describe each passage, based on patterns in the enormous amount of text it was trained on. It does not know your participants, your field or your research question beyond what you tell it. That is the root of both its usefulness (it has seen a lot of language) and its limits (it has seen very little of your context).

Using AI to code interview transcripts

Interviews are where the stakes are highest, because the meaning is often in what people imply, hesitate over or come back to twenty minutes later. A workflow that keeps you in charge looks like this:

  1. Read before anything is coded. Read all your transcripts yourself, or at the very least a good sample, before you see any AI output. Familiarisation is the closeness the hand-coders are defending, and no tool does it for you.
  2. Code one or two transcripts by hand first. This gives you your own sense of what matters, and a benchmark to judge the AI against. If you are new to coding, our guide to coding an interview transcript walks through it.
  3. Run the AI pass. Ask for codes with definitions and the exact quote behind each one. A code without its quote can't be checked.
  4. Audit line by line. For every code, read the passage. Keep it, rename it, move it, or delete it. Then read the transcript for passages the AI skipped.
  5. Compare with your hand-coded transcripts. Where you disagree with the AI, you have learned something about your data or about the tool. Write it down.
  6. Build the themes yourself. AI suggestions for groupings can be a prompt to think, but deciding what the patterns mean, and what story they tell about your research question, is the analysis.

Umer and colleagues (2026) used almost exactly this workflow on Roman Urdu interviews with people with bipolar disorder and their carers. ChatGPT produced structured codes quickly, but a bilingual researcher auditing every code line by line found it flattened emotional and cultural meaning. It read a phrase about caring for the sick simply as “patient care,” missing the unspoken expectations placed on daughters-in-law. The AI saved time. The researcher supplied the analysis.

Using AI to code open-ended survey responses

Open-ended survey answers are a better fit for AI assistance, for two reasons. The volume is often large (hundreds or thousands of responses), and each response is short, so the task is closer to sorting content into categories than to interpreting a life story. Even the critics draw this distinction. Jowsey et al. (2025) contrast reflexive thematic analysis with word-counting forms of content analysis, which they note can be automated.

The risk with surveys is different: AI tends to amplify the common answer. The letter from Jowsey et al. warns that models are predisposed to reproduce dominant language and patterns, which risks quieting marginal voices. In a survey that can mean the one response that reframes the whole question gets folded into a generic category. So for surveys, spend your review time on the outliers and on responses the AI placed in catch-all categories like “other” or “general comments.” Our guide to analysing open-ended survey responses covers the rest of the process.

AI drafts the codes, you keep the interpretation: an AI first pass of codes on the left and the researcher's review on the right, with codes renamed, kept or rewritten. Start free with thematicanalysis.ai for your first 3 participants.

thematicanalysis.ai was built for exactly this kind of review. It drafts codes from your interview transcripts or survey responses, and every code comes with a definition and the verbatim quote it came from, so you can check each one against the data before it goes anywhere near your findings. You edit, move or remove codes and rename themes, and the interpretation stays with you. If you want to see how its first pass compares with your own, the first three participants are free.

What the research shows

The studies comparing AI and human coding point the same way more often than the online debate suggests.

Studies comparing AI-assisted and human qualitative coding
StudyWhat they testedWhat they found
Morgan (2023)ChatGPT themes compared with an earlier manual analysisGood at concrete, descriptive themes; struggled with subtle, interpretive ones
Hamilton et al. (2023)ChatGPT vs human coders on guaranteed income interviewsAI faster and scalable; humans more transparent, nuanced and sensitive to context
De Paoli (2024)Inductive thematic analysis of semi-structured interviews with an LLMAn LLM can carry out parts of an inductive analysis, but with clear limits, and writing prompts that work is neither easy nor obvious
Barany et al. (2024)Human-only, AI-only and two hybrid ways of building a codebookHybrids rated best and applied most reliably; the AI-only codebook ranked lowest
Umer et al. (2026)ChatGPT coding Roman Urdu mental health interviews, audited by a bilingual researcherEfficient surface coding; missed emotional, idiomatic and cultural meaning; human oversight essential

Two findings recur. First, AI is reliable on what the text says and unreliable on what it means. Second, the best results come when a person and a model each do what they are good at. Barany and colleagues found the fully human codebook missed some themes the hybrid approaches caught, and the fully automated one was the outlier on every measure. For a wider look at validity, see whether AI thematic analysis is valid for academic research.

The case against: interpretive authority

The Reddit thread's worry has a serious academic version. In December 2025, Qualitative Inquiry published an open letter from Tanisha Jowsey, Virginia Braun, Victoria Clarke, Deborah Lupton and Michelle Fine, endorsed by 419 experienced qualitative researchers from 32 countries, rejecting generative AI for reflexive qualitative research. Their reasons are worth stating in full, because they are stronger than “AI makes mistakes.”

  • AI cannot make meaning. A language model predicts plausible text. It does not understand the world or the data, so at best it simulates analysis. Reflexive analysis is meaning-making, so it cannot be done by something that does not make meaning.
  • Qualitative research is a human practice. It is done by people, with people, for people. The authors argue that AI is inappropriate at every phase of reflexive analysis, including initial coding, because coding is already interpretation.
  • The harms are real. The environmental cost of data centres and the exploitation of data workers, often in the Global South, are ethical reasons to refuse, separate from the methodological ones.

Two further points deserve attention. The letter notes that even researchers who keep a human “in the loop” warn that wanting the AI to be reliable weakens our ability to critique it. That is automation bias, and it is the most practical risk for a tired student at midnight. Nguyen and Welch (2025), writing in Organizational Research Methods, add that uncritical use brings epistemic risks, because the essence of qualitative analysis is interpreting meaning, which they see as a human capability.

And the hand-coders' core claim is true. Working through every line is one of the ways researchers come to know their data. If a tool lets you skip reading your transcripts, you will know them less well, and it will show in your analysis and your viva.

The counter-argument

The letter drew published replies. Susanne Friese answered it directly in 2025, and in 2026 Friese, Kien Nguyen-Trung, Steve Powell and David Morgan followed with Beyond Binary Positions in Qualitative Inquiry. They accept the ethical concerns but contest the conclusion that AI is inherently incompatible with meaning-based research. Drawing on those replies, the evidence above and our own view, the case for researcher-led use runs like this.

  • The AI doesn't make meaning. The researcher does. The model proposes; the researcher decides. In Friese's 2025 reply, meaning emerges from the researcher working with the data and the tool, with the researcher holding interpretive authority throughout.
  • It is scaffolding, like tools we already accept. CAQDAS software, memos, diagrams and a colleague's second opinion all shape analysis without taking it over. Friese and colleagues argue that, used critically and under close researcher leadership, AI can sit in the same category.
  • Hand coding does not guarantee closeness. This one is our view rather than the papers'. Coding transcript eighteen of twenty at 1am is not intimacy with the data, it is endurance. Closeness comes from reading, re-reading, memoing and arguing with your own interpretation. A tool that removes clerical work can leave more time for that.
  • The evidence favours hybrids, not either extreme. The studies above keep finding that human review of AI output does better than AI alone, and in some cases catches things humans alone missed.
  • Abstinence is not the only ethical response. Friese argues that environmental and labour concerns call for responsible use and stronger regulation, not blanket refusal. You may or may not agree, but it is a position you can defend.

Where does that leave the Reddit thread? Both sides agree on more than it looks. The critics are right that AI cannot interpret and that uncritical use produces thin analysis. The defenders are right that a researcher who reads the data, audits every code and builds the themes themselves has not handed over interpretive authority. The disagreement that remains is about whether initial coding is already interpretation, and for reflexive approaches, that is a real methodological question you should settle with your supervisor rather than with a tool vendor.

Where the line sits

A practical way to split the work, for methods where AI assistance is acceptable:

Tasks suited to AI assistance compared with tasks the researcher should keep
Reasonable to hand to AIKeep for yourself
A first pass of descriptive codesReading and knowing the data
Pulling the quote behind each codeDeciding whether the code fits the quote
Spotting repeated wording across many responsesDeciding what the repetition means
Suggesting possible groupings to react toBuilding and naming themes
A second perspective to check your coding againstReflexivity, positionality and the final argument

Seven rules for keeping interpretive authority

  1. Read your data before the AI does. Familiarisation is not optional, and it is the part hand-coders are right to protect.
  2. Never accept a code without reading its quote. If a tool can't show you the passage behind a code, don't use that code.
  3. Code a sample by hand and compare. One or two transcripts is enough to calibrate how far to trust the tool on your data.
  4. Disagree on purpose. For each theme the AI suggests, ask what it leaves out and whose voice it smooths over. Look hardest at emotion, irony, culture and silence.
  5. Write memos about your changes. A record of what you kept, rewrote and rejected is your audit trail, and it is evidence of your interpretive work. Our guide to reflexivity explains how to use it.
  6. Build the themes yourself. Treat AI groupings as prompts, not findings. The difference between a code and a theme is where your analysis lives.
  7. Disclose exactly what you did. Name the tool, the stage, and how you checked it. Our guide to reporting AI-assisted analysis has wording for your methods section.

Which analysis methods it fits

How well AI-assisted coding fits common qualitative analysis methods
MethodFit with AI-assisted coding
Qualitative content analysisGood fit, especially for descriptive, manifest-content coding of large datasets
Codebook and framework approachesGood fit for applying and refining a codebook, with human checks
Reflexive thematic analysisContested. Braun and Clarke co-authored the 2025 rejection. Agree your approach with your supervisor first
Grounded theory, IPA, phenomenologyContested for the same reasons; initial coding is itself interpretive in these traditions

If you haven't chosen a method yet, our overview of thematic analysis frameworks and our comparison of thematic analysis software can help.

Frequently asked questions

Can AI code qualitative data?

AI can produce a usable first pass of descriptive codes for interview transcripts and open-ended survey responses. Studies such as Morgan (2023) and Hamilton et al. (2023) found it handles concrete, descriptive content well but misses subtler, interpretive meaning. The researcher still has to check every code against the data and do the interpretation.

Does using AI mean losing interpretive authority?

Only if you accept its output without scrutiny. Interpretive authority means the researcher decides what the data means. If you read the data yourself, check each AI-suggested code against its quote, change or reject codes that don't hold up, and build the themes yourself, the interpretation stays yours. Critics such as Jowsey et al. (2025) argue that for reflexive approaches even initial coding should stay human.

Is AI-assisted coding acceptable for reflexive thematic analysis?

It is contested. In a 2025 open letter, 419 qualitative researchers including Braun and Clarke rejected generative AI for reflexive thematic analysis at every phase, including initial coding. Others, such as Friese et al. (2026), argue that researcher-led use is compatible with reflexive work. If your study uses reflexive TA, agree the approach with your supervisor before you start.

Is AI better for survey responses or interview transcripts?

It is generally better suited to large sets of short open-ended survey responses, where the task is closer to categorising content and the volume makes hand coding slow. Long interview transcripts carry more context, emotion and implication, which is exactly where AI coding is weakest, so they need closer human review.

How do I check AI-generated codes?

Read the source passage for every code, not just the code label. Code one or two transcripts yourself before looking at the AI output and compare the two. Look for codes that are too literal, for meaning the AI flattened, and for passages it skipped. Keep a note of what you changed and why.

Do I have to disclose AI use in my dissertation?

Yes. Most universities and journals now require you to say which tool you used, for which part of the analysis, and how you checked its output. Our guide to reporting AI-assisted analysis has wording you can adapt.

References

  • Barany, A. et al. (2024). ChatGPT for education research: exploring the potential of large language models for qualitative codebook development. In Artificial Intelligence in Education (AIED 2024). Author PDF
  • De Paoli, S. (2024). Performing an inductive thematic analysis of semi-structured interviews with a large language model: an exploration and provocation on the limits of the approach. Social Science Computer Review, 42(4), pp. 997–1019. doi:10.1177/08944393231220483
  • Friese, S. (2025). Response to: “We reject the use of generative artificial intelligence for reflexive qualitative research.” SSRN. doi:10.2139/ssrn.5690262
  • Friese, S., Nguyen-Trung, K., Powell, S. and Morgan, D.L. (2026). Beyond binary positions: making space for critical and reflexive GenAI integration in qualitative research. Qualitative Inquiry. doi:10.1177/10778004261429393
  • Hamilton, L., Elliott, D., Quick, A., Smith, S. and Choplin, V. (2023). Exploring the use of AI in qualitative analysis: a comparative study of guaranteed income data. International Journal of Qualitative Methods, 22. doi:10.1177/16094069231201504
  • Jowsey, T., Braun, V., Clarke, V., Lupton, D. and Fine, M. (2025). We reject the use of generative artificial intelligence for reflexive qualitative research. Qualitative Inquiry. doi:10.1177/10778004251401851
  • Morgan, D.L. (2023). Exploring the use of artificial intelligence for qualitative data analysis: the case of ChatGPT. International Journal of Qualitative Methods, 22. doi:10.1177/16094069231211248
  • Nguyen, D.C. and Welch, C. (2025). Generative artificial intelligence in qualitative data analysis: analyzing, or just chatting? Organizational Research Methods, 29(1), pp. 3–39. doi:10.1177/10944281251377154
  • Umer, M. et al. (2026). Using ChatGPT for thematic analysis of qualitative interviews in cultural research: a methodological investigation. Asian Journal of Psychiatry. doi:10.1016/j.ajp.2026.105071