Qualitative Methodology Guide
How to Analyse Qualitative Data: A Step-by-Step Guide
Whether you're still planning your study or already sitting on a folder of interview transcripts, “how do I actually analyse this” is one of the most common questions in qualitative research, and one of the least explained. This guide walks through the real, well-documented steps, the main methods you can choose between, and the honest realities (time, subjectivity, sample size) that most methods courses skip over.

Executive Summary: Analysing Qualitative Data in Brief
Qualitative data analysis is an interpretive, systematic process moving through four shared stages: data organisation, close coding, pattern generation (categories or themes), and contextual synthesis back to literature. The method you choose depends on your research question:
What “analysing” qualitative data actually means
Qualitative data analysis is the process of organising, coding, and interpreting non-numerical data, interview transcripts, open survey responses, focus group recordings, documents, so that patterns of meaning become visible and defensible. It's fundamentally different from quantitative analysis: you're not calculating a mean or running a significance test, you're building an interpretive argument, one that still has to be systematic, transparent, and traceable back to your original data if it's going to hold up under scrutiny (Bingham, 2023).
That last part is the bit most beginners underestimate. “Qualitative” doesn't mean “unstructured.” A rigorous qualitative analysis needs to be credible, consistent, and free enough from unexamined bias that another researcher could follow your reasoning from raw data to final conclusion (Bingham, 2023).
Choosing a method: what each one is, and when it fits
Thematic analysis
- What it is: An approach for identifying patterns of meaning, themes, across a dataset, most commonly following Braun and Clarke's (2006) six-phase process of familiarisation, coding, generating themes, reviewing, defining, and writing up.
- The process: You code the data closely, cluster related codes around a shared interpretive idea, then test and refine those candidate themes against the full dataset before naming and writing them up.
- Considerations: It doesn't require a fixed theoretical lens, works with almost any qualitative data type, and can run inductively (patterns drawn from the data) or deductively (patterns tested against an existing framework).
- When it's chosen: When you want a flexible, widely understood method that works across disciplines, and when your question is genuinely “what patterns of meaning exist here,” not “what new theory explains this” or “how often does X appear.”
- Limitations: Its flexibility means comparatively few built-in procedural guardrails, so first-time researchers can under-structure their analysis without realising it. Reflexive versions of the method also deliberately reject inter-rater reliability checks, which can surprise reviewers trained in other traditions. Our guide on the difference between a code and a theme covers the single most common mistake at this stage: mistaking a topic summary for a genuine theme.
Grounded theory
- What it is: A methodology aimed at generating a new theory, grounded directly in the data, to explain a social process the existing literature doesn't already account for (Chun Tie, Birks and Francis, 2019).
- The process: Data collection and analysis happen concurrently and iteratively rather than sequentially. Coding moves through stages, usually called open (or initial), axial (or focused), and selective (or theoretical) coding depending on which genre you follow, using the constant comparative method throughout: comparing new data against existing codes, and codes against each other, until categories stabilise. Theoretical sampling then directs which new data to collect next, based on gaps the analysis has already surfaced, continuing until theoretical saturation, the point where new data stops adding anything to the developing categories.
- Considerations: Grounded theory has three distinct genres, classic (Glaser), evolved (Strauss and Corbin), and constructivist (Charmaz), which differ in how much the researcher's own interpretation shapes the resulting theory. Choosing one and staying consistent with it matters for methodological coherence.
- When it's chosen: When your research question genuinely asks “what's actually going on here, and can we build a new explanation for it,” particularly in an under-theorised area, and when your project timeline allows the iterative collect-analyze-collect cycle the method requires.
- Limitations: It's a demanding methodology for a first qualitative project. Concurrent data collection and analysis, theoretical sampling, and the requirement to reach genuine saturation all take considerably more time and methodological confidence than a more linear method, and coding qualitative data well is time-consuming regardless of which method you choose (Timonen, Foley and Conlon, 2018).
Qualitative content analysis
- What it is: A method for systematically classifying and interpreting text data through coding, with three distinct approaches: conventional (codes derived entirely from the data), directed (codes drawn partly from existing theory before analysis begins), and summative (starting with keyword counts, then interpreting the context around them) (Hsieh and Shannon, 2005).
- The process: Depending on the approach, you either code inductively as you read (conventional), apply a predetermined coding scheme derived from theory (directed), or count and then contextually interpret specific terms (summative), then group codes into categories that describe the manifest or latent content of the text.
- Considerations: Content analysis sits closer to description than deep interpretation compared to thematic analysis or grounded theory, though directed and summative approaches bring in more structure and quantification.
- When it's chosen: When your research question is closer to “what's actually present in this text and how often or in what form,” especially with large volumes of text, existing theoretical categories to test, or a need for some quantifiable element alongside the qualitative interpretation.
- Limitations: The summative approach in particular risks reducing meaning to word frequency, missing context or connotation the researcher needs to interpret carefully rather than count. Directed content analysis also carries a real risk of researchers finding what their existing framework predisposes them to find (Hsieh and Shannon, 2005).
Case study research
- What it is: An in-depth examination of a single bounded case, an organisation, event, programme, or individual, using multiple sources of data to build a rich, contextual understanding, rather than seeking patterns across many participants (Greenhalgh, 2025).
- The process: You define the boundaries of your case clearly at the outset, then draw on multiple data sources (interviews, documents, observation) to build a detailed, triangulated account, typically analysed thematically or narratively once collected.
- Considerations: Greenhalgh (2025) notes that case study research is widely used but often poorly understood, partly because it's practiced differently across disciplines, so being explicit about which case study tradition you're following matters more here than in most other methods.
- When it's chosen: When your question is genuinely about one specific, bounded instance in depth, rather than about generalising across a population, and when multiple data sources are available for that single case.
- Limitations: Findings are harder to generalise beyond the specific case by design, and the method's flexibility across disciplines means there's less methodological consensus to lean on when justifying your choices to a reviewer unfamiliar with your specific tradition.
Narrative analysis
- What it is: An approach that treats participants' accounts as stories, examining structure, sequence, and how something is told, not just its content.
- The process: Rather than fragmenting data into codes early, narrative analysis typically preserves whole accounts for longer, examining plot, characters, turning points, and how a participant constructs meaning through the way they tell their story.
- Considerations: This method resists the early fragmentation that thematic analysis and grounded theory both rely on, since breaking a story into isolated codes too early can destroy exactly what the method is trying to study.
- When it's chosen: When how someone tells their story, sequence, tone, self-positioning, matters as much as what they say, common in identity research, illness narratives, and life-history studies.
- Limitations: It typically works with fewer participants studied in greater depth, which narrows generalisability, and the close, whole-account reading it requires makes it a slower method per participant than thematic or content analysis.
If you're still unsure which fits, working backwards from your research question is usually faster than comparing method definitions in the abstract: patterns across many people points to thematic analysis, a new explanatory theory points to grounded theory, frequency or manifest content points to content analysis, deep insight into one bounded instance points to case study, and how a story is told points to narrative analysis.
How many interviews or sources is enough?
There's no single number, but the empirical research gives real ranges rather than guesswork. Studies that have tested when new codes stop appearing (a concept called saturation) generally find it happens within 9 to 17 interviews for most projects with a reasonably focused research question, and separately, that most new codes tend to show up within the first six interviews of a study, with very few genuinely new ones appearing after twelve (Hennink and Kaiser, 2022; Guest, Bunce and Johnson, 2006). Braun and Clarke's own guidance for thematic analysis specifically suggests roughly 6 to 10 interviews for a smaller undergraduate project, 10 to 20 for a master's-level study, and 30 or more for a PhD, scaling with the depth expected at each level.
Treat these as sanity checks rather than fixed targets. The honest answer to “how many is enough” is: enough that your themes are well-supported and further data isn't changing your picture, not a number you hit and stop.
The parts nobody tells you about upfront
It takes longer than you think, every time.
Reading transcripts closely enough to code them well, then reviewing and refining themes across several rounds, is slow, careful work with no real shortcut through the reading itself (Bingham, 2023). If your project timeline budgets a week for analysis, budget more.
It's subjective, and that's not a flaw to apologise for.
Two careful researchers coding the same transcript will produce genuinely different, both-defensible codes, because interpretation is part of the method, not contamination of it. What matters is that your reasoning is documented and consistent, not that it matches what someone else would have written.
Software organises your data. It doesn't analyse it for you.
Tools like NVivo, ATLAS.ti, and Dedoose are genuinely useful for managing large datasets, but a peer-reviewed review of NVivo specifically is blunt about the boundary: these programs “do not... replace the need for the human researcher” (Dhakal, 2022). If you're deciding between qualitative software options, our guide on using NVivo for thematic analysis covers what the software actually does phase by phase, and where its real learning curve sits.
Where AI genuinely helps, and where it doesn't
Coding transcripts by hand, especially the open-coding phase described above, is exactly the kind of slow, repetitive first pass that AI tools can meaningfully speed up. The research is consistent that AI-assisted coding works best as a supervised first draft, not a finished analysis: comparative studies against human coders find real agreement on straightforward, descriptive coding, but a persistent gap on interpretive, latent-level meaning, the kind of reading-between-the-lines that a genuine theme requires (Wachinger et al., 2025). Our guide on AI thematic analysis and academic validity covers exactly what the evidence supports and where it doesn't.
Stuck on the coding step alone? thematicanalysis.ai/analyze turns your transcripts into a first-pass codebook in minutes, not weeks.
The short version
Analysing qualitative data means moving systematically from organising raw material, through careful coding, to identifying patterns, and finally connecting those patterns to the wider literature, documenting your reasoning at every step along the way. There's no single correct method, thematic analysis, grounded theory, content analysis, and case study all suit different questions, but they share the same underlying discipline: close reading, transparent decisions, and a trail back to the data that anyone could follow. Do that well, and it doesn't matter that it took longer than you hoped. That's what makes it credible.
Frequently asked questions
What does analysing qualitative data actually mean?+
Qualitative data analysis is the systematic process of organising, coding, and interpreting non-numerical data (such as interview transcripts, focus groups, or documents) to build a credible, traceable interpretive argument. Unlike quantitative analysis, it does not calculate statistical tests, but requires rigorous, defensible reasoning back to verbatim evidence.
How do I choose between thematic analysis, grounded theory, and content analysis?+
Work backwards from your research question: choose thematic analysis if you seek patterns of meaning across many accounts; grounded theory if you need to build a new explanatory theory for an under-theorised social process; qualitative content analysis if you need structured classification or frequency-informed counts; case study research for an in-depth bounded instance; and narrative analysis if how a story is told matters as much as what is said.
How many interviews are enough for qualitative research saturation?+
Empirical saturation studies (Hennink & Kaiser, 2022; Guest et al., 2006) show that the majority of new codes appear within the first 6 to 12 interviews, with saturation typically achieved within 9 to 17 interviews for a focused study. Braun and Clarke recommend roughly 6–10 interviews for undergraduate projects, 10–20 for master's theses, and 30+ for doctoral dissertations.
Does qualitative data software (like NVivo) analyse the data for you?+
No. CAQDAS tools like NVivo, ATLAS.ti, and Dedoose organise, tag, and query text, but as peer-reviewed literature confirms, they do not replace human interpretive synthesis. The researcher remains entirely responsible for developing insights, reviewing themes, and writing the narrative.
Where does AI genuinely help in qualitative coding?+
AI tools excel at the time-consuming open-coding first pass, transforming raw transcripts into a structured, quote-grounded codebook in minutes. However, research (Wachinger et al., 2025) shows AI functions best as a supervised first draft, requiring human researchers to refine latent meanings, resolve ambiguities, and ensure reflexive validity.
References
- Bingham, A.J. (2023). From data management to actionable findings: a five-phase process of qualitative data analysis. International Journal of Qualitative Methods, 22. doi:10.1177/16094069231183620
- Braun, V. and Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3(2), pp. 77–101. doi:10.1191/1478088706qp063oa
- Chun Tie, Y., Birks, M. and Francis, K. (2019). Grounded theory research: a design framework for novice researchers. SAGE Open Medicine, 7. doi:10.1177/2050312118822927
- Dhakal, K. (2022). NVivo. Journal of the Medical Library Association, 110(2), pp. 270–272. doi:10.5195/jmla.2022.1271
- Greenhalgh, T. (2025). Case studies: a guide for researchers, educators, and practitioners. (open access via PMC). PMC12458855
- Guest, G., Bunce, A. and Johnson, L. (2006). How many interviews are enough? An experiment with data saturation and variability. Field Methods, 18(1), pp. 59–82. doi:10.1177/1525822X05279903
- Hennink, M. and Kaiser, B.N. (2022). Sample sizes for saturation in qualitative research: a systematic review of empirical tests. Social Science & Medicine, 292, 114523. doi:10.1016/j.socscimed.2021.114523
- Hsieh, H.F. and Shannon, S.E. (2005). Three approaches to qualitative content analysis. Qualitative Health Research, 15(9), pp. 1277–1288. doi:10.1177/1049732305276687
- Timonen, V., Foley, G. and Conlon, C. (2018). Challenges when using grounded theory: a pragmatic introduction to doing GT research. International Journal of Qualitative Methods, 17(1). doi:10.1177/1609406918758086
- Wachinger, J., Bärnighausen, K., Schäfer, L.N., Scott, K. and McMahon, S.A. (2025). Prompts, pearls, imperfections: comparing ChatGPT and a human researcher in qualitative data analysis. Qualitative Health Research, 35(9), pp. 951–966. doi:10.1177/10497323241244669