Qualitative methods guide
Qualitative Content Analysis: What It Is, When to Use It, and How to Do It
You have a pile of interview transcripts, or a few hundred open-ended survey answers, and your supervisor has said “just do a content analysis.” Then you open the literature and find three different definitions, two competing step lists and a debate about whether counting is allowed. This guide sorts that out. It covers what qualitative content analysis is, the three ways people actually do it, how to tell whether it suits your question, the process from raw text to finished categories with one example carried all the way through, and where the method falls short.

What is qualitative content analysis?
Qualitative content analysis is a research method for interpreting text by coding it systematically and grouping the codes into a small number of clearly defined categories (Hsieh & Shannon, 2005). It is used on interview transcripts, open-ended survey answers, documents and published studies, and follows three phases: preparation, organising and reporting (Elo & Kyngäs, 2008).
In practice you break the text into meaning units, give each a short code, group similar codes, and keep grouping until the categories describe the phenomenon. There are three approaches:
- Conventional: categories come out of the data. Use it when little is known.
- Directed: you start with categories from an existing theory and test or extend it.
- Summative: you count particular words or content, then interpret how they are used.
Its strength is transparency: another researcher can follow your rules and see how you got from a quote to a category. Its weakness is depth. It describes well but explains less than grounded theory or interpretive thematic analysis.
What qualitative content analysis is
Content analysis started life as a counting method. Early twentieth century researchers tallied words in newspapers and propaganda, and a quantitative version of it is still used in media studies. The qualitative version keeps the discipline of the original (explicit categories, written rules, a clear trail) but cares about what the words mean in context rather than how often they appear.
The definition most dissertations cite comes from Hsieh and Shannon (2005): a research method for the subjective interpretation of the content of text data through the systematic classification process of coding and identifying themes or patterns. Notice the two words pulling against each other there. Subjective, because you are interpreting. Systematic, because you do it by rules that someone else could check. Mayring (2000) leans harder on the systematic side and describes step-by-step models and rules for assigning text to categories, applied without rushing to quantify.
One distinction you will be asked about in a viva is manifest versus latent content. Manifest content is what the text plainly says. Latent content is what it implies: the hesitation before an answer, the joke that covers a complaint, the topic a participant keeps steering away from. Elo and Kyngäs (2008) ask you to decide before you start whether you will analyse only the manifest content or the latent content too, and to say so in your write-up. Latent analysis gets you further but invites over-interpretation, so it needs more evidence per claim.
The output is a category system. Each category has a name, a definition, the codes inside it, and example quotes. Depending on the aim, the categories may sit in a hierarchy (main category, generic categories, sub-categories) that you can draw as a tree. Our thematic map generator draws that kind of hierarchy for you as a figure.
The three approaches: conventional, directed and summative
Hsieh and Shannon's 2005 paper in Qualitative Health Research is the reason most people now talk about three types. The difference between them is where your codes come from. To make that concrete, imagine the same study run three ways: twelve interviews with international students about their first term at a UK university.
| Approach | Codes come from | The student study would ask |
|---|---|---|
| Conventional (inductive) | The data, during analysis | “What was the first term like for you?” Categories are built from whatever students raise. |
| Directed (deductive) | An existing theory or earlier findings, before analysis | Starts from a model of acculturation stress (for example language, homesickness, discrimination, practical stress) and codes against it. Anything that doesn't fit gets a new code. |
| Summative | Keywords chosen from interest or the literature | Counts how often students say “lonely” against softer words like “quiet” or “on my own,” then asks who uses which, and when. |
Conventional content analysis suits a topic where theory is thin. You read everything several times, highlight the exact words that seem to carry key ideas, label them, and let the labels settle into categories. The advantage, in Hsieh and Shannon's words, is getting information from participants without imposing preconceived categories. Any theory you know about goes in the discussion chapter, where you compare your categories against it.
Directed content analysis is for extending or testing something that already exists. You turn the theory into a coding scheme with an operational definition for each category, then code the data against it. Hsieh and Shannon describe two ways to start: highlight everything that looks relevant first and code it afterwards (safer, you miss less), or code straight into the predefined categories and park whatever doesn't fit for later. Elo and Kyngäs call the coding scheme a categorisation matrix, which can be structured (only the theory's categories) or unconstrained (the theory's categories plus new ones built inductively).
Summative content analysis begins with counting and then moves to interpretation. If you stopped at the counts it would be a quantitative study. It becomes qualitative when you go back into the text and ask why a word is used, by whom and in what circumstances. Hsieh and Shannon's own illustration is how often clinicians, patients and families say “death” and “dying” versus euphemisms like “passing.”
In student projects the conventional approach is by far the most common. If you aren't sure which one you are doing, you are probably doing conventional content analysis. The same data-first versus theory-first choice runs through every qualitative method, and our guide to inductive vs deductive coding explains how to decide.
When to use it (and when not to)
Content analysis is a good fit when your research question starts with “what”: what do nurses report about night shifts, what concerns do parents raise in open survey comments, what do these policy documents say about inclusion. It handles large amounts of text well and gives a clear, describable answer. It is common in nursing, health, education and communication research, partly because examiners in those fields value its step-by-step transparency.
Reach for it when:
- You want to describe a phenomenon, not build a theory of it.
- You have a lot of fairly short text, such as open-ended survey answers, where deep interpretation of each response isn't possible anyway.
- You want to report how widespread each category is across participants or documents.
- You have an existing framework to test or extend (the directed approach).
- You are reviewing literature and want to categorise what studies have found (see synthesising findings across studies).
Pick something else when your question is about process or explanation (“how does trust develop between…”), which is grounded theory territory, or about lived experience and meaning, which points to IPA or phenomenology. Hsieh and Shannon are direct about this: the most a conventional content analysis delivers is concept development or a model, not a theory. If you are still choosing, our overview of qualitative data analysis methods compares the main options.
The comparison you will most likely have to defend is against thematic analysis. Vaismoradi, Turunen and Bondas (2013) note that the two share a lot and are often confused. The practical difference is this. Content analysis stays closer to the surface of the text, uses tightly defined categories, and is comfortable reporting frequencies. Braun and Clarke's reflexive thematic analysis treats themes as the researcher's interpretation of shared meaning and does not use counts as evidence that a theme matters. If you want the longer version, our guide to reflexive TA versus other qualitative methods goes further.
| Aspect | Qualitative content analysis | Reflexive thematic analysis |
|---|---|---|
| Main output | Categories with definitions and coding rules | Themes: interpretive stories about shared meaning |
| Level | Mostly manifest content, latent optional | Semantic and latent meaning |
| Counting | Common and often expected | Not treated as evidence of importance |
| Second coder | Often used to check consistency | Not required; reflexivity is the quality check |
| Key sources | Hsieh & Shannon (2005); Elo & Kyngäs (2008); Mayring (2000) | Braun & Clarke (2006, 2021) |
| Best for | Describing what is in a large body of text | Interpreting what experiences mean |
The process, step by step
Elo and Kyngäs (2008) organise both the inductive and the deductive routes into three phases: preparation, organising and reporting. That structure is the one most methods sections borrow, and it is the one used here. The steps below are for the conventional (inductive) route. Differences for the directed route are noted where they matter.
Phase 1: Preparation
1. Choose your unit of analysis. This is the chunk of text you treat as one thing: a whole interview, one answer, a paragraph, a sentence. Elo et al. (2014) put the trade-off neatly. Too broad and a unit carries several meanings at once. Too narrow and the text fragments into pieces that no longer mean anything. For interviews, most students treat each transcript as the unit of analysis and code at the level of meaning units inside it.
2. Decide on manifest, latent, or both. Write the decision down. It changes what counts as evidence later.
3. Get a sense of the whole. Read every transcript through, once without a pen. Elo and Kyngäs suggest keeping a few questions in mind as you read: who is telling this, where and when it happened, what is happening, and why. Only then start marking.
Phase 2: Organising
4. Open coding. Go through the text and write notes and headings in the margin for everything relevant to your question. Graneheim and Lundman (2004) give a useful sequence here. Lift out a meaning unit (the words that belong together), shorten it into a condensed meaning unit that keeps the core, and give it a code. Hsieh and Shannon suggest open coding three or four transcripts, settling on preliminary codes, then coding the rest (and recoding the first ones) with new codes added whenever something doesn't fit. A code labels one idea in the text; a category groups many codes. The code vs theme guide covers the same distinction in thematic analysis, and how to code an interview transcript covers code types and codebooks.
5. Group codes into categories. Put codes that belong together into sub-categories. A category is a group of content that shares something. The point is to decide which codes belong together and which don't, so each category should be internally consistent and clearly different from its neighbours.
6. Abstract. Keep grouping upward: sub-categories into generic categories, generic categories into main categories. Name each level using content-characteristic words. Stop when further merging would lose something that matters to your question. Elo and Kyngäs admit that part of this step rests on the researcher's insight and is hard to describe, which is exactly why you keep notes as you go.
7. Define every category. Write a definition and pick one or two anchor quotes for each. Mayring calls these anchor examples. They are what a second coder would use to check whether a new passage belongs.
Directed route: steps 4 to 7 run the other way round. You build the categorisation matrix from the theory first, write the definitions and coding rules, pre-test the matrix on a transcript or two, then code. Data that won't fit gets abstracted inductively into new categories, and those new categories are often the most interesting part of the findings.
Phase 3: Reporting
8. Report the categories and how you got them. Present the category system (a table or tree diagram works well), describe what each category contains, and support it with quotes. Then describe the analysis itself in enough detail that a reader could judge it. Elo et al. (2014) stress that readers judge the trustworthiness of a content analysis by how clearly the path from data to categories is shown, so this part carries real weight.

A worked example: from quote to category
Back to the twelve international students, analysed with a conventional approach and manifest content only. The research question is: What challenges do international students describe in their first term? Here is how four excerpts move through the steps.
| Meaning unit | Condensed | Code |
|---|---|---|
| “I worked out that after rent I had about forty pounds a week, so I stopped going when people went out for food.” (P3) | Little money left after rent, so skipped social meals | Cost keeps me out of social life |
| “My part-time job was capped at twenty hours and the shifts were always the evenings the society met.” (P8) | Work hours clash with society meetings | Paid work competes with belonging |
| “Nobody told us you could email the lecturer. At home you would never do that.” (P5) | Didn't know contacting lecturers was allowed | Unwritten rules about staff contact |
| “The first essay came back with ‘too descriptive’ and I honestly didn't know what that meant.” (P11) | Feedback language not understood | Unfamiliar academic expectations |
After coding all twelve transcripts the student has 41 codes. Grouping and abstraction then goes like this:
| Codes (examples) | Sub-category | Main category |
|---|---|---|
| Cost keeps me out of social life · Paid work competes with belonging · Choosing cheaper housing far from campus | Money shaping social life | Money worries (9 of 12 participants) |
| Sending money home · Fear of visa rules if hours exceeded | Financial obligations and limits | |
| Unwritten rules about staff contact · Unfamiliar academic expectations · Not knowing what “independent study” means | Hidden academic norms | Learning the unwritten rules (10 of 12) |
The definition written for the second main category might read: “Learning the unwritten rules: accounts of expectations that staff or home students treated as obvious but that were never stated, covering contact with staff, assessment language and study habits. Excludes difficulties with English itself, which are coded under Language.” That last sentence is a coding rule. It tells a second coder where the boundary is, and it is the kind of detail that makes content analysis checkable.
Notice what the analysis doesn't do. It doesn't explain why money and hidden rules combine to push some students to the edge of university life, or build a model of how belonging develops. Those would be natural questions for the discussion chapter, or for a follow-up study using a different method.
Should you count?
Reporting that a category appeared in 9 of 12 interviews is normal in content analysis and often expected. It helps a reader see whether a category describes most of the sample or a vocal few. Hsieh and Shannon suggest rank-ordering code frequencies in directed studies rather than running statistical tests, because qualitative samples are rarely designed for them.
Two cautions. First, count participants, not mentions. One person who talks about rent eleven times is still one person. Second, a rare category can matter more than a common one. If one student describes being racially abused on a night bus, that belongs in your findings even though the count is one.
Limitations you should name in your thesis
Examiners like to see that you know where your method is weak. The literature is fairly specific about content analysis:
- It describes more than it explains. Content analysis has no built-in technique for connecting concepts to each other (Elo & Kyngäs, 2008). You get categories, and perhaps a hierarchy, but not a causal or process account.
- Conventional analysis can miss the context. Hsieh and Shannon warn that without a full understanding of the context you may fail to spot key categories, producing findings that don't represent the data. They also note it is easily confused with grounded theory or phenomenology, so say clearly which one you are doing.
- Directed analysis leans toward confirmation. Starting from a theory is an informed bias. You are more likely to find support for it than evidence against it, and interview probes built from the theory can cue participants to agree. An audit trail and a second reviewer of your definitions help.
- Summative analysis can stay shallow. Word counts show usage, not meaning, unless you do the interpretive work that follows.
- Abstraction often stops too early. Elo et al. (2014) observe that a long list of overlapping categories usually means the grouping is unfinished. The results then read like a list of what participants said instead of an analysis of it.
- Latent content invites over-reading. Pauses, laughter and silences are data, but the further you move from the words, the more evidence you need for each claim.
- It looks easy and isn't. Elo and Kyngäs point out that its reputation as a simple method leads researchers to underestimate it. The mechanics are simple. Getting clean, non-overlapping categories out of messy text takes several rounds.
Most of these are handled the same way: define categories tightly, keep an audit trail of your decisions, have someone check a sample of your coding against your definitions, and show your working in the write-up. Elo et al. (2014) include a checklist of questions for each phase that is worth running through before submission.
Getting the first pass done faster
Look back at the worked example. The judgement calls (what counts as a meaning unit, where one category ends and another begins, whether “money worries” is really about money) are yours and should stay yours. The slow part is everything around them: reading twelve transcripts line by line, writing out 41 codes, copying quotes into a spreadsheet, and counting which participant said what. That can take weeks, and it is usually the part where students run out of time.
thematicanalysis.ai runs a qualitative content analysis framework built on Hsieh & Shannon and Elo & Kyngäs. You paste or upload your interview transcripts, survey answers or published papers and it does the conventional route: open codes grounded in the text, grouped into categories, with the verbatim quote behind every code and a note of how many sources support each category. You then read it the way a careful supervisor would, merge, split and rename, and export it as a Word report. On a paid pass you can also export to NVivo or MAXQDA and carry on coding there.
Try it on one transcript and see whether the categories it finds match the ones you would have found. The first three sources are free.
Writing it up in your methods section
Name the approach, cite the process you followed, state your unit of analysis and whether you analysed latent content, and say how you checked the analysis. Something like this, adapted to your study:
Data were analysed using conventional qualitative content analysis (Hsieh & Shannon, 2005), following the preparation, organising and reporting phases described by Elo and Kyngäs (2008). Each transcript was treated as a unit of analysis, and analysis focused on manifest content. Transcripts were read in full several times before open coding. Meaning units were condensed and coded (Graneheim & Lundman, 2004), and codes were grouped into sub-categories and abstracted into main categories. Each category was given a written definition, coding rules and anchor quotes. A second researcher coded a 20% sample against the category definitions, and disagreements were resolved through discussion.
If you used software or AI for any part of the coding, say which part and how you checked it. Our guide on reporting AI-assisted analysis has wording you can adapt. For how the methods section fits with the findings chapter, see qualitative analysis for a dissertation, and if your analysis includes latent content, a short positionality statement helps examiners see where your interpretation comes from.
Frequently asked questions
What are the steps of qualitative content analysis?
Following Elo and Kyngäs (2008): (1) choose a unit of analysis, (2) decide whether to analyse manifest content only or latent content too, (3) read all the data to get a sense of the whole, (4) open code the text, (5) group codes into sub-categories, (6) abstract them into main categories, (7) write a definition, coding rules and anchor quotes for each category, and (8) report the categories and the analysis process.
What are the three types of qualitative content analysis?
Hsieh and Shannon (2005) describe conventional content analysis, where categories come from the data; directed content analysis, where an existing theory supplies the initial codes; and summative content analysis, which counts keywords or content and then interprets how they are used in context.
What is qualitative content analysis in simple terms?
It is a way of reading a body of text (interviews, open survey answers, documents) and sorting what it says into a small set of clearly defined categories, following rules you could explain to someone else. Hsieh and Shannon (2005) define it as the subjective interpretation of text through a systematic process of coding and identifying patterns.
What is the difference between content analysis and thematic analysis?
They overlap a lot. Content analysis usually stays closer to what the text says, produces categories with clear definitions, and often reports how many participants or sources support each one. Reflexive thematic analysis (Braun and Clarke) treats themes as the researcher's interpretive story about patterns of meaning and does not treat counting as evidence. If your examiner expects a descriptive, rule-based category system, content analysis fits. If they expect an interpretive account, thematic analysis fits better.
Is qualitative content analysis inductive or deductive?
It can be either. Conventional (inductive) content analysis builds categories from the data. Directed (deductive) content analysis starts from an existing theory or earlier findings and uses them as the first coding scheme. Elo and Kyngäs (2008) describe both routes through the same three phases: preparation, organising and reporting.
How many categories should I end up with?
There is no fixed number. Hsieh and Shannon cite Morse and Field's suggestion of 10 to 15 clusters at the grouping stage, which you then abstract into fewer main categories. Elo et al. (2014) warn that a long list of categories is usually a sign the abstraction is unfinished and categories overlap.
Can I use qualitative content analysis for a literature review?
Yes. The unit of analysis becomes the reported findings of each study instead of an interview transcript. You code what each paper found, group the codes into categories, and report how many studies support each category.
Is it acceptable to use AI for qualitative content analysis?
Many universities accept AI-assisted coding if you disclose it, check every code against the source text, and make the final analytic decisions yourself. Treat the AI output as a first pass that you audit, not as the analysis. Check your own department's policy before you submit.
References
- Elo, S. and Kyngäs, H. (2008). The qualitative content analysis process. Journal of Advanced Nursing, 62(1), pp. 107–115. doi:10.1111/j.1365-2648.2007.04569.x
- Elo, S., Kääriäinen, M., Kanste, O., Pölkki, T., Utriainen, K. and Kyngäs, H. (2014). Qualitative content analysis: a focus on trustworthiness. SAGE Open, 4(1). doi:10.1177/2158244014522633
- Graneheim, U.H. and Lundman, B. (2004). Qualitative content analysis in nursing research: concepts, procedures and measures to achieve trustworthiness. Nurse Education Today, 24(2), pp. 105–112. doi:10.1016/j.nedt.2003.10.001
- Hsieh, H.-F. and Shannon, S.E. (2005). Three approaches to qualitative content analysis. Qualitative Health Research, 15(9), pp. 1277–1288. doi:10.1177/1049732305276687
- Mayring, P. (2000). Qualitative content analysis. Forum Qualitative Sozialforschung / Forum: Qualitative Social Research, 1(2), Art. 20. doi:10.17169/fqs-1.2.1089
- Vaismoradi, M., Turunen, H. and Bondas, T. (2013). Content analysis and thematic analysis: implications for conducting a qualitative descriptive study. Nursing & Health Sciences, 15(3), pp. 398–405. doi:10.1111/nhs.12048