How to do thematic analysis for a dissertation
Most guides to thematic analysis stop at the six phases. That leaves the part students actually get stuck on: what a code looks like on a real transcript, how you decide six candidate themes should really be three, and what to type into your methodology chapter once you're done. This guide runs one worked example the whole way through, from raw interview text to the paragraphs you submit.

Three decisions to make before you code anything
Every one of these ends up as a sentence in your methodology chapter, so settle them before the coding starts rather than reverse-engineering a justification at 2am.
Which framework. Braun and Clarke's reflexive thematic analysis is the default in most social science and health departments, and it's the one your examiners will recognise fastest. If you're synthesising published studies for a review chapter rather than analysing your own data, Thomas and Harden's thematic synthesis fits better. Our guide to the five frameworks covers the rest.
Inductive or deductive. Are codes coming from the data, or from a codebook you built out of existing theory? Most dissertations end up somewhere in between, which is fine, but you have to say so. Inductive vs. deductive coding walks through how to word it.
Semantic or latent coding. Semantic coding stays with what participants said. Latent coding reads for assumptions underneath it. Undergraduate work is usually semantic, and that is a defensible choice as long as you name it.
The worked example
Say your research question is: how do students in paid work manage the demands of full-time study? You've run six semi-structured interviews, 40 minutes each, transcribed. That's a typical masters-level dataset. Everything below uses the same six transcripts.
Phases 1 and 2: reading, then coding your transcripts
Familiarisation means reading all six transcripts before coding any of them. Resist the urge to start highlighting on the first read. You're looking for the shape of the dataset, and you cannot see that from transcript one.
Then code. A code is a short label, two to six words, attached to a specific chunk of text. Here's a real-looking extract with the codes a first pass would produce:
“I do Thursday and Friday nights at the bar, so Friday lectures are basically gone. I've watched maybe two of them live all term. And you can't tell your tutor that, can you, because then it sounds like you're not taking it seriously. So you just say you were ill.”
P03, second-year, 16 hours/week
- shift work displacing timetabled teaching — paid hours directly replace attendance, not just study time.
- catching up asynchronously — recordings used as a substitute for live sessions.
- concealing employment from staff — work is hidden rather than disclosed.
- fear of appearing uncommitted — the reason given for hiding it.
Four codes from four sentences. That density is normal and it is why coding six transcripts by hand takes most people a full week. Two habits will save you in the viva. Keep every code attached to the verbatim quote that produced it, because you will need to defend where each one came from. And code the whole dataset before you start grouping anything, because themes built from the first two transcripts tend to survive contact with the other four whether or not the data supports them.
Phase 3: from codes to candidate themes
Six transcripts will give you somewhere between 60 and 120 codes. Now you cluster them. A theme is not a topic, and this is the single most common thing examiners push back on. “Time management” is a topic. “Time is managed by quietly downgrading what counts as attendance” is a theme, because it makes a claim.
Grouping the example dataset produced four candidates:
- Work quietly rewrites what counts as attendance (12 codes, all six participants)
- Disclosure feels risky, so it doesn't happen (9 codes, five participants)
- Financial necessity is treated as a personal failing (7 codes, four participants)
- Peer networks substitute for institutional support (4 codes, two participants)
Phases 4 and 5: reviewing and naming
Reviewing is where candidates get cut, merged, or promoted. Two tests do most of the work. Does the theme hold together internally, with its codes telling one story? And is it distinct from the others, or is it really a sub-theme wearing a hat?
Candidate 4 in the example fails on evidence. Four codes from two participants out of six isn't a theme across this dataset; it's an observation worth a sentence in the discussion, or a recommendation for further research. Candidates 2 and 3 turn out to be closely linked, since both are about how students account for their work to other people, but they survive as separate themes because one is about what students hide and the other is about how they judge themselves. Three final themes from six interviews is a healthy result.
Naming matters more than students expect. A theme called “Challenges” tells a reader nothing. “Attendance as the first thing to go” tells them exactly what they are about to read.
How many themes should a dissertation have?
Three to six for most undergraduate and masters projects. Braun and Clarke don't set a number, and anyone who quotes you a rule is inventing it, but the practical constraint is your word count. Each theme needs a definition, two or three supporting quotes, and analytic commentary that does more than paraphrase the quotes. That's roughly 600 to 800 words per theme. In a 3,000-word findings chapter, seven themes means seven thin ones, and thin themes read as an analysis that never finished.
Fewer than three usually means you've grouped at too high a level and should split.
Writing the methodology paragraph
Your methodology chapter has to state the approach, the framework with its citation, how coding was done, and how themes were arrived at. Adapt this to what you actually did, and check it against your own supervisor's expectations:
Interview data were analysed using reflexive thematic analysis (Braun & Clarke, 2006, 2019). Following familiarisation with all six transcripts, semantic codes were generated inductively across the full dataset, with each code anchored to a verbatim extract. Codes were then collated into candidate themes, which were reviewed against both the coded extracts and the dataset as a whole (Phase 4). Candidate themes supported by fewer than three participants were not retained as themes but are noted in the discussion. Three themes were defined and named in Phase 5.
If you used software of any kind, name it here in the same paragraph. That includes AI-assisted coding.
Writing the findings chapter
One section per theme, in a deliberate order. Open the chapter with a short paragraph naming all three themes so the reader knows the structure, and a table mapping themes against participants is worth the space it takes.
Within each theme, the pattern is: define the theme in a sentence, present a quote, then say what the quote is doing. That third step is where marks live. Compare these two:
Weak. “P03 said they had watched two lectures live all term. This shows that work affects attendance.”
Better. “P03 described having watched ‘maybe two’ live lectures in a term. What is notable is not the absence itself but how it is framed: attendance is presented as something already conceded, requiring no decision. Four of the six participants described timetabled teaching in similar terms, as the first commitment to be given up rather than one weighed against others.”
The second version quotes precisely, interprets, and tells you how widely the pattern held. Do that for every quote you use.
How long it takes
For six interviews done by hand, budget two to three weeks of part-time work: two days on familiarisation, five to seven on coding, three or four on building and reviewing themes, and the rest on writing up. Coding is the bottleneck, and it is the part that expands when you are tired, because a tired coder produces inconsistent codes and then has to redo them.
Using AI for the first coding pass
A tool can generate that first set of codes in minutes instead of a week, and that is a genuine saving on the most mechanical stage of the process. It is not a finished analysis. The interpretive work in Phases 4 and 5, deciding what a theme is claiming and whether your data really supports it, is the part being examined, and it stays yours.
Three conditions before you use one. Check your institution's policy on AI assistance, since they differ sharply and some require declaration on the cover sheet. Review and rename every code yourself against the transcript it came from. And anonymise your transcripts before they go anywhere, which your ethics approval almost certainly requires anyway.
The tool on this site does the first pass in either direction: your own transcripts, one participant per box, or the findings of published studies if you're writing a review chapter. Every code comes back with the verbatim quote that produced it, so you can check each one rather than trusting it. Three sources are free.
What examiners push back on
- Themes that are topics. If a theme name could be a lecture title, it isn't making a claim yet.
- Quotes left to speak for themselves. A quote followed by “this shows that” and a restatement is description, not analysis.
- Themes resting on one participant. Report how many participants contributed to each theme, and say so plainly when a theme is thinly supported.
- No account of your own position. Reflexive TA expects a short reflexivity statement. If you were a working student yourself, that belongs in the methodology, not hidden.
- Nothing that contradicts the story. Real datasets contain disagreement. An analysis where every participant agrees usually means the analysis stopped early.
Try it on your own studies — free
Paste the findings of 3–15 studies, choose a framework — Braun & Clarke and more — and watch codes cluster into themes with a verbatim quote behind every one. First 3 studies free, no signup.
Start your free analysis