Skip to content
← Guides & helpGuides8 min readBy The CiteDash team

Thematic analysis step by step

Braun and Clarke's six phases of thematic analysis explained step by step, with a worked coding example and the quality checks examiners apply.

Thematic analysis is probably the most common qualitative method in graduate theses, and one of the most commonly mishandled. The name sounds simple: read the data, find the themes. In practice, the difference between a findings chapter that survives a viva and one that gets picked apart is whether your themes are patterns of meaning built through systematic coding, or just topic headings with quotes pasted underneath.

This guide walks through the six phases described by Virginia Braun and Victoria Clarke, whose 2006 paper in Qualitative Research in Psychology is the standard citation for the method, with a worked coding example and the quality checks examiners actually apply. It is written for a thesis context: interviews, focus groups, or open-ended survey responses, usually analysed by one researcher who must defend every decision alone.

What is thematic analysis?

Braun and Clarke define thematic analysis as a method for identifying, analysing and reporting patterns of meaning, called themes, across a qualitative dataset. Two parts of that definition carry weight. First, it is a method, not a methodology: unlike grounded theory or interpretative phenomenological analysis, it is not tied to a single theoretical framework, so you can use it from different epistemological positions, but you must state the position you are using it from. Second, the unit of analysis is a pattern of shared meaning organised around a central concept, not a topic. Barriers to exercise is a topic. Exercise felt like another job I was failing at is a pattern of meaning.

Since 2006 the authors have developed the method into what they now call reflexive thematic analysis, emphasising that themes do not emerge from data like fossils from rock: they are actively generated by the researcher through engagement and interpretation. That framing has real consequences for how you report the analysis, which the next section covers.

Know which school of thematic analysis you are in

The methodological literature distinguishes three broad approaches that all trade under the thematic analysis name. Coding reliability approaches treat coding as a measurement task: a structured codebook fixed early, multiple coders, agreement statistics. Codebook approaches, such as framework analysis and template analysis, use a structured codebook but keep an interpretive stance. Reflexive thematic analysis treats coding as an organic, evolving interpretive practice by a researcher whose subjectivity is a resource to be disciplined through reflexivity, not a bias to be eliminated.

Pick one and be consistent, because examiners flag hybrids that make no methodological sense. The classic is claiming reflexive thematic analysis while reporting intercoder reliability scores, a combination Braun and Clarke explicitly argue against for their approach. Your methodology chapter should name the school, cite it, and explain why it fits your research question, and your findings chapter should then behave like the school you named.

Decisions to make before you code

Four decisions shape everything downstream, and each should be stated in your methods chapter rather than left implicit:

  • Inductive or deductive: are codes built up from the data, are you bringing a framework of prior concepts to the data, or are you running a declared hybrid of both?
  • Semantic or latent: are you coding the explicit surface meaning of what participants said, or the assumptions and ideas underneath the surface?
  • What counts as a theme for you: a pattern of shared meaning with a central organising concept, not simply anything several participants mentioned.
  • What the dataset is: which transcripts, which questions, whether field notes and marginalia are in or out. Changing the dataset mid-analysis quietly invalidates the earlier phases.

Phase 1: familiarisation

Familiarisation means immersing yourself in the data before trying to label it. Transcribe the recordings yourself or at least check every transcript against audio, read each transcript at least twice, and keep informal notes on anything striking: recurring phrases, contradictions, jokes, silences. Transcription feels clerical but is analytically useful, because hesitations and self-corrections are often where the interesting meaning lives. If you gathered open-ended responses through a form built in Collect, export the free-text answers and treat each respondent's set as a document in the same way.

Resist the urge to start coding on the first read. Notes at this stage are memos to yourself, not codes: observations like participants keep describing autonomy in negative terms. They will make the next phase faster and considerably richer. Decisions about what data you have available start much earlier than analysis, of course; the guide on collecting research data for a thesis covers sampling, consent, and storage.

Phase 2: generating initial codes, with a worked example

A code is a short label that captures something analytically interesting about a segment of data in relation to your research question. Coding in this phase is systematic: you work through every transcript, giving each one full and equal attention, tagging everything relevant rather than cherry-picking vivid passages. Suppose your study asks how remote workers experience the boundary between work and home, and a participant says: I answer emails from bed, honestly. My manager says switch off, but then messages at nine at night, so what am I supposed to do.

That one extract could legitimately carry several codes: work intrudes into domestic space; mixed messages from management; responsibility for boundaries pushed onto the worker. Notice the codes are more specific than a topic (boundaries) and less abstract than a theme. Expect a first pass across a full dataset to generate dozens or even a few hundred codes, and expect to merge and rename constantly. Keep a running code list with a one-line definition for each, because drift in what a code means between transcript one and transcript twelve is the most common silent failure in solo coding.

Phase 3: generating candidate themes

Themes are built, not found: you construct them by clustering codes whose meanings point at a shared central concept. Export or print the code list, group codes that speak to each other, and ask what the organising idea of each cluster is. In the remote-work example, codes about intrusion, mixed managerial signals, and self-blame might cluster into a candidate theme like boundary work has been outsourced to the worker. Some codes will not fit anywhere; park them in a miscellaneous pile rather than forcing them into a cluster where they dilute the concept.

A thematic map helps at this stage: a one-page sketch of candidate themes, subthemes, and the codes feeding each. It is a working document, and its job is to be wrong in instructive ways. Keep the dated versions; they become part of your audit trail and often a figure in the methodology chapter.

Phase 4: reviewing themes

Review runs at two levels. First, against the coded extracts: read every extract sitting under each candidate theme and check that the theme genuinely holds them together. Second, against the entire dataset: reread the transcripts and check that the themes capture the important patterns and that nothing significant is left homeless. Themes get split, merged, renamed, and discarded here, and losing a candidate theme is progress rather than failure. A findings chapter with three coherent, well-evidenced themes beats one with seven overlapping ones every time.

Two tests are worth applying to each survivor. Can you state the theme's central organising concept in a single sentence? And can you say clearly how it differs from its neighbours? If two themes keep borrowing each other's extracts, they are probably one theme wearing two names.

Phase 5: defining and naming themes

Write a short definition of each theme: its central concept, its scope, what it includes and excludes, and how it relates to the research question and to the other themes. If you cannot write the definition, the theme is not ready, and it is better to discover that now than under questioning.

Then name it. Good theme names carry meaning rather than labelling territory: It never switches off is a theme name; Technology is a topic heading. A short participant quote often makes an honest, vivid name, with an analytic subtitle if your department prefers formality. Avoid names that simply restate your interview questions, because that pattern suggests the analysis never left the topic guide.

Phase 6: writing up the analysis

The write-up is an analytic narrative, not an annotated quote collection. Each theme section should make an argument about the data, using extracts as evidence for analytic points, with your interpretation doing the connective work between them. The weakest findings chapters fall into a quote-then-paraphrase rhythm, restating each extract in slightly duller words. The strongest tell the reader what the pattern is, show it with well-chosen extracts from across the dataset, and say why it matters for the research question.

Your methodology and discussion chapters also make claims about the literature, and those need real, checkable citations. If you draft those sections in CiteDash's Thesis Editor, drafting is grounded in held full-text sources and the Fact Checker verifies claims against those sources before they reach you, with citations that resolve to real papers rather than free text. That covers the literature side; the analysis of your own transcripts remains your interpretive work, and should be reported as such.

Quality checks examiners apply to thematic analysis

Before submission, audit the analysis against the checks that come up in vivas:

  • Themes are patterns of shared meaning with a central concept, not topic summaries or interview-question headings.
  • Extracts are drawn from across the dataset, not from three favourite participants.
  • The stated approach matches the reporting: no reliability statistics inside a reflexive analysis, no claimed codebook you cannot produce.
  • The analysis answers the research question rather than summarising everything participants said.
  • A reflexivity statement says who you are in relation to the data and how that shaped interpretation.
  • The audit trail exists: dated code lists with definitions, thematic maps, and memos that show how the themes evolved.

A last word on honesty

Braun and Clarke's original paper closes with a 15-point quality checklist that is worth printing and keeping above your desk, not least because examiners have read it too. Almost every item on it reduces to the same demand: that the written account match the analysis you actually performed. Thematic analysis done step by step, with the decisions recorded as they were made, meets that demand automatically. Thematic analysis reconstructed backwards from a set of plausible-looking themes does not, and the seams show under questioning. Do the phases in order, keep the paper trail, and the viva becomes a conversation about your findings instead of an audit of your method.

Ready to try it on your own thesis?

Get Started Free

Do this in CiteDash

More guides