Skip to content
← Guides & helpGuides8 min readBy The CiteDash team

Coding interview transcripts: a practical method

A practical method for coding interview transcripts: codebooks, first and second cycle coding, memos, second coders, and a clean audit trail.

Twelve interviews, each an hour long, transcribed into a few hundred pages: this is the point where many qualitative theses stall. Coding is how you get from that pile to findings you can defend. It means labelling segments of transcript so that patterns become visible, comparable, and auditable. Done systematically, coding produces an analysis an examiner can trace from any claim in your findings chapter back to specific lines of specific transcripts. Done impressionistically, it produces themes you cannot justify under questioning, supported by whichever quotes you happened to remember.

This guide sets out a practical method: preparing transcripts, building a codebook, first and second cycle coding with worked examples, memos, second-coder options, saturation, and the audit trail. It stays neutral between analytic traditions where it can, and flags where reflexive thematic analysis, framework analysis, or grounded theory would diverge, because the right procedure depends on the methodology you have committed to.

What a code actually is

In Johnny Saldana's widely used coding manual, a code is described as a word or short phrase that assigns a summative, salient, essence-capturing attribute to a portion of data. Three properties follow from that definition. A code is short: a label, not a summary sentence. A code is analytic: it says something relevant to your research question, not merely which topic a passage touches. And a code is reusable: the same code applied across transcripts is precisely what makes patterns visible, which is why uncontrolled synonyms are so corrosive.

Keep the ladder of abstraction straight, because examiners probe it. Codes label segments of data. Categories group related codes. Themes, in thematic approaches, are patterns of shared meaning built out of categories and codes. A findings chapter that presents raw codes as themes has skipped two rungs of the ladder, and it reads that way.

Prepare your transcripts before you code

Decide your transcription convention and apply it uniformly. Verbatim transcription keeps every hesitation, false start, and laugh; intelligent verbatim lightly cleans the speech while preserving meaning. The choice depends on whether the texture of talk matters to your research question, and it belongs in your methods chapter either way. Add line numbers or timestamps so extracts can be cited precisely, both in your own coding and in the findings chapter.

Then anonymise: replace names with pseudonyms, mask identifying details such as employers and places, and keep the key linking pseudonyms to identities in a separate, secured location, as your ethics approval almost certainly requires. Check your consent coverage before analysis rather than after, because what participants agreed to constrains what you can quote and at what length. If you are still at the design stage, run recruitment and consent through a proper instrument; CiteDash's Collect handles consent-first research forms, and the consent records belong in the same audit trail as the analysis.

Choose your coding approach and say so

Inductive coding builds codes upward from the data. Deductive coding brings a predefined framework of codes to the data. Most real projects are a declared hybrid: a handful of a priori codes from the literature or the interview guide, plus explicit openness to codes the data insists on. Any of these is defensible; an undeclared mixture is not, because the reader cannot tell which findings were sought and which were found.

The choice must also align with your named methodology. Reflexive thematic analysis expects coding to be evolving and interpretive. Framework analysis expects a working analytical framework applied systematically, with charting. Grounded theory has its own sequence of open, focused, and theoretical coding tied to constant comparison. Nothing derails a viva faster than a methods chapter citing one tradition while the findings chapter performs another, so name the approach, cite it, and follow its logic.

Build a codebook

A codebook is the contract that keeps coding consistent: a structured list of codes, each defined well enough that the same segment would be coded the same way in month one and month four. For codebook and coding-reliability approaches it is central; even in reflexive work, a light version guards against definition drift without freezing interpretation. A workable codebook entry has six parts:

  • Name: short and specific, such as boundary intrusion rather than problems.
  • Definition: one or two sentences on the concept the code captures.
  • Inclusion criteria: what must be present in a segment for the code to apply.
  • Exclusion criteria: the near-miss cases the code does not cover, with a pointer to the code that does.
  • Example: a verbatim extract that clearly belongs.
  • Near-miss example: an extract that looks close but is excluded, and why.

A worked codebook entry

Name: mixed managerial signals. Definition: participant describes management communicating expectations about availability that conflict with stated policy or with other messages from the same management. Inclusion: explicit contradictions between policy and practice attributed to a manager or employer. Exclusion: conflicting expectations coming from clients or family, which belong under competing external demands. Example: my manager says switch off, but then messages at nine at night. Near miss: my clients expect answers on weekends, excluded because no manager is involved, coded instead as competing external demands.

Notice what the near-miss buys you: the boundary between two adjacent codes is now written down, so future-you, or a second coder, does not redraw it differently on a tired Friday. Every dispute you settle in the codebook is a dispute you do not have to settle from memory in the viva.

First cycle coding: a worked example

First cycle coding is the initial pass that fractures transcripts into labelled segments, and several styles can legitimately coexist within it. Descriptive coding labels what a segment is about. In vivo coding uses the participant's own words as the label, which is valuable when the phrasing itself is the finding. Process coding uses gerunds to capture actions and dynamics: justifying, deflecting, negotiating.

Take this excerpt from a hypothetical study of remote work: I answer emails from bed, honestly. My manager says switch off, but then messages at nine at night, so what am I supposed to do. A descriptive code might be after-hours email. An in vivo code: what am I supposed to do. A process code: deflecting responsibility for boundaries. The same lines can carry all three, and choosing among them is an analytic decision worth a memo. Work through every transcript systematically; let segments carry multiple codes where they earn them, and give uncodeable-but-interesting passages a holding code rather than silence, so nothing relevant simply disappears.

Second cycle coding: from codes to categories

Second cycle coding reorganises the first-cycle output into a smaller, more conceptual set. Pattern coding is the workhorse: grouping first-cycle codes that share an underlying logic and naming the group for the logic rather than the topic. Expect to merge synonymous codes, split overloaded ones, retire orphans that appeared once and mean little, and build a shallow hierarchy: categories on top, codes beneath, extracts at the bottom, every level traceable to the one below.

Resist the seduction of counting. Code frequency describes your dataset, and it can be worth reporting as description, but it is not proof of importance: a code applied once can matter more than one applied forty times, because qualitative significance lies in meaning and in relationship to the research question, not in tally marks. If you report counts, frame them accordingly, and never let a table of frequencies stand in for analysis.

Memos: the thinking that makes coding defensible

Memos are dated notes to yourself about analytic decisions: why a code was split, what a category actually means, a hunch about a pattern worth chasing, a doubt about your own influence on a particular interview. Write them continuously, from the first transcript onward, because they are the difference between I felt the themes emerged and here is the documented reasoning by which I built them.

Memos pay twice. During analysis, they stop definition drift and surface connections between codes you would otherwise lose between sessions. At write-up, a surprising amount of the findings chapter assembles itself from memo fragments, and the methodology chapter's claims about your audit trail are only true if the memos exist. Five minutes of memo after each coding session is the cheapest insurance in qualitative research.

Second coders, intercoder agreement, and saturation

Whether you involve a second coder depends on your school, not on a universal rule. Coding-reliability approaches expect it: a second coder codes a sample of transcripts against the codebook, agreement is computed (percent agreement at minimum; Cohen's kappa is the standard chance-corrected statistic), disagreements are discussed, and the codebook is revised until the definitions hold. Codebook approaches often use a lighter version of the same discipline. Reflexive thematic analysis explicitly rejects agreement statistics, on the argument that coding is interpretive and a second reader is a source of richer interpretation rather than a reliability instrument.

For a solo thesis, the honest options are: a full reliability procedure with a volunteer second coder; structured discussion of a transcript sample with a peer or supervisor, reported as exactly that; or a defended solo reflexive analysis with a strong audit trail. Any of these can pass a viva. Describing one while having done another cannot, because the codebook, the memos, and the story have to agree.

A related honesty question is when to stop. Saturation, the point at which further data yields no new codes or no new understanding of existing ones, is the most cited and most abused stopping rule in qualitative research; the methodological literature distinguishes code saturation, where no new codes appear, from meaning saturation, where no new depth accrues, and the second arrives later. The defensible thesis position is to document rather than declaim: track when new codes stop appearing as you move through transcripts, record the observation in a dated memo, and report it with appropriate modesty, as evidence about your dataset rather than a certificate the sample size earned.

The audit trail your examiner expects

Assemble the audit trail as you go; reconstructing it at submission time is miserable and unconvincing. It should contain:

  • Raw data: recordings and transcripts, stored per your ethics approval.
  • The codebook, in dated versions that show how definitions evolved.
  • The coded dataset: which extracts carry which codes.
  • Memos, dated, from first transcript to final theme.
  • The mapping from each claim in the findings chapter to its codes and extracts.
  • Consent records and the anonymisation key, stored separately from the data.

From codes to chapter

Two closing notes on the write-up. First, your findings chapter interprets your own data, but your methodology and discussion chapters make claims about the literature, and those need real citations: drafting them in the Thesis Editor keeps that work grounded, with claims verified against held full-text sources before they reach you and citations that resolve to real papers. Second, if AI tools assisted any part of your transcription or analysis, disclose it: the guide on writing an AI use disclosure statement covers the phrasing, and CiteDash generates an AI-use disclosure as part of thesis assembly. A clean coding method deserves a write-up that is equally easy to audit.

Ready to try it on your own thesis?

Get Started Free

Do this in CiteDash

More guides