Skip to content
← Guides & helpGuides8 min readBy The CiteDash team

How to justify your sample size

How to justify your sample size to examiners: power analysis for quantitative studies, saturation and information power for qualitative.

Somewhere in your viva, an examiner will ask how you arrived at your sample size. It is a favourite question because it is diagnostic: a candidate who can answer it understands the logic of their whole design, and a candidate who cannot has usually inherited a number from convenience and hoped nobody would ask.

The good news is that sample size justification is one of the most learnable parts of research design. There is a standard playbook for quantitative studies, a defensible set of frameworks for qualitative ones, and an honest way to handle the common situation where practical constraints, not calculations, set your ceiling. This guide covers all three, plus exactly what to write in the methodology chapter.

Why examiners care about sample size justification

For quantitative work, sample size determines statistical power: the probability that your study detects an effect of a given size if that effect is real. An underpowered study is not just weaker, it is uninformative in a specific way. A non-significant result tells you almost nothing, because the study could not have detected the effect even if it existed, and any significant result it does produce is more likely to overestimate the true effect. Examiners increasingly raise the question at the proposal stage, not just the final viva, so the justification is worth writing early.

For qualitative work, the concern is adequacy rather than power: whether the data are rich and varied enough to support the claims being made. In both traditions, the examiner is really asking one question. Did you decide your sample size for reasons, before collecting data, or did you rationalise it afterwards? The justification must therefore appear in the methodology chapter as a decision made in advance, with its inputs stated.

Quantitative studies: the a priori power analysis

The standard justification for a quantitative sample is an a priori power analysis, run before data collection. It takes four inputs and returns the sample size you need:

  • The statistical test you plan to run, because power calculations are test-specific: a two-group comparison, a correlation, a regression, and a factorial ANOVA all have different arithmetic.
  • The significance level (alpha), conventionally set at .05 in most fields.
  • The desired power, conventionally .80, meaning you accept a one-in-five chance of missing a true effect; higher-stakes work often specifies .90.
  • The expected effect size, which is the input that actually requires judgement, and the one the next section is about.

Running and reporting the power calculation

Free software runs the calculation; G*Power is the usual choice in many departments, and most statistics packages include equivalents. The calculation itself takes minutes. Report all four inputs and the resulting number in your methods chapter, then recruit to that number plus a margin for dropouts and unusable responses, stating the margin too. An examiner who sees a complete, pre-specified power analysis usually moves on to other questions, because the paragraph answers the challenge before it is made. Keep the calculation output with your research records, because the viva question is easier to answer with the original in hand.

How to choose an effect size honestly

Alpha and power are conventions; the effect size is a judgement, and it is where power analyses go wrong. Guessing a large effect because it makes the required sample conveniently small is transparent to examiners. Defensible sources, in rough order of strength:

  • Meta-analyses in your area, which pool effect estimates across studies and are the strongest anchor when one exists.
  • Effect sizes reported in the specific studies your design replicates or extends, adjusted downward, since published effects tend to skew optimistic.
  • The smallest effect size of interest: the smallest effect that would actually matter in your context, which reframes the question from prediction to relevance and is increasingly the recommended approach.
  • Conventional benchmarks (small, medium, large) as a last resort, clearly labelled as conventions rather than estimates.

A caution about pilot studies

Pilot studies deserve a specific warning. A small pilot estimates the effect size so imprecisely that powering a main study on it alone is unreliable, and the practice is widely cautioned against in the methodological literature. Use pilots to test procedures, instruments, and recruitment, and lean on the published literature or the smallest effect of interest for the effect size input. Whichever source you use, cite it. An effect size with a reference behind it is a justification; an effect size without one is a guess.

Why rules of thumb are weak sample size justifications

Every field circulates rules of thumb: thirty per group, ten participants per survey item, ten events per predictor in a regression. They persist because they are easy, and examiners discount them because they are not justifications. They are folklore with occasional statistical ancestry, and a rule of thumb knows nothing about your effect size, your design, or your measurement reliability, which are precisely the things that determine the sample you need.

If a rule of thumb is genuinely load-bearing in your field, cite it as corroboration alongside a real calculation, not instead of one. A sentence noting that the power analysis result is consistent with the disciplinary guideline is fine. The guideline alone, offered as the entire justification, invites exactly the viva question this post opened with.

Qualitative sample sizes: saturation and information power

Qualitative sample size cannot be computed, but it can be justified, and the word saturation on its own is no longer enough, because it is routinely asserted without evidence. Two frameworks give you something concrete to write. The first is a documented saturation process: state in advance how you will recognise saturation, for example no new codes arising across consecutive interviews, then show in the methods or an appendix how the criterion was met, with the code log to back it.

The second is information power, a planning framework built on a simple idea: the more information a sample holds that is relevant to the study, the fewer participants are needed. Rating your study against its five dimensions, and citing the framework, converts a bare participant count from an apparent accident into a reasoned design decision. State the planned range in advance, then report the final number and why it sufficed. The dimensions to rate are:

  • Aim: a narrow aim needs fewer participants than a broad one.
  • Specificity: participants with dense, specific experience of the phenomenon carry more information than a mixed group.
  • Theory: a study backed by established theory needs less data than one building from scratch.
  • Dialogue: strong interviewer technique and rich interviews reduce the number needed.
  • Analysis: in-depth analysis of individual cases needs fewer participants than cross-case pattern analysis.

Mixed methods and multi-strand sample sizes

Mixed methods theses need a justification per strand, and the strands answer to different standards: a power analysis for the quantitative arm, a saturation or information power argument for the qualitative arm. Do not let one strand's logic excuse the other. A large survey does not justify three interviews, and rich interviews do not rescue an underpowered experiment.

Sequential designs add one wrinkle worth stating in the methods chapter. When the qualitative strand samples from the quantitative one, interviewing survey outliers for instance, say how many you selected, by what criterion, and why that subset carries the information the explanation needs. The subsample has its own logic, and it deserves its own sentence.

When practical constraints limit your sample size

Sometimes the honest answer is that access, budget, or a rare population capped your sample below what the power analysis requested. The defensible move is transparency in three steps. First, still report the a priori calculation, so the reader knows what the design wanted. Second, state the constraint plainly and what you did to mitigate it: longer recruitment, wider criteria, a more sensitive design, a within-subjects switch if one was available. Third, report a sensitivity analysis: given the sample you actually achieved, the smallest effect your study could reliably detect, so the reader can calibrate the conclusions.

One practice to avoid: post hoc observed power calculated from your own results, which examiners recognise as uninformative because it is a restatement of the p-value. A sensitivity analysis answers the same worry legitimately. And write the limitation into the discussion chapter yourself, on your terms, before an examiner writes it for you on theirs. A limitation you name and quantify is a mark of competence; one the examiner finds first is a mark against the thesis.

What to write in your methodology chapter

Before the paragraph itself, a reality check: a sample size justification only holds if the analysis that follows matches the plan it was computed for. Plan the pipeline before recruiting: How to collect research data for a thesis covers instruments, consent, and getting clean data out the other end. When results come in, Data Analysis takes your CSV, runs named statistical tests, and generates the charts your results chapter needs; the full workflow is walked through in Dissertation statistics: CSV to Results chapter. Running the test you pre-specified, rather than the one that looks best afterwards, is what keeps the power analysis meaningful.

The sample size paragraph itself has a standard anatomy, whatever your tradition:

  • The target number and how it was determined, with the calculation inputs or framework named.
  • The source for the key judgement: the cited effect size, or the information power rating.
  • The recruitment margin for attrition and exclusions, stated in advance.
  • The final achieved sample and the reasons for any gap.
  • For qualitative work: the saturation criterion and where the evidence for it lives.

Answering the sample size question in the viva

Prepare a ninety-second answer with three beats. The decision: what number you targeted and the calculation or framework that produced it. The judgement: where the effect size or the information power rating came from, with the citation. The reality: what you achieved, what limited it, and what the sensitivity analysis says about what the study could detect. Delivered in that order, the answer demonstrates design thinking even when the final sample was smaller than you wanted.

Examiners do not require perfect samples. They require evidence that you understood the relationship between your sample and your claims, sized the study deliberately, and calibrated your conclusions to what the data could support. A justified small sample with honest limitations reads far better than a large sample with no reasoning attached, and the difference between the two is a paragraph you can write this week.

Ready to try it on your own thesis?

Get Started Free

Do this in CiteDash

More guides