Skip to content
← Guides & helpGuides9 min readBy The CiteDash team

How to write a literature review in medicine

Write a medical literature review examiners trust: pick the review type, build a PICO search, appraise study designs, and synthesise by outcome.

A medical literature review is judged on whether someone else could repeat it. That is the distance between the review you wrote as an undergraduate, which summarised what you had read, and the chapter your examiners want, which states what the evidence currently supports, how much confidence that evidence can carry, and precisely where the space sits that your study occupies.

This guide covers the parts that are specific to medicine: choosing a review type, turning a clinical question into a search a librarian would recognise, appraising study designs rather than counting them, and synthesising by outcome instead of paper by paper. The general AI-assisted workflow is covered in how to write a literature review with AI, so this assumes you have that and need the clinical layer on top.

Carry one rule through the whole chapter: in medicine, a claim with no study design attached is not a finding. Every sentence you write about effectiveness, risk or prevalence should be traceable to a design, a population and an outcome measure.

What examiners expect from a medical literature review

The chapter has to do four things at once. It has to show command of the clinical field, demonstrate that you can separate strong evidence from weak evidence, justify the design of your own study, and stay current enough that a reviewer cannot name a recent trial you missed.

The currency requirement is the one that catches people. Clinical evidence moves, and a chapter written a year before submission can be overtaken by a single large trial. Build a refresh into your plan: rerun the search close to submission, record the date you ran it, and add a line to your methods stating when the search was last updated.

Command of the field is easier than it looks once you stop trying to cover everything. A narrow review that handles one question thoroughly reads as expertise. A broad review that touches forty topics reads as a reading list with paragraph breaks.

Justifying your own design is the connective tissue between chapters. If the review shows that every existing study of your outcome relied on self-reported measures, your decision to use an objective measure stops being a preference and becomes an argument. Write the review so that a reader arriving at your methods chapter already expects the design you chose.

Choose your review type before you start searching

The label you choose commits you to a method, and examiners will hold you to it. Decide first, write it into your protocol, and do not quietly change type halfway through because the searching got hard.

For a doctoral thesis the choice usually turns on scope and time rather than ambition. A full systematic review on a broad question can consume a year by itself, which is defensible when the review is the thesis and reckless when it is chapter two. Settle this with your supervisor before you commit, because downgrading later is awkward to explain and upgrading later means re-screening everything you have already been through.

  • Narrative review: expert overview of a field, no exhaustive search claim. Appropriate for a background chapter, weak if you present it as comprehensive.
  • Scoping review: maps what kinds of evidence exist and where the gaps are. Useful when the question is broad or the literature is young.
  • Systematic review: pre-specified protocol, exhaustive search, formal appraisal, reproducible screening. The strongest claim and the heaviest workload.
  • Rapid review: systematic methods with declared shortcuts, such as one screener or a limited date range. Declare every shortcut you took.
  • Umbrella review: a review of existing systematic reviews, used when the primary literature has already been synthesised several times.

Protocols, registration and reporting standards

If you are writing a systematic review, registration on PROSPERO before screening starts is the norm in health research, and reporting follows PRISMA. Both are about locking your decisions in before you know what the results look like, which is the whole point of the method.

The practical benefit is that a protocol turns a vague search into a set of instructions you can hand to a second screener. Write down your question, databases, search strings, date limits, language limits, inclusion and exclusion criteria, appraisal tool and synthesis approach before you screen a single title. If you are running the screening and extraction in CiteDash, the PRISMA walkthrough covers how the counts and the flow diagram fall out of the process rather than being reconstructed at the end.

Turn the clinical question into PICO, then into a search string

PICO is not a formatting exercise. It is how you convert a question into search terms and, later, into the columns of your evidence table.

Take a worked example. If your question is whether early mobilisation reduces delirium in older adults after hip fracture surgery, the components fall out cleanly.

  • Population: adults aged 65 and over, post hip fracture surgery.
  • Intervention: early mobilisation, however each study defines the timing.
  • Comparator: usual care or delayed mobilisation.
  • Outcome: incidence and duration of postoperative delirium, plus whatever secondary outcomes you care about.
  • Study design limit, if you want one: randomised trials and prospective cohorts only.

Building a search a librarian would recognise

Each PICO element becomes a block of synonyms joined with OR, and the blocks are joined with AND. Combine controlled vocabulary with free text, because indexing lags and recent papers may not carry the subject headings yet. Truncation catches variant endings, and phrase searching keeps multi-word concepts intact.

Two habits separate a defensible search from an improvised one. First, save the exact string for every database, along with the date you ran it and the number of records returned, because that goes into your methods and your flow diagram. Second, do not filter your outcome term into the search unless you have to: outcomes are often reported without appearing in titles or abstracts, so an outcome block can silently drop relevant trials.

Once the string is stable, run it wide. Literature Finder searches an internal corpus alongside live OpenAlex, PubMed, Semantic Scholar and arXiv, which is enough to build the core of a clinical search. If your handbook or your subject librarian requires additional databases, run those separately in their own interfaces and record them as separate lines in your methods, with their own strings and counts.

Screening with inclusion criteria you can defend

Screen in two stages: titles and abstracts first, then full text for everything that survives. Record a reason for every full-text exclusion, because those reasons populate the bottom of your PRISMA flow diagram and are the first thing a methodologically minded examiner will look at.

Write criteria that a second person could apply without asking you what you meant.

  • Population limits: age bands, diagnosis, care setting, comorbidity exclusions.
  • Intervention limits: dose, timing, delivery, and what counts as a close enough variant.
  • Design limits: which designs you accept and whether you allow non-randomised evidence.
  • Outcome limits: which measured outcomes make a study eligible, and which reporting formats you can extract from.
  • Practical limits: language, date range, publication type, and whether conference abstracts count.

Appraise the evidence instead of describing it

A design hierarchy is a starting heuristic, not a verdict. A small, poorly concealed randomised trial can be less informative than a large, well-controlled cohort study. Say what makes each body of evidence trustworthy or shaky, in your own words, using an established instrument so your judgements are anchored to something.

The standard tools in health research are RoB 2 for randomised trials, ROBINS-I for non-randomised studies of interventions, the Newcastle-Ottawa Scale for observational designs, CASP checklists across several designs, and GRADE for rating certainty across the whole body of evidence rather than individual studies. Pick one appraisal instrument per design type and apply it consistently.

Appraisal is the reason abstracts are not enough. Allocation concealment, attrition, blinding and analysis decisions live in the methods and the supplementary material. This is also why grounding in CiteDash works the way it does: discovery ranges wide across metadata and abstracts, but a claim is only verified against full text you actually hold, whether that is an open access paper or a PDF you uploaded.

Report the appraisal in the running text, not only in an appendix table. A sentence noting that blinding was not feasible and outcome assessment was unblinded in three of the five trials tells an examiner more than a column of scores, and it hands you the language you will need in your discussion when you compare your own findings against the existing evidence.

Synthesise by outcome, not study by study

The most common failure in a medical literature review is the study parade: a paragraph per paper, each opening with an author name and a year, no argument connecting them. It is easy to write and painful to read, and it hides whether you understood the field.

Organise instead around outcomes, populations or mechanisms. Under each outcome, state what the evidence shows, then account for the studies that disagree and explain why they might. Differences in dose, timing, measurement instrument, follow-up length or population usually explain more of the disagreement than quality alone.

An evidence matrix makes this tractable. Put studies in rows and your extraction fields in columns: design, sample, setting, intervention detail, comparator, outcome measure, effect direction, appraisal verdict. Synthesis Lab builds that matrix across your saved papers so you can read down a column and see the pattern instead of holding twenty PDFs in your head. Contradiction is content, not an embarrassment, and a paragraph explaining why two trials disagree is worth more than five that summarise agreement.

Write each synthesis section to a repeating shape: what the evidence supports, how strong that evidence is, what disagrees and why, and what remains untested. Four moves per outcome produce a chapter that argues rather than summarises, and they make the closing paragraph almost automatic, because the untested items have been accumulating as you went.

Retractions, corrections and preprints in clinical literature

Retracted trials keep getting cited long after retraction, usually because the citing author copied a reference from another paper without rechecking it. In medicine that is a serious error, because the claim you are propagating may have been withdrawn for data fabrication. Check every reference you inherit rather than every reference you personally found, since the inherited ones are the risk.

CiteDash flags retracted papers with a badge and blocks them from being cited, and the free retraction check will screen a DOI if you are working outside the app. Also watch for expressions of concern and corrections, which do not remove a paper but change how much weight it can carry.

Preprints are useful for currency and dangerous for authority. They have not been peer reviewed, and the version you read may differ from the version eventually published. If you cite one, label it as a preprint in the text, note the version and date, and check before submission whether a peer-reviewed version has appeared with different numbers.

Build the screen into your workflow rather than saving it for the end. Checking a DOI as you add a paper to your library costs seconds. Auditing two hundred references in the week before submission costs a weekend, which is why it usually gets skipped.

Citation style and the last checks

Medical writing usually runs on numbered styles. Vancouver and AMA are both among the twelve styles Reference Manager supports, so pick whichever your handbook or target journal specifies and stay in it. Numbered styles are unforgiving about ordering, so let the tool renumber when you move text rather than editing numbers by hand.

Before the chapter goes to your supervisor, work through this list.

  • Every effectiveness claim names a design and a population, not just an author and a year.
  • Search strings, databases and dates are recorded in the methods, with counts that match your flow diagram.
  • Every included study has an appraisal verdict, and the weak ones are described as weak in the text.
  • No reference was copied from another paper without you opening the source.
  • Preprints are labelled, retractions are screened, and corrections are accounted for.
  • The synthesis is organised by outcome, and each section ends with what is still unknown.

Common mistakes in a medicine literature review

Most of the marks lost in this chapter come from a short list of habits rather than from ignorance of the field.

  • Calling a narrative review systematic because it felt thorough.
  • Reporting effect sizes without the confidence intervals or the sample sizes that make them interpretable.
  • Treating a meta-analysis as settled without checking heterogeneity and how the included studies were appraised.
  • Reviewing the evidence and then choosing a study design that does not follow from the gap you identified.
  • Leaving the search unrepeated for a year and letting an examiner find the trial you missed.

Ready to try it on your own thesis?

Get Started Free

Do this in CiteDash

More guides