Skip to content
← GlossarySampling & data collection

Secondary Data

Existing data originally collected by someone else or for another purpose, reused and reanalyzed to answer a new research question.

Secondary data are data the researcher did not collect: government statistics, archived survey datasets, organizational records, or data published by other researchers. Reusing them saves time and money, allows access to larger or longer-running datasets than a student could gather alone, and avoids burdening new participants. The tradeoff is fit: the variables, definitions, and time periods were chosen for someone else's purposes, so they may only approximate the concepts your question needs, and quality must be assessed rather than assumed.

In a thesis, secondary data demand their own methodological rigor. Describe the source, how and when the data were originally collected, and any known limitations. Explain how the available variables map onto your concepts, and note where the fit is imperfect. Treating the dataset critically, rather than as given truth, is what separates competent secondary analysis from convenience.

Writing the thesis this term belongs to?

CiteDash takes a thesis from research question to a compiled document, with AI that cites only real papers and verifies every claim against its source.

Try CiteDash freeNo sign in required
Guides & help