EpidemiologyCDC WONDER: a goldmine for student epidemiologists
Most medical students believe that research starts with data collection, and that data collection means questionnaires, consent forms, an ethics submission and six months of chasing patients in an outpatient department. That is one way to do research. It is also the slowest way, and it is the reason so many student projects die somewhere between the proposal and the first hundred completed forms.
There is another way, and it is the one that has produced the majority of the papers our members have published in the last three rounds. It is called secondary data analysis, and the single most productive source for it is CDC WONDER — Wide-ranging ONline Data for Epidemiologic Research, a free public query system maintained by the United States Centers for Disease Control and Prevention.
What is actually in there
CDC WONDER is not one dataset. It is a front end onto roughly twenty of them. The one that matters most for a first paper is the Multiple Cause of Death file, which contains a de-identified record for every death registered in the United States from 1999 to the most recent finalised year. Every record carries the underlying cause of death coded to ICD-10, up to twenty contributing causes, and the decedent's age, sex, race and ethnicity, state and county of residence, level of urbanisation, and place of death.
That is a population of well over seventy million records, already cleaned, already coded to an international standard, already denominator-matched against census population estimates so that the system returns age-adjusted rates rather than raw counts. There is no ethics submission, because there is no identifiable human subject; there is no data-collection phase, because the collection was finished before you asked.
Alongside mortality, WONDER carries natality (every US birth certificate), cancer incidence, infant deaths linked to birth records, and several notifiable-disease surveillance streams. The query interface is identical across all of them, so the skill transfers.
Why it produces publishable papers
Editors publish descriptive epidemiology when it answers a question nobody has answered in that exact configuration before, and when the trend it reports is one a clinician can act on. WONDER makes both achievable for a student because the axis of novelty is combinatorial. Take a cause of death, cross it with a demographic stratum, cross it with a time window, and you have a cell that may genuinely never have been described.
Consider what our members have published: mortality trends in accidental poisoning across twenty-five years; deaths from polyneuropathies and peripheral nervous system disorders by census region; cardiovascular mortality among adults with concurrent diabetes stratified by rurality. None of those required a single new observation. Each required a well-posed question, a defensible extraction, and a competent analysis.
The methodological ceiling is higher than people assume, too. Once you have a time series of age-adjusted rates you can run Joinpoint regression to detect the years at which the trend changed direction, report annual percentage change with confidence intervals, and fit a forecasting model to project forward. That is the difference between a descriptive note and a paper a journal wants.
The five decisions that determine your paper
First, the cause. Choose an ICD-10 grouping that is clinically coherent and reasonably common — too rare and your cells fall below the CDC's suppression threshold, too broad and your finding says nothing. Rates built on fewer than twenty deaths are flagged as unreliable and should not be reported as point estimates; plan your strata so this does not happen.
Second, the time window. 1999 is the floor because ICD-10 replaced ICD-9 in US death certification that year, and a series that crosses the coding change is not a series. State your endpoint as the most recent finalised year and say so explicitly.
Third, the stratification. Two or three strata make a paper; six make a data dump. Pick the ones with a mechanism behind them — if you stratify by urbanisation you should be able to say in the discussion why access to care might differ.
Fourth, the rate. Report age-adjusted rates per 100,000 standardised to the 2000 US standard population, which is what WONDER produces by default and what every comparable paper uses. Crude rates across decades are uninterpretable because the population aged.
Fifth, the comparison. A number alone is not a finding. A number that rose while a comparable number fell is a finding.
Where students go wrong
The commonest error is treating a query result as an analysis. Exporting a table and describing it in prose is not a paper; it is an appendix. The analysis is the trend model, the stratum comparison, and the interpretation.
The second commonest is ignoring the suppression and unreliability rules and reporting a rate built on eleven deaths as though it were solid. Reviewers check this, and it is the fastest route to a desk rejection.
The third is failing to describe the extraction reproducibly. Your methods section must state the database and its version, the exact query date, the ICD-10 codes used, the years included, whether the cause was underlying or multiple-cause, and the standard population. A reader must be able to reproduce your table without emailing you.
The fourth is skipping the limitations that are intrinsic to death-certificate data: cause of death is assigned by a certifier who may be wrong, coding practice drifts over time, and race and ethnicity on death certificates are known to be misclassified for some groups. Naming these does not weaken your paper. Omitting them does.
A realistic first timeline
In our programme a CDC WONDER project runs across roughly ten weeks of part-time work. Two weeks to settle the question and read the twenty papers that bound it. One week to build and validate the extraction. Two weeks on analysis, including the trend modelling. Three weeks on the first full draft. Two weeks on internal review and revision before submission.
That is genuinely achievable alongside clinical rotations, which is the entire point. The barrier to a first publication has never been intelligence or access. It has been a question small enough to finish and a dataset that does not require permission. WONDER solves the second. Your mentor is there for the first.
