In 1986, researchers compared two treatments for kidney stones using the records of a British hospital group. Open surgery succeeded in 78% of its cases; the newer, less invasive procedure succeeded in 83%. The natural reading is that the newer procedure is better, and the natural reading is wrong. Split the patients by stone size and open surgery wins among small stones, 93% to 87%, and wins again among large stones, 73% to 69%. The treatment that is better for every patient is worse on paper.

This reversal, an association that holds in every subgroup yet flips when the subgroups are combined, is called Simpson’s paradox, and it is not a curiosity. It is the sharpest available demonstration of why AP Statistics insists so firmly on the vocabulary of lurking variables, confounding, and the limits of observational data.

How the reversal happens

Nothing in the arithmetic is exotic. An overall success rate is a weighted average of subgroup rates, weighted by how many patients each subgroup contributes. In the kidney-stone data, surgeons steered the difficult large stones toward open surgery and the easy small stones toward the new procedure. Open surgery’s overall figure was therefore an average dominated by hard cases, and the new procedure’s by easy ones. Each treatment’s overall rate says as much about its caseload as about its quality.

The instrument below holds the four subgroup success rates fixed at their published values. The slider controls only the case mix: what fraction of the difficult cases each treatment receives.

Each column is one treatment; the darker band is its large-stone caseload and the lighter band its small-stone caseload, with the marker showing the resulting overall success rate. The subgroup rates printed beside the bands never change. At an even case mix the overall comparison agrees with the subgroups, and surgery leads. Drag the slider toward the historical value of about 77% and the marker order reverses while every printed rate stands still. The paradox is not in the treatments; it is in the weighting, and the slider is the lurking variable made into a physical object.

The statistical moral

The paradox settles a question students often ask about the design unit: why so much ceremony about random assignment? Because random assignment is precisely the device that severs the link the slider controls. When treatments are assigned by coin flip, difficult cases distribute themselves evenly, the case mixes match, and the overall comparison means what it appears to mean. When treatments are assigned by human judgment, as in any observational study, the assignment mechanism is free to correlate with severity, and the aggregate numbers inherit that correlation. The kidney-stone surgeons were making sensible medical decisions; the data simply recorded their sensible decisions as a statistical illusion.

The exam vocabulary maps onto the picture exactly. Stone size is a lurking variable, associated both with the treatment received and with the outcome, which is the definition of confounding. The corrected analysis, comparing like with like inside each subgroup, is stratification. And the one-sentence conclusion the rubric wants is a causal disclaimer: because this is an observational study, the overall association between treatment and success cannot support a causal claim.

A last caution keeps the lesson honest: the paradox does not say that subgroup rates are always the truth and aggregates always the lie. Splitting data by a variable that is a consequence of the treatment, rather than a pre-existing condition, can manufacture reversals just as spurious in the other direction. Which level of the data answers the question depends on what caused what, and that dependence, formalized, is the modern field of causal inference. The AP course’s insistence on naming the study design before interpreting any number is the first chapter of that field.

The reversal has appeared in consequential places: in the 1973 Berkeley graduate admissions data, where the university appeared to favor men overall while individual departments slightly favored women who had applied to the most competitive programs, and in batting averages, where one player can trail another in both halves of a season yet lead for the year. A worthwhile exercise with the slider: find the exact case mix at which the overall rates tie, and note that nothing medical happens there at all; the tie is a fact about arithmetic, which is the entire point.