Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Simpson's paradox is a reversal: an association in every subgroup flips or vanishes when subgroups are pooled. Both numbers are correct; one answers the question.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Simpson’s paradox occurs when an association observed within every subgroup of a dataset reverses or disappears once the subgroups are combined. A treatment can succeed more often than its comparator among small stones and among large stones, yet succeed less often overall, because the two treatments were given to different mixes of patients. Both numbers are correct; only one answers the causal question.
The reversal is not an arithmetic error, a coding mistake, or a small-sample artefact. It follows deterministically from weighted averages with unequal weights, and it can be reproduced exactly in a spreadsheet. What the reversal does not tell you is which of the two answers to report — that is decided by the causal structure behind the data, not by the data.
Edward H. Simpson gave the effect its name through a 1951 paper in the Journal of the Royal Statistical Society, Series B, titled “The Interpretation of Interaction in Contingency Tables”. Simpson did not discover the reversal. Karl Pearson noted it in 1899 and George Udny Yule described it in 1903, which is why statisticians also call it the Yule-Simpson effect, and why “Simpson’s paradox” is a standard example of Stigler’s law of eponymy.
Colin R. Blyth’s 1972 paper in the Journal of the American Statistical Association, “On Simpson’s Paradox and the Sure-Thing Principle”, supplied the word “paradox” and set the reversal against Leonard Savage’s sure-thing principle — the decision axiom that an act preferred in every state of the world should be preferred without knowing the state. Blyth’s contribution was to show that the sure-thing principle governs decisions under a fixed causal setup, not comparisons of observed conditional frequencies across subgroups of differing composition. Applying it to raw subgroup rates is precisely where researchers go wrong.
A group’s overall success rate is a weighted average of its rates within strata, and the weights are that group’s own distribution across those strata. Two groups with different distributions are therefore averaged with different weight vectors, and the comparison of two weighted averages need not follow the comparison of the underlying rates.
Two conditions are jointly necessary for a full reversal. First, the stratifying variable must be associated with the outcome, so that stratum-specific rates actually differ. Second, the stratifying variable must be associated with group membership, so that the weight vectors differ. Remove either condition and the crude comparison agrees in sign with the stratified one.
The second condition is worth stating precisely, because it explains why randomised trials rarely display a reversal: the crude risk difference equals a weighted average of stratum-specific risk differences only when the covariate distribution is the same in both arms. Randomisation makes those distributions equal in expectation, so a reversal in a well-conducted, adequately sized trial reflects chance imbalance rather than structure. In an observational cohort, unequal covariate distributions across exposure groups are the normal condition, not the exception.
Charig and colleagues reported in the British Medical Journal in 1986 on four approaches to removing kidney stones, and the comparison of two of them — open surgery and percutaneous nephrolithotomy — is the most widely reproduced clinical instance of Simpson’s paradox. Open surgery succeeded in 273 of 350 cases (78%); percutaneous nephrolithotomy succeeded in 289 of 350 (83%). Read at that level, the less invasive procedure wins.
Stratify by stone diameter and the ordering flips in both strata. For small stones, open surgery succeeded in 81 of 87 cases (93.1%) against 234 of 270 (86.7%) for percutaneous nephrolithotomy. For large stones, open surgery succeeded in 192 of 263 (73.0%) against 55 of 80 (68.8%). Open surgery is better for small stones, better for large stones, and worse overall.
The mechanism sits in the denominators. Large stones — the harder cases — made up most of the open-surgery series (263 of its 350 cases, 75%) and a small minority of the percutaneous series (80 of its 350, 23%); the two series also came from different periods, open surgery from 1972–80 and percutaneous nephrolithotomy from 1980–85, as the procedure was phased in. Stone size predicts both which procedure a patient received and whether it succeeded, and does not lie on the pathway from treatment to outcome, which makes it a textbook confounder. Here the stratified estimate is the one to report.
Bickel, Hammel and O’Connell published the best-known non-clinical instance in Science in 1975, under the title “Sex Bias in Graduate Admissions: Data from Berkeley”. Across the 1973 graduate cohort, roughly 44% of 8,442 male applicants and roughly 35% of 4,321 female applicants were admitted. Department by department, the disparity largely vanished, and after pooling with proper weighting the authors reported a small bias in favour of women. Women applied disproportionately to departments with low admission rates for all applicants.
Berkeley is a genuine aggregation reversal and a poor advertisement for the conclusion usually drawn from it. The stratified estimate is the causally correct one only if department is a confounder — a common cause of applicant sex and admission. If instead the field a woman applies to is itself shaped by the discrimination under study, department is a mediator, and conditioning on it removes part of the effect being estimated rather than a bias. Bickel and colleagues raised the upstream question themselves — they attributed the aggregate disparity not to the admissions committees but to “prior screening at earlier levels of the educational system”, and warned that no demonstrable bias in admissions is no grounds for concluding there is none elsewhere in the educational process. The textbook retelling usually drops that caveat.
That ambiguity is the honest general lesson. The data cannot tell you whether department is a confounder or a mediator. Only an argument about how the world works can.
Simpson’s paradox presents as a choice between two numbers, and no property of the numbers resolves it. The crude estimate and the stratified estimate are both unbiased answers to different questions, and the question a researcher wants answered is a causal one.
The formal tool is a directed acyclic graph plus Judea Pearl’s back-door criterion: a set of covariates is sufficient for adjustment when it blocks every back-door path from exposure to outcome and contains no descendant of the exposure. Applied to stone size, the criterion licenses stratification. Applied to a post-treatment variable, it forbids it.
This is why “we also present the stratified analysis” is not a safe default. Stratifying on a collider — a variable caused by both exposure and outcome — manufactures an association inside strata where none exists in the population, producing a reversal that moves away from the truth rather than towards it. A reversal is evidence that something structural is happening. It is not evidence that the finer-grained number is the better one.
The crude estimate is correct whenever the stratifying variable lies on the causal pathway from exposure to outcome. Consider a drug that lowers mortality by lowering blood pressure. Stratify on post-treatment blood pressure and the drug will look inert within every stratum, because within a stratum the mechanism has been held fixed. The unstratified comparison is the total effect, and the total effect is what a clinician needs.
The same logic governs stratification on any post-baseline variable: treatment adherence, dose received, tumour response, time to a landmark. Each can produce a striking reversal and each is a descendant of the exposure. Timing is the most useful practical filter available — a variable measured after exposure is a candidate mediator or collider, not a confounder, and a reversal it produces is an artefact of the conditioning rather than a discovery.
Reversals produced by conditioning on survival, on remaining in follow-up, or on being admitted to hospital belong to the same family, and they overlap with immortal time bias whenever the stratifying variable encodes elapsed time.
Confounding is a causal structure; Simpson’s paradox is an observable arithmetic symptom. Neither implies the other, and treating the two as synonyms causes two distinct errors.
Confounding usually does not produce a reversal. Most confounding inflates or attenuates an estimate without crossing the null, so the absence of a reversal is worthless as evidence that a cohort is unconfounded. A crude and an adjusted odds ratio that agree in sign can still differ by a factor that changes the clinical conclusion entirely.
Reversals also occur without confounding. A stratifier that is a mediator or a collider will produce one, and so, in a milder form, will non-collapsibility of the effect measure. Writing “Simpson’s paradox was present, so we adjusted” describes a symptom and asserts a diagnosis, which is the same defect behind the common complaint that my effect disappeared after adjusting.
Simpson’s paradox has a continuous form, usually called the reversal paradox, in which a regression slope fitted to pooled data has the opposite sign to the slopes fitted within clusters. Scatter plots showing within-group lines sloping one way and an overall line sloping the other are the visual signature.
Three named versions recur in the literature. William Robinson’s 1950 paper in the American Sociological Review, “Ecological Correlations and the Behavior of Individuals”, found the correlation between the proportion of foreign-born residents and literacy positive across the 48 US states and the District of Columbia and negative at the individual level, which established the ecological fallacy. Frederic Lord’s 1967 note in Psychological Bulletin described two statisticians reaching opposite conclusions about weight gain from identical data — one using change scores, one using analysis of covariance — a disagreement now known as Lord’s paradox and resolved only by stating which causal quantity is wanted.
The third version is routine and unnamed: a mixed model whose random intercepts absorb between-cluster variation reports a within-cluster slope, while a naive pooled regression reports a blend of within- and between-cluster effects. Separating the two by group-mean centring is standard practice, and the underlying failure of independence is the one described in pseudoreplication.
Non-collapsibility is a mathematical property of the odds ratio and the hazard ratio, and it produces stratum-versus-crude discrepancies with no confounding present at all. If a covariate is a genuine risk factor for the outcome but is distributed identically across exposure groups, the crude odds ratio still differs from the common stratum-specific odds ratio. Greenland, Pearl and Robins set out the distinction in “Confounding and Collapsibility in Causal Inference”, Statistical Science, 1999, and it remains the standard reference for why “the estimate changed on adjustment” is not proof of confounding.
The direction matters for interpretation. Non-collapsibility of the odds ratio attenuates the crude estimate towards the null relative to the conditional estimate; it does not carry the estimate across the null. A sign reversal in an odds ratio therefore still requires confounding, mediation or collider conditioning to explain it, while a change in magnitude may require no bias explanation whatsoever.
The hazard ratio behaves similarly and worse. Because a hazard ratio conditions on survival to each time point, unmeasured heterogeneity in frailty depletes the higher-risk arm first, so hazard ratios drift towards the null over follow-up even under a constant underlying effect. That mechanism is one of the standard explanations offered when survival curves cross.
Pavlides and Perlman quantified a base rate in “How Likely Is Simpson’s Paradox?”, The American Statistician, 2009. Placing a uniform distribution over the probability simplex of 2x2x2 contingency tables, they computed the proportion exhibiting a reversal as approximately 0.0166 — about one table in 60.
That figure is a mathematical base rate under a specific prior, not an empirical incidence, and it should not be quoted as the frequency in real research. Real datasets are not uniform draws from a simplex: the two conditions required for a reversal are exactly the two conditions that define a confounder, so the paradox is systematically over-represented wherever allocation is non-random. Registry, claims and electronic health record analyses concentrate the risk, because treatment there is assigned by clinical judgement on the same variables that predict outcome.
The practical implication is that one table in 60 is a floor for the settings that matter, and the frequency in any specific analysis depends entirely on how strongly the covariate predicts both allocation and outcome.
Simpson’s paradox is detectable from a submitted manuscript far more often than authors expect, because the evidence is usually printed in the paper’s own tables.
Recompute the crude table from the subgroup table. Subgroup results are typically reported as counts, or as rates with denominators. Summing them reproduces the pooled comparison, which can then be checked against the abstract. A discrepancy in sign between the recomputed pooled estimate and every reported subgroup estimate is a reversal, full stop.
Compare the crude estimate with a Mantel-Haenszel estimate. The Cochran-Mantel-Haenszel procedure pools the stratum-specific estimates themselves rather than the raw counts, so a difference between exposure groups in how they are distributed across strata cannot pull the summary the way it pulls the crude estimate. A large gap between the crude and Mantel-Haenszel estimates identifies aggregation effects even where the sign does not flip, and the gap is computable from a published stratified table without any raw data.
Inspect the allocation margins, not only the outcome margins. The tell is a covariate whose distribution differs sharply between exposure groups — 75% of one arm and 23% of the other in the kidney stone data. Baseline tables report this directly, and standardised mean differences above about 0.1 are the conventional flag.
Look for silent pooling. Multi-centre trials analysed without a centre term, cohorts assembled from several waves or registries, sequencing runs merged across batches, and meta-analyses that add raw numerators and denominators across trials rather than combining effect estimates all create the conditions for a reversal, and none of them announce themselves in the results section.
Check whether the stratifier precedes the exposure. A reversal produced by stratifying on a post-baseline variable is a different finding from one produced by stratifying on a baseline confounder, and the manuscript rarely says which it has.
Reviewers rarely use the phrase “Simpson’s paradox” in a report. The comments arrive in this form instead.
“The direction of the effect in Table 3 is inconsistent with the direction reported in the abstract, and the discrepancy is not addressed.” “The crude and adjusted estimates differ substantially; the authors should state which they consider to estimate the target effect, and why.” “Data were pooled across sites without adjustment for site, and site-specific results are not shown.” “The subgroup results contradict the primary analysis. Please provide the stratum-specific numbers with denominators.” “Treatment allocation appears strongly related to disease severity; the unadjusted comparison is not interpretable.” “Adjustment for [post-treatment variable] may have removed the effect of interest.”
Two of these decide papers. A reviewer who finds subgroup results that all point the opposite way to the headline claim will generally recommend rejection rather than revision, because the finding under review has not merely been imprecisely estimated — it has been stated in the wrong direction. And a reviewer who suspects the stratified analysis was chosen because it gave the wanted answer will say so in the confidential comments to the editor, where the author never sees it. Reviewers in epidemiology and outcomes research are trained to make exactly this check, and it is fast.
Report the reversal explicitly rather than choosing one estimate and staying quiet. A pooled estimate presented alone, when the subgroup estimates run the other way, reads as concealment once a reviewer recomputes the table — and reviewers do recompute the table.
Give the crude and adjusted estimates together, with the stratum-specific numerators and denominators, so a reader can reproduce both. State which estimate you regard as answering the research question and give the causal argument for that choice, naming the role you are assigning to the stratifying variable: confounder, mediator, collider or effect modifier. State when the stratifying variable was measured relative to exposure. If the variable might plausibly be a mediator, present both the total and the adjusted effect and say what each means, rather than asserting one.
Distinguish a reversal from effect modification. Stratum-specific effects that differ in magnitude but agree in sign are heterogeneity, and heterogeneity is a finding to characterise, not a bias to remove. Language discipline matters here, and the boundary between describing an association and asserting an effect is where most of these manuscripts fail: see correlative, not causal and the Table 2 fallacy for the two adjacent failures.
Simpson’s paradox has no statistical resolution, and claims to the contrary deserve scepticism. Judea Pearl argues that the paradox dissolves entirely once causal assumptions are made explicit, because the back-door criterion then names the correct adjustment set. Critics respond that the criterion relocates the difficulty rather than removing it, since the graph itself is an untestable assumption and the observed reversal supplies no evidence about which graph is right.
Both positions are defensible and the practical consequence is the same: the choice of estimate rests on subject-matter argument a reader can dispute, so the argument must be written down.
Two further caveats apply. The 0.0166 base rate from Pavlides and Perlman describes random tables under a uniform prior and does not estimate how often reversals occur in published research; we are not aware of an empirical incidence figure for published research, and the denominator — analyses in which a reversal could have been detected — is not observable from the published record. And a reversal seen in small strata may be sampling noise: before rebuilding an analysis around one, check that the stratum-specific confidence intervals actually exclude the pooled point estimate, which in thinly populated strata they frequently do not. Statistical objections of this kind are among the most common reasons manuscripts are sent back.
Confounding · Collider bias · Mediator versus confounder · Table 2 fallacy · My effect disappeared after adjusting
Checked before submission by causal language discipline, which compares crude and adjusted estimates reported in the same manuscript, flags subgroup results that contradict the abstract, and names the causal role being assumed for each adjustment variable.
Simpson’s paradox in statistics is a reversal of association under aggregation: a relationship that holds within every subgroup changes direction or disappears when the subgroups are combined into one table. The reversal arises from weighted averaging with unequal weights, and choosing between the two results requires a causal argument rather than a statistical test.
Simpson’s paradox happens when combining groups flips a result. If one treatment works better for small stones and better for large stones, but surgeons gave it mostly to the hard cases, its overall success rate can look worse. Nothing was miscalculated. The groups being averaged were not comparable in composition.
Charig and colleagues reported in the British Medical Journal in 1986 that open surgery for kidney stones succeeded in 78% of cases against 83% for percutaneous nephrolithotomy, yet open surgery won within both stone-size strata: 93.1% against 86.7% for small stones, and 73.0% against 68.8% for large stones. Stone size drove allocation.
Simpson’s paradox is called a paradox because it appears to violate the intuition, formalised as Savage’s sure-thing principle, that something better in every subgroup must be better overall. Colin Blyth supplied the label in 1972. The intuition holds for decisions under a fixed causal setup and fails for comparisons of subgroup rates across groups of differing composition.
The Yule-Simpson effect is another name for Simpson’s paradox, crediting George Udny Yule, who described the reversal in 1903, and Karl Pearson, who noted it in 1899, alongside Edward H. Simpson’s 1951 paper on interaction in contingency tables. The phenomenon named is identical: association within strata reversing on aggregation.
A trend that reverses after grouping means the groups differ in composition on a variable related to the outcome, so the pooled comparison averages unlike things. The reversal itself is arithmetic. Whether the grouped or the ungrouped estimate answers your question depends on whether that variable causes the exposure or is caused by it.
Aggregation bias and the reversal paradox are alternative names for Simpson’s paradox, the second used mainly for continuous data, where a regression slope fitted to pooled observations has the opposite sign to the slopes fitted within clusters. The ecological fallacy, described by William Robinson in 1950, is the same effect at the level of geographic units.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect