Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Confounding occurs when a common cause of both exposure and outcome distorts their association. It exists in the data before analysis, unlike collider bias.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Confounding occurs when a variable causes both the exposure and the outcome, creating an association between them that does not reflect an effect of one on the other. Age confounds the relationship between grey hair and heart disease: age causes both, so the two appear associated even though neither causes the other.
That definition is short, and almost every practical difficulty in observational research comes from applying it correctly. This page covers the three conditions a confounder must satisfy, how to choose an adjustment set, the difference between confounding and the biases that resemble it, why adjustment can make an estimate worse, how confounding by indication defeats the obvious remedy, how the problem is detected in a submitted manuscript, how to bound the part of it you cannot remove, and what reviewers say when it has been handled badly.
A variable is a confounder of the relationship between an exposure and an outcome when all three of the following hold.
It causes the exposure. Not merely correlates with it. A variable associated with the exposure for some other reason — because both share a cause, or because the exposure causes it — is not a confounder.
It causes the outcome, by a path that does not run through the exposure.
It does not lie on the causal pathway between exposure and outcome. A variable that the exposure causes, which in turn causes the outcome, is a mediator rather than a confounder, and adjusting for it removes part of the effect you are estimating.
All three are required. Two of the three describe something else, and the something else usually needs opposite treatment.
The three conditions are statements about the world, not about the dataset, so they are settled by subject-matter argument and by the temporal order of measurement rather than by any procedure run on the data. A fourth condition is implicit and frequently decisive in practice: the variable must have been recorded, on the scale at which it acts, before any analytic method can use it. Disease severity satisfies the three conditions in almost every study of a prescribed drug; it fails the fourth one in most administrative claims databases, where the reason a clinician chose a treatment is not a field.
The most common working definition among researchers — a variable associated with both the exposure and the outcome — is wrong, and it is wrong in a way that causes real errors.
Association is symmetric; causation is not. A variable can be associated with both the exposure and the outcome because it is caused by both, which makes it a collider, and adjusting for it introduces bias rather than removing it. Or it can be associated with both because it sits on the pathway between them, which makes it a mediator, and adjusting for it removes the effect.
Three variables can have identical statistical relationships with your exposure and outcome and require three different treatments — adjust, do not adjust, and do not adjust — depending on causal facts that no correlation matrix contains. This is the single most important thing to understand about confounding: it is defined causally, not statistically, and no diagnostic run on the data can identify it.
The formal version of that impossibility is worth stating precisely, because it closes off a whole family of proposed fixes. A chain A → C → Y and a fork A ← C → Y imply exactly the same joint distribution over the three variables — they are Markov-equivalent — so a mediator and a confounder are indistinguishable to any test, any information criterion and any machine-learning model that sees only those three columns. The arrow direction has to come from outside the data, which is why the tool for recording it is a directed acyclic graph rather than a correlation matrix.
Confounding, mediation and collider bias can all be present in a single observational cohort, and telling them apart is a matter of when each variable was determined and what caused it. Consider a study of whether a drug reduces mortality in an observational cohort.
Patients with more severe disease are both more likely to receive the drug — because clinicians treat sicker patients more aggressively — and more likely to die. Severity causes the exposure, causes the outcome, and does not sit between them. It is a confounder, and unadjusted analysis will make the drug look harmful.
Numbers make that concrete. Take a constructed cohort of 1,000 patients. Among the 400 with severe disease, 300 receive the drug and 100 do not, with mortality of 30% treated and 40% untreated. Among the 600 with mild disease, 100 receive the drug and 500 do not, with mortality of 5% treated and 10% untreated. Within each stratum the drug reduces risk — a risk ratio of 0.75 in the severe stratum and 0.50 in the mild stratum. Pooled without stratification, 95 of 400 treated patients die (23.8%) against 90 of 600 untreated (15.0%), a crude risk ratio of 1.58. Confounding by severity has not merely attenuated the estimate; it has reversed its direction, which is the same arithmetic that produces Simpson’s paradox.
Now consider blood pressure measured after treatment starts. The drug lowers blood pressure, and blood pressure affects mortality. This is a mediator: adjusting for it removes exactly the mechanism by which the drug works, and the drug will appear to do nothing.
Now consider hospital admission during follow-up. Both severe disease and drug side effects can cause admission. Restricting the analysis to admitted patients conditions on a collider and can produce an association with no causal basis at all.
Severity, post-treatment blood pressure and admission may all be associated with both exposure and outcome in the data. Only one belongs in the adjustment set.
Confounding by indication is the form that defeats the standard remedy, because the variable driving treatment choice is a clinical judgement rather than a recorded measurement. A prescriber weighs frailty, trajectory, comorbidity, patient preference and an impression formed in the consulting room; the database records an ICD code and a dispensing date. Adjustment reaches the code, not the judgement.
Three well-documented reversals show the size of the problem. Observational cohorts reported lower coronary heart disease incidence among postmenopausal women taking hormone therapy; the Women’s Health Initiative randomised trial of estrogen plus progestin (JAMA 2002;288:321–333) reported a hazard ratio of 1.29 for coronary heart disease. Observational studies associated high beta-carotene intake with lower lung cancer risk; the ATBC trial (New England Journal of Medicine 1994;330:1029–1035) found 18% higher lung cancer incidence in the beta-carotene arm, and CARET (New England Journal of Medicine 1996;334:1150–1155) found 28% higher incidence with beta-carotene plus retinol.
The mechanism behind the third case is the cleanest demonstration in the literature that adherence itself carries prognosis. In the Coronary Drug Project (New England Journal of Medicine 1980;303:1038–1041), five-year mortality among patients randomised to placebo was 15.1% for those who took at least 80% of their capsules and 28.2% for those who took less, and adjustment for 40 recorded baseline characteristics did not remove the gap. People who take their pills differ from people who do not, in ways no covariate list captures — which is why “adherent users versus non-users” is a comparison with a known bias, not a sensitivity analysis.
This is also the standing objection to effect estimates drawn from registry and electronic health record data, where the untreated comparator is often everyone else in the database and the indication is unrecorded by construction.
Design removes confounding before the data exist, which is the only route that reaches unmeasured variables. Four design options are available, and they differ in what they give up to do it.
Randomisation removes confounding by every variable, measured and unmeasured, because treatment assignment is made independent of all patient characteristics. This is the reason randomised trials are the reference standard, and the reason no observational method fully substitutes for one.
Randomisation does not guarantee balance in any single trial — chance imbalance occurs — but it guarantees that imbalance is random rather than systematic, which is what allows valid inference. This is why significance tests on baseline characteristics are uninformative in a randomised trial and are discouraged by the CONSORT explanation and elaboration document: any baseline difference in a properly randomised trial arose by chance, so a p-value tests a hypothesis known in advance to be true. Stratified randomisation and minimisation reduce chance imbalance on prognostic variables prospectively, and variables used to stratify allocation should then appear in the analysis model.
Restriction eliminates confounding by a variable by admitting only one level of it: studying only non-smokers removes smoking as a confounder entirely. It narrows generalisability and can introduce selection problems.
Matching pairs exposed and unexposed subjects on confounders. In cohort studies it removes confounding by the matched variables; in case-control studies it requires matched analysis, because matching on a confounder in a case-control design introduces a selection effect that must then be handled analytically.
Active-comparator, new-user design is the pharmacoepidemiological remedy for confounding by indication. Instead of comparing users with non-users, it compares patients initiating one drug against patients initiating an alternative drug for the same indication, with follow-up starting at initiation for both. The design does not measure the indication; it holds it approximately fixed by construction, and it simultaneously removes the prevalent-user and immortal time problems that the user-versus-non-user comparison creates.
Analysis can only rearrange information already in the dataset, and every method below shares that ceiling. Five approaches are in routine use, and they differ in how they model the confounders, not in which confounders they can reach.
Stratification estimates the association within levels of the confounder and combines them. Transparent and limited by the number of variables that can be stratified simultaneously.
Regression adjustment includes confounders as covariates. Flexible, and it carries an assumption people routinely forget: the functional form must be correct. Adjusting for age as a linear term when its effect is curved leaves residual confounding by age even though age is “in the model”.
Propensity scores model the probability of exposure given covariates, then match, stratify or weight on that score. The advantage is that they handle many confounders through one dimension and make covariate balance checkable — conventionally by standardised mean differences, with values below 0.10 widely treated as acceptable balance — a convention traceable to Normand and colleagues (Journal of Clinical Epidemiology 2001;54:387–398). Austin’s balance-diagnostics paper (Statistics in Medicine 2009;28:3083–3107) notes there is no consensus on the value and recommends judging a standardised difference against its sampling distribution under a correctly specified model rather than against a fixed cut-off. The limitation is identical to regression: only measured confounders enter the model.
Inverse probability weighting reweights the sample so that exposure is independent of measured confounders, and extends naturally to time-varying confounding, where a confounder is affected by prior exposure — a situation standard regression handles incorrectly. Marginal structural models fitted by inverse probability of treatment weighting were introduced for exactly this case by Robins, Hernán and Brumback (Epidemiology 2000;11:550–560), and are the standard answer when a covariate is both a confounder for later treatment and a consequence of earlier treatment.
Standardisation and the parametric g-formula take the opposite route: model the outcome given exposure and covariates, then average the predictions over the covariate distribution of a chosen target population. Standardisation returns a marginal effect rather than a conditional one, which matters when the effect measure is non-collapsible, as odds ratios and hazard ratios are. Doubly robust estimators combine an exposure model and an outcome model and remain consistent if either one is correctly specified.
Instrumental variables and related approaches can address unmeasured confounding, in exchange for assumptions that are themselves untestable and often implausible.
Weighting and matching both require positivity: every subject must have a non-zero probability of each exposure level given their covariates. Propensity scores near 0 or 1 produce extreme weights, unstable estimates and, when trimmed, an estimand defined by whoever survived the trim rather than by the research question.
Method choice determines which confounders are addressed and which failure mode you inherit, and the differences are not a matter of sophistication.
| Method | Removes confounding by | Leaves untouched | Principal failure mode |
|---|---|---|---|
| Randomisation | All variables, measured and unmeasured | Post-randomisation selection and dropout | Chance imbalance in a single trial; non-adherence |
| Restriction | The restricted variable, completely | Every other variable | Lost generalisability; restricted variable may be a collider |
| Matching (cohort) | The matched variables | Unmatched and unmeasured variables | Discarded unmatched subjects change the estimand |
| Stratification | Measured variables, few at a time | Everything not stratified | Sparse strata; cannot scale past a handful of variables |
| Regression adjustment | Measured variables, as modelled | Unmeasured variables; misspecified form | Wrong functional form leaves residual confounding |
| Propensity score | Measured variables entering the score | Unmeasured variables | Positivity violations; instruments amplify bias |
| Inverse probability weighting | Measured, including time-varying | Unmeasured variables | Extreme weights from near-positivity violations |
| Instrumental variables | Unmeasured confounding of exposure and outcome | Violations of the exclusion restriction | Weak instruments; untestable assumptions |
Every analytic approach can only address confounders you measured, measured well, and modelled correctly.
Residual confounding — from unmeasured variables, mismeasured variables, or incorrect functional form — is the standing objection to every observational estimate, and no amount of methodological sophistication removes it. A propensity score model is not more protected against unmeasured confounding than a regression; it simply presents the same limitation differently.
Measurement error in a confounder is the part most often overlooked, because the variable is present in the model and therefore looks handled. Adjusting for a mismeasured confounder removes only the confounding carried by the measured proxy and leaves the remainder; Fewell, Davey Smith and Sterne showed by simulation in the American Journal of Epidemiology (2007) that adjustment for imperfectly measured confounders can leave substantial bias even when every relevant variable is nominally in the model. Self-reported smoking status adjusted as a three-level variable does not remove confounding by pack-years.
This is why the phrase “adjusted for potential confounders” is weak in a manuscript. It describes a procedure without naming what was adjusted for, why those variables, or what remains unaddressed.
Choose an adjustment set from an assumed causal structure, written down before fitting anything. A directed acyclic graph is the standard tool: draw the variables and the arrows you believe exist, and the graph identifies which sets of variables suffice to identify the effect.
The rule the graph applies is Pearl’s back-door criterion: a set of variables suffices if it blocks every path from exposure to outcome that begins with an arrow into the exposure, and if it contains no descendant of the exposure. Greenland, Pearl and Robins set this out for epidemiologists in “Causal diagrams for epidemiologic research” (Epidemiology 1999;10:37–48), and the R package dagitty and its browser interface at dagitty.net enumerate the sufficient sets automatically once the graph is drawn. A graph frequently returns more than one sufficient set, which is useful: agreement between estimates from two disjoint adjustment sets is evidence that the assumed structure is not badly wrong.
Where the full graph is beyond reach, VanderWeele’s modified disjunctive cause criterion (European Journal of Epidemiology 2019) gives a defensible default: adjust for every pre-exposure covariate that is a cause of the exposure, of the outcome, or of both; exclude any variable known to be an instrument for the exposure; and additionally include any variable that is a good proxy for an unmeasured common cause of exposure and outcome. It is a fallback, and it is far better than the four criteria below.
What not to use:
Statistical significance in bivariate tests. Association with the outcome is not causal role, and this criterion will select colliders and mediators.
Change-in-estimate criteria. Including a variable because it moves the coefficient by more than some threshold selects exactly the variables that most distort the estimate, which includes mediators and colliders — both of which move estimates substantially. The 10% threshold in wide use traces to Maldonado and Greenland’s simulation study of confounder-selection strategies (American Journal of Epidemiology 1993;138:923–936), which found the change-in-estimate criterion performed best with a low cut-point. The authors warned against applying that finding outside their simulation framework, and Greenland has since described the 10% default as epidemiologic folklore rather than a result.
Stepwise selection. Automates the above and produces adjustment sets no causal argument supports.
Everything available. Guarantees any mediators and colliders in the dataset are included.
Adjustment adds bias, rather than removing it, in four identifiable situations, and all four are common enough to check for by name.
Overadjustment for a mediator. Conditioning on a variable on the causal pathway removes the indirect effect and returns a direct effect that answers a different question. Where the exposure is structural and the mediator is clinical — socioeconomic position and stage at diagnosis, for instance — most of the effect lives in the mediator, and the adjusted estimate approaches null. See my effect disappeared after adjusting.
Collider stratification. Conditioning on a common effect of exposure and outcome opens a path that was closed, manufacturing an association between causes that are independent in the population.
M-bias. Adjusting for a pre-exposure covariate can still create bias when that covariate is a common effect of two unmeasured variables, one causing the exposure and the other causing the outcome. Greenland described this structure in Epidemiology (2003), and it is the counterexample to the belief that any variable measured before exposure is safe to adjust for.
Bias amplification. Adjusting for a variable that strongly predicts the exposure but has no independent path to the outcome — a prescriber preference, a formulary rule, a distance to clinic — leaves any unmeasured confounding in place while shrinking the exposure variation that identifies the effect, so the existing bias grows. Myers and colleagues showed by simulation in the American Journal of Epidemiology (2011) that conditioning on a perfect or near-instrument can increase both bias and variance — while concluding that in most of their scenarios the increase was small relative to total estimation error, so controlling unmeasured confounding should usually still take priority. Instruments are routinely added to propensity score models by researchers being thorough.
A fifth failure is one of interpretation rather than estimation. Coefficients for the covariates in an adjusted model are not causal effects of those covariates, because the adjustment set was chosen to identify the exposure effect and not theirs — the Table 2 fallacy, named by Westreich and Greenland in the American Journal of Epidemiology (2013;177:292–298).
Confounding is one of three structurally distinct biases that produce a wrong association, and only confounding responds to adjustment.
Selection bias arises from who is in the study rather than from a common cause within it. Adjustment does not fix it, and it can be present in a study with no confounding at all.
Information bias arises from mismeasurement of exposure, outcome or covariates. Mismeasured confounders produce residual confounding even when the right variables were chosen.
Immortal time bias arises from how follow-up time is allocated relative to exposure definition, and is not addressed by adjusting for anything.
Recognising which one you have matters, because confounding is the only one adjustment addresses, and defaulting to “we adjusted for confounders” when the problem is selection or time allocation leaves the real bias untouched.
Two further phenomena are frequently mislabelled as confounding. Simpson’s paradox is an arithmetic consequence of pooling across unbalanced strata, and it can be produced by a confounder, a collider or a mediator — the reversal itself tells you nothing about which. Non-collapsibility is not a bias at all: an odds ratio or hazard ratio changes when covariates are added even with no confounding present, because the conditional and marginal quantities differ mathematically, so a moving hazard ratio is not by itself evidence that a confounder was found.
Confounding is detected in a manuscript by reading the covariate list against the timing of each measurement and by checking what the authors say the list is for, not by any statistic reported in the paper.
Check when each adjusted variable was measured. A covariate recorded after exposure onset is a mediator, a descendant of one, or a collider until the authors argue otherwise. Adherence, dose changes, first follow-up laboratory values and post-treatment biomarkers all belong in this category.
Check whether the selection rule is stated at all. “Variables were included if p < 0.20 in univariable analysis” and “all available covariates were included” are both statistical rules applied to a causal question, and both are quotable in a review as written.
Compare the crude and adjusted estimates. Both should be reported. A large move demands an explanation — which variable moved it, and by what causal argument — and a suspiciously small move in a setting with obvious confounding by indication suggests the confounder was not really measured.
Look for the indication. In any comparative effectiveness paper, ask what made a clinician choose the exposure and whether that information is in the dataset. If it is not, no covariate list closes the gap and the limitation belongs in the discussion in those words.
Look for a negative control. An association observed where no causal effect is possible reveals confounding directly. Lipsitch, Tchetgen Tchetgen and Cohen set out the method in Epidemiology (2010;21:383–388), and Jackson and colleagues applied it in the International Journal of Epidemiology (2006), finding influenza vaccination in seniors associated with lower all-cause mortality outside influenza season — a period in which the vaccine cannot act, so the association measured the health of the people who get vaccinated.
Check the causal language against the design. An observational estimate described with “reduces”, “protects against” or “leads to” is claiming an identified effect, and the identification argument has to be somewhere in the methods. See correlative, not causal.
Residual confounding can be bounded even when it cannot be removed, and a bound is a far stronger limitations paragraph than an acknowledgement.
The Cornfield inequality. Cornfield and colleagues showed in the Journal of the National Cancer Institute (1959;22:173–203) that for an unmeasured binary confounder to fully account for an observed relative risk, that confounder must be associated with the exposure at least as strongly as the observed relative risk itself. With a smoking–lung cancer relative risk near 9, any hypothetical constitutional factor would have to be roughly nine times more prevalent in smokers, which no candidate variable came close to.
The E-value. VanderWeele and Ding generalised the argument in the Annals of Internal Medicine (2017;167:268–274). For an observed risk ratio above 1, the E-value is RR + √(RR × (RR − 1)): it is the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need with both the exposure and the outcome to explain away the estimate. An observed risk ratio of 1.5 has an E-value of 2.37; a risk ratio of 2.0 has an E-value of 3.41. Reporting the E-value for the confidence limit nearest the null is the informative version, because it says what would be needed to move the result to non-significance.
Quantitative bias analysis. Where a plausible prevalence and effect size for the unmeasured confounder can be assumed, a bias analysis returns the estimate corrected under those assumptions, with the assumptions stated. Probabilistic versions place distributions on the bias parameters and return an interval that includes systematic as well as random error — which is what a confidence interval from the primary model conspicuously does not.
Confounding rarely arrives in a review under that name. Reviewers write the specific version, and these comments decide papers.
“The adjustment set is not justified.” “Residual confounding cannot be excluded.” “Covariates appear to have been selected by statistical criteria.” “Adjustment for [variable] may have removed the effect of interest, as it plausibly lies on the causal pathway.” “Confounding by indication is likely and is not addressed.”
The last is specific to comparative effectiveness work and worth anticipating: treatment was chosen for reasons frequently recorded nowhere in the data, and no adjustment reaches an unrecorded indication.
Four more appear often enough to prepare for. “Several adjusted covariates were measured after baseline; please justify their inclusion or restrict the model to pre-exposure variables.” “Please report unadjusted alongside adjusted estimates.” “The causal directed acyclic graph underlying the adjustment set is not presented.” “Please provide an E-value or equivalent sensitivity analysis for unmeasured confounding.” The last of these is now a routine request in epidemiology and health services research, and it is far easier to answer before submission than in revision.
Report confounding by naming the variables, the causal argument for each, and the residual that remains, in that order.
Name the confounders and why each was selected, by causal role rather than by availability. State the functional form used for continuous confounders. Report unadjusted and adjusted estimates so readers can see the movement. Address residual confounding specifically rather than generically — which variables you could not measure, and in which direction they would likely bias the estimate.
This is what the reporting guideline asks for in plain terms. STROBE item 16(a) reads: “Give unadjusted estimates and, if applicable, confounder-adjusted estimates and their precision (eg, 95% confidence interval). Make clear which confounders were adjusted for and why they were included.” The second sentence is the one manuscripts skip, and it is the one reviewers quote back.
Where possible, quantify it. A quantitative bias analysis or an E-value states how strong an unmeasured confounder would need to be to explain away the observed effect, which is far more informative than a sentence acknowledging that unmeasured confounding may exist. A limitations section that names an unmeasured confounder, states the direction it would push the estimate, and gives the strength it would require is doing analysis; one that says results should be interpreted with caution is not.
Collider bias · Mediator versus confounder · Directed acyclic graphs · Propensity score matching · Table 2 fallacy · My effect disappeared after adjusting · Epidemiology review
PerfectPaper reads every adjusted covariate against the point in the study at which it was measured, and reports which ones the exposure plausibly caused. Checked before submission by causal language discipline, which reports how the adjustment set was selected and flags covariates plausibly acting as mediators or colliders.
A variable that causes both the exposure and the outcome and does not lie on the pathway between them. Its presence makes exposure and outcome appear associated when that association reflects the confounder rather than an effect.
Grey hair and heart disease appear associated because age causes both. Adjusting for age removes the association, showing it was confounded rather than causal.
By design through randomisation, restriction, matching or an active-comparator new-user design, or by analysis through stratification, regression adjustment, propensity scores, inverse probability weighting or standardisation. Analytic methods address only measured confounders.
Confounding remaining after adjustment, because a confounder was unmeasured, measured imprecisely, or modelled with the wrong functional form. It is the standing limitation of observational research.
Not necessarily. Colliders and mediators are also associated with both, and adjusting for either introduces bias. Confounding is defined by causal direction, which no statistical test can determine.
Not with standard adjustment. Instrumental variable methods and related designs can address unmeasured confounding under assumptions that are themselves untestable, and an E-value or quantitative bias analysis can bound its likely influence.
Yes in expectation, for measured and unmeasured variables alike, because assignment is independent of patient characteristics. Chance imbalance can still occur in an individual trial, but it is random rather than systematic.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect