Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Collider bias is created by adjustment rather than removed by it. Conditioning on a variable caused by both exposure and outcome induces an association that is not real.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Collider bias is a distortion introduced by adjusting for a variable that is caused by both the exposure and the outcome. Unlike confounding, which exists in the data before you analyse it, collider bias is created by the analysis: conditioning on a collider induces an association between exposure and outcome that does not otherwise exist, and can reverse the sign of a real effect.
The word “collider” describes the shape of the causal diagram. Two arrows point into the variable — they collide there — rather than out of it. That single structural difference reverses what adjustment does. The definition also covers a variable caused by a cause of the exposure and a cause of the outcome, which is why a covariate measured before the exposure is not automatically safe.
Collider bias needs only two unrelated causes and one selection rule to appear, and the standard illustration is university admission.
Talent and effort both contribute to university admission, and suppose the two are unrelated in the general population — knowing an applicant is talented tells you nothing about how hard they work.
Now look only at admitted students. Among them, talent and effort will appear negatively correlated: a student admitted with low talent must have had high effort, and vice versa, or they would not have been admitted. The correlation is real in the admitted sample and absent in the population. Nothing produced it except the act of looking only at admitted students.
The arithmetic is worth doing once, because the size of the artefact surprises people. Take 10,000 applicants, half high in talent, half high in effort, the two independent, so that 2,500 applicants sit in each of the four cells. Admit anyone high in either, and 7,500 are admitted. Among admitted applicants with high talent, 2,500 of 5,000 have high effort — 50%. Among admitted applicants with low talent, 2,500 of 2,500 have high effort — 100%. A 50-percentage-point gap has appeared in a population where the true gap is zero, and it was manufactured entirely by the admission rule.
That is collider bias in its entirety. Conditioning on a common effect of two variables creates an association between them.
Collider bias has a graphical definition that makes it decidable rather than a matter of judgement. On a path such as A → C ← Y, the node C where two arrowheads meet is a collider for that path. Under Pearl’s d-separation rules the path is blocked by default, and conditioning on C — or on any descendant of C — opens it, creating a non-causal association between A and Y. Cole, Platt, Schisterman, Chu, Westreich, Richardson and Poole set the mechanism out with worked numbers in “Illustrating bias due to conditioning on a collider” (International Journal of Epidemiology 2010;39:417–420).
Two consequences follow that manuscripts routinely miss. First, collider status is path-specific: the same variable can be a confounder on one path and a collider on another, so “is X a collider?” is not answerable without the graph. Second, the descendant clause means a downstream proxy behaves like the collider itself — adjusting for a correlate of hospital admission rather than admission is not a defence.
Elwert and Winship gave the phenomenon its broadest name in “Endogenous selection bias: the problem of conditioning on a collider variable” (Annual Review of Sociology 2014;40:31–53), which treats regression adjustment, sample selection and missing-data exclusion as three surfaces of one operation. Hernán, Hernández-Díaz and Robins had already shown the same structure underlying study selection, writing the selection indicator as a node S and noting that restricting the analysis to S = 1 is a conditioning operation (Epidemiology 2004;15:615–625). See directed acyclic graphs for how the rules are applied to a specific adjustment set.
The birth weight paradox is the standard published demonstration that adjustment can invert a real effect. Among US infants born in 1991, infants of smokers had higher risks of both low birth weight and infant mortality than infants of nonsmokers — yet among low-birth-weight infants, mortality was lower for infants of smokers, at a relative rate of 0.79.
Hernández-Díaz, Schisterman and Hernán reconstructed the structure in “The birth weight ‘paradox’ uncovered?” (American Journal of Epidemiology 2006;164:1115–1120). Birth weight is affected by smoking and shares common causes with mortality, such as birth defects. Stratifying on low birth weight therefore compares infants who are small because their mothers smoked against infants who are small for reasons carrying much higher mortality, and the comparison flatters smoking. Their conclusion is stated as a rule: “adjustment for birth weight is unwarranted when the analytical goal is to estimate overall effects of prenatal variables on infant mortality.”
Whitcomb, Schisterman, Perkins and Platt then quantified it by simulation in Paediatric and Perinatal Epidemiology (2009;23:394–402), showing that when the birth weight–mortality relation carries substantial uncontrolled confounding, the bias in a birth-weight-adjusted estimate is “sufficient to yield opposite causal conclusions, i.e. a factor that poses increased risk appears protective.” Nothing in the fitted model announces this. The stratified estimate has an ordinary confidence interval, ordinary residuals and ordinary fit statistics.
Confounding is widely understood, routinely discussed, and conventionally acknowledged somewhere in an observational paper’s limitations. Collider bias has three properties that make it more damaging in practice.
It is created by the researcher. Confounding is a property of the world you are trying to overcome. Collider bias is introduced by an analytic choice, usually one made in good faith and often one that looks like careful adjustment.
It can reverse direction. Confounding typically inflates or attenuates. Collider bias can manufacture an association where the true effect is zero, or flip a protective effect into an apparently harmful one.
Adjusting for more variables makes it worse. The instinct when a reviewer questions confounding is to adjust for more. If any of those additions is a collider, the estimate degrades while appearing more rigorous.
A fourth property is institutional rather than statistical: confounding has a defensive vocabulary that authors and reviewers share, and collider bias does not. “We adjusted for potential confounders” is accepted as a response to a confounding objection; there is no equivalent sentence that answers a collider objection, because the answer has to be a graph.
Collider bias is not uniformly large, and the magnitude depends on which of two structures produced it. Distinguishing them is what turns a reviewer’s objection into an arguable claim.
Greenland compared the two in “Quantifying biases in causal models: classical confounding vs collider-stratification bias” (Epidemiology 2003;14:300–306), and found that bias from stratifying on a variable affected by both exposure and disease “may often be comparable in size with bias from classical confounding”, whereas other biases from collider stratification “may tend to be much smaller”.
The smaller class is M-bias, in which the collider precedes the exposure and is linked to it and to the outcome only through unmeasured causes. Liu, Brookhart, Schneeweiss, Mi and Setoguchi simulated 178 scenarios in American Journal of Epidemiology (2012;176:938–948) and found M-bias of −2% to −5% in realistic settings, exceeding 15% only when the relative risks linking the collider to the unmeasured factors reached 8 or more. Their operational conclusion matters as much as the number: when a collider is also an important confounder, controlling the confounding takes precedence over avoiding M-bias.
Selection colliders are the class that bites. Munafò, Tilling, Taylor, Evans and Davey Smith showed in “Collider scope” (International Journal of Epidemiology 2018;47:226–235) that even modest influences on selection into, or attrition from, a study generate biased estimates of both phenotypic and genotypic associations. Griffith and colleagues then documented it in a live dataset, reporting in Nature Communications (2020;11:5749) that UK Biobank participants tested for COVID-19 were highly selected relative to the wider cohort on genetic, behavioural, cardiovascular, demographic and anthropometric traits.
Collider bias takes six recognisable forms in submitted work, and only the first of them looks like adjustment.
Adjustment for a common effect. The textbook case: a covariate caused by both exposure and outcome enters the model. Schisterman, Cole and Platt call the related error of adjusting for an intermediate variable, or a descending proxy for one, overadjustment bias, and separate it from merely unnecessary adjustment, which harms precision but not validity (Epidemiology 2009;20:488–495). The distinction is useful in a response letter: not every superfluous covariate is a bias.
Selection into the sample. The most consequential form, and the one that does not look like adjustment at all. If being in your study depends on both the exposure and the outcome, you have conditioned on a collider before fitting anything. Hospital-based studies are the classic example: admission depends on many exposures and on disease severity, so associations among admitted patients can differ systematically from those in the population. Berkson demonstrated exactly this in Biometrics Bulletin (1946;2:47–53), which is why the whole family is a form of selection bias.
Survival to measurement. If subjects must survive to be measured, and both exposure and outcome affect survival, the measured sample is conditioned on a collider. This is why studies of risk factors among survivors of a disease can produce paradoxical results — the obesity paradox and similar findings have been attributed to exactly this structure.
Index event selection. Cohorts that begin at a disease event — first myocardial infarction, first hospitalisation, entry onto a transplant list — condition on having reached that event, which both the exposure and the outcome influence. Secondary analyses of routine data are especially exposed, because the cohort is defined by having generated encounters; see registry data limitations.
Loss to follow-up. If dropout depends on both exposure and outcome, complete-case analysis conditions on remaining in the study, which is a collider.
Stratification and restriction. Both are conditioning. Presenting results within strata of a collider carries the same bias as adjusting for it, and a subgroup table is not exempt because it reports no coefficient.
Collider bias is detected by inspecting the causal role and the measurement timing of every covariate, and by reading the eligibility criteria as conditioning statements. No statistical test finds it.
Draw the causal structure before choosing the adjustment set, and for each covariate ask a single question: could the exposure or the outcome cause this variable? If either could, it is not a confounder, and adjusting for it is a decision requiring separate justification.
Timing is the most useful practical heuristic. A variable measured after the exposure is a candidate collider or mediator, not a confounder. Post-baseline variables in a cohort study, anything measured during follow-up, and anything measured at the same time as the outcome all warrant scrutiny. The heuristic is one-directional: it catches post-exposure colliders and misses M-bias entirely, since an M-bias collider is measured at baseline like any confounder.
Three further checks are worth running on your own manuscript before a reviewer runs them on you. Write the graph in dagitty syntax and let the tool at dagitty.net enumerate the minimal sufficient adjustment sets, then check that your model’s covariate list is one of them (Textor, van der Zander, Gilthorpe, Liśkiewicz and Ellison, International Journal of Epidemiology 2016;45:1887–1894). Check descendants as well as colliders, since a proxy opens the same path. And read your eligibility criteria as if they were covariates: a criterion such as “patients who completed at least three cycles” is a conditioning statement written in prose.
For selection, ask whether participation, enrolment, survival to measurement, or response depends on both the exposure and the outcome. If it does, the bias is present regardless of which covariates you chose.
Stepwise selection, change-in-estimate criteria, and “adjust for everything available” all choose covariates by statistical association with the outcome.
Colliders are, by construction, associated with the outcome. So are mediators. These procedures therefore actively select the variables that introduce bias, and they do so more reliably the larger the dataset, because association is easier to detect. A change-in-estimate rule is the worst of the three for this purpose: it retains whichever variables move the coefficient most, and a collider moves it a great deal — see my effect disappeared when I added a covariate.
Machine-learning covariate selection has the same problem in a less visible form: predictive performance is unrelated to causal role, and a model that predicts the outcome well may do so partly through colliders. The same applies to a propensity score built from every available baseline field; balance diagnostics on that score will look excellent, because balance is a statistical property and collider status is not. Propensity score matching inherits whatever the covariate list contains.
The reporting habit that follows from all of this is small and concrete: name the rule that produced the adjustment set. “Covariates were selected from a directed acyclic graph specified before analysis” and “covariates with p < 0.20 in bivariate analysis were retained” are different studies, and only one of them can answer a collider objection.
Collider bias has no diagnostic test, and that is the central difficulty: the biased estimate is a perfectly ordinary number, and model fit statistics do not distinguish it from an unbiased one. No residual plot, information criterion, or goodness-of-fit statistic is sensitive to the causal role of a covariate.
The circumstantial signals worth attending to are an effect that appears or reverses on adjustment, a result that contradicts randomised evidence, a paradoxical protective effect in a subgroup selected on illness, and an association present only among people selected into a restricted sample.
Two further signals are specific to the manuscript rather than the estimate. A covariate table in which every coefficient is interpreted causally suggests the adjustment set was never tied to a single estimand — the Table 2 fallacy. And an analysed denominator smaller than the enrolled denominator, with no accounting for the difference, means the sample was conditioned on completeness.
Each of these has innocent explanations too. The point is not that they prove collider bias but that they should prompt you to re-examine the causal structure rather than to trust the adjusted number because it is adjusted.
Conditioning on a collider is sometimes correct, and stating when sharpens rather than weakens the rest of the argument. Three cases recur.
Mediation analysis. Estimating a controlled direct effect requires conditioning on the mediator, and the mediator is usually a collider with respect to unmeasured common causes of mediator and outcome. The conditioning is intended; what is required is the identifying assumption stated explicitly — no unmeasured mediator–outcome confounding — rather than silence. The distinction between the two roles is set out in mediator versus confounder.
A variable that is both collider and confounder. Liu and colleagues’ 2012 simulation is the reference point: with M-bias running −2% to −5% in the scenarios they judged realistic, and confounding frequently much larger, controlling the confounding is the better trade. Say that you made the trade and why, and give the estimate both ways.
Prediction rather than estimation. A collider is a legitimate and often excellent predictor. A risk model that includes one is not wrong; a sentence that describes its coefficients as effects is. Keep the estimand attached to the model.
The boundary case that is not an exception: selecting on the exposure alone, or on the outcome alone, is ordinary design and does not condition on a collider. A cohort restricted to one occupational group and a case-control study assembled on outcome status are both legitimate; what bites is selection driven by both sides at once.
Collider bias is prevented at design, by fixing the adjustment set from causal structure before any model is fitted, and mitigated afterwards only where the selection mechanism was measured.
Choose the adjustment set from a causal diagram, drawn before analysis, and report it. The graph makes the assumptions inspectable, which is the whole value — readers can disagree with an arrow, which they cannot do with an unstated intuition. Publish the dagitty model code in the supplement so a reviewer can recompute the sufficient sets rather than argue from a picture.
Prefer baseline variables. Confounders precede the exposure; a set restricted to genuinely pre-exposure variables cannot contain colliders of the exposure-outcome relationship. Treat this as a heuristic, not a theorem, since M-bias colliders are pre-exposure by construction.
Handle selection explicitly. Inverse probability of selection weighting can address selection collider bias when the selection mechanism is measured. Where it is not, say so as a limitation and describe the likely direction.
Report sensitivity. Show the estimate under different adjustment sets, including the minimal sufficient set from the graph, so readers can see how much the conclusion depends on contested choices. Where the exposure definition interacts with follow-up time, check separately for immortal time bias, which arises from how person-time is classified rather than from the adjustment set, so a sufficient-set analysis will not surface it.
Collider bias frequently cannot be removed, because the selection mechanism was never measured, and the honest response is to bound it rather than to acknowledge it. A limitations paragraph that names the bias and stops is the version reviewers reject.
Bound the bias. Smith and VanderWeele give selection bounds in Epidemiology (2019;30:509–516) that state the minimum strength of selection required to explain away an observed risk ratio, with implementations in the R package EValue. A sentence of the form “selection would have to be associated with both exposure and outcome at a risk ratio of at least 2.4 to move this estimate to the null” is an analysis; “results should be interpreted with caution” is not. Guidance on writing that paragraph is in how to write a limitations section.
Reason about direction, and commit to it. The admissions arithmetic above shows the direction is often deducible from the structure alone: conditioning on a common effect of two positively acting causes induces a negative association between them. State which way the bias would push your estimate and why.
Change the estimand rather than the model. An estimate conditional on being selected is a valid estimate for the selected population; it is only wrong when reported as a population estimate. Restating the estimand is sometimes the whole correction.
Report the analysis you cannot defend anyway. Where the graph admits no sufficient adjustment set, that is a design finding. It points towards a negative-control outcome, an instrumental variable, or a different data source — not towards a longer covariate list.
Reviewers usually name the consequence before they name the structure, and the phrasings below are the ones that recur on observational manuscripts.
“The adjustment set is not justified.” “Covariate X is measured after exposure and may be affected by it.” “The study population is selected on factors related to both exposure and outcome.” “The observed association may reflect selection rather than causation.” “Adjustment appears to have been performed using statistical criteria.”
The versions that decide papers are more specific still. “The reported effect reverses on adjustment for a post-baseline variable; please justify the causal role of that variable.” “The covariates in the adjusted model do not correspond to any sufficient adjustment set implied by the DAG in Figure 1.” “Collider bias is acknowledged in the limitations but its direction and plausible magnitude are not stated.” “Conditioning on the stratifying variable may induce the reported subgroup difference.” The last of these is the one authors least expect, because a subgroup table does not feel like an adjustment.
If your effect appeared or reversed after adjustment, expect this specifically, and expect it to be the comment that decides the paper. The response that works is a graph plus a bound, not a longer covariate list — which is also the substance of the reviewer objection that a finding is correlative rather than causal.
Report collider bias by naming the structure, the direction and the magnitude, in that order.
State the estimand in one sentence before the model. Name the rule that produced the adjustment set, and publish the graph in reproducible form. Say which variables you considered and excluded as colliders or mediators, and why — an excluded-variable sentence is as informative as an included one. Where selection may act, describe the selection mechanism, report the analysed denominator against the enrolled denominator, and give a bound or a weighted analysis rather than an acknowledgement. Where the reported estimate is conditional on selection, say which population it describes. Where results differ across groups defined by a selection-related variable, report them separately rather than aggregating, which is also the substance of health equity reporting.
Confounding · Mediator versus confounder · Selection bias · Directed acyclic graphs · Table 2 fallacy · My effect disappeared when I added a covariate · Epidemiology review
PerfectPaper reads the adjustment set against the stated estimand and the measurement timing of each covariate, and reports where a variable in the model is plausibly caused by both the exposure and the outcome, naming the likely direction of the resulting bias. It reads eligibility criteria the same way, as conditioning statements written in prose. The companion check on causal language discipline then tests whether the wording of the conclusion matches what that adjustment set can support.
A confounder causes both the exposure and the outcome, and adjusting for it removes bias. A collider is caused by both, and adjusting for it creates bias. The arrows point in opposite directions, and so does the consequence of conditioning. Collider status is also path-specific: one variable can be a confounder on one path and a collider on another.
Yes. Conditioning on a collider can turn a positive association negative or manufacture one where the true effect is zero, which is why it produces confident, coherent, entirely wrong conclusions. The published demonstration is the birth weight paradox: among US infants born in 1991, mortality among low-birth-weight infants was lower for infants of smokers, at a relative rate of 0.79.
Frequently. When inclusion in the sample depends on both the exposure and the outcome, the study has conditioned on a collider by construction, before any covariate is chosen. Hernán, Hernández-Díaz and Robins formalised this by writing selection as a node S and treating restriction to S = 1 as a conditioning operation (Epidemiology 2004;15:615–625).
Choose the adjustment set from an assumed causal structure rather than from model fit or data availability, prefer variables measured before the exposure, and consider whether sample selection itself depends on both exposure and outcome. Enumerating the minimal sufficient adjustment sets with dagitty and checking that your model matches one of them takes minutes and pre-empts the most common reviewer objection.
Not reliably. Adding covariates reduces confounding only if those covariates are confounders. Adding colliders or mediators introduces bias, so a more heavily adjusted model can be further from the truth than a simpler one. Greenland found that stratifying on a variable affected by both exposure and disease produces bias often comparable in size to classical confounding (Epidemiology 2003;14:300–306). Schisterman, Cole and Platt separate overadjustment bias, which is a validity problem, from unnecessary adjustment, which only harms precision (Epidemiology 2009;20:488–495).
No. Colliders are identified by causal structure, not by any statistical property, and a collider-biased estimate looks like an ordinary estimate with ordinary fit statistics. What you can test is how much selection would have to act to explain the result away, using the selection bounds of Smith and VanderWeele (Epidemiology 2019;30:509–516).
Because the exposure may have caused it. Anything the exposure causes is either a mediator or a descendant of one, and if the outcome also influences it, it is a collider. Neither belongs in a confounding adjustment set. The heuristic is one-directional: it will not catch M-bias, where the collider precedes the exposure. That form is usually the smaller worry — Liu and colleagues simulated 178 scenarios and found M-bias of −2% to −5% in the realistic ones, rising above 15% only when the collider’s links to the unmeasured factors reached a relative risk of 8 (American Journal of Epidemiology 2012;176:938–948).
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect