Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
A confounder causes the exposure; a mediator is caused by it. Adjusting for a confounder removes bias, adjusting for a mediator removes the effect you are measuring.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A confounder causes both the exposure and the outcome and lies outside the causal pathway. A mediator is caused by the exposure and in turn causes the outcome, so it lies on the pathway. Adjusting for a confounder removes bias; adjusting for a mediator removes part of the effect you are trying to measure.
The same statistical operation — adding one column to a regression — can therefore correct an estimate or destroy it, depending on a fact no model can see: the direction of the arrow between the exposure and that variable. Schisterman, Cole and Platt gave the second case its standard definition in Epidemiology (2009;20:488–495) — the term overadjustment was already in use but, as they note, defined inconsistently — defining it as control for an intermediate variable — or for a proxy descended from one — that lies on the causal path from exposure to outcome, and distinguishing it from unnecessary adjustment, which costs precision without moving the estimate. No residual plot, no information criterion and no p-value separates the two cases, because both produce a perfectly well-behaved model.
Ask one question: does the exposure cause this variable?
If yes, it is a mediator or a descendant of one — do not adjust for it unless you specifically want the direct effect, and say so if you do.
If no, and the variable causes both exposure and outcome, it is a confounder — adjust for it.
If no, and the variable is caused by both the exposure and the outcome, it is a collider, and adjusting for it manufactures an association that does not exist in the population — see collider bias.
If no, and the variable causes only the outcome, it is a competing risk factor. Adjusting for it is optional: it buys precision in a linear model and is not required for validity.
Temporal order helps but does not settle it. A variable measured before the outcome can still be a mediator if the exposure preceded it. The failure mode is measurement date standing in for causal order: pack-years recorded at enrolment, body mass index recorded at baseline and medication use recorded at the index visit all sit downstream of exposures that began years earlier. The reliable version of the test asks when the process that set the variable’s value occurred relative to the exposure, which is a question you answer on a directed acyclic graph drawn before the analysis, not in the data dictionary.
| Variable role | Does the exposure cause it? | Adjust for it? | Consequence of the wrong choice |
|---|---|---|---|
| Confounder | No | Yes | Omitting it leaves bias of unknown sign and unknown size |
| Mediator | Yes | Only for a declared direct effect | Adjusting subtracts the indirect effect and understates the total effect |
| Collider | Yes, and the outcome causes it too | No | Adjusting creates an association where none exists |
| Mediator–outcome confounder affected by the exposure | Yes | Neither adjusting nor omitting works | Adjusting blocks part of the effect; omitting leaves confounding |
| Competing risk factor | No | Optional | Omitting costs precision, not validity |
| Proxy or descendant of a mediator | Yes, indirectly | No | Partial attenuation that looks like a modest confounding correction |
The last row is the one that escapes review most often. A variable does not have to be the mediator to behave like one: any measured descendant of a mediator carries part of the exposure’s signal, so adjusting for hospital length of stay, discharge destination or number of follow-up visits attenuates a treatment effect even though none of them is the mechanism anyone would name.
Adjusting for a mediator replaces the total effect with an estimate close to a direct effect, so the coefficient shrinks by roughly the amount that travelled through the mediator. Robins and Greenland set out the identification conditions for direct and indirect effects in Epidemiology (1992;3:143–155), and the decomposition they formalised is the reason the two numbers are not interchangeable: on the difference scale the total effect equals the natural direct effect plus the natural indirect effect, and on the ratio scale it is their product.
Take a constructed illustration with no data behind it. An exposure carries a total risk ratio of 1.60 for the outcome. Of that, the path running through the mediator contributes a natural indirect risk ratio of 1.30, and the remainder, 1.23, is the natural direct effect — about 62% of the excess risk is mediated. Put the mediator in the model and the printed estimate is 1.23 with an ordinary confidence interval and ordinary diagnostics. A reader shown only that number sees a weak association and no trace of the 1.60 that existed before the column was added.
Two further facts make the shrinkage worse than it looks. First, the attenuation is routinely reported as evidence that the original association was confounded, which is the single most common misreading of this situation — see my effect disappeared after adjusting. Second, adding the mediator to a regression does not generally return the natural direct effect at all. What it estimates is closer to a controlled direct effect, and the two coincide only when there is no exposure–mediator interaction and no unmeasured confounding of the mediator–outcome relationship. Cole and Hernán made the second half of that point in the International Journal of Epidemiology (2002;31:163–165) — stratifying on the intermediate is unbiased only if there is no unmeasured confounding of the mediator–outcome relationship as well as of the exposure–outcome relationship — in a paper titled “Fallibility in estimating direct effects” for a reason.
The clearest example: adjusting the association between socioeconomic position and cancer survival for stage at diagnosis. Late-stage presentation is part of how disadvantage produces worse survival, not a nuisance variable standing beside it. Adjusting for stage therefore removes some of the very effect the paper set out to measure and leaves an estimate answering a question nobody asked.
The same pattern appears whenever an access, treatment or behaviour variable sits between a structural exposure and a health outcome.
The recurring shape has three parts: a structural exposure (income, insurance status, rurality, educational attainment, exposure to racism), a service or behaviour variable in the middle (time to presentation, stage, receipt of guideline-concordant treatment, adherence, hospital volume, referral to specialist care), and a clinical outcome. Every one of the middle variables is a mechanism the paper exists to describe. A model that adjusts for treatment received answers “among patients who received the same treatment, does disadvantage still predict death?” — a legitimate question, and almost never the one the abstract claims to have answered. Reporting a mediator-adjusted disparity as the disparity understates it, which is the substance of health equity reporting.
Administrative data makes this near-automatic. Registries and electronic health records capture what happens after care begins far more completely than what preceded it, so the covariates most readily to hand in registry analyses are precisely the ones most likely to sit on the pathway. The convention of listing every available variable as a “potential confounder” then converts data availability into an estimand.
A mediator is frequently a collider as well, so adjusting for it does not merely subtract the indirect effect — it opens a new biasing path through variables you never measured. Hernández-Díaz, Schisterman and Hernán set out the canonical case in the American Journal of Epidemiology (2006;164:1115–1120): among low-birth-weight infants, those born to mothers who smoked have lower infant mortality than those born to mothers who did not, which read naively says maternal smoking protects the smallest babies.
The structure explains it. Birth weight is caused by maternal smoking, so it is a mediator. Birth weight is also caused by other, largely unmeasured causes of infant mortality, such as congenital anomalies. Restricting or adjusting to low birth weight conditions on a common effect of smoking and of those unmeasured causes, so within the low-birth-weight stratum the non-smokers’ small babies are enriched for the more lethal explanations. The protective association is an artefact of the adjustment, not a finding about tobacco.
The general lesson is about assumption strength. The total effect requires no unmeasured confounding of the exposure–outcome relationship. A direct effect additionally requires no unmeasured confounding of the mediator–outcome relationship — a strictly stronger condition, and one randomisation does nothing to deliver, because randomising the exposure does not randomise the mediator. A trial that adjusts for a post-randomisation variable forfeits the balance randomisation bought.
Adjusting for a mediator is legitimate when the direct effect is the estimand you actually want, you name it as such in the abstract, and you defend the assumptions a direct effect requires. Three situations recur in submitted work.
Mechanism questions. Whether an exposure acts other than through a named pathway is a direct-effect question by construction, and a null direct effect alongside a non-null total effect is a substantive result worth reporting.
Surrogate and biomarker questions. Whether a drug’s effect on mortality is fully explained by its effect on a laboratory value is the question behind every surrogate endpoint validation, and it is answered by decomposing the effect, not by asserting it.
Policy questions about a fixed intervention. “What would the disparity be if everyone received guideline-concordant treatment?” is a controlled direct effect with a policy interpretation, and it is defensible so long as the paper says that is the quantity, sets the mediator to a stated value, and does not present the result as the observed disparity.
Do a mediation analysis and name it as one, reporting direct and indirect effects separately. What is not acceptable is presenting a mediator-adjusted coefficient as the total effect.
The method to describe is no longer the causal-steps procedure of Baron and Kenny (Journal of Personality and Social Psychology 1986;51:1173–1182), which remains in wide use in psychology and health-services research despite having no formal identification conditions and no way to accommodate an exposure–mediator interaction. VanderWeele’s four-way decomposition (Epidemiology 2014;25:749–761) splits a total effect into a controlled direct effect, a reference interaction, a mediated interaction and a pure indirect effect, and makes the interaction visible rather than assuming it away.
Identification rests on four no-unmeasured-confounding conditions, and a Methods section that names them is more persuasive than one that names software: no unmeasured exposure–outcome confounding, no unmeasured mediator–outcome confounding, no unmeasured exposure–mediator confounding, and no mediator–outcome confounder that is itself affected by the exposure. Implementations exist in the R packages mediation (Tingley and colleagues, Journal of Statistical Software 2014;59(5)) and CMAverse (Shi, Choirat, Coull, VanderWeele and Valeri, Epidemiology 2021;32:e20–e22).
Report the sensitivity analysis with the estimate. Mediational E-values (Smith and VanderWeele, Epidemiology 2019;30:835–837) give the minimum strength of unmeasured mediator–outcome confounding required to explain away a direct or indirect effect, and one number does more for a limitations paragraph than a paragraph of acknowledgement. Give the proportion mediated only when the total effect’s confidence interval excludes the null; it is a ratio with the total effect in its denominator, so it becomes unstable and can exceed 1 or turn negative as that denominator approaches zero.
A variable that confounds the mediator–outcome relationship and is itself affected by the exposure cannot be handled by adding or omitting a column, because both choices are wrong. Adjust for it and you block part of the effect you are estimating; leave it out and you leave confounding in place. This is the fourth identification condition above, and it is the one violated most quietly.
The standard instance is a repeated exposure with a time-varying clinical marker: CD4 count both predicts subsequent antiretroviral treatment and responds to prior treatment, so it is simultaneously a confounder of later treatment and a consequence of earlier treatment. Any longitudinal design in which the exposure is repeated and the intermediate is remeasured is a candidate. The remedy is a g-method — a marginal structural model fitted with inverse probability of treatment weighting, or g-estimation of a structural nested model — rather than a better covariate list. VanderWeele and Tchetgen Tchetgen extended mediation itself to this setting in the Journal of the Royal Statistical Society Series B (2017;79:917–938).
This is a different defect from immortal time bias, which concerns how follow-up time is allocated relative to the exposure definition rather than which columns sit in the model, and the two frequently appear in the same cohort paper.
Mediator adjustment is detected by reading the covariate list against the study timeline. No diagnostic identifies it, and no goodness-of-fit statistic prefers the correct model.
Date every covariate. For each variable in the adjustment set, establish when the process that produced its value occurred relative to the exposure. STROBE item 16(a) asks authors to give unadjusted and confounder-adjusted estimates and to “make clear which confounders were adjusted for and why they were included”; most papers supply the list and skip the justification, and the STROBE checklist is where a reviewer will point.
Read the crude-to-adjusted change. A large one-step collapse toward the null — say a hazard ratio near 1.9 falling to roughly 1.1 as a single post-exposure covariate enters the model — is the signature. Confounding control usually nudges an estimate; a mediator can take most of it away at once.
Watch for “independent of”. “The association was independent of stage” is a direct-effect claim written in total-effect prose, and it is one of the phrasings that causal language review exists to catch.
Watch for the extra sensitivity model. “We additionally adjusted for treatment received” is presented as reassurance, and the additional variable is frequently the mediator, so the sensitivity analysis is answering a different question rather than confirming the first answer.
Check whether a stratification is a downstream stratification. Results stratified by treatment received, by adherence category or by response status are mediator adjustment under another name.
Check the other coefficients in the table. A model whose adjustment set is correct for the exposure is not automatically correct for any other variable in it, which is the Table 2 fallacy — and it is how a mediator gets reported with its own effect estimate and its own interpretation.
Reviewers rarely write “you adjusted for a mediator”. They write the specific version, and these are the comments that send a paper back.
“Stage at diagnosis is plausibly on the causal pathway from socioeconomic position to survival; adjusting for it changes the estimand from a total to a direct effect, and the abstract reports a total effect.” “Please justify each covariate with respect to its timing relative to the exposure, or provide a directed acyclic graph.” “The attenuation after adjustment is interpreted as confounding; an equally consistent explanation is that the added variable is a mediator, and the manuscript does not distinguish them.” “If a direct effect is intended, state the assumption of no unmeasured mediator–outcome confounding and provide a sensitivity analysis for it.” “Adjustment for a post-randomisation variable forfeits the balance randomisation provided.” “The proportion mediated is reported without a confidence interval and the total effect includes the null.”
The comment that costs the most time is the last kind of request rather than the first: asking the authors to restate the estimand forces the title, abstract, discussion and every table caption to be rewritten, where refitting one model would have taken an afternoon.
Leaving the variable out is unavailable more often than methods papers assume, and there are five specific workarounds.
A reviewer or editor demands the adjustment. Report both models side by side with named estimands, and say in the response to reviewers which question each answers. Adding the extra model is a small change; conceding that it replaces the original one is not.
Only post-exposure variables were recorded. In registry and claims data the pre-exposure covariates frequently do not exist. Estimate what the data supports, call it a direct effect, and state which confounders of the exposure–outcome relationship remain unmeasured rather than implying the adjustment set was chosen.
The variable is a mediator for one exposure and a confounder for another. No single adjustment set serves every coefficient in one table. Fit one model per exposure with its own adjustment set and report the exposure of interest only, which also removes the Table 2 fallacy from the same manuscript.
The mediator is measured with error. Measurement error in a mediator attenuates the estimated indirect effect and leaves part of it inside the “direct” estimate, so a mediator-adjusted coefficient is not a clean direct effect even when every identification condition holds. State the direction of that residual rather than treating the decomposition as exact.
The mediator is entangled with the exposure definition. Where exposure is defined by having received a treatment, restricting or adjusting on receipt is not an analytic choice but a design feature. Use a pre-specified landmark or a time-varying exposure definition, and report how many participants the landmark excludes.
Report the estimand in the same sentence as the number, every time it appears. “The total effect of area deprivation on 5-year mortality was HR 1.42” is reportable; “the adjusted hazard ratio was 1.42” is not, because a reader cannot recover which question it answers.
Give three labelled rows where both estimands are of interest: the unadjusted estimate, the confounder-adjusted total effect, and the mediator-adjusted direct effect, with the mediator named in the row label. Publish the directed acyclic graph that produced the adjustment set and reference it in the Methods, so a reviewer can dispute the graph rather than guess at it. Where a formal mediation analysis is reported, give the natural direct and natural indirect effects with confidence intervals, name the software and version, state the four identification conditions, and give a mediational E-value.
Then write the limitations section about the specific assumption at risk. “Residual confounding is possible” applies to every observational paper ever published and therefore distinguishes none of them. “The direct effect assumes no unmeasured confounding of the stage–survival relationship; comorbidity burden is unmeasured here and would bias the direct effect away from the null” is a statement a reader can weigh.
Confounding · Collider bias · My effect disappeared after adjusting · Health equity reporting · Selection bias · Epidemiology review
PerfectPaper reads the adjustment set against the study timeline, flags each covariate the exposure plausibly causes, and reports where the estimand the model returns stops matching the effect the abstract claims. Checked before submission alongside causal language discipline, which catches “independent of” and “after accounting for” attached to a variable on the pathway.
A confounder causes the exposure; a mediator is caused by it. Adjusting for a confounder removes bias, while adjusting for a mediator removes part of the causal effect being estimated.
Only if you want the direct effect rather than the total effect, and only if you say so explicitly. Presenting a mediator-adjusted estimate as the total effect misstates what was measured.
Ask whether the exposure causes it. If the exposure precedes and produces the variable, it is on the pathway and adjusting for it will attenuate the effect. Measurement date is not the test: a variable recorded at baseline is still a mediator when the exposure began years before enrolment.
The effect estimate shrinks toward null, often substantially, and is frequently misread as evidence that the original association was confounded. Schisterman, Cole and Platt call this overadjustment bias (Epidemiology 2009;20:488–495), and no model diagnostic reveals it.
Yes, and it is the hardest case in the table. A mediator–outcome confounder that is itself affected by the exposure cannot be handled by adjusting or by omitting, because adjusting blocks part of the effect and omitting leaves confounding; the remedy is a g-method such as a marginal structural model.
No, but one variable is often both. A mediator is caused by the exposure and causes the outcome; a collider is caused by two variables that point into it. Birth weight is a mediator of maternal smoking and a collider with unmeasured causes of infant mortality, which is why adjusting for it produces the birth weight paradox.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect