Skip to content

SOLUTIONS

What is regression to the mean?

Regression to the mean is the tendency of an extreme measurement to be followed by a less extreme one, because selection captures transient error alongside real signal.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is regression to the mean?

Regression to the mean is the tendency for an extreme measurement to be followed by a less extreme one, because the first value combined a stable signal with transient error unlikely to repeat. Francis Galton named the effect in 1886, reporting that a child’s height deviated from the population average by only about two-thirds of the mid-parent deviation.

Regression to the mean is not a biological or clinical process. It is an arithmetic consequence of selecting on a measurement that contains noise, and it produces apparent improvement in any group chosen because its baseline values were unusual. This page covers the mechanism, how to calculate the expected size of the effect, where it enters clinical and laboratory research, why randomisation removes it from some comparisons and not others, how it is detected in a manuscript, and what reviewers write when it has not been addressed.

The mechanism: selection on a measurement that contains error

Regression to the mean arises because an observed value is the sum of a person’s stable true value and a transient component — biological variation on the day, instrument noise, observer effect, digit preference. Selecting the highest observed values selects people whose true values are high and people whose transient component happened to be large. On remeasurement the true component persists and the transient component is redrawn, so the group’s mean moves towards the population mean.

The size of the movement is fixed by one quantity: the correlation between the two measurements. For a bivariate normal pair with common mean μ and correlation ρ, the conditional expectation is E[Y | X = x] = μ + ρ(x − μ). An individual selected at x therefore regresses by (1 − ρ)(x − μ) in expectation. Written in standard units, a person 2 SD above the mean on a measure with test-retest correlation 0.80 is expected 1.6 SD above on the retest.

Two implications follow directly. First, ρ here is the reliability of the measurement, so regression to the mean is a direct function of measurement error: halving the error variance halves the regression. Second, when ρ = 1 there is no regression at all, and when ρ = 0 the selected group’s expected follow-up value is the population mean — the entire apparent gain is spurious.

How much regression to expect

The expected regression for a group selected above a threshold can be computed before any data are collected, which makes it a quantity to report rather than a hazard to acknowledge. For a normal measurement with mean μ and standard deviation σ, selecting everyone above a cut-off c gives a selected mean of μ + σλ, where λ is the inverse Mills ratio φ(z)/(1 − Φ(z)) evaluated at z = (c − μ)/σ. The expected fall at remeasurement is (1 − ρ)σλ.

A worked example. Suppose systolic blood pressure in the source population has mean 130 mmHg and standard deviation 20 mmHg, and a trial enrols everyone whose single screening reading is at least 160 mmHg. Then z = 1.5, λ = φ(1.5)/(1 − Φ(1.5)) = 0.130/0.0668 = 1.94, and the enrolled group averages 130 + 20 × 1.94 ≈ 169 mmHg. If the correlation between single casual readings taken at separate visits is 0.70, the expected mean at the second visit is 130 + 0.70 × 38.8 ≈ 157 mmHg — an apparent fall of about 12 mmHg produced by nothing but the enrolment rule.

A fall of that size is on the same scale as the reductions reported for antihypertensive monotherapy. A single-arm study using that enrolment rule would therefore report a clinically impressive result before its intervention did anything at all.

Regression to the mean is not a force acting on individuals

Regression to the mean describes the expected behaviour of a selected group, not a pull exerted on any person, and confusing the two produces a specific class of wrong conclusion. No individual is under pressure to become average; a person selected at the extreme is simply more likely than not to have been measured on an unrepresentative day.

The corollary that surprises people is that regression to the mean does not shrink the population. If it did, variability would contract with every round of measurement and everyone would eventually be identical. The distribution is stationary: for every selected high scorer who falls back, someone previously near the mean rises. Regression is symmetric and appears only because the group was selected.

Horace Secrist’s 1933 book The Triumph of Mediocrity in Business is the canonical failure to see this. Secrist tracked firms with extreme profits, found them less extreme later, and read an economic law into the arithmetic. Harold Hotelling’s review compared the demonstration to proving the multiplication table by arranging elephants in rows. Galton himself initially read his stature data the same way, as a hereditary tendency towards mediocrity, before recognising it as a property of imperfect correlation.

Where regression to the mean enters clinical and laboratory research

Regression to the mean enters a study through the selection rule, so the tell is always an eligibility criterion, a subgroup definition, or a sampling step that uses the outcome variable itself. Five patterns account for most of what appears in submitted manuscripts.

Threshold eligibility on the primary outcome. “HbA1c ≥ 8.0% at screening”, “at least four exacerbations in the preceding year”, “pain score ≥ 7 on an 11-point scale”. Each guarantees a selected group whose mean will fall regardless of treatment.

Single-arm before-and-after designs. Uncontrolled phase II studies, quality-improvement projects, service evaluations and historically controlled series have no arm in which regression can be observed and subtracted.

Responder and non-responder analyses. Defining responders by change from baseline splits the sample on a quantity that is mathematically correlated with baseline, so “responders” are enriched for people whose baseline was inflated by error.

Baseline-severity subgroup effects. The finding that “patients with the highest baseline values improved most” is the expected result of regression to the mean and is almost never evidence of effect modification when the subgroups were defined by the baseline value.

Extreme-value discovery. Selecting the largest observed effects in a high-dimensional screen — the highest fold changes, the top-ranked variants, the strongest correlations — selects for upward error, so replication estimates shrink. Multiple-testing correction controls the error rate of that screen but does nothing to the magnitude of the surviving estimates, which remain inflated by exactly the selection that made them significant.

Why randomisation protects the between-arm comparison but not the within-arm one

Randomisation removes regression to the mean from the treatment contrast in a controlled trial, and removes it from nothing else. When enrolment selects on a baseline threshold, both arms are selected identically, both regress by the same expected amount, and the difference between arms is unbiased. The regression cancels in the subtraction.

The within-arm change does not cancel. Reporting that “systolic pressure fell by 18 mmHg in the treatment arm” as a result attributable to treatment is invalid even in a properly randomised trial, because part of that 18 mmHg is regression shared with the control arm. Only the between-arm difference estimates the effect.

Uncontrolled designs have no subtraction available. This is one reason placebo-controlled and no-treatment-controlled comparisons diverge: Hróbjartsson and Gøtzsche’s 2001 New England Journal of Medicine analysis of trials randomising patients to placebo or to no treatment found little evidence of clinically important placebo effects on objective outcomes. If the placebo response on such outcomes is small, then much of what an uncontrolled study attributes to placebo is better read as regression to the mean and natural history than as a response to the ritual of treatment. Trial-design reporting is checked in more detail in clinical trial review.

The correlation between baseline and change is negative by construction

The correlation between a baseline value and the subsequent change from that baseline is negatively biased by arithmetic alone, and it is one of the most commonly reported artefacts of regression to the mean in the applied literature. Because the baseline value X appears with a negative sign inside the change Y − X, the two are coupled even when the change is entirely random.

For measurements with equal variance and correlation ρ, the correlation between baseline and change is exactly −√((1 − ρ)/2). At ρ = 0.70 that is −0.39; at ρ = 0.50, −0.50. A manuscript reporting “greater improvement in patients with worse baseline values, r = −0.4, p < 0.001” has usually reported the reliability of its own instrument, not a clinical finding.

Oldham’s 1962 paper in the Journal of Chronic Diseases proposed the standard remedy: regress the change (Y − X) on the average (X + Y)/2 rather than on X, because the difference of the two measurement errors is uncorrelated with their average. Oldham’s method answers a subtly different question and has its own critics, but a manuscript that plots change against baseline without acknowledging the coupling has produced no evidence at all. This is a common route to an effect that appears or disappears with adjustment.

Change scores, ANCOVA, and Lord’s paradox

Analysis of covariance on the follow-up value with baseline as a covariate is the standard analysis for a randomised pre-post trial, and Vickers and Altman set out the reasoning in BMJ 2001;323:1123. ANCOVA adjusts each participant’s follow-up score for their own baseline, is unaffected by chance baseline imbalance, and does not inherit regression to the mean the way a comparison of change scores does. It is also more precise: the gain in power over a change-score analysis grows as the baseline-follow-up correlation falls.

Change-score and ANCOVA analyses answer the same question in a randomised trial and can be expected to agree. In observational data they answer different questions and routinely disagree — this is Lord’s paradox, from Frederic Lord’s 1967 example of boys and girls eating in the same university dining halls, in which the change-score comparison shows no difference between the sexes while the baseline-adjusted comparison shows a large one. Neither number is wrong; they estimate different contrasts, and choosing between them requires a causal argument rather than a statistical one. The choice depends on whether baseline is a confounder of the exposure-outcome relationship or a variable the exposure has already affected, in which case adjusting for it can open a collider path.

How regression to the mean is detected in a manuscript

Regression to the mean is detected structurally rather than statistically: no test identifies it, and the biased estimate looks like an ordinary number with ordinary confidence limits. A reviewer works through the design description looking for the selection rule.

Read the eligibility criteria first and ask whether any criterion is stated as a threshold on the outcome variable or on something strongly correlated with it. Then ask how many measurements the threshold was applied to — a single screening reading is the highest-risk case, and papers rarely say. Then check whether the value used to determine eligibility is the same value reported as baseline; when it is, the baseline is upward-biased by construction and every change from it is inflated.

Next, look at the results structure. A single-arm pre-post table, a subgroup analysis split by baseline tertiles, a responder definition based on change, a scatterplot of change against baseline, or a discussion sentence beginning “patients with more severe disease at entry benefited most” each signal the same problem. In a high-dimensional analysis, look for discovery effect sizes reported without shrinkage and for a replication cohort whose estimates are uniformly smaller, which is the expected pattern rather than evidence of a failed replication.

Finally, search the manuscript for the phrase itself. A paper that selects on a threshold and never uses the words “regression to the mean” has not considered it; a paper that names it once in the limitations without quantifying it has acknowledged it without addressing it. Both are common, and the second reads worse to a methodologically alert reviewer than the first.

What a peer reviewer says when regression to the mean is mishandled

Reviewers rarely write “regression to the mean” as their opening phrase. The comments arrive as objections to the design or the interpretation, and phrasings of the following kind are typical of the statistical objections raised against a study that has selected on an extreme baseline.

“Participants were enrolled on the basis of a single elevated measurement, so the observed improvement cannot be distinguished from regression to the mean.” “The uncontrolled design does not permit attribution of the within-group change to the intervention.” “The reported association between baseline severity and change is expected under mathematical coupling and does not demonstrate effect modification.” “Please report the between-group difference rather than the within-group changes.” “The subgroup defined by baseline value is not a valid basis for claiming a differential treatment effect.” “How many measurements were averaged to determine eligibility, and was the same measurement used as the baseline?” “The authors should quantify the expected regression given the selection threshold and the reliability of the instrument.”

The last comment is the one that is hard to answer after the fact, because it requires a reliability estimate the study may never have collected. Anticipating it at the design stage takes one additional screening visit; answering it at revision often cannot be done.

Design remedies

Regression to the mean is removed or reduced at the design stage far more effectively than in analysis, and three measures do most of the work.

Average several baseline measurements. Because regression is proportional to (1 − ρ), and the reliability of a mean of k measurements follows the Spearman–Brown relation kρ/(1 + (k − 1)ρ), averaging three readings with single-reading reliability 0.70 raises reliability to 0.875 and cuts the expected regression from 0.30 to 0.125 of the selection deviation — a reduction of about 58% in exchange for two extra readings.

Separate the selection measurement from the baseline measurement. Use one set of readings to decide eligibility and an independent later set as the analysis baseline. The second set is not conditioned on the threshold, so the group’s true mean is estimated without the upward bias, and the change from it is not inflated.

Randomise, and use a concurrent control. Randomisation cancels regression in the between-arm contrast exactly. Where randomisation is impossible, a concurrent comparison group selected by the same threshold serves the same function, which is why an interrupted time series with a comparison series is stronger than a simple before-and-after. Preclinical work faces the same requirement when animals are allocated by tumour volume or another baseline reading, and a caliper measurement on a small tumour is a low-reliability measurement by any standard.

Analysis remedies and their limits

Regression to the mean can be partly corrected in analysis, and every correction depends on a reliability estimate that the study must supply. Barnett, van der Pols and Dobson set out the practical options in International Journal of Epidemiology 2005;34:215–220, which remains the standard applied reference.

For a group selected above a threshold, the expected regression (1 − ρ)σλ can be estimated and subtracted, with a confidence interval that propagates uncertainty in ρ. This requires an external or internal estimate of the between-occasion correlation; using the study’s own selected sample to estimate ρ underestimates it, because selection restricts the range.

For high-dimensional discovery, shrinkage estimators are the appropriate correction. Empirical Bayes and conditional-likelihood methods pull selected effect sizes back towards the null by an amount that depends on the selection threshold and the standard error. In genome-wide association work, where the same bias is known as the winner’s curse, these corrections are the standard way to produce a discovery estimate that a replication cohort has some prospect of reproducing. Reporting a discovery effect size and a shrunk estimate side by side is more informative than reporting either alone, and is expected in genomics review.

The limit of all analytic correction is that it addresses the magnitude of the bias, not its existence. A corrected single-arm estimate is still an estimate from a design with no counterfactual, and a reviewer is entitled to say so.

What regression to the mean is not

Regression to the mean is frequently conflated with four adjacent phenomena, and distinguishing them determines which remedy applies.

It is not confounding. Confounding comes from a common cause of exposure and outcome and is addressed by adjustment; regression to the mean comes from selection on a noisy measurement and adjustment does not touch it.

It is not the placebo effect. Both produce apparent improvement in an untreated group, but the placebo effect is a response to treatment context, while regression to the mean would occur if the participants were never seen again. Only a no-treatment control separates them.

It is not regression dilution. Regression dilution is the attenuation of a slope caused by measurement error in the predictor. It shares a cause with regression to the mean — imperfect reliability — but describes bias in an estimated relationship rather than movement of a selected group.

It is not a time-related bias. Immortal time bias and lead time bias arise from how follow-up time and diagnosis timing are handled, not from selection on an extreme value, and no amount of repeated baseline measurement addresses them.

How to report it

State the selection rule precisely, including how many measurements the eligibility threshold was applied to and whether the same measurement served as the analysis baseline. Report the reliability of the primary measure, from the study itself or from a named external source, because every quantitative statement about regression depends on it.

Report the between-group difference as the effect estimate and present within-group changes as description only, labelled as such. If the design is uncontrolled, calculate the expected regression from the threshold and the reliability, report it alongside the observed change, and let the reader see the comparison. If a baseline-defined subgroup appears to benefit more, say explicitly that the pattern is expected under regression to the mean, and give the analysis that distinguishes the two or state that the data cannot distinguish them. Stating that limitation plainly is stronger than a general sentence about generalisability, and it forecloses the overclaim that reviewers respond to most sharply.

Related

Confounding · Collider bias · Table 2 fallacy · My effect disappeared after adjusting

Checked before submission by causal language discipline, which flags within-group change presented as a treatment effect and baseline-defined subgroup claims.

Review my manuscript

Frequently asked questions

What does regression to the mean mean in statistics?

Regression to the mean means that a group selected for extreme values on a measurement containing error will, on remeasurement, average closer to the population mean. The expected movement is (1 − ρ) times the selection deviation, where ρ is the correlation between the two measurements.

Can you give a simple example of regression to the mean?

Patients enrolled because a single blood pressure reading exceeded 160 mmHg will average lower at a second visit even with no treatment, because some were captured on an unrepresentative day. With a population standard deviation of 20 mmHg and reliability 0.70, the expected fall is about 12 mmHg.

Why does regression to the mean happen?

Regression to the mean happens because an observed value contains a stable component and a transient one. Selecting extreme observations selects for both, and only the stable component persists at remeasurement. The transient component is redrawn from its distribution, so the group’s average moves back towards the population mean.

What is regression towards mediocrity?

Regression towards mediocrity is Francis Galton’s 1886 term for what is now called regression to the mean. Galton reported that a child’s height deviated from the population average by roughly two-thirds of the mid-parent deviation, and initially read this as a hereditary tendency rather than a consequence of imperfect correlation.

Does regression to the mean mean that extremes disappear over time?

No. Regression to the mean does not compress a population, because the effect is symmetric: individuals near the average move outward as often as selected extremes move inward. The spread of the distribution is stationary. The movement is visible only when a group has been selected on an extreme measured value.

What causes regression to the mean in repeated measurements?

Measurement error and within-person biological variability cause regression to the mean in repeated measurements. The magnitude is set entirely by the test-retest correlation: at a correlation of 1.0 there is no regression, and at 0.0 a selected group’s expected follow-up value is the population mean.

Is regression to the mean a real effect or a statistical artefact?

Regression to the mean is a real, predictable property of imperfectly correlated repeated measurements, not a measurement mistake. It becomes an artefact only when the resulting movement is interpreted as a treatment effect, a hereditary law, or a differential response by baseline severity.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.