Skip to content

SOLUTIONS

What is an effect size?

An effect size states how large a difference, ratio or association is, on a scale a reader can judge. A p-value cannot supply that number, and never has.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is an effect size?

An effect size is a number that states how large a difference, ratio or association is, on a scale a reader can interpret — 4.2 mmHg, a hazard ratio of 0.68, a Cohen’s d of 0.41. An effect size answers how much, whereas a p-value answers only how surprising the data would be if nothing were happening.

Those two questions are separable, and conflating them is among the statistical faults reviewers raise most often. A study can be overwhelmingly significant and describe a change too small for any patient to notice, or non-significant while pointing at a difference that would change practice if it held. This page covers the families of effect size and what each one is on the scale of, why standardisation both helps and misleads, why Cohen’s conventional benchmarks fail in most applications, why small significant studies systematically overstate magnitude, how the problem is detected in a real manuscript, and the phrasings reviewers use when it has been mishandled.

What an effect size measures that a p-value does not

An effect size and a p-value answer different questions, and only one of them is about magnitude. A p-value is a joint function of three things: the size of the underlying effect, the variability of the measurements, and the sample size. Because sample size enters, a p-value can be driven to any value by collecting more data, whatever the magnitude of the effect.

Consider a trial of 20,000 participants that reports p < 0.001 for a mean systolic blood pressure reduction of 0.3 mmHg. The p-value is correct and the finding is uninteresting: 0.3 mmHg is smaller than the measurement error of a cuff and smaller than the within-person variation between two consecutive readings. Now consider a 24-participant crossover study reporting p = 0.09 for a 9 mmHg reduction with a 95% confidence interval of −1.4 to 19.4 mmHg. That result is compatible with a clinically decisive effect and with nothing, and the correct summary is the interval rather than the verdict.

The effect size is the quantity that survives translation. Doubling the sample changes the p-value; it does not change what the population difference is, only how precisely you have estimated it.

The four families of effect size

Effect sizes fall into four families, and the first decision in reporting one is which family the question belongs to.

Differences on the original measurement scale. A mean difference of 4.2 mmHg, a median difference of 3 days in hospital, a risk difference of 1.8 percentage points, a difference in fluorescence intensity in arbitrary units. These are unstandardised, and where the scale is meaningful to the reader they are the most informative choice available.

Standardised differences. Cohen’s d expresses a mean difference in units of a pooled standard deviation. Hedges’ g applies a small-sample correction to d — the multiplier is approximately 1 − 3/(4·df − 1), which shrinks d by about 4% at n = 10 per group and under 1% at n = 50 per group. Glass’s Δ divides by the control group’s standard deviation alone, which is the right choice when the intervention plausibly changes the variance as well as the mean.

Ratios. Risk ratio, odds ratio, rate ratio, hazard ratio and fold change are multiplicative. A ratio of 1 means no effect, ratios are asymmetric around 1, and their confidence intervals must be computed on the log scale and back-transformed.

Association and variance explained. Pearson’s r, R², eta-squared, partial eta-squared, omega-squared and Cramér’s V describe how much of the variation in one variable travels with another. Squared measures are not on the same scale as unsquared ones: r = 0.30 and R² = 0.09 are the same fact stated twice.

A fifth, less used family describes overlap directly. The probability of superiority — the chance a randomly drawn treated observation exceeds a randomly drawn control observation — is Φ(d/√2) under normality, so d = 0.5 corresponds to about 64%, and d = 0.2 to about 56%. Non-parametric versions include Cliff’s δ and the Vargha–Delaney A statistic, which are appropriate for ordinal outcomes where a mean difference is not defined.

Standardisation helps comparison and hides the units

Standardising a difference divides it by a standard deviation, and the standard deviation is a property of the sample studied, not of the intervention. Two identical interventions can therefore report different effect sizes.

A 5 mmHg reduction in a general-population sample with a standard deviation of 20 mmHg gives d = 0.25. The same 5 mmHg reduction in a narrowly selected trial cohort with a standard deviation of 10 mmHg gives d = 0.50. Nothing about the drug changed. Restriction of range works in the opposite direction for correlations, attenuating r when the sample spans less of the underlying variable than the population does.

Two operational consequences follow. First, report the unstandardised effect whenever the scale carries meaning — millimetres of mercury, days, grams, counts — and add the standardised version only if a meta-analytic audience needs it. Second, treat any cross-study comparison of d as conditional on the two samples having comparable dispersion, and say so rather than assuming it.

Cohen’s benchmarks and why they mislead

Jacob Cohen proposed 0.2, 0.5 and 0.8 as small, medium and large values of d in Statistical Power Analysis for the Behavioral Sciences, and he framed them as conventions for use when no better basis for judgement exists, relative to the typical magnitudes of the behavioural sciences of his era. The benchmarks were never offered as a field-independent scale.

Applied outside that framing, they invert the substantive judgement in both directions. In a large cardiovascular prevention trial, a standardised effect near 0.1 can correspond to thousands of events prevented across a treated population and be worth the intervention. In a single-cell assay with an artificial dynamic range, d = 1.2 can reflect a change of no biological consequence, or a technical artefact of normalisation.

The alternative is not another set of thresholds. The alternative is an anchor argued in the manuscript: a minimal important difference from prior work, a comparison to the effect of an established treatment on the same outcome, the population-level implication of the difference, or a defined biological threshold. A stated anchor is arguable, and reviewers can disagree with it. “A medium effect by conventional standards” gives them nothing to engage with, which is why it reads as evasion.

Effect size estimates from small samples are biased upward

An effect size estimated from a small sample and filtered through a significance threshold is systematically too large. The mechanism is a selection effect rather than a computational error: with low power, only the larger sample estimates clear the threshold, so the published estimates are drawn from the upper tail of the sampling distribution.

Gelman and Carlin (2014) formalised this as the type M (magnitude) error — the expected factor by which a statistically significant estimate exaggerates the true effect — alongside the type S (sign) error, the probability that a significant estimate has the wrong sign. Both can be computed before data collection from a plausible effect size, and both grow sharply as power falls. Ioannidis (2008) made the same argument for the epidemiological literature under the heading of inflated discovered associations.

The empirical footprint is visible in replication work. The Open Science Collaboration’s 2015 replication of 100 psychology studies reported a mean original effect size of r = 0.403 against a mean replication effect of r = 0.197, with 97% of original studies significant and 36% of replications so. Shrinkage of that order is what an underpowered, threshold-filtered literature predicts.

Three practical responses exist. Use bias-corrected estimators where the correction is known — Hedges’ g rather than d, omega-squared rather than eta-squared. Design analyses with a defensible effect size rather than the one your pilot returned, since a pilot estimate is drawn from the same biased distribution. And treat a first, small, significant result as an upper bound rather than a point estimate, which is the honest framing when a reviewer asks whether the experiment has been replicated.

Precision belongs to the estimate

An effect size reported without an interval is half a result. The confidence interval states which population values the data are compatible with, and the width of that interval is frequently the more decisive fact.

CONSORT 2010 makes this explicit for trials. Item 17a requires, verbatim, “For each primary and secondary outcome, results for each group, and the estimated effect size and its precision (such as 95% confidence interval).” The APA Publication Manual directs authors across the psychological literature to report effect sizes and, wherever possible, confidence intervals, and the APA Journal Article Reporting Standards list both among the elements a quantitative results section is expected to contain. The ICMJE Recommendations instruct authors not to rely solely on statistical hypothesis testing such as P values, which they describe as failing to convey important information about effect size and precision of estimates.

Two intervals with the same point estimate can support opposite conclusions. A hazard ratio of 0.85 with a 95% interval of 0.78 to 0.93 supports a modest, well-estimated benefit. A hazard ratio of 0.85 with an interval of 0.42 to 1.71 supports nothing at all, and describing the second as “a 15% reduction in hazard” is the sentence that draws the sharpest reviewer comment. Note also that a ratio measure is uninterpretable without its baseline: a hazard ratio of 0.68 against a 2% annual event rate and against a 40% annual event rate describe very different clinical situations, and the absolute figures should appear beside the ratio. The same applies to competing risks, where a cause-specific hazard ratio and a subdistribution hazard ratio answer different questions and are routinely reported as if interchangeable.

Statistical significance, practical significance, and the anchor in between

Practical significance is a claim about the world, and it requires a benchmark external to the dataset. The standard instrument is the minimal important difference: the smallest change on an outcome that patients or clinicians regard as worthwhile.

Named examples make the idea concrete. The 2014 European Respiratory Society and American Thoracic Society technical standard on field walking tests places the minimal important difference for six-minute walk distance in chronic respiratory disease at roughly 25 to 33 m, so a trial reporting a statistically significant 12 m gain has demonstrated a real but sub-threshold change and should say so. In event-driven work the anchor is often the number needed to treat: an intervention that moves a 4.0% event rate to 3.0% delivers a 25% relative reduction, a 1.0 percentage point absolute reduction, and a number needed to treat of 100 — three descriptions of one result, with very different rhetorical weight.

Minimal important differences are themselves contested. Estimates for the same instrument differ by anchor method, by baseline severity, and between improvement and deterioration, and some widely quoted values rest on single small studies. Cite the source and the population for any threshold you invoke rather than presenting it as a constant.

Effect sizes in laboratory and omics work

Preclinical and high-throughput studies report effect sizes under different names, and the same reasoning applies to each.

Fold change is a ratio effect size, usually reported as log2 fold change in transcriptomic and proteomic work. Raw fold changes computed from low-count features are unstable, which is why DESeq2 shrinks log2 fold changes with an empirical Bayes prior and limma moderates variance estimates across features. Reporting unshrunken fold changes for lowly expressed genes reliably produces a multiple-testing and effect-inflation objection.

Percentage change in an image-derived measurement is an effect size on an arbitrary scale, and its interpretability depends entirely on the normalisation and the segmentation decisions upstream of it. A 40% increase in mean fluorescence intensity means little without the exposure settings, the background subtraction and the number of fields, which is the substance of most image quantification comments.

Difference between conditions in cultured cells is frequently reported with a denominator that counts wells or images rather than independent biological units. The effect size may be estimated correctly while its precision is fabricated — the defining feature of pseudoreplication, and the reason p-values shrink as more cells are counted without any change in the underlying magnitude.

Batch structure can manufacture an effect size outright. When condition is confounded with processing day, plate or sequencing run, the reported difference includes the batch difference, and no correction applied after the fact fully separates them. Batch effect scrutiny is therefore an effect size question, not only a normalisation question.

How mishandled effect sizes are detected in a manuscript

Detecting a mishandled effect size is a reading procedure rather than a statistical test, and it runs on the manuscript text as submitted.

Check that every claim in the abstract carries a magnitude. Abstract sentences of the form “treatment significantly improved X” with no number, matched against a results table where the difference is small, is a reliable signal.

Check the units of every reported effect size. A d, an r, an η² and a fold change reported in one results section without stating which is which, or a partial eta-squared labelled as eta-squared, indicates the values were taken from software output rather than chosen.

Check whether an interval accompanies every point estimate, and whether the discussion argues from the point estimate while the interval crosses the null.

Check the direction of the standardisation. If d is large while the raw difference is small, look for a restricted sample or a within-subject standard deviation used where a between-subject one belongs. Both inflate d without any change in the phenomenon.

Check that a ratio has its baseline. Relative risks, hazard ratios and fold changes presented without absolute rates or baseline levels cannot be judged, and the omission usually favours the authors’ preferred reading.

Check whether the interpretation rests on Cohen’s labels. The words “small”, “medium” and “large” appearing without an anchor from the literature indicate the practical-significance argument was never made.

Check the sample size against the claimed magnitude. A large standardised effect from a handful of units per group is a candidate type M error, and the discussion should acknowledge that rather than treat the estimate as settled. This is where an underpowered design and an overstated conclusion meet.

What a peer reviewer says when it has been mishandled

Reviewer phrasings for a mishandled effect size are stable across fields, and recognising them in advance is the point of reading them here.

“The authors report p-values throughout but no effect sizes or confidence intervals.” “The reported difference is statistically significant but its clinical importance is not established.” “The abstract states that the intervention improved outcomes without stating by how much.” “Effect sizes are described as medium and large by conventional criteria; the authors should justify these thresholds in the context of this outcome.” “Given the sample size, the observed effect is likely to be an overestimate, and the discussion does not acknowledge this.” “The odds ratio is reported for a common outcome and is being interpreted as a risk ratio.” “The confidence interval includes values consistent with no effect, and the conclusion should be tempered accordingly.”

The comment that most often decides a paper is the last one in substance rather than the first: a reviewer who accepts your numbers but rejects your reading of them. Anticipating it is a matter of aligning the discussion with the interval instead of the point estimate, which is also what an overclaim check looks for. When a reviewer has already written that the effect size is not meaningful, the response required is an anchor, not a defence of the p-value.

How to report an effect size well

Report the unstandardised effect with its units first, then the interval, then the standardised version if the audience needs it: “mean difference 4.2 mmHg (95% CI 1.1 to 7.3), d = 0.31”. Name the estimator explicitly — Hedges’ g rather than “effect size”, omega-squared rather than “variance explained” — and state the denominator used for any standardisation.

Give the baseline alongside every ratio, and the absolute difference alongside every relative one. State the anchor by which you are judging importance and cite where it comes from. Where the study was powered on an assumed effect, report that assumption and compare it to what was observed, rather than converting the observed result into retrospective power, which is a deterministic function of the p-value and adds no information.

Finally, write the discussion from the interval. If the interval spans trivial and substantial values, the sentence that follows is “the data are compatible with effects ranging from negligible to clinically important”, and reviewers of statistical objections reward that sentence far more than they reward confidence.

Contested ground and honest limits

Several points in the reporting of effect sizes are genuinely unsettled, and a manuscript that pretends otherwise reads worse than one that names them.

Whether standardised effect sizes should be used at all in clinical research is disputed: they permit meta-analysis across instruments and simultaneously invite comparisons across incomparable populations. Which variance-explained measure to prefer is disputed, with partial eta-squared dominant in practice and criticised for not being comparable across designs with different factor structures. Minimal important differences differ by derivation method for the same instrument. And the magnitude of literature-wide inflation is estimated rather than known — replication programmes measure it in particular fields under particular selection rules, and those estimates do not transfer cleanly to yours.

None of that licenses silence. Each is a bounded uncertainty that can be stated in a sentence with its source, and stating it is what distinguishes a discussion section a reviewer trusts from one they audit.

Related

When a reviewer says the effect size is not meaningful · Post-hoc power · Sample size too small · Pseudoreplication · Confounding · Statistical objections

Checked before submission by overclaim detection, which compares the magnitude in the results table against the strength of the claim in the abstract and discussion, and names the sentences that outrun the interval.

Review my manuscript

Frequently asked questions

What does effect size mean in statistics?

Effect size means the magnitude of a difference, ratio or association, expressed on a scale that can be judged — a mean difference of 4.2 mmHg, a hazard ratio of 0.68, a Cohen’s d of 0.41. Effect size is independent of sample size, which is precisely what distinguishes it from a p-value.

What is an effect size in plain terms?

An effect size is the answer to “how much?”. Where a p-value reports how unlikely the data would be under a null hypothesis, an effect size reports how large the difference actually is. A study can have a very small p-value and an effect size too small to matter to anyone.

What is an effect size measure?

An effect size measure is a specific statistic used to express magnitude: Cohen’s d, Hedges’ g and Glass’s Δ for standardised differences; risk ratio, odds ratio and hazard ratio for ratios; Pearson’s r, R², eta-squared and omega-squared for association; Cliff’s δ for ordinal data. Each is on a different scale and must be named explicitly.

How would you define effect size?

Effect size is defined as a quantitative statement of the magnitude of a phenomenon, reported with the units or scale on which it was measured and with an interval showing how precisely it was estimated. Effect size does not become larger with more participants; only the precision of its estimate improves.

What counts as an effect size in a paper?

Any number that states magnitude counts: an absolute risk difference, a number needed to treat, a log2 fold change, a percentage change in fluorescence intensity, a regression coefficient with its units. A test statistic such as t or F is not an effect size, and neither is a p-value.

What is an example of an effect size?

An intervention that moves an event rate from 4.0% to 3.0% has a risk difference of 1.0 percentage point, a risk ratio of 0.75 and a number needed to treat of 100. All three are effect sizes describing one result, and the choice among them changes how large the finding sounds.

What is a good effect size?

There is no size that is good in the abstract, because importance requires an external anchor: a minimal important difference, the effect of an established treatment on the same outcome, or a defined biological threshold. Cohen’s conventional labels of 0.2, 0.5 and 0.8 were offered as a last resort when no such anchor exists, and were never intended as a field-independent scale.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.