Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Statistical power is the probability a design detects an effect of a stated size. Power is fixed before data exist, which is why power computed afterwards is circular.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Statistical power is the probability that a study will detect an effect of a specified size, if an effect of that size is real. Power is a property of a design — sample size, variability, the effect assumed, and the Type I error rate — fixed before any data exist. Trials are conventionally designed at 80% or 90% power.
That definition is compact, and nearly every practical difficulty comes from the phrase “of a specified size”. Power is never a single number attached to a study; it is a number attached to a study and an assumed effect. Change the assumed effect and the power changes, without a single observation moving. This page covers the four quantities that determine power, where the 80% convention comes from, worked examples for continuous and time-to-event endpoints, the design choices that move power more than adding participants does, what low power does to a published literature, how underpowering is detected in a manuscript, and what reviewers say when it has been handled badly.
Statistical power describes a design, not a result. Power answers the question “if the true effect were δ, how often would a study of this size and this variability produce a significant test?” — a question that can be answered completely before recruitment opens and that the data, once collected, do not update.
Power is the complement of the Type II error rate: power = 1 − β, where β is the probability of failing to reject a null hypothesis that is false by exactly δ. The two error rates are not symmetric in how they are treated. The Type I error rate α is a property of the test alone and holds whatever the truth is. Power depends on a quantity nobody knows, which is why a power calculation is an argument about an assumption rather than a computation from data.
This asymmetry is the root of most confusion in the literature. A trial reported as “90% powered” was 90% powered to detect the effect its authors assumed. If the true effect is half that, the same trial had roughly 37% power, and no amendment to the manuscript changes the arithmetic.
A power calculation ties together four quantities, and fixing any three determines the fourth: the effect size to be detected, the variability of the outcome, the sample size, and the Type I error rate.
The trade among them is not linear, and the non-linearity is what surprises people. For a two-sample comparison of means, the required sample size per group is proportional to the square of the variability divided by the square of the effect. Halving the effect you want to detect quadruples the sample size. A study that needs 60 participants per arm to detect a 0.5 standard-deviation difference needs about 240 to detect 0.25, and about 960 to detect 0.125.
Variability enters the same way, which is why measurement quality is a power intervention. Reducing the residual standard deviation by 30% — through a better instrument, a more standardised protocol, or repeated measurement of the same participant — reduces the required sample size by about half, and does so without recruiting anyone.
The Type I error rate enters through the critical value rather than by squaring, so it is the weakest of the four levers, but it is not negligible when multiplicity is involved. That case is set out below under design choices.
The 80% and 90% statistical power conventions are codified in regulatory guidance rather than derived from any statistical theorem — and the guidance records the convention rather than creating it. ICH E9, Statistical Principles for Clinical Trials, section 3.5, states plainly: “The probability of type II error is conventionally set at 10% to 20%; it is in the sponsor’s interest to keep this figure as low as feasible especially in the case of trials that are difficult or impossible to repeat.”
The same section opens with the requirement the calculation exists to satisfy: “The number of subjects in a clinical trial should always be large enough to provide a reliable answer to the questions addressed.” It also names the source of the assumed effect: “The treatment difference to be detected may be based on a judgement concerning the minimal effect which has clinical relevance in the management of patients or on a judgement concerning the anticipated effect of the new treatment, where this is larger.”
One clause in ICH E9 section 3.5 is quoted far less often and is the one reviewers increasingly ask about: “It is important to investigate the sensitivity of the sample size estimate to a variety of deviations from these assumptions.” A single number with no sensitivity range does not meet the guidance it is usually written to satisfy. In non-clinical fields, no equivalent convention has any authority behind it at all — 80% has simply been inherited, and a laboratory study designed at 80% power for an effect nobody has measured is following a habit rather than a standard.
Statistical power for a two-arm parallel trial with a continuous primary outcome has a closed form. Take a two-sided 5% Type I error rate and a standardised effect of Δ = δ/σ.
The normal approximation gives the required number per group as n = 2(z<sub>0.975</sub> + z<sub>1−β</sub>)² / Δ². At 80% power the bracket is (1.96 + 0.8416)² = 7.85, so n ≈ 15.7/Δ² per group. At 90% power it is (1.96 + 1.2816)² = 10.51, so n ≈ 21.0/Δ² per group.
For Δ = 0.5 — a moderate effect — that is 63 per group at 80% power and 84 at 90%; standard tables give 64 and 85, the difference being the t-distribution correction the normal approximation omits. For Δ = 0.25 the requirement is 251 per group, and for Δ = 0.2 it is 393.
For a binary outcome, the same structure applies with the variance written in terms of the two proportions: n = (z<sub>0.975</sub> + z<sub>1−β</sub>)²[p₁(1−p₁) + p₂(1−p₂)]/(p₁−p₂)². A trial expecting to reduce an event rate from 20% to 15% needs 7.85 × 0.2875 / 0.0025 ≈ 903 participants per arm at 80% power. The 25% relative reduction sounds modest; the 1,806 participants it requires is the number that determines whether the trial happens.
In a time-to-event trial, statistical power is driven by the number of events observed, not by the number of participants enrolled. Enrolling more participants raises power only through the events they contribute, which is why an event-driven trial specifies a target event count and analyses when that count is reached rather than on a calendar date.
Schoenfeld’s formula makes the relationship explicit: the required number of events is d = 4(z<sub>0.975</sub> + z<sub>1−β</sub>)²/(log HR)². For a hazard ratio of 0.75 at 90% power and a two-sided 5% Type I error rate, that is 4 × 10.51 / (log 0.75)² = 42.0 / 0.0828 ≈ 508 events. At 80% power the same hazard ratio needs about 379.
The squared logarithm is where trial programmes are decided. Moving the target hazard ratio from 0.75 to 0.85 raises the event requirement from about 508 to about 1,590 — a threefold increase from a change most clinicians would describe as small. This is also why a trial with a lower-than-expected event rate is in trouble even when recruitment succeeds: the participants arrived, the events did not, and power fell with them. Trials that switch to a composite endpoint mid-course are usually solving this problem, and the switch is visible on the registry record. Registration and endpoint checking is the pass that catches it, and clinical trial review covers the wider reporting set.
Several routine design and analysis decisions change the sample size a given statistical power requires by more than any plausible recruitment increase would, and each is made before a participant is enrolled.
Dichotomising a continuous variable. Cohen’s 1983 analysis in Applied Psychological Measurement showed that dichotomising one of two bivariate normal variables at the mean reduces the correlation to 0.798 of its original value, and so the variance accounted for to 0.637r² — a loss equivalent to discarding roughly 36% of the sample, or needing about 1.6 times as many participants for the same power. Dichotomising both variables costs more still: the surviving correlation is (2/π)·arcsin(r), which at r = 0.4 leaves about 0.43r² — equivalent to discarding well over half the sample. (Cohen’s own figure for this case, 0.405r², assumed the single-variable factor simply squares; MacCallum, Zhang, Preacher and Rucker, Psychological Methods, 2002, show that it does not.) Converting a continuous endpoint into “responder / non-responder” is the most common way a study discards power it already had.
Adjusting for baseline. Analysis of covariance with the baseline value as a covariate reduces the residual variance by a factor of (1 − ρ²), where ρ is the baseline-outcome correlation. At ρ = 0.7 the required sample size falls to about 51% of the unadjusted figure. Change-from-baseline analysis, by contrast, is only more efficient than the raw endpoint when ρ exceeds 0.5.
Clustering. In a cluster-randomised trial the required sample size is multiplied by the design effect, 1 + (m − 1)ρ, where m is the cluster size and ρ the intracluster correlation. With 50 participants per practice and an intracluster correlation of 0.02 — a typical primary-care value — the design effect is 1.98, so the trial needs almost twice the participants of an individually randomised equivalent. Ignoring clustering does not merely lose power; it inflates the Type I error rate, which is the pseudoreplication problem in a trial setting.
Multiplicity. Controlling the family-wise error rate lowers α, and lowering α raises the required sample size. A Bonferroni correction across 10 primary comparisons sets α at 0.005, whose critical value 2.807 gives a bracket of 13.31 against 7.85 — about 70% more participants per group for the same power. At the genome-wide threshold of 5 × 10⁻⁸ the bracket is 39.6, roughly five times the single-test requirement. See multiple comparisons objections for how this is argued in review.
Interactions and subgroups. Brookes and colleagues (Journal of Clinical Epidemiology, 2004) showed that a trial with 80% power for an overall effect has about 29% power to detect an interaction of the same magnitude, and that sample sizes must be inflated roughly fourfold to recover it. Nearly every subgroup analysis in the literature is underpowered by construction, and a non-significant interaction test in a trial powered for a main effect is close to uninformative.
Attrition. Inflating the calculated sample size by 1/(1 − d) for an anticipated dropout proportion d is the standard remedy: 20% loss means recruiting 25% more. This works only if dropout is unrelated to outcome; when it is related, the problem is bias rather than power, and no inflation factor addresses bias.
Low statistical power does more than produce false negatives, and the additional damage is the part most authors have not internalised. An underpowered study that reaches significance necessarily overstates the effect.
The arithmetic is forced. At 20% power for a two-sided 5% test, a result is significant only if the estimate exceeds 1.96 standard errors from zero, while the true effect sits at about 1.12 standard errors. Every publishable estimate from that design therefore overstates the truth by at least 75%, and the average significant estimate overstates it by more. Gelman and Carlin (Perspectives on Psychological Science, 2014) formalised this as Type M, or magnitude, error, alongside Type S, or sign, error. At 6% power the sign error becomes material: roughly one significant result in five points in the wrong direction.
Button and colleagues (Nature Reviews Neuroscience, 2013) estimated a median statistical power of 21% across 49 meta-analyses covering 730 studies in neuroscience. Under a prior odds of one true hypothesis in ten, 21% power yields a positive predictive value of 0.021/(0.021 + 0.05) ≈ 30% — meaning roughly seven in ten significant findings from such a field would be false, before any consideration of flexible analysis.
This is the mechanism behind replication failure in fields with small samples. A replication powered for the published effect is powered for an inflated effect, so it is underpowered for the real one, and its failure is then read as evidence against the original finding rather than as an artefact of the same design problem twice.
Post hoc power — statistical power computed after a study using the effect observed in that study — is a monotone transformation of the p-value and contains no information the p-value does not. Hoenig and Heisey named the problem precisely in The American Statistician (2001): a non-significant result mathematically must yield low observed power, and a significant one high observed power, so observed power cannot serve as evidence about the result it is computed from.
The circularity is worth stating twice because it is intuitive-sounding and wrong. A power calculation asks how often a design would detect an effect of size δ. Setting δ equal to the effect you measured assumes the truth of the estimate whose reliability is under dispute.
What answers the underlying question is a confidence interval, which states directly which effects the data can and cannot exclude. A design-based sensitivity statement — “with 12 per group this study could detect a difference of 1.2 standard deviations at 80% power” — is legitimate, because it is computed from the design rather than from the result. Reviewer asked for a post hoc power calculation sets out how to decline the request without stonewalling the reviewer.
Power problems are visible in the text without access to the data, and a small number of checks find most of them.
Read the assumed effect against the cited literature. The commonest tell is an assumed effect larger than anything the cited prior studies observed. If the sample size paragraph assumes a 50% relative reduction and the cited pilot reported 18%, the calculation was reverse-engineered from an achievable enrolment. A round assumed effect with no citation — “a 30% difference”, “an effect size of 0.5” — is the same signal in a different form.
Compare the calculation’s n with the analysed n. A calculation stating 120 per arm alongside a results table reporting 94 and 89 means the trial was not delivered as designed, and the discrepancy is frequently unmentioned. Compare both with the registered target enrolment, which is public.
Check the endpoint identity. Power is often calculated on one outcome and the headline claim made on another. A calculation for six-month mortality supporting a discussion built on a biomarker change is a mismatch, not a detail.
Check the unit of analysis. If the calculation is in participants and the model treats repeated measurements as independent observations, the reported precision is not the designed precision. See my n is not independent, why my p-value shrinks with more cells, and the pseudoreplication check.
Check for the missing parameters. A cluster trial with no stated intracluster correlation, a survival trial with no target event count, or a non-inferiority trial with no justified margin each has an incomplete calculation regardless of the number it produces.
Check whether absence of evidence became evidence of absence. A non-significant primary result described in the abstract as “no difference” or “comparable” is a claim the design cannot support unless equivalence was tested against a pre-specified margin. The overclaim check is built for this pass.
Reviewers rarely write “the study is underpowered” without qualification. The phrasings that appear in real reports are more specific, and each points at a different remedy.
“The sample size calculation assumes an effect substantially larger than any reported in the cited literature; please justify the assumption.” “No sample size calculation is reported for the primary endpoint.” “The number analysed differs from the number the calculation specifies, and the discrepancy is not explained.” “Power was calculated for the primary outcome, but the conclusions rest on secondary endpoints for which no calculation is provided.” “The intracluster correlation used is not stated or sourced.” “Subgroup differences are reported without a test for interaction, for which this trial is not powered.” “The authors interpret a non-significant result as evidence of no effect; an equivalence framing with a pre-specified margin would be required to support that claim.” “Please report a post hoc power calculation” — which is itself a mistaken request, though the concern behind it is usually real.
If the comment you received was about size rather than about power arithmetic, reviewer says my sample size is too small separates the three distinct objections that sentence carries.
Report a statistical power calculation as an argument, not as a number. Name the primary endpoint, the assumed effect and its explicit source, the assumed variability or event rate and its source, the Type I error rate including any multiplicity adjustment, the target power, the resulting group size, and the attrition inflation applied. CONSORT 2010 item 7a asks for exactly one thing — “How sample size was determined” — and most reported paragraphs answer only the last clause of it.
Add the sensitivity range ICH E9 section 3.5 asks for: what the required sample size would have been under a smaller assumed effect or a higher variance. Two extra sentences convert a number into a defensible argument.
Where the study is finished and smaller than designed, state the achieved precision rather than recomputing power. The confidence interval on the primary effect, and the smallest effect the interval excludes, tell a reader what the study can support. If the effect is precisely estimated but small, that is a separate conversation about clinical relevance — see reviewer says the effect size is not meaningful.
Statistical power is a frequentist, pre-data construct, and several of its limitations are genuine rather than rhetorical.
Power says nothing about whether the assumed effect is worth detecting. The minimal clinically important difference that feeds the calculation is a judgement, often made by the people with the strongest interest in a small one, and it is rarely elicited from patients. Two trials with identical power can differ entirely in whether their target effect matters.
Power says nothing about whether an adequately powered trial will succeed. Phase III trials are almost all designed at 80% or 90% power, and a large share of them nonetheless return a null primary result. That is not a failure of the calculation. A trial powered at 90% for a hazard ratio of 0.75 delivers exactly what it promised if the true hazard ratio turns out to be 0.95: the design was correct and the assumption was wrong. Power protects you against sampling noise around an effect you have assumed; it offers no protection at all against having assumed the wrong effect, and the second risk is by far the larger one in most programmes.
Power is also not the only basis on which a study can be sized. Precision-based sample size determination targets a confidence interval width rather than a hypothesis test, and suits estimation questions better. Bayesian assurance, introduced by O’Hagan, Stevens and Campbell (Pharmaceutical Statistics, 2005), integrates power over a prior distribution for the effect instead of conditioning on one value, and typically returns a number far below the nominal power — which is informative precisely because it is discouraging. Neither approach is standard in journal reporting, and a manuscript using one should say what it did and why.
Finally, power is defined for a pre-specified analysis. A study with flexible outcome definitions, optional covariates or undeclared interim looks has no single power value at all, because it has no single test.
Pseudoreplication · Post hoc power · Sample size objections · Multiple comparisons · All research methods concepts · Statistical reviewer objections
Statistical power means the probability that a study will produce a significant result if the effect it was designed to detect is genuinely present. Power is determined by sample size, outcome variability, the assumed effect and the Type I error rate, and is fixed by the design before any data are collected.
Power in statistics is 1 − β, the complement of the Type II error rate: the probability of correctly rejecting a false null hypothesis when the true effect equals a specified value. Power is always relative to that specified effect, so a study has many power values, one for each effect considered.
Statistical power is calculated from four inputs: the effect to be detected, the variability of the outcome, the sample size and the Type I error rate. For a two-group comparison of means at 80% power and a two-sided 5% test, the requirement is approximately 15.7/Δ² per group, where Δ is the standardised effect.
Conventional levels are 80% and 90%. ICH E9 section 3.5 states the type II error probability is “conventionally set at 10% to 20%” and advises keeping it as low as feasible for trials difficult to repeat. No theorem selects 80%; it is an inherited convention rather than a statistical requirement.
An underpowered study is one whose design has a low probability of detecting the effect of interest. Underpowering produces false negatives and, less obviously, inflates any effect that does reach significance: at 20% power, every significant estimate overstates the true effect by at least 75%.
No. Statistical power is a probability; sample size is one of four inputs that determine it, alongside outcome variability, the assumed effect and the Type I error rate. Two studies with identical sample sizes can have very different power if their outcomes differ in variability or their target effects differ in size.
Not usefully from the observed effect. Power computed from the result is a monotone transformation of the p-value and adds no information, as Hoenig and Heisey showed in 2001. A design-based statement of what effect the study could have detected is legitimate; a confidence interval on the observed effect is more informative still.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect