Skip to content

SOLUTIONS

Reviewer asked for a post hoc power calculation

Power computed from your observed effect just restates your p-value. Why it answers nothing, what the reviewer actually wants, and what to send instead.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Reviewer asked for a post hoc power calculation

Post hoc power — a power calculation computed from the effect you observed — is circular. It is a strictly decreasing function of your p-value: a non-significant result must return low power, a significant one high power. It tells the reviewer nothing new. What it looks like is evidence. What it is, is your p-value in different units.

You should still answer the comment, because the concern behind it is usually legitimate. It is just not answerable by the calculation requested.

Why observed power says nothing

Observed power is calculated by assuming the true effect equals the one you measured. But your measured effect is itself uncertain, and it is the thing under dispute. Assuming it is correct in order to evaluate whether your study could have detected it takes as given exactly what the reviewer is questioning.

The arithmetic consequence is mechanical: observed power and the p-value are one-to-one, and the mapping does not depend on how many subjects you enrolled. If p is 0.049 your observed power will be near 50%, always, regardless of the science. Reporting that number as reassurance is reporting the p-value twice.

The relationship is not an approximation. For a two-sided test at alpha = 0.05, observed power is Φ(|Z| − 1.96) + Φ(−|Z| − 1.96), where Z is the test statistic that produced your p-value. Sample size, variance and effect size enter only through Z. Two studies with the same p-value have the same observed power whether one enrolled 18 animals and the other 1,800 patients. Hoenig and Heisey set this out in “The Abuse of Power” (The American Statistician 2001;55:19–24), the paper most reviewers are half-remembering when they ask for the calculation and the paper you should cite when you decline it.

The same result has been published repeatedly in clinical and pharmacy literature, which is useful when the reviewer’s field is not statistics. Goodman and Berlin argued in Annals of Internal Medicine (1994;121:200–206) that power is a pre-experimental concept and that confidence intervals are the post-experimental tool; Levine and Ensom titled their Pharmacotherapy paper (2001;21:405–409) “Post hoc power analysis: an idea whose time has passed?”; Senn’s BMJ letter (2002;325:1304) is headed “Power is indeed irrelevant in interpreting completed studies.” Greenland went one step further in Annals of Epidemiology (2012;22:364–368), showing that non-significance combined with high power does not imply support for the null over the alternative — so even the version of the argument that runs in the reviewer’s favour does not work.

Two things follow for the response letter. A reviewer who knows this result will read a supplied observed-power figure as a sign that the objection was not understood, and CONSORT 2010 says in as many words that the calculation has little merit — the quotation is at the end of this page, and it is the citation to lead with.

The arithmetic, with numbers

Observed power at alpha = 0.05, two-sided, maps onto the p-value as follows. Nothing else about the study changes these figures.

Two-sided p-value Observed power at alpha = 0.05 Verdict a naive reading invites
0.001 90.8% “Well powered”
0.01 73.1% “Adequately powered”
0.05 50.0% “Borderline”
0.10 37.6% “Underpowered”
0.20 24.9% “Badly underpowered”
0.50 10.4% “Hopeless”
0.80 5.7% “No power at all”

Read the table as a warning rather than a lookup. The third column contains no information the second does not, and the second contains none the first does not. A referee who asks you to compute the middle column, having already read the left one, is asking for a restatement and will treat it as a finding. That is the specific harm: the number acquires rhetorical weight it has not earned, and once it is in the manuscript an editor may act on it.

Note the floor. Observed power can never fall below alpha, which is why a p-value of 0.80 still returns 5.7% rather than zero. A “power” figure that is arithmetically incapable of dropping below 5% is not measuring anything about your design.

What the reviewer is actually worried about

The request for post hoc power is usually a proxy for one of three substantive concerns, and identifying which one you have received determines the entire response.

The result is null and the reviewer suspects the study could not have seen a real effect. This is a question about precision, and the honest answer is a confidence interval plus a statement of what the design could detect. It is a fair question and you should answer it fully.

The paper reads non-significance as absence of effect. The reviewer wants the claim rewritten, not a calculation. Handle it in the discussion and abstract, and say what the data are compatible with rather than what they failed to show; how to write a limitations section covers the paragraph that usually needs rewriting.

The result is positive on a small sample and the reviewer suspects the effect is inflated. This is the strongest version of the concern and post hoc power actively conceals it, because a significant result mechanically returns high observed power. Statistically significant estimates from low-powered designs are biased upward in magnitude — Gelman and Carlin call this a type M (magnitude) error, alongside type S (sign) errors, in “Beyond Power Calculations” (Perspectives on Psychological Science 2014;9:641–651). The remedy is a design analysis using a plausible external effect size, not a calculation fed by your own estimate. Where the concern is really about whether the finding holds up at all, reviewer says the experiment was not replicated is the adjacent comment.

What to give instead

A confidence interval on the effect. This is the direct answer to the question the reviewer is actually asking, which is what effect sizes your study can rule out. A 95% interval from a 2% decrease to a 30% decrease tells the reader the data are compatible with a trivial effect and with a large one — which is honest, informative, and impossible to extract from a power figure. Report it on the scale readers care about, and pair it with a minimal clinically important difference so the interval can be read against something. The relevant question is whether the interval excludes an effect worth having, which is a property of the interval and not of the p-value; what is a confidence interval covers the interpretation reviewers most often want corrected.

An equivalence or non-inferiority framing, if your claim is that there is no difference. A non-significant p-value is not evidence of no effect. If you want to claim equivalence, test for it against a pre-specified margin and say what margin you chose and why. The standard procedure is the two one-sided tests of Schuirmann (Journal of Pharmacokinetics and Biopharmaceutics 1987;15:657–680), operationally equivalent to checking whether the 90% confidence interval lies entirely inside the margin at alpha = 0.05. Declaring the margin after seeing the interval is the failure mode reviewers catch; what is non-inferiority sets out how margins are justified.

A prospective calculation, if you did one. Report the assumed effect, its source, alpha, power and the resulting group size. A study powered for a pre-specified effect is defensible even when the observed effect turns out smaller — that is what prospective calculation is for. Give the source of the assumed effect explicitly: a named prior study, a pilot, a regulatory margin, or a smallest effect of interest agreed in advance. “Based on previous work” without a citation invites the follow-up question.

A precision statement. “With n=12 per group, this study could detect a difference of 1.2 standard deviations with 80% power” is a legitimate sensitivity statement, because it is computed from the design rather than from the result. Frame it as what the design could detect, not as how powerful the study turned out to be. Lenth’s “Some Practical Guidelines for Effective Sample Size Determination” (The American Statistician 2001;55:187–193) is the standard reference for stating the detectable effect rather than retrospective power, and gives you something to cite that is not a rebuttal.

A design analysis, when the finding is positive and the sample is small. Take an effect size you would consider plausible from outside your own data, and report the probability that an estimate reaching significance at your sample size would exaggerate the true effect, and the probability it would have the wrong sign. This answers the inflation concern directly and turns a defensive paragraph into an analysis.

A worked example a reviewer can check

Consider a two-arm trial with 40 participants per group, a continuous outcome, an observed mean difference of 1.8 points and a pooled standard deviation of 5.0 points. The standard error is 1.12, t = 1.61 on 78 degrees of freedom, and p = 0.11. The reviewer writes: “The study appears underpowered; please provide a post hoc power calculation.”

Observed power for that result is 36%. It is 36% because p is 0.11. Had the same trial enrolled 400 per group and returned p = 0.11 for a difference of 0.57 points, observed power would still be 36%. The number describes the p-value, not the trial.

Now the informative version. The 95% confidence interval is −0.43 to 4.03 points. If the minimal clinically important difference on this scale is 3.0 points, the data are compatible with a 4-point benefit and with a small harm: the trial has not excluded a clinically important effect, so it is inconclusive rather than negative, and the discussion must say so. The design-based statement is that with 40 per group and a standard deviation of 5.0 points, the trial had 80% power to detect a difference of 3.2 points — marginally larger than the effect that would matter, which is a real and reportable design limitation.

Equivalence fails too, and it is worth saying why in the response. The 90% confidence interval runs −0.06 to 3.66 points, which does not lie inside a ±3.0-point margin, so the two one-sided tests do not support equivalence at that margin. Three sentences of arithmetic have told the reviewer everything the requested calculation could not: what the trial can exclude, what it cannot, and which of those is the problem.

How to word the response

Be direct and unapologetic, and cite something. A workable shape:

We have not included a post hoc power calculation. Because such a calculation uses the observed effect as the assumed true effect, it is a deterministic function of the reported p-value and provides no additional information. In its place we now report the 95% confidence interval for each primary comparison (Table 2), which conveys the range of effects compatible with our data, including the smallest effect the study can exclude.

Note what that paragraph does. It declines the calculation, gives the reason in one clause, and puts a replacement analysis on the table in the same breath — which is a defensible statistical position rather than a refusal, and is the form an editor can act on without adjudicating a methods dispute.

If you ran a prospective calculation, lead with it instead, because it converts the exchange from a disagreement into a correction:

The sample size was fixed in advance: 40 per group provides 80% power at alpha = 0.05 (two-sided) to detect a 3.0-point difference, assuming a standard deviation of 5.0 points taken from Smith et al. (2021). We have added this to the Methods (page 6). We have not added a retrospective calculation based on the observed difference, since for a completed study that quantity is a function of the p-value alone (Hoenig and Heisey, 2001); the corresponding confidence interval is now reported in Table 2.

Three phrasings to avoid. Do not write “the study was underpowered to detect the observed effect”, which is the circularity restated as an admission. Do not write “achieved power was 36%, indicating that a larger sample is needed”, which invites a rejection you did not have to accept. Do not write “no significant difference was found, indicating the treatments are equivalent” anywhere in a paper where the reviewer has already raised power — that sentence is what produced the comment. General structure for the letter is in how to write a response to reviewers.

If the reviewer insists

A reviewer who repeats the request after a clear explanation is usually asking to be shown that the concern was taken seriously, not to be persuaded on the statistics. Three options, in order of preference.

Provide the confidence interval prominently and add the requested calculation in supplementary material with a one-line note on its interpretation. This satisfies the letter of the request without giving the number rhetorical weight.

Offer the substitute the reviewer will recognise as responsive: a sensitivity analysis reporting the effect detectable at 80% power given the realised sample size and observed variance, labelled as such. A reviewer asking for “power” generally wants a statement about what the study could see, and this is that statement computed legitimately. Naming it “sensitivity analysis” or “detectable effect size” rather than “post hoc power” matters, because the labels carry different methodological commitments.

Or ask the editor to adjudicate. If your response letter has already explained the circularity clearly, editors will usually side with the statistics. Where the journal’s own submission template contains a field for post hoc power, fill it and add the interpretive footnote rather than leaving it blank — a blank field routes to a technical check, and a completed field with a caveat routes to the editor.

When a post hoc calculation is legitimate

A calculation performed after data collection is legitimate whenever the effect size entering it comes from outside the data being analysed. That single test separates the defensible cases from the circular one.

Detectable effect size at the realised sample size. Using the achieved n and the observed variance, report the smallest effect the design could detect at 80% power. Variance is a nuisance parameter rather than the quantity in dispute, so reusing it is not circular in the way reusing the effect is. State it as a design property.

Power at a pre-specified effect, computed with the realised n. If enrolment fell short of target, recomputing power for the originally specified effect at the sample actually obtained is a valid and often requested amendment. Report both figures and the shortfall.

Power for a future study. Using this study’s variance estimate to size the next experiment is exactly what pilot data are for. Say that the effect estimate is not being carried forward, and use a smallest effect of interest instead; carrying a small study’s point estimate into the next calculation is how underpowered literatures perpetuate themselves.

Design analysis under an assumed effect. Type S and type M calculations require you to name a plausible effect from prior literature or theory, which is what makes them informative about exaggeration.

The one calculation that fails the test is the one the reviewer asked for: power evaluated at your own observed effect. Everything above changes the input; nothing above changes the timing.

How to catch the problem in your own draft

Search the manuscript for the phrases that mark it before a reviewer does. “Observed power”, “achieved power”, “post hoc power”, “retrospective power” and “the study was underpowered to detect the observed difference” all indicate a circular calculation already in the text — often added in response to an earlier reviewer at a different journal.

Then search for the claim that invites the comment. “No significant difference”, “comparable between groups”, “similar in both arms” and “there was no effect of” are the standard triggers when they appear without an interval. A null result described in words but not bounded by numbers reads to a statistical reviewer as a claim the data have not established, and the power comment is the usual way that objection arrives; the underlying issue is often reviewer says the effect size is not meaningful rather than power at all.

Check finally that your denominator is the one you think it is. Power arguments collapse when the analysed n includes technical replicates, repeated measures on the same subject or multiple cells from one animal, because the independent units are fewer than the count in the table — see what is pseudoreplication. A power calculation on the wrong unit is wrong whichever direction it points.

Where the reporting guidance stands

CONSORT 2010 addresses this directly. The Explanation and Elaboration document (Moher et al., BMJ 2010;340:c869), under item 7a on sample size, states: “There is little merit in a post hoc calculation of statistical power using the results of a trial; the power is then appropriately indicated by confidence intervals (see item 17).” That sentence is quotable in a response letter and carries more weight with an editor than a methodological argument in your own words; the CONSORT checklist covers what item 7a asks you to report instead.

ARRIVE 2.0 takes the same line for animal work by asking only for the prospective version. Item 2b reads: “Explain how the sample size was decided. Provide details of any a priori sample size calculation, if done.” The words are “a priori”; nothing in the ARRIVE guidelines asks for a retrospective figure.

State the limit of this accurately rather than overstating it. There is no cross-journal prohibition on post hoc power, individual journals differ in whether their statistical guidance mentions it at all, and reviewers who request it are usually following custom rather than a stated policy. What exists is a consistent methodological literature spanning 1994 to 2014 — Goodman and Berlin, Hoenig and Heisey, Lenth, Levine and Ensom, Senn, Greenland, Gelman and Carlin — and one explicit sentence in CONSORT. That is enough to decline on, and it is more than most reviewers expect you to have.

Before you resubmit

If the underlying concern was that your study is small, reviewer says my sample size is too small covers the three things that comment can mean. If the concern was that non-significance was reported as absence of effect, check the discussion — that is a claim problem rather than a power problem, and the overclaim check is the pass for it. If the comment arrived from a trials submission alongside protocol and registration queries, clinical trial review covers the cluster it usually travels in.

More objections in this family: responding to statistical reviewer comments.

Related

What is statistical power · What is a confidence interval · What is effect size · Reviewer says the sample size is too small · Statistical review capacity

PerfectPaper flags a power figure computed from the observed effect, checks whether each null claim in the discussion is supported by an interval that excludes a meaningful effect, and reports the smallest effect the design could detect at the sample size actually analysed.

Review my manuscript

Frequently asked questions

What is post hoc power?

Statistical power calculated after a study using the effect size observed in that study, rather than an effect size specified in advance. Because it assumes the observed effect is the true one, it is determined entirely by the test statistic that produced your p-value, and carries no information independent of it.

Why is post hoc power considered inappropriate?

Because it is circular. It assumes the truth of the estimate whose reliability is in question, and it maps one-to-one onto the p-value whatever the sample size. A non-significant result mathematically must yield low observed power, which is why it cannot serve as evidence about that result. Hoenig and Heisey set out the argument in The American Statistician (2001;55:19–24).

How do I respond to a reviewer who asks for a post hoc power calculation?

Decline the specific calculation, explain in two sentences that it is a deterministic function of the reported p-value, and supply the confidence interval for each primary comparison in its place. Cite CONSORT 2010 item 7a, which states there is little merit in a post hoc power calculation and that power is then appropriately indicated by confidence intervals.

What should I report instead of post hoc power?

A confidence interval on the effect size, which shows the range of effects your data are compatible with. If claiming no difference, use an equivalence test against a pre-specified margin. If describing the design, state what effect the study was capable of detecting rather than how powerful it proved to be.

Is a prospective power calculation different?

Entirely. A prospective calculation uses an effect size specified before the data exist, usually from prior literature or a minimum effect of interest, and it legitimately justifies a sample size. The problem is only with calculations that take the observed effect as input.

Does low observed power mean my study was underpowered?

Not on its own, because observed power at alpha = 0.05 is 50% whenever p equals 0.05 and falls as p rises, whatever the sample size. Whether the study was underpowered is answered by the width of the confidence interval relative to the smallest effect worth detecting, not by a power figure derived from the result.

Can I decline to run the post hoc power calculation at all?

Declining a specific analysis is acceptable when you explain why and offer an alternative that addresses the underlying concern. What editors do not accept is silence, or a changed manuscript with no explanation. A refusal with a reason and a substitute analysis is ordinary scientific disagreement, and for this particular calculation the reason is short enough to state in two sentences.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.