Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
The comment usually means one of three different things. Work out which before you respond, because the answer to each is different and only one needs new data.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
“The sample size is inadequate” is three different objections wearing one sentence: the study was underpowered for the effect you claim, the n you reported counts the wrong thing, or the reviewer does not believe a result this large from a group this small. Only the first needs more data. Work out which one you received before you write a word of response.
Three signals separate the three objections: whether the reviewer named a number, what the rest of the report asks for, and which verb the comment uses. Diagnose from all three before drafting, because the three answers share no material.
Read the rest of the review. The other comments usually disambiguate it. A reviewer who also asks about variance, confidence intervals, or replication means power. A reviewer who asks what n refers to, or how many animals contributed, means the unit of analysis. A reviewer who calls the effect surprising, or asks for independent validation, does not think your statistics are wrong — they think the result might not be real.
Check whether they gave a number. A reviewer who says “n=3 is insufficient for this claim” is making a power argument. One who says “the sample size is inadequate” with no number is more often expressing doubt about the finding.
Check where the comment sits. A sample-size objection filed under major points, or repeated in the summary paragraph, is a condition of acceptance. The same words in a numbered minor comment usually ask for a methods sentence, not a new experiment. Editors read the summary paragraph; readers of your response letter read it against that paragraph, which is why the response to reviewers has to answer the version stated at the top.
| What the reviewer wrote | Which objection it is | What answers it |
|---|---|---|
| “n=3 is insufficient to support a claim of this magnitude” | Power | The prospective calculation, or a confidence interval plus a minimum detectable effect |
| “The sample size calculation does not state the assumed effect or its source” | Power | The six inputs restated in the methods |
| “The study appears underpowered given the variability in Figure 2” | Power | The standard deviation used, and the interval on the primary comparison |
| “What does n refer to here?” | Unit of analysis | A methods sentence naming the experimental unit |
| “The authors treat each cell as an independent observation” | Unit of analysis | Reanalysis at the animal level, or a random effect for animal |
| “n differs between the methods and the figure legends” | Unit of analysis | Reconciled denominators at every stage |
| “The sample size is inadequate” with no number attached | Belief | Orthogonal evidence, or a narrower claim |
| “Please validate this finding by an independent method” | Belief | A second method, a dose response, or a rescue experiment |
A power objection is answered by the prospective calculation if you performed one, and by a confidence interval plus a minimum detectable effect if you did not. The honest response depends on which of those two you are in.
If you did, report it: the effect size assumed, where that assumption came from, the alpha, the power, and the resulting group size. Reviewers accept a study powered for a pre-specified effect even when the observed effect is smaller. That acceptance is the entire point of specifying the effect in advance — the calculation is a statement about the design, and the design did not change when the data came in.
If you did not, do not compute power after the fact from your observed effect. That calculation is circular — it re-expresses your p-value as a power figure and adds nothing. Hoenig and Heisey showed the arithmetic in The American Statistician (2001;55:19–24): observed power is a deterministic function of the p-value, so a non-significant result must return observed power below 50% whatever the science, and a significant one must return at least 50%. Give a confidence interval instead. A 95% interval on your effect tells the reviewer exactly what your data can and cannot exclude, which is what they were actually asking. This is covered in full on reviewer asked for a post hoc power calculation.
The objection is often correct on the merits. Button and colleagues (Nature Reviews Neuroscience, 2013) estimated a median statistical power of 21% across 49 meta-analyses covering 730 studies in neuroscience. A reviewer working in a field with that distribution is applying a well-calibrated prior, not being obstructive, and a response that treats the comment as hostile reads badly against it. See what is statistical power for the arithmetic that produces those numbers.
A defensible sample size justification names six things: the primary outcome, the assumed effect, the source of that assumption, the assumed variability or event rate, the alpha and target power, and the inflation applied for expected attrition. A paragraph that gives only the resulting number answers none of them.
Reporting guidelines ask for exactly this and no more. CONSORT 2010 item 7a asks “How sample size was determined”. ARRIVE 2.0 item 2b asks how the sample size was decided, with details of any a priori calculation “if done” — the phrase permits an honest statement that no calculation was performed. STROBE item 10 asks you to explain how the study size was arrived at, and does not require a power calculation at all; “all eligible records in the registry between 2015 and 2022” is a complete and honest answer to it. See the ARRIVE guidelines and the STROBE checklist for the surrounding items.
A worked version, for a two-group comparison of a continuous outcome: the primary outcome is tumour volume at day 21; the smallest difference worth detecting is 0.5 standard deviations, taken from the between-group difference reported in the pilot cohort described in the methods; alpha is 0.05 two-sided; target power is 80%. Lehr’s rule of 16 gives n per group ≈ 16/d² = 16/0.25 = 64. Inflating for 15% expected attrition gives 64/0.85 = 76 per group. Every number in that paragraph is auditable, and a reviewer who disagrees can disagree with a specific input rather than with the conclusion.
Where no prospective calculation exists, the legitimate substitute is a sensitivity power analysis: the minimum effect the design could have detected, computed from the sample size and alpha rather than from the result. “With n=12 per group, this study could detect a difference of 1.2 standard deviations with 80% power” is a statement about the design and is defensible. “Observed power was 43%” is a statement about the result and is not. Frame it as what the design could detect, never as how powerful the study turned out to be.
If the claim being defended is that two conditions do not differ, no sample size argument rescues it, because a non-significant p-value is not evidence of no effect. Test for equivalence against a pre-specified margin — two one-sided tests, or a formal non-inferiority design — and state the margin and its justification. Where the outcome is clinical, that margin should be anchored to a minimal clinically important difference rather than chosen for convenience.
The unit-of-analysis version is the most consequential of the three and the one most often misread as a power complaint. A reviewer who spots it is not asking for more data — they are asking you to reanalyse what you have at the correct level.
If you measured 150 cells from three mice and reported n=150, your sample size is three. The cells are informative about each animal’s value but they are not independent observations, and the reported p-value assumes an independence the data do not have.
The size of the error is governed by the intraclass correlation coefficient, ICC or rho, the share of total variance sitting between animals rather than between cells within an animal. Kish’s design effect converts it into a penalty: DEFF = 1 + (m − 1)rho, where m is the number of measurements per unit. At 50 cells per animal and an ICC of 0.2, which is modest for a phenotype that differs between animals at all, DEFF is 10.8 and the naive standard error is understated by a factor of 3.3. That is why a finding at p < 0.001 on the inflated n can be p > 0.2 at the correct one, and why a p-value that shrinks every time you add cells is a symptom rather than a success.
Journals increasingly define n rather than leaving it to the author. The British Journal of Pharmacology states the rule plainly in its guidance for authors and peer reviewers (Curtis et al., 2018;175:987–993): group size is the number of independent values, so technical replicates of one sample are not separate observations and must not be counted as n. ARRIVE 2.0 item 1b asks authors to name the experimental unit explicitly and offers “a single animal, litter, or cage of animals” as its examples. Quoting the journal’s own definition in your response is more persuasive than arguing the general principle.
The remedy is usually to summarise to a per-animal value and test on those, or to fit a model with a random effect for animal. Both are reanalyses, not new experiments. Report which package and which random effect structure you used, and expect the p-value to move by orders of magnitude rather than at the margin. Reviewer says my n is not independent covers the full pattern, and the pseudoreplication agent finds it before submission.
There is a quieter variant worth checking before you conclude the reviewer means power: the n that does not reconcile. Methods say 24 animals, Table 1 sums to 22, and Figure 3’s legend reports n=8 per group across three groups. A reviewer reading three different numbers writes “the sample size is unclear”, which looks like a power comment and is a counting comment. Reconcile the denominators at every stage — enrolled, treated, analysed, plotted — and state the reason for each exclusion in the same sentence as the number.
Sometimes the statistics are fine and the reviewer doubts the result. The tells are requests for independent validation, for an orthogonal method, or for a replication in a second model.
The doubt is frequently well-founded rather than temperamental. At 20% power for a two-sided 5% test, a result is significant only if the estimate exceeds 1.96 standard errors from zero, while the true effect sits at about 1.12 standard errors, so every publishable estimate from that design overstates the truth by at least 75%. Gelman and Carlin (Perspectives on Psychological Science, 2014) formalised this as Type M, or magnitude, error, alongside Type S, or sign, error. A reviewer who says “this effect is implausibly large for n=4” is describing Type M error in ordinary language, and the correct answer is not a defence of the p-value. See what is effect size for how to report the magnitude in a way that survives the objection.
More data will not fix this and neither will more statistics. What helps is evidence from a different direction: the same conclusion by another method, a dose response, a rescue experiment, or a mechanism. An orthogonal result is one whose failure modes do not overlap with the first assay’s — quantitative PCR and RNA sequencing share an RNA extraction and a reverse-transcription step and are less orthogonal than they look, whereas a genetic and a pharmacological perturbation converging on the same phenotype are genuinely independent lines. Reviewer wants orthogonal validation covers the most common shape of that request.
If you have none of those, narrowing the claim is a legitimate response — a smaller claim the data clearly support is more publishable than a large one under doubt. Moving from “X drives Y” to “X is associated with Y in this model, with the magnitude estimated imprecisely” concedes nothing you had evidence for.
The sentence “the sample size is too small to support these conclusions” can arrive on three manuscripts in the same week and demand three unrelated revisions. The three composites below show why the diagnosis, not the sentence, determines the work.
A randomised trial of 42 patients reporting a 9-percentage-point absolute reduction in the primary endpoint, with a 95% interval from −2 to +20. The rest of the report asked for the sample size calculation and for intervals on the secondary endpoints. That is a power objection. The revision restated the prospective calculation — assumed a 30-percentage-point difference from a named prior trial, alpha 0.05 two-sided, 80% power, 33 per arm, inflated to 39 for dropout, with recruitment closing at 42 of the planned 78 — reported the interval for every endpoint, and rewrote the abstract to say the trial could not distinguish a trivial reduction from a large one. No new patients.
A preclinical paper reporting n=180 neurons from six mice, with a treatment effect at p < 0.0001. The same reviewer asked, three comments later, how many animals contributed. That is a unit-of-analysis objection wearing power’s clothes. The revision summarised to six per-animal means, fitted a linear mixed model with a random intercept for animal as a confirmatory analysis, reported the new p-value of 0.03, and replaced the 180-point scatter with per-cell points overlaid by per-animal means. No new mice, and the conclusion survived at a smaller stated magnitude.
A three-replicate biochemistry paper claiming a sixfold change with tight error bars and no other statistical comments in the report, alongside a request for validation in a second cell line. That is disbelief. The revision added the second line, which showed a 2.4-fold change, and the claim moved from “sixfold” to “reproducible across two lines, with magnitude varying between them”. The reviewer’s objection was correct and the paper improved.
Reviewers rarely write “the study is underpowered” as their first sentence on this. They write the specific version, and these are the phrasings that decide papers.
“The sample size calculation does not state the assumed effect or its source.” “No justification is provided for the study size; please explain how the sample size was arrived at, even if no formal calculation was performed.” “Please clarify the experimental unit.” “The authors treat each cell as an independent observation.” “n differs between the methods section and the figure legends.” “With three biological replicates the authors cannot support a claim of no difference between conditions.” “The effect size reported is implausibly large given the sample; please discuss the possibility that it is inflated.” “A post hoc power calculation has been added; this is a function of the reported p-value and does not address the concern.” “The limitations acknowledge a small sample but the abstract retains the original claim.”
The last two are the ones authors create during revision rather than receive at first review, and both are avoidable. Adding observed power to a response letter converts a survivable comment into an escalation, and leaving the abstract untouched converts a statistical fix into evidence that you were not listening.
Answer in the register of the diagnosis: a design statement for power, a reanalysis statement for the unit of analysis, and a claim statement for disbelief. Each of these is short on purpose.
For power, when a prospective calculation exists:
The sample size was fixed prospectively. We assumed a 12-point difference in the primary endpoint, taken from [source], with alpha 0.05 two-sided and 80% power, giving 40 participants per arm, inflated to 46 for anticipated dropout. We have added these inputs to the Methods and now report 95% confidence intervals for all endpoints (Table 2), which state the range of effects the trial can and cannot exclude.
For the unit of analysis:
We agree that the experimental unit is the animal rather than the cell. We have reanalysed the data at the animal level, summarising the 30 cells measured per animal to a single value and testing on the six resulting values, and confirmed the result with a linear mixed model including a random intercept for animal. The revised p-value is 0.03; Figure 3 now shows individual cells with per-animal means overlaid, and the Methods state the experimental unit explicitly.
For disbelief:
The reviewer is right that three replicates cannot establish the magnitude of this effect. We have added an independent line of evidence [name it] and have narrowed the claim in the abstract and discussion from a quantitative statement of magnitude to a statement of direction that both assays support.
Adding samples is the cleanest answer when the objection is genuinely about power and the material is obtainable — and the disclosure obligation is what makes it defensible. Say clearly that the additional samples were collected after the initial analysis, give the new total, and report whether the conclusion changed.
Undisclosed data addition after seeing a result is a serious problem, because testing repeatedly as data accumulate and stopping when p crosses 0.05 raises the false-positive rate above the nominal 5%. Simmons, Nelson and Simonsohn’s 21-word disclosure — “We report how we determined our sample size, all data exclusions (if any), all manipulations, and all measures in the study” — exists to make that history visible in a single sentence. If the peeking was planned, name the design: a group-sequential trial with a pre-specified alpha-spending function, such as an O’Brien–Fleming or Pocock boundary, controls the error rate across interim looks and should be described as what it is. What is p-hacking covers the failure mode that undisclosed addition belongs to.
Say so directly, and say why: the model is rare, the cohort is closed, the funding period ended. A named constraint is a fact a reviewer can weigh; “further work is needed” is not. Then do three things. Report the effect size with its interval rather than only the p-value. State explicitly what the study can and cannot support. And limit the claims in the abstract and discussion to match.
Write the limitation as an analysis rather than an apology. “The 95% CI ranges from a 4% absolute reduction to a 9% absolute increase, so this study cannot distinguish benefit from harm at magnitudes that would change practice” names the boundary; “the sample size was relatively small, so results should be interpreted with caution” names nothing. See how to write a limitations section for the ordering rule — sample size is a weak lead limitation precisely because the reader can already see the precision in your interval.
Reviewers are generally more receptive to an acknowledged limitation than to a defended one. The revision that fails is the one that argues the sample was adequate while leaving the original claim untouched.
Check that the claims moved with the analysis. A revision that fixes the statistics and leaves the abstract asserting the original result draws the same objection from a reviewer who now believes you were not listening — the overclaim check is built for exactly that pass.
Five checks, in order. Does every reported n name its unit? Do the denominators reconcile across methods, tables and figure legends? Does the abstract’s strongest sentence survive the revised analysis? Is every interval reported alongside the effect it belongs to, rather than a p-value alone? And has any observed power figure been removed from the manuscript and the response letter? Preclinical manuscripts should also check the in vivo rigour items, where the experimental unit, randomisation and blinding are judged together.
See responding to statistical reviewer comments for the other objections in this family.
Post hoc power · n is not independent · Statistical power · Confidence intervals · Effect size not meaningful
PerfectPaper separates the three sample-size objections before you draft a response: it reads the reported n against the level the treatment was applied at, reconciles the denominators across methods, tables and figure legends, and flags every claim in the abstract that the revised analysis no longer supports.
Diagnose the comment first. If it is about power, give the prospective calculation or a confidence interval and a minimum detectable effect. If it is about the unit of analysis, reanalyse at the correct level. If it is about disbelief, add an orthogonal line of evidence or narrow the claim. Only the first of the three can require new data.
Whether n=3 is too small depends entirely on the claim. Three biological replicates support a large, consistent effect and are conventional in much of molecular biology. They do not support a small effect, an interaction, or a claim of no difference. The number is not the problem; the match between the number and the claim is.
Provide a confidence interval on your observed effect and explain that a power calculation computed from observed data is circular. A design-based sensitivity statement — the smallest effect the study could have detected at 80% power — is also legitimate. Reviewers generally accept this, because the interval answers what they wanted to know: what effect sizes your study can rule out.
If you can, it is often the cleanest answer — but say clearly that the additional samples were collected after the initial analysis, and report whether the conclusion changed. Undisclosed data addition after seeing a result is a serious problem; disclosed addition is ordinary practice.
Because the reviewer is counting independent units and you are counting measurements. If 300 cells came from six animals, your sample size is six, and the extra observations inflate significance without adding evidence. Answer that version by reanalysing at the animal level, not by collecting more cells.
Name the constraint — closed cohort, rare model, ended funding period — then report the effect with its interval, state what the study can and cannot support, and cut the abstract and discussion back to match. An acknowledged limitation with a stated boundary is far more survivable than a defence of the original claim.
Say so explicitly, in the methods and the abstract. An exploratory study reporting effect sizes and intervals for hypothesis generation is a legitimate contribution. An exploratory study written as though it were confirmatory is the one that draws this objection.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect