Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
What reviewers mean by the twelve most common statistical objections, what they want to see, and how to answer when you can fix it and when you cannot.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Most statistical reviewer comments are shorter than the problem they point at. “The sample size is inadequate” can mean the study is underpowered, that the reported n counts the wrong unit, or that the reviewer does not believe the effect. Each reading calls for a different response, and answering the wrong one is how a revision fails twice.
This cluster covers the twelve statistical objections that recur most, what each usually means, what evidence resolves it, and how to respond when you cannot run another experiment.
Six further statistical objections arrive often enough to belong in the same list, and each is answerable by reanalysis, by reporting, or by narrowing a claim rather than by new data.
“The statistical methods are not described in sufficient detail.” The missing item is almost always specific: the software version, the random-effect structure of a mixed model, how ties or missing values were handled, or which of several reported analyses was the primary one. Write against the SAMPL guidelines (Lang and Altman, International Journal of Nursing Studies 2015), which enumerate what a statistical methods paragraph must name.
“Missing data are not addressed.” A complete-case analysis is a decision, not a default, and reviewers increasingly ask you to say so. Report how many observations were lost from each variable, whether loss differed by arm, and what assumption your handling makes. Where randomised participants were dropped for non-adherence, the objection is really about intention to treat and the fix is an analysis set, not a sentence.
“A significant effect in one group and a non-significant effect in the other is not an interaction.” Gelman and Stern named this in The American Statistician (2006;60:328–331): the difference between significant and not significant is not itself statistically significant. Nieuwenhuis, Forstmann and Wagenmakers reviewed 513 articles in Nature, Science, Nature Neuroscience, Neuron and The Journal of Neuroscience (Nature Neuroscience 2011;14:1105–1107); of the 157 articles where the error could occur, 78 tested the interaction correctly and 79 did not. In a separate set of 120 cellular and molecular papers they found no study that used the correct procedure. The remedy is one test of the interaction term.
“The comparison is within arms rather than between them.” Bland and Altman set out why testing baseline against follow-up separately in each randomised group, then contrasting the two verdicts, is misleading (Trials 2011;12:264). The comparison the design supports is between groups, adjusted for baseline.
“The reported statistics do not reconcile.” Nuijten and colleagues ran statcheck over more than 250,000 p-values in eight psychology journals from 1985 to 2013 and found that half of papers using null-hypothesis testing contained at least one p-value inconsistent with its own test statistic and degrees of freedom, and one in eight contained a gross inconsistency capable of changing the conclusion (Behavior Research Methods 2016;48:1205–1226). Recompute every p-value in the manuscript from the statistic and degrees of freedom you report before you resubmit.
“The analysis was not pre-specified.” The reviewer is asking whether the model, the outcome and the subgroup were chosen before or after the data were seen. Say which, per analysis. Labelling exploratory work as exploratory is a complete answer; presenting it as confirmatory is what a reviewer means by p-hacking.
The ICMJE Recommendations set the reporting floor most journals adopt, and the statistics clause is short enough to read in full. It states: “Describe statistical methods with enough detail to enable a knowledgeable reader with access to the original data to judge its appropriateness for the study and to verify the reported results. When possible, quantify findings and present them with appropriate indicators of measurement error or uncertainty (such as confidence intervals). Avoid relying solely on statistical hypothesis testing, such as P values, which fail to convey important information about effect size and precision of estimates. References for the design of the study and statistical methods should be to standard works when possible (with pages stated). Define statistical terms, abbreviations, and most symbols. Specify the statistical software package(s) and versions used. Distinguish prespecified from exploratory analyses, including subgroup analyses.”
Three obligations in that clause generate most statistical review comments. Verifiability requires enough method that a reader could rerun the analysis. Quantification requires an effect size with a confidence interval rather than a p-value alone. The prespecified-versus-exploratory distinction requires a per-analysis label that most manuscripts never supply. A response letter that quotes the clause and shows where the revision now satisfies it is harder to argue with than one that promises to have been careful.
Sort every statistical comment into one of four kinds before planning the revision, because the four differ in what they require and in how long they take. The sorting takes an hour and saves a round.
Reporting gaps. The analysis was right and the description was incomplete. Tells: the reviewer asks what something means, which version, which package, what n refers to. The fix is text, and it belongs in the next submission in full.
Analysis gaps. The data are adequate and the model applied to them is not. Tells: a named alternative method, a request for correction, a question about clustering or variance. The fix is reanalysis of existing data, and the honest report says whether the conclusion changed.
Design gaps. The experiment as performed cannot support the claim. Tells: requests for an independent replication, a different comparator, or a randomised comparison. The fix is new data or a narrower claim, and pretending otherwise reads as evasion. These sit closer to methods and design objections than to statistics.
Claim gaps. The numbers are fine and the sentences describing them are not. Tells: “the authors overstate”, “causal language is used for an observational design”, “the abstract does not reflect Table 3”. The fix is the abstract and discussion, and it is the cheapest of the four to make and the most common to skip. Causal wording specifically is covered by correlative, not causal.
Six reanalyses resolve the majority of statistical objections without a single additional sample, and naming them in the response letter is more persuasive than arguing that the original analysis was acceptable.
Reanalyse at the correct independent unit — per animal, per donor, per independently derived culture — or fit a random effect for it. Report the effect size with a 95% interval alongside every p-value, which answers precision questions that a power figure cannot. Apply, or explicitly decline with a named family, the correction the reviewer wants, and put the corrected table in supplementary material either way; FDR versus family-wise error sets out which target fits which claim. Refit under the alternative test or distributional model and state whether the conclusion holds. Run the analysis under a second plausible handling of missing data as a sensitivity analysis. And narrow the claim in the abstract and discussion to what survives all of the above.
The one objection that resists this list is pseudoreplication, and only in the sense that reanalysis at the true unit sometimes removes the result. That is still the correct answer; the alternative is publishing an inference the design does not support. Hurlbert’s survey of 176 experimental papers (Ecological Monographs 1984;54:187–211) found pseudoreplication in 27% overall and in 48% of those applying inferential statistics, so a reviewer raising it is working from a well-founded prior rather than pedantry.
Take the comment the opening paragraph names — “the sample size is inadequate” — on a manuscript reporting a 34% reduction in a marker, p = 0.008, with figure legends stating n = 120.
Read as power. The reviewer doubts the study could detect the effect claimed. The response reports the prospective calculation if one exists — assumed effect, its source, alpha, power, resulting group size — and otherwise gives the 95% interval, which here might run from a 9% to a 52% reduction. That interval is the honest statement of what the data exclude, and it is what statistical power questions are usually reaching for.
Read as unit of analysis. The n = 120 is 120 cells from four animals. The true n is 4. The response reanalyses per animal, reports the wider interval, and says plainly whether significance survives. No new animals are involved, and the reported n falls by a factor of 30.
Read as disbelief. The other comments ask for orthogonal validation or a second model. Neither more data nor more statistics addresses this; evidence from a different direction does, and where none exists, a narrower claim is the legitimate response.
Three responses, one sentence of review. Sending the power answer to a unit-of-analysis reviewer produces a second round in which the reviewer now believes you did not understand the objection, which is a harder position than the one you started in.
Decline a specific analysis when it is wrong, name the reason, and supply the analysis that answers the underlying concern. Three requests recur in that category.
A post hoc power calculation computed from the observed effect is circular by construction — Hoenig and Heisey showed it is a deterministic function of the p-value (The American Statistician 2001;55:19–24), so a non-significant result must return low observed power whatever the science. Offer the confidence interval instead.
A Bonferroni correction applied across every test in the paper, including unrelated hypotheses, controls nothing meaningful and inflates type II error. Name the family each correction was applied over and explain why these tests are not one family.
A formal normality test as the basis for choosing a test is a weak instrument: underpowered at the sample sizes where the assumption matters and oversensitive where it does not. Supply a QQ plot and the sample size instead.
In all three, offer the requested output in supplementary material with a one-line note on its interpretation. That satisfies the letter of the request without giving a misleading number rhetorical weight, and it removes the editor’s reason to treat the exchange as a refusal.
Answer each statistical comment in a fixed four-part shape: restate the concern in your own words, state what you did, give the number, and say where in the revised manuscript it appears. A workable instance reads: “The reviewer is correct that the analysis treated individual cells as independent. We have reanalysed at the animal level (n = 4 per group); the reduction is 31% (95% CI 4% to 50%, p = 0.03) and Figure 2B, the abstract and the discussion have been amended accordingly.”
Three details decide how that letter lands. Give the number in the letter, not only a pointer to a line in the manuscript. Say explicitly when a reanalysis changed the result, because a reviewer who discovers it themselves reads it as concealment. And where you disagree, disagree in a sentence with a citation rather than a paragraph of justification. How to write a response to reviewers covers the letter as a whole.
Naming a statistician who performed the reanalysis, in the letter and in the author list or acknowledgements, changes how the rest of the response is read. Journals differ widely in how much dedicated statistical review they can call on, and a manuscript that arrives with the statistics already independently checked is treated differently by an editor who has none available.
Every statistical concession has to reach the abstract, the results text, the figure legends and the discussion, or the revision draws the same objection a second time. This is the single most common way a technically correct revision fails.
The specific failures are easy to enumerate and easy to check. A methods section stating false discovery rate control while a figure legend still reports nominal p. An abstract asserting the original effect after the reanalysis widened the interval to include no effect. A discussion describing an effect as “marked” that the results report as a 2-point shift on a 100-point scale, where the answer is a published minimal clinically important difference or a quieter adjective. An adjusted model whose covariate coefficients are discussed as though each were a causal effect, which is the Table 2 fallacy. And a paper that now adjusts for confounding but still uses the verb “reduces” in the title.
Read the page matching the objection, not the closest general advice. Reviewer comments are terse because the reviewer assumes you share their vocabulary, and the most common revision failure is answering a sentence rather than the concern behind it.
If you are revising, the overclaim check is worth running before you resubmit: a revision that fixes the statistics while leaving the abstract’s original claims intact draws the same objection a second time, from a reviewer who now believes you were not listening.
If you have not submitted yet, these are the objections worth pre-empting. PerfectPaper’s standing review covers statistical models, confidence-interval coherence and causal inference, and custom reviewer agents cover the field-specific conventions a general statistical reader misses.
PerfectPaper reads the reported n, the figure legends, the methods and the abstract against each other, and reports where the stated sample size, the test applied and the strength of the claim stop agreeing. Where the manuscript is high-dimensional, the multiple testing agent checks correction, threshold consistency and enrichment background; where measurements come from a hierarchy, a dedicated pseudoreplication pass reconstructs what each n counts from the methods and the figure legends.
Answer the concern rather than the sentence. State what the reviewer’s objection would mean if correct, give the specific evidence bearing on it, and say plainly where you disagree and why. Editors accept reasoned disagreement; they do not accept a comment being ignored.
Often not. Many statistical objections are resolved by reanalysis at the correct unit, by reporting an effect size with its interval, or by narrowing a claim to what the data support. New data is the last resort, not the first.
Say so, with a citation, and offer the analysis you believe is correct alongside. A post hoc power calculation computed from an observed effect is the most common example: it is circular by construction, and offering a confidence interval instead usually satisfies the underlying concern.
If the objections are substantive and you are not confident answering them, yes. Name that person in the response letter, and in the author list or the acknowledgements as their contribution warrants. Reviewers read a named statistician as evidence the concern was worked through rather than argued around.
Sample size, non-independent replicates, multiple comparisons, test choice, effect magnitude and post hoc power are the six that recur most and have their own pages here. Six more complete the set: insufficient methods detail, unaddressed missing data, the significant-versus-non-significant interaction error, within-arm rather than between-arm comparison, p-values that do not reconcile with their test statistics, and analyses that were not pre-specified.
Reanalyse at the correct independent unit, report effect sizes with confidence intervals, apply or explicitly decline the requested correction with the family named, refit under the alternative test, add a sensitivity analysis for missing data, and narrow the claim to what survives. Say in the letter whether any of those changed the conclusion.
Four parts and rarely more than a short paragraph: restate the concern, state what you did, give the number with its interval, and cite the location in the revised manuscript. Long justifications read as defensiveness; a number and a location read as compliance.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect