Skip to content

SOLUTIONS

Reviewer questions my choice of statistical test

Normality, pairing and unequal variance are the three grounds behind most of these comments. What each requires, and when switching tests is the wrong fix.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Reviewer questions my choice of statistical test

A reviewer who questions your choice of statistical test is almost always standing on one of three grounds: the data are not plausibly normal, the design is paired and the analysis is not, or the variances differ between groups. Each ground has a specific remedy, and identifying which one you received matters more than reaching for a non-parametric test reflexively.

“A t-test is not appropriate here” is the usual wording. There is also a fourth possibility worth checking first, because it outranks all three: the test may be fine and the problem may be that it was applied at the wrong level.

Check the level before the test

The level of analysis outranks the choice of test, so establish the level first. If your comparison treats cells, wells or tumours as independent observations, no change of test fixes it. A Mann-Whitney on pseudoreplicated data is exactly as wrong as a t-test on it, because the problem is the independence assumption both share. Rule this out first — reviewer says my n is not independent covers it, and it is the more serious objection when both apply.

The diagnostic is mechanical rather than statistical. Count the number of times the treatment was independently applied, not the number of measurements produced. Three dishes of transfected cells imaged at 50 fields each give n = 3 for a treatment applied per dish, whatever the figure legend says. The signature in the results is a p-value that falls as the number of measured units rises while the number of experiments stays constant, which is the subject of my p-value shrinks when I add more cells and the defining feature of pseudoreplication.

Two remedies exist, and both keep the measurements. Aggregate to the independent unit and analyse the aggregate means, which is simple and defensible and gives up power you did not really have. Or fit a mixed model with a random intercept for the independent unit, which uses the within-unit measurements to estimate their own variance component and is the better answer when units contribute unequal numbers of observations. Neither is a change of test in the sense the reviewer probably meant.

The comment forms, and what each one is really asking

Reviewers rarely name the assumption they are worried about. They write one of a small number of sentences, and each maps onto a different remedy.

“A t-test assumes normally distributed data; the distributions in Figure 2 are clearly skewed” is the normality ground, and the remedy is a distributional model or a demonstration of robustness rather than an automatic rank test. “The samples are not independent — the same animals appear in both groups” is the pairing ground, and it is usually correct. “Error bars differ substantially between groups; an equal-variance test is not justified” is the unequal-variance ground, and Welch’s test answers it in one line. “The authors report a t-test on n = 240 cells from 4 mice” is the level objection wearing the clothes of a test objection.

Two further forms are worth recognising because they are not really about the two-group test at all. “Multiple pairwise t-tests were performed without correction” is a familywise error comment, covered by reviewer wants multiple comparison correction. “A linear model is inappropriate for a bounded outcome” or “counts should not be analysed as continuous data” is a model-specification comment, covered by reviewer says the model is not appropriate. Answering a specification comment with a rank test does not address it, and reviewers notice.

The rest of the family, including objections about power, effect size and correction, sits under responding to statistical reviewer comments.

If the ground is normality

Normality is the ground most often raised and most often misdiagnosed. Two things are worth knowing before you switch tests.

The t-test assumes normality of the sampling distribution, not of the data, and with reasonable group sizes the central limit theorem makes it robust to substantial non-normality. For n above roughly 15 per group with no extreme skew, the t-test is usually defensible and more powerful than its non-parametric alternative.

Normality tests are a poor basis for the decision. Shapiro-Wilk is underpowered at small n, where the assumption matters most, and over-sensitive at large n, where it matters least. Reviewers who ask for one are asking for a weak instrument. A histogram or QQ plot in supplementary material is more informative and more persuasive. The deeper problem is procedural: Rasch, Kubinger and Moder showed in Statistical Papers (2011;52:219–231) that pre-testing the assumptions of a two-sample t-test and choosing the follow-up test on the result gives worse error control than simply running the robust test from the start.

Where the data genuinely are skewed — concentrations, times to event, counts — the better answer is often a transformation or a model matched to the distribution rather than a rank test. Log-transformation for multiplicative data, a Poisson or negative binomial model for counts, and a gamma model for positive continuous outcomes all preserve interpretability that a rank test discards. Two specifics make the transformation route defensible in review. Back-transforming a mean difference on the log scale returns a ratio of geometric means, not a difference of arithmetic means, which Bland and Altman set out in their BMJ statistics note on transformation when comparing two means (1996;312:1153); state the estimand in those terms and the reviewer has nothing left to query. And overdispersion is the usual reason a Poisson model fails — when the variance exceeds the mean, a negative binomial or quasi-Poisson model is the correction, not a switch to ranks.

Non-parametric tests remain the right choice for genuinely ordinal outcomes and for small samples with clear non-normality. Note that Mann-Whitney tests stochastic dominance rather than a difference in means, so report a median difference or a Hodges-Lehmann estimate rather than describing it as a mean difference. Divine, Norton, Barón and Juarez-Colunga make the sharper version of this point in The American Statistician (2018;72:278–286): the Wilcoxon–Mann–Whitney procedure is not a test of medians, and only becomes one under a location-shift assumption most datasets do not satisfy. Fagerland’s paper in BMC Medical Research Methodology (2012;12:78) adds the large-sample case, arguing that in large studies with skewed data the rank test is not the safe default it is assumed to be, because its behaviour depends on the shapes being compared and it can disagree with a comparison of means that the paper’s conclusion actually rests on. If the manuscript’s claim is about an average, a test of stochastic dominance is a different claim — see what is effect size for how to report the one you mean.

If the ground is pairing

Pairing is the least arguable version of this comment, and the one most likely to help you. If measurements are naturally paired — before and after in the same animal, treated and control halves of one sample, matched littermates — an unpaired test discards the pairing and is both wrong and less powerful.

The remedy is a paired test or a model with a subject random effect. It usually strengthens the result, which makes this the one version of this comment that is good news.

The paired form depends on the outcome type, and picking the wrong member of the family draws a second comment. A continuous outcome takes a paired t-test, or a Wilcoxon signed-rank test where the within-pair differences are badly skewed. A binary outcome takes McNemar’s test, and analysing paired binary data with a chi-squared test on a 2 × 2 table is a distinct and common error. More than two conditions per subject takes a repeated-measures model or a mixed model with a random intercept per subject, not a series of paired comparisons.

Pairing also shows up as a design the analysis has flattened. Crossover trials, split-plot layouts, matched case-control designs and littermate-controlled animal experiments all carry a blocking factor that must appear in the model. Where discordance between paired units is itself the finding, that is a separate question about measurement reliability rather than about the test.

One trap sits next to this one. Analysing change from baseline with an unpaired test on the change scores is not the same as adjusting for baseline, and when allocation was not randomised on the baseline value the two can point in different directions because of regression to the mean. Analysis of covariance on the follow-up value with baseline as a covariate is the standard remedy, and it is more powerful than a test on change scores whenever baseline and follow-up are correlated.

If the ground is unequal variance

Use Welch’s t-test. It does not assume equal variances, sacrifices almost no power when variances are in fact equal, and is a reasonable default even without a reviewer asking. Methodologists in several fields now argue for it as the default — Delacre, Lakens and Leys in psychology, Ruxton in ecology, and Rasch, Kubinger and Moder for the two-sample case generally.

Do not use a preliminary variance test to decide. That two-stage procedure distorts the error rate of the test that follows it.

The published argument is specific enough to cite in a response letter. Delacre, Lakens and Leys, in the International Review of Social Psychology (2017;30:92–101), argue that Welch’s test should be the default because it controls the Type I error rate when variances and sample sizes are both unequal while giving up little power when they are not. Ruxton made the same case for the ecological literature in Behavioral Ecology (2006;17:688–690), calling the unequal-variance test an underused alternative to both Student’s t-test and the Mann-Whitney U test. Zimmerman’s analysis in the British Journal of Mathematical and Statistical Psychology (2004;57:173–181) is the reference for why the pre-test is the problem: conditioning the choice of test on a Levene or Bartlett result leaves the conditional error rate of the selected test wrong in both directions.

Mechanically, Welch’s test uses the same difference in means but a separate-variances standard error and the Welch–Satterthwaite approximate degrees of freedom, which are non-integer and usually smaller than n₁ + n₂ − 2. Reporting t(31.4) = 2.81, p = 0.008 with fractional degrees of freedom is the visible signature that a Welch test was run, and it is worth showing for that reason. The extension to more than two groups is Welch’s ANOVA with Games-Howell post-hoc comparisons rather than one-way ANOVA with Tukey. Where the variances differ and the distributions are non-normal, the Brunner-Munzel test addresses the nonparametric Behrens-Fisher problem directly (Biometrical Journal 2000;42:17–25) and is the correct alternative to Mann-Whitney, which assumes equal distributions under the null.

Bartlett’s test is additionally sensitive to non-normality, so a significant Bartlett result on skewed data is evidence of skew as much as of heteroscedasticity. Where variance rises with the mean — a very common pattern in concentration and count data — a variance-stabilising transformation or a model with a mean-variance relationship built in addresses the cause rather than the symptom, and it keeps a confidence interval you can interpret on the original scale.

Which test the design implies

The design determines the test, so read the design back out of the methods before choosing. The table below maps the four most frequently queried situations onto the default analysis and the substitute that reviewers flag.

Design Default analysis Commonly reported instead
Two independent groups, continuous outcome Welch’s t-test Student’s t-test with unstated equal-variance assumption
Two measurements on the same subject Paired t-test or mixed model Unpaired t-test on the two columns
Two independent groups, binary outcome Chi-squared or Fisher-Irwin, with a risk difference Chi-squared on paired data, or a t-test on proportions
Multiple measurements per animal or dish Mixed model with a random intercept per unit t-test with n set to the number of measurements
Right-skewed positive continuous outcome Log-scale model or gamma GLM Mann-Whitney reported as a difference in means
Counts with variance above the mean Negative binomial or quasi-Poisson Poisson model, or ANOVA on raw counts

For the binary row specifically, Campbell’s analysis of 2 × 2 tables in Statistics in Medicine (2007;26:3661–3675) recommends the N − 1 chi-squared test when the smallest expected count is at least 1, and the Fisher-Irwin test otherwise — a more useful rule than the widely taught “expected count below 5” threshold, and one that stops Fisher’s exact test being used where it is unnecessarily conservative.

When the obvious remedy is unavailable

The remedy is sometimes blocked, and the response then has to name the constraint rather than pretend it away. Four blocks recur.

The sample is too small for any asymptotic test. With three or four independent units per group, no test has meaningful power and a rank test cannot even produce a p-value below 0.05 at some sample sizes — Mann-Whitney with n = 3 per group has a minimum attainable two-sided p of 0.10. Report the effect estimate with its interval and say the study was not designed to test that comparison. Do not answer with a retrospective power calculation; see reviewer asks for post-hoc power and reviewer says my sample size is too small.

The distribution has no standard model. Zero-inflated, bounded, censored or bimodal outcomes are not served by any of the usual switches. A permutation test on the statistic you actually care about — the difference in means, the difference in medians, the ratio — makes no distributional assumption and is exact under the randomisation, and a bias-corrected bootstrap interval gives the effect estimate the rank test does not. Both are computable from the data already collected.

Re-running the experiment is not possible. Where the objection is about design rather than arithmetic — a missing blocking factor, a level that cannot be reconstructed — say so explicitly, present the analysis that the data will support, and move the claim to match it. A narrowed claim survives review; a wide claim defended with the wrong test does not, which is the subject of how to write a limitations section.

The reviewer has asked for a test you believe is wrong. This happens most often as a request for a rank test on data where a transformation is better, or a request for a normality test. Run the requested analysis, report it, and then give the reason you consider the original primary. A reviewer who receives both numbers is in a position to agree with you; a reviewer who receives an argument and no number is not.

Worked example: one comparison under four tests

Consider a constructed example rather than a published one: a two-arm animal study with 22 animals in the control arm and 19 in the treated arm, and a plasma cytokine concentration as the outcome — a positive, right-skewed quantity with standard deviations of 14.2 and 31.6 pg/mL. The manuscript reports Student’s t-test, p = 0.041, and concludes that treatment raises the cytokine.

Three things follow from the numbers already stated. The standard deviations differ by a factor of 2.2, and the larger variance sits in the smaller group, which is the configuration in which Student’s test is anti-conservative — its nominal 5% error rate is too generous, so p = 0.041 is optimistic. Welch’s test on the same data returns p = 0.056, and the qualitative conclusion changes at a conventional threshold. A log-scale analysis, appropriate for a positive skewed concentration, returns p = 0.019 with a ratio of geometric means of 1.62 (95% CI 1.09 to 2.41). Mann-Whitney returns p = 0.021 with a Hodges-Lehmann shift, and answers a question about stochastic dominance rather than about mean concentration.

The response that works reports all four, names the log-scale analysis as primary because the outcome is multiplicative and the model fits the observed mean-variance relationship, and states plainly that the choice was made after review rather than pre-specified. The response that fails picks whichever test gives p below 0.05 and reports only that one, which is p-hacking whether or not it was intended as such — and a reviewer who asked for a second test is already looking for exactly this.

Responding

Front-load the number, not the argument. Whichever remedy applies, run the alternative and report whether the conclusion changes. That sentence — “the conclusion is unchanged under a Welch test (p = 0.008)” — resolves most of these comments in one line, and it is more convincing than an argument that the original test was acceptable.

If the conclusion does change, that is important information and it belongs in the paper. Report both, and say which you consider primary and why.

Where you decline to switch, give the reason: robustness at your sample size, the interpretability of a mean difference, a distributional model that fits better. Reviewers accept a reasoned choice. What draws a second round is a switch made silently, or an assumption defended without evidence.

Three practical points shape the letter itself. Put the sensitivity result in the manuscript, not only in the response — a supplementary table showing the primary comparison under Student’s, Welch’s and a rank test is a permanent answer to a question that a later reader will also ask. Change only the comparisons the reviewer named, because re-running every test in the paper under a new default invites a fresh round of comparison. And keep the claim in the abstract synchronised with whatever analysis you have made primary; a conclusion sentence that still reflects the withdrawn test is the most common cause of a third round. The letter format itself is covered in how to write a response to reviewers.

How to report the test so it is not queried again

Report the test, its assumption and its estimand together, in that order. The ICMJE Recommendations set the standard for the Methods section directly: “Describe statistical methods with enough detail to enable a knowledgeable reader with access to the original data to judge its appropriateness for the study and to verify the reported results.” A sentence reading “groups were compared by t-test” does not meet that bar, because it does not say which t-test.

A reportable sentence names five things: the test, the variant, the unit of analysis, the software with its version, and the effect estimate with an interval. “Treatment groups were compared by Welch’s two-sample t-test on log-transformed concentrations, with the animal as the unit of analysis (R 4.4.1); the ratio of geometric means was 1.62 (95% CI 1.09 to 2.41).” The SAMPL guidelines of Lang and Altman, published as a chapter in Guidelines for Reporting Health Research: A User’s Manual (2014), are the general-purpose checklist for this reporting, and the trial and observational equivalents are the CONSORT checklist and the STROBE checklist.

State whether the analysis was pre-specified. A test chosen after seeing the data is not disqualified by that fact, but a paper that describes a post-hoc choice as though it were planned is making a claim that the record will not support. Where several tests are reported, say which is primary before reporting any of them, so the reader does not have to infer it from which p-value the discussion emphasises.

Related

Reviewer says my n is not independent · Reviewer says the model is not appropriate · Reviewer wants multiple comparison correction · What is statistical power · Responding to statistical reviewer comments

PerfectPaper reads the design described in the methods against the test reported in each figure legend, and reports where the unit of analysis, the pairing structure or the variance assumption stops matching the test that was run.

Review my manuscript

Frequently asked questions

What do I do when a reviewer says I used the wrong statistical test?

Identify which of four grounds the comment rests on — wrong unit of analysis, non-normality, ignored pairing, or unequal variance — because each has a different remedy. Then run the alternative, report whether the conclusion changes, and state which analysis you consider primary and why.

How do I respond to a reviewer who says a t-test is not appropriate?

Run the test the reviewer implies, report the result in one sentence alongside the original, and say whether the conclusion is unchanged. Where you decline to switch, give the specific reason — robustness at your sample size, the interpretability of a mean difference, or a distributional model that fits better.

The reviewer says my data are not normal — do I have to run a normality test?

Not usually, and formal normality tests are a poor guide — underpowered where the assumption matters and oversensitive where it does not. Show the distribution instead, and note the group sizes; that answers the comment better than a p-value about normality.

The reviewer wants a non-parametric test because my sample is small. Should I switch?

Often yes when the data are clearly non-normal, but note the trade: rank tests answer a different question and give you no interpretable effect size in original units. For skewed continuous data, a transformation or a matched distributional model frequently serves better.

The reviewer asked for Welch’s t-test. Should I make it my default?

Welch’s is a t-test that does not assume equal variances between groups. It sacrifices very little power when variances are equal and protects against inflated error when they are not, which is why methodologists in psychology and ecology now argue for it as the routine default.

The reviewer asked me to switch tests and the result became non-significant. What now?

Report both, state which analysis you consider primary and why, and adjust the claims to match the more defensible one. Concealing the discrepancy is the one response that reliably ends badly.

The reviewer objected to my test and to my unit of analysis. Which do I answer first?

The unit of analysis, every time. Test choice and independence are separate issues, and no change of test repairs an analysis run at the wrong level. Settle the level first, then choose the test the design implies.

Last updated September 10, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.