Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
When correction is genuinely required, when it is arguable, and how to answer a reviewer who asks for it on analyses where it would be the wrong choice.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A reviewer asking for multiple comparison correction is raising one of three situations, and each needs a different answer. Sometimes correction is plainly required and you should apply it. Sometimes it is arguable and you should justify the choice you made. Sometimes correction would be the wrong move on that set of analyses, and you should say so.
Deciding which you have is the whole task. The rest is arithmetic.
Multiple comparison correction exists because the chance of at least one false positive rises with the number of tests, not with the strength of any one of them. Under a global null hypothesis, 20 independent tests at α = 0.05 return at least one significant result 64% of the time, because 1 − 0.95²⁰ = 0.64. At 50 tests it is 92%. Nothing about any individual p-value changes; the reference class does.
FDA works the same arithmetic in Multiple Endpoints in Clinical Trials: Guidance for Industry (final guidance, October 2022), using the regulatory convention of a one-sided 2.5% error rate per endpoint. On that convention, two independent endpoints where success on either alone would establish an effect roughly double the error rate, from 2.5% to about 5%. Three independent endpoints put it near 7%, and ten near 22%. Correlated endpoints still inflate the rate, though by less than independent ones — which makes correlation a reason to choose a method that exploits it, never a reason to skip correction.
What the arithmetic does not tell you is which tests belong in the count. That is the judgement the reviewer is really asking you to defend, and it is not arithmetic at all.
Genome-scale or high-dimensional testing. If you tested thousands of genes, peaks, proteins or metabolites and called some significant, false discovery rate control is not optional and is not arguable. If none of 20,000 genes were truly differential, roughly 1,000 of them would still clear a nominal p < 0.05, so a list selected on nominal p carries false positives on that scale by construction. State the method and the threshold. Bonferroni over 20,000 genes puts the per-test threshold at 2.5 × 10⁻⁶; the genome-wide significance threshold of 5 × 10⁻⁸ conventional in genome-wide association studies is the same calculation over approximately one million independent common variants. High-dimensional work is the one setting where nobody disputes the requirement, which is why an uncorrected omics table is among the fastest routes to a statistical rejection.
A family of tests supporting one conclusion. If your claim is “treatment affects the immune compartment” and you tested twelve cell populations, those twelve tests support one claim and the family-wise error rate applies to the family. The tell is the sentence in your abstract: if a single one of the twelve, whichever it turned out to be, would have produced that sentence, the twelve are one family. This is the most common genuine version of the objection in flow cytometry work, where a 14-colour panel yields dozens of gated populations and the manuscript reports the three that moved.
Multiple pairwise comparisons after an omnibus test. Comparing every group against every other after an ANOVA inflates error unless corrected — Tukey, Dunnett against a single control, or Holm. The count grows fast: k groups give k(k − 1)/2 pairwise comparisons, so six groups give 15 tests, not six. If the comparisons that matter are all against one control, the count is k − 1 = 5, and choosing the design that produces five tests instead of 15 is worth more power than any correction method will give back. The non-parametric equivalent is Dunn’s test after a Kruskal-Wallis, and a manuscript that runs Kruskal-Wallis and then unadjusted Mann-Whitney tests has the same defect in a different notation. See wrong statistical test.
Multiple primary endpoints in a trial. If two endpoints can each independently establish success, the alpha must be split or otherwise controlled. This is a regulatory expectation, not a preference. FDA’s position is that where a trial carries more than one primary or secondary endpoint, the endpoints belonging to each family, and the analyses that will test them, are specified in advance rather than settled once the data are in. Alpha splitting (0.04 and 0.01, say) and hierarchical gatekeeping — test the second endpoint only if the first succeeds — are the two standard structures, and a trial with several endpoints and no stated structure has not controlled anything. Pre-specification is checkable against the registry record; see trial registration and endpoints.
Repeated evaluation of the same endpoint over time or at interim looks. ICH E9 lists repeated evaluation over time and interim analyses as sources of multiplicity alongside multiple endpoints and multiple treatment comparisons. Testing the same outcome at weeks 4, 8, 12 and 24 and reporting the week that reached significance is a multiplicity problem even though there is only one endpoint. The remedy ICH E9 points to is design rather than correction: collapse the repeated measurements into a summary measure such as an area under the curve, so there is one test instead of four. Pre-specifying a single primary timepoint achieves the same reduction.
The family is the set of tests over which the error rate is controlled, and choosing it — not choosing between Holm and Bonferroni — is where manuscripts and reviewers actually disagree. Correction is a function of two inputs, the method and the family, and only one of them is usually stated.
Three questions settle most cases. First, would any single positive result in the set, on its own, have supported the conclusion you drew? If yes, the set is one family. Second, was the set defined before you looked at the data, or is it the subset that survived? A family assembled after the results are known is not a family; it is a selection, and it is the mechanism behind p-hacking whether or not anyone intended it. Third, would a reader who saw only the significant tests be misled about how many were run? If yes, the count belongs in the paper regardless of what you do about it.
Two families that are genuinely separate can coexist in one manuscript, and saying so explicitly is a stronger position than correcting across everything. What no manuscript can do is leave the family undefined and then assert that correction was unnecessary. ICH E9 requires the opposite in confirmatory work: “any aspects of multiplicity which remain after steps of this kind have been taken should be identified in the protocol; adjustment should always be considered and the details of any adjustment procedure or an explanation of why adjustment is not thought to be necessary should be set out in the analysis plan.” The clause that helps you is the last one. A stated explanation of why adjustment is not necessary is an accepted output. Silence is not.
Pre-specified secondary endpoints reported descriptively. If they are labelled secondary, interpreted as supportive rather than confirmatory, and not used to claim success, many statisticians accept reporting them uncorrected with that framing made explicit. The boundary FDA draws is between analyses that characterise an effect already demonstrated — time of onset, distribution of effect sizes, effects on the components of a composite endpoint — which it treats as descriptive, and analyses that set out to demonstrate an additional effect, which need to sit inside the pre-specified testing strategy. The same line works in a non-regulatory paper: characterising an established effect is description, establishing a new one is a test.
Distinct questions in one paper. Two unrelated hypotheses tested on the same cohort are not obviously one family. The relevant question is whether a reader would treat any single positive as establishing the claim. A paper reporting a treatment effect on tumour volume and, separately, a correlation between two baseline markers has two questions, not one family of two tests.
Exploratory analyses labelled as such. Exploration generates hypotheses. Correcting an exploratory screen into oblivion can be as misleading as not correcting a confirmatory one. Label it, report effect sizes, and do not describe the results as findings. Rothman argued the general form of this position in Epidemiology (1990;1:43–46), “No adjustments are needed for multiple comparisons”, and Perneger made the applied case in the BMJ (1998;316:1236–1238), “What’s wrong with Bonferroni adjustments” — that adjusting inflates type II error, that the adjusted result depends on which other tests happen to be in the paper, and that the null hypothesis being controlled (no effect anywhere) is rarely one anybody cares about. Citing one of these is legitimate. Citing one of them to avoid correcting a genome-wide screen is not, and a statistical reviewer will know the difference.
The honest position in every arguable case is the same: state what you did, state why, and let the reader judge. Reviewers rarely object to a justified choice. They object to a silent one.
Two situations recur.
Correcting across independent questions. If a reviewer asks you to Bonferroni-correct every p-value in the paper including unrelated analyses, that is not standard practice and it would make the paper less informative. Say that the tests address distinct hypotheses, not one family, and that correction across them would inflate type II error without controlling anything meaningful. The concrete penalty is worth stating: correcting 40 p-values that span four unrelated questions applies a threshold of 0.00125 to each, which drops power far below what the study was designed for. A study powered at 80% for a single comparison retains nothing like that power once its alpha is divided by 40, and the shortfall is not recoverable by any choice of method. See statistical power.
Correcting a single pre-specified primary comparison because other analyses appear elsewhere in the paper. A pre-specified primary endpoint is one test. Descriptive or secondary analyses reported alongside it do not retroactively make it a family. The defence is documentary rather than statistical: point to the protocol, the registry entry, or the analysis plan showing that the primary comparison was fixed before the data were seen, and note that the secondary material is reported without significance claims. A reviewer who is asking because the paper does not make the pre-specification visible is asking a fair question in an unfair form, and the fix is to make it visible.
In both cases, answer the objection rather than the request. A reply that says only “we disagree” fails; a reply that names the family, gives the reason, and shows the corrected numbers anyway usually succeeds. Writing the response letter matters as much as being right.
Benjamini-Hochberg controls the false discovery rate and is the default for high-dimensional work: it accepts a known proportion of false positives among discoveries, which is the right trade when you are generating a list to follow up. Benjamini and Hochberg introduced it in the Journal of the Royal Statistical Society: Series B (1995;57:289–300). It assumes independence or positive dependence among the tests; the Benjamini-Yekutieli variant (Annals of Statistics 2001;29:1165–1188) holds under arbitrary dependence, paying a factor of roughly ln(m) in the threshold, which for 20,000 tests is about 10.
Bonferroni or Holm control the family-wise error rate, appropriate when a single false positive would be damaging — a small number of confirmatory tests, or a regulatory setting. Holm is uniformly more powerful than Bonferroni and there is no reason to prefer plain Bonferroni. Holm’s step-down procedure appeared in the Scandinavian Journal of Statistics (1979;6:65–70); Hochberg’s step-up procedure (Biometrika 1988;75:800–802) is more powerful still under positive dependence, but that dependence structure is an assumption rather than a given, and a confirmatory setting is the wrong place to assume it without argument.
Dunnett when every comparison is against one control, which is common in dose-response designs and more powerful than correcting all pairwise comparisons. Dunnett published the procedure in the Journal of the American Statistical Association (1955;50:1096–1121). It uses the correlation induced by sharing a control arm, which is exactly the structure a generic correction throws away.
Permutation-based maxT — the Westfall-Young approach — estimates the null distribution of the maximum statistic by resampling, so it accounts for the actual correlation among tests rather than assuming the worst case. It is the most powerful family-wise option when tests are strongly correlated and you can afford the resampling.
| Method | Controls | Use when |
|---|---|---|
| Benjamini-Hochberg | False discovery rate | High-dimensional screens generating a candidate list |
| Benjamini-Yekutieli | False discovery rate under arbitrary dependence | Tests with unknown or negative dependence |
| Holm | Family-wise error rate | A small confirmatory set; strictly better than Bonferroni |
| Hochberg | Family-wise error rate | Positively dependent confirmatory tests |
| Dunnett | Family-wise error rate | Every comparison is against one shared control |
| Tukey | Family-wise error rate | All pairwise comparisons among groups are of interest |
| Westfall-Young maxT | Family-wise error rate | Strongly correlated tests, resampling is affordable |
| Hierarchical gatekeeping | Family-wise error rate | Ordered endpoints where later tests depend on earlier success |
State the method, the threshold, and the family it was applied over. The family definition is the part reviewers most often find missing.
Multiple comparison correction changes the significance threshold and leaves the effect estimate exactly where it was. A hazard ratio of 1.8 is 1.8 before and after Benjamini-Hochberg. What changes is whether you are entitled to call it a finding.
Two consequences follow. First, reporting adjusted p-values beside unadjusted confidence intervals is internally inconsistent, because a 95% interval that excludes the null while the adjusted p exceeds the threshold tells the reader two different things. Either widen the intervals to match the correction — Benjamini and Yekutieli’s false-coverage-rate intervals (Journal of the American Statistical Association 2005;100:71–81) do this for selected parameters — or report unadjusted intervals and say plainly that they are unadjusted. Second, selection inflates the estimates themselves. The effects that clear a multiplicity threshold are, on average, larger than their true values, because clearing the threshold required an upward error. This is why replication effect sizes shrink, and why an effect size taken from a selected screen should be described as an upper bound rather than an estimate. Neither problem is solved by the correction; both need to be said out loud.
Multiple comparison correction has no clean answer when the correct fix is unavailable to you — the statistical analysis plan was never written, the experiment cannot be repeated, or applying the correction the reviewer wants removes every result in the paper. Three moves are available, in order of preference.
Reframe the paper to match its evidence. A screen that generates candidates is a legitimate contribution when it is presented as one. Change the claim from “X regulates Y” to “we identify N candidates, of which Y showed the largest effect”, report the corrected values, and describe the follow-up that would confirm it. Reviewers accept demotions far more readily than they accept defended overclaims; this is the same failure mode as an overclaim anywhere else in the manuscript.
Add an independent confirmatory test on the specific hypothesis. One targeted experiment on one hypothesis is a family of one, and it does not inherit the multiplicity of the screen that generated it. This is often cheaper than the reviewer’s alternative, and it converts a contested statistical argument into an empirical answer.
Report both and let the reader choose. Give the uncorrected values in the main text with the correction stated, the corrected values in a supplementary table, and the count of tests performed. A reader who disagrees with your family definition can still evaluate the work. What you cannot do is retrofit a pre-specification — describing an analysis as pre-specified when the record does not support it is a misstatement, and registry records are public.
Multiplicity defects are caught by counting tests and reconciling thresholds across the manuscript, not by reading the statistics section alone. Four checks find most of them.
Count the tests actually performed, not the ones reported. Add up every comparison in every figure panel, including the ones drawn without an asterisk. The figures usually contain more tests than the methods section describes.
Reconcile the threshold across text, tables, figure legends and supplement. A methods section stating FDR control while a figure legend reports nominal p is the single most common finding in this class, and it is entirely mechanical to detect.
Check that the stated family matches the stated claim. If the abstract makes one claim and the analysis reports fourteen tests bearing on it, the family is fourteen whatever the methods section says.
Check that the correction was applied to the right set of p-values. Applying Benjamini-Hochberg after filtering to the genes that already passed a nominal cutoff controls nothing, because the input distribution has been truncated. The same error appears when correction is applied within each figure panel separately and the panels address one claim.
Reviewers rarely write “correct for multiple comparisons” as their first sentence. They write the specific version, and these are the recurring forms.
“The number of statistical tests performed is not stated.” “Figure 3 reports 24 pairwise comparisons without adjustment; please apply an appropriate correction or justify its absence.” “The methods describe FDR control at q < 0.05, but Figure 2 and Table S4 report nominal p-values.” “It is unclear which endpoint was primary and whether the reported comparisons were pre-specified.” “The authors report significance for the subgroup analysis; was a test of interaction performed?” “Adjustment appears to have been applied within each panel rather than across the family of tests supporting the stated conclusion.” “Please state the family over which correction was applied.” “The screen identifies 312 differentially expressed genes at nominal p < 0.05; how many survive correction?”
The last two decide papers. A reviewer asking which family you used has already accepted that correction is a judgement; a reviewer asking how many survive has already guessed the answer.
A multiplicity comment is answered in one of two directions — apply the correction, or name the family and decline — and each has a mechanical follow-through that decides whether the reviewer is satisfied.
If you are applying correction, apply it everywhere consistently and check that the text, tables and figures agree afterwards. A methods section stating FDR control while a figure legend reports nominal p is a common and avoidable finding.
If you are declining, name the family, explain why these tests do not constitute one, and offer the corrected values in supplementary material so the reviewer can see the result either way. That last step resolves most disputes without conceding the argument.
Do not offer post hoc power as a defence of a result that failed correction; it is computed from the observed p-value and adds no information. Do not describe a result that failed correction as “trending towards significance”. And do not answer a multiplicity comment by adding more tests, which is the response that most reliably produces a second round.
Multiple testing correction · FDR versus family-wise error · p-hacking · Statistical reviewer objections
For omics manuscripts, the multiple testing agent checks correction, threshold consistency across the manuscript, and enrichment background — the three places this goes wrong before a reviewer ever sees it.
PerfectPaper counts the comparisons in every figure panel, reconciles the significance threshold across the methods, tables, legends and supplement, and reports where the stated family does not match the claim the abstract makes.
Correction is required when several tests support one conclusion, when you ran many tests and reported the significant ones, or when more than one endpoint could independently establish success. High-dimensional screens always require it. Two genuinely separate questions tested on the same cohort usually do not.
Applying it is one of two acceptable answers; the other is naming the family, explaining why these tests do not form one, and supplying the corrected values in supplementary material anyway. What is not acceptable is leaving the family undefined and asserting that correction was unnecessary, because that answers nothing the reviewer asked.
The family is the set of tests over which any single positive would have supported the conclusion you drew. Define it before looking at the results, state it in the methods, and report how many tests it contains. ICH E9 asks for either the adjustment procedure or a stated explanation of why adjustment is unnecessary.
Secondary endpoints belong inside the correction if any of them could be used to claim success. Secondaries that are pre-specified, labelled as secondary, and interpreted as supportive rather than confirmatory can be reported uncorrected, provided the manuscript says that is what they are.
Bonferroni is conservative, and Holm controls the same family-wise error rate with strictly more power, so a reviewer naming Bonferroni is usually naming the idea rather than the procedure. The larger question is whether family-wise control is the right target at all: false discovery rate control fits discovery work, and family-wise control fits confirmatory tests where one false positive would be damaging.
Name the family, show that the tests address distinct hypotheses rather than one claim, state that correcting across them would inflate type II error without controlling anything meaningful, and supply the corrected values in supplementary material anyway. A justified refusal with the numbers attached resolves most of these.
Report that outcome plainly, with effect sizes and intervals, and reframe the work as exploratory. A screen that generates candidates for follow-up is a legitimate paper. A screen presented as confirmatory findings that do not survive correction is not, and defending it takes longer than demoting it.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect