Skip to content

SOLUTIONS

Statistical review capacity

Most journals cannot get a statistician for most papers, so statistical problems are found by subject reviewers who were not asked to look — or after publication.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Statistical review capacity

Nearly every journal wants statistical review. Very few can get it for most papers. In a 2020 survey of 107 biomedical journals, 23% used specialised statistical review for all articles and 34% rarely or never used it. The pool of statisticians willing to review is far smaller than the pool of subject specialists, they are asked by many journals, and the work is unpaid.

The practical consequence is that statistical problems are found by subject reviewers who were not asked to look for them, by a statistician after acceptance when an editor gets nervous, or by readers after publication. All three are worse than finding them at submission, and the third is how corrections happen.

How often papers actually get a statistical reviewer

Statistical review reaches a minority of biomedical papers, and that is a measured figure rather than an impression. Tom Hardwicke and Steven Goodman contacted 364 journals, received 127 replies and analysed 107 eligible responses — 29% of those surveyed, spanning 57 biomedical fields and answered mostly by editors in chief (PLOS ONE 2020;15(10):e0239598). Of the responding journals, 23% used specialised statistical review for all articles, 9% used it for 50–99% of articles, 34% used it for 10–50%, and 34% rarely or never used it.

The distribution is concentrated by journal type. Clinical and hybrid journals, 88 of the 107, reported a median of 30% of articles receiving statistical review; the remaining 19 journals reported a median of 2%. A manuscript outside clinical medicine is, on those figures, unlikely to be read by anyone whose brief is the analysis.

The picture has not improved over a generation. Goodman, Altman and George contacted 171 journals in the mid-1990s and heard back from 114 (Journal of General Internal Medicine 1998;13(11):753–756); about a third of the responders guaranteed statistical review for all accepted manuscripts, and roughly half conducted it at the editor’s discretion. Two decades of rising submission volume and rising analytical complexity left the share of journals reviewing every paper roughly where it started.

The barriers editors named deserve separating. Among the 36 journals that rarely or never used statistical review, 14 said statistical review was not required for the article types they handle, 9 cited a lack of resources or access to statistical reviewers, and 8 said they considered ordinary peer review adequate. The first is a scope judgement. The second is a capacity problem. The third is a belief about what subject reviewers catch, and the next section is the argument against it.

When in the process statistical review happens

Statistical review happens at four distinct points, and the point determines how much it can change. In the same 2020 survey, among the 71 journals that used statistical review for more than 10% of articles, 35% solicited it contemporaneously with peer review, 27% after peer review but before the editorial decision, 17% on an ad hoc basis, and 6% only after an editorial decision had been made.

That last 6% is the arrangement described above: the statistician consulted once the paper is effectively accepted. By then the argument is fixed, the figures are drawn, and a finding that the analysis population was never defined arrives as an obstruction rather than as review. It is the single worst place to put the scarce specialist, and it is where a nervous editor most often puts them.

Contemporaneous review, the modal arrangement, has the opposite failure. The statistician reads at the same time as subject reviewers who may recommend rejection, so specialist attention is spent on manuscripts that will never be published — and on a journal rejecting most of what it receives, that is most of the statistical reviewing effort.

Placing statistical review after subject review and before the decision spends the specialist only on manuscripts with a future. It lengthens the elapsed time to first decision, which is the trade an editor is actually making, and it is the same sequencing argument that governs desk screening.

What subject reviewers do and do not catch

Subject specialists reliably catch statistics that contradict domain knowledge: an effect size implausible for the biology, a control that behaved impossibly, a survival curve that does not match clinical experience.

They much less reliably catch the errors that require reading the analysis as an analysis — an n that counts measurements rather than independent units, a covariate that is a mediator, a hazard ratio reported where hazards are plainly not proportional, multiplicity that was never addressed, a post hoc subgroup presented as pre-specified.

Those are not gaps in expertise so much as gaps in attention. A reviewer asked to assess whether a finding advances the field is not reading the methods with a statistician’s suspicion, and it is unreasonable to expect otherwise.

Four of those five classes have a determinate signature that does not require expertise to spot, only the instruction to look. Pseudoreplication shows up as degrees of freedom larger than the number of animals, dishes or patients described in the methods. Non-proportional hazards show up as Kaplan–Meier curves that cross beneath a single reported hazard ratio. Unaddressed multiplicity shows up as a results section containing more hypothesis tests than the methods names corrections for, with no correction method stated. A retrofitted subgroup shows up as an outcome in the abstract that does not appear in the trial registration. None of those four checks needs a statistician; all four need someone whose task that evening was the analysis.

The disclaimer is the reliable tell that no such person read the paper. “The statistical methods appear appropriate, although I am not a statistician” is a familiar sentence in biomedical peer review, and it should be read by an editor as an explicit statement that the statistics were not reviewed rather than as reassurance that they were.

Why the reviewer form decides what gets caught

The reviewer form is part of the mechanism, because reviewers answer the questions they are asked. A form whose open comment prompt reads “please comment on the significance and originality of the work” produces comments on significance and originality. A form with a discrete statistics field — is the analysis population defined, is the unit of analysis the independent unit, is a correction named where multiple tests were performed — produces answers to those questions from reviewers who would not otherwise have raised them.

This is the lowest-effort intervention available to an editor and the one most often skipped, because changing the form is an editorial-board decision rather than a per-manuscript one. A structured statistics section on the form also converts the “I am not a statistician” disclaimer into something usable: a reviewer who cannot judge model specification can still report that the reported n changes between the abstract and Table 1.

The limit is real and worth stating. A form prompt raises detection of determinate reporting problems and does not turn a subject specialist into a statistical reviewer for questions of specification, identification or test choice. It moves the floor, not the ceiling.

The triage question

Most journals cannot get statistical review for every paper, so the real question is which papers need it most.

The honest answer is: papers whose conclusions rest on an analysis rather than on an observation. A paper reporting that a mutation abolishes a phenotype in a clean genetic system needs less statistical scrutiny than one reporting a modest survival difference in an observational cohort with covariate adjustment.

Triage by how much work the statistics are doing, not by field or by whether numbers appear.

Operationally that produces three tiers. Tier one, statistical review needed: any conclusion resting on an adjusted or modelled comparison — observational cohorts with covariate adjustment, propensity score analyses, time-to-event models, mediation analyses, non-inferiority designs, prediction models, and any paper making a subgroup or interaction claim or testing many outcomes at once. Tier two, a structured check is enough: a comparison whose direction is legible in the raw data, where the model reports precision rather than producing the finding. Tier three, reporting review only: descriptive and methodological papers making no inferential claim.

The test that assigns a manuscript to a tier takes one minute. Delete the model and look at the summary statistics: if the conclusion still stands, the statistics are reporting a result; if it disappears, the statistics are producing one, and a specialist should read it. That question also separates the papers where a wrong modelling choice changes the answer from those where it changes only the confidence interval.

A worked triage: two manuscripts

Two manuscripts arriving the same morning show how far apart the tiers sit in practice.

The first reports that a knockout abolishes a migratory phenotype. Three independent lines, each assayed in three biological replicates, produce a difference visible in the raw scatter — untransformed values that do not overlap between genotypes. The analysis is a t test. If the t test were removed entirely the reader would reach the same conclusion from the figure. The statistical risk here is confined to the unit of analysis: whether n = 9 counts replicates or lines, which is pseudoreplication and is checkable against the methods without a statistician. Tier two.

The second reports that a drug is associated with a 19% reduction in mortality in a registry cohort of 18,400 patients, adjusted hazard ratio 0.81, with the unadjusted comparison showing no difference. Every part of that sentence is the analysis: the adjustment set decides the estimate, the proportional hazards assumption decides whether one hazard ratio describes the follow-up, the handling of deaths from other causes decides whether the competing risks are being treated as censoring, and the gap between the crude and adjusted results is exactly the space where a specification error lives. Removing the model removes the finding. Tier one, and it is the paper an editor should spend their one available statistician on.

Notice what triage did not use: field, journal section, sample size, or whether the authors listed a statistician among the co-authors. The second manuscript is larger, more clinically consequential and more likely to be cited, and those are the reasons an editor gives afterwards for wishing it had been reviewed.

What is checkable without a statistician

A meaningful subset of statistical problems is checkable without a statistician, and knowing which subset is what makes triage possible.

Determinate: whether the reported n is defined and what it counts; whether percentages match their numerators; whether numbers reconcile across abstract, tables and figures; whether a correction method is named where many tests were run; whether the reported test is named at all; whether confidence intervals are reported alongside p-values; whether the analysis population is stated.

Requires judgement: whether the model specification is right for the design; whether the adjustment set is causally defensible; whether a violated assumption invalidates the conclusion.

The first list is where most published statistical problems actually live. Errors of omission and inconsistency are far more common than errors of sophisticated method choice, and they are exactly the ones a determinate check finds.

The prevalence has been measured. Nuijten and colleagues ran the R package statcheck over more than 250,000 p-values reported in eight major psychology journals between 1985 and 2013 (Behavior Research Methods 2016;48(4):1205–1226) and found that half of published papers using null-hypothesis significance testing contained at least one p-value inconsistent with its own test statistic and degrees of freedom, and that one in eight contained a grossly inconsistent p-value that may have affected the statistical conclusion. Those are arithmetic inconsistencies inside the reported numbers, detectable without the data and without a specialist.

Several other checks are similarly mechanical. statcheck recomputes a p-value from the reported test statistic and degrees of freedom. The GRIM test — granularity-related inconsistency of means, described by Nick Brown and James Heathers in 2016 — exploits the fact that a mean of N integer-valued responses must be an integer divided by N, so a reported mean of 3.48 from a sample of 20 is impossible rather than merely surprising. Degrees of freedom can be read back against the stated sample size. A confidence interval can be recomputed from a reported effect and standard error. Percentages can be divided back into their denominators. Each is a question with a yes or a no, and each is a question a statistician should never have been the one to ask.

What a determinate check cannot settle

A determinate check settles arithmetic and reporting, and settles nothing about identification. Whether the adjustment set closes the backdoor paths and opens no new ones, whether an instrument is valid, whether a surrogate endpoint stands in for the clinical one, whether the estimand answers the question the abstract asks — none of those is recoverable from the manuscript’s internal consistency, because a fully self-consistent paper can be built on the wrong model.

Three failures in particular survive every mechanical check. A covariate that is a mediator produces a coherent table, a defensible-looking method and an estimate of the wrong quantity. A hazard ratio computed under violated proportionality produces a valid confidence interval around a number that describes no period of the follow-up. And a set of adjusted coefficients presented as though each were a causal effect — the Table 2 fallacy — is a presentation choice that no consistency check flags.

The boundary matters for what a journal should promise. A determinate check establishes that a manuscript is statistically reviewable: its numbers agree, its n is defined, its tests are named, its analysis population is stated. It does not establish that the analysis is right, and a journal that describes an automated check as statistical review has quietly redefined the term rather than solved the shortage.

Reducing the demand

Journals reduce demand for statistical review with two levers.

Screen before review. If a manuscript reaches a statistician with unreconciled numbers and an undefined n, the scarce specialist spends their time on things a check could have surfaced. Desk screening at scale covers the boundary. A statistical reviewer’s first paragraph is the diagnostic: if it lists missing reporting items rather than engaging the analysis, that reviewer was used as a proofreader, and they will decline the next invitation.

Tell authors what will be checked. A journal that publishes its statistical reporting expectations gives authors something to submit against. Authors are generally not concealing anything; they do not know what is expected, and an expectation that is never written down cannot be met.

Both levers have named instruments, which is the difference between a policy and an intention. The SAMPL guidelines — Statistical Analyses and Methods in the Published Literature, Lang and Altman, International Journal of Nursing Studies 2015;52:5–9, and in the EASE Science Editors’ Handbook of 2013 — set out the minimum a methods section must report, item by item, and are directly usable as a submission checklist. The ICMJE Recommendations state the standard a journal is holding authors to: “Describe statistical methods with enough detail to enable a knowledgeable reader with access to the original data to judge its appropriateness for the study and to verify the reported results”, and, in the same section, “Avoid relying solely on statistical hypothesis testing, such as P values, which fail to convey important information about effect size and precision of estimates.”

A journal that publishes both, and says which items it screens before review, converts part of its statistical-review demand into author work done before submission. The recurring failure it removes is not fraud but ignorance of the expectation — including p-hacking by authors who genuinely did not know that reporting the subgroup that worked required saying so.

What editors do when no statistician is available

An editor with no statistical reviewer available has six options, and five of them are better than sending the paper out unreviewed.

Recruit a statistical editor to the board rather than inviting one per paper. A board member reviewing on a rota answers a standing obligation instead of an ad hoc favour, which converts an invitation that may be declined into a scheduled duty — and the 1998 survey found staff or board biostatisticians concentrated in the largest journals, which is precisely the gap a smaller journal can close by appointment rather than by invitation.

Triage rather than review. A statistician spending 15 minutes assigning tiers across a week’s submissions protects more papers than the same statistician spending five hours on one manuscript, because tier assignment is where their expertise is least substitutable.

Ask the subject reviewer a specific statistical question. “Is the unit of analysis in Figure 3 the animal or the cell?” is answerable by a cell biologist. “Please comment on the statistics” is not, and produces the disclaimer.

Require the analysis code, the analysis population definition and a named analyst. A paper whose author contributions statement names who performed the analysis, and whose code is deposited, can be checked by a reader after publication even when it was not checked before. This is a weaker remedy honestly labelled, not a substitute.

Invite statisticians who are never invited. Invitation lists are built from corresponding authors of related papers, which structurally excludes the biostatisticians who appear third or fourth on those author lists and the methods faculty who publish in statistics journals rather than clinical ones — the same concentration problem described in why it is so hard to find peer reviewers. Widening the list to methods faculty and to statisticians who are not the corresponding author reaches people who are asked far less often, and who are correspondingly likelier to say yes.

The sixth option, publishing without statistical review and relying on post-publication correction, is the current default at a third of journals, and the correction record is what it produces.

What a statistical reviewer actually writes

A statistical reviewer’s comments look different from a subject reviewer’s, and an editor who has read a few can recognise their absence.

“The unit of analysis is the cell, but cells within a dish are not independent; please analyse at the level of the biological replicate and report the number of independent experiments.” “The analysis population is not defined; please state whether the primary analysis is intention-to-treat and how many randomised participants were excluded and why.” “Twenty-eight comparisons are reported and no correction is named; please state the correction or designate the analyses as exploratory.” “The Kaplan–Meier curves cross at approximately 14 months, so a single hazard ratio does not describe the follow-up; please report a time-varying effect or a restricted mean survival time.” “The adjustment set includes a variable on the causal pathway between exposure and outcome; adjusting for it estimates a direct effect rather than the total effect the abstract describes.” “The abstract reports a subgroup finding that does not appear in the trial registration; please identify it as post hoc.” “Table 2 reports adjusted coefficients for every covariate and the discussion interprets several of them causally, which the design does not support.”

Two features distinguish those comments. Each names a location in the manuscript, and each states what would resolve it. A statistical review that says the analysis is “inadequate” without either is as unusable to an author as no review at all — and a reviewer who is asked for a post hoc power calculation in response has been given the wrong remedy for the right concern.

What PerfectPaper contributes

PerfectPaper’s standing review includes specialists in statistical models and regression, causal inference and confounding, confidence-interval and p-value coherence, and cohort accounting, with code execution available for recomputing reported values. Field-specific conventions that generalist statistical review misses are covered by custom reviewer agentspseudoreplication and competing risks among them, both of which are common in the literature and rarely caught by subject reviewers.

The recomputation is the part that matters for capacity. PerfectPaper recomputes reported values rather than reading them: percentages against their numerators, p-values against their reported estimates and confidence intervals, totals against their tables, and the analysed n against the n in the abstract and the flow diagram. Every finding is anchored to a location in the manuscript, so an author or an editor sees which number disagrees with which, not that something disagrees somewhere.

PerfectPaper does not replace a statistician on a paper whose conclusions turn on a contested modelling choice. It reduces how often a statistician is needed for a paper whose problem is that the n was never defined.

More in this cluster: the peer review capacity problem.

Talk to us about journal use

Frequently asked questions

Why do so few papers get statistical review?

The pool of statisticians willing to do unpaid review is far smaller than the pool of subject specialists, and they are invited by many journals at once. Most editors can secure statistical review only for a minority of submissions. In the 2020 Hardwicke and Goodman survey of 107 biomedical journals, 34% rarely or never used specialised statistical review.

What proportion of journals use a statistical reviewer for every paper?

23% of the 107 journals responding to the 2020 survey used specialised statistical review for all articles, against about a third of the 114 responders to the equivalent 1998 survey. Clinical and hybrid journals reported a median of 30% of articles receiving statistical review; other journals reported a median of 2%.

Which papers most need a statistical reviewer?

Papers whose conclusions depend on the analysis rather than on the observation — observational comparisons, adjusted models, survival analyses, anything with multiplicity or subgroup claims. Papers resting on a clean qualitative result need it least. The one-minute test: delete the model and see whether the conclusion survives.

What statistical problems can be found without a statistician?

Undefined or inconsistent n, percentages that do not match their numerators, numbers that disagree between abstract and tables, unnamed tests, missing correction where many tests were run, and unstated analysis populations. These account for a large share of real published problems: statcheck found at least one internally inconsistent p-value in half of published psychology papers using significance testing.

How can a journal need less statistical review than it needs now?

Two levers reduce the demand. Screening reporting items before review keeps the scarce specialist off manuscripts whose numbers do not yet reconcile, and publishing the statistical reporting expectation moves that work to the author before submission. The SAMPL guidelines of Lang and Altman are a published, item-by-item instrument a journal can adopt without writing its own.

What should an editor do when no statistical reviewer is available?

Use the statistician for triage rather than full review, ask subject reviewers specific statistical questions instead of general ones, add a structured statistics section to the reviewer form, require the analysis code and a named analyst, and invite biostatisticians who never appear on corresponding-author invitation lists.

Can automated checking replace statistical review?

No, for questions of model specification and causal identification. It can substantially reduce how often statistical review is needed for problems of omission and inconsistency, which is where most errors actually are.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.