Skip to content

SOLUTIONS

What is p-hacking?

P-hacking is trying analytic options until one crosses p < 0.05 and reporting only that one. Four common flexibilities together give a 60.7% false-positive rate.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is p-hacking?

P-hacking is the practice of trying several analytic options — different outcomes, covariates, exclusions, subgroups or stopping points — and reporting only the version that crossed a significance threshold. P-hacking rarely involves dishonesty; each individual choice is usually defensible. Simmons, Nelson and Simonsohn showed in 2011 that four ordinary flexibilities, used together, produce a false-positive rate of 60.7%.

The number above is the whole argument. A researcher exploiting four common degrees of freedom is more likely to find a significant effect that does not exist than to correctly conclude that nothing is there, while every sentence of the resulting paper remains literally true.

What p-hacking is, precisely

P-hacking is selection of an analysis conditional on its result. The defining feature is not the number of tests run but the fact that which analysis reaches the reader was decided after the p-value was seen, and that the selection was not disclosed.

Three things follow from that definition, and each is routinely misunderstood.

P-hacking does not require running many tests. One analysis, chosen from two after inspecting both, is p-hacking. Conversely, a hundred prespecified tests reported in full is not p-hacking — it is a multiple-comparisons problem, which is a different defect with a different remedy.

P-hacking does not require intent. The 2011 simulations assume a researcher who believes every choice is made on methodological grounds. Motivated reasoning supplies the belief; the selection does the damage.

P-hacking is invisible in the reported result. A p-hacked p-value of 0.03 is arithmetically correct. Nothing in the number, the model output or the fit statistics distinguishes it from an honest one. The evidence for p-hacking lives in what was not reported.

Why analytic flexibility inflates false positives

Analytic flexibility inflates false positives because each additional analytic path is an additional opportunity for noise to cross the threshold, and the researcher keeps only the maximum. Simmons, Nelson and Simonsohn quantified this with 15,000 simulated samples per scenario, drawn from a null distribution in which no effect exists, using a baseline two-condition design with 20 observations per cell.

Researcher degree of freedom False-positive rate at p < .05
Two dependent variables correlated at r = .50 9.5%
Adding 10 more observations per cell if the first test fails 7.7%
Controlling for gender, or for a gender × treatment interaction 11.7%
Dropping, or not dropping, one of three conditions 12.6%
All four combined 60.7%

Optional stopping deserves separate attention because researchers consistently underestimate it. A researcher who begins with 10 observations per condition and tests for significance after every additional observation, stopping at 50 per condition, obtains p < .05 about 22% of the time when the null is true. The intuition that “a few extra participants cannot matter much” is wrong by a factor of four.

The same paper demonstrated the effect on real participants. Twenty undergraduates listened to either “When I’m Sixty-Four” by The Beatles or an instrumental control track; controlling for father’s age, those who heard the Beatles song were nearly a year and a half younger by date of birth — adjusted means of 20.1 against 21.5 years — F(1, 17) = 4.92, p = .040. Getting younger is not possible. The authors then disclosed what the report had omitted: 34 participants collected across three conditions, and a battery of additional measured variables, of which the paper had presented only the subset that worked.

The forms p-hacking takes

P-hacking takes a small number of recurring forms, and naming them is more useful than the general term because each leaves a different trace in a manuscript.

Outcome switching. A study prespecifies one primary outcome, finds nothing, and reports a secondary or newly constructed outcome in the abstract. This is the dominant form in clinical research and the one journals police most directly.

Optional stopping. Data collection continues, or halts, based on interim significance rather than a prespecified rule.

Covariate flexibility. Covariates are added or removed until the coefficient of interest becomes significant. This form also risks introducing collider bias or removing a genuine effect by adjusting for a mediator, so the harm compounds.

Exclusion flexibility. Outlier rules, attention checks, protocol-violation criteria and quality filters are set after seeing which cutoff yields significance. Reaction-time research is the canonical case: fast-response and slow-response cutoffs vary widely across published studies, and no convention decides between them, so the cutoff is chosen paper by paper and its justification is rarely independent of the data it is applied to.

Subgroup search. The overall effect is null, so the paper reports the effect in women, or in the high-baseline tertile, without correction and without prespecification.

Model-form flexibility. Transformation, link function, random-effects structure, or the decision to dichotomise a continuous variable at the median. Choosing among plausible statistical tests after seeing which one clears the threshold is p-hacking even when the chosen test is the more appropriate one.

Silent condition dropping. A three-arm experiment reports two arms.

P-hacking, HARKing and fraud are three different things

P-hacking, HARKing and data fabrication are frequently conflated, and the conflation makes conversations about all three less useful.

P-hacking selects the analysis after seeing the result. The data are real, the test is correctly computed, the reporting is incomplete.

HARKing — hypothesising after the results are known — selects the hypothesis after seeing the result. A paper that presents an exploratory finding as the study’s original prediction has HARKed even if only one analysis was ever run. The two often travel together: the p-hacked analysis needs a hypothesis, and the hypothesis is written to fit.

Fabrication and falsification invent or alter the data themselves. John, Loewenstein and Prelec surveyed 2,155 academic psychologists in 2012 and found self-admission rates of 63.4% for failing to report all of a study’s dependent measures and 55.9% for deciding whether to collect more data after checking significance, against 0.6% for falsifying data. The questionable practices are orders of magnitude more common than fraud on every estimate the paper reports, and respondents rated them as substantially more defensible.

Exploratory analysis is not on this list, because exploratory analysis is legitimate. Searching a dataset for patterns is how hypotheses are generated. The defect appears only when the search is presented as a test.

The garden of forking paths: p-hacking without a decision

The garden of forking paths, a phrase from Andrew Gelman and Eric Loken’s 2013 analysis, describes the case where a researcher runs exactly one analysis and the false-positive rate is still inflated. The choices were contingent on the data — had the numbers come out differently, a different but equally defensible analysis would have been chosen — even though only one path was ever walked.

This matters because it defeats the most common defence. “I only ran one test” is not evidence against p-hacking; it is evidence about how many tests were executed, not how many were available.

The scale of the available paths was measured directly by Silberzahn and colleagues in 2018. Twenty-nine analyst teams, comprising 61 researchers, received the same dataset and the same question: are football referees more likely to give red cards to dark-skin-toned players? The teams used 21 unique combinations of covariates and returned odds ratios from 0.89 to 2.93, median 1.31. Twenty teams (69%) found a statistically significant positive effect and nine did not. Neither the analysts’ prior beliefs nor their statistical expertise explained the spread, and peer ratings of analysis quality did not either.

One dataset, one question, competent analysts acting in good faith, and the answer ranged from a null to a near-tripling of odds. That spread is the size of the garden.

How p-hacking is detected in a real manuscript

P-hacking is detected in a manuscript by comparing what was planned against what was reported, and by looking for internal inconsistencies that a single prespecified analysis would not produce. No statistical test on the reported numbers can do it. These are the checks that work.

Registration against report. For a registered study, open the registry entry and compare the primary outcome, the timepoint and the analysis population against the abstract. A primary outcome that appears in the registry but not in the paper, or an abstract outcome absent from the registry, is the strongest available signal.

Degrees of freedom against stated sample size. The denominator degrees of freedom of an F test fix how many observations actually entered that model, once the estimated parameters are counted back in. An analysis of covariance with one factor, one covariate and an intercept reported as F(1, 17) analysed 20 cases. When the methods section says 34 participants were recruited, that gap is the finding, and it is often the only place exclusions appear at all.

Covariate sets that change between tables. If Table 2 adjusts for age, sex and site while Table 3 adjusts for age, sex, site and baseline severity, the manuscript must say why. Unexplained drift between adjustment sets is the printed residue of a search.

Exclusion rules that appear after the results. An outlier criterion stated in a footnote, with no citation and no justification independent of this dataset, is a candidate. So is any exclusion described with a threshold that is not a round number or a published convention.

Analyses labelled neither prespecified nor exploratory. CONSORT 2010 item 18 asks for “Results of any other analyses performed, including subgroup analyses and adjusted analyses, distinguishing pre-specified from exploratory”. A manuscript that does not make that distinction has not met the checklist, and the omission is checkable in seconds.

Reported p-values against their own test statistics. Nuijten and colleagues recomputed 258,105 p-values across 30,717 psychology articles published between 1985 and 2013 using the statcheck algorithm. Among articles containing null-hypothesis tests, 49.6% contained at least one p-value inconsistent with its reported test statistic and degrees of freedom, and 12.9% contained a gross inconsistency that changed whether the result was significant. Gross inconsistencies were more common among results reported as significant.

PerfectPaper reads the complete manuscript rather than an abstract or a sampled section, so the registration statement, the methods, the figure legends and every table are available to the same pass — which is what makes the comparison checks above possible at all.

What corpus-level detection can and cannot establish

Corpus-level detection of p-hacking works by examining the distribution of published p-values rather than any single paper, and its findings are more contested than the summaries suggest. The logic is sound: under a true null, p-values are uniformly distributed, so an excess of values just below 0.05 implies selection. Two methods dominate — the caliper test, which compares the density immediately below a threshold with the density immediately above it, and p-curve, introduced by Simonsohn, Nelson and Simmons in 2014, which reads the shape of the significant p-value distribution as evidence for or against a real effect.

The limitation is worth stating plainly, because it is a case study in the very problem. Head and colleagues text-mined millions of p-values from PubMed Central in 2015 and reported a bump just below 0.05, concluding that p-hacking is widespread. Hartgerink reanalysed the same data in 2017 and found the conclusion turned on two analytic choices: a bin width of 0.005 comparing 0.04–0.045 against 0.045–0.05, and the exclusion of p = .045 and p = .05. Shifting the bin boundaries so that the heavily-reported two-decimal values sit at the top of each bin, and retaining the excluded values, left no evidence of left-skew p-hacking at all — at bin widths of 0.00125, 0.005 and 0.01 alike. Separately, researchers report p-values to two decimal places far more often than to three, which manufactures peaks at .01, .02, .03, .04 and .05 with no p-hacking involved.

Corpus methods can establish that selection is present in a literature. They cannot establish that a particular paper was p-hacked, and their own results depend on analytic choices that a critic can reasonably contest.

Evidence that pre-commitment changes what gets published

Pre-commitment to an analysis before seeing the data changes published results by a margin large enough to be visible without any statistical subtlety. Three bodies of evidence converge.

Trial registration. Kaplan and Irvin examined large NHLBI-funded cardiovascular trials and found that 17 of 30 (57%) published before 2000 reported a significant benefit, against 2 of 25 (8%) published after 2000, χ² = 12.2, p = 0.0005. None of the pre-2000 trials had been prospectively registered; all of the post-2000 trials had. After 2000 the relative risks clustered far more tightly around 1.0, and neither comparator choice nor industry co-sponsorship explained the shift. Registration did.

Registered Reports. Scheel, Schijen and Lakens compared the first hypothesis of every published Registered Report in psychology as of November 2018 (N = 71) against a random sample of hypothesis-testing studies from the standard literature (N = 152). Positive results appeared in 96% of standard reports and 44% of Registered Reports. In a Registered Report, peer review and the acceptance decision occur before the results exist.

Outcome monitoring. The COMPare project prospectively checked 67 trials published in the New England Journal of Medicine, The Lancet, JAMA, the BMJ and the Annals of Internal Medicine against their registered protocols, in real time, as each trial appeared. Fifty-eight of the 67 were misreported: prespecified outcomes went unreported, and outcomes that no registration had named were added without being labelled as new. Nine trials reported every outcome correctly, which is the useful part of the result — perfect reporting is achievable, and most of these trials did not achieve it.

Whether these gaps are entirely attributable to p-hacking is genuinely arguable — registration coincided with other reforms, and Registered Reports may attract different hypotheses. The direction and the magnitude are not arguable.

What reviewers say when p-hacking is suspected

Reviewers rarely write the words “p-hacking”. They write comments that mean it, and recognising the phrasing lets an author address the substance before submission rather than in a rebuttal.

“The primary outcome reported here does not match the registered primary outcome, and no explanation is given.” “It is unclear whether the exclusion criteria were specified before or after the analysis.” “Please state the stopping rule and whether interim analyses were conducted.” “The covariate set differs between Tables 2 and 3; please justify.” “The subgroup finding is presented as a primary result but appears to be exploratory and is uncorrected.” “The reported degrees of freedom are inconsistent with the stated sample size.” “Given the number of comparisons undertaken, the significance of this finding is difficult to interpret.”

The most damaging version is the shortest: “This reads as post hoc.” It is difficult to rebut, because the reviewer is not disputing an analysis — the reviewer has stopped believing the reported sequence of events. Answering it well requires evidence rather than assurance, which is why a response to reviewers on this point should point to a timestamped protocol, an analysis-code commit history, or a full report of every analysis run.

How to protect an analysis before the data arrive

Protecting an analysis against p-hacking is done before data collection, because after the data arrive the only remaining options are disclosure and honesty about what happened.

Preregister the analysis, not just the study. A registration naming the primary outcome but leaving the model, the covariates, the exclusion rules and the handling of missing data unspecified leaves most of the garden intact. Name the estimand, the model, the covariates, the exclusion criteria, the transformation and the correction procedure.

Fix the stopping rule in advance. Either a fixed sample size justified by a power calculation, or a formal group-sequential design with alpha spending, such as O’Brien-Fleming or Pocock boundaries. Both are legitimate; informal peeking is not. A sample size chosen for feasibility reasons is acceptable if the reason is stated — reviewer objections to sample size are answered by a stated rule, not by post-hoc justification.

Specify the correction procedure with the hypothesis count. Deciding between Bonferroni and Benjamini-Hochberg false discovery rate control after seeing which one preserves the finding is itself a degree of freedom. In high-dimensional work, where the count runs to thousands, multiple-testing control in omics has to be part of the design rather than a paragraph in the results.

Separate the confirmatory and exploratory sections of the paper before running either. The distinction cannot be made retrospectively with any credibility, and reviewers know it.

Report a multiverse or specification curve where the choices are genuinely arguable. Steegen and colleagues proposed multiverse analysis in 2016 and Simonsohn, Simmons and Nelson formalised specification curve analysis in 2020. Both report the estimate under every defensible combination of analytic choices instead of one. A finding that survives 200 specifications is a different claim from one that survives the specification its author chose.

How to report an exploratory analysis credibly

An exploratory analysis is reported credibly by labelling it exploratory in the same sentence that states its result, not in a limitations paragraph three pages later. Reviewers and readers assign the label to whatever is nearest the number.

Report the denominator. “We examined 14 secondary outcomes; one reached nominal significance” is a fundamentally different statement from “cognitive score improved significantly”, and only the first lets a reader calibrate. Give the uncorrected p-value and the corrected one where a correction applies.

State what would confirm it. An exploratory finding with a named confirmatory design attached — the sample size, the outcome, the threshold — reads as a hypothesis. The same finding with a mechanistic story attached reads as an overclaim, and overclaiming relative to the evidence is the objection that follows.

Do not convert an exploratory finding into the paper’s title. A title that states the subgroup effect commits the manuscript to a claim the analysis does not support, and it is the first thing a reviewer compares against the registration.

Report effect sizes with intervals rather than significance alone, since a p-hacked estimate is biased upward in magnitude as well as in significance. This is why effect sizes that are statistically significant but not meaningful cluster in exactly the literatures where selection is strongest, and why failures to replicate tend to return smaller estimates rather than opposite ones.

Contested points and limitations

Several claims about p-hacking are treated as settled and are not, and a manuscript that acknowledges the disagreement is more credible than one that does not.

How much published research is affected is unknown. The corpus-level estimates disagree with each other, are sensitive to bin width and rounding conventions, and cannot be checked against ground truth. Anyone quoting a single percentage of the literature as p-hacked is quoting a number that no method currently supports.

Whether all optional stopping is a defect is disputed. Under a Bayesian analysis, evidence can be evaluated as data accumulate without the alpha inflation that afflicts repeated frequentist testing, and some methodologists argue the problem is specific to the frequentist framework rather than to the behaviour. The counter-argument is that the reported inference in almost every affected paper is frequentist.

Preregistration does not eliminate p-hacking. Registrations vary widely in specificity, deviations are common and inconsistently reported, and a vague registration can lend unearned authority to an analysis chosen after the fact. Preregistration converts a hidden degree of freedom into a visible one; it does not remove it.

Correction is not always required. Prespecified, hypothesis-driven comparisons in a confirmatory design do not automatically demand adjustment, and some statisticians argue that mechanical correction of every reported test is itself a distortion. The disagreement is real. What is not disputed is that the number of comparisons undertaken must be disclosed either way.

Not every result just below 0.05 is p-hacked. True effects of modest size produce p-values near the threshold routinely. A single paper reporting p = 0.043 is evidence of nothing. The signal is distributional across a literature, or comparative against a registration within a paper.

The American Statistical Association named the practice directly when it issued its 2016 statement on p-values, with then-president Jessica Utts observing that editorial bias toward significant results “leads to practices called by such names as ‘p-hacking’ and ‘data dredging’ that emphasize the search for small p-values over other statistical and scientific reasoning.” The statement’s fourth principle is the operative one: proper inference requires full reporting and transparency.

Related

Multiple comparisons · Statistics objections · Post-hoc power · Confounding · Pseudoreplication · The Table 2 fallacy

Checked before submission by trial registration and endpoint discipline, which compares the outcomes a manuscript reports against the outcomes its registration statement names, and flags analyses that appear only after the results.

Review my manuscript

Frequently asked questions

What does p-hacking mean?

P-hacking means selecting an analysis because of the result it produced. A researcher tries several defensible options — outcomes, covariates, exclusions, stopping points — and reports the one that crossed the significance threshold, without disclosing the others. The reported p-value is arithmetically correct and substantively meaningless, because the selection is not accounted for.

What is p-hacking in statistics?

In statistics, p-hacking is the inflation of the type I error rate caused by choosing an analysis conditional on its p-value. Under the null hypothesis, each additional analytic path is another chance for noise to cross 0.05. Four common flexibilities combined raise the false-positive rate from a nominal 5% to 60.7%.

Is p-hacking the same as data dredging?

P-hacking and data dredging describe the same underlying practice, and the American Statistical Association named both together when it issued its 2016 statement on p-values. Data dredging usually emphasises searching many variables for a pattern; p-hacking usually emphasises trying analytic variations on one hypothesis. The remedy for both is disclosure of everything examined.

What is a simple example of p-hacking?

A researcher collects 20 participants per group, finds p = 0.11, adds 10 more per group, and now finds p = 0.04 — reporting only the final sample. That single flexibility carries a false-positive rate of 7.7% rather than 5%. Repeated testing after every added observation raises it to roughly 22%.

What is p-hacking in a clinical trial?

In a clinical trial, p-hacking most often takes the form of outcome switching: the registered primary outcome shows no effect, so a secondary or newly defined outcome appears in the abstract instead. The COMPare project checked 67 trials in five major journals against their registrations and found 58 of them misreported, with prespecified outcomes missing and unregistered outcomes added.

How would you define p-hacking to a non-statistician?

P-hacking is asking a dataset the same question many different ways and reporting only the phrasing that gave the answer you wanted. Nothing in the arithmetic is wrong. What is missing is the count of how many phrasings were tried, without which a reader cannot judge whether the answer means anything.

Is p-hacking the same as cherry-picking results?

P-hacking is cherry-picking applied to analyses rather than to studies or findings. Cherry-picking can mean citing only supportive literature or reporting only the experiments that worked; p-hacking specifically means selecting among analyses of one dataset. All are forms of selective reporting, and all are addressed by disclosing the full set from which the reported item was chosen.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.