Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Publication bias arises when whether a study is published depends on what it found. Positive results appear more often and sooner, so the literature overstates effects.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Publication bias is the distortion that arises when whether a study reaches the literature depends on what it found. Trials with statistically significant or favourable results are published more often, sooner, and more prominently than those without, so the visible evidence base overstates effects. Meta-analyses built on that visible base inherit the overstatement and cannot detect it from the included studies alone.
Robert Rosenthal named the underlying phenomenon the “file drawer problem” in 1979: the studies that never appear are not a random sample of the studies that were run, and no amount of care in synthesising what did appear recovers what did not. This page covers the mechanism, the measured size of the effect in named literatures, funnel plots and the tests built on them, adjustment methods and their failure modes, how publication bias is detected while reading an actual manuscript, and what reviewers write when a paper handles it badly.
Publication bias is produced by a chain of decisions in which the study’s result influences whether the next step happens. Three gates matter, and they are not equally important.
The author gate is the largest. Surveys of investigators with unpublished work consistently find that the dominant reason given is the authors’ own judgement that a null result was not interesting or not worth writing up — not rejection by a journal. A null trial that is never drafted cannot be rejected, and it leaves no trace in any editorial record.
The editorial gate operates on perceived importance. A negative result is rarely rejected for being negative in so many words; it is rejected for lacking novelty, impact or a clear advance, which correlates with direction of effect. See why papers get rejected and desk rejected without review for how that judgement is actually worded.
Time-lag bias is a third gate that is often missed. Cohorts of trials followed forward from ethics approval — the design used by Stern and Simes and by the Cochrane methodology reviews that followed — consistently find that trials with statistically significant results reach publication substantially sooner than those without. The size of the gap varies by cohort and by field, and it is measured in years rather than months. A meta-analysis run at any given moment therefore sees a version of the literature that is skewed even if every trial is eventually published.
Publication bias concerns whether a study becomes available. Several related biases concern what happens to a study that is available, and conflating them leads people to apply the wrong remedy.
Outcome reporting bias occurs within a published paper: the trial appears, but the outcomes that reached statistical significance are reported fully while others are dropped, downgraded from primary to secondary, or described only as “not significant”. Chan and colleagues compared trial protocols with the resulting publications and found that roughly half of efficacy outcomes and around two-thirds of harm outcomes per trial were incompletely reported, and that significant outcomes were markedly more likely to be reported in full. Cross-checking against registered endpoints is the specific remedy, and it does not require a meta-analysis.
Small-study effects is the correct name for the pattern seen in a funnel plot: smaller studies giving systematically different results from larger ones. Publication bias is one cause. Genuine clinical heterogeneity — smaller trials recruiting higher-risk patients in whom the treatment works better — is another, as is poorer methodological quality in small studies, and chance.
Citation bias, language bias and duplicate publication bias affect which available studies a reviewer finds. Trials reporting positive results are cited more, are more likely to appear in English-language journals, and are more likely to be published more than once, so a search that stops at English-language indexed journals compounds the primary bias.
Selective analysis reporting overlaps with multiple comparisons: one analysis of many is reported, chosen after seeing which produced a small p-value.
Publication bias has been quantified directly wherever a complete registry of studies exists to compare against the published record. Four measurements are worth knowing by name.
Antidepressant trials, Turner and colleagues, New England Journal of Medicine, 2008. Using FDA registration records for 12 antidepressants and 12,564 patients, the authors identified 74 registered trials, of which 31% were never published. Separate meta-analyses of the FDA dataset and the published dataset differed by 32% in effect size overall, with per-drug inflation ranging from 11% to 69%. This is the single most-cited demonstration because the denominator — every trial the regulator saw — was knowable.
Preclinical stroke research, Sena and colleagues, 2010. In a survey of animal stroke studies, only about 2% of publications reported no significant effect, a rate no plausible experimental programme produces. The authors estimated that publication bias accounted for roughly a third of the reported efficacy. Preclinical literatures are especially exposed because registration is rare rather than absent: dedicated registries exist — preclinicaltrials.eu and the German Animal Study Registry — but uptake remains low, so in practice there is usually no denominator against which to name the missing studies; see preclinical review and in vivo rigour.
Registered Reports in psychology, Scheel, Schijen and Lakens, 2021. Comparing 71 published Registered Reports with 152 hypothesis-testing studies sampled from the standard literature, the authors found positive results in 96% of standard articles and 44% of Registered Reports — a 52-percentage-point gap. The comparison is against a random sample of the standard literature rather than against matched studies, so the two sets differ in more than the timing of the publication decision — but no other difference plausibly accounts for a gap that size, and the authors point to reduced publication bias and Type I error inflation as the likeliest explanation.
Cross-disciplinary trend, Fanelli, 2012. Sampling papers across disciplines and countries, Fanelli found that the frequency of papers reporting support for the tested hypothesis rose from roughly 70% at the start of the 1990s to roughly 86% by 2007, with the increase present in most disciplines. The number is not a bias estimate on its own — some rise is expected as fields learn which hypotheses are worth testing — but no field has a hypothesis-generation process good enough to justify a figure that high or a trend in that direction.
A funnel plot displays each study’s effect estimate on the horizontal axis against its precision — conventionally the standard error, plotted with zero at the top — on the vertical axis. Because small studies estimate the effect imprecisely, they scatter widely at the bottom; large studies cluster near the pooled estimate at the top. Under no bias the scatter is a symmetrical inverted funnel.
Asymmetry — a visible gap where small studies with small or negative effects should sit — is the classic signal. Contour-enhanced funnel plots, introduced by Peters and colleagues, add statistical-significance contours to the plot so that a reader can see whether the missing region is the non-significant region, which points at publication bias, or a region that has no particular significance status, which points at heterogeneity.
Three limitations are worth stating plainly. Funnel plot asymmetry is not a test for publication bias; it is a description of small-study effects with several possible causes. Visual assessment is unreliable, and published re-readings of the same plots by different assessors disagree. And a funnel plot with fewer than about ten studies conveys almost nothing — the scatter of a handful of points is indistinguishable from any shape you care to see in it.
Tests for funnel plot asymmetry ask whether effect size is associated with precision across the included studies. Four are in common use, and choosing between them is not arbitrary.
Egger’s regression test (Egger and colleagues, BMJ, 1997) regresses the standardised effect estimate on its precision and tests whether the intercept differs from zero. It is the default in most software.
Begg and Mazumdar’s rank correlation test (1994) correlates effect size with variance non-parametrically. It has lower power than Egger’s test and is now rarely preferred.
Harbord’s and Peters’ tests exist because Egger’s test misbehaves with binary outcomes: the log odds ratio and its standard error are mathematically correlated, which inflates the false-positive rate when events are rare or effects are large. For odds ratios, one of these variants should be used instead.
Two conventions govern interpretation. The Cochrane Handbook’s rule of thumb is that tests for funnel plot asymmetry should be used only when at least ten studies are included, because below that the power is too low to separate chance from real asymmetry. And the conventional threshold is p < 0.10, not 0.05, precisely because the tests are underpowered — a compensation that is often silently dropped in manuscripts reporting “no evidence of publication bias, p = 0.31”.
A non-significant asymmetry test in eight studies is not evidence of absence. It is an uninformative test that should not have been run.
Methods that attempt to correct a pooled estimate for publication bias exist, and every one of them rests on an assumption about the missing studies that the data cannot verify.
Trim and fill (Duval and Tweedie, 2000) removes asymmetric studies, estimates the number missing, imputes mirror-image studies and recomputes the pooled estimate. Simulation work — Terrin and colleagues among others — shows it performs poorly under between-study heterogeneity, where it can add studies that were never missing and shift an unbiased estimate. Trim and fill is best presented as a sensitivity analysis, never as a corrected result.
PET-PEESE (Stanley and Doucouliagos) uses the regression of effect on standard error to extrapolate to a hypothetical study of infinite precision. The extrapolation is far outside the observed data and is sensitive to the largest study in the set.
Selection models — Vevea and Hedges’s weight-function models, and Copas’s selection model — write down an explicit model for the probability that a study with a given p-value is published, then estimate the effect under that model. The honesty of the approach is its strength: the assumption is visible and can be varied. The weakness is that the selection function is not identified by the data and must be assumed.
p-curve (Simonsohn, Nelson and Simmons, 2014) examines the distribution of significant p-values, on the reasoning that a real effect produces right-skew towards small p-values while selective reporting of a null effect produces a flat or left-skewed curve. p-curve addresses evidential value across a set of significant findings, not the pooled magnitude, and it is sensitive to which results are entered.
The excess significance test (Ioannidis and Trikalinos) compares the number of significant findings observed with the number expected given the studies’ power. A large excess implies missing null results somewhere in the chain.
Agreement between two or more of these methods is more informative than any one of them. Disagreement is itself a reportable finding.
Detecting publication bias while reading a submitted manuscript is a document-comparison task before it is a statistical one. Six checks do most of the work, and four of them require no meta-analysis at all.
Compare the registry record with the paper. Retrieve the ClinicalTrials.gov, ISRCTN or PROSPERO entry named in the methods. Check the registered primary outcome against the reported primary outcome, the registered sample size against the analysed sample, and the registration date against the enrolment start date. A primary outcome that appears in the paper as a secondary outcome, or a registered outcome that has vanished, is the highest-yield single finding available.
Count the outcomes. A trial protocol listing nine outcomes and a paper reporting four, with all four significant, is selective reporting on its face regardless of what the discussion says.
Check the search strategy in a systematic review. A search restricted to PubMed, restricted to English, or with no attempt at trial registries, conference abstracts, regulatory documents or contact with authors, has a known direction of error. The review’s own inclusion flow chart usually reveals this in one line.
Look at the number of included studies before believing any asymmetry test. If the meta-analysis pools seven studies and reports Egger’s test, the test is uninformative, and the manuscript’s conclusion drawn from it is unsupported.
Read the funnel plot rather than the sentence about it. Manuscripts routinely include a plot showing visible asymmetry and a sentence stating that the plot is symmetrical.
Check whether harms are reported at the same resolution as benefits. Asymmetry between a fully specified efficacy analysis and a one-paragraph safety summary is a reporting-bias signal within a single paper.
PerfectPaper performs the registry-to-manuscript comparison and the outcome census as part of a structured review of clinical and epidemiological manuscripts, reporting what it found in the registry entry alongside the claim in the text rather than asserting misconduct. See clinical trial review and epidemiology review for how those checks are scoped.
Reviewers rarely write the phrase “publication bias” when they mean it. The comments below are composed examples in the register these comments arrive in, grouped by what the reviewer has noticed.
On an underpowered asymmetry test: “The authors report no evidence of publication bias on the basis of Egger’s test with eight included studies. This test has negligible power at that number and cannot support the conclusion drawn.”
On an absent assessment: “No assessment of risk of bias due to missing results is presented. Please address this domain explicitly, per PRISMA 2020.”
On trim and fill presented as a correction: “The trim-and-fill adjusted estimate is presented alongside the primary estimate without qualification. Trim and fill is unreliable under the level of heterogeneity reported here (I² = 71%) and should be described as a sensitivity analysis.”
On the search: “The search was limited to English-language publications indexed in MEDLINE. Given the known association between direction of effect and language of publication, this restriction should be justified or the limitation stated.”
On a registry discrepancy: “The registered primary outcome (ClinicalTrials.gov NCT…) differs from the primary outcome reported here. Please explain when and why the outcome was changed.”
On conclusions drawn past the evidence: “Small-study effects are evident in Figure 3, yet the discussion treats the pooled estimate as unbiased.”
On the discussion section: “The limitations paragraph acknowledges publication bias in general terms without stating its likely direction or magnitude for this synthesis.”
That last comment is the most common and the most easily pre-empted, and it is the same failure mode as a generic limitations paragraph anywhere else — see overclaim checking.
Publication bias is a structural problem, and the interventions that measurably change it operate on structures rather than on individual authors’ diligence.
Prospective registration is the precondition for everything else, because it creates a denominator. The ICMJE has required registration before enrolment of the first participant as a condition of consideration since 2005, and registration is what makes the registry-versus-paper comparison above possible.
Registered Reports move the accept decision to before data collection, on the basis of the question and the method. The 96%-versus-44% gap measured by Scheel and colleagues is the clearest available evidence that much of the excess of positive findings comes from the publication process rather than from nature — the authors themselves put it as the most plausible explanation for a gap that large, not as a demonstrated cause.
Mandatory results posting. The FDA Amendments Act of 2007 requires summary results for applicable trials to be posted to ClinicalTrials.gov, generally within twelve months of primary completion. Compliance audits have repeatedly found substantial non-posting, and the published estimates vary with the cohort and the definition of an applicable trial, so quote a specific rate only alongside the cohort it came from.
Preprints and data sharing shorten the time-lag gate and make unreported outcomes recoverable; see data availability and accessions.
Formal assessment tools. ROB-ME, the Cochrane tool for risk of bias due to missing evidence, replaces the informal funnel-plot ritual with signalling questions: how many studies are known to be missing despite having measured the outcome, whether the search was comprehensive, and whether statistical or graphical methods suggest results are missing because of what they showed. GRADE treats publication bias as a domain that can lower certainty by one or two levels independently of the other domains.
State the assessment as a specification, not as a courtesy. A sentence that satisfies most reviewers has four parts: what you did to find unpublished work, how many studies you pooled and therefore which asymmetry methods were appropriate, what the methods showed, and in which direction and by roughly how much a plausible amount of missing evidence would move your conclusion.
Report the funnel plot when you have ten or more studies and say what the contours show. Report the asymmetry test with the threshold you used and why. Present any adjusted estimate as a sensitivity analysis with its assumption named. If you pooled fewer than ten studies, write that tests for funnel plot asymmetry were not performed because the number of studies was insufficient — that sentence is stronger than an uninformative p-value, and reviewers read it that way.
For a primary study rather than a synthesis, the equivalent commitment is a registered protocol, a reported outcome list that matches it, and a results section that reports the pre-specified analysis whether or not it worked. A null primary outcome reported cleanly is publishable work; see when your effect size is called not meaningful and replication objections for how to frame it.
Publication bias cannot be measured from the published literature. Every method described on this page infers something about studies that are absent from the properties of studies that are present, and that inference is not identified by the data — it is identified by an assumption about the selection process. Two analysts applying different, equally defensible assumptions to the same meta-analysis can reach different adjusted estimates.
Funnel plot asymmetry has at least four causes and the plot does not distinguish them. Trim and fill is unreliable under heterogeneity, which is present in most clinical meta-analyses. Asymmetry tests are underpowered where they are most needed. Registries close the loop only for registered study types, which excludes most preclinical, observational and qualitative research entirely.
The honest position is that publication bias is a risk to be characterised and bounded, not a quantity to be measured and subtracted. Reviewers respond well to that framing and badly to an adjusted number presented as a correction.
Confounding · Collider bias · Multiple comparisons objections · Sample size objections · Journal selection
Checked before submission by trial registration and endpoint review, which compares the registered outcomes against the reported ones and names each discrepancy with its registry identifier.
Publication bias means that whether a study appears in the literature depends on what the study found. Statistically significant and favourable results are written up, submitted and accepted more often and sooner, so the published record is a biased sample of the research actually conducted.
Publication bias in research is the systematic difference between the studies that were conducted and the studies a reader can find. Because null and unfavourable results are disproportionately absent, pooled estimates from published studies overstate effects, and no analysis of the included studies alone can recover the missing ones.
The file drawer problem is another name for publication bias, coined by Robert Rosenthal in 1979. The image is of null results sitting in investigators’ filing cabinets, never written up. Author decisions, not journal rejections, are the largest contributor to that drawer.
Publication bias inflates the pooled effect in a meta-analysis, because the studies available to include are skewed towards larger and more favourable estimates. Turner and colleagues found a 32% difference in effect size between the FDA-registered and the published antidepressant trial data. Thirty-one percent of the registered trials were never published, and 11 more were published in a way the authors judged to convey a positive outcome; both mechanisms feed the gap.
Publication bias happens because the result influences three sequential decisions: whether the authors draft the paper, whether editors and reviewers judge it important, and how quickly it moves through the process. Investigator surveys consistently identify the authors’ own judgement that a null finding is uninteresting as the dominant cause.
Publication bias is detected by comparing the registered study record with the published paper, auditing the search strategy of a review, inspecting a funnel plot when ten or more studies are pooled, and running an asymmetry test such as Egger’s regression at p < 0.10. No method identifies it with certainty.
Publication bias cannot be corrected, only bounded. Trim and fill, PET-PEESE and selection models each estimate an adjusted effect under an assumption about the missing studies that the data cannot verify. Present any adjusted estimate as a sensitivity analysis, never as a corrected result.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect