Skip to content

SOLUTIONS

What is a confidence interval?

A confidence interval is a range produced by a method whose coverage holds over repeated studies. The 95% describes that method, not the single interval reported.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is a confidence interval?

A confidence interval is a range of values, computed from data by a procedure whose coverage is guaranteed in the long run: a 95% confidence interval comes from a method that captures the true parameter in 95% of repeated studies. The interval shows which effect sizes are compatible with your data, and how precisely the effect was estimated.

That definition places the 95% on the procedure rather than on the particular interval in front of you, and almost every reporting error in the literature comes from moving it. This page covers what the coverage statement actually promises, how an interval is built, the five misinterpretations named in the methodological literature, when the standard formula returns the wrong interval, the arithmetic checks that expose a mishandled interval in a manuscript, and what reviewers write when they find one.

What the 95% refers to

The 95% in a 95% confidence interval is a property of the method, assessed over hypothetical repetitions of the study, not a probability attached to the single interval you computed. Run the same design and the same interval formula 100 times on data generated by the same underlying truth, and roughly 95 of the intervals produced will contain that truth. Which 95 is unknowable, and the one you have is either right or wrong.

This is the frequentist coverage guarantee, and it is deliberately modest. It says nothing about the parameter being random; the parameter is fixed and the interval is what varies from study to study. A Bayesian credible interval makes the statement researchers usually want — a 95% posterior probability that the parameter lies in the stated range — but it requires a prior distribution, and it is a different object with different assumptions. Reporting a frequentist interval and interpreting it as a credible interval is the most common conceptual slip in applied work, and it is the first of the misinterpretations catalogued below.

How a confidence interval is constructed

A confidence interval is constructed by taking a point estimate, attaching a measure of its sampling variability, and extending outward far enough that the resulting range covers the truth at the stated rate. In its most common form — the Wald interval — the recipe is estimate ± (critical value × standard error), with the critical value 1.96 for a 95% interval under a normal approximation, or the corresponding t quantile when the standard error is itself estimated from a small sample.

Three ingredients therefore determine the interval: the estimate, the standard error, and the reference distribution. Each can be wrong independently. A correct estimate with an underestimated standard error yields an interval that is too narrow and a coverage rate below 95%. A correct standard error referred to a normal distribution when a t distribution with 8 degrees of freedom was required yields the same failure in a smaller way — 1.96 instead of 2.31, an interval about 15% too narrow.

Not every interval is built this way. Score-based intervals invert a hypothesis test rather than adding a margin to an estimate. Profile-likelihood intervals trace the likelihood function and need not be symmetric. Bootstrap intervals resample the data and read percentiles off the resulting distribution. These alternatives exist because the Wald recipe fails in identifiable situations, described further down this page.

What a confidence interval tells you that a p value does not

A confidence interval carries the magnitude and the units of the effect, and a p value carries neither. “Mortality was lower in the treated group (p = 0.03)” tells a reader that some difference was detected; “absolute risk difference −4.2 percentage points, 95% CI −7.9 to −0.5” tells the reader that the data are compatible with anything from a trivial half-point reduction to a substantial eight-point one, which is a different clinical conclusion at each end.

The two are arithmetically linked for the same test on the same scale: a 95% Wald interval excludes the null value precisely when the corresponding two-sided p value is below 0.05. That correspondence is why an interval is strictly more informative — it contains the test result and adds the effect size and its precision. The International Committee of Medical Journal Editors states the preference directly in its Recommendations, under Statistical methods: “When possible, quantify findings and present them with appropriate indicators of measurement error or uncertainty (such as confidence intervals)”, followed by the instruction to avoid relying solely on statistical hypothesis testing, such as P values, which fails to convey important information about effect size and precision of estimates.

The correspondence is not universal. Fisher’s exact test paired with a Wald interval, a score interval paired with a Wald test, and any other mismatch between the interval method and the test method can put a p value of 0.04 alongside an interval that includes the null. That combination is not automatically an error, but it obliges the methods section to name both procedures rather than only one.

The five misinterpretations named in the literature

Greenland and colleagues catalogued 25 numbered misinterpretations of statistical tests, p values, confidence intervals and power in the European Journal of Epidemiology in 2016 (31:337–350). Five concern confidence intervals directly, and each is stated there as a false proposition:

    1. “The specific 95% confidence interval presented by a study has a 95% chance of containing the true effect size.”
    1. “An effect size outside the 95% confidence interval has been refuted (or excluded) by the data.”
    1. “If two confidence intervals overlap, the difference between two estimates or studies is not significant.”
    1. “An observed 95% confidence interval predicts that 95% of the estimates from future studies will fall inside the observed interval.”
    1. “If one 95% confidence interval includes the null value and another excludes that value, the interval excluding the null is the more precise one.”

Number 21 has a usable quantitative correction. Cumming and Finch showed in American Psychologist (2005) that two independent 95% intervals can overlap by roughly half the average margin of error and still correspond to a p value near 0.05 — so “the error bars touch, therefore the groups do not differ” is wrong in the direction that costs findings, not the direction that manufactures them. Number 22 matters most in meta-analysis, where the confidence interval around a pooled estimate answers a different question from the prediction interval for a future study, and the prediction interval is usually far wider.

Width, sample size and precision

The width of a confidence interval is governed by the standard error, which for an ordinary mean shrinks with the square root of the sample size. Quadrupling n halves the interval width; doubling n narrows it by about 29%. This is the arithmetic behind every underpowered study’s wide interval, and it is why a proposal to “add a few more animals” rarely rescues an inconclusive result.

Two rules of thumb are worth carrying. First, the Cochrane reconstruction: for a large-sample 95% Wald interval, the standard error is (upper limit − lower limit) ÷ 3.92, which lets you recover the standard error from any published interval and check the reported p value against it. Second, the rule of three, from Hanley and Lippman-Hand in JAMA (1983): if an adverse event occurred zero times in n patients, the upper limit of the 95% interval for its rate is approximately 3/n. Zero events in 30 patients is compatible with a true rate as high as 10%, which is the honest reading of most small safety cohorts.

A wide interval is a statement about precision, not a failure of the analysis, and it should be reported as such rather than retrofitted with a power calculation after the fact — see what to do when a reviewer asks for post-hoc power and when a reviewer says the sample size is too small.

When the standard formula returns the wrong interval

The Wald interval is the default in most software and it fails in several well-characterised situations.

Proportions near 0 or 1. Brown, Cai and DasGupta showed in Statistical Science (2001; 16:101–133) that the Wald interval for a binomial proportion has chaotic coverage that does not settle down even at large n, and can fall far below the nominal 95%. They recommend the Wilson score interval or the Jeffreys interval for small samples and Agresti–Coull for larger ones. A published proportion whose interval extends below 0% or above 100% is a Wald interval that should have been a Wilson one.

Ratio measures. Odds ratios, risk ratios and hazard ratios are estimated on the log scale, so their intervals are symmetric in logs and asymmetric when printed. An odds ratio of 1.50 reported as 95% CI 1.20 to 1.80 is arithmetically suspicious, because the log-scale midpoint of 1.20 and 1.80 is √(1.20 × 1.80) = 1.47, not 1.50.

Sparse data and separation. In logistic regression with few events per covariate, the Wald standard error inflates as the coefficient grows, and the interval can widen absurdly or cover the null while the likelihood ratio test rejects it. Profile-likelihood intervals, or Firth’s penalised likelihood, behave correctly where the Wald interval does not.

Correlated observations. Any interval computed as though observations are independent when they are not is too narrow. Repeated measurements from the same animal, multiple cells from one culture, and patients clustered within centres all inflate the effective replicate count and shrink the interval accordingly — the mechanism described in pseudoreplication and objected to as n is not independent. Cluster-robust standard errors fix this only when the number of clusters is large; Cameron and Miller’s practitioner’s guide in the Journal of Human Resources (2015) sets out the small-cluster corrections, including the wild cluster bootstrap and referring the statistic to a t distribution with G−1 degrees of freedom.

Non-normal or skewed statistics. Bootstrap intervals help here, but the plain percentile bootstrap is biased when the statistic’s sampling distribution is skewed; the bias-corrected and accelerated (BCa) variant exists for that reason and should be named in the methods when used.

Confidence intervals quantify random error only

A confidence interval quantifies sampling variability and nothing else. It contains no allowance for confounding, selection, measurement error, protocol deviation, or the analytic choices made before the interval was computed. A 95% interval from a badly confounded observational study is a precise statement about the wrong quantity, and increasing the sample size narrows it around the same bias.

This is worth stating explicitly in a manuscript because narrow intervals read as authoritative. A hazard ratio of 1.8 with a 95% CI of 1.6 to 2.0 in a cohort with unmeasured confounding is not stronger evidence of causation than the same estimate with a wider interval; it is a better-estimated association of unknown causal status. Where the analysis conditioned on a variable caused by both exposure and outcome, collider bias shifts the interval bodily and the coverage guarantee says nothing about it. The same caution applies to intervals printed beside covariates in a regression table, which is the Table 2 fallacy.

Bias analysis is the honest complement. An E-value, a simple bias parameter, or a quantitative sensitivity analysis extends the interval to reflect systematic error, and doing so changes many “significant” findings into ranges that include the null.

How mishandled intervals are detected in a manuscript

Confidence-interval errors are detectable by arithmetic on the reported numbers, which is what makes them one of the more tractable statistical review tasks. The checks below run on a results table alone, with no access to the data.

The point estimate must sit inside the interval, and for a log-scale Wald interval it must equal the geometric mean of the limits: OR = √(lower × upper). A discrepancy beyond rounding means the estimate and the interval came from different models, which is usually a copy error between analysis versions.

The interval must agree with the p value. Recover the standard error as (upper − lower) ÷ 3.92, form z = estimate ÷ SE, and compare against the reported p. An interval excluding the null beside p = 0.09, or an interval spanning the null beside p = 0.01, is an inconsistency requiring an explanation.

Interval limits must respect the parameter’s range. Proportions below 0 or above 1, correlation limits outside −1 to 1 (which happens when a Fisher z transformation was skipped), and negative lower bounds on inherently positive quantities all indicate the wrong interval method.

Interval widths must scale with group size. A subgroup of 12 patients whose interval is as tight as the interval for the 400-patient full cohort indicates the wrong denominator, an interval computed on the wrong scale, or clustered observations counted as independent.

Within-group intervals must not be used to infer a between-group difference. Reporting that the treated group improved (95% CI excluding zero) while the control group did not (95% CI including zero) is not a comparison; the interval for the difference has to be computed directly, and it is frequently much less impressive.

Baseline-to-follow-up intervals within randomised arms deserve particular scepticism, because a within-arm change interval answers a question randomisation was not designed to answer.

Four of the six are closed-form tests on numbers already present in the table, so an automated check can run them directly; the last two require reading what the text claims the intervals show. The harder judgement is the one no arithmetic reaches: whether the prose describing an interval is more confident than its limits allow, the pattern handled by the overclaim check.

What a peer reviewer says when confidence intervals are mishandled

Reviewers rarely write “your confidence interval is wrong.” The phrasings below are typical of how the objection is worded in review reports, and each one names a specific defect described above.

“Effect estimates should be presented with confidence intervals rather than p values alone.” “The reported confidence interval is inconsistent with the p value given in the same row.” “The authors interpret a non-significant result as evidence of no effect; the confidence interval extends to a clinically important benefit.” “Overlapping error bars are not a valid test of the difference between groups; please report the interval for the difference.” “The intervals appear to assume independence between repeated measurements from the same subject.” “Confidence intervals for proportions near the boundary should use an exact or score-based method.” “The precision implied by these intervals is not credible given the number of events.”

Two of these decide papers. The first is the “absence of evidence” objection — Altman and Bland’s BMJ note of 1995 remains the standard citation — where an interval running from 0.65 to 1.05 is described as showing no effect, when it is equally compatible with a 35% reduction. The second is the independence objection, which cannot be answered by re-analysis in a revision if the design itself lacked independent replicates; see my replicates disagree and when a reviewer says the wrong statistical test was used.

How to report a confidence interval

Report the interval level, both limits, the estimate, and the units, in that sentence. “95% CI” rather than “CI”; both limits separated by “to” rather than a hyphen, because a hyphen is ambiguous when a limit is negative; and the same number of decimal places on the estimate and on the limits.

The reporting guidelines are explicit. CONSORT 2025, the updated 30-item checklist for randomised trials, requires that results for each outcome be given for each group with the estimated effect size and its precision, such as a 95% confidence interval; for binary outcomes it asks for both absolute and relative effect sizes. STROBE item 16a requires unadjusted estimates and, where applicable, confounder-adjusted estimates with their precision, so that readers can see how far and in which direction adjustment moved the estimate.

Interpret both limits in the discussion, not just whether the interval crosses the null. State what the lower limit would mean if it were the truth, and what the upper limit would mean, and whether either would change practice — that is the reading that distinguishes an inconclusive study from a genuinely null one, and it is closely related to whether the effect size is meaningful. Where many intervals are reported, say how many and whether any adjustment was made; twenty independent 95% intervals will exclude the null about once by chance, which is the interval-scale version of the multiple comparisons problem.

Contested: should they be called compatibility intervals

A significant body of methodological opinion holds that the word “confidence” is itself the problem. Amrhein, Greenland and McShane argued in Nature in March 2019 (567:305–307), in a comment supported by more than 800 signatories, that authors should interpret estimates rather than tests and explicitly discuss the lower and upper limits of what they term compatibility intervals — the interval containing the parameter values most compatible with the data under the model, with no implied confidence in any one of them.

The relabelling is not universally accepted, and no major reporting guideline adopts it: ICMJE, CONSORT 2025 and STROBE all still say “confidence interval”. Its practical value is that it forecloses misinterpretation 19: a reader who sees “values compatible with these data” is less likely to read a probability statement into it than one who sees “we are 95% confident.” Its limitation is that compatibility is still conditional on the entire model — the distributional assumptions, the adjustment set, and the absence of systematic error — so the renamed interval carries exactly the same blind spot regarding bias described above. Using the term without explaining it in a manuscript submitted to a clinical journal will often prompt a query rather than approval, so if you adopt it, define it at first use and report the level as usual.

Related

Effect size not meaningful · Sample size too small · Post-hoc power · Multiple comparisons · Pseudoreplication · Confounding

Checked before submission by statistical objections, which recomputes the standard error implied by each reported interval and flags limits that disagree with the accompanying p value or with the prose describing them.

Review my manuscript

Frequently asked questions

What does a 95% confidence interval mean?

A 95% confidence interval is a range produced by a method that captures the true parameter in 95% of repeated studies. The 95% describes the long-run performance of the procedure, not the probability that this particular interval contains the truth. The interval you have either contains the true value or does not.

What is a confidence interval in statistics?

A confidence interval in statistics is an estimated range for an unknown parameter, built from a point estimate and its standard error, carrying a stated coverage rate. Common forms include the Wald interval (estimate ± 1.96 standard errors for 95%), the Wilson score interval for proportions, and bootstrap intervals for non-normal statistics.

How do you interpret a confidence interval in a research paper?

Interpret both limits, not merely whether the interval crosses the null. A 95% interval of −7.9 to −0.5 percentage points is compatible with a trivial reduction and with a substantial one, so the study is imprecise rather than conclusive. State what each limit would mean clinically, and report the interval alongside the estimate and its units.

What is the difference between a confidence interval and a p value?

A confidence interval carries the effect size, its units and its precision; a p value carries none of these. For the same test on the same scale, a 95% interval excludes the null exactly when the two-sided p value is below 0.05, so the interval contains the test result and adds information the p value discards.

Can you explain a confidence interval in plain language?

A confidence interval is the set of values for an effect that your data do not rule out, given your model. Narrow means the study pinned the effect down; wide means it did not. The stated level, usually 95%, refers to how often the method gets it right across many studies, not to this one result.

What does the 95% in a 95% confidence interval refer to?

The 95% refers to the coverage of the procedure over hypothetical repetitions of the study: repeat the design and the interval formula many times, and about 95% of the intervals generated will contain the true parameter. It is not a 95% probability that the true value lies inside the single interval reported.

Why is a confidence interval preferred to reporting significance alone?

A confidence interval shows magnitude and precision, which a significance verdict discards. The ICMJE Recommendations ask authors to quantify findings with indicators of uncertainty such as confidence intervals and to avoid relying solely on hypothesis testing, and CONSORT 2025 and STROBE both require effect estimates with their precision in results tables.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.