Skip to content

SOLUTIONS

Reviewer says the effect is not clinically meaningful

Statistical significance is not magnitude. This objection is about whether the size of your effect matters, and more data makes it worse rather than better.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Reviewer says the effect is not clinically meaningful

“The effect is not clinically meaningful” is the one statistical objection where a larger study makes your position worse. The reviewer is not disputing that the difference is real — they are saying that a difference this size does not matter, and a bigger sample would only establish a trivial effect more precisely.

The clinical-meaningfulness objection is also the one most often answered badly, because the instinct is to defend the p-value, and the p-value is not what is being questioned. This page covers what the reviewer means, the arithmetic that makes more data counterproductive, the three defences that work, what to do when no published threshold exists for your outcome, how the problem is detected in a submitted manuscript, and the sentences reviewers actually write.

What the reviewer means

Significance and magnitude are independent. With enough observations, a difference of 0.3 points on a 100-point scale reaches p < 0.001 while remaining invisible to any patient. The reviewer has looked at your effect size and concluded it falls below the threshold where anyone would act differently.

Common phrasings: “the difference, while statistically significant, is small”; “is a 2 mmHg reduction clinically relevant?”; “the confidence interval includes values of no practical importance”; “this is unlikely to change practice”.

What every one of those sentences shares is an unstated threshold. The reviewer is comparing your estimate against a number they have in their head — from their own practice, from a guideline, from a published minimal clinically important difference for the instrument — and has not written it down. Your first move is to make that comparison explicit, because a threshold on the page is arguable and a threshold in a reviewer’s head is not. Name the quantity, name the threshold you are measuring it against, and name where the threshold came from.

The objection is not a complaint about statistics. It is a complaint about the sentence in your abstract, and it survives every reanalysis you could perform on the same data.

Why a bigger study makes this objection stronger

A larger sample moves the p-value and leaves the effect where it was, which is precisely the reviewer’s point. The arithmetic is worth doing once, because seeing it kills the instinct to promise more data in a response letter.

For a two-group comparison of a continuous outcome, the sample size per arm at which the expected result sits exactly on a given significance threshold is n = 2(z/d)², where d is the standardised difference. (That is the 50%-power boundary; a conventionally powered study needs more, which only sharpens the point.) Take the 0.3-point difference above on an instrument with a standard deviation of 10 points: d = 0.03. Reaching p < 0.05 two-sided (z = 1.96) takes about 8,500 participants per arm. Reaching p < 0.001 (z = 3.29) takes about 24,000 per arm. At 48,000 participants the difference is still 0.3 points, and the 95% confidence interval has tightened around a number nobody would act on.

That is the trap in this objection and the reason it differs from every other statistical complaint in the family. When a reviewer says your sample size is too small, more data is sometimes the answer. When a reviewer says the effect is not meaningful, more data makes the sceptical reading unavoidable: a precisely estimated trivial effect is a stronger demonstration of triviality than an imprecise one. The correct currency is the magnitude and its interval, which is what an effect size reports and a p-value does not.

Which version of the objection did you receive?

Five distinct objections travel under the phrase “not clinically meaningful”, and they take five different revisions. Diagnose from the reviewer’s exact wording before drafting, because the material for each is unrelated to the others.

What the reviewer wrote Which objection it is What answers it
“The difference, while statistically significant, is small” Point estimate below an unstated threshold A named external threshold your effect exceeds, or a concession and a reframe
“The confidence interval includes values of no practical importance” Imprecision, not magnitude Both interval limits read against the threshold, with what each would mean
“Is a 2 mmHg reduction clinically relevant?” The unit itself is not interpretable to the reader What one unit means in practice, plus a population-scale or benchmark argument
“This is unlikely to change practice” The endpoint is not one clinicians act on An endpoint argument, or a narrowed claim about mechanism
“The abstract calls the improvement marked; Table 2 reports 2 points on a 100-point scale” Overclaim, not magnitude A rewritten adjective in the abstract and discussion

The last row is the most common of the five in submitted work, and the cheapest to fix. Nothing is wrong with the analysis; the discussion has described a modest number in language the results do not support. Fixing the sentence removes the objection entirely, which is why the check belongs before submission rather than after.

The three legitimate responses

Argue the magnitude matters, with an external benchmark. The strongest version cites a minimal clinically important difference established elsewhere for your outcome and population. If your effect exceeds it, say so and cite it. Many patient-reported outcome instruments have published thresholds; using one turns your defence from assertion into evidence.

Argue that a small effect matters at scale or in aggregate. A small shift in a common exposure can outweigh a large shift in a rare one. This is legitimate for population-level interventions and it needs to be made explicitly, with the population-level arithmetic shown, rather than implied.

Concede and reframe. If the effect genuinely is small and no benchmark supports its importance, say so and reposition the contribution: as evidence about mechanism, as a bound on the effect’s size, or as a negative result worth recording. A precisely estimated small effect is a real contribution when presented as one.

Each of the three has conditions it must meet to survive a second round, and the next three sections give them.

Benchmarking against a published threshold

A published threshold is the strongest answer available, and it holds only if the threshold was derived for your instrument and version, in a comparable population, in the same direction of change, over a comparable follow-up interval, and from data other than the data you are defending. A borrowed number that fails any one of those is an invitation for the reviewer to write a second, sharper comment.

The construct has a definition. Jaeschke, Singer and Guyatt introduced it in Controlled Clinical Trials in 1989 (volume 10, pages 407–415) as “the smallest difference in score in the domain of interest which patients perceive as beneficial and which would mandate, in the absence of troublesome side effects …, a change in the patient’s management.” Note what that wording contains: patients, a domain, and a management consequence. A threshold derived without asking a patient anything is measuring something else. What a minimal clinically important difference is covers the derivation methods in full.

Three checks separate a defensible citation from a decorative one.

Check that the published estimate is not one value out of a wide range. Copay and colleagues applied several standard anchor-based and distribution-based calculations to a single cohort of lumbar spine surgery patients (The Spine Journal, 2008) and found that the candidate thresholds a single dataset yields differ substantially depending on which method is used. The widely cited 12.8-point value for the Oswestry Disability Index is one member of that set, not a constant of nature, and quoting it without acknowledging the spread asserts more precision than the literature contains.

Check that the derivation study reports an anchor–instrument correlation and an interval. Devji and colleagues published a credibility instrument for anchor-based thresholds in The BMJ in 2020 (369:m1714), whose five core criteria require that the anchor be rated by the patient, be interpretable and relevant to the patient, correlate satisfactorily with the instrument, that the estimate be precise, and that the chosen anchor category reflect a small but important difference. A derivation paper reporting none of the correlation and none of the interval cannot support a claim of clinical importance downstream.

Check that you are comparing like with like. A within-person minimal important change describes what one patient must gain to notice a benefit; your trial reports a between-group difference in mean change. They sit on the same scale and are not the same quantity, and placing one beside the other is the single most common way this defence collapses in review. Where the threshold is well established the comparison is usually clean: the 2014 European Respiratory Society and American Thoracic Society technical standard puts the minimal important difference for six-minute walk distance at 30 m (95% CI 25 to 33 m), so a significant 12 m gain is a real but sub-threshold change and the paper should say so in those words.

Arguing that a small effect matters at scale

A small per-person effect can carry large population consequences, and the argument is legitimate when the exposure is common, the intervention is deliverable at population scale, and you show the arithmetic instead of gesturing at it. Geoffrey Rose set out the reasoning in “Sick individuals and sick populations” (International Journal of Epidemiology 1985;14:32–38), having laid the groundwork in BMJ 1981;282:1847–1851: a measure that shifts a whole distribution slightly can prevent more events than one that treats the small tail intensively, while offering almost nothing to any individual who takes part.

Show it in numbers. An absolute risk reduction of 0.2 percentage points is a number needed to treat of 500 — unimpressive at the bedside. Applied across 5 million exposed people, it is 10,000 events prevented. Both descriptions are true, and the manuscript that only asserts “these findings have important public health implications” has given the reviewer nothing to check.

The argument has three failure modes, and reviewers know all of them. It fails when the intervention is not actually deliverable at population scale — a specialist procedure delivered to a few thousand patients per year does not get to borrow population arithmetic. It fails when the estimate is not transportable: a per-person effect measured in a selected cohort and multiplied by a national denominator assumes the effect is the same in people who look nothing like the study sample. And it fails when the effect is not causal, because scaling an association multiplies the confounding along with the effect. Where the manuscript’s verbs already outrun its design, this defence makes the causal-language problem worse rather than better.

State the denominator, the time horizon and the assumed uptake explicitly. “Applied to the 5 million adults in this age band, at the observed adherence of 61%, over five years” is a claim a reviewer can dispute on its inputs. “At population level, this effect is substantial” is not.

Conceding the magnitude and reframing the contribution

Conceding is a strategy, not a surrender, and it is the right one whenever no benchmark supports the effect and the interval is narrow. A precisely estimated small effect answers a question the field had open: how big this effect is not. Say that, and the contribution becomes the precision rather than the direction.

Three reframes work in practice. The bound. “This trial excludes benefits larger than 3.2 points, well below the 10-point threshold established for this instrument” is a finding, and it is more useful to the next investigator than a significance claim. The mechanism. A small but reliably detected effect can establish that a pathway operates in humans even when the magnitude is clinically inert, provided the paper claims exactly that and no more. The formal equivalence frame. Where the point of the study is that two things do not differ importantly, say so with a pre-specified margin rather than a non-significant p-value: two one-sided tests, or a full non-inferiority design, with the margin justified against a clinical threshold rather than chosen for convenience.

The reframe has to reach the abstract, the discussion and the title, not just the response letter. Reviewers read the revised manuscript against the letter, and a paper whose letter concedes the magnitude while its abstract keeps the original adjective converts a survivable comment into evidence that the authors were not listening. Write the concession as an analysis: a limitations section that names the boundary carries far more weight than one saying results should be interpreted with caution.

Worked example: one 6-point difference, three defensible papers

Consider three manuscripts reporting the same 6.0-point between-group difference on a 0–100 patient-reported instrument with a standard deviation of 12 points, against a published within-person threshold of 10 points. The number is identical in all three; the defensible claim is different in each.

The large, precise trial. With 550 participants per arm the difference is 6.0 points, 95% CI 4.6 to 7.4, p < 0.001. The entire interval sits below the 10-point threshold. No benchmark argument is available, and more participants would only shrink the interval further inside the region the reviewer is objecting to. The correct revision concedes and bounds: the trial has established that the intervention’s effect on this instrument does not reach the threshold, which is a publishable and useful result stated that way.

The small, imprecise trial. With 25 participants per arm the same 6.0-point difference carries a 95% CI of roughly −0.7 to 12.7. The data are compatible with no effect and with an effect above the threshold simultaneously, so the honest reading interprets both limits and claims neither. What must not appear in the revision is an observed-power calculation, which is a deterministic function of the p-value and adds nothing — see reviewer asked for a post hoc power calculation. The confidence interval is the object to discuss.

The distribution argument. A between-group mean of 6 points below a within-person threshold of 10 does not establish that few patients benefited, because a treatment shifts a distribution rather than moving everyone by the mean. The defensible move is a responder analysis with the threshold pre-specified and derived independently, reported across a range — at 5, 10 and 20 points — so the reader can see whether the conclusion turns on the cut-point. Snapinn and Jiang set out what dichotomising sacrifices in Trials in 2007 (8:31): the approach gives up statistical efficiency, and a proportion above an arbitrary cut-point is not itself a measure of clinical relevance. Report it alongside the continuous result, never instead of it.

What not to do

Do not answer with the p-value. “The difference was highly significant (p < 0.001)” restates what the reviewer already accepted and signals you have missed the point.

Do not collect more data. More observations narrow the interval around the same small effect. That strengthens the reviewer’s argument, not yours.

Do not switch to a relative measure to make it look bigger without also reporting the absolute one. A relative risk reduction of 50% on a baseline risk of 0.4% is an absolute reduction of 0.2 percentage points. Reporting only the relative figure when the absolute is small is a known presentational problem and reviewers who are looking for it will find it. The same result stated as a number needed to treat of 500 makes the magnitude unmissable, which is exactly why relative-only reporting attracts the objection.

Do not derive the threshold from the data it judges. Estimating a minimal important difference in your own cohort, with your own anchor question, and then using it to declare your own result meaningful is circular, and it is easy to miss because it looks like diligence. The threshold must come from elsewhere, be pre-specified, and appear in the methods rather than surfacing for the first time in the results.

Do not go looking for a subgroup where the effect is larger. A subgroup selected after seeing the data produces an inflated estimate by construction, and a reviewer who was already unconvinced by the primary magnitude will read the subgroup as confirmation. A selected estimate overstates the truth by construction, which is the same mechanism that inflates p-hacked effects.

When no published threshold exists for your outcome

No validated threshold exists for most outcomes, and saying so plainly is a stronger position than borrowing an unrelated one. Five moves substitute for the benchmark, and they can be used together.

State what one unit means in the world. Metres walked, exacerbations avoided per patient-year, mmHg, days of hospitalisation, steps per day. A reader who cannot convert your unit into an experience cannot judge your magnitude, and the conversion is your job rather than theirs.

Use a distribution-based figure as a floor, never as a definition of importance. Norman, Sloan and Wyrwich found minimal important difference estimates clustered near half a standard deviation across 62 effect sizes (Medical Care, 2003; mean 0.495, SD 0.155). Half a standard deviation of a 12-point SD is 6 points — a statement about what the instrument can distinguish, not about what a patient would want. The smallest detectable change is the harder floor: with an SD of 12 and a reliability of 0.90, the standard error of measurement is 12 × √0.10 = 3.8 points, and the 95% smallest detectable change is 2.77 × 3.8 ≈ 10.5 points. An effect smaller than that cannot be distinguished from measurement error in an individual patient, however precisely the group mean is estimated.

Give both interval limits a clinical reading. “The interval runs from a 1-point to an 11-point improvement; at the lower limit no patient would notice, at the upper limit most would” answers the reviewer’s actual question and requires no threshold at all.

Report the effect on more than one scale. Absolute difference, relative difference, number needed to treat, and where the outcome is time-to-event, an absolute time gain rather than only a hazard ratio, which contains no time units and cannot be read against a clinical threshold.

Say that no validated threshold exists, and name what you checked. “No minimal important difference has been published for this instrument in this population; the values available were derived in post-surgical cohorts over 12-week intervals and are not transportable here” is a specification a reviewer can accept. “The clinical relevance of this difference is unclear” is not.

How the problem is detected in a manuscript

Detection is a comparison exercise, and it takes six checks that a reader can run without recomputing anything.

Compare the confidence interval to the threshold, not the p-value to 0.05. If the interval spans the threshold, the paper cannot conclude the effect is clinically important whatever the p-value says. This is the check most often skipped when a result is otherwise positive.

Compare the adjective to the number. Find every occurrence of substantial, marked, dramatic, clinically relevant and meaningful, and check each against the number in the results table. The gap between them is where this objection is generated.

Check the scale range and version. A threshold stated for a 0–100 transformation applied to a 0–10 raw score, or an instrument threshold applied to one subscale, is a units error that survives review with some regularity because both numbers look plausible.

Check whether the reported measure is relative only. A results section reporting hazard ratios, odds ratios or percentage reductions with no absolute counterpart anywhere is reporting the flattering half of the result.

Check where the threshold first appears. A responder definition present in the results but absent from the protocol, the registration record and the methods was chosen after the data were seen.

Check that the abstract, results and discussion describe the same magnitude. A modest number in Table 2 and a strong claim in the last paragraph of the discussion is the pattern reviewers flag, and the two are usually three sections apart.

What to report

Give both absolute and relative effects, with confidence intervals. Where a benchmark exists, state it and where it comes from. Where the outcome is a scale, state what a unit means in practice. And make sure the abstract and discussion describe the effect in the same terms as the results — this objection often arrives because a modest effect was described as substantial or marked three sections later.

The reporting guidelines ask for exactly this. CONSORT 2010 item 17a requires, verbatim, “For each primary and secondary outcome, results for each group, and the estimated effect size and its precision (such as 95% confidence interval)”, and the 2025 update carries the requirement forward, asking for both absolute and relative effect sizes where the outcome is binary. STROBE item 16a requires unadjusted estimates and, where applicable, confounder-adjusted estimates with their precision. A results section built on p-values alone fails the item on every line, and a manuscript that satisfies it has already answered most of this objection before a reviewer raises it.

That last point is worth a specific pass before resubmitting: the overclaim check is built to find exactly the gap between a reported number and the adjective applied to it in the abstract, and the same discipline applies sentence by sentence in the abstract.

What reviewers actually write

Reviewers rarely write “the effect size is not clinically meaningful” as their opening sentence. They write the specific version, and these are the phrasings that decide papers.

“The between-group difference is 2.1 points on a 100-point scale; please justify the claim that this is clinically important.” “The confidence interval for the primary outcome includes values that would not change management.” “The authors cite a minimal clinically important difference derived in a different population and follow-up interval without justifying its applicability here.” “A within-patient threshold is compared to a between-group difference in means; these are not the same quantity.” “Only relative risk reductions are reported; please add absolute differences and the number needed to treat.” “The responder threshold does not appear in the registered protocol.” “The abstract describes the improvement as marked, which the results do not support.” “The trial was powered to detect the stated threshold and the observed difference is below it; describing the result as clinically meaningful is not consistent with the design.” “Statistical significance is not in question; clinical relevance is.”

The last two are the ones authors most often create during revision rather than receive at first review. Both come from fixing the analysis and leaving the claim.

How to word the response

Answer in the register of the diagnosis: a benchmark statement, an aggregate statement with arithmetic, or a concession that reaches the abstract. All three are short on purpose, and none of them mentions the p-value.

For the benchmark:

The observed difference of 34 m in six-minute walk distance exceeds the minimal important difference of 30 m (95% CI 25 to 33 m) established in the 2014 ERS/ATS technical standard, which we now cite in the Methods and Discussion. We have added the absolute difference with its 95% confidence interval to Table 2 and state in the Discussion where the lower limit of that interval sits relative to the 30 m threshold.

For the aggregate argument:

We agree the per-person effect is small. We have added the population arithmetic to the Discussion: an absolute reduction of 0.2 percentage points corresponds to a number needed to treat of 500, and, applied to the 5 million adults in this age band at the observed 61% adherence, to approximately 6,100 events over five years. We state the assumptions behind that projection and the reasons the estimate may not transport.

For the concession:

The reviewer is correct that the effect is below the threshold for this instrument. We have removed the claim of clinical importance from the abstract, title and discussion, and now present the result as a bound: this trial excludes effects larger than 3.2 points, which constrains what the intervention can achieve on this outcome. The precision of that bound is the contribution we now claim.

Each of these belongs in the response letter and in the manuscript, in the same words. How to write a response to reviewers covers the surrounding structure.

For clinical trial manuscripts

If your primary endpoint is a response rate or survival difference, the related question is whether the endpoint itself is meaningful — a surrogate endpoint improvement that has not been shown to track a clinical one attracts this objection in a stronger form. Response criteria and endpoint definitions covers the endpoint side.

Two trial-specific traps compound it. A composite endpoint can reach significance on its least serious component, so a composite result described as a reduction in a serious outcome is a magnitude claim the components do not support; report the components separately. And where the effect is expressed as a hazard ratio, the reviewer cannot read a clinical magnitude from it at all, because a hazard ratio carries no time units — an absolute time gain over a stated horizon is what permits the comparison against a threshold. Whether the endpoint is a legitimate stand-in for the outcome patients care about is the subject of what a surrogate endpoint is, and the wider reporting pass sits in peer review for clinical trial manuscripts.

Before you resubmit

Run five checks, in order. Does every effect estimate appear with an interval rather than a p-value alone? Is an absolute measure reported wherever a relative one is? Does every threshold you cite name its derivation study, its population and its direction? Does the strongest adjective in the abstract survive comparison with the number in the results table? And has the claim in the title moved with the analysis?

More in this family: responding to statistical reviewer comments.

Related

What is an effect size · Minimal clinically important difference · Confidence intervals · Sample size too small · Post hoc power · Statistical objections

PerfectPaper reads the reported effect, its confidence interval and any cited threshold together, and reports every place the manuscript describes a magnitude the numbers do not support — including the adjectives in the abstract that sit three sections away from the table they refer to.

Review my manuscript

Frequently asked questions

How do I respond when a reviewer says the effect is not clinically meaningful?

Work out which version you received first. If the point estimate is below a threshold, cite an external threshold your effect exceeds or concede and reframe. If the interval is wide, interpret both limits. If the objection is really about an adjective in the abstract, change the adjective. Do not answer with the p-value and do not promise more data.

What counts as a clinically meaningful effect size?

There is no universal figure; it is set by the outcome, the instrument and the population. The usable form of the answer is a minimal clinically important difference — the smallest change patients or clinicians would regard as meaningful — published for your instrument and population. Citing one, and checking that it was derived in a comparable population, in the same direction, and independently of the data you are defending, is the strongest available answer to this objection.

How do I argue a small effect is important?

Either exceed an established benchmark for your outcome, or show that the effect operates at a scale where a small per-person change produces a large aggregate one. Both need to be argued explicitly with numbers rather than asserted: name the denominator, the time horizon and the assumed uptake.

Does reporting the effect differently answer this objection?

Only if you report more, not less. Give both the relative and the absolute effect, always. A relative measure alone can make a trivial absolute difference appear substantial, and reviewers increasingly treat relative-only reporting as a presentational red flag. CONSORT 2025 asks for both for binary outcomes, and a number needed to treat makes the absolute magnitude unmissable.

Will collecting more data answer the objection?

No, and it will hurt. A larger sample estimates the same small effect more precisely, which makes the reviewer’s point more firmly than your own. At a standardised difference of 0.03, reaching p < 0.001 takes roughly 24,000 participants per arm and the difference is still 0.3 points.

How do I answer this objection when no published threshold exists for my outcome?

Say so, and name what you checked. Then state what one unit means in practice, use a distribution-based figure such as the smallest detectable change as a floor rather than a definition of importance, give both confidence-interval limits a clinical reading, and report the effect on more than one scale.

Can a paper survive review when the effect really is small?

Yes, when the effect is presented accurately. A precisely estimated small effect bounds what the intervention can do, which is useful. What is not publishable is a small effect described as though it were large.

Last updated September 10, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.