Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
A minimal clinically important difference is the smallest score change patients experience as beneficial. It belongs to a population and context, not the instrument.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A minimal clinically important difference (MCID) is the smallest change in a patient-reported outcome score that patients themselves experience as beneficial rather than as noise. Jaeschke, Singer and Guyatt introduced the term in 1989 and estimated it at about 0.5 points per item on seven-point scales. An MCID belongs to a population, an instrument and a moment, never to a questionnaire alone.
That definition is compact, and nearly every dispute about MCIDs in peer review comes from treating the number as a fixed property of the instrument. This page covers where the definition came from, how anchor-based and distribution-based estimates are produced, why the same instrument yields values differing several-fold, the difference between a within-person threshold and a between-group difference, how the concept is detected in a submitted manuscript, and what reviewers write when it has been mishandled.
Jaeschke, Singer and Guyatt published the formal definition in Controlled Clinical Trials in 1989 (volume 10, pages 407–415), where they also introduced the abbreviation MCID. Their wording: “the smallest difference in score in the domain of interest which patients perceive as beneficial and which would mandate, in the absence of troublesome side effects …, a change in the patient’s management.”
Two features of that sentence are routinely lost. The first is perceive: the standard is the patient’s experience, not a clinician’s judgement and not a statistical test. The second is mandate a change in management: the threshold was conceived as a decision boundary for care, not as a label to attach to a trial result.
The original estimate came from studies of dyspnoea, fatigue and emotional function in chronic heart and lung disease using the Chronic Respiratory Questionnaire, on which patients rating themselves slightly better had gained a mean of roughly 0.5 per item on a seven-point Likert scale. That figure is still quoted as though it were a constant of the instrument. It was an average across three small studies, in patients with chronic heart and lung disease, over one recall interval.
Later writers dropped the word “clinically” and moved to minimal important difference (MID), on the reasoning that “clinically” invites clinician judgement into a quantity the definition assigns to the patient. The two labels are used interchangeably in practice, which is one reason a manuscript has to say what it means rather than rely on the acronym.
Terminology for score-interpretation thresholds is genuinely unharmonised. Derivation studies variously report a minimal clinically important difference, a minimal important difference, a minimal important change, a minimal clinically important improvement, a minimal detectable change, a smallest worthwhile effect, a patient acceptable symptom state and a substantial clinical benefit — sometimes for the same quantity under different names, more often for different quantities under names a reader cannot tell apart. No single term has been adopted as a standard, and none looks likely to be.
The practical consequence is that the acronym carries no information on its own. A reader who sees “MCID = 10” cannot tell whether ten points is a within-person change threshold, a between-group difference, the edge of measurement error, or a level of symptom the authors regard as acceptable. Until the manuscript says which, the number is not checkable, and a reviewer is entitled to say so.
The variants are not synonyms. Minimal important change (MIC) is explicitly a within-individual quantity over time. Minimal clinically important improvement (MCII) is directional, applying only to getting better. The United States Food and Drug Administration avoids all of them: its April 2023 draft guidance Patient-Focused Drug Development: Incorporating Clinical Outcome Assessments Into Endpoints For Regulatory Decision-Making uses meaningful score difference (MSD) instead. A manuscript that cites an “MCID” without saying which construct it means has not yet said anything checkable.
An MCID is a property of a population, an intervention, a follow-up interval, a direction of change and a derivation method — jointly. Change any one and the number moves, often by more than the number itself.
Follow-up interval is the dependence that surprises authors most, and the mechanism is not subtle. An anchor-based threshold is read off the patient’s answer to “compared with before, how are you now?” — and the standard behind that answer moves as recovery proceeds. Early after a major operation, patients who call themselves slightly better have often gained very little on the instrument, because escape from the immediate post-operative state is itself the improvement they are reporting. A year later the same rating category corresponds to a substantially larger score gain, because the patient is now comparing against a recovered baseline and expects more. A threshold derived at three months and applied at two years is therefore not the same quantity, even with the same instrument, the same operation and the same anchor question. Devji and colleagues’ credibility instrument treats the elapsed interval between baseline and follow-up as a criterion in its own right, on the grounds that a transition rating over a long recall period is a poorer record of change than one over a short period. The interval is therefore part of what a borrowed threshold carries with it, and a manuscript that reuses a threshold across a different follow-up owes the reader an argument that it transfers.
Direction is the second dependence. A threshold for improvement is not automatically the mirror of a threshold for deterioration: patients do not weigh a loss and a gain of equal size equally, and most derivation studies estimate only the improvement side. A paper that applies an improvement threshold to worsening scores has assumed a symmetry it did not test.
The third is method dependence, and here the demonstration is unusually clean. Copay and colleagues (The Spine Journal, 2008) applied several common anchor-based and distribution-based calculations to one cohort of 454 lumbar spine surgery patients and reported that the candidate values spanned fivefold for the Oswestry Disability Index and tenfold for back pain. They settled on 12.8 points for the ODI, 4.9 for the SF-36 physical component summary, 1.2 for back pain and 1.6 for leg pain — but the range they published is the more useful result, because it shows what “the MCID for the ODI” conceals.
Anchor-based estimation maps change in the instrument onto an external variable whose meaning is already interpretable — most often a patient-reported global rating of change, asked as something like “compared with when you started, is your pain much better, a little better, about the same, a little worse or much worse?” The MCID is then read off as the score change among patients selecting the smallest meaningful category.
Every step of that mapping is a discretionary choice, and each is a place a manuscript can go wrong: which anchor, which category counts as minimal, whether “a little worse” is treated as the mirror of “a little better”, and whether the estimate is the mean change in that category, an ROC-derived cut-point maximising the Youden index, or a predictive-modelling estimate.
Devji and colleagues published a credibility instrument for exactly this in The BMJ in 2020 (369:m1714). It has five core criteria: the anchor is rated by the patient; the anchor is interpretable and relevant to the patient; the MID estimate is precise; the correlation between anchor and instrument is satisfactory; and the chosen anchor threshold reflects a small but important difference. Two further criteria apply to transition-rating anchors, covering the elapsed interval and the correlation of the rating with baseline, follow-up and change scores.
Applied at scale, those criteria are unflattering. Systematic surveys of the anchor-based literature — Carrasco-Labra and colleagues catalogued several thousand estimates across hundreds of instruments for Journal of Clinical Epidemiology in 2021 — find that a large share of published thresholds cannot be assessed against them at all, because the derivation study never reports the correlation between the anchor and the instrument, or reports a point estimate with no interval around it. The survey found the anchor–instrument correlation weak or undocumented for 89 per cent of the 5,324 estimates it catalogued, and the estimate imprecise for 47 per cent. Both omissions are fatal to reuse: an anchor uncorrelated with the score is measuring something else, and a threshold without an interval hides how little the estimate is pinned down.
Two practical consequences follow for an author citing someone else’s threshold. If the derivation paper does not report the anchor–instrument correlation, the threshold cannot be defended as anchor-based in any meaningful sense, and saying so in the limitations is better than being told. If it reports no interval, the reader should be given the range of values other derivations produced for the same instrument instead of a single borrowed number.
Present state bias is the tendency of a patient’s transition rating to be weighted more heavily by how they feel at follow-up than by how much they have actually changed since baseline. Terluin and colleagues have described and simulated the effect in the methodological literature on minimal important change, and the name is now standard in that literature.
The practical consequence is that a global rating of change is not a clean record of change. It is a blend of change and current state, and the blend varies with the length of the recall interval and with how ill the patient is at the moment of asking. A second and largely separate distortion comes from the composition of the sample: the proportion of participants who improved shifts ROC-derived cut-points directly, because the cut-point that maximises the Youden index depends on the prevalence of the class being detected. Neither problem is visible in the reported threshold, and both are recoverable from a well-described methods section, which is why the description is what a reviewer asks for.
Estimators differ in which of these they resist. Mean change in the smallest meaningful anchor category and ROC-based cut-points are simple and transparent but inherit the proportion-improved dependence. Predictive-modelling and latent-variable approaches can adjust for the proportion improved and for the unreliability of the anchor itself, in exchange for stronger modelling assumptions that the derivation paper then has to state.
A related trap is well enough known to have its own paper title: Terluin and colleagues, in Quality of Life Research in 2021, “Assessing baseline dependency of anchor-based minimal important change (MIC): don’t stratify on the baseline score!” (30(10):2773–2782) Stratifying by baseline to test whether the threshold depends on severity produces spurious dependency through regression to the mean — patients scoring low at baseline have nowhere to go but up, whether or not the treatment did anything. A manuscript reporting that its MCID “differs by baseline severity” on the strength of a baseline-stratified analysis has usually found the artefact rather than the dependency.
Distribution-based estimation derives a threshold from the variability of the scores rather than from anything a patient said. The three common forms are half a standard deviation, one standard error of measurement, and the effect-size conventions.
The half-standard-deviation rule has a real empirical basis. Norman, Sloan and Wyrwich, reviewing 38 studies yielding 62 effect sizes for Medical Care in 2003, found MID estimates clustered tightly around 0.5 SD (mean 0.495, SD 0.155), and argued the regularity reflects a general psychophysical limit: human discrimination across many tasks runs to roughly one part in seven, which is close to half a standard deviation. Their own framing was a “threshold of discrimination”, not a definition of importance.
The arithmetic connecting the two distribution rules is worth stating, because it explains why they so often agree. The standard error of measurement is SD × √(1 − reliability). At a reliability of 0.75 the SEM is exactly 0.5 SD, so the “half an SD” and “one SEM” criteria coincide at a reliability typical of many well-behaved instruments and diverge everywhere else.
The regulatory position on all of this is unambiguous in direction, whatever the exact wording. The FDA’s 2023 draft guidance treats distribution-based methods as informative about measurement variability but insufficient on their own to establish a meaningful score difference, because nothing in the calculation consults a patient. That is the right way to read them generally: half a standard deviation tells you what the instrument can distinguish, not what a patient would want. A manuscript reporting only a distribution-based threshold has reported measurement precision and labelled it clinical importance.
The honest use of a distribution-based figure is as a floor. If the anchor-based threshold a paper cites is smaller than one standard error of measurement, that is worth saying out loud, because the paper is proposing to detect something below the resolution of its own instrument.
The smallest detectable change is the smallest difference in an individual’s score that exceeds measurement error, conventionally computed at 95 per cent confidence as 1.96 × √2 × SEM, or approximately 2.77 SEM. The minimal important change is what patients regard as worth having. They are separate quantities, and their ordering matters.
When the MIC is smaller than the smallest detectable change — a common situation for individual patients on short instruments — a change can be important and still be indistinguishable from noise in one person, even though the mean change in a group is estimated far more precisely. Copay and colleagues confronted this directly in their spine cohort and treated the minimum detectable change as the defensible threshold, effectively raising the bar to the point where classifying an individual patient could be justified rather than asserted.
The two quantities also answer to different fixes. A smallest detectable change that is too large is a measurement problem, addressed by a more reliable instrument, more items or repeated measurement. A minimal important change that is too small to be detected in one patient is not a measurement problem at all — the instrument is working, and the honest conclusion is that individual-level responder classification is beyond what this instrument can support, however well the group-level estimate behaves.
The distinction is the reason a manuscript that uses an MCID to sort individual participants into responders and non-responders owes the reader the instrument’s SEM. Without it, the classification’s error rate is unknowable. This is the same class of problem as post-hoc power: a number computed after the fact that appears to license a conclusion it cannot support.
A within-person MCID and a between-group treatment effect are different quantities on the same scale, and equating them is the most consequential error in this area. An MCID derived from individual change describes what one patient must gain to notice a benefit. A randomised trial reports the difference between arms in mean change — an average over people, most of whom experienced something other than the average.
A between-group mean difference well below the within-person MCID can still correspond to a substantial share of patients crossing it, because the treatment shifts a distribution rather than moving every patient by the mean. Conversely, a between-group difference above the MCID does not establish that most patients benefited: a large gain in a minority can carry a mean past a threshold nobody near the middle of the distribution reached.
The FDA’s 2023 draft guidance separates the two uses of a threshold cleanly — judging the expected effect for an average patient in a target population, and classifying individual patients as having changed meaningfully — and treats them as different claims requiring different support. Using one threshold for both takes on assumptions worth stating: that the threshold does not depend on where the patient started, and that it is the same for getting better and getting worse. Neither holds automatically, and the follow-up and direction dependences described above are exactly the ways they fail. Manuscripts that slide between the two uses tend to attract the same objection reviewers raise about effect sizes that are not meaningful.
Responder analysis dichotomises a continuous outcome at an MCID and reports the proportion exceeding it in each arm. Snapinn and Jiang set out the case against in Trials in 2007 (8:31): the approach sacrifices a potentially large amount of statistical efficiency, and fails at its own objective, since a proportion above an arbitrary cut-point is not itself a measure of clinical relevance. Dichotomising a continuous variable discards information, and the power lost has to be bought back with participants — which is how a threshold intended to make a result interpretable ends up interacting with sample size objections.
The regulatory position is a compromise rather than a prohibition, and its three parts are worth borrowing whether or not a submission is going to a regulator. Prespecify a single threshold and justify it. Derive it from data other than the trial being analysed. And show the treatment effect across a range of thresholds as well as at the chosen one, so a reader can see whether the conclusion turns on the cut-point.
The independent-derivation requirement is the one manuscripts fail most often, and the failure is easy to miss because it looks like diligence: the authors estimate the MCID in their own cohort, using their own anchor question, and then use it to declare their own result meaningful. A threshold fitted to the data it judges is not a threshold. Prespecification of endpoints and thresholds is also what trial registration and endpoint checks are looking for.
The sensitivity analysis is the easiest of the three to supply and the most often omitted. Reporting the responder proportions at, say, half the chosen threshold, the threshold itself, and twice it takes three lines of a table and pre-empts the reviewer’s obvious question. Where the difference between arms survives across that range, the paper is stronger for having shown it; where it does not, the authors learn that before an editor does.
An MCID answers “how much better”, and a patient acceptable symptom state answers “good enough” — the score level at or beyond which patients consider their condition satisfactory. They are different constructs and they produce different numbers on the same scale.
Tubach and colleagues estimated both for pain and function in rheumatic disease, work carried forward in the OMERACT–OARSI literature, and the pattern that matters is structural rather than numerical. On a 0–100 pain scale the improvement threshold sits somewhere in the mid-teens, while the acceptable-state level sits near the middle of the scale — so the two constructs answer different questions and can disagree about the same patient in both directions. A patient can achieve the improvement threshold and remain far from an acceptable state, if they started very badly. A patient who started close to the acceptable level can reach it without ever achieving the improvement threshold.
For a trial, that means the two produce different responder counts from the same data, and neither is wrong. Reporting only one of them describes half the result, and choosing between them after seeing which gives the better answer is the same problem as choosing the threshold after seeing the data. Where both are available for an instrument, report both.
Detection of MCID misuse is a documentary exercise, and it begins with the citation attached to the threshold. Six checks find most of it.
Trace the threshold to its source study and compare populations. An ODI threshold derived in elective lumbar fusion does not transfer to conservative management of acute back pain, and the transfer is rarely argued.
Check the scale range and version. A threshold stated for a 0–100 transformation applied to a 0–10 raw score, or an MCID for a full instrument applied to one subscale, is a units error that survives peer review with some regularity because both numbers look plausible.
Compare the threshold against the reported confidence interval. If the trial’s 95 per cent confidence interval for the treatment effect spans the MCID, the paper cannot conclude the effect is meaningful, whatever the p-value says. This is the check most often skipped when a result is otherwise positive.
Identify whether the threshold is within-person or between-group, then check which quantity the paper compares it to. A within-person MIC placed alongside a difference in mean change is the error described above.
Look for a threshold that appears only in the results. A responder definition absent from the protocol, the registration record and the methods, but present in the results, was chosen after the data were seen.
Check whether a single number is presented where the literature has a range. For the ODI, the SF-36 and most pain scales, published estimates differ several-fold, and a manuscript citing one without acknowledging the spread is asserting more precision than exists.
PerfectPaper reads the threshold, its cited derivation study and the reported effect estimate together, and reports where the comparison being made is not the comparison the threshold supports. That check runs alongside the wider clinical trial reporting review and the overclaim check.
“The MCID cited was derived in a different population and time frame; please justify its applicability here.”
“The authors compare a between-group difference in means to a threshold derived from within-patient change. These are not the same quantity.”
“The confidence interval for the treatment effect includes values well below the stated MCID, so the claim of clinical importance is not supported by these data.”
“Published MCID estimates for this instrument span a wide range. Please state which estimate was used, how it was derived, and how sensitive the conclusions are to that choice.”
“The responder threshold does not appear in the trial registration or the protocol. Please clarify when it was specified.”
“Please report the exact wording and recall period of the anchor question, and its correlation with the change score.”
“The minimal important change reported here is smaller than the instrument’s smallest detectable change, so individual participants cannot be reliably classified as responders.”
“The trial was powered to detect the MCID and the observed difference is smaller than the MCID; describing it as clinically meaningful is not consistent with the design.”
The last is the one that decides papers, because it is unanswerable by revision. When a threshold has been chosen well and the effect falls below it, the honest move is to say so — a limitation stated plainly reads better to editors than a redefinition, and it is easier to defend in a response to reviewers.
Four elements make a threshold reproducible, and stating them takes a paragraph. Perspective: whose judgement of “meaningful” — the patient’s, a clinician’s, or an observer’s, such as a parent or carer reporting on someone else. Application: individual-level or group-level; within-person, within-group or between-group; improvement, worsening or both. Magnitude: whether the threshold marks the edge of measurement error, the edge of importance, a substantial rather than minimal difference, or an acceptable state. Methodology: the approach (qualitative, anchor-based, distribution-based or combined), the specific estimator (mean change, an ROC cut-point, a regression or predictive model, half a standard deviation), and the context in which it was derived — population, intervention, follow-up interval and study design.
Add three things beyond that list. Report an interval for the threshold, because a bare point estimate implies a precision that derivation studies rarely have and readers rarely check. Report the proportion of participants exceeding the threshold in each arm alongside the continuous analysis, rather than instead of it, so the reader gets both the efficient estimate and the interpretable one. And plot the cumulative distribution of change by arm: it shows the whole shift, exposes whether the effect is concentrated in a subgroup, and lets a reader apply their own threshold instead of accepting yours.
If the threshold is borrowed rather than derived, one further sentence is owed: what population it was derived in, and why that population is close enough to this one for the number to transfer. This is the sentence most often missing, and it is the one a methods reviewer looks for first.
The concept is contested among methodologists, and a manuscript that acknowledges this is on stronger ground than one that treats a cited number as settled.
The half-standard-deviation regularity is an empirical observation about a literature, not a principle, and Norman and colleagues framed it as a threshold of discrimination rather than a definition of clinical importance. Anchor-based methods, the ones regulators prefer, all rest on retrospective global ratings that are known to be contaminated by present state bias and by imperfect recall over the interval. The newer latent-variable estimators correct known biases but require longitudinal factor or item-response models that most derivation studies do not fit, and their assumptions are harder for a reviewer to check than the biases they remove.
There is also a defensible position that thresholds should always be reported alongside, and never in place of, the full distribution of change. Any single cut-point misdescribes some patients by construction: two people either side of it may differ by one point, and the dichotomy declares one a responder and the other not. That is a reason to show the distribution, not a reason to abandon thresholds — a paper that reports both is making a claim a reader can audit.
None of this makes MCIDs useless. It makes an unattributed MCID useless, which is a different claim and a much easier one to act on.
Confounding · When a reviewer says the effect size is not meaningful · Post-hoc power · Statistical objections · Response criteria in oncology · Research methods concepts
MCID stands for minimal clinically important difference: the smallest change in a patient-reported outcome score that patients perceive as beneficial. Jaeschke, Singer and Guyatt introduced the abbreviation in 1989. Related labels include minimal important difference, minimal important change, and the FDA’s meaningful score difference, which are not interchangeable.
The minimum clinically important difference is the smallest score change patients themselves regard as worthwhile, rather than the smallest change that is statistically detectable. It is estimated by anchoring score change to a patient-reported rating of improvement, and it varies with population, intervention, follow-up interval and derivation method.
Jaeschke, Singer and Guyatt defined it in 1989 as the smallest difference in score in the domain of interest which patients perceive as beneficial and which would mandate a change in the patient’s management. The definition rests on patient perception, not on clinician judgement or a significance test.
A clinically meaningful difference is a change large enough that patients notice and value it. In trials it is operationalised as a threshold, typically anchored to patient ratings of improvement, against which the observed effect is compared. Thresholds for within-person change and for between-group differences are separate quantities.
A minimally important change is the smallest within-individual score change that patients, on average, regard as important. Published values differ substantially for the same instrument, because the threshold depends on the population, the follow-up interval, the direction of change and the estimator used. Thresholds can also differ with the length of follow-up over which they were derived, because a patient’s retrospective rating of change is a weaker record of change over a long recall period — which is why the credibility instrument of Devji and colleagues treats the elapsed interval as a criterion in its own right.
No. Statistical significance concerns whether an observed difference is compatible with chance; a minimal clinically important difference concerns whether the difference is large enough to matter to patients. A large trial can produce a significant result far below the threshold, and a small trial can miss significance for a difference above it.
Anchor-based methods map score change onto a patient rating of change, then take the mean change, an ROC cut-point, or a predictive-modelling estimate in the smallest meaningful category. Distribution-based methods use half a standard deviation or one standard error of measurement, and regulators treat these as insufficient on their own.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect