Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
A non-inferiority trial tests whether a new treatment is worse than an active control by more than a prespecified margin. The margin, not the p-value, carries the claim.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A non-inferiority trial is a randomised trial designed to show that a new treatment is not worse than an established active control by more than a prespecified margin. It does not test for equality, and it does not test for superiority. Its conclusion is directional: whatever disadvantage the new treatment carries is small enough to be clinically acceptable.
The design exists because placebo control is often unethical once an effective treatment exists, and because a new treatment can be worth having for reasons other than greater efficacy — fewer adverse effects, oral rather than intravenous administration, no monitoring requirement, lower burden on the patient. This page covers the margin and how it is chosen, the constancy and assay-sensitivity assumptions the design rests on, why intention-to-treat is not the conservative analysis here, how a mishandled non-inferiority claim is detected in a manuscript, and what reviewers say when it has been mishandled.
The null hypothesis of a non-inferiority trial is that the new treatment is worse than the active control by at least the margin. Rejecting that null licenses the conclusion that the true difference is smaller than the margin — nothing more.
This inversion is the source of most errors in the literature. In a superiority trial, the null is “no difference”, and a non-significant result means the trial failed to demonstrate an effect. In a non-inferiority trial, the null is an assertion of meaningful inferiority, and the burden of proof runs the other way. A trial that fails to demonstrate non-inferiority has not demonstrated inferiority, and a superiority trial that finds no significant difference has not demonstrated non-inferiority. Reinterpreting a null superiority result as evidence of comparability is a distinct error, and it survives peer review often enough to be worth checking on every manuscript that reports a negative primary outcome.
The formal test is conducted one-sided at 2.5%, which corresponds to reading the lower (or upper, depending on the direction of benefit) bound of a conventional two-sided 95% confidence interval against the margin. Nothing else in the analysis carries the claim.
The non-inferiority margin, usually written as M2 or delta, is the largest difference in the unfavourable direction that the investigators are willing to call clinically acceptable. Everything the trial can conclude is conditional on that number.
A margin chosen too wide makes non-inferiority trivially easy to demonstrate, and in the limiting case a treatment with no effect at all would pass. A margin chosen too narrow makes the trial unaffordably large and may be indefensible against ordinary measurement noise. Because the margin is chosen before the data exist and is not estimated from them, it is the one quantity in a non-inferiority trial that peer review can evaluate on its own merits — and it is the quantity most often reported without any justification.
The margin must appear in the protocol and the registration record, not only in the published sample-size paragraph. A margin that first appears in a manuscript’s discussion section is not a prespecified margin, whatever the sentence around it claims. Trial registration and endpoint discipline is the check that catches this, because public registries retain a version history and a margin added mid-trial is visible in it.
The FDA’s final guidance Non-Inferiority Clinical Trials to Establish Effectiveness, issued in November 2016, sets out a two-step construction that has become the reference framework.
M1 is the entire effect of the active control relative to placebo, estimated from historical placebo-controlled trials of that control — conventionally taken not as the point estimate but as the lower bound of the confidence interval around a meta-analytic estimate, so that sampling error in the historical evidence is discounted rather than assumed away. M1 is a statistical quantity derived from external data.
M2 is the largest loss of that effect the field is willing to accept, and it is a clinical judgement rather than a calculation. M2 must be smaller than M1, and a common convention sets it at half of M1, preserving at least 50% of the control’s established effect. That 50% figure is a convention adopted for its conservatism, not a quantity derived from anything; a manuscript that cites it should say so rather than present it as a standard.
Two analytical routes then exist. The fixed-margin approach, often called the 95%–95% method, fixes M1 as the lower bound of the confidence interval around the historical estimate, sets M2 as a preservation fraction of it, and requires the current trial’s 95% interval to exclude the margin — two layers of discounting, and the more conservative option. The synthesis method combines the historical and current estimates into a single test statistic, is more efficient, and buys that efficiency by treating the historical effect as directly transportable. The two methods can disagree on the same dataset, and a manuscript should name which one it used.
Constancy is the assumption that the active control’s effect relative to placebo in the current trial would be the same as the effect observed in the historical placebo-controlled trials from which M1 was derived. The entire non-inferiority argument depends on it, and the current trial contains no placebo arm with which to check it.
Constancy fails in specific, recognisable ways. Standard of care changes, so background therapy differs between the historical trials and the present one. Diagnostic criteria shift, so the disease being treated is not the same disease. Enrolment moves to different regions or care settings with different event rates. Outcome definitions are revised. Adherence in a modern trial differs from adherence in a trial conducted twenty years earlier. Each of these erodes the control’s effect in the present trial, and an eroded control effect makes non-inferiority easier to demonstrate — the failure is not neutral, it is biased toward the sponsor’s hypothesis.
A manuscript that names the historical trials used to set M1, states their dates, and argues explicitly that the populations and background care are comparable has done the work. A manuscript that cites a margin as “clinically accepted” has not, and the reviewer has no way to evaluate the claim.
Assay sensitivity is the ability of a trial to detect a difference between treatments if one exists. ICH E10, Choice of Control Group and Related Issues in Clinical Trials, made it the pivotal concept for active-control designs, and it behaves in the opposite direction to intuition.
A superiority trial that is sloppy — imprecise outcome measurement, heterogeneous population, poor adherence, high dropout — tends toward the null and therefore fails. A non-inferiority trial that is sloppy tends toward the null as well, and the null is what it is trying to establish. Everything that degrades a trial’s ability to see a difference makes non-inferiority easier to declare. This is the reason a non-inferiority result cannot be interpreted from internal evidence alone, and the reason regulators attend so closely to trial conduct quality in this setting.
Three-arm trials that include the new treatment, the active control and a placebo resolve the problem directly, because the control-versus-placebo comparison demonstrates assay sensitivity within the trial that depends on it. They are frequently unethical and therefore rare, which leaves the constancy argument doing the work in almost every published non-inferiority trial.
Non-inferiority is established when the confidence interval for the treatment difference lies entirely on the acceptable side of the margin. Three further readings are possible from the same interval, and a well-written results section distinguishes them.
If the interval excludes the margin and also excludes zero in the favourable direction, the new treatment is both non-inferior and superior. If the interval excludes the margin but includes zero, non-inferiority is established and superiority is not. If the interval includes the margin, non-inferiority is not established, regardless of where the point estimate sits — a point estimate favouring the new treatment with an interval crossing the margin is an inconclusive trial, not a positive one. If the interval lies entirely beyond the margin, the new treatment is inferior by a clinically meaningful amount.
The point estimate is not the finding. A recurring pattern in non-inferiority abstracts is a difference of, say, 1.2 percentage points reported against a 10-point margin and described as showing the treatments are “similar”, when the interval spans −11 to +13 and the trial has established nothing at all. This is the single most common misreading of the design, and it is visible from the abstract alone.
In a superiority trial, intention-to-treat analysis is conservative: non-adherence, crossover and protocol violations dilute any true difference and pull the estimate toward the null, making the trial harder to win. In a non-inferiority trial, the same dilution pulls the estimate toward the trial’s own alternative hypothesis, and intention-to-treat becomes anti-conservative.
ICH E9, Statistical Principles for Clinical Trials, therefore calls for confirmatory trials to plan both a full-analysis-set and a per-protocol analysis so that any difference between them can be discussed explicitly — advice that matters most in this design, because here the intention-to-treat analysis is the anti-conservative one. The CONSORT extension for non-inferiority and equivalence trials asks authors to state which analysis populations were used and to interpret the result against the non-inferiority hypothesis. Discordance is informative rather than embarrassing: if non-inferiority holds under intention-to-treat but not per-protocol, the likely explanation is that poor adherence in both arms has flattened a real difference.
Per-protocol analysis is not itself safe, because exclusion of participants after randomisation breaks the randomisation and can introduce selection effects of exactly the kind described under confounding. Neither population is the conservative one. Reporting both, with the exclusions enumerated, is what the guidance asks for and what reviewers of clinical trial manuscripts look for first.
Required sample size in a non-inferiority trial scales approximately with the inverse square of the margin. Halving the margin quadruples the number of participants needed, holding the event rate, power and one-sided alpha constant.
This is why margin choice is an operational decision as well as a scientific one, and why a margin that looks generous is often a budget artefact rather than a clinical judgement. A trial powered for a 10-percentage-point margin in a binary outcome needs roughly four times as many participants at 5 points. A margin worth questioning on these grounds usually announces itself the same way: a suspiciously round number, and a justification paragraph that cites no external estimate of the control’s effect.
A related consequence: non-inferiority trials with narrow, defensible margins are large, and large trials detect small differences in secondary and safety outcomes that smaller trials would miss. Interpreting those incidental findings requires the same multiplicity discipline as any other prespecified secondary analysis, and a difference that is statistically detectable in a 6,000-patient non-inferiority trial may still be clinically negligible.
Testing for superiority after non-inferiority has been established is permitted without an adjustment for multiplicity, because the two hypotheses form a closed testing procedure: superiority is a strictly stronger claim than non-inferiority, so the sequence is hierarchical and the family-wise error rate is preserved. The superiority claim should be made in the intention-to-treat population.
Testing for non-inferiority after superiority has failed is permitted only if the margin was prespecified. A margin selected after the results are known is not evidence, and a paper that introduces one at that point has performed an unregistered analysis, whatever its confidence interval shows. The distinction is easy to state and easy to check: the margin either appears in the registration record and protocol or it does not.
Sponsors sometimes prespecify both hypotheses at the design stage precisely so that this sequence is legitimate. That is good practice, and it should be visible in the statistical analysis plan rather than asserted retrospectively in a response letter. If a reviewer raises it, the response should point to the dated document, not to a rationale constructed afterwards.
Biocreep is the gradual erosion of efficacy that occurs when each new treatment is tested for non-inferiority against the last one to pass, rather than against the original agent whose superiority over placebo was demonstrated.
If treatment B is shown non-inferior to A within a margin that permits losing half of A’s effect, and treatment C is then shown non-inferior to B within the same margin, C may retain little of the effect that made A worth using — while every individual trial in the chain was correctly conducted. The mechanism is arithmetic, not misconduct.
The standard defence is to require the active control to be the best available treatment with a well-characterised placebo-controlled effect, and to derive M1 afresh from the original placebo-controlled evidence for each new trial rather than inheriting a margin from the previous one. A manuscript that justifies its margin by citing the margin used in an earlier non-inferiority trial has, structurally, done the thing that produces biocreep, and it is worth naming explicitly in review.
A non-inferiority margin expressed as a hazard ratio or a risk ratio does not correspond to a fixed absolute difference. A hazard ratio margin of 1.3 permits a much larger absolute excess of events when the control event rate is 20% than when it is 2%, so a trial that enrols a lower-risk population than the historical trials has quietly widened its own margin in absolute terms.
Time-to-event margins carry a further assumption. A single hazard ratio has a stable interpretation only under proportional hazards; when the hazards are non-proportional the estimated ratio depends on follow-up duration and on the censoring distribution, and a margin defined on that scale inherits the instability. Crossing survival curves are the visible symptom, and restricted mean survival time differences are the usual remedy, which requires the margin to be respecified on a time scale. Where participants can experience an event that precludes the primary outcome, the analysis also needs competing-risks handling before the margin means anything.
State the scale, the assumption, and the check. A manuscript that reports a hazard ratio against a hazard ratio margin without a proportional-hazards assessment has left the central claim unverified.
Non-inferiority does not establish equivalence. Equivalence requires bounding the difference in both directions, and it is the appropriate design for bioequivalence and for some device comparisons; non-inferiority bounds one direction only.
Non-inferiority does not establish that a treatment works. It establishes a relationship to a control, and if that control has no effect in the population studied, the relationship is uninformative. This is why the assay-sensitivity argument is not a formality.
Non-inferiority on efficacy does not license a claim of overall benefit. The justification for a non-inferiority design is normally an advantage elsewhere — tolerability, route of administration, monitoring burden, feasibility in the settings where the disease is actually treated — and that advantage has to be demonstrated, not asserted. A trial that establishes non-inferior efficacy and reports a numerically favourable but underpowered safety comparison has demonstrated one thing and implied another. That gap is where overclaiming tends to enter the abstract.
Detection is a document-level task rather than a statistical one, because the numbers in a mishandled non-inferiority paper are usually correct. The following signals recur.
The margin appears in the sample-size paragraph and nowhere else. It is not restated in the results, not marked on the forest plot or CI figure, and not referenced in the conclusion. The reader cannot check the claim against the interval without doing arithmetic the authors should have done.
The margin has no external anchor. The justification reads “considered clinically acceptable” or “consistent with previous trials in this indication”, with no citation to a placebo-controlled estimate of the control’s effect and no statement of M1.
The registration record disagrees with the paper. The registry entry specifies superiority, specifies a different margin, specifies a different primary outcome, or was amended after the first participant was enrolled. Registry version histories are public and dated.
Only intention-to-treat is reported. No per-protocol analysis appears, or one appears in a supplement with no comment on concordance.
Adherence and dropout are high and undiscussed. Loss to follow-up above roughly 10%, or substantial crossover, with no sensitivity analysis — in this design that pattern favours the conclusion drawn.
The abstract’s language outruns the interval. “Comparable”, “similar efficacy”, “as effective as”, “no significant difference” in a paper whose confidence interval includes the margin.
Non-inferiority is claimed on a secondary outcome while the primary outcome was analysed for superiority and failed.
The control arm underperforms its historical benchmark. The event rate in the control arm is markedly better or worse than in the trials used to set M1, which is direct evidence against constancy and is almost never discussed when it occurs.
PerfectPaper reads the registration record, the statistical analysis plan reference, the sample-size paragraph and the reported intervals together, and reports where a non-inferiority claim is not supported by the margin the manuscript itself declares. The check belongs to the wider clinical trial reporting lane, and it exists because the defect is one of correspondence between sections rather than of calculation within one.
“The non-inferiority margin is not justified; please state how it was derived from the effect of the active control relative to placebo.” “The margin appears to preserve less than half of the control’s established effect.” “The confidence interval for the primary outcome includes the prespecified margin, so non-inferiority is not established; the conclusion should be revised.” “Only an intention-to-treat analysis is presented. Given the design, a per-protocol analysis is required and the two should be compared.” “There is no evidence that the active control was effective in this population; assay sensitivity cannot be assumed.” “The registration record describes a superiority trial.” “The authors conclude the treatments are equivalent; the design supports a one-sided claim only.” “Loss to follow-up exceeded 15% and would bias this design toward the stated conclusion.”
Two of these are the hardest to answer. The margin-justification comment is not satisfiable by adding a sentence — if the margin cannot be anchored to a placebo-controlled estimate, the trial’s central claim is unanchored and the honest revision is a change of conclusion. The assay-sensitivity comment is the most difficult, because the evidence that would answer it does not exist within the trial; the answerable version is a constancy argument naming the historical trials and comparing populations, event rates and background care.
Both are methods objections rather than statistics objections in practice, even when a statistical reviewer raises them, and framing the response as a design argument rather than a recalculation usually goes better.
Report the margin, its derivation, and the historical evidence it came from, in the methods, with the M1 estimate and the preservation fraction stated as numbers. Name the analysis method as fixed-margin or synthesis. Prespecify both analysis populations and report both, with a CONSORT flow diagram accounting for every exclusion. Present the treatment difference with its two-sided 95% confidence interval and the margin on the same figure. State the direction of the test and the one-sided alpha. Address constancy explicitly by comparing the current population and control-arm event rate with the historical trials. Confine the conclusion to what the interval supports, and if the trial was inconclusive, say inconclusive rather than similar.
The CONSORT extension for non-inferiority and equivalence trials, published by Piaggio and colleagues in JAMA in 2006 and updated in 2012, specifies the reporting items, and journals that require CONSORT for randomised trials will expect a non-inferiority manuscript to meet it. Following it is the simplest available protection against the reviewer comments above, and against the effect-size objection that follows a margin nobody can defend.
Confounding · Competing risks · Statistics objections
Checked before submission by the clinical trials review lane, which compares the declared margin, the registration record and the reported confidence interval, and names the specific claim each one supports.
Non-inferiority means a new treatment has been shown to be worse than an active control by no more than a prespecified margin. The claim is one-sided and bounded: it rules out a clinically meaningful disadvantage, and it says nothing about equality, superiority, or whether the treatment works in absolute terms.
The purpose of a non-inferiority trial is to evaluate a treatment whose advantage lies outside efficacy — better tolerability, oral administration, less monitoring — against an established treatment when a placebo arm would be unethical. The trial asks whether the efficacy sacrificed for that advantage stays within an acceptable bound.
A superiority trial tests a null hypothesis of no difference and wins by rejecting it. A non-inferiority trial tests a null hypothesis of meaningful inferiority and wins by rejecting that instead. Failing a superiority trial does not establish non-inferiority, and the margin must be prespecified for the second claim.
A non-inferiority margin is the largest disadvantage in efficacy, fixed before the data are collected, that investigators are willing to accept in exchange for the new treatment’s other advantages. It is usually set to preserve a stated fraction of the active control’s established effect against placebo, commonly half.
Compare the confidence interval for the treatment difference with the margin, not the p-value. Non-inferiority holds only if the entire interval falls on the acceptable side of the margin. An interval that includes the margin means the trial is inconclusive, however favourable the point estimate looks.
Use a non-inferiority design when an effective treatment already exists, a placebo arm would be unethical, the active control has a well-characterised effect from placebo-controlled trials, and the new treatment offers a genuine non-efficacy advantage. Without a quantified control effect, no defensible margin can be derived.
No. Non-inferiority bounds the difference in one direction only, so it cannot exclude the possibility that the new treatment is better. Equivalence requires bounding the difference in both directions and is a separate design, used for bioequivalence and some device comparisons.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect