Skip to content

SOLUTIONS

What is a surrogate endpoint?

A surrogate endpoint stands in for an outcome patients feel. It is valid only when treatment effects on the surrogate predict treatment effects on the real outcome.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is a surrogate endpoint?

A surrogate endpoint is a measurement used in place of the clinical outcome a trial actually cares about, chosen because it arrives sooner or is easier to measure. Tumour shrinkage stands in for survival; bone mineral density stands in for fracture. The substitution is valid only when treatment effects on the surrogate predict treatment effects on the outcome.

That last sentence carries the entire difficulty. A surrogate can correlate beautifully with the true outcome across patients and still mislead completely about what a drug does, because the two properties are different and only one of them is usually measured. This page covers the regulatory definition, the Prentice criterion, the difference between individual-level and trial-level surrogacy, the four ways surrogates fail, the oncology evidence base, how the problem is detected in a manuscript, and what reviewers write when it is handled badly.

The regulatory definition

United States law defines a surrogate endpoint in section 507(e)(9) of the Federal Food, Drug, and Cosmetic Act as “a marker, such as a laboratory measurement, radiographic image, physical sign, or other measure, that is not itself a direct measurement of clinical benefit, and—(A) is known to predict clinical benefit and could be used to support traditional approval of a drug or biological product; or (B) is reasonably likely to predict clinical benefit and could be used to support the accelerated approval of a drug or biological product in accordance with section 506(c).”

Two tiers sit inside that one sentence, and the distinction is the most useful thing in the definition. A surrogate that is known to predict clinical benefit supports traditional approval. A surrogate that is only reasonably likely to predict it supports accelerated approval, which is conditional and carries a confirmatory-trial obligation. The FDA publishes both categories in its Table of Surrogate Endpoints That Were the Basis of Drug Approval or Licensure, split into adult non-cancer, adult cancer, paediatric non-cancer and paediatric cancer sections, and updates it every six months under section 507.

Authors writing for a clinical journal are held to a standard the statute makes explicit: the burden of showing which tier a surrogate occupies sits with whoever uses it, and “reasonably likely” is not a synonym for validated.

Surrogate endpoint, biomarker, intermediate endpoint

A biomarker is any objectively measured characteristic — a lab value, an image, a genotype. A surrogate endpoint is the small subset of biomarkers that a trial has substituted for a clinical outcome and is prepared to make a treatment decision on. Most biomarkers are not surrogate endpoints, and calling one a surrogate is a claim about evidence, not a description of the assay.

A clinical endpoint measures how a patient feels, functions or survives. Overall survival, symptomatic fracture, stroke, and validated patient-reported symptom scales are clinical endpoints. Overall survival is the reference standard in oncology precisely because it needs no interpretation: it is unambiguous, unblinded assessment cannot bias it, and no adjudication committee is required.

An intermediate clinical endpoint is a third category worth keeping distinct. Progression-free survival and disease-free survival are not biomarkers — they are clinical events with a time axis — but they are also not the outcome of ultimate interest. The FDA treats intermediate clinical endpoints as their own basis for accelerated approval. Blurring the three terms is common in manuscripts and is the source of a specific reviewer objection: a paper that calls circulating tumour DNA clearance “a validated surrogate” when what exists is a correlation study has changed the meaning of the word.

The Prentice criterion

Ross Prentice set out the operational definition in Statistics in Medicine in 1989 (volume 8, pages 431–440): a surrogate endpoint is a response variable for which a test of the null hypothesis of no relationship to treatment is also a valid test of that hypothesis for the true endpoint. In practice the criterion decomposes into four conditions — treatment affects the surrogate, treatment affects the true endpoint, the surrogate is prognostic for the true endpoint, and, the demanding one, the effect of treatment on the true endpoint is fully captured by the surrogate.

The fourth condition is what makes the criterion restrictive. Full capture means that conditioning on the surrogate leaves no treatment effect on the true endpoint at all. Any mechanism by which the drug reaches the outcome without passing through the surrogate — a second pathway, an off-target harm, a toxicity — violates it.

Two practical consequences follow, and both are routinely missed. Verifying full capture requires a trial powered on the true endpoint, which is the trial the surrogate was supposed to make unnecessary. And failure to reject the conditional-independence test is weak evidence, because that test has low power: a trial can be far too small to detect the residual treatment effect it is claiming does not exist. A manuscript reporting “the Prentice criterion was satisfied” from a single randomised trial is reporting a non-significant result in an underpowered test, and should say so.

Individual-level versus trial-level surrogacy

Marc Buyse and Geert Molenberghs formalised the split in Biometrics in 1998, and it is the most useful distinction in the field. Individual-level surrogacy asks whether the surrogate predicts the outcome within a patient, after adjusting for treatment — measured as an association, often reported as R²indiv. Trial-level surrogacy asks whether the treatment effect on the surrogate predicts the treatment effect on the outcome across randomised trials — measured by regressing effect estimates on effect estimates across a meta-analysis, reported as R²trial.

Only trial-level surrogacy licenses the substitution a trial makes. A marker can be strongly prognostic and useless as a surrogate: prognosis is a statement about patients, surrogacy is a statement about interventions, and a drug may move the marker by a mechanism unconnected to the outcome.

The reporting standard that follows is concrete. A trial-level validation needs a set of randomised trials, not a cohort; it needs the number of trials and the number of patients contributing; it needs a confidence interval or prediction interval around R²trial, which is usually wide when fewer than ten trials contribute; and it needs the population, line of therapy and drug class specified, because trial-level surrogacy established for cytotoxic chemotherapy does not transfer to checkpoint inhibitors, whose delayed separation of survival curves changes the relationship. Citing a single correlation coefficient with no denominator and no interval is a common defect in this section of a manuscript.

Why a correlated marker still misleads

Four distinct mechanisms break the link between a surrogate and the outcome it stands for, and naming which one is plausible is more useful than a general caveat.

The drug moves the marker by a pathway that does not reach the outcome. In the ILLUMINATE trial (New England Journal of Medicine, 2007), torcetrapib raised HDL cholesterol by 72.1% at twelve months and increased all-cause mortality, hazard ratio 1.58 (95% CI 1.14 to 2.19, P=0.006). HDL was and remains a strong risk marker; raising it pharmacologically by inhibiting cholesteryl ester transfer protein was not the same intervention as having a high HDL.

The marker captures only part of the mechanism. In the Cardiac Arrhythmia Suppression Trial (New England Journal of Medicine, 1991), encainide and flecainide suppressed ventricular ectopy after myocardial infarction — the surrogate did exactly what it was meant to — and produced 43 deaths from arrhythmia on drug versus 16 on placebo (P = 0.0004). Ectopy suppression was real; proarrhythmia ran alongside it and outside it.

The intervention harms the outcome through a channel the marker cannot see. Riggs and colleagues (New England Journal of Medicine, 1990) gave 202 postmenopausal women with existing vertebral fractures 75 mg of sodium fluoride daily for four years and raised lumbar spine bone mineral density by about 35%. Non-vertebral fractures were 72 in the treated group versus 24 on placebo. The new bone was cancellous, cortical density fell 4% at the radius, and the densitometer read the wrong quantity.

Measurement of the surrogate is itself affected by treatment. Where response or progression is assessed by imaging in an open-label trial, a treating investigator’s judgement about equivocal scans is an unblinded measurement of the surrogate. Overall survival is immune to this; radiographic progression is not, which is why blinded independent central review exists and why its absence belongs in the limitations.

Surrogate endpoints in oncology

Oncology runs on surrogates, and the trial-level evidence supporting them is weaker than practice implies. Prasad, Kim, Burotto and Vandross (JAMA Internal Medicine, 2015) reviewed 36 trial-level meta-analyses reporting 65 surrogate–survival correlations and found that 34 of 65 (52%) were of low strength, r ≤ 0.7, with roughly a quarter strong. Haslam, Hey, Gill and Prasad (European Journal of Cancer, 2019) applied the same approach to a larger body of published trial-level analyses and reached the same headline: strong surrogate–survival correlations were the minority, and the answer depended on which surrogate, which tumour type and which line of therapy was being examined rather than holding across oncology. Gyawali, Hey and Kesselheim (JAMA Internal Medicine, 2019) looked downstream at what confirmation produced, and found that confirmatory trials for 19 of 93 cancer drug indications granted accelerated approval demonstrated an improvement in overall survival.

The bevacizumab breast cancer case is the worked example most oncologists reach for. E2100 (Miller et al., New England Journal of Medicine, 2007) randomised 722 patients and reported median progression-free survival of 11.8 versus 5.9 months with paclitaxel plus bevacizumab, hazard ratio 0.60, and objective response 36.9% versus 21.2% — with overall survival of 26.7 versus 25.2 months, hazard ratio 0.88, P=0.16. The FDA granted accelerated approval on the progression endpoint in 2008 and withdrew the metastatic breast cancer indication in 2011 after confirmatory trials produced smaller progression effects and no survival benefit.

Not every oncology surrogate sits in the same place. Pathological complete response in some neoadjuvant settings and minimal residual disease in some haematological malignancies have stronger evidence than objective response rate does in most solid tumours, and each is specified for a defined population and drug class. The general lesson is that surrogacy is a property of a triplet — surrogate, population, intervention class — never a property of the marker alone. Response criteria carry their own definitional pitfalls, covered separately under RECIST response criteria.

What a surrogate endpoint is legitimately for

Surrogate endpoints earn their place in three situations, and stating which one applies is stronger than defending the surrogate in the abstract.

Screening decisions in early-phase work. A single-arm phase II study using objective response rate to decide whether a compound merits a randomised trial is using the surrogate for exactly what it supports: a go/no-go judgement inside a development programme, not a claim of patient benefit.

Outcomes with long latency. Where the clinical event takes fifteen years — fracture in early osteoporosis, cardiovascular death in primary prevention, progression in indolent lymphoma — a trial on the true endpoint may answer a question about a drug class that has already been superseded by the time it reports. The trade being made is explicit uncertainty for feasibility.

Rare diseases with few events. Where the eligible population worldwide numbers in the hundreds, a survival-powered trial cannot be run, and a surrogate with mechanistic support plus a registry follow-up commitment is the honest structure.

Surrogate endpoints are not appropriate as the basis for a definitive claim of clinical benefit in a common disease where a survival trial is feasible, and they are not appropriate for comparative-effectiveness claims between agents with different toxicity profiles, because the toxicity channel is exactly what a surrogate cannot see. Where the question is about how long patients live, the endpoint is how long patients live; time-to-event complications such as competing risks and crossing survival curves are separate problems and do not justify substituting a marker.

How surrogate misuse is detected in a manuscript

Detection is textual, not statistical, and it works by comparing four places in the paper against each other.

The registered primary endpoint versus the reported one. Open the ClinicalTrials.gov or ISRCTN record cited in the methods and read the primary outcome measure and its time frame. A paper whose registry entry names progression-free survival at 12 months and whose abstract leads with overall survival — or the reverse — has an endpoint discrepancy that belongs in the text, with the date of the protocol amendment that caused it. This check is mechanical and is the single highest-yield one available; it is covered in depth under trial registration and endpoint discipline.

The sample size calculation versus the conclusion. A power calculation stating detectable hazard ratios for progression, followed by a discussion concluding that the drug “prolongs survival”, is a mismatch between what the trial was built to detect and what it claims. Overall survival in such a trial is a secondary, usually immature, usually underpowered endpoint, and a non-significant survival result there is not evidence of no effect.

The title and abstract conclusion versus the results table. Surrogate slippage happens in the compression. The results table says “median PFS 8.1 versus 6.2 months”; the abstract conclusion says “improves outcomes”; the title says “benefit”. Each step drops a qualifier, and the title is what gets cited.

The validation citation, followed to source. Where a manuscript asserts that its surrogate is validated, read the cited paper and check three things: whether it reports trial-level or only individual-level association, how many randomised trials contributed, and whether the population and drug class match. A citation supporting R²indiv used to justify a trial-level substitution is a category error, and it is frequent.

Two further tells are worth noting. A composite endpoint containing a surrogate component — for example a composite of death, myocardial infarction and revascularisation, where revascularisation is discretionary and unblinded — inherits the surrogate’s weaknesses while presenting as a clinical endpoint, and requires component-wise reporting. And a discussion that argues from mechanism to benefit (“given the established role of this pathway, the observed response translates into…”) is making the inference the data did not, which is the same defect discussed under correlative not causal and causal language discipline.

What reviewers say when it is mishandled

“The primary endpoint is a surrogate and the conclusions are framed in terms of clinical benefit.” “No trial-level validation is provided for this surrogate in this population.” “The reference cited in support of surrogacy reports patient-level association only.” “Overall survival was a secondary endpoint and the trial was not powered for it; the non-significant result should not be presented as absence of harm.” “Progression was assessed by investigators in an open-label design without central review.” “The registered primary outcome differs from the one reported here and the change is not explained.” “The title should reflect the endpoint actually measured.”

The last of those looks cosmetic and decides papers. Editors read the title as the claim; a title that says “improves survival” over a progression-free survival result is the version of the error that survives into citation, into guidelines and into practice, and it is the one an editor can insist on fixing without renegotiating the science. Anticipating it takes a sentence; a related pattern is set out under overclaim checking.

How to report a surrogate endpoint honestly

Name the surrogate as a surrogate in the abstract, in the same sentence as the effect estimate, and keep the qualifier attached wherever the number travels. Report the true endpoint alongside it whenever it was collected, including when the result is null and including when follow-up is immature — with the number of events, because an immature survival analysis with 41 events is a different object from one with 400.

State the surrogacy evidence specifically: which validation study, trial-level or individual-level, how many randomised trials, which population, which drug class, and the interval around the estimate. Where no trial-level validation exists in your setting, say that plainly rather than citing an adjacent one; an attributed absence is a stronger sentence than a stretched citation.

Specify the assessment mechanics, because they determine how much the surrogate can be biased: assessment schedule and whether it was identical across arms, whether central review was blinded, how censoring was handled, and what the sensitivity analyses showed. Unequal assessment intervals between arms manufacture progression-free survival differences on their own.

Finally, describe the direction and magnitude of the residual uncertainty rather than acknowledging it. “If the treatment effect on progression does not translate, the survival benefit could be null; the confirmatory trial reads out in 2028” is a specification. “Surrogate endpoints have limitations” is not.

Contested points, stated plainly

Whether progression-free survival should be accepted for regulatory approval is genuinely disputed, and a manuscript that pretends otherwise reads as naive to a methodological reviewer. The case against rests on the trial-level correlation figures above and on the record of confirmatory trials. The case for rests on unblinded crossover, which contaminates overall survival in trials where control-arm patients receive the experimental agent on progression, and on subsequent lines of therapy that dilute a real first-line effect. Both arguments are legitimate; the crossover argument is verifiable in a specific trial by reporting the crossover rate, and a paper that invokes it without that number has asserted rather than argued.

The strength of the correlation evidence is also contested on method. Trial-level R² estimates are sensitive to which trials enter the meta-analysis, to whether effect estimates are weighted, and to the small number of trials available in most settings — which is why prediction intervals matter more than point estimates. Critics of the surrogacy literature and its defenders are frequently disagreeing about the estimator rather than the drug.

A third open question is whether surrogacy established in one drug class transfers to another with a different mechanism. The empirical answer, for checkpoint inhibitors relative to cytotoxics, appears to be no; the theoretical answer has been clear since Prentice, because full capture is a mechanism-specific property. Manuscripts still cite chemotherapy-era validations for immunotherapy trials.

Related

Confounding · Lead-time bias · Immortal time bias · Effect size not clinically meaningful · Peer review for clinical trials

PerfectPaper reads the whole manuscript rather than the abstract, so a claim of survival benefit built on a progression endpoint is caught in the three places it usually appears — the title, the abstract conclusion and the discussion — and not only where the endpoint is defined in the methods.

Checked before submission by trial registration and endpoint discipline, which compares the reported primary endpoint against the registry record and flags conclusions framed in terms of an outcome the trial did not measure.

Review my manuscript

Frequently asked questions

What does surrogate endpoint mean in a clinical trial?

A surrogate endpoint in a clinical trial is a substitute measure — a lab value, an image finding, or an earlier clinical event — recorded in place of the outcome that matters to patients, because it arrives sooner or more often. The substitution holds only if the treatment effect on the substitute predicts the treatment effect on the real outcome.

What is a surrogate marker?

A surrogate marker is the same thing as a surrogate endpoint, stated in biomarker language: a measured characteristic standing in for a clinical outcome. Not every biomarker is a surrogate marker. Calling one a surrogate asserts that evidence links treatment effects on it to treatment effects on the outcome, which is a claim requiring a citation.

What is an example of a surrogate endpoint?

Bone mineral density is a surrogate endpoint for fracture. Riggs and colleagues (New England Journal of Medicine, 1990) raised lumbar spine density by about 35% with sodium fluoride in 202 women and recorded 72 non-vertebral fractures on treatment against 24 on placebo. The surrogate moved in the intended direction while the outcome moved the other way.

What is the difference between a surrogate endpoint and a clinical endpoint?

A clinical endpoint measures how a patient feels, functions or survives — death, stroke, symptomatic fracture. A surrogate endpoint measures something else and is accepted as a stand-in. Overall survival is a clinical endpoint; tumour response is a surrogate. The difference is not precision but relevance: the surrogate requires an argument, the clinical endpoint does not.

Is progression-free survival a surrogate endpoint?

Progression-free survival is used as a surrogate for overall survival and is more precisely an intermediate clinical endpoint, since it counts events rather than measuring a biomarker. Its trial-level correlation with survival is weak in many settings: Prasad and colleagues (JAMA Internal Medicine, 2015) found 34 of 65 oncology surrogate–survival correlations at r ≤ 0.7.

What makes a surrogate endpoint valid?

Validity requires trial-level evidence that treatment effects on the surrogate predict treatment effects on the outcome across randomised trials, in the same population and drug class. Prentice’s 1989 criterion demands that the surrogate fully capture the treatment effect. Patient-level correlation, however strong, does not establish validity and is the substitution most often made in manuscripts.

Why do trials use surrogate endpoints instead of survival?

Trials use surrogate endpoints to reach an answer sooner, with fewer patients, where the clinical outcome takes years or the disease is rare enough that a survival-powered trial cannot be run. The trade-off is explicit: the FD&C Act distinguishes surrogates “known to predict” clinical benefit from those only “reasonably likely to predict” it, and the second class supports conditional approval only.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.