Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
A composite endpoint counts a participant as having an event when any one of several specified outcomes occurs, such as death, myocardial infarction or stroke.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A composite endpoint is a single trial outcome that counts a participant as having had an event when any one of several specified component outcomes occurs — for example death, myocardial infarction or stroke, whichever comes first. Composite endpoints raise the event rate, so a trial reaches adequate power with fewer participants, and they avoid testing several outcomes separately.
That definition is uncontroversial. Everything difficult about composite endpoints follows from one consequence of it: the analysis treats a death and an unplanned revascularisation as the same event, and the reader of the abstract usually cannot tell which one moved the result.
Composite endpoints solve a power problem and a multiplicity problem in the same move. The FDA’s Multiple Endpoints in Clinical Trials guidance for industry, finalised in October 2022, defines a composite endpoint as one “defined as the occurrence or realization in a subject of any one of the specified components”, and gives the reason plainly: “the incidence of each of the events may be too low to allow a study of reasonable size to have adequate power; the composite endpoint can provide a substantially higher overall event rate.”
The multiplicity argument is the second half. If cardiovascular death, myocardial infarction and stroke were three separate primary endpoints, the trial would need to spend its alpha across three tests. Combined into one variable tested once, the FDA notes, “no multiplicity problem will occur for this endpoint.” That is a genuine statistical gain, not a rhetorical one — a trial can be smaller, shorter, or both.
The third reason is scientific rather than statistical. A treatment such as an antiplatelet agent in coronary artery disease is expected to prevent several related events, and the trialist frequently does not know in advance which one will move. The FDA describes exactly this case: prevention of “myocardial infarction, stroke, or death”, “possibly without knowledge of which event(s) may be affected.”
A composite endpoint is almost always analysed as time to the first occurrence of any component, and that single convention generates most of the interpretive difficulty. The FDA states it directly: “an event is usually defined as the first occurrence of any of the designated component events.”
Three things follow, and none of them is obvious from a Kaplan–Meier curve. First, every component carries equal weight: a participant who dies contributes one event, and a participant who has an unplanned hospitalisation contributes one event, and the analysis cannot distinguish them. Second, a participant who is hospitalised in month three and dies in month seven contributes only the hospitalisation — the death falls outside the primary analysis and does not appear in the component breakdown. Third, recurrent events are discarded entirely; four heart failure hospitalisations count once.
The FDA is explicit about the second point when discussing decomposition: for participants with more than one event type, “events occurring after the first composite event (e.g., end-stage renal disease or death occurring after a doubling of serum creatinine) would not be counted in the decomposition.” A published component table is therefore a table of first events, not of all events, and the two can differ materially. Time-to-event definitions also carry their own hazards around when follow-up starts; see immortal time bias.
The central criticism of composite endpoints is that their components frequently differ in importance to patients and in how strongly treatment affects them, and that these two gradients run in opposite directions. Ferreira-González and colleagues examined this in BMJ in 2007 in a systematic review of randomised cardiovascular trials reporting composite primary endpoints.
The pattern they described has two halves that compound each other. The less severe components — hospitalisation, revascularisation, symptomatic worsening — occur far more often than death, so they supply most of the events in the composite. And the treatment effect was typically larger on those less important components than on mortality. The composite therefore moves largely on the strength of the events patients care about least, while the headline result is read as though it were a statement about the most severe component.
Neither half is a defect in the arithmetic. Both are consequences of pooling outcomes of unequal severity into one count and reporting the pool as a single effect estimate.
The FDA guidance names the same failure mode as an interpretive limit: the effect on the composite “will not be a reasonable indicator of the effect on all of the components or an accurate description of the drug’s benefit if the clinical importance of different components is substantially different and the treatment effect is chiefly on the least important event.”
Decomposition means reporting how the first composite events divided among the components, and it is the first thing an informed reader looks for. The canonical shape of the problem is a renal trial whose primary endpoint is the first occurrence of doubling of serum creatinine, end-stage renal disease or death — as in RENAAL (Brenner et al., New England Journal of Medicine, 2001). Doubling of creatinine is the most frequent of the three: in the FDA’s decomposition of RENAAL, 52% of first composite events were doublings of serum creatinine, 29% were deaths and 19% were end-stage renal disease events. A reader who sees only the composite hazard ratio has no way to recover that split.
Two cautions attach to any such table. It counts first events only, so deaths following a non-fatal component are absent; the FDA recommends additionally analysing “all events for the event type of interest (even those that occur after events of other event types)”. And component-level tests are structurally weak: “testing for individual component endpoints is likely to be underpowered as the sample size or total number of events is usually planned for testing the composite endpoint.”
This creates a trap for authors. A non-significant component result is not evidence of no effect on that component, and writing it up as though it were invites the post hoc power objection. The FDA’s own rule on adjustment turns on intent: examining components to understand the composite requires no multiplicity adjustment, while claiming a distinct effect on a component does.
The published record on how composite outcomes are actually reported is poor, and it has been examined systematically. Cordoba, Schwartz, Woloshin, Bae and Gøtzsche reviewed randomised trials with composite outcomes in BMJ in 2010 and reported three recurring defects: authors almost never justified why those particular components were combined; the components combined were frequently of very unequal clinical importance; and the composite was described inconsistently across the abstract, the methods and the results of the same paper.
That last finding is the most practically useful one for an author. The defect is usually not a wrong analysis; it is that the composite is defined one way in the protocol, described more loosely in the methods, and reported in the abstract as if it were the most severe of its components. Pre-specification discipline is the remedy, and it is checkable against the registry record — see trial registration and endpoint discipline.
CONSORT 2010 item 6a already required completely defined, pre-specified primary and secondary outcome measures, including how and when they were assessed, and its successor, the CONSORT 2025 statement (Hopewell et al., published simultaneously in JAMA, BMJ and The Lancet in 2025), keeps outcome definition as a core reporting requirement. A composite is completely defined only when each component’s diagnostic criteria and ascertainment window are stated.
“Major adverse cardiac events” and “MACE” name no fixed set of components, and treating the acronym as self-explanatory is a recognised defect rather than a stylistic quibble. Kip, Hollabaugh, Marroquin and Williams set out the consequences in the Journal of the American College of Cardiology in 2008, in a paper whose subject was precisely that the coronary-intervention literature used the same three letters for materially different endpoints.
Their demonstration was to apply several different published MACE definitions to one cohort of patients treated with drug-eluting stents. The definitions did not agree: the same patients, the same follow-up, and a different answer depending on which components the acronym was taken to include. Their recommendation was to evaluate safety and effectiveness outcomes separately rather than to rely on a non-standardised label.
The practical rule for a manuscript is that the acronym must never appear without its expansion at first use, including whether it contains all-cause or cardiovascular death, whether it includes revascularisation, and whether periprocedural myocardial infarction counts and under which universal-definition type. Two papers reporting “MACE” are not comparable until those questions are answered, and a meta-analysis that pools them without answering them is pooling different quantities.
A composite endpoint of non-fatal events analysed alongside death raises a competing-risks problem that a standard Cox model does not handle. When death is a component, the competing-risks problem does not arise by construction, because death cannot preclude the endpoint — it is the endpoint. When death is excluded from the composite, participants who die can no longer experience the component events, and censoring them treats death as though it were uninformative loss to follow-up, which it is not.
This is the most common technical error attached to composites of hospitalisation, progression or morbidity that leave mortality out. Cause-specific hazards and subdistribution hazards answer different questions, and the choice must be stated rather than defaulted into by the software. The consequences are set out in more detail under competing risks reporting, and unexpected crossing curves are frequently the visible symptom — see my survival curves cross.
Hierarchical composite endpoints rank the components by clinical priority instead of counting whichever came first, and they are the main methodological answer to the gradient problem. The win ratio, introduced by Pocock, Ariti, Collier and Wang in the European Heart Journal (2012;33:176–182), forms matched pairs of treated and control participants, compares each pair on the most important outcome first (cardiovascular death), and only if that is tied moves to the next (heart failure hospitalisation). The win ratio is wins divided by losses; the authors illustrated it with EMPHASIS-HF, PARTNER B and CHARM, and the FDA’s 2022 multiple-endpoints guidance cites the paper when endpoints “are ordered based on clinical importance”.
Two related methods share the logic. The Finkelstein–Schoenfeld test (Statistics in Medicine, 1999) combines mortality with a longitudinal measure through pairwise comparisons on the same hierarchical principle, and the win ratio can be understood as a way of reporting that comparison as an interpretable ratio. Desirability of outcome ranking (DOOR) generalises the idea to an ordinal scale of overall patient outcome, in which each participant is placed on a ranked scale that combines efficacy and harm; Evans and colleagues introduced it for antibacterial trials in Clinical Infectious Diseases in 2015, and the approach has since been applied well beyond cardiology.
State the limitations honestly, because reviewers do. A win ratio is not a hazard ratio and has no risk-difference interpretation; its value depends on the follow-up distribution and on how ties are handled; and the unmatched form can be sensitive to censoring patterns that differ between arms. Hierarchical endpoints fix the weighting problem and introduce an estimand that is harder to explain to a clinician.
“Composite endpoint” and the “composite variable strategy” of ICH E9(R1) are different concepts that share a word, and conflating them produces confused methods sections. ICH E9(R1), the addendum on estimands and sensitivity analysis adopted on 20 November 2019, uses “composite variable strategy” to describe one of five ways of handling intercurrent events: “An intercurrent event is considered in itself to be informative about the patient’s outcome and is therefore incorporated into the definition of the variable.”
The worked case in the addendum is treatment discontinuation for toxicity being scored as treatment failure, or death being assigned a value on a physical-functioning scale that reflects the absence of function. ICH E9(R1) also observes that “progression-free survival in oncology trials measures the treatment effect on a combination of the growth of the tumour and survival” — a construct that is simultaneously a composite endpoint and a composite variable strategy, which is precisely why the terms get muddled.
The distinction that matters in writing: a composite endpoint answers “what counts as an event”; a composite variable strategy answers “what do we do when something happens that changes the meaning of the measurement”. A modern methods section should specify both, and specifying one is not specifying the other.
A composite endpoint interacts badly with a non-inferiority margin, and this is the least-discussed of its problems. A non-inferiority margin is justified against the effect size of the active control on the endpoint in question; adding low-severity, high-frequency components inflates the overall event rate, which makes a fixed relative margin correspond to a larger absolute number of severe events that may be tolerated as “non-inferior”.
The asymmetry compounds it. In a superiority trial, adding a component the treatment does not affect biases the estimate towards the null and reduces the sponsor’s power. In a non-inferiority trial, the same dilution moves the estimate towards the margin and works in the sponsor’s favour. Where a composite is used as a non-inferiority endpoint, the margin justification must be tied to the composite as constituted, not borrowed from a trial that used a narrower one.
Reviewers of trials read these design choices as claims about clinical equivalence, and overstatement here is treated severely; the general pattern is covered under overclaiming and causal language discipline.
Detection is a small set of concrete checks against the manuscript text, not a matter of judgement. Run them in this order.
Compare the abstract’s endpoint wording to the methods. If the abstract says “reduced cardiovascular events” and the methods define a composite including revascularisation, the abstract has silently promoted the composite to its most severe component.
Look for the component table. If results are reported only for the composite, the paper is not decomposable and the FDA’s recommendation that “results for each component event should therefore be individually examined and should be included in study reports” is unmet.
Read the component table for direction and share. Compute what proportion of first events came from the least severe component. If the composite is significant and the mortality component points the other way, the discussion must address it explicitly.
Check the registry record. Compare the registered primary outcome to the reported one, component by component, and note any addition, removal or change of ascertainment window.
Check the ascertainment. Look for a clinical events committee, blinded adjudication, and pre-specified diagnostic criteria for each component. Soft components — hospitalisation, revascularisation, worsening symptoms — are the ones most sensitive to unblinded ascertainment, and they are also the ones that usually drive the result.
Check the competing-risks handling if death is not a component, and check whether recurrent events were discarded when the disease is one of recurrence.
“The composite is driven by the least severe component.” “The choice of components is not justified, and their clinical importance is not comparable.” “Please report each component separately, with event counts and effect estimates.” “The abstract conclusion refers to cardiovascular events, but the composite includes unplanned revascularisation, which is ascertainment-sensitive in an open-label design.” “The composite as reported differs from the registered primary outcome.”
Two more appear frequently and are harder to answer after the fact. “Deaths occurring after a non-fatal component are not accounted for in the analysis” is a comment about the time-to-first-event convention that cannot be resolved without re-analysis. “Non-inferiority was declared on a composite whose margin was derived from a differently constituted endpoint” is a design objection, and it is generally fatal at review.
If a reviewer’s objection is that the composite was driven by the least important component, the honest response is a component table and a narrowed claim, not a defence of the composite. Guidance on framing that reply without conceding more than the data require is under how to write a response to reviewers, and the broader class of statistical objections under reviewer objections on statistics.
Define every component with its diagnostic criteria, its ascertainment window, and who adjudicated it. State whether death is all-cause or cause-specific, and whether periprocedural events count. Name the counting rule — time to first event, recurrent events, or hierarchical — in the methods rather than leaving it to be inferred from the software.
Report the composite and every component in the same table, with event counts and effect estimates, and say in the text which component contributed the largest share of first events. Where mortality is a component, report it separately as a pre-specified secondary endpoint as well; the FDA guidance describes this arrangement for clinically critical endpoints too infrequent to serve as a primary.
Match the claim to the endpoint. If the composite includes revascularisation, the conclusion is about the composite, not about “cardiovascular events” in general. A single sentence stating what the composite does and does not license — written before the abstract is drafted — prevents most of the review comments above.
Competing risks · Trial registration and endpoints · Causal language · When an effect size is not meaningful · Clinical trial manuscripts
Checked before submission by trial registration and endpoint discipline, which compares the reported primary outcome to the registered one component by component and flags composites whose abstract claim is broader than the endpoint supports.
A composite endpoint means a single outcome variable that registers an event whenever any one of several specified component events occurs in a participant. The FDA’s October 2022 multiple-endpoints guidance defines it as the occurrence in a subject of any one of the specified components, most often analysed as time to the first such occurrence.
A composite outcome is the same construct as a composite endpoint: several distinct clinical events combined into one variable, counted as occurring when the earliest of them occurs. The two terms are used interchangeably in trial reports, with “outcome” more common in epidemiology and “endpoint” more common in regulatory and cardiovascular trial writing.
Cardiovascular death, non-fatal myocardial infarction and non-fatal stroke combined as time to first event is the standard example. Outside cardiology, the RENAAL trial’s renal composite of doubling of serum creatinine, end-stage renal disease or death is often cited, and it shows the usual pattern: the least severe component supplies most of the first events — 52% of RENAAL’s first composite events were doublings of serum creatinine, against 29% deaths and 19% end-stage renal disease.
The first occurrence of any component counts as one event, and every component counts equally. A death and an unplanned hospitalisation each contribute a single event, later events in the same participant are not counted, and a death that follows a non-fatal component is absent from the component breakdown entirely.
A composite endpoint is defined by listing its components with diagnostic criteria and ascertainment rules, then analysed as time to first occurrence of any of them, usually by Kaplan–Meier and Cox methods. Hierarchical alternatives such as the win ratio rank components by clinical priority rather than counting whichever event arrived first.
MACE, or major adverse cardiac events, is a composite endpoint whose components are not standardised. Kip and colleagues showed in the Journal of the American College of Cardiology in 2008 that applying different published MACE definitions to the same patients produces different results, so the acronym must always be expanded at first use — including whether death is all-cause or cardiovascular, and whether revascularisation is counted.
A hierarchical composite endpoint ranks its components by clinical importance and compares participants on the most important one first, breaking ties with the next. The win ratio (Pocock et al., European Heart Journal 2012;33:176–182) and the Finkelstein–Schoenfeld test are the standard implementations, and neither yields a hazard ratio or a risk difference.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect