Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Crossing curves mean the hazard ratio you reported is an average of effects that reversed direction, and the log-rank test is at its least powerful exactly there.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Two Kaplan-Meier curves that cross are telling you the treatment effect changed direction over time. That is a finding, not a plotting problem — and it means two things you reported are misleading: the hazard ratio, which has averaged an effect that reversed, and the log-rank p-value, which is at its least sensitive precisely when curves cross.
Cox regression assumes proportional hazards: that the ratio between groups is roughly constant over follow-up. Crossing curves are the clearest possible violation of that assumption. A single hazard ratio summarising them is an average over a period when one group did better and a period when the other did, weighted in a way that depends on when events happened rather than on anything clinically meaningful.
That weighting is not a nuisance you can name and move past. The partial likelihood of Cox’s 1972 model gives each event time an influence tied to the risk set at that instant, so under non-proportional hazards the estimate absorbs the censoring pattern and the length of follow-up you happened to accrue. Extend accrual by twelve months and the same underlying biology returns a different number. A hazard ratio estimated across crossing curves is not an effect size another trial could reproduce; it is a summary of one dataset’s event times.
The log-rank test compounds it. It weights all event times equally and sums differences, so an early advantage and a late disadvantage partly cancel. Curves that visibly separate twice can produce a comfortably non-significant p-value.
The arithmetic of that cancellation is why a trial can look null and not be. Work through an illustrative case: an arm carries an excess of ten deaths across the first six months, then avoids fourteen over the following two years. The log-rank statistic adds those observed-minus-expected contributions with equal weight, so most of the late advantage is spent paying off the early deficit and the residual is small relative to its own standard error. The curves are visibly apart at two years and the p-value is comfortably above 0.05. Checkpoint-inhibitor trials in solid tumours have repeatedly produced this shape — in IMvigor211 (Powles et al., Lancet 2018;391:748–757) the primary overall-survival comparison in the PD-L1 IC2/3 population was not significant on its log-rank analysis (HR 0.87, 95% CI 0.63–1.21, p=0.41), even though the atezolizumab advantage persisted late: 24-month overall survival was 23% versus 13% at a median 33 months of follow-up (van der Heijden et al., European Urology 2021;80:7–11). When the hazards are not proportional, the log-rank test is no longer the most powerful test available, and statistical power calculated on a proportional-hazards assumption overstates what the study could have detected.
There is a deeper reason the hazard ratio misbehaves here, and it applies even without visible crossing. Hernán set it out in “The hazards of hazard ratios” (Epidemiology 2010;21:13–15): a hazard at time t is defined only among those who have not yet had the event by t, so a hazard ratio at 24 months compares two groups that randomisation no longer makes exchangeable — the more effective arm has carried its frailer patients further into follow-up. Depletion of susceptibles drags hazard ratios toward 1 over time even when the treatment effect is constant. Crossing curves are the visible end of that continuum; the milder cases are invisible by eye, which is why the formal check below belongs in every survival analysis rather than only in the ones that look wrong.
Crossing in the right-hand tail of a Kaplan-Meier curve is usually small numbers, and the numbers at risk beneath the axis are how you tell. Pocock, Clayton and Altman set out the reporting standard in Lancet 2002;359:1686–1689: print the numbers at risk under the time axis, and do not read the far right of a survival plot as though it carried the same information as the left.
The mechanism is the estimator’s own arithmetic. Kaplan-Meier steps are multiplicative: at a death with 12 patients still at risk, the estimate falls by one twelfth of its current value — about 4 percentage points when survival sits near 50%. With 5 at risk, the same single death moves the curve 10 percentage points. Two or three late events can therefore carry one curve across the other and back without anything having changed in the biology. The pointwise confidence intervals from Greenwood’s formula widen accordingly, and in that region they almost always overlap.
Three checks separate a real reversal from tail noise. First, locate the crossing time against the numbers at risk: a working rule that survives review is to treat crossing as unresolved once fewer than roughly 10–15% of each arm’s randomised patients remain at risk, and to say so rather than interpret it. That is a rule of thumb, not a threshold with published operating characteristics, and it should be labelled as one in the text. Second, ask whether the curves separate again after crossing or merely touch and part; a single touch is not a reversal. Third, check whether one curve is flat because of administrative censoring rather than because patients are surviving — a plateau made of censoring marks is not a plateau.
A treatment with early harm and late benefit. Surgery, intensive induction, transplant. Early mortality followed by a durable advantage is a real and clinically important pattern. Allogeneic transplant is the textbook shape: peri-procedural mortality concentrated in the first 100 days, then a lower long-run relapse hazard. The hazard ratio here is the worst possible summary, because its value depends almost entirely on how long follow-up ran relative to the early excess.
Delayed effect. Immunotherapy is the canonical case: curves overlap for months and then separate. Reported as a single hazard ratio, a genuine long-term benefit is diluted by the flat early period. The dilution is proportional to the delay: a six-month lag inside a trial with median follow-up of eighteen months puts a third of the observation period into a window where the treatment effect is, by construction, zero.
A cured fraction. Where a subgroup is effectively cured, the curve plateaus and the hazard falls toward zero in one group but not the other. Mixture cure models — the formulation traces to Berkson and Gage in 1952 — split the population into a cured fraction and an uncured fraction with its own survival distribution, and they estimate the two separately instead of averaging them into one ratio. If your discussion already argues that a subset of patients never recur, a cure model states that argument formally rather than leaving it in the prose.
Unmodelled heterogeneity. A mixture of responders and non-responders can produce crossing at the population level even where each subgroup is proportional. Frailty models make the unobserved variation explicit as a random effect on the hazard; the practical point is that the crossing may be an aggregation artefact, so the question to answer in the paper is whether a pre-specified biomarker or risk stratum separates the two populations.
Small numbers at risk late. Curves at the tail are estimated from few patients and wander. Apparent crossing among the last handful is often noise — check the numbers at risk before interpreting it.
Differential censoring or a competing risk. Curves can also cross because one arm loses patients to a competing event or to informative censoring rather than to the endpoint under study. Kaplan-Meier treats a competing death as though the patient could still have the event, which inflates one arm’s apparent event probability; the correct object is a cumulative incidence function. Where deaths from other causes are common, treat this as a competing-risks question before treating it as a proportional-hazards question, because the two have different remedies.
Crossing curves come from six recognisable patterns, and the pattern determines the analysis rather than the other way round. Identify yours before choosing a method, because a late-weighted test applied to tail noise and a cure model applied to a delayed effect are both wrong in ways a reviewer will name.
| Pattern | Curve signature | Analysis that fits |
|---|---|---|
| Early harm, late benefit | Early separation one way, sustained reversal after a defined point | Piecewise hazard ratio with a pre-specified cut point, plus RMST at a horizon past the crossover |
| Delayed effect | Curves superimposed for months, then a widening gap that persists | Pre-specified late-weighted log-rank test; RMST at a long horizon; landmark analysis |
| Cured fraction | One curve plateaus and stays flat with events continuing in the other | Mixture cure model reporting the cured fraction and the uncured survival separately |
| Unmodelled heterogeneity | Crossing at the population level, proportional within strata | Stratified model, or a frailty term, with the stratifying variable pre-specified |
| Late-tail noise | Crossing only after most patients are censored or gone | Report numbers at risk and decline to interpret the tail |
| Differential censoring or competing risk | One arm flat with dense censoring marks; other-cause events common | Cumulative incidence function and a cause-specific or subdistribution model |
Check the assumption formally. Scaled Schoenfeld residuals against time, and a plot of them; a trend indicates non-proportionality. Do this routinely rather than only when curves visibly cross, because milder violations do not show up by eye. The residuals come from Schoenfeld (Biometrika 1982;69:239–241) and the scaled test of trend from Grambsch and Therneau (Biometrika 1994;81:515–526), implemented as cox.zph in the R survival package and as estat phtest in Stata. Two cautions belong with the output. The test depends on the time transform — identity, log, rank and Kaplan-Meier transforms can disagree on the same data, so state which one you used. And its power is low with few events, so a non-significant result in a trial with 60 deaths is not evidence that hazards are proportional; report the plot alongside the p-value and let the reader see the slope.
Report restricted mean survival time. RMST is the area under the curve up to a chosen horizon, and it remains interpretable when hazards are not proportional. The difference in RMST between arms has a direct meaning — average time gained over that period — and does not assume anything about the hazard shape. Royston and Parmar set out the design and analysis case in BMC Medical Research Methodology 2013;13:152, and Uno and colleagues in Journal of Clinical Oncology 2014;32:2380–2385. Two specifications make it reviewable. The horizon τ must be pre-specified and no longer than the smallest of the largest observed follow-up times across arms, because the area cannot be estimated past the data; picking τ after inspecting the curves is choosing the answer. And the difference should be reported in its own units — “2.1 months of additional survival over 36 months” — which lets a reader test it against a minimal clinically important difference in a way a hazard ratio never permits. rmst2 in the R package survRM2 is the reference implementation.
Or model the time-dependence. Fit an interaction between treatment and time, or split follow-up into intervals and report a hazard ratio for each. Pre-specify the cut points if you can; chosen after seeing the data, they need to be labelled exploratory. A landmark analysis is the disciplined version: fix a landmark time, exclude patients with an event before it, and analyse from the landmark onward. Do that badly and you manufacture immortal time bias, so the landmark must be set in the protocol and the excluded patients counted in the flow diagram. Flexible parametric survival models with time-dependent effects, and Aalen’s additive hazards model, both estimate a time-varying difference without forcing it through a constant ratio.
Or use a weighted test. If you expect a delayed effect, a weighted log-rank test that emphasises later times is more powerful than the standard one. Pre-specify the weighting, because choosing it afterwards is choosing your p-value. The Fleming-Harrington G(ρ,γ) family is the standard machinery, weighting each event time by S(t)^ρ(1−S(t))^γ: G(0,0) is the ordinary log-rank test, G(1,0) up-weights early event times, and G(0,1) up-weights late ones, which is the form a delayed effect calls for. One implementation detail catches people out: the rho argument of survdiff in the R survival package reaches only the G(ρ,0) half of that family, so it will give you early weighting and cannot give you the late-weighted test — that needs a package written for non-proportional hazards. Where the shape is genuinely unknown in advance, the max-combo test evaluates several weightings together and controls the type I error over the set — which is the honest form of the manoeuvre, because running four weighted tests and reporting the smallest p-value is p-hacking, and any such set needs explicit multiplicity control if it is run at all.
RMST is the usual answer, and there are four situations where you cannot simply switch to it. Each has a defensible fallback, and each requires you to say in the paper which one you took.
The log-rank test is the locked primary analysis. Report it as specified — do not substitute a test that was not pre-specified and present it as primary. Add RMST, weighted tests or piecewise estimates as sensitivity analyses, labelled post hoc, and state the crossing as an interpretive limitation of the primary result rather than as a reason to re-run it. Regulatory and journal reviewers accept a pre-specified test that lost power; they do not accept a swapped estimand described as though it had always been the one.
Follow-up is truncated in one arm. When one arm’s longest observed time is much shorter, τ is forced down, and an RMST difference over a short horizon may not capture the reversal at all. Say what τ was and why it was bounded, and report the piecewise hazard ratios as the complement, since they can describe the period beyond τ that RMST cannot.
You only have published curves, not patient-level data. For a meta-analysis or a re-analysis of someone else’s trial, reconstruct pseudo-individual data from the digitised curve and the numbers at risk using the algorithm of Guyot, Ades, Ouwens and Welton (BMC Medical Research Methodology 2012;12:9). RMST, weighted tests and time-varying models can then be computed. State that the data are reconstructed and that the reconstruction depends on the published numbers at risk being complete.
The crossing comes from channelling rather than from treatment. In registry and electronic health record analyses, early excess events in one arm often reflect who received that treatment first rather than what it did. No survival-analysis method repairs that; the fix is at the design layer — active comparator, new-user design, a clone-censor-weight approach — and a statistical remedy applied on top of a confounded contrast produces a better-behaved estimate of the wrong quantity.
Say the hazards are not proportional and show the evidence. Reporting a hazard ratio while a figure on the same page shows crossing curves is the version that draws a reviewer comment, because the figure contradicts the statistic.
Then report a measure that survives the violation, and describe the pattern in words: early excess mortality, separation after eight months, and what that means for a patient deciding.
Three concrete requirements follow from that. Name the estimand explicitly — the ICH E9(R1) addendum on estimands and sensitivity analysis, adopted in November 2019, makes this the expected form, and “difference in restricted mean survival time at 36 months” is an estimand while “the treatment effect” is not. Make the figure carry the evidence: numbers at risk beneath the axis at regular intervals, censoring marks visible, and follow-up shown only as far as the data support. And put the direction in the abstract, because an abstract that reports a hazard ratio for crossing curves and nothing else will be quoted that way for years, whatever the discussion says.
Write the limitation as a specification rather than an apology. “Hazards were non-proportional (scaled Schoenfeld residuals, p=0.004); we therefore report the difference in restricted mean survival time at 36 months as the primary measure and present the hazard ratio only for comparability with prior trials” is a sentence a reviewer accepts. “Results should be interpreted with caution” is not. The same discipline applies across the rest of a limitations section.
Reviewers rarely write “the proportional hazards assumption is violated” as their opening line. They write the specific version, and these are the comments that decide papers.
“Figure 2 shows the curves crossing at approximately eight months, yet a single hazard ratio is reported; please justify the proportional hazards assumption or report an alternative measure.” “No test of proportional hazards is presented.” “The numbers at risk are not shown, so the reader cannot judge whether the late separation is supported.” “The weighted log-rank test appears to have been selected after the primary analysis; please state where it was pre-specified.” “The abstract reports a non-significant p-value while the discussion argues for a durable late benefit; these cannot both be the conclusion.” “Median survival is reported for a curve that plateaus above 50%, so the median is not estimable in that arm.” “Deaths from other causes are frequent in this cohort and Kaplan-Meier has been used where a cumulative incidence function is required.”
The last two are the ones authors most often fail to anticipate, and they belong to the wider family of statistical objections that arrive when a summary statistic and a figure disagree. A reviewer who writes “the wrong test was used“ is usually not disputing the arithmetic; they are saying the estimand does not match the picture.
Competing risks and time-related bias checks whether the interpretation in the text matches the estimand the model actually produces, and reports where a cause-specific model is discussed as though it estimated absolute risk. If deaths from other causes are common in your cohort — usual in cancer populations — that agent also covers whether Kaplan-Meier was the right estimator at all, which is a separate and equally consequential question.
PerfectPaper reads the survival figure, the numbers at risk and the reported hazard ratio together, and reports where a curve shape contradicts the statistic drawn from it. For randomised work, clinical trial review additionally checks that the analysis presented matches the one the protocol specified, and that a composite endpoint is not hiding components whose effects run in opposite directions — the same reversal problem, one level down.
Related symptom: my p-value got smaller when I measured more cells.
Crossing means the direction of the treatment effect changed during follow-up. Curve crossing is a direct violation of the proportional hazards assumption, so a single hazard ratio averages opposing effects and misrepresents both.
Only with the violation stated and an alternative measure alongside. A hazard ratio under crossing curves has no stable interpretation, because its value depends on the distribution of event times rather than on a constant effect.
Restricted mean survival time is the usual answer: the area under each survival curve up to a pre-specified horizon, with the between-arm difference read as time gained. RMST assumes nothing about the hazard shape, which is exactly what crossing curves take away from you.
Yes — show the evidence rather than asking the reader to trust the figure. Plot scaled Schoenfeld residuals against time with a test of trend, and run that check routinely, since violations that matter are often not visible in the curves at all.
Usually not. Late crossing among few remaining patients is normally noise, because each event moves the curve by a large step once the risk set is small. Check the numbers at risk; if they are small, say so rather than interpreting the tail.
Because the log-rank test weights every event time equally and sums the differences, so an early disadvantage cancels a late advantage. A delayed effect is the common case: the treatment does nothing for the first several months, and that flat window contributes noise to the sum rather than signal.
A late-weighted log-rank test from the Fleming-Harrington G(ρ,γ) family, most often G(0,1), or a max-combo test if the timing is uncertain — pre-specified in the protocol in either case. Report the difference in restricted mean survival time alongside it, because the weighted test gives a p-value while RMST gives a number of months.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect