Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Earlier detection lengthens the interval between diagnosis and death without delaying death, making screened patients appear to survive longer when nothing changed.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Lead-time bias is the apparent survival gain produced by diagnosing a disease earlier without changing when the patient dies. Survival is measured from diagnosis, so moving the diagnosis date earlier lengthens measured survival by exactly the amount of time gained in detection — even if the disease course and the date of death are entirely unaltered.
The distinctive difficulty is that the resulting number is large, statistically significant and entirely real, yet answers no question anyone asked. In the Mayo Lung Project, median survival among participants with resected early-stage lung cancer was 16.0 years in the screened arm and 5.0 years in the usual-care arm — an eleven-year difference — while lung cancer mortality was 4.4 deaths per 1,000 person-years in the screened arm against 3.9 in usual care (Marcus et al., Journal of the National Cancer Institute 2000;92:1308–1316).
A patient whose cancer is detected by symptoms at year five and who dies at year eight has three years of measured survival. Screen the same patient and detect the same cancer at year three, with no change to treatment or outcome, and they now have five years of measured survival.
Nothing about the disease changed. The clock started earlier.
Zelen and Feinleib gave this the formal treatment in “On the theory of screening for chronic diseases” (Biometrika 1969;56:601–614), and their two definitions are the ones to use in a manuscript. Lead time is the interval between the date a screening test detects the disease and the date it would have been diagnosed from symptoms. Sojourn time, or the preclinical screen-detectable period, is the whole window during which the disease is detectable by the test but has not yet produced symptoms. Lead time is bounded above by sojourn time: a test cannot advance diagnosis by more than the length of the detectable-but-silent phase.
The consequence is an identity, not an effect. Measured survival among screen-detected cases equals true survival from the biological onset of a fatal course, plus the lead time. Set the treatment effect to exactly zero and the identity still delivers a longer survival curve for the screened group, a lower hazard ratio, and separation on a Kaplan–Meier curve that will pass a log-rank test at any sample size you like. A larger study makes a lead-time artefact more significant, never less.
Lead-time bias means five-year survival is not a valid endpoint for evaluating screening. Any screening programme that detects disease earlier will improve survival statistics automatically, which is why survival comparisons between screen-detected and clinically detected cases are uninformative about benefit.
The valid endpoint is mortality in the population offered screening, which counts deaths against everyone eligible and is unaffected by when diagnosis occurred.
Welch, Schwartz and Woloshin measured how badly the two endpoints diverge in practice. Across the 20 most common solid tumour types in the SEER programme from 1950 to 1995, five-year survival rose for every single one — from 3 percentage points for pancreatic cancer to 50 points for prostate cancer — while mortality fell for 12 and rose for the other 8. The correlation between a tumour’s change in five-year survival and its change in mortality was Pearson r = 0.00 and Spearman r = −0.07; the correlation between change in survival and change in incidence was Pearson r = +0.49 (JAMA 2000;283:2975–2978). Survival tracked how many people were being diagnosed, not how many were dying.
The same paper’s argument is not common knowledge among the clinicians who act on it. In a randomised survey of 412 US primary care physicians, Wegwarth and colleagues presented one screening test supported by a five-year survival rise from 68% to 99% and another supported by a cancer mortality fall from 2 to 1.6 per 1,000; 69% recommended the test backed by the irrelevant survival statistic against 23% for the test backed by the relevant mortality statistic (P < 0.001), and 47% said that finding more cancers in a screened population “proves that screening saves lives” (Annals of Internal Medicine 2012;156:340–349).
Lead time is measured in years, not months, for the cancers people actually screen for, which is why it overwhelms plausible treatment effects rather than nudging them.
Draisma and colleagues modelled the Rotterdam section of the European Randomized Study of Screening for Prostate Cancer (ERSPC), covering 42,376 men and 1,498 prostate cancers, and reported that a single PSA test at age 55 carried an estimated mean lead time of 12.3 years (range 11.6–14.1), falling to 6.0 years (range 5.8–6.3) at age 75. Overdetection moved in the opposite direction, from 27% at age 55 to 56% at age 75 (JNCI 2003;95:868–878). A twelve-year lead time means a screen-detected prostate cancer diagnosed at 55 will clear a five-year survival endpoint, a ten-year endpoint, and most of the way to a fifteen-year endpoint before the counterfactual diagnosis would even have occurred.
Sojourn time is shorter for breast cancer and, importantly, is not one number. In a Swedish screening cohort covering 1996–2010, sensitivity-adjusted mean sojourn times for small invasive breast cancers differed by mammographic appearance: 4.26 years (95% CI 3.50–5.26) for powdery and crushed-stone calcifications, 3.76 years (3.15–4.53) for stellate masses, and 2.65 years (2.06–3.55) for circular masses (Chang et al., Cancers 2020;12:1855). Any correction that assumes a single mean sojourn time is assuming away that heterogeneity.
The Mayo Lung Project is the cleanest published demonstration that survival and mortality can move in opposite directions in the same randomised trial.
Marcus and colleagues randomised 9,211 male smokers between 1971 and 1983 to chest x-ray plus sputum cytology every four months for six years, or to usual care — in which participants were advised at entry to have the same tests annually — and followed them to a median of 20.5 years. Survival among participants diagnosed with lung cancer before 1 July 1983 was significantly better in the screened arm (two-sided P = 0.0039), and median survival for resected early-stage disease was 16.0 years screened against 5.0 years unscreened. Lung cancer mortality was 4.4 deaths per 1,000 person-years (95% CI 3.9–4.9) in the screened arm and 3.9 (95% CI 3.5–4.4) in usual care, P = 0.09 — numerically higher in the arm with the better survival curve. The authors’ own conclusion names the mechanism: “Similar mortality but better survival for individuals in the intervention arm indicates that some lesions with limited clinical relevance may have been identified in the intervention arm.”
Contrast this with the National Lung Screening Trial, which enrolled 53,454 high-risk participants and reported 247 lung cancer deaths per 100,000 person-years with low-dose CT against 309 with chest radiography, a 20.0% relative mortality reduction (95% CI 6.8–26.7, P = 0.004; New England Journal of Medicine 2011;365:395–409). Both trials produced better stage distributions and better survival in the screened arm. Only one of them produced a mortality difference, and only the mortality difference is evidence.
Infant neuroblastoma screening is the case where the survival–mortality gap was tested twice, prospectively, in cohorts of half a million and of two and a half million children, and closed to nothing both times.
Woods and colleagues offered urinary catecholamine screening at three weeks and six months to all 476,654 children born in Quebec between May 1989 and April 1994, achieving 92% participation, and found standardised ratios for neuroblastoma death against five unscreened comparison populations of 1.11, 0.90, 1.40, 0.96 and 1.39, every confidence interval spanning 1 (NEJM 2002;346:1041–1046). Schilling and colleagues offered screening at one year of age to 2,581,188 children in six German states, with 1,475,773 actually screened — 61.2% of the children born between July 1994 and October 1999 — and 2,117,600 unscreened controls; stage 4 incidence was 3.7 per 100,000 screened against 3.8 per 100,000 control, deaths were 1.3 per 100,000 against 1.2, and the estimated overdiagnosis rate was 7 cases per 100,000 children (95% CI 4.6–9.2) (NEJM 2002;346:1047–1053).
Both programmes detected many more cases, at earlier stages, with excellent survival among those detected. Neither shifted the rate at which children died. A manuscript that reports only the first half of that sentence is reporting lead time.
Lead time, length time and overdiagnosis are frequently packed into a single limitations sentence, and separating them changes what a reviewer expects you to do about each.
Lead time is a property of when a given case is detected. It inflates survival for every screen-detected case, including cases where screening genuinely helps.
Length-time bias is a property of which cases a periodic test tends to catch. Tumours with long sojourn times spend more calendar time in the detectable-but-silent window and are therefore over-represented among screen-detected cases; those same tumours are, on average, the slower and more survivable ones. The Swedish sojourn-time figures above are the mechanism: a test run every two years preferentially harvests the 4.26-year lesions and preferentially misses the 2.65-year ones.
Overdiagnosis is the limiting case of both. A cancer that would never have produced symptoms in the patient’s lifetime has an effectively infinite lead time, contributes a guaranteed survivor to the numerator, and cannot be identified in any individual patient. The German neuroblastoma estimate of 7 per 100,000 is what an overdiagnosis rate looks like when it is measured against a concurrent unscreened population rather than asserted.
All three are forms of selection bias acting on who enters the analysed case series and on what date, which is why none of them is fixed by adjusting for tumour stage, grade, receptor status or comorbidity.
Lead-time bias and immortal time bias both make an intervention look better by manipulating a clock, but they manipulate different ends of it and have different remedies.
Lead-time bias moves the start of follow-up earlier for the exposed group, because diagnosis is the origin and the intervention changed the diagnosis date. Immortal time bias misallocates person-time after the origin, by counting a period during which the patient could not have had the outcome as exposed time.
The remedy differs accordingly. Immortal time is a design-and-coding problem, fixed by a landmark analysis with a pre-specified landmark or by treating exposure as time-varying. Lead time cannot be coded away, because the origin itself is contaminated; it is fixed by changing the endpoint or by modelling the sojourn distribution explicitly.
If your study compares survival between screen-detected and symptom-detected cases, or reports improving survival over a period when detection became earlier, lead time is the first explanation to exclude.
Stage-shift arguments are not sufficient on their own: earlier stage at diagnosis is exactly what lead time produces.
Six further patterns should trigger the check, and all six are readable from the manuscript without access to the data.
The time origin is set by the intervention. Any survival analysis whose t = 0 is the date of diagnosis, date of biopsy, date of first abnormal test or date of registry entry has a clock that the exposure moved. Report which calendar dates define t = 0 in each arm.
Incidence rose without a fall in late-stage incidence. A screening programme that truly advances lethal disease should remove cases from the advanced-stage category, not only add them to the early-stage one. Rising total incidence with flat advanced-stage incidence is the published signature of lead time plus overdiagnosis, and it is what both neuroblastoma trials found.
The comparator is historical. Survival for cases diagnosed in 2015–2020 against cases diagnosed in 2000–2005, in a disease whose detection practice changed in between, measures the change in detection.
Interval cancers are excluded. Cases arising between scheduled screens belong in the screened arm’s denominator. Dropping them leaves a screen-detected series compared against people who never attended, which adds self-selection to lead time.
The design is a single-arm case series. A cohort of screen-detected cases with no unscreened comparator has no quantity in it that can be corrected; the only honest report is stage distribution and case-fatality among those detected, with no comparative claim.
The apparent benefit is implausibly large. A case-fatality hazard ratio of 0.3 to 0.4 for screen detection is the range in which lead time and length time account for a large share of the apparent effect rather than a marginal one: in the West Midlands series an uncorrected 0.34 rose to 0.68 once lead time, length time and self-selection had been removed. See the corrected numbers in the next section for the scale of the deflation.
Lead-time correction is possible, is published, and moves estimates by a large amount — which is the reason to report it rather than to treat mortality as the only option.
Duffy and colleagues set out a tractable method in “Correcting for lead time and length bias in estimating the effect of screen detection on cancer survival” (American Journal of Epidemiology 2008;168:98–104), assuming an exponential distribution for the preclinical screen-detectable period and adding a sensitivity analysis over two latent tumour categories to bound length bias.
Applied to 26,766 breast cancers in women aged 50–74 in the West Midlands from 1988 to 2004, the correction cascade ran as follows (Lawrence et al., Breast Cancer Research and Treatment 2009;116:179–185). The uncorrected relative risk of 10-year breast cancer fatality for screen-detected against symptomatic disease was 0.34 (95% CI 0.31–0.37). Correcting for lead time alone moved it to 0.49 (95% CI 0.45–0.53). Adding length-bias sensitivity analysis gave a range of 0.49–0.59 with a median of 0.51. Correcting further for self-selection, by using interval cancers as the symptomatic comparator, gave a median of 0.68. Roughly half the apparent survival advantage was lead time, length time and self-selection; a real advantage remained.
Two features of that cascade are worth carrying into your own manuscript. First, the correction has a direction you can state in advance: it always shrinks the apparent benefit, so an uncorrected estimate is an upper bound, not an unbiased one. Second, the correction depends on an assumed sojourn distribution, and the Swedish estimates of 2.65 to 4.26 years across mammographic subtypes show that the assumed parameter is not a constant of nature. Report the corrected estimate across a range of assumed mean sojourn times, and name the value at which your conclusion would reverse.
Mortality in the population offered screening is the correct endpoint, and most authors writing about screen-detected disease do not have a randomised trial to hand. Five substitutes exist, in descending order of strength.
Report incidence-based mortality over the whole eligible population. Deaths from the disease among everyone offered or eligible for detection, divided by that whole population, is the endpoint’s observational form. It requires a population denominator, not a case series, and registry linkage often supplies one — while carrying the registry and electronic-record caveats about who generates a record at all.
Report stage-specific incidence trends. Advanced-stage incidence is the sensitive marker. If earlier detection is genuinely intercepting lethal disease, the rate of advanced-stage presentation in the eligible population must fall; if it is flat while total incidence rises, the additional cases were not on their way to becoming the advanced ones.
Use interval cancers as the symptomatic comparator. Comparing screen-detected cases against cancers arising between screens, within the same invited population, removes the self-selection component — the difference between 0.51 and 0.68 in the West Midlands series.
Present a lead-time-corrected estimate with an explicit sensitivity range. State the assumed sojourn distribution, its mean, its source, and the corrected estimate at the lower and upper plausible means.
If none of the above is possible, bound the claim rather than softening it. Write the direction and the magnitude: “an uncorrected 10-year case-fatality ratio of 0.34 is an upper bound on benefit; published corrections for lead time, length time and self-selection in comparable breast series have moved similar estimates to roughly 0.68.” That is a limitations paragraph with a number in it, which is what a limitations section is for.
Do not substitute a surrogate endpoint — stage at diagnosis, tumour size, nodal status, or a biomarker — for mortality and present it as evidence of benefit. Stage at diagnosis is the variable lead time directly manipulates.
Reviewers usually do not write “lead-time bias” first. They write the specific version, and these are the comments that hold papers.
“Survival is measured from date of diagnosis, which the intervention itself advances; please report mortality in the eligible population.” “The improvement in five-year survival across the study period coincides with the introduction of screening and cannot be separated from it.” “The stage shift reported in Table 2 is the expected consequence of earlier detection and is not independent evidence of benefit.” “Interval cancers appear to be excluded from the screened group; please report them in the invited denominator.” “Total incidence rose over the study period while advanced-stage incidence was unchanged; the authors should address overdiagnosis explicitly.” “Lead-time bias is acknowledged in the discussion but its magnitude and direction are not estimated.” “The comparison group was diagnosed in an earlier calendar period under different detection practice.” “No correction for lead time is applied; at minimum, report the estimate under a plausible range of mean sojourn times.”
The comment about magnitude and direction is the one most often provoked by a limitations paragraph written to pre-empt it. A generic acknowledgement of lead-time bias, with no direction and no magnitude, reads to a methods reviewer as an admission that the analysis was not done.
Report lead-time bias by naming the clock, the denominator and the endpoint, in that order.
State the calendar event that defines t = 0 in each group, and say whether the intervention could have moved it. Give the denominator: everyone eligible, everyone invited, everyone attending, or only those diagnosed — these are four different studies. Name the primary endpoint and justify it, using disease-specific mortality in the eligible population where the design allows and saying plainly where it does not. Report interval cancers. Report advanced-stage incidence alongside total incidence. Where a correction is feasible, give the corrected number and its sensitivity range rather than an adjective.
Keep the causal language matched to the endpoint you actually measured; “screening improved survival” and “screening reduced mortality” are different claims, and only the second is a benefit claim. Following STROBE item 13 for participant flow makes the denominators visible, which is where most lead-time problems become legible to a reader.
Length-time bias · Overdiagnosis · Immortal time bias · Selection bias · Epidemiology review
PerfectPaper reads the time origin, the denominator and the stated endpoint together, and reports where a survival comparison involving screen-detected disease is measuring lead time rather than benefit. Checked by competing risks and time-related bias, which reports survival comparisons involving screen-detected disease where lead time is unaddressed, and by causal language discipline, which flags benefit claims made from a survival endpoint.
Because detecting disease earlier lengthens the interval from diagnosis to death even when the date of death is unchanged. Mortality in the screened population avoids this because it does not depend on diagnosis date. Across the 20 most common solid tumours from 1950 to 1995, change in five-year survival correlated with change in mortality at Pearson r = 0.00.
Lead time is about when disease is detected in a given patient. Length time is about which patients are detected — screening preferentially finds slow-growing disease with better prognosis. Sojourn time links them: the same 2.65-to-4.26-year range that sets how much a diagnosis can be advanced also sets which lesions a periodic test tends to catch.
Estimates of mean lead time can be used to correct survival comparisons, but the correction depends on assumptions about the preclinical period. Using mortality as the endpoint is more reliable than adjusting survival. Duffy’s exponential-sojourn correction moved a West Midlands breast series from a fatality ratio of 0.34 to 0.49, and to a median of 0.68 once self-selection was also removed.
No. A shift toward earlier stage is what lead time produces by itself, so stage shift alone is not evidence of mortality benefit. The stronger stage-based signal is a fall in advanced-stage incidence across the whole eligible population, which lead time does not produce.
Lead time is measured in years. Modelling of the Rotterdam ERSPC data put the mean lead time for a single PSA test at 12.3 years at age 55 and 6.0 years at age 75. Mean sojourn times for small invasive breast cancers in a Swedish cohort ranged from 2.65 to 4.26 years depending on mammographic appearance.
Check what calendar event defines t = 0 and whether the intervention could have moved it; check whether the comparator is historical or unscreened; check whether interval cancers are excluded; and check whether advanced-stage incidence fell while total incidence rose. No statistical test identifies lead-time bias.
No. Overdiagnosis is the limiting case, in which the lead time is effectively infinite because the disease would never have produced symptoms at all. Lead-time bias also affects cases that would have become symptomatic, and it inflates their measured survival even when screening genuinely helps them.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect