Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Screening preferentially detects slow-growing disease, which spends longer in a detectable preclinical state. Screen-detected cases have better prognosis by selection.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Length-time bias is the tendency of screening to preferentially detect slow-growing disease. A tumour that spends years in a detectable but asymptomatic state has many opportunities to be caught by a periodic screen; one that progresses to symptoms in months has few. Screen-detected cases are therefore enriched for indolent disease with better prognosis — before any treatment effect.
Length-time bias is length-biased sampling applied to sojourn time — the interval during which a disease is detectable by the screening test but has not yet produced symptoms. Zelen and Feinleib set out the formal framework in Biometrika (1969;56:601–614), modelling disease as a passage from a preclinical detectable state into a clinical state and deriving mean lead time as a function of observable quantities. Sojourn time is the quantity screening samples on, and it is not the quantity anyone reports.
The arithmetic is unforgiving. If sojourn times in the population have density f and finite mean μ, the cases picked up by a single prevalence screen do not follow f. They follow the length-biased density t·f(t)/μ, because a case is available to be caught in proportion to how long it sits in the detectable window. The mean sojourn time among prevalence-detected cases is therefore E[T²]/E[T] = μ(1 + CV²), where CV is the coefficient of variation of the sojourn distribution. For an exponential sojourn distribution CV = 1, so the screen-detected mean is exactly 2μ — twice the population mean, from the distributional form alone, with no assumption about tumour biology.
Two consequences follow that manuscripts routinely get backwards. First, the size of the bias tracks the heterogeneity of sojourn time, not its average: a disease whose detectable window is nearly the same length in every patient generates almost no length bias, while one mixing six-month and ten-year windows generates a great deal. Second, the bias does not disappear as the screening interval shortens. Shortening the interval raises the detection probability for every sojourn length at once; as long as fast disease can still surface symptomatically between rounds, the ratio that drives the enrichment survives.
Length-time bias and lead-time bias are two different distortions of the same survival comparison, and a paper can control one while leaving the other untouched. Lead-time bias concerns when a given case is detected. Length-time bias concerns which cases are detected at all.
They compound. A screening programme both moves diagnosis earlier and selects the slower disease, so survival comparisons between screen-detected and clinically detected cases are affected twice over.
The practical test is what a fix would have to change. Correcting lead time requires an estimate of how far back the diagnosis date moved; correcting length time requires an estimate of how the composition of the detected group differs in underlying prognosis. Subtracting an estimated lead time from screen-detected survival — the most common correction in submitted work — leaves length-time bias entirely intact, because it adjusts a date rather than a mixture.
Length-time bias never arrives alone. Any comparison of screen-detected against symptom-detected cases carries lead time, length time, and self-selection — the tendency of screening attenders to be healthier, more adherent and lower-risk than non-attenders — and the three are multiplicative rather than alternatives.
Lawrence and colleagues quantified the stack on a population registry series (Breast Cancer Research and Treatment 2009;116:179–185): 26,766 breast cancers in women aged 50–74 in the West Midlands, 1988–2004, including 10,100 screen-detected cases, 6,009 interval cancers and 9,853 in women who had not attended screening. The uncorrected relative risk of 10-year breast-cancer fatality for screen-detected versus symptomatic cases was 0.34 (95% CI 0.31–0.37). Correcting for lead time moved it to 0.49 (95% CI 0.45–0.53). Sensitivity analyses for length bias gave a corrected range of 0.49–0.59 with a median of 0.51. Correcting additionally for self-selection, by using interval cancers as the comparator, gave a median of 0.68.
Read the sequence rather than the endpoint: 0.34 → 0.49 → 0.51 → 0.68. Roughly half of the apparent survival advantage in the naive comparison was bias, and the authors — who concluded a real advantage remained after correction — could only say so because they estimated each component explicitly. A paper that reports 0.34 and stops has not reported a smaller effect than they did; it has reported a different quantity.
Length-time bias makes screening look effective in exactly the analysis most authors reach for first. Comparing outcomes between screen-detected and interval or symptom-detected cancers will favour screening even if screening changes nothing, because the two groups differ in the biology of the disease detected, not only in when it was found.
The extreme form is overdiagnosis: disease so slow it would never have caused symptoms in the patient’s lifetime, which contributes a perfect survival record to the screened group.
Note what this does to the estimand. A hazard ratio comparing screen-detected with symptom-detected cases is a comparison of two differently-composed mixtures, so its proportional-hazards assumption is violated by construction: the indolent fraction concentrates in one arm and the hazard ratio drifts with follow-up time. Curves drawn this way often converge or cross late, which is a crossing-survival-curves problem with a structural cause rather than a modelling one, and a Kaplan–Meier curve stratified by mode of detection cannot be repaired by choosing a different model. Case fatality and 5-year survival are affected identically; the defect is in the comparison, not the summary statistic.
The Mayo Lung Project is the cleanest published demonstration that better survival and no mortality benefit can coexist. Marcus and colleagues reported extended follow-up in the Journal of the National Cancer Institute (2000;92:1308–1316) for a randomised trial of 9,211 male smokers conducted between 1971 and 1983, in which the intervention arm was offered chest x-ray and sputum cytology every 4 months for 6 years.
At a median follow-up of 20.5 years, lung cancer mortality was 4.4 deaths per 1,000 person-years (95% CI 3.9–4.9) in the intervention arm and 3.9 per 1,000 person-years (95% CI 3.5–4.4) in usual care, two-sided P = .09 — no benefit, with the point estimate on the wrong side. Yet among participants diagnosed with lung cancer before 1 July 1983, survival was better in the intervention arm, two-sided P = .0039, and median survival for patients with resected early-stage disease was 16.0 years in the intervention arm versus 5.0 years in usual care.
Sixteen years versus five, and no lives saved. The trialists’ own reading is the sentence to quote: “Similar mortality but better survival for individuals in the intervention arm indicates that some lesions with limited clinical relevance may have been identified in the intervention arm.” An 11-year median survival gap in resected early-stage disease is what length-time bias plus overdiagnosis looks like when the randomised mortality endpoint is available to contradict it. In a non-randomised study the same 11 years would have been the headline result.
Neuroblastoma screening in Germany produced a stark contrast in case fatality by mode of detection with a null mortality result attached. Schilling and colleagues offered urine screening at approximately one year of age to 2,581,188 children in 6 of 16 German states (New England Journal of Medicine 2002;346:1047–1053), with 2,117,600 children in the remaining states as controls; 1,475,773 children, or 61.2% of the birth cohort, were screened.
Neuroblastoma was detected by screening in 149 children, of whom 3 died — about 2%. Fifty-five children with negative screening tests were later diagnosed, of whom 14 died — about 25%. On mode-of-detection reasoning, screening cut case fatality more than tenfold.
It did not. The incidence of stage 4 neuroblastoma was 3.7 per 100,000 screened children (95% CI 2.7–4.7) against 3.8 per 100,000 controls (95% CI 2.9–4.6), and deaths among children with neuroblastoma were 1.3 per 100,000 screened (95% CI 0.7–1.8) against 1.2 per 100,000 controls (95% CI 0.7–1.7). The authors estimated overdiagnosis at 7 cases per 100,000 children (95% CI 4.6–9.2). Screening found a different population of tumours — the biologically favourable ones that regress — while the aggressive tumours arrived symptomatically between screens and in the control states at the same rate.
Length-time bias is detected by asking what defines the comparison groups, not by any statistical test. If your study compares outcomes by mode of detection, length time applies. If it reports that screen-detected cancers are lower grade or slower growing and treats that as evidence of benefit, that is the bias rather than the effect.
The design that avoids it is a randomised comparison of the offer of screening, analysed by mortality across everyone randomised, regardless of what was detected or how.
Four specific tells recur in submitted manuscripts.
The paper adjusts for stage, grade or proliferation index and reports the association attenuated. Tumour stage, histological grade and Ki67 are consequences of the sojourn time that produced the detection, so conditioning on them is conditioning on a mediator of the selection itself. The adjusted estimate is not an unbiased estimate, and interpreting the covariate coefficients as effects is a Table 2 fallacy.
Interval cancers are used as a favourable comparator without acknowledging what they are. Interval cancers are the short-sojourn tail by construction, and they carry aggressive biology accordingly — in a population-based Norwegian series of 282 cases, 200 screen-detected and 82 interval, blood vessel invasion was the strongest single factor predicting interval presentation (Journal of Clinical Pathology 2017;70:313–319). Comparing screen-detected against interval cases removes self-selection and maximises length-time contrast.
Case fatality falls while population mortality does not. Compute both. Divergence between them is the signature, and it is the same signature the Mayo Lung Project and the German neuroblastoma programme produced.
The stage distribution shifts but advanced-stage incidence does not fall. Welch and colleagues applied this test to SEER data for women aged 40 and over (New England Journal of Medicine 2016;375:1438–1447): after screening mammography spread, small tumours rose from 36% to 68% of detections and large tumours fell from 64% to 32%, but the incidence of large tumours fell by only 30 cases per 100,000 women while small-tumour detection rose by 162 per 100,000, implying 132 per 100,000 overdiagnosed. A favourable stage distribution with a flat advanced-stage incidence rate is arithmetic evidence of length-biased detection, not of earlier detection.
Reviewers rarely write the words “length-time bias” first. They write the specific version, and these comments decide screening papers.
“Survival is compared by mode of detection; screen-detected and symptom-detected cancers are not exchangeable populations, and the comparison is uninterpretable as a measure of screening effect.” “The authors report that screen-detected tumours were lower grade and node-negative more often, and present this as evidence of benefit; this is the expected consequence of length-biased detection.” “Please report incidence-based mortality in the whole invited population rather than case fatality among detected cases.” “Adjustment for stage does not address length-time bias, because stage is downstream of the sojourn time that determined detection.” “The comparator is non-attenders, so the estimate carries self-selection as well as length and lead time; please state the direction and, if possible, the magnitude of each.” “The abstract states that screening improved 5-year survival; the trial’s pre-specified endpoint is disease-specific mortality, which is not reported.” “No estimate of overdiagnosis is provided, although the excess incidence in the screened arm persists after the final screening round.”
The last of those is the one authors most often fail to anticipate. Persistent excess cumulative incidence in the screened arm, after screening has stopped and enough follow-up has elapsed for the lead time to be exhausted, is the standard excess-incidence estimator of overdiagnosis — and reviewers of screening work expect to see it computed, not discussed.
When randomisation is unavailable, length-time bias is bounded rather than removed, and the bound is reported as a number. Four options exist, in descending order of strength.
Change the endpoint to incidence-based mortality in the whole invited population. Count deaths from the disease among everyone invited or eligible, not among cases, over a fixed period. This is the observational analogue of the intention-to-treat principle and it discards mode of detection entirely, which is precisely why it works.
Apply an explicit correction with sensitivity analysis over the length-bias parameter. Duffy and colleagues published a workable method (American Journal of Epidemiology 2008;168:98–104): a lead-time correction assuming an exponential preclinical screen-detectable period, combined with a sensitivity analysis that posits two latent tumour categories — one more prone to screen detection and correspondingly less prone to death from the cancer — and varies the mixture across plausible magnitudes. They demonstrated it on 25,962 West Midlands breast cancer cases diagnosed 1988–2004. The output is a range, and the range is the finding.
Use interval cancers as the symptomatic comparator to strip out self-selection, and state the trade-off. Interval cases attended screening, so they match attenders on health-seeking behaviour; they are also the short-sojourn extreme, so the residual length-time contrast is larger than against unscreened symptomatic cases. Reporting both comparators, as Lawrence and colleagues did, brackets the answer instead of choosing a convenient end of it.
Report excess incidence with sufficient post-screening follow-up. Cumulative incidence in screened versus unscreened populations, followed until the lead time is exhausted, estimates overdiagnosis directly. Follow-up that stops while screening is still running will attribute lead time to overdiagnosis and inflate it.
Two things are worth stating about the limits of modelling. Ryser and colleagues showed that model misspecification can substantially bias mean sojourn time estimates, and that clinical follow-up after the last screening round is what makes the indolent fraction precisely estimable — in the Canadian National Breast Screening Study 2, 1980–1985, that fraction was not precisely identifiable (American Journal of Epidemiology 2019;188:197–205). And no amount of covariate adjustment substitutes for any of the above: length-time bias is a form of selection bias, and like every member of that family it is not fixed by a longer covariate list or a propensity score model built on post-detection variables.
Length-biased sampling appears wherever the probability of being observed is proportional to the duration of a state, and cancer screening is only its most-discussed instance. Any cross-sectional sample of an ongoing process oversamples long episodes: sampling ICU patients on a given day oversamples long stays, and sampling current unemployment spells oversamples long spells, by the same t·f(t)/μ weighting.
The clinical research version with the largest footprint is prevalent-user bias in pharmacoepidemiology. Ray described it in American Journal of Epidemiology (2003;158:915–920): studies that enrol people already taking a drug select survivors of the early treatment period, so an agent whose hazard is front-loaded looks safe, and baseline covariates measured at entry are themselves consequences of the drug. The new-user design — restricting analysis to people observed from the start of the current course — removes it. The structural relationship to immortal time bias is close but not identical: immortal time misallocates follow-up, while prevalent-user and length-time bias misallocate membership.
The same weighting affects prevalence-based cohorts of any chronic condition, registry cohorts assembled from people currently under care (registry and EHR analyses are the usual site), and any surrogate endpoint defined by detection rather than by disease onset.
Report length-time bias by naming the comparison, the direction and the magnitude, in that order.
State the primary endpoint and whether it is mortality in an invited population or survival among detected cases; do not let the abstract use the second to make a claim about the first. Report sojourn time or programme sensitivity if you estimated it, with the method and its assumptions named. Report interval cancer rates by round, since they bound how much fast disease the programme is missing. Where survival is compared by mode of detection at all, present the uncorrected estimate, the lead-time-corrected estimate and the length-bias sensitivity range side by side. Give the direction length-time bias would push your estimate and why, rather than writing that results “should be interpreted with caution” — a limitations section that names a bias without naming its direction reads as an admission rather than an analysis.
Keep the causal language matched to the design. “Screen-detected cancers had better survival” is a description of a selected mixture; “screening improved survival” is a claim the design does not support, and the gap between them is the single most common overclaim in screening manuscripts. Observational screening comparisons describe association among differently-composed groups, which is the substance of the correlative-not-causal objection when it lands on this literature.
Lead-time bias · Overdiagnosis · Selection bias · Immortal time bias · Epidemiology review
PerfectPaper reads the endpoint, the comparison groups and the covariate list together, and reports where a survival advantage attributed to screening is an artefact of which cases the screen was able to find. Checked before submission by causal language discipline, which flags a mode-of-detection survival comparison presented as a screening effect.
Because slow disease remains detectable but asymptomatic for longer, so a periodic screen has more chances to catch it. Fast-progressing disease often appears symptomatically between screening rounds. Formally, the probability of screen detection rises with sojourn time, so the detected cases follow the length-biased density t·f(t)/μ rather than the population density f.
Length time is about which cases are detected — screening selects indolent disease. Lead time is about when a case is detected — the diagnosis date moves earlier. Both inflate survival in screen-detected groups. Subtracting an estimated lead time corrects only the date and leaves the length-time enrichment untouched.
Randomise the offer of screening and compare mortality across all randomised participants, rather than comparing outcomes between screen-detected and symptom-detected cases. Where randomisation is unavailable, use incidence-based mortality in the whole invited population, or report a length-bias sensitivity range using the two-latent-category approach of Duffy and colleagues (American Journal of Epidemiology 2008;168:98–104).
Overdiagnosis is the extreme end of the same phenomenon: disease so slow-growing it would never have caused symptoms. Length-time bias describes the general enrichment for indolent disease. Overdiagnosis is estimated as excess cumulative incidence in the screened group after screening stops and the lead time is exhausted; the German neuroblastoma programme put it at 7 cases per 100,000 children (95% CI 4.6–9.2).
Large enough to reverse the conclusion. In the Mayo Lung Project, median survival for resected early-stage lung cancer was 16.0 years in the screened arm versus 5.0 years in usual care, while lung cancer mortality was 4.4 versus 3.9 deaths per 1,000 person-years — no benefit. For an exponential sojourn distribution, mean sojourn among prevalence-detected cases is exactly twice the population mean.
No. Stage, histological grade and proliferation index are consequences of the sojourn time that allowed detection, so adjusting for them conditions on a mediator of the selection rather than removing it. A reviewer who raises length time will not be satisfied by a longer covariate list, and the adjusted coefficients should not be read as effects.
The reviewer means the denominator should be everyone invited or eligible, not everyone diagnosed. Survival among detected cases depends on which cases were detected, which is the bias. Disease-specific mortality across an invited population is unaffected by mode of detection, which is why randomised screening trials pre-specify it as the primary endpoint.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect