Skip to content

SOLUTIONS

Research methods concepts, explained plainly

Short, precise explanations of the bias, design and inference terms that decide whether a paper survives review — each with how to detect it in your own work.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Research methods concepts, explained plainly

Research methods concepts are the named failure modes of study design, measurement and inference, and every term here names a specific way a study can produce a confident wrong answer. Each page states what the concept is, how it arises, how to detect it in your own manuscript, and what a reviewer will say if you miss it.

Fifteen concepts are grouped below into three families: causal structure, time and selection, and inference and error. The grouping is by mechanism rather than by field, because the same structure arrives in a mouse experiment, a registry analysis and a randomised trial under three different names, and the remedy follows the mechanism rather than the name it was given locally.

What every concept on this page has in common

Every concept on this page describes a study that returns a wrong number and shows no sign of it. The estimate carries an ordinary confidence interval, the model returns ordinary diagnostics, and nothing in the output separates it from a correct result. That is what makes these fifteen worth learning as a set rather than looking up one at a time.

Three properties recur. The defect is created at design and found at review, so no choice made during analysis can undo it. A larger sample makes it worse in the only sense that matters, narrowing the interval around a wrong value. And the direction of the error is usually knowable in advance from the structure alone, which is why a limitations paragraph that names the direction and the likely magnitude reads as an analysis, while one that asks readers to interpret the findings with caution reads as an admission.

The magnitude is not hypothetical. Samy Suissa followed 979 Saskatchewan residents for a year after a first hospitalisation for chronic obstructive pulmonary disease and analysed the same inhaled-corticosteroid data twice. The time-fixed analysis used by the cohort studies of the day returned an adjusted rate ratio of 0.69 (95% CI 0.55–0.86); a time-dependent analysis of the identical data returned 1.00 (95% CI 0.79–1.26) (American Journal of Respiratory and Critical Care Medicine 2003;168:49–53). No covariate was added and no patient was removed; only the allocation of person-time before the first dispensing changed. A protective effect became a null effect because of a clock.

Causal structure

Causal-structure concepts decide what a variable is doing in the causal diagram, and therefore whether adjusting for it removes bias, creates bias, or deletes part of the effect being estimated. The three roles — common cause, intermediate, common effect — are not properties of the variable but of its position relative to the exposure and the outcome in the specific question being asked.

The practical consequence is that “we adjusted for potential confounders” is not a statement about validity until the adjustment set is justified variable by variable. A directed acyclic graph is the standard instrument for that justification, because it forces each variable’s role to be declared before the model is fitted rather than defended after a reviewer asks. Westreich and Greenland set out the reporting consequence in the American Journal of Epidemiology (2013;177:292–298): an adjustment set chosen to estimate one exposure’s effect does not license causal interpretation of the other coefficients in the same table, because each of those covariates has its own confounding structure that the model was never built to handle.

Simpson’s paradox is the case where the structural question has a genuine answer and the textbook retelling skips it. Bickel, Hammel and O’Connell reported in Science in 1975 that roughly 44% of 8,442 male applicants and roughly 35% of 4,321 female applicants were admitted to Berkeley’s 1973 graduate cohort, with the disparity largely vanishing department by department. Whether the stratified estimate or the pooled one is correct depends entirely on whether department is a common cause of applicant sex and admission, or an intermediate on the path being studied.

Time and selection

Time-and-selection concepts decide the answer through the study’s boundaries — who entered, who stayed, and when the clock started — before any variable is measured. These are the biases that adjustment does not touch, because the variable that would fix them takes the same value for everyone in the dataset.

Selection has a single structural definition, given by Hernán, Hernández-Díaz and Robins in “A structural approach to selection bias” (Epidemiology 2004;15:615–625): the study conditions on a common effect of two variables, one on the exposure side and one on the outcome side. That definition unifies Berkson’s bias, the healthy worker effect, differential loss to follow-up and non-response into one question you can ask of any manuscript, and it explains why the remedy is a design change or a bound rather than a longer covariate list.

The three screening concepts form a family that a stage-adjusted analysis cannot separate. Lead-time bias concerns when a case is found, length-time bias concerns which cases are found, and overdiagnosis concerns cases that never needed finding at all; a screening study that reports improved five-year survival without reporting disease-specific mortality has answered none of the three. Immortal time is arithmetic rather than sociology: person-time that carries zero events by construction is added to the exposed denominator, so the exposed rate falls. Registry and electronic health record analyses concentrate all of these at once, which is why registry data limitations is worth reading alongside them.

Inference and error

Inference-and-error concepts decide what repeated testing, non-independent observations and underpowered designs do to a p-value that still prints as significant. These are the failures that live in the statistics section rather than the methods section, and they are the most detectable of the three families from the manuscript alone.

Benjamini and Hochberg defined the false discovery rate in 1995 (Journal of the Royal Statistical Society B) as E[V/R], the expected proportion of rejected hypotheses that are true nulls. Choosing between that and family-wise error control is a claim-level decision, not a software default: control the family-wise rate when any single reported finding must stand on its own, and the false discovery rate when the claim is about a list. Genomics conventions and confirmatory trial conventions differ for exactly that reason, and the choice is the substance of multiple testing in omics and of the reviewer comment collected at reviewer: multiple comparisons.

Power fails in two directions. Button and colleagues estimated a median statistical power of 21% across 49 meta-analyses covering 730 studies in neuroscience (Nature Reviews Neuroscience, 2013), which sharply reduces the probability that any one significant finding in such a literature reflects a true effect, before flexible analysis is considered. Power computed after the fact from the observed effect is worse than uninformative: Hoenig and Heisey showed in The American Statistician (2001) that observed power is a monotone transformation of the p-value, so a non-significant result must yield low observed power by arithmetic and cannot be evidence about itself.

Pseudoreplication is the easiest of the four to see. Stuart Hurlbert introduced the term in Ecological Monographs (1984;54:187–211) in a survey where pseudoreplication occurred in 27% of the studies examined, and in 48% of those that applied inferential statistics at all; Lazic later read a single issue of Nature Neuroscience and classified 12% of its papers as pseudoreplicated with a further 36% suspected (BMC Neuroscience 2010;11:5). If your n counts wells, cells, sections or images rather than animals, litters or independent cultures, the objection at reviewer: n is not independent is the one you will receive.

The three questions that locate most methods objections

Three questions locate most methods objections before a reviewer asks them, and each maps onto one family above.

What is the unit of analysis, and how many independent ones are there? Count the units that were randomised, housed, treated or sampled independently, then compare that count with the n in the statistical test. A gap between the two is pseudoreplication, and it is the reason a p-value shrinks when more cells are counted while nothing about the biology changes.

What is each covariate doing in the causal diagram? Take the adjustment set one variable at a time and say whether it causes the exposure, is caused by it, or is caused by both it and the outcome. A variable in the third category is a collider, and adjusting for it manufactures an association — which is also the usual explanation when an effect disappears after adjusting.

When did follow-up start, and was everyone eligible at that moment? Write down time zero for each arm and check that exposure status was determined using information available then. Any criterion requiring survival, adherence or a second visit uses future information to define entry, which creates immortal time on one side and selection on the other.

A fourth question applies whenever more than one hypothesis was tested: how many comparisons were made in total, including those that produced nothing? The denominator, not the correction procedure, is the part reviewers query first, and an undisclosed denominator is the operational definition of p-hacking.

Concepts that are routinely confused with each other

Eight pairs on this page are confused often enough that the confusion has its own signature in review, and each pair separates on a single question.

Pair The question that separates them What goes wrong if reversed
Confounder versus mediator Does the variable cause the exposure, or is it caused by it? Adjusting for a mediator deletes part of the effect you set out to estimate
Confounder versus collider Common cause, or common effect? Adjusting for a collider creates an association that was not there
Selection bias versus confounding Did the study condition on a common effect, or did a common cause exist first? A longer covariate list is offered in answer to a selection objection
Lead-time bias versus overdiagnosis Would the disease have progressed? Survival gains are claimed for cases that never needed detection
Length-time bias versus lead-time bias Which cases are found, or when they are found? Stage adjustment is presented as having addressed both
Regression to the mean versus treatment effect Were participants selected on an extreme baseline value? An uncontrolled before-and-after change is reported as an effect
False discovery rate versus family-wise error Is the claim the list, or any single item on it? A confirmatory endpoint is reported under a false discovery rate threshold
Pseudoreplication versus a small sample Do the measurements come from independent units? A precise interval is reported around an effect estimated from one animal

The last row is the one that reaches print most often, because the two look identical in the output and opposite in the design. The usual presenting symptom is replicates that disagree with each other, and the distinction between technical and biological replication is where the answer starts.

A worked example: one registry cohort, four concepts

Consider a retrospective cohort assembled from a national registry, comparing patients who started a drug after discharge from hospital with patients who did not, and reporting a 31% lower mortality rate in the treated group across 28 secondary outcomes.

Time zero is the first defect. If exposure is defined by having filled a prescription within 90 days of discharge, then every treated patient survived to fill it, and the survival between discharge and the fill has been credited to the drug. Reallocating that person-time is the standard remedy, and a landmark analysis with a pre-specified landmark is the standard implementation.

Selection is the second. Registry membership requires a healthcare encounter, and the comparator “everyone else in the database” is a group defined by having generated encounters — itself a function of both illness and the exposure. This is a structural problem with the comparator rather than a covariate problem, so propensity score matching will balance the measured variables without touching it.

Multiplicity is the third: 28 secondary outcomes tested at the conventional 0.05 threshold produce roughly 1.4 false positives under the complete null, and reporting the two that reached significance without the denominator is not a reporting omission but a change of claim. The fourth is the Table 2 fallacy — the covariate coefficients in the mortality model will be read as effects of age, comorbidity and deprivation by at least one reader, and by at least one reviewer as evidence that the authors did not distinguish an adjustment set from a set of causal estimates.

None of these four are visible in the effect estimate, the confidence interval or any goodness-of-fit statistic. All four are visible in the methods section.

What is detectable from the manuscript alone

Many of these problems are detectable in a finished manuscript without access to the data. Undefined units of analysis, missing correction, adjustment sets containing mediators, and time-related exposure definitions are all visible in the text, because each of them is a sentence the authors either wrote or did not write.

The reporting checklists are the fastest route to those sentences, and each names the item to quote. STROBE, a 22-item guideline for observational studies published in October 2007, asks at item 13 for numbers at every stage from potentially eligible to analysed, which is how denominators that do not reconcile become visible — see the STROBE checklist. CONSORT 2025, published on 14 April 2025 with 30 items, up from 25 in CONSORT 2010, asks at item 21d which analyses were pre-specified rather than post hoc, and at item 21b who is included in each analysis and in which group — the two items that decide whether intention to treat was actually applied. PRISMA 2020, 27 items across seven sections, asks at item 14 for the reporting-bias assessment that publication bias requires. ARRIVE 2.0, 21 items split into an Essential 10 and a Recommended Set of 11, asks at sub-item 3c for the exact value of n in each experimental group for each analysis, which is the single most efficient pseudoreplication check available to a reader.

Two things are not detectable from the text: whether an unmeasured variable exists, and whether the data are what the authors say they are. A bound is the appropriate response to the first, and a bounded, quantified limitation reads as an analysis where an acknowledgement reads as an admission.

What reviewers write when these concepts are missed

Reviewers rarely name the concept. They write the specific version of it, and these are the sentences that decide papers.

“The number of animals is not stated per group; the n appears to be the number of images.” “Follow-up begins at diagnosis for the unexposed group and at treatment initiation for the exposed group.” “The adjustment set includes variables on the causal pathway between exposure and outcome.” “Twenty-eight outcomes are analysed and two are reported as significant, with no adjustment for multiplicity and no statement of the total number of comparisons.” “The improvement in the high-baseline group is consistent with regression to the mean.” “The abstract reports five-year survival; disease-specific mortality is not reported.” “The coefficients in Table 2 are discussed as if each were a causal effect.” “Controls do not appear to be drawn from the source population that gave rise to the cases.”

If you arrived here holding one of those sentences, the objection is already located: reviewer objections on statistics and reviewer objections on methods are organised by what the reviewer wrote rather than by what the concept is called.

What this collection covers, and what sits next to it

The concepts collected here are design and inference structures for original research articles: the ones that make an estimate wrong. The collection does not cover reporting format, which the checklist pages above handle, and it does not cover writing tasks such as drafting a cover letter or a response to reviewers.

Two adjacent families are worth naming because they are frequently filed here by mistake. Measurement and modelling artefacts — batch effects, overfitting — degrade an estimate through the data-generating and model-fitting process rather than through causal structure, and their remedies are design randomisation and held-out validation rather than adjustment. Reporting and interpretation quantities — effect size, confidence intervals — describe an estimate that is already either structurally sound or structurally wrong; a narrow interval reports precision, never validity.

How to use this

Each page is written to be read in two minutes by someone who has just encountered the term in a review, and to be precise enough to act on. Where a concept has a check that catches it before submission, the page names it.

If you arrived from a reviewer comment rather than from a term, reviewer objections is organised by what the reviewer said. If you arrived from something odd in your own results, analysis symptoms is organised by what you observed.

If you are drafting rather than responding, read the three questions above against your own methods section before the first submission, and read the family that matches your design in full: causal structure for observational comparisons, time and selection for anything with follow-up or screening, inference and error for anything with more than one test. The same structures recur across fields under different local names, so the family a concept belongs to matters more than the label a reviewer reached for.

PerfectPaper reads a manuscript for these structures directly — the unit of analysis against the stated n, the adjustment set against each variable’s causal role, time zero against the eligibility criteria, and the number of comparisons against the correction applied — and reports which of them the text does not resolve.

Review my manuscript

Frequently asked questions

A reviewer named a concept I do not recognise. Where do I start?

Start with the concept page itself, then read its detection section against your own methods before drafting a reply. Research methods concepts are named structures in design, measurement and inference — confounding, selection, immortal time, multiplicity, pseudoreplication and their relatives — and each names a mechanism precise enough to imply both the check and the remedy. The reviewer comment is almost always about one specific sentence in your manuscript, and locating that sentence matters more than the general definition.

How do I work out which of these applies to my study?

Work from the design rather than from the vocabulary. An observational comparison of two groups raises causal structure first; anything with follow-up, screening or a time-defined exposure raises time and selection first; anything testing more than one hypothesis raises inference and error first. The three families above are grouped by mechanism for exactly this reason, because the same structure arrives in a mouse experiment and a registry analysis under two different names.

My paper is nearly submitted. Which checks are worth running now?

Four checks are worth running before submission, and each maps onto a family above: the unit of analysis against the n reported in each test, the role of every covariate in the adjustment set, the definition of time zero against the eligibility criteria, and the total number of comparisons made against the correction applied. Each is answerable from the manuscript in a few minutes, and each corresponds to a sentence you either wrote or did not write.

The reviewer said the design is biased but did not name the bias. What is it likely to be?

Confounding, selection bias and time-related biases such as immortal time account for most misleading observational results, and an unnamed objection usually points at one of the three. The distinguishing question is whether a common cause existed before the study, whether the study conditioned on a common effect, or whether exposure status was assigned using information from after time zero. Only the first is addressed by adjustment, which is why the other two are the more dangerous readings of an unnamed comment.

My statistics are clean and the model fits. Could one of these still apply?

Yes, and that is the defining property of the set. Every concept here describes a study that returns a wrong number and shows no sign of it: the interval is ordinary, the diagnostics are ordinary, and nothing in the output separates the result from a correct one. A narrow confidence interval reports precision, never validity, so goodness of fit is not evidence against any objection on this page.

Can these problems be seen in a finished manuscript, or do you need the data?

Many are visible in the text alone. Undefined units of analysis, missing correction, adjustment sets containing mediators, and time-related exposure definitions all show up in the methods section without access to the data. Two things do require more than the text: whether an unmeasured variable exists, and whether the data are what the authors say they are. A quantified bound is the appropriate response to the first.

Everything here comes from epidemiology. Does any of it apply in my field?

Pseudoreplication was introduced in ecology by Hurlbert in 1984 and is now discussed across neuroscience and cell and molecular biology as well; multiplicity governs genomics; regression to the mean applies to any before-and-after comparison in a group selected on an extreme baseline value. The structures are field-independent even where the vocabulary is not, which is why the family a concept belongs to is more useful than the label a reviewer reached for.

Last updated September 10, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.