Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Symptom-first diagnosis for analysis surprises: an effect that vanishes on adjustment, curves that cross, replicates that disagree, a p-value that keeps falling.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
An analysis that looks wrong rarely announces itself as an error message; it announces itself as a surprise. The effect vanishes when you add a control. Two survival curves cross. The knockout does not reproduce the knockdown. The p-value improves every time you count more cells. The software ran without complaint, the arithmetic is correct, and the surprise is the only signal you get.
Each of those has a name, and knowing the name is most of the fix. This page indexes them by what you observed rather than by what the phenomenon is called, because you cannot look up a term you do not have yet.
Seven observations cover the analysis surprises this site indexes. Find the one that matches what is on your screen; the linked page carries the mechanism, the diagnostic and the remedy.
My effect disappeared when I added a covariate — could be confounding correctly controlled, or a mediator or collider that should never have entered the model. The three are indistinguishable by model fit and distinguishable by causal structure.
My survival curves cross — the hazard ratio is averaging effects that reversed direction, and the log-rank test is at its least powerful exactly there.
My knockdown and knockout phenotypes disagree — genetic compensation, off-target effects, or acute versus chronic loss. Each licenses a different claim.
My p-value gets smaller every time I measure more cells — the diagnostic signature of pseudoreplication, and one of the few statistical errors detectable without statistics.
My replicates disagree — technical spread is a measurement problem; biological spread is usually the finding.
My association reverses when I split the data — an aggregation reversal, produced deterministically by unequal weights across strata. Both the crude and the stratified number are arithmetically correct, and which one answers your question is decided by causal structure rather than by the data.
My extreme group improved before the treatment could act — selection on a noisy measurement guarantees apparent improvement at remeasurement, with a size fixed by the test-retest correlation and computable before any data are collected.
Triage a surprising result in four steps, in this order: reproduce the number, name the single thing that changed, classify that thing as procedural or biological, and only then interpret the biology. Jumping to step four is the failure mode the rest of this page exists to prevent.
1. Reproduce the number from raw data. Re-run the analysis from the source file rather than from the processed table you have been looking at. A surprising number that will not reproduce is a data-handling bug, not a phenomenon. The three recurring sources are a join that silently duplicated rows, a filter applied twice so that an exclusion was compounded, and a grouping variable whose factor levels reordered between sessions so that the reference category changed underneath the model.
2. Name the single thing that changed. Write it down as one sentence: “the coefficient halved when I added maternal education”, or “the p-value fell from 0.09 to 0.004 when I went from 60 cells to 300”. A surprise you can state as a one-line contrast between two analyses is diagnosable. A general feeling that the numbers look odd is not.
3. Classify the change as procedural or biological. More cells, a different covariate, a different clone, a different day, a different exclusion rule and a different software version are all procedural. A different dose, a different genotype and a different tissue are biological. This is the decision that routes you to the right page above.
4. Check the unit of analysis and the denominators before interpreting anything. Ask what one row of your analysis dataset is, and whether that row is the thing that received the treatment. Then reconcile the numbers: the N in the abstract, the column totals in Table 1, the analysed N in the primary model and the flow diagram frequently disagree. STROBE item 13(a) asks for numbers at each stage — potentially eligible, examined for eligibility, confirmed eligible, included, completing follow-up and analysed — and item 13(b) asks for the reasons for non-participation at each stage.
The literature on each of these is written for people who already know what to search. If you have not met the term “collider” you will not find the collider page, however good it is — and the moment you most need it is the moment you are staring at a coefficient that halved for no reason you can see.
The rule of thumb worth carrying: a result that changes when you change something procedural, rather than something biological, is telling you about the procedure. More cells, a different covariate, a different clone, a different day — if any of those moves the answer substantially, the movement is the finding to chase first.
The rule has a corollary that saves months. Procedural sensitivity is cheap to test and biological explanation is expensive, so test the procedure first even when the biological story is the more attractive one. Adding a random effect for animal, recomputing on per-animal means, refitting without the suspect covariate, or plotting the outcome against run date each takes minutes. Commissioning a mechanistic experiment to explain an artefact takes a year and returns nothing.
Charig and colleagues reported in the British Medical Journal in 1986 on three treatments for kidney stones (BMJ 1986;292:879–882), and the open-surgery versus percutaneous comparison in their table is the standard clinical instance of a reversal. Open surgery succeeded in 273 of 350 cases (78%); percutaneous nephrolithotomy succeeded in 289 of 350 (83%). At that level the less invasive procedure wins.
Stratify by stone diameter and the ordering flips in both strata. For small stones, open surgery succeeded in 81 of 87 cases (93.1%) against 234 of 270 (86.7%). For large stones, open surgery succeeded in 192 of 263 (73.0%) against 55 of 80 (68.8%). Open surgery is better for small stones, better for large stones, and worse overall.
The mechanism sits in the denominators. Large stones — the harder cases — went disproportionately to open surgery (263 of its 350 cases, 75%), and small stones disproportionately to percutaneous nephrolithotomy (270 of its 350, 77%). Nothing here is a coding error, and no diagnostic run on the pooled model would flag it. What resolves the case is an argument about the world: stone size causes both the treatment assigned and the chance of success, and does not sit on the pathway from treatment to outcome, which makes stratification correct here. Had the stratifying variable been measured after treatment, stratifying would have been the error instead.
Selection on a noisy measurement produces improvement on its own, and the size is computable in advance. Suppose systolic blood pressure in the source population has mean 130 mmHg and standard deviation 20 mmHg, and a study enrols everyone whose single screening reading is at least 160 mmHg. The enrolled group averages about 169 mmHg, because selecting above a cut-off selects people whose true value is high together with people whose reading happened to be high that morning.
If the correlation between single casual readings at separate visits is 0.70, the expected mean at the second visit is about 157 mmHg — a fall of roughly 12 mmHg produced by the enrolment rule and nothing else. A fall of that size is large enough to be read as a treatment effect, so a single-arm study with that eligibility criterion reports an impressive result before its intervention has done anything.
The general form is worth carrying: the expected regression is (1 − ρ) times the amount by which the selected group sits above the population mean, where ρ is the test-retest correlation of the measurement. Reducing measurement error shrinks the artefact faster than proportionally, because (1 − ρ) is the error’s share of the total variance. A randomised concurrent control removes it from the between-arm comparison, because both arms regress equally, and removes it from nothing else — within-arm change from baseline still carries it in full.
Standard error falls with the square root of the number of observations the test believes it has. Counting 60 cells and then 300 from the same three animals multiplies n by five in the formula, which divides the standard error by about 2.2 — a 55% reduction — while the number of independently treated units stays at three. Doubling the cells alone shrinks the standard error by about 29%.
That is the whole mechanism, and it explains why the p-value improves in proportion to microscope time. The extra cells carry information about variation within an animal, which is not the variation the comparison between groups is made against. Three animals per group is the evidence whether you image 60 cells or 6,000.
The self-test takes ten seconds: ask what happens to your n if you go back to the microscope tomorrow. If counting more cells raises n, your n counts cells. If n can only rise by treating another animal or running another independent culture, it counts independent units and you are fine.
Not everything on this list is an error. Crossing survival curves can be genuine early harm and late benefit. Disagreeing biological replicates can be real heterogeneity of response. A knockout differing from a knockdown can be compensation, which is itself publishable.
The task is to establish which, and to say so in the paper rather than smoothing it over. A disclosed and explained oddity is easier to defend in review than a result that looks impossibly clean.
Three features distinguish a real surprise from an artefact in practice. First, a real effect survives a procedural stress test: it is still there under a different exclusion rule, a different covariate set, and analysis at the correct unit. Second, it has a direction that a mechanism predicted in advance, rather than one supplied after the fact. Third, it reproduces in units the first analysis did not use — another cohort, another clone, another cage, another run. An anomaly that survives all three is a finding to write up. An anomaly that disappears under any one of them was procedural, and the procedure is what you learned about.
Reviewers rarely name the phenomenon. They write the specific version, and these are the comments that decide papers.
“The unadjusted and adjusted estimates differ substantially; please state which covariates were responsible and justify the adjustment set by causal role rather than by availability.” “The figure shows the survival curves crossing, yet a single hazard ratio is reported and interpreted; please address proportional hazards.” “The stated n appears to be the number of cells rather than the number of animals; please clarify the independent experimental unit for each panel.” “One representative experiment is shown; please report all independent replicates and the variability between them.” “The association reverses on stratification and the manuscript reports only the stratified estimate; please justify the choice.” “Eligibility required a screening value above a threshold, and within-arm change from baseline is interpreted as a treatment effect.”
Two more are worth anticipating, because both are triggered by the attempt to pre-empt them. “The limitations section acknowledges the discrepancy but does not state its likely direction or magnitude” — a generic acknowledgement reads as an admission without an analysis. And “the reported variability between biological replicates is implausibly small for this assay” — extremely clean data draws scrutiny rather than deflecting it.
The obvious remedy fails often enough that the fallback deserves stating explicitly, and in every case the fallback is a smaller claim rather than a different statistic.
More animals are not available to you. You cannot manufacture independent units after the fact. Report per-animal means with the effect size and its confidence interval, label the analysis exploratory, state the number of independent units plainly, and describe the design that would be adequately powered. What you cannot do is keep the 300-cell p-value.
Batch is fully confounded with condition. If every treated sample ran on Tuesday and every control on Thursday, condition and run day are the same variable, and no correction method separates them. Batch-correction tools will still run and will still return a clean-looking result, which is the trap: the output looks corrected and remains confounded. The honest options are a bridging experiment in which a shared reference sample is run in every batch, a replication in a design where batch and condition are crossed, or a stated limitation that the comparison cannot be interpreted as it stands.
The pre-registered analysis is the one the data violated. Report the pre-specified analysis first, in full, and report the assumption failure with its evidence — scaled Schoenfeld residuals against time for proportional hazards, a variance-components summary for clustering. Then report the corrected analysis, clearly labelled post hoc. Substituting the corrected analysis silently converts a defensible paper into an undisclosed deviation.
Samples are exhausted and the experiment cannot be repeated. Say so, name the experiment that would have resolved the ambiguity, and confine the claim to what the surviving data support. A bounded claim with an explicit reason is publishable; an unbounded claim resting on data nobody can check is what gets found later.
Report an unresolved surprise by naming the observation, the candidate explanations, the evidence that discriminates between them, and the claim that survives — in that order, in the results and again in the limitations.
State the observation as a number with its unit, not as an impression: “the coefficient fell from 1.84 to 1.02 on adjustment for stage”, not “the effect attenuated”. Name the candidate explanations, and say which the data favour and why. Where the data cannot discriminate, say so in a sentence and give the direction each explanation would push the estimate, so a reader can see how much the conclusion depends on the choice. Then keep the abstract consistent with the weaker claim, because the abstract is where the mismatch gets noticed. Reporting both analyses and explaining the difference is stronger than reporting the convenient one, and it removes the single most damaging discovery a reviewer can make.
PerfectPaper’s standing review covers statistics, methodology and evidence, and custom reviewer agents add the checks that catch these specific patterns — pseudoreplication, batch effects, competing risks and causal language among them.
PerfectPaper reads the figure legends, the methods and the reported denominators together, and reports where a stated n exceeds the number of independently treated units, where an adjustment set contains a variable the exposure causes, and where a figure contradicts the statistic printed beside it. Those are the three symptom classes on this page that change a conclusion rather than a decimal place.
If a reviewer has already raised one of these, the response side is covered in statistical reviewer comments and methods and design comments.
Reproduce the number from the raw data before anything else, because a surprise that will not reproduce is a data-handling bug rather than a phenomenon. Then name the single thing that changed between the two analyses, and ask what one row of your analysis dataset represents.
Ask whether the thing that changed was biological or procedural. If adding cells, changing a covariate or switching clones moved the answer, investigate the procedure before interpreting the biology.
Report the one whose assumptions match your design, and say what changed and why. Where you cannot tell which is correct, report both — a discrepancy disclosed is a limitation, a discrepancy hidden is a problem.
Because at least one analysis choice is doing work the data cannot support — usually the unit of analysis, the adjustment set, or an exclusion rule applied after seeing the outcome. Sensitivity to a procedural choice is itself a result to report, alongside the range the estimate takes across defensible choices.
Yes. Disclosed anomalies are limitations; discovered ones are credibility problems, and reviewers who find something unreported treat the rest of the paper differently.
Extremely clean data with no variation between biological replicates invites questions, because biological systems vary. Reporting the spread you actually observed is more persuasive than a figure implying there was none.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect