Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Whether disagreement is a problem depends on what kind of replicate disagreed. Technical spread is a measurement issue; biological spread is usually the finding.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Before troubleshooting, establish which kind of replicate disagreed, because the two have opposite implications. Technical replicates that disagree indicate a measurement problem you should fix. Biological replicates that disagree are usually reporting real variation — and that variation is the thing your statistics exist to characterise, not an obstacle to it.
Conflating them leads to the most common bad response: repeating the experiment until three of them agree, and reporting those. The second most common is the reverse error — treating genuine measurement noise as biology and writing a discussion paragraph about heterogeneity that a better pipetting technique would have deleted.
The unit of replication decides whether disagreement is noise, a finding, or a design error, so name the unit before you name the problem. Lazic, Clarke-Williams and Munafò set out the three units cleanly in “What exactly is ‘N’ in cell culture and animal experiments?” (PLOS Biology 2018;16:e2005282): the biological unit is the entity your conclusion is about, the experimental unit is the entity independently allocated to a condition, and the observational unit is the thing you actually measured. Sample size is the number of experimental units, not the number of measurements.
Their three criteria for genuine replication are worth applying literally to the disagreement in front of you: independent allocation to conditions, independent application of the treatment, and no mutual influence between units. Three wells split from one flask, treated from one master mix, on one plate, satisfy none of them. Three mice from one cage, dosed via one shared water bottle, fail the second and third.
The failure is common enough to have been counted. In a random sample of 200 published animal experiments from 2011 to 2016 in which an intervention was applied to parents and the effect measured in offspring — a design where the correct unit is not a matter of opinion — Lazic and colleagues found that only 22% (95% CI 17–29%) replicated the correct entity–intervention pair, 46% (95% CI 38–53%) showed pseudoreplication, and 32% (95% CI 26–39%) gave too little information to judge. If your three disagreeing values are three observational units from one experimental unit, the disagreement is measurement scatter and the honest n is one. That is the substance of pseudoreplication and of the independence objection.
Technical disagreement is variation between repeated measurements of the same biological material. Same sample, same day, measured more than once, with different answers. This is instrument, handling or preparation noise, and it is worth chasing.
Usual sources: pipetting at the bottom of a volume range, incomplete mixing, plate position and edge effects, freeze-thaw cycles, degraded reagent, a standard curve extrapolated beyond its range, or a detector near its limit of quantification.
The useful diagnostic is whether the spread scales with the mean. Constant absolute spread suggests additive noise like background; constant relative spread suggests multiplicative noise like pipetting error, and often calls for analysis on a log scale. Run that diagnostic explicitly rather than by eye: plot the standard deviation of each replicate set against its mean across the full dynamic range. A flat cloud is additive; a line through the origin is multiplicative, and the coefficient of variation is then the right summary. A U-shape, with the coefficient of variation rising at both the top and the bottom of the range, is the signature of an assay being read outside its linear region at one end and near the limit of quantification at the other.
Quantify the noise before you argue about it. The MIQE guidelines (Bustin et al., Clinical Chemistry 2009;55:611–622) mark “PCR efficiency calculated from slope”, “Number and stage (reverse transcription or qPCR) of technical replicates” and “Repeatability (intraassay variation)” as essential reporting items, with “Number and concordance of biological replicates” and “Reproducibility (interassay variation, CV)” listed as desirable. Most fields have an equivalent convention; if yours does not, report the coefficient of variation anyway, because a reviewer cannot evaluate “the replicates were consistent” and can evaluate a number.
A qPCR technical triplicate returning Cq values of 24.1, 24.3 and 26.8 is a pipetting or mixing failure, not a biological observation, and the arithmetic shows why. At an amplification efficiency of 100%, template doubles each cycle, so a 2.5-cycle gap between two aliquots of the same tube implies 2^2.5 ≈ 5.7-fold different starting template in material that was identical before it was dispensed. No biology occurred between the pipette tip and the well.
The correct response is mechanical: check the volume against the pipette’s working range, re-vortex the master mix and the template, look at whether the outlier well sits on a plate edge or in a column loaded last, and re-run. If the spread persists after those checks, the assay is being read near its limit of quantification and the answer is more input, not more replicates. Reporting the mean of 24.1, 24.3 and 26.8 as one data point carries a fold-error into every downstream comparison, and averaging is exactly what hides it.
The same logic applies wherever a measurement is repeated on one preparation: duplicate ELISA wells, repeated flow cytometry acquisitions from one stained tube, repeated fields imaged from one coverslip, and repeated peptide injections from one digest. In every case the repeat estimates the instrument, and the instrument is not the subject of your paper.
Biological disagreement is variation between independent biological units measured once each. Different animals, different donors, independently derived cultures, different days — with different answers. Usually this is not an error.
Biological systems vary. Animals differ in weight, microbiome, cage and stress. Primary cells differ by donor. Cell lines drift across passages. A treatment effect that is large in one animal and small in another is a real observation about heterogeneity of response, and it is often more interesting than the mean.
The correct handling is to characterise it, not eliminate it: report the individual values, not just the mean with error bars, and let the reader see the spread. If the effect direction is consistent while magnitude varies, say so — that is a stronger and more honest description than a bar chart implying uniformity.
Direction-consistent, magnitude-variable is a specific and defensible claim, and it is worth stating in those words. Four donors all moving the same way by 1.4-, 2.1-, 2.3- and 6.0-fold support “the response is consistent in direction and varies roughly four-fold in magnitude across donors”. They do not support “treatment produces a 2.95-fold increase”, which is the mean of numbers that were never drawn from one thing. Where the spread is large enough that the mean describes nobody, report the range and say what the effect size means at each end of it, because a reviewer weighing clinical or biological relevance needs the smallest effect you observed, not the average.
Resist one specific temptation: splitting biological replicates into “responders” and “non-responders” after seeing the data. A post hoc split on the outcome guarantees that the groups differ on the outcome, and the apparent gap between them shrinks on repeat measurement for the ordinary reason that regression to the mean predicts. A responder subgroup is a hypothesis, defined by a pre-specified threshold and tested in fresh units, not a result.
A SuperPlot is a figure that shows cell-level variability and experiment-level reproducibility in the same panel, which removes most of the ambiguity that makes reviewers ask what the points are. Lord, Velle, Mullins and Fritz-Laylin describe the construction in “SuperPlots: Communicating reproducibility and variability in cell biology” (Journal of Cell Biology 2020;219(6):e202001064): plot every observational unit as a small, faint point, colour or shape them by their biological replicate, overlay the per-replicate summary as a large marker, and run the statistics on the large markers only.
The plot is honest about the thing that trips up the analysis: the small points are usually hundreds of cells and the large points are usually three experiments, and only the second number is the sample size. Averaging within each biological replicate and testing across replicate means is the simplest correct analysis and needs no extra machinery. A random-intercept mixed model, with biological replicate as the grouping factor, uses the same structure and additionally reports how much of the total variance sits between units rather than within them — the intraclass correlation. A high intraclass correlation is the formal statement that your cells are not independent, which is the mechanism behind a p-value that shrinks as you count more cells.
A batch effect is a systematic shift in measurements that tracks a processing group rather than a biological one, and it is the third explanation for replicates that disagree. Where disagreement tracks the day, the operator, the reagent lot or the sequencing run, you may have a batch effect rather than biological variation. The diagnostic is whether the grouping of results matches an experimental grouping rather than a biological one.
This matters most when batch aligns with condition. If every treated sample ran on Tuesday and every control on Thursday, the difference between conditions and the difference between days are the same variable and cannot be separated afterwards by any method. The batch effects agent checks this specifically, including whether batch was modelled and whether the design permits separation at all.
Detect it before you model it. Record processing date, operator, lot number, instrument and plate position as columns in the data, then colour a principal-components plot or a sample-correlation heatmap by each of them in turn; a batch effect looks like samples clustering by a metadata column that has nothing to do with biology. Leek and colleagues make the general case in “Tackling the widespread and critical impact of batch effects in high-throughput data” (Nature Reviews Genetics 2010;11(10)), noting that batch effects arise from laboratory conditions, reagent lots and personnel differences and become a major problem specifically when they correlate with the outcome of interest. Where batch is recorded and not confounded with condition, include it as a blocking factor or a random effect, or remove it with an established method such as ComBat or surrogate variable analysis, and report which you used and what changed. Where the design is high-dimensional, batch structure also inflates the number of apparently significant features, which is a multiple-testing problem as much as a batch one.
Do not discard the outlier replicate because the other two agree. Exclusion needs a pre-specified rule and a documented technical reason — a failed control, a visibly compromised sample — not disagreement with its neighbours. Post hoc exclusion of inconvenient replicates is a serious problem and it is detectable. ARRIVE 2.0 item 3a asks authors to “Describe any criteria used for including or excluding animals (or experimental units) during the experiment, and data points during the analysis”, and item 3b to “For each experimental group, report any animals, experimental units, or data points not included in the analysis and explain why” — see the ARRIVE guidelines and in vivo rigour. Statistical outlier tests do not rescue a post hoc decision either, and at n = 3 they will not even engage. Grubbs’ test assumes a single outlier drawn from an otherwise Gaussian distribution and has almost no power at three points, and the ROUT method of Motulsky and Brown (BMC Bioinformatics 2006;7:123) is explicit about the same limit: with one or two residual degrees of freedom their method “never found an outlier no matter how far it was from the other points”. Even on larger sets, ROUT at the authors’ recommended Q of 1% falsely flags one or more outliers in about 1–3% of experiments where all scatter is genuinely Gaussian. A test can support a rule you set in advance; it cannot manufacture one afterwards, and it cannot adjudicate a triplicate.
Do not keep repeating until three agree. Selecting the experiments that agree is selecting on the outcome, and the resulting spread is not an estimate of anything. The mechanism is arithmetic: if you run six experiments and report the three with the smallest spread, the reported standard deviation is the minimum of a set of sample standard deviations, which is biased low by construction, and every confidence interval and p-value computed from it is wrong in a known direction. Discarding experiments that “did not work” without a pre-specified failure criterion is one of the recognised forms of p-hacking.
Do not report only the representative one. If three experiments disagreed, a single “representative” panel misrepresents the data. This is exactly the practice that produces the replication objection in review. Where the panel is an image, the accompanying quantification must come from all replicates and not from the displayed field, which is the routine finding in image quantification review.
Do not write “n = 3” and stop. Three what, allocated how, measured how many times each. The same string covers three mice, three wells and three reads of one well, and those are three different papers. State the experimental unit in the figure legend every time.
When repeating the experiment is impossible, report the disagreement and downgrade the claim to what the observed spread supports. Irreplaceable material makes this a real situation rather than an excuse: a patient cohort already consumed, a discontinued reagent lot, a primary line that stopped dividing, an animal cohort already sacrificed under a protocol you cannot reopen quickly, a grant period that ended.
The available moves, in order of strength. First, use the variance you actually observed to state precision honestly — a confidence interval computed from three biological replicates is wide, and a wide interval that contains a large effect is a legitimate result reported as such. Second, validate on a different axis rather than repeating the same measurement: an independent assay, an independent readout, or an independent model bearing on the same claim is worth more than a fourth replicate of the same protocol, which is the logic behind orthogonal validation. Third, state the limitation as a specification — how many independent units, what the observed spread was, and what the spread does and does not permit you to conclude — rather than as an apology; how to write a limitations section covers the form. Fourth, if the claim depends on a difference the data cannot resolve, weaken the claim rather than the standard of evidence.
What does not work is post hoc power analysis. Computing power from the effect you observed is a monotone transformation of the p-value and adds no information about whether your three disagreeing replicates were enough; the post hoc power objection covers why reviewers reject it. If the honest answer is that the design was underpowered, say so in those words and state the number of independent units the design would have needed, which is a design statement rather than a defence.
Reviewers rarely write “your replicates disagree”. They write the specific version, and the specific versions are predictable enough to pre-empt. “Please state whether n refers to biological or technical replicates.” “The error bars are not defined; specify whether they are SD, SEM or a confidence interval, and state n.” “Figure 3B shows three points but the methods describe three wells from a single culture.” “One replicate was excluded; please state the pre-specified exclusion criterion.” “The representative blot does not correspond to the quantification; show all replicates.” “Samples appear to cluster by processing date rather than by treatment group.” “The variability across donors is substantial and is not discussed.”
The last one decides papers more often than it looks, because it is usually triggered by a figure that concealed the spread rather than by the spread itself. A reviewer who can see four values with a four-fold range and a consistent direction will read a bounded, real result. A reviewer who sees a bar with an error bar, then finds the range in a supplementary table, reads a paper that tried to hide something. Anticipating the objection is the same work as anticipating a request for more controls: decide what the data support, then present it in the form that shows it.
What to report when replicates disagree is a short fixed list, and it is the same list whether the disagreement turned out to be technical or biological. The number of independent experiments, what makes them independent, the individual values rather than only summary statistics, and the variability honestly described. Where magnitude varied but direction held, say that explicitly.
In practice that is six items in the figure legend and methods: the experimental unit; the number of independent experimental units, per group, as an exact number rather than a range; the number of technical replicates per unit and how they were combined; what the error bars represent; the exclusion criterion and every excluded unit with its reason; and the test, with the level at which it was applied. ARRIVE 2.0 item 3c states the third of these as a bare requirement — “For each analysis, report the exact value of n in each experimental group” — and “n = 3–5 per group” does not satisfy it.
Disclosed variability reads as a bounded result. Three experiments implied to be identical read as a claim that has not been checked, and the second is the one that draws a second round of review.
Related symptoms: my p-value gets smaller with more cells, my knockdown and knockout disagree.
PerfectPaper reads the figure legends, the methods and the reported n together, and flags where the unit tested is not the unit the conclusion is about. Preclinical manuscripts are additionally checked by pseudoreplication review before submission.
Replicates disagree for one of three reasons: measurement noise between repeats of the same sample, genuine biological variation between independent units, or a technical grouping such as day, operator or reagent lot that tracks the results. Identify which before changing anything, because the first is a fault to fix and the second is a result to report.
A technical replicate measures the same biological sample more than once and estimates measurement noise. A biological replicate is an independent biological unit and estimates biological variation. Only the second supports generalisation.
Only with a pre-specified rule and a documented technical reason, disclosed in the paper. Excluding a replicate because it disagrees is selecting on the outcome.
No threshold exists, and any journal-wide number would be arbitrary across assays and organisms. What matters is that the variation is reported and that the claim is consistent with it. Large variation with a consistent direction supports a weaker but real claim.
Yes, wherever the number is small enough to plot. A bar with an error bar over three or four points hides the very thing under discussion, and a scatter or dot plot with the summary overlaid answers the reviewer’s next question before it is asked.
An operator batch effect is the explanation, and it is a design fact rather than a personal one. Document it, and if operator is confounded with condition — one person ran the treated samples and another the controls — the comparison cannot be interpreted until the design is corrected.
Add more biological replicates if the disagreement is biological and the interval is too wide to support the claim; adding technical replicates in that situation narrows nothing that matters, because it estimates the instrument rather than the population. If the disagreement is technical, fix the measurement first, since more replicates of a poorly controlled assay average noise rather than remove it.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect