Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Pseudoreplication counts measurements from one experimental unit as independent. It inflates significance by orders of magnitude and is detectable without statistics.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
Pseudoreplication is treating multiple measurements from the same experimental unit as independent observations. Fifty cells measured from one mouse produce fifty numbers but one independent observation, because those cells share that animal’s genetics, treatment and environment. Analysing them as fifty inflates the apparent sample size, shrinks the standard error, and produces p-values that can be wrong by orders of magnitude.
Stuart Hurlbert introduced the term in ecology in “Pseudoreplication and the design of ecological field experiments” (Ecological Monographs 1984;54:187–211), a survey of the field-experiment literature in which 48% of the studies that used inferential statistics were pseudoreplicated. The problem it names is not confined to ecology: it is at its most acute wherever the gap between the number of measurements and the number of independent units is largest, which in practice means cell and molecular biology. Lazic read a single issue of Nature Neuroscience and classified 12% of its papers as pseudoreplicated, with a further 36% suspected but undecidable because the papers did not report enough to tell (BMC Neuroscience 2010;11:5).
Independence is a property of the experimental design, not of the measurement. The independent unit is whatever level satisfies two conditions: the treatment was applied to it independently, and biological variation occurs between units of that kind.
If you inject six mice — three treated, three control — the treatment was applied six times, independently, and the animals vary biologically from one another. The animal is the unit. Whether you measure one cell per animal or ten thousand, the experiment contains six independent observations.
The cells are not worthless. They give you a more precise estimate of each animal’s value, which reduces measurement noise. What they do not do is tell you anything about how animals differ from each other, and the comparison between treated and control animals is a comparison across that variation.
The unit is not always the animal, and the first condition decides which level it is. A drug delivered in the drinking water of a cage of five mice was applied once, to the cage: the cage is the unit and the five animals are subsamples. A compound given to a pregnant dam and assessed in her pups makes the litter the unit, not the pup. ARRIVE 2.0 makes this a reporting requirement rather than a matter of taste — item 1b of the Essential 10 asks authors to state “The experimental unit (e.g. a single animal, litter, or cage of animals)”. Naming the unit once in the methods prevents most of what follows; see the ARRIVE guidelines and in vivo rigour.
Standard error falls with the square root of the sample size. This is why the error is not a technicality.
Report 150 measurements from 6 animals as n = 150 and the standard error is computed as though you had 150 independent animals. Relative to the correct calculation on 6, it is smaller by a factor of five. Because the test statistic is the effect divided by the standard error, the statistic is five times larger than it should be.
In practice this moves results across several orders of magnitude of p-value, so a finding reported at p < 0.001 on the inflated count can be comfortably non-significant when analysed on the units. This is why reviewers treat the objection as serious rather than pedantic: unlike most statistical criticisms, which shift a result at the margin, this one can reverse it.
The quantity that governs the size of the error is the intraclass correlation coefficient, ICC or rho, the share of total variance sitting between units rather than within them. Kish’s design effect converts it into a penalty directly: DEFF = 1 + (m − 1)rho, where m is the number of measurements per unit. Take 25 cells per animal and an ICC of 0.5, which is unremarkable for a phenotype that differs between animals: DEFF is 13, and 150 measurements carry the information of about 11 independent observations. ICC = 0 is the only value at which measurements are worth their face value, and it means the animals do not differ from one another at all.
Aarts and colleagues measured the consequence for the false positive rate. Reviewing 314 neuroscience papers they found 53% reported nested data, and simulation put the probability of declaring a null effect significant at as much as 80% against a nominal 5% (Nature Neuroscience 2014;17:491–496). A test running at 80% type I error is not a weakened test; it is a coin weighted to say yes. The same inflation corrupts any statistical power calculation built on the measurement count, which is why justifying a sample size after the fact from the number of cells measured is worthless — see post-hoc power.
Journals increasingly define n in their instructions rather than leaving it to the author. The British Journal of Pharmacology states the rule in its guidance for authors and peer reviewers (Curtis et al., British Journal of Pharmacology 2018;175:987–993): “group size is the number of independent values, so one sample run five times is n = 1, not n = 5.” The same document settles the canonical case in one sentence — “Five cells from one mouse given the same treatment is not n = 5, and these five cells should be regarded as technical replicates, with n = 1 consensus value taken forward into statistical analysis” — and sets 5 as the minimum n for a dataset subjected to statistical analysis at that journal, with smaller groups to be labelled exploratory or preliminary.
Two things follow for a manuscript. At a journal with such a policy the objection is a compliance check rather than a referee’s opinion, and an editor can act on it without a statistical reviewer. And a design of three animals with fifty cells each does not meet a minimum of five independent values however many measurements it contains, which makes it a sample size problem revealed by fixing the pseudoreplication rather than created by it.
Pseudoreplication takes a small number of recurring forms, and naming which one you have is usually enough to identify the fix.
Cells from an animal. The canonical case. Treatment applied to animals, measurements taken on cells.
Wells from a culture. A flask split into six wells is one independent culture. The wells are technical replicates, however separately they are subsequently handled. Splitting a flask on Monday and treating the wells on Friday does not make them independent, because they still carry one passage history and one set of culture conditions.
Sections from a block, fields from a slide, images from a section. Each level is nested within the one above, and the treatment was applied at the top. Image analysis is where the counts grow fastest — twenty fields per section, three sections per animal, and n reads 60 for four animals; see image quantification reporting.
Two tumours in one mouse. Two eyes, two limbs, bilateral implants. These are paired within an animal, and the pairing must be modelled.
Repeated measures on a subject. Multiple time points, multiple readings, multiple assays on one sample. Adjacent time points on one subject correlate more strongly than distant ones, so this form usually needs a covariance structure rather than a single random intercept.
Pooled replicate experiments. Three independent experiments, each producing 100 cells, pooled into a single test on 300. Between-experiment variation is usually the largest source of variability in cell biology, and pooling erases it entirely. It is also the form most likely to be a disguised batch effect, because the day of the experiment carries reagent lot, passage number and operator with it.
Cages, litters and tanks. Housing is a treatment when the treatment arrives through housing. Diet, water-borne compound, temperature, tank and enrichment are applied to the enclosure, and the enclosure is the unit.
Passages of one cell line. Three passages of one immortalised line are three measurements of a single genotype, not three biological replicates. Genuine replication at that level requires independently derived lines, independent donors, or an independent clonal derivation.
Patients within clinics, students within schools. The same structure in human research, where it is usually called clustering and is better recognised.
Reads, cells and spots in sequencing. Thousands of single cells from one animal are thousands of measurements of one unit, which is why pseudobulk aggregation to the donor before differential testing is now the standard analysis; see multiple testing in omics. The same applies to events recorded per sample in flow cytometry.
Pseudoreplication can be diagnosed with one question, asked of your own manuscript or of one you are reviewing.
Could you increase your n tomorrow without running a new experiment?
If going back to the microscope and counting more cells raises your n, then your n counts measurements, and it is not a sample size. If n can only rise by treating another animal, deriving another independent culture, or repeating the experiment, then it counts independent units.
The test takes ten seconds and settles the question without any statistical knowledge. It is also the test to apply to a manuscript you are reviewing.
Pseudoreplication is found by reconciling three numbers that a paper reports in three different places, not by any statistical test.
Compare the n in the figure legend with the animal count in the methods. A legend reading n = 120 in a study whose methods describe eight mice is the whole finding. Where the legend gives no n at all, treat that as the same problem until the authors say otherwise.
Read the degrees of freedom. A t-test on six animals has 4 degrees of freedom; reported as t(148) it was run on 150 measurements. Degrees of freedom are the single most reliable tell because they are printed in the result and cannot be reconciled with a small animal count by any reading.
Look for a mismatch between the design sentence and the analysis sentence. “Mice were randomised to treatment or vehicle” followed by “cells were compared by unpaired t-test” is randomisation at one level and inference at another.
Check the scale of the claimed precision. Standard errors an order of magnitude smaller than the visible spread of the plotted points, or error bars narrower than the plotting symbol, usually mean the denominator counts measurements.
Ask what is nested inside what. Sections in blocks, blocks in animals, animals in cages, cages in cohorts. Every nesting level between the treatment and the measurement is a level at which variance is going unmodelled, and a paper that names none of them has not considered the question.
Pseudoreplication has a symptom that can be noticed during analysis rather than during design: if your p-value improves every time you measure more cells from the same samples, the analysis is counting measurements as independent.
Real evidence does not accumulate from measurement effort alone. If it appears to, the standard error is being computed at the wrong level. Measuring more cells per animal reduces the uncertainty in each animal’s own value, and that reduction stops mattering once within-animal noise is small relative to between-animal variation — usually after a few dozen cells. A p-value still falling at cell number 400 is arithmetic, not evidence, and the phenomenon has its own page: my p-value gets smaller with more cells.
Pseudoreplication has four standard remedies, and the first one is correct in the great majority of designs.
Summarise to the independent unit. Take the mean of the measurements within each animal, then run the comparison on those means. Six numbers, not 150. This is simple, transparent, immediately understood by any reviewer, and correct.
The objection people raise is that it discards information. It discards the within-animal precision, which was never contributing to the between-animal comparison anyway. Use the median rather than the mean where the within-unit distribution is skewed or contains a small number of extreme cells, and say which you used.
Fit a mixed-effects model with a random effect for animal, culture or experiment. This uses every measurement while estimating variance components at each level, and it handles unbalanced designs — different numbers of cells per animal — better than summarising does. Intervals widen relative to the naive analysis, and that widening is the correction working.
State the software, the random effect structure, and how degrees of freedom were determined. In R that means naming the package, lme4 with lmerTest or nlme, giving the formula in full, and reporting whether denominator degrees of freedom came from the Satterthwaite or the Kenward–Roger approximation; Kenward–Roger is the more conservative and is the usual choice when the number of units is small. In SAS the equivalent is PROC MIXED with DDFM=KR, in Stata mixed. A model reported as “a mixed model was used” is not reproducible and reviewers say so.
Use generalized estimating equations or cluster-robust standard errors where the question is about the population-average effect and the variance structure itself is not of interest. GEE with an exchangeable working correlation and the animal as the cluster gives a valid marginal estimate, but its sandwich standard errors are unreliable when the number of clusters is small — under roughly 30 to 40 clusters, prefer a mixed model or summarising.
Use the pairing where units are paired within a subject. Two tumours per animal call for a paired analysis or an animal random effect, not an unpaired test on twelve tumours.
Analyse experiments as blocks. Where an experiment was repeated, include experiment as a factor or random effect rather than pooling observations across repeats. Choosing a fixed rather than a random effect is defensible when there are only two or three repeats, since a variance component estimated from three levels is barely estimated at all.
Summarising to the unit fails in specific, recognisable situations, and each has an accepted alternative.
Too few units to fit a random effect. Three independent experiments will not support a well-identified variance component, and lme4 frequently returns a singular fit with the animal variance estimated at exactly zero. Report the singular fit rather than deleting the random effect, and fall back to the summarised analysis on three values, stating the model that failed and why.
The outcome is a proportion or a count. Averaging binary cell-level calls per animal and running a t-test on the proportions is usually acceptable and always transparent; a generalized linear mixed model with a binomial or negative binomial family is the alternative when the per-animal denominators differ widely.
The question is genuinely at the measurement level. A study of how a property is distributed across cells within an individual is a within-unit question, and the cells are the right unit for it. What that design cannot support is a between-group claim, so keep the two analyses and their two claims separate in the text.
Only one independent unit exists. A single patient-derived line, a single donor, a single field site. No analysis rescues an n of 1, and the honest presentation is descriptive, with the limitation stated as a specification of what the experiment can establish rather than an apology — see how to write a limitations section.
The data are already pooled and the grouping was not recorded. Where the experiment of origin was never tracked, the nesting cannot be reconstructed after the fact. Say so, analyse at the level you can defend, and record the grouping variable next time; it is one column.
A figure drawn from nested data must show both levels. Plot the individual measurements lightly — they convey the distribution and the sample handling — and overlay the per-unit means as larger points, with the statistics computed on those.
This design has a name and a published recipe. Lord, Velle, Mullins and Fritz-Laylin called it a SuperPlot (Journal of Cell Biology 2020;219:e202001064) and give tutorials for building one, colour-coding each independent experiment so that the reader can see which cells came from which repeat. A SuperPlot makes both the cell-level variability and the experimental reproducibility visible in one panel.
Many journals now require individual data points rather than bar charts with error bars, and this presentation satisfies that while making the nesting structure visible. A reader can see at a glance how many animals contributed and how variable they were, which is exactly what a bar chart conceals. State in the legend which level the error bars and the p-value were computed at, because a SuperPlot drawn correctly and analysed at the cell level is still pseudoreplicated.
Correcting pseudoreplication sometimes removes significance. This is the uncomfortable case, there are three legitimate responses to it and one that is not, and it is worth being direct about all four.
Report it as a non-significant difference with the effect size and confidence interval, which bounds what the experiment can say. Reframe the result as preliminary and state what a properly powered design would require. Or collect more independent units — more animals, more independent cultures — which is the only kind of additional data that helps.
The arithmetic of that last option is worth doing before committing to it. Because the design effect scales with the number of units and not with the number of measurements, adding six more animals at 10 cells each buys more information than adding 500 cells to the animals you have. Aarts and colleagues reach the same conclusion from the power side: collecting more independent units beats collecting more observations per unit.
What you cannot do is keep the analysis at the measurement level. Reviewers who raise this objection check the revision specifically for a mixed model added to the supplement while the abstract retains the original p-value.
Pseudoreplication is not an objection to taking many measurements, and reviewers occasionally overreach with it.
Technical replicates are good practice. Running a sample in triplicate detects pipetting error and instrument drift, and the British Journal of Pharmacology guidance says so explicitly while still requiring one consensus value per unit to go forward. Measuring 200 cells per animal instead of 5 gives a better estimate of each animal’s value, and where within-animal noise is large that is worth doing.
Nor is the objection an objection to nested designs. Nesting is how biology works, and a properly analysed nested experiment is stronger than a flat one, not weaker, because it separates the variance you care about from the variance you do not.
Where the objection does not reach at all is the distinct question of whether an experiment was repeated. Independent replication and independent units are different requirements, and a paper can satisfy one and fail the other — see reviewer says the experiment was not replicated. Choosing the wrong test for correctly counted units is a third, separate problem.
Reviewers rarely use the word pseudoreplication. They write the specific version, and these are the forms it takes.
“What does n refer to?” “The authors treat individual cells as independent observations.” “These appear to be technical rather than biological replicates.” “The analysis does not account for clustering within animals.” “It is unclear how many independent experiments contributed to this figure.”
The first of those is the most common form and the easiest to prevent: state, in every figure legend, what the n counts.
Answering the comment takes three moves and nothing else: name the experimental unit, re-run the analysis at that level, and update every number that changed — abstract, results text, figures and supplement together. Do not argue that the effect is obvious from the cell-level data, and do not present both analyses and let the reader choose. If the corrected result is weaker, say what it is and what it bounds; a revision that reports a smaller effect honestly is accepted far more often than one that relitigates the objection. See how to write a response to reviewers.
Statistical power · Confounding · My p-value gets smaller with more cells · My replicates disagree · Reviewer says my n is not independent
Caught before submission by the pseudoreplication agent, which reads every figure legend, determines what each reported n counts, reconstructs the experiment’s structure from the methods, and reports each mismatch with its remedy.
An independent biological unit subject to the treatment — a separate animal, donor, or independently derived culture. Multiple measurements from one such unit are technical replicates and do not increase the independent sample size. ARRIVE 2.0 item 1b asks authors to state the experimental unit explicitly, giving “a single animal, litter, or cage of animals” as its examples.
Three. The cells give a more precise estimate of each animal’s value but add no independent observations, because they share everything about that animal. The British Journal of Pharmacology puts it as a rule: “Five cells from one mouse given the same treatment is not n = 5.”
A mixed-effects model lets you use all the measurements, but inference is made at the level of the independent units, so intervals reflect the true sample size. It is the correct analysis, not a way to preserve the original p-value. Report the random effect structure and whether degrees of freedom came from Satterthwaite or Kenward–Roger.
Structurally yes. Patients within clinics and cells within animals are the same problem, and both are addressed by analysis at the cluster level or by a random effect for cluster. Survey statistics calls the penalty the design effect, DEFF = 1 + (m − 1)rho, and it is the same quantity in both fields.
Because standard error falls with the square root of n. Overstating n by a factor of twenty-five understates the standard error fivefold, which multiplies the test statistic by five and moves p-values across orders of magnitude. In simulation, Aarts and colleagues found the false positive rate reaching 80% against a nominal 5% (Nature Neuroscience 2014;17:491–496).
No. Between-experiment variation is usually the largest source of variability, and pooling removes it from the analysis entirely. Include experiment as a factor, or summarise each experiment to one value.
State the number of independent units and what they are, and separately the number of measurements, for example: “n = 6 animals per group; 25 cells measured per animal; statistics computed on per-animal means.” Add the test, the model and the degrees of freedom, so that a reader can check the level at which inference was made without reading the methods.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect