Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Pseudoreplication is the one statistical objection that changes p-values by orders of magnitude, and the remedy is almost always reanalysis rather than new experiments.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A reviewer saying your n is not independent claims your reported sample size counts measurements rather than independent units, so your test assumes an independence your data do not have. The objection arrives in many phrasings — “the authors treat each cell as an independent observation”, “what does n refer to here?”, “these are technical replicates” — and carries one meaning.
The objection is the most consequential statistical one you can receive, because unlike most it changes results by orders of magnitude rather than at the margin. It is also, usually, fixable without a single new experiment.
The comment is also common enough that experienced reviewers look for it by default. Lazic examined a single issue of Nature Neuroscience as a case study and found that 12% of the papers had pseudoreplication, with a further 36% suspected of it but impossible to judge because the reporting gave insufficient information (BMC Neuroscience 2010;11:5). Aarts and colleagues, reviewing 314 papers drawn from five neuroscience journals, found that a little over half involved nested data of some kind (Nature Neuroscience 2014;17:491–496). A referee raising the point is applying a standing expectation, not singling your manuscript out.
Independence is a property of the experimental design, not of the measurements. The independent unit is whatever level the treatment was applied at and biological variation occurs at.
If you treated six mice and measured 25 cells from each, you have six independent units and 150 measurements. The 150 cells give you a more precise estimate of each mouse’s value — genuinely useful — but they do not tell you about mouse-to-mouse variation, which is the variation your comparison is against. A t-test on 150 cells uses a standard error computed as though you had 150 independent animals, and it will be far too small.
The same structure appears as: wells split from one culture, sections from one block, images from one slide, two tumours in one animal, repeated measures on one subject, and replicates pooled across experiments.
Two conditions decide the level, and both must hold. The treatment must have been applied to that level independently, and biological variation must occur between units of that level. A compound delivered in the drinking water of a cage of five mice was applied once, to the cage, so the cage is the unit. A compound given to a pregnant dam and assessed in her pups makes the litter the unit, not the pup. ARRIVE 2.0 turns this into a reporting requirement rather than a matter of taste: item 1b of the Essential 10 asks authors to state “The experimental unit (e.g. a single animal, litter, or cage of animals)”, and naming the unit once in the methods pre-empts most of what follows; see the ARRIVE guidelines.
The wording of the comment tells you which level the reviewer suspects, and therefore which number in your manuscript to check first. Reviewers rarely write the word pseudoreplication; they write the specific version.
| Reviewer’s wording | What is being claimed | What to check first |
|---|---|---|
| “What does n refer to here?” | The legend does not say what n counts | Every figure legend, against the animal or culture count in the methods |
| “The authors treat each cell as an independent observation” | Measurements were counted as units | Reported degrees of freedom against the number of animals |
| “These appear to be technical rather than biological replicates” | Wells or aliquots came from one source | Whether cultures were independently derived or split from one flask |
| “The analysis does not account for clustering” | Nesting exists and was not modelled | Whether any random effect or cluster term appears in the methods |
| “It is unclear how many independent experiments contributed” | Repeats were pooled | Whether experiment of origin was recorded and modelled |
| “Please clarify the experimental unit” | ARRIVE or journal compliance | Methods statement naming the unit, per ARRIVE 2.0 item 1b |
Each row points at a different fix, and answering the wrong one wastes a revision round. “Please clarify the experimental unit” can sometimes be answered with a sentence in the methods; “the authors treat each cell as an independent observation” never can.
The standard error shrinks with the square root of n. Going from a true n of 6 to a claimed n of 150 shrinks it fivefold, which moves p-values by orders of magnitude. A finding at p < 0.001 on the inflated n can be p > 0.2 at the correct one.
That asymmetry is why reviewers treat this comment as serious rather than pedantic, and why “we have added more cells” is the wrong response — it makes the reported n larger while leaving the real one unchanged.
The size of the error is governed by the intraclass correlation coefficient, ICC or rho, the share of total variance sitting between units rather than within them. Kish’s design effect converts it into a penalty: DEFF = 1 + (m − 1)rho, where m is the number of measurements per unit. At 25 cells per animal and an ICC of 0.5 — unremarkable for a phenotype that differs between animals — DEFF is 13, so 150 measurements carry the information of about 11 independent observations, not 150. ICC = 0 is the only value at which measurements are worth their face value, and it means the animals do not differ from one another at all.
Aarts and colleagues put a number on the consequence: in their simulations, ignoring the nesting drove the chance of declaring a null effect significant far above the nominal 5%, reaching the order of 80% in the most strongly clustered designs they modelled (Nature Neuroscience 2014;17:491–496). A test running at that false-positive rate is not a weakened test; it is a coin weighted to say yes. Eisner reached the same conclusion for cellular physiology and titled the paper accordingly — “Pseudoreplication in physiology: More means less” (Journal of General Physiology 2021;153:e202012826).
Settle the question before you draft a response, using three numbers your manuscript already reports in three different places. No statistical test identifies pseudoreplication.
Compare the legend n with the animal count in the methods. A legend reading n = 120 in a study whose methods describe eight mice is the whole finding. Where the legend gives no n at all, the reviewer is entitled to assume the worst.
Read the degrees of freedom. A two-sample t-test on six animals, three per group, has 4 degrees of freedom; reported as t(148) it was run on 150 measurements. Degrees of freedom are the most reliable tell because they are printed in the result and cannot be reconciled with a small animal count by any reading.
Apply the one-question test. Could you increase your n tomorrow without running a new experiment? If going back to the microscope and counting more cells raises your n, then your n counts measurements and it is not a sample size. This takes ten seconds and needs no statistical knowledge — it is the same test set out on what is pseudoreplication.
Check the p-value’s behaviour. If your p-value improved every time you measured more cells from the same samples, the standard error was being computed at the wrong level; that symptom has its own page, my p-value gets smaller with more cells.
Journals increasingly settle the definitional part for you. The design and analysis guidance the British Journal of Pharmacology directs its authors, reviewers and editors to defines group size as the number of independent values, not the number of measurements taken, so a single sample assayed five times counts once (Curtis et al., British Journal of Pharmacology 2018;175:987–993). At a journal with such a policy the comment is a compliance check rather than a referee’s opinion, and an editor can act on it without a statistical reviewer.
Pseudoreplication has four standard remedies and one prohibition, and the first remedy is correct in the great majority of designs. Choose by design structure rather than by which one preserves the p-value.
Summarise to the independent unit. Take the mean of your 25 cells per mouse, then run the comparison on six values. Simple, transparent, and immediately understood by any reviewer. You lose the visual richness of a 150-point plot, but you can still show the cells with the per-animal means overlaid — many journals now prefer this figure. Use the median rather than the mean where the within-unit distribution is skewed or holds a few extreme cells, and say which you used.
Fit a mixed-effects model with a random effect for animal, culture, or experiment. This uses all the data while estimating variance at the right level, and it is the better answer when your groups are unbalanced or you have several nesting levels. Say which package and which random effect structure you used. In R that means naming lme4 with lmerTest, or nlme, giving the formula in full, and stating whether denominator degrees of freedom came from the Satterthwaite or the Kenward–Roger approximation; Kenward–Roger is the more conservative and is the usual choice when units are few. In SAS the equivalent is PROC MIXED with DDFM=KR, in Stata mixed. “A mixed model was used” is not reproducible, and reviewers say so on the second round.
For paired structures, use the pairing: two tumours per animal call for a paired analysis or an animal random effect, not an unpaired test on twelve tumours.
Use cluster-robust standard errors or generalized estimating equations where the target is the population-average effect and the variance structure is not itself of interest. GEE with an exchangeable working correlation and the animal as the cluster gives a valid marginal estimate, but its sandwich standard errors are unreliable when clusters are few — under roughly 30 to 40 clusters, prefer a mixed model or summarising.
Do not pool across independent experiments and test on the combined observations. Between-experiment variation is usually the largest source of variability, and pooling hides it entirely. It is also the form most likely to be a disguised batch effect, because the day of an experiment carries reagent lot, passage number and operator with it.
Summarising to the unit fails in a small number of recognisable situations, and each has an accepted alternative that a reviewer will accept if you name it.
Too few units to fit a random effect. Three independent experiments will not support a well-identified variance component, and lme4 frequently returns a singular fit with the between-unit variance estimated at exactly zero. Report the singular fit rather than quietly deleting the random effect, and fall back to the summarised analysis on three values, stating the model that failed and why.
The outcome is a proportion or a count. Averaging binary cell-level calls per animal and testing the proportions is transparent and usually acceptable; a generalized linear mixed model with a binomial or negative binomial family is the alternative when the per-animal denominators differ widely.
The question is genuinely at the measurement level. A study of how a property is distributed across cells within one individual is a within-unit question, and the cells are the right unit for it. What that design cannot support is a between-group claim, so keep the two analyses and their two claims visibly separate in the text.
Only one independent unit exists. A single patient-derived line, a single donor, a single field site. No analysis rescues an n of 1, and the honest presentation is descriptive, with the limitation written as a specification of what the experiment establishes rather than an apology — see how to write a limitations section.
The grouping was never recorded. Where experiment of origin was not tracked, the nesting cannot be reconstructed after the fact. Say so, analyse at the level you can defend, and record the grouping variable next time; it is one column.
Reanalysis at the correct unit sometimes removes the result. This is the hard case and it is worth being straight about.
If the effect survives at the correct unit with a wider interval, report it and say so. If it does not survive, you have three honest options: report it as non-significant with the effect size and interval, reframe the finding as preliminary and say what would test it properly, or collect more independent units. What you cannot do is keep the original analysis.
Reviewers who raise this comment are watching for exactly that. A revision that adds a mixed model in the supplement while the abstract keeps the original p-value gets rejected, and deservedly.
If you choose the third option, do the arithmetic first. Because the design effect scales with the number of units and not with the number of measurements, six more animals at 10 cells each buys more information than 500 more cells from the animals you already have. Aarts and colleagues reach the same conclusion from the power side: optimising a nested design nearly always means collecting more truly independent observations rather than more observations per object. Any statistical power calculation built on the measurement count is void for the same reason, which is why a sample size justified after the fact from the number of cells measured persuades nobody — see post-hoc power.
Answer this objection in three moves and nothing else: name the experimental unit, re-run the analysis at that level, and update every number that changed. A worked shape, for a study of six animals per group with 25 cells each:
We agree. The experimental unit is the animal, and the original analysis treated the 150 cells per group as independent observations. We have re-analysed by taking the per-animal mean of the 25 cells and comparing the six values per group (Welch’s t-test, t(9.4) = 2.41, p = 0.038); the mixed-effects model with a random intercept for animal, fitted in
lme4with Kenward–Roger degrees of freedom, gives the same conclusion and is now reported in Supplementary Table 3. The difference is 18% (95% CI 1.2% to 34%) rather than the 19% previously reported, and the p-value moves from p < 0.001 to p = 0.038. Figure 2b now plots cells in grey with per-animal means overlaid, and all figure legends state that n is animals. The abstract, Results and Methods have been updated accordingly.
Four features make that paragraph work, and each is checkable by a reviewer in one pass. It concedes rather than argues. It names the unit as a fact about the design. It gives the number that changed and the number that did not. And it lists every location in the manuscript that was touched, so the reviewer does not have to hunt for an unrevised copy of the old p-value. See how to write a response to reviewers for the letter structure around it, and a worked response letter for the surrounding format.
Two things not to do. Do not argue that the effect is obvious from the cell-level data — the reviewer is not disputing that the cells look different. And do not present both analyses and invite the reader to choose; a manuscript reporting two incompatible p-values for one comparison has made the reviewer’s point for them.
Correcting the unit changes at least five places in a manuscript, and reviewers check all of them on the second round.
Every figure legend. State the number of independent units and what they are, and separately the number of measurements: “n = 6 animals per group; 25 cells measured per animal; statistics computed on per-animal means.” Add the test and the degrees of freedom so the level of inference is checkable without reading the methods.
Every figure. Plot the measurements lightly and overlay the per-unit means as larger points, colour-coded by independent experiment. Lord, Velle, Mullins and Fritz-Laylin published this design as the SuperPlot (Journal of Cell Biology 2020;219:e202001064). A SuperPlot drawn correctly but still analysed at the cell level is still pseudoreplicated, so state in the legend which level the error bars and the p-value came from. Where the measurements are image-derived, the legend also has to say how many fields per section and how many sections per animal contributed, because that is the next thing a reviewer asks for.
The abstract. Any effect size, interval or p-value quoted there must be the corrected one. This is the single most common place the old number survives.
The methods. One sentence naming the experimental unit, one naming the model and its random effect structure, one naming the software and version.
The sample size statement. Six animals is a defensible n; it is also small, and correcting the unit often converts a statistics objection into a sample size objection. Anticipate that in the same revision rather than discovering it in round three.
Reviewers occasionally push this comment past what it supports, and three specific pushbacks are legitimate if stated without heat.
Taking many measurements is not the error. Running a sample in triplicate detects pipetting error and instrument drift, and measuring 200 cells per animal rather than 5 gives a better estimate of each animal’s value. The Curtis guidance requires one consensus value per unit to go forward into the analysis; it does not discourage technical replication.
Nested designs are not the error either. Nesting is how biology works, and a correctly analysed nested experiment is stronger than a flat one because it separates the variance you care about from the variance you do not.
And the objection does not reach the separate question of whether an experiment was repeated. Independent replication and independent units are different requirements, and a paper can satisfy one and fail the other — see reviewer says the experiment was not replicated. A third and distinct objection is that the test itself was the wrong one for correctly counted units, covered at reviewer says I used the wrong statistical test.
Where a reviewer is wrong on the facts — the cultures were independently derived, the model already carried a random effect they missed — say so once, point to the line, and offer the clarifying sentence you have added. Do not spend a paragraph on it.
Pseudoreplication is the objection the pseudoreplication agent exists for. It reads every figure legend, determines what each n counts, reconstructs the physical structure of the experiment from the methods, and reports every place the stated n exceeds the number of independent units. It is worth a slot on any manuscript where measurements come from a hierarchy — which is nearly all in vivo and cell biology work.
PerfectPaper reconciles the n in each figure legend against the animal, donor and culture counts in the methods, and reports each place the reported sample size counts measurements rather than independent units. The same reconciliation catches the sequencing and cytometry version, where thousands of single cells from one donor become thousands of rows in a differential test and pseudobulk aggregation to the donor is the standard fix.
Related objections: sample size is too small, and the full family on responding to statistical reviewer comments.
Treating measurements that are not independent as though they were. Measuring 50 cells from one animal produces 50 numbers but one independent observation, because those cells share that animal’s genetics, treatment, handling and environment.
The number of units the treatment was independently applied to, and across which biological variation occurs. Usually animals, donors, independently derived cultures, or independent experiments — not the measurements taken from them.
Concede, name the experimental unit, re-run the analysis at that level, and update the abstract, results, figures, legends and supplement together. Give both the number that changed and the number that did not, and list every location you revised so the reviewer can check them in one pass.
Only if the dishes are genuinely independent — separately derived cultures rather than aliquots split from one suspension. Wells split from a common source share whatever happened to that source, so they are technical replicates.
A mixed model lets you use all the data, which is not the same as keeping the p-value. The model estimates variance at the correct level, so the inference reflects your true number of independent units. The interval usually widens; that is the point.
Only on the facts. If the cultures were independently derived, or the model already carried a random effect the reviewer missed, say so once and point to the line. Arguing that the cell-level difference is visually obvious concedes the statistical point while appearing to dispute it.
Structurally yes. Patients within clinics, students within schools and cells within animals are the same problem, and the remedies are the same family: analysis at the cluster level, or a model with a random effect for cluster.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect