Skip to content

SOLUTIONS

What is a batch effect?

A batch effect is systematic technical variation between samples processed together. It biases results when batch correlates with the condition studied.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

What is a batch effect?

A batch effect is systematic technical variation introduced when samples are processed in separate groups — different days, reagent lots, sequencing runs, plates, operators or sites — that shifts measured values independently of biology. Batch effects distort high-throughput results whenever the grouping is correlated with the condition under study, and can produce differences that look biological but are not.

Batch effects were characterised as a widespread problem across high-throughput technologies by Leek and colleagues in their 2010 Nature Reviews Genetics review, “Tackling the widespread and critical impact of batch effects in high-throughput data”. The years since have produced a large correction literature and much less change in how most manuscripts describe their processing. This page covers what counts as a batch, why the effect is structured rather than random, the single condition that makes it fatal, how it is detected in a submitted manuscript, what the named correction methods assume, and what reviewers write when it has been handled badly.

What counts as a batch

A batch is any group of samples that shared a processing circumstance which other samples did not share. The list is assay-specific and longer than most methods sections acknowledge.

Bulk RNA sequencing: extraction day, RNA extraction kit lot, library preparation kit lot, the person who prepared the libraries, the multiplexing pool, the flow cell, the lane within a flow cell, and the sequencer instrument. Read length and chemistry version change quantification outright.

Single-cell RNA sequencing: the 10x Chromium chip and channel, the chemistry version (v2 libraries recover markedly fewer transcripts per cell than v3), dissociation protocol and duration, time from tissue collection to loading, ambient RNA background, and the reference genome and aligner version used for counting.

Methylation arrays: the Illumina array chip itself and the sample’s physical position on it. The HumanMethylation450 BeadChip carries 12 samples per chip in a six-row, two-column arrangement, and the EPIC array carries eight; position on the chip is a documented source of variation on both, which is why chip and position belong in the model rather than in the supplementary table.

Mass spectrometry proteomics and metabolomics: the TMT or iTRAQ plex, the LC column, a guard column change, the order of injection within a run, and instrument drift across a run. Signal drift within a single continuous batch is measurable, and it is the reason pooled quality-control samples are injected at intervals.

Anything on plates: row and column position, edge wells, the plate itself, and the day the plate was read.

Multi-site studies: the site, which bundles every one of the above into a single variable that is also correlated with the population recruited.

Why a batch effect is not random noise

A batch effect is structured, not stochastic, and that structure is what makes it dangerous. Random measurement noise is independent across samples and features; it widens confidence intervals and reduces power, and averaging more samples reduces it. Batch effects do neither.

Batch effects act coherently across thousands of features at once, which is why they dominate the leading principal components of almost every high-dimensional dataset. They are also feature-specific in magnitude and direction: a reagent lot change may raise the measured abundance of GC-rich transcripts and lower others, so a single global scaling factor cannot remove it. Benjamini and Speed documented GC-content-dependent coverage bias in high-throughput sequencing in 2012, and library preparation batches differ in exactly that bias.

Batch effects also affect variance, not only means. Two batches can share a mean and differ in dispersion, which alters the standard error of every downstream test even when centring the data appears to have fixed the problem. A method that shifts means and leaves variance untouched leaves half of the effect in place.

The practical consequence is uncomfortable: collecting more samples within the same confounded design increases the precision of a biased estimate. This is the same failure described in my p-value shrinks as I add more cells — more measurements of the same non-independent units narrow the interval around the wrong number.

The condition that makes a batch effect fatal

A batch effect harms an analysis only when batch is correlated with the biological variable of interest. That correlation, not the size of the technical variation, decides whether a result survives.

A dataset can have enormous batch structure and still support valid inference, provided cases and controls are distributed evenly across every batch: the batch term absorbs the technical shift, the contrast of interest is estimated within batches, and the penalty is a modest loss of precision. The same dataset with the same technical magnitude, but with all cases on plate 1 and all controls on plate 2, supports no inference at all about that contrast.

Batch is therefore a confounder in the strict causal sense when it influences both which condition a sample carries — usually through collection logistics — and the measured outcome. It satisfies every part of that definition, and it is the one confounder in genomics that investigators can eliminate entirely at the design stage and routinely do not.

Partial correlation is the common real case and the hardest to reason about. When batch and condition are 70% aligned rather than fully aligned, a correction method will attribute some genuine biological signal to batch and some batch signal to biology, and no diagnostic run on the data can tell you the split.

A worked example: population differences that were processing dates

Spielman and colleagues reported in 2007 that a substantial proportion of genes were differentially expressed between HapMap lymphoblastoid cell lines of European and Asian ancestry. Akey, Biswas, Leek and Storey responded the same year in Nature Genetics, pointing out that the European and Asian samples had been processed at different times, so ancestry was confounded with processing date; comparing two processing dates within a single population produced more differentially expressed genes than the population comparison itself.

The instructive detail is that nothing about the original analysis was careless in the conventional sense. The arrays worked, the normalisation was standard, the statistics were correct, and the finding was biologically plausible — expression differences between human populations do exist. The design was the defect, and the design was invisible in the reported analysis because processing date was not in the model and was not in the paper.

That is the shape of the problem. A batch effect does not announce itself as an artefact; it arrives as a clean, well-powered, mechanistically sensible result.

How a batch effect is detected in a real manuscript

Detection of a batch effect starts with metadata, not with the expression matrix, and an experienced reviewer works in a fixed order.

Look for the processing variables in the methods. If the manuscript does not state when samples were extracted, how many library preparation batches were used, how many sequencing runs, and how samples were allocated to them, the first finding is that batch is unassessable — which is a reportable defect, not a neutral omission.

Cross-tabulate batch against condition. A two-by-two table of batch by group answers the fatal-condition question directly. If that table has structural zeros, no correction method rescues the comparison.

Colour the principal components by technical variables. Project samples onto PC1 and PC2 and colour by run, plate, date and operator in turn. If a technical variable separates the samples on a leading component, the effect is large relative to biology. This is a screening tool rather than a test: absence of separation on PC1 does not mean absence of a batch effect on the features that matter.

Test the association formally. Regress each of the top principal components on each recorded technical covariate and report the association, or use guided principal component analysis (Reese et al., Bioinformatics, 2013), which produces a delta statistic with a permutation-based null built specifically for batch detection. Principal variance component analysis apportions total variance across biological and technical factors and is readable by non-specialists.

Look at relative log expression plots. RLE plots, described by Gandolfo and Speed in 2018, show per-sample deviations from a feature-wise median; systematic shifts in the median or interquartile range of whole groups of samples are a direct visual signature of a batch, and they reveal the variance differences that a PCA scatter plot hides.

For single-cell data, quantify mixing. kBET (Büttner et al., Nature Methods, 2019) tests whether the batch composition of a cell’s local neighbourhood matches the global composition. The local inverse Simpson’s index introduced with Harmony (Korsunsky et al., Nature Methods, 2019) reports the effective number of batches in a neighbourhood. Both must be read alongside a conservation metric, for the reason given further down this page.

One heuristic is worth stating for anyone reviewing genomics submissions, and it is implemented in the genomics review lane: if the sample table and the batch table cannot be reconstructed from the manuscript and its deposited metadata, treat every downstream claim as unverified rather than as unsupported. The distinction changes what you ask the authors to do.

Design choices that prevent batch effects

Design is the only intervention that removes a batch effect rather than modelling it, and four choices do nearly all the work.

Randomise sample-to-batch allocation. Assign samples to extraction days, plates, positions and sequencing runs by a randomised schedule, and record the schedule. Randomisation protects against technical variables nobody thought to record, which is the class that correction methods cannot address by construction.

Block deliberately when randomisation is not available. If samples arrive over two years, ensure every batch contains a balanced mix of conditions. Blocking beats randomisation at small n, because a randomised allocation of 12 samples can produce a badly unbalanced draw by chance.

Carry bridging samples. Include the same reference material, or the same set of technical replicates, in every batch. Bridging samples make the batch shift directly estimable rather than inferred, and they are the only mechanism that lets you calibrate batches processed years apart. Commercial reference RNA, pooled quality-control samples in metabolomics, and a common reference channel in TMT proteomics all serve this function.

Record everything, including what seems irrelevant. Reagent lot numbers, instrument identifiers, operator, and processing timestamps take minutes to record and cannot be reconstructed afterwards. Depositing them alongside the data is now expected by many journals; the data availability and accessions lane checks whether deposited metadata actually contains the processing variables an analysis claims to have adjusted for.

Batch as a model term versus batch correction as a preprocessing step

Including batch as a term in the statistical model and removing the batch effect from the data before analysis are not two routes to the same answer, and conflating them is the most common analytical error in this area.

Batch as a covariate. Adding batch as a fixed effect in a linear model, a generalised linear model, or a mixed model with batch as a random effect estimates the biological contrast within batches and propagates the uncertainty of the batch estimate into the standard errors. Degrees of freedom are spent, and the reported confidence interval reflects that.

Batch correction as preprocessing. Estimating batch parameters, subtracting them, and passing the adjusted matrix to a downstream test that treats it as raw data hides the fact that those parameters were estimated at all. The downstream test then behaves as though it has more information than it does.

The limma package draws this distinction explicitly: the documentation for removeBatchEffect states that the function is not intended for use before linear modelling, and that batch factors should instead be included in the linear model, with the corrected matrix reserved for visualisation and clustering. That guidance is widely quoted and widely ignored, and a corrected matrix appearing as the input to a differential expression call is a specific, checkable reviewer finding.

The named correction methods and what each assumes

Each batch-correction method encodes a different assumption about what the batch did, and the assumption determines whether the method is appropriate.

ComBat (Johnson, Li and Rabinovic, Biostatistics, 2007) fits a location-and-scale model per feature and shrinks the batch parameters towards a common distribution using empirical Bayes. That shrinkage is what makes it usable with few samples per batch. ComBat assumes batch effects are additive in location and multiplicative in scale on the analysed scale, and it requires batch to be known.

ComBat-seq (Zhang, Parmigiani and Johnson, 2020) reworks the same idea for RNA-seq count data using a negative binomial model, so that adjusted values remain integers suitable for count-based differential expression tools. Applying the original ComBat to log-transformed counts and then running a count-based test is a mismatch worth flagging.

Surrogate variable analysis (Leek and Storey, PLoS Genetics, 2007) estimates unmeasured sources of heterogeneity directly from the data and supplies them as covariates. Surrogate variable analysis addresses the batch variables you did not record, which is its whole point, and it depends on those surrogate variables being separable from the biological signal.

Remove unwanted variation (Gagnon-Bartsch and Speed, Biostatistics, 2012) uses negative control features assumed unaffected by the biology to estimate the unwanted variation. Its validity rests entirely on the control set being genuinely unaffected, which is an assumption about biology, not about statistics.

Functional normalisation (Fortin et al., Genome Biology, 2014) uses the control probes built into Illumina methylation arrays to capture technical variation, and was designed for studies in which global methylation differences between groups are expected and should not be normalised away.

None of these methods can distinguish batch variation from biological variation that happens to align with batch. All of them remove biological signal in proportion to how confounded the design is.

Why correcting first and testing second inflates confidence

Applying a batch-correction method and then running an ordinary differential test on the corrected data produces anti-conservative p-values when the design is unbalanced. Nygaard, Rødland and Hovig demonstrated this in Biostatistics in 2016, in a paper whose title states the finding plainly: methods that remove batch effects while retaining group differences may lead to exaggerated confidence in downstream analyses.

The mechanism is not the loss of degrees of freedom that tool documentation warns about — the authors show that matters mainly for small batches. Their finding is a systematic inflation of the F-statistic, induced when estimation errors from the batch adjustment are applied across the whole dataset, and it is as harmful at large sample sizes as at small. The authors’ recommendation is to include batch in the model rather than to correct and then test.

This matters most for the studies most likely to reach for correction — retrospective, unbalanced, assembled from cohorts collected at different times — because the imbalance that motivates correction is the same imbalance that breaks it. Combined with the multiplicity burden of testing tens of thousands of features, described in the multiple testing in omics lane, an inflated per-feature null propagates into a false discovery rate estimate that is itself wrong.

Batch effects in single-cell data are a different problem

Single-cell integration optimises a different objective from bulk batch correction, and treating the two as the same produces a specific class of error. Bulk correction aims to remove a shift between groups of samples. Single-cell integration aims to align cell populations across datasets so that the same cell type from two experiments occupies the same region of a shared embedding, while keeping distinct cell types apart.

Mutual nearest neighbours (Haghverdi et al., Nature Biotechnology, 2018) identifies pairs of cells that are each other’s nearest neighbours across batches and uses those pairs to estimate the correction, which allows batches with non-identical cell type composition. Harmony (Korsunsky et al., Nature Methods, 2019) iterates soft clustering and linear correction in a reduced-dimensional space. Deep generative approaches such as scVI model batch as a covariate in a latent variable model.

The failure mode specific to this class is over-integration: forcing perfect batch mixing merges genuinely distinct populations, and it does most damage to the rare and the novel — the cell states that motivated the experiment. Any mixing metric can be maximised by destroying biology, so a mixing score reported without a conservation score is uninterpretable. Luecken and colleagues made this explicit in their Nature Methods benchmark in 2022, which scored 68 method-and-preprocessing combinations across 85 batches on both batch-removal and biological-conservation metrics, and found that method ranking depends on which of the two objectives is weighted and on the complexity of the integration task.

A second single-cell-specific issue is that batch is usually the sample, and cells are not independent replicates of that sample. That is a pseudoreplication problem layered on top of the batch problem, and it needs a pseudobulk or mixed-model treatment that no integration method provides.

When a batch effect cannot be corrected at all

A batch effect is uncorrectable when batch and condition are completely confounded, and the honest response is to say so rather than to apply a method.

If every case was processed in run 1 and every control in run 2, the batch parameter and the condition parameter are the same parameter. No statistical method separates them, because the separation is not identifiable from the data. ComBat applied to such a design returns a corrected matrix that looks reasonable and has had the biological effect removed along with the technical one — or, if the biological covariate is protected during correction, has had the technical effect preserved and relabelled as biology.

Three responses are defensible. Report the comparison as confounded and describe what would be needed to resolve it. Re-process a balanced subset of samples in a single batch, even a small one, and use it to check the direction of the main finding. Or validate the specific claims on an independent platform and an independent sample set, which is what orthogonal validation means in this context and why reviewers ask for it here in particular.

What is not defensible is applying a correction method to a fully confounded design and reporting the output without stating that the design was confounded.

What a peer reviewer says when a batch effect is mishandled

Reviewer comments about batch effects are formulaic enough to anticipate, and anticipating them is the cheapest revision available to an author.

“Samples appear to have been processed in two batches that correspond to the case and control groups; the reported differences cannot be distinguished from technical variation.”

“The authors do not report how samples were allocated to sequencing runs. Please provide the full sample-to-batch mapping.”

“ComBat was applied prior to differential expression testing. Please justify this rather than including batch as a covariate in the design matrix, and report the results of the covariate-based analysis.”

“Principal component analysis is shown coloured by phenotype only. Please provide the same plot coloured by processing date, plate and sequencing run.”

“The authors state that data were corrected for batch effects. Which variables were treated as batch, and how were the batch parameters estimated?”

“Cells from the same donor are treated as independent observations, and donor is also the batch.”

That last comment most often ends a single-cell submission, and it is the same objection described in reviewers saying n is not independent and in my replicates disagree. A reviewer who writes it has usually already decided the analysis needs redoing rather than rewording.

What to report

Report the batch structure as data, not as reassurance. A methods section that satisfies a methodologically literate reviewer contains six specific items.

The processing variables recorded, named individually: extraction date, kit lot, library preparation batch, plate and position, sequencing run, lane, instrument, site, operator.

The allocation mechanism: randomised, blocked, or convenience order, stated plainly. “Samples were processed in the order received” is an acceptable sentence and a far better one than silence.

The cross-tabulation of batch against every experimental factor, in the supplement, as a table.

The diagnostic evidence: which technical variables associated with which principal components, with the statistic and the p-value, or the guided principal component analysis delta and its permutation p-value.

The handling: the exact model formula including the batch terms, or the correction method with its version and every non-default argument. If a corrected matrix was used for anything, state which figures and analyses used it and which used the uncorrected data.

A sensitivity analysis: the primary result with and without batch handling, and where feasible with a leave-one-batch-out repetition. Divergence between them is informative, and reporting that divergence is stronger than concealing it.

Contested and unresolved

Three things about batch effects are genuinely unsettled, and saying so is more useful than implying the field agrees.

Whether to correct at all, or only to model. The Nygaard et al. result argues for including batch in the model. Researchers assembling data from many published sources often have no shared design matrix to model within, and correction is the only option available to them. No consensus rule states where that boundary sits.

How much biological signal correction removes. Every method removes some, the amount depends on the degree of confounding, and no published diagnostic reliably quantifies it on real data where the truth is unknown. Simulation studies answer the question only under their own generative assumptions.

Which single-cell integration method to use. Benchmarks including Luecken et al. (2022) rank methods differently by task, metric weighting and preprocessing, and the ranking is not stable across datasets. A method that wins on one atlas is not thereby the right choice for a two-batch experiment with a rare population of interest.

Unmeasured batch variables remain the standing limitation across all of this, in the way residual confounding is the standing limitation in observational epidemiology. Surrogate variable analysis partly addresses it and does not eliminate it. A study that reports its batch structure fully is not thereby unbiased; it is inspectable, which is the achievable standard.

Related

Confounding · Pseudoreplication · My effect disappeared after adjusting

Checked before submission by the batch effects reviewer lane, which reconstructs the sample-to-batch mapping from the manuscript, flags factors confounded with processing groups, and names where a corrected matrix was used as the input to a statistical test.

Review my manuscript

Frequently asked questions

What does “batch effect” mean?

A batch effect means systematic technical variation between groups of samples that were processed together — the same day, plate, reagent lot, instrument or site. The variation shifts measured values with no biological cause, and it becomes a bias in the reported results when the processing groups line up with the experimental groups.

What is a batch effect in simple terms?

A batch effect is a measurement difference caused by how and when samples were handled rather than by what the samples are. Twenty samples run in January and twenty run in March will differ measurably even if they are biologically identical, and that difference is indistinguishable from biology if January’s samples were all controls.

What is a batch effect in RNA-seq data?

A batch effect in RNA sequencing is technical variation attributable to extraction day, RNA kit lot, library preparation batch, multiplexing pool, flow cell, lane or instrument. It affects thousands of transcripts coherently, differs in size and direction from transcript to transcript, and typically dominates the first two principal components of the expression matrix.

What is a batch effect in single-cell RNA sequencing?

A batch effect in single-cell RNA sequencing is technical variation between samples, chips, chemistry versions or dissociation runs that separates cells of the same type into distinct clusters by their sample of origin. Integration methods such as Harmony, mutual nearest neighbours and scVI address it, at the risk of merging genuinely distinct rare populations.

What is an example of a batch effect?

Spielman and colleagues reported large gene expression differences between European and Asian HapMap cell lines in 2007. Akey, Biswas, Leek and Storey pointed out the same year in Nature Genetics that the two ancestry groups had been processed at different times, and that comparing two processing dates within a single population yielded more differentially expressed genes than the population comparison did.

What is a batch effect in a methylation array study?

A batch effect on an Illumina methylation array is technical variation attributable to the chip a sample was run on and its physical position on that chip, along with scan date and reagent lot. The 450k BeadChip holds 12 samples per chip and the EPIC array holds eight, so chip and position belong in the analysis model.

What is batch variation in a proteomics or metabolomics experiment?

Batch variation in mass spectrometry is systematic drift and offset attributable to the TMT plex, the LC column, injection order within a run, and instrument state over time. Pooled quality-control samples injected at intervals make the drift estimable, and a common reference channel makes separate plexes comparable to one another.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.