Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
Orthogonal validation is a demand about the measurement, not the analysis. Which second method counts depends on which failure mode the reviewer has in mind.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
A reviewer asking for orthogonal validation — “the finding should be validated by an independent approach” — is not commenting on your statistics or your controls. The comment is about the measurement itself: they are entertaining the possibility that your assay produced this result for a technical reason rather than a biological one.
The useful question is which technical failure they have in mind, because a second method that shares the same failure mode validates nothing.
Orthogonal means the second method does not share the first one’s principal source of error — not merely that it is a different technique.
Two antibody-based assays against the same epitope are not orthogonal — they share the antibody’s specificity problem. An antibody-based assay and a mass spectrometry measurement are, because they fail for unrelated reasons. Two sequencing-based quantifications sharing a library preparation are not orthogonal to each other; sequencing and a hybridisation assay are.
Before choosing, name the failure mode the reviewer is worried about. Then pick a method that cannot fail that way.
Orthogonality is a property of a pair of methods with respect to a named failure mode, not a property of a technique in isolation. A western blot and a targeted mass spectrometry run performed on aliquots of one lysate are orthogonal with respect to antibody specificity and not orthogonal with respect to anything that happened during lysis — proteolysis, phosphatase activity, incomplete solubilisation of a membrane fraction. Everything downstream of a single sample preparation inherits that preparation’s artefacts. When the doubt is about the sample rather than the detection chemistry, orthogonality has to be pushed upstream to independently collected biological material.
The International Working Group for Antibody Validation set out five conceptual validation strategies in Nature Methods in 2016 (Uhlén and colleagues, “A proposal for validation of antibodies”): genetic strategies, orthogonal strategies, independent antibody strategies, expression of tagged proteins, and immunocapture followed by mass spectrometry. Four of the five are ways of breaking a shared failure mode, and the paper is worth naming in a response letter because it gives an editor a published standard to judge your answer against rather than a matter of taste.
A one-sentence test settles most cases. Write out “this result would be wrong if __“, fill the blank with the most plausible technical explanation, then ask whether the proposed second method would still return the same answer if that sentence were true. If it would, the second method is not orthogonal to the concern, however different the instrument looks.
Reviewers write one sentence about independent validation while thinking about one of roughly six specific failures, and the remedy differs for each.
Specificity and cross-reactivity. The reagent reports something other than the intended target — an antibody binding a paralogue, a probe cross-hybridising, a small molecule hitting a second kinase. Break it with a genetic control (knockout, knockdown with rescue) or a reagent-independent readout such as mass spectrometry.
Amplification and normalisation artefact. The number is a ratio produced by a pipeline with several tunable steps: reverse transcription efficiency, PCR efficiency, reference-gene choice, library depth, spike-in scaling. Break it with a method that does not amplify, or with an absolute measurement.
Technical confounding with the comparison. Treatment and control were processed on different days, plates, chips or flow cells, so the difference is inseparable from the run. A second method run on the same layout reproduces the artefact perfectly. This is a batch effect problem, and the orthogonal answer has to change the batch structure, not the assay.
Indirect readout. The assay measures a proxy — luminescence for viability, a reporter for transcription, a dye for calcium — and the perturbation may act on the proxy. Break it with a readout on a different physical principle.
Off-target perturbation. The phenotype came from a manipulation that did more than intended: seed-based short interfering RNA effects, CRISPR cutting at a second locus, a drug at a concentration well above its reported potency. Break it with a second, non-overlapping reagent plus a rescue.
Analysis-choice dependence. Threshold, peak caller, gating strategy or segmentation decided the result. A second wet-lab method addresses this only incidentally; what the reviewer needs is the result under alternative pre-specified analysis choices, which for imaging is the subject of image quantification reporting.
Protein abundance — western blot validated by mass spectrometry, targeted proteomics, or a functional readout. A second antibody helps only if it targets a different epitope, and say so if it does. Targeted proteomics in a selected- or parallel-reaction-monitoring mode identifies the protein by peptide mass and fragmentation rather than by affinity, which is the reason it answers a specificity objection; report the peptides monitored, since a single shared peptide between paralogues reintroduces the ambiguity the experiment was meant to remove.
Transcript abundance — RNA-seq validated by qPCR is conventional, though the two share reverse transcription, so it is weaker than it appears. A hybridisation method or a protein-level readout is stronger. If qPCR is what you have, report it to the MIQE standard (Bustin and colleagues, Clinical Chemistry, 2009) — amplification efficiency, reference-gene selection and its justification, and RT conditions — because those are precisely the parameters the shared failure mode runs through.
Genomic occupancy — ChIP-seq validated by CUT&RUN, or by a reporter assay testing whether the bound site does anything. CUT&RUN (Skene and Henikoff, eLife, 2017) still uses an antibody but replaces crosslinking, sonication and immunoprecipitation, so it is orthogonal to fragmentation and background-capture artefacts and not orthogonal to epitope specificity. The ENCODE and modENCODE ChIP-seq guidelines (Landt and colleagues, Genome Research, 2012) remain the reference an occupancy reviewer is likely to have in mind, including antibody characterisation and replicate reproducibility.
Interaction — co-immunoprecipitation validated by proximity ligation, by a biophysical binding measurement, or by a structural constraint. In situ proximity ligation (Söderberg and colleagues, Nature Methods, 2006) adds a distance constraint of roughly 40 nm in intact cells, so it removes the post-lysis-association artefact that is the standard objection to a co-immunoprecipitation. A binding measurement on purified components — surface plasmon resonance, isothermal titration calorimetry, microscale thermophoresis — removes the cellular context instead, which is a different orthogonality and answers a different doubt.
Localisation — imaging validated by fractionation. Fractionation must be shown to have worked: report the marker proteins for each compartment in the same blot, since a fractionation with no purity controls is not evidence that outranks the image it was meant to check.
Screen hits — always validated individually, by a method that does not share the screen’s readout. A counter-screen in the same plate format with the same detection chemistry confirms the screen, not the biology.
Extracellular vesicles and other operationally defined entities — where the object is defined by the method that isolates it, orthogonality is mandatory rather than optional. The MISEV2018 guidelines (Théry and colleagues, Journal of Extracellular Vesicles, 2018) require characterisation by more than one technique, including one giving single-particle information, for exactly this reason.
A reviewer asking for orthogonal validation of a knockdown phenotype is almost always naming off-target activity, and there are three distinct answers with different strengths.
Suppose short hairpin RNA against gene X reduces protein X by 80% and halves colony formation. The weakest answer is a second short hairpin RNA, because hairpins designed against the same transcript can share seed sequences and therefore share off-targets. A stronger answer is a chemically distinct reagent class — an antisense oligonucleotide, or CRISPR interference at the promoter — since the off-target spectrum is generated by different rules. The strongest answer is a rescue: re-express an RNA interference–resistant coding sequence and show the phenotype returns to baseline, which is orthogonal to off-target activity by construction because the only thing restored is the intended target.
A knockout is a different experiment, not a stronger version of the same one. Deleterious mutations can trigger transcriptional adaptation that knockdowns do not — Rossi and colleagues reported this in Nature in 2015, and El-Brolosy and colleagues traced a mechanism to mutant messenger RNA degradation in Nature in 2019 — so a knockout that fails to reproduce a knockdown phenotype is evidence of something, but not automatically evidence that the knockdown was an artefact. When your two genetic approaches disagree, say what the disagreement is and do not average it away; that specific situation is treated in when your knockdown and knockout differ.
Often. If the reviewer doubts that your measured increase in protein X is real, a second abundance measurement addresses it. But if what they doubt is that the increase matters, a functional experiment — does perturbing X change the phenotype — answers a more useful question and frequently satisfies the objection more convincingly than a technical replication would.
Read the rest of the review to tell these apart. A reviewer who also asks about mechanism wants function. One who asks about antibody validation wants a second measurement.
Three signals in the review discriminate reliably. Words about reagents — clone, catalogue number, validation, specificity, RRID and resource reporting — mean the doubt is about the measurement. Words about consequence — necessary, sufficient, causal, epistasis, dose-dependence — mean the doubt is about the claim, and the same reviewer will usually have written a separate comment asking that the mechanism be shown. Words about magnitude — modest, within the range of variability, borderline — mean the doubt is about the statistics, and no second method fixes an underpowered comparison.
A functional answer also lets you satisfy two comments with one experiment, which matters when a revision has a deadline. A rescue experiment is simultaneously an orthogonal control for off-target activity and a sufficiency test.
Design the validation as an independent experiment, because a second method run on the same aliquots on the same day inherits everything the reviewer is worried about.
Use independently collected biological material. New animals, new passages, new patient samples, new differentiations. Re-measuring the stored lysates from the original experiment tests the detection step and nothing before it.
Break the batch structure deliberately. Randomise treatment and control across runs, plates and operators, rather than processing all controls first. A validation whose layout mirrors the original layout can only agree.
Count the replicates the reviewer will count. Three technical measurements from one dish is one biological replicate; presenting it as n = 3 is pseudoreplication, and a validation experiment that commits it invites a second round of the same objection with a sharper tone.
Pre-specify the threshold and the analysis. Decide before unblinding what counts as agreement — a direction, a fold-change floor, a statistical criterion. Otherwise the validation becomes a search for a version that matches.
Blind the readout where it is subjective. Manual scoring, image segmentation choices and gating are the readouts where blinding changes results most, and stating that the analyst was blinded is a sentence editors and reviewers register.
Report the validation at full scale. Show every replicate, including the ones that disagreed, and include the raw blots or images in a supplementary file. A validation reported as a single representative panel is read as a selected panel.
Two methods agree when they agree in direction, in magnitude and across a range — not when both produce a significant p-value in the same direction on one target.
Reporting a correlation coefficient between two assays measuring the same quantity across a panel of targets is the most common weak claim in this space, because a high correlation is compatible with a large constant offset or a systematic proportional bias. Method-comparison statistics exist for this: the Bland–Altman difference plot (Bland and Altman, The Lancet, 1986) displays the bias and the limits of agreement directly, and Deming regression is the appropriate fit when both methods carry measurement error. For a validation of one target rather than a panel, the equivalent is to show the two methods across a dose or time series and demonstrate that the shapes match, not only the endpoints.
Dynamic range is the trap that catches otherwise careful validations. A qPCR assay resolving a 1.3-fold difference and an RNA-seq experiment sequenced to a depth at which that difference is inside the noise are not in conflict; the second method simply cannot see what the first reported. Say which method has the resolution for the effect size in question before you use it as a check, and if it does not, the comparison is uninformative rather than confirmatory — a question of statistical resolution as much as of chemistry.
Report the disagreement, then design the experiment that discriminates between the two explanations. Concealing it is the failure mode that turns a revision into a rejection, because a reviewer who later obtains the omitted result reads it as selective reporting rather than as an inconclusive experiment.
Disagreement has a small number of explanations and they are separable. The two methods may be measuring different things — total protein versus a modified form, steady-state transcript versus nascent transcription, occupancy versus function. One method may lack the resolution for the effect size. The biological material may differ between the two experiments in a way nobody recorded. Or one result is wrong. Work through those in that order, because the first three are frequent and the fourth is the one everybody assumes.
Write the disagreement into the manuscript with its interpretation attached. “Transcript-level and protein-level measurements diverge for X, consistent with post-transcriptional regulation, which we did not test” is a publishable sentence. “Results were confirmed by an independent method” over a supplementary figure that shows partial agreement is not. The general handling of conflicting measurements is covered in when replicates disagree.
State what the alternative would be, why it is not available, and what the limitation means for the claim. Then check whether the claim can be narrowed so the validation is no longer load-bearing — a result described as “detected by ChIP-seq” is a smaller and safer claim than one described as established occupancy.
Narrowing is more powerful than most authors expect, because it converts an unmet demand into an accurate sentence. Every claim in a paper sits somewhere on a ladder — detected, associated, correlated with, required for, sufficient for, causal — and orthogonal validation is load-bearing only at the rungs above association. Moving down one rung usually removes the objection entirely, and an overclaim pass exists to find the sentence that sits a rung higher than the data.
Independent evidence that requires no bench time is worth exhausting first. Public repositories — the ENCODE portal, the Gene Expression Omnibus, PRIDE, the Human Protein Atlas, DepMap — frequently contain a measurement of your target made by a different method in a different laboratory, and a re-analysis of one of those datasets is orthogonal in exactly the sense being demanded. State the accession, the analysis you ran and its limitations. This answer is weaker than a new experiment and is very often accepted, particularly for a secondary claim.
Where the barrier is that the method does not exist in your system — no validated antibody for the species, no genetic tractability, a tissue that cannot be fractionated — name that constraint specifically in the limitations section rather than in the response letter alone. A constraint stated only to the reviewer disappears from the published record, and the next reader raises it again.
Reviewers rarely use the word “orthogonal” first, and recognising the phrasing variants tells you which failure mode is in play.
“The conclusion rests entirely on a single antibody; please confirm with an independent approach.” “The qPCR validation uses the same RNA and the same reverse transcription step as the sequencing, so it is not independent.” “Have the authors confirmed these interactions by a method that does not involve lysis?” “A second, non-overlapping short interfering RNA and a rescue experiment are required before the phenotype can be attributed to the intended target.” “The screen hits should be confirmed in a secondary assay with a distinct readout.” “Can the authors exclude that this reflects a difference between the batches rather than between the conditions?” “The authors state that the result was validated, but the validation experiment appears to have been performed on the same samples.”
The last of those is the comment that most often follows a revision, and it is the reason the design section above matters more than the choice of technique. Answering it in the response letter means saying which specific failure mode the second method cannot share and why — a structure worth carrying into every point of the response to reviewers, where the strongest form is to restate the reviewer’s concern as a failure mode, name the method chosen, and state in one sentence what that method cannot do wrong.
For antibody and reagent provenance, data availability and resource identification checks clones, RRIDs, authentication and validation reporting — the details whose absence prompts this objection. If the underlying doubt is about the claim rather than the measurement, the overclaim check will find the sentence that outran the evidence.
More in this family: reviewer comments on methods and design.
Reviewer wants more controls · Reviewer says the finding is correlative, not causal
PerfectPaper reads each measurement claim against the method that produced it and reports where a stated validation shares a failure mode with the original assay — the same RNA preparation, the same epitope, the same batch structure — before a reviewer does.
Confirming a result with a method that does not share the first method’s main source of error. Two techniques that could both fail for the same reason are not orthogonal, however different they look. Orthogonality is defined against a specific failure mode, so the same pair of methods can be orthogonal for specificity and not orthogonal for sample preparation.
Only partially. Both depend on reverse transcription and on the same RNA preparation, so they share failure modes. It is conventional and often accepted, but a protein-level or hybridisation-based readout is a stronger answer. If qPCR is the available option, report efficiency and reference-gene justification to the MIQE standard, since those are the parameters the shared failure runs through.
If it targets a different epitope, and you say so, it addresses specificity meaningfully. Two antibodies against the same epitope share the problem the reviewer is raising. A genetic control — signal lost in a knockout or knockdown line — is stronger than any second antibody, and is one of the five validation strategies named by the International Working Group for Antibody Validation in 2016.
Frequently, and it is often the better answer, because it tests whether the measured change matters rather than only whether it is real. Which the reviewer wants is usually clear from their other comments: reagent vocabulary means they doubt the measurement, causal vocabulary means they doubt the claim.
Say so explicitly, state what it would have established, and narrow the claim so the validation is not load-bearing. An acknowledged limitation is more acceptable than a defended gap. Check the public repositories first — a re-analysis of an independent dataset measuring your target by another method is orthogonal evidence that requires no new experiment.
Name the failure mode you understand them to be raising, name the method you chose, and state in one sentence what that method cannot do wrong. Then show that the validation used independently collected material and a different batch structure, because a second method run on the original aliquots answers only the detection step.
Report both results and say what would explain the divergence — different molecular species, insufficient resolution in one assay, unrecorded differences in the material, or one result being wrong. Then narrow the claim to what both methods support. Omitting the disagreeing experiment is the response that converts a revision into a rejection.
Last updated September 10, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect