Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
A standing objection in molecular biology. What counts as causal evidence, and why rewriting the claim is often the correct answer rather than a retreat.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
The correlative-not-causal objection says your data show that X and Y travel together while your text says X produces Y. The reviewer has found the gap between those two sentences, and is right that the results support only the first. The objection is about the verb, and it closes either by a perturbation experiment or a more exact verb.
The complaint is not about the quality of the data, and it has a fixed structure in experimental biology: a perturbation is missing, or a perturbation exists but does not establish what the text claims. Which of the two you are facing determines everything downstream — the first calls for bench work or a rewrite, the second for a control.
Causal evidence in experimental biology means a perturbation of X whose consequences for Y you measured, in a system where nothing else was deliberately changed. In practice, reviewers are looking for some subset of four things.
Loss of function. Remove or inhibit X and Y decreases. The strongest common form, and the one most often expected. Loss of function establishes that X is necessary for Y in your system and says nothing about whether X alone is enough to produce it.
Gain of function. Supply or activate X and Y increases. Weaker alone, because overexpression can act through routes the physiological protein does not — mass action driving a low-affinity interaction, mislocalisation when a targeting sequence is saturated, or squelching of a shared cofactor. State the expression level relative to endogenous protein; a construct running at 30-fold endogenous is a different experiment from one at 2-fold.
Rescue. Remove X, lose Y, restore X, recover Y. This is the one that closes the argument, because it controls for the off-target effects that undermine loss of function alone. A reviewer who has already accepted your knockdown and still objects is usually asking for this. A rescue only does that work if the rescue construct is immune to the original reagent: silent mutations in the siRNA target site, a codon-altered cDNA, or a species orthologue that the sgRNA cannot cut. A rescue with a construct the reagent still targets is not a rescue.
Dose or time dependence. Y tracks the level of X, or follows it in time. Supportive rather than decisive, but genuinely strengthening — and the temporal version is the stronger of the two, because a response that precedes the change in X excludes the ordering your claim requires.
If you have none of these, you have a correlation, however many correlations you have. Twelve cell lines, two patient cohorts and a single-cell atlas all showing the same association are twelve, two and one instances of the same evidentiary type, and no number of them crosses into the next one. Independent replication answers a different objection — see reviewer says the experiment was not replicated — and stacking correlated samples answers neither, which is the independence problem wearing a different hat.
The verb in your title is a claim about evidentiary type, and reviewers read it that way. Most correlative-not-causal comments are triggered by one word.
| Verb | What it asserts | Minimum evidence a reviewer expects |
|---|---|---|
| Is associated with | X and Y co-vary | The measurement, with its dispersion and n |
| Marks, identifies | X predicts Y | The association plus a stated discrimination metric |
| Is required for | Removing X removes Y | Loss of function, ideally two independent reagents |
| Is sufficient for | Supplying X produces Y | Gain of function in a background lacking X |
| Promotes, enhances | X increases Y | Loss of function, with magnitude reported |
| Drives, causes | X produces Y and is the operative agent | Loss and gain of function, plus a rescue |
| Mediates | X carries the effect of an upstream signal to Y | The above, plus epistasis placing X between them |
“Required for” and “sufficient for” are not weaker than “drives” — they are more precise, and precision is what the objection is asking for. A paper that establishes necessity and says so is complete. A paper that establishes necessity and writes “drives” has an unsupported sentence in it, and the reviewer found it before you did.
Correlation and causation come apart in a cell for three reasons, and each one has a different experimental answer.
A shared upstream regulator. A transcription factor or a stress programme raises X and raises Y independently, so they co-vary with no arrow between them. This is confounding in its structural sense, and the reason molecular biology answers it by intervening rather than by adjusting: perturbing X breaks the shared parent’s influence on X while leaving its influence on Y intact, which no covariate list can do.
Reverse causation. Y raises X. Feedback is the rule rather than the exception in signalling, so a protein induced by a process is easy to mistake for a protein that produces it. Time-resolved measurement after an acute stimulus is the cheap discriminator: measure X and Y at five or six points in the first hours, and see which moves first.
The readout is a proxy. X may genuinely cause the thing you measured without causing the thing you named. Phospho-signal, reporter output and marker expression each stand in for a process, and the substitution is where the claim quietly widens. Writing the causal structure down as a directed acyclic graph forces the proxy into the open, and it also separates a variable on the path from a variable beside it — the mediator versus confounder distinction that decides whether adjusting for something helps or destroys the effect you are measuring.
Observational human work cannot perturb anything, which is why it developed a different apparatus. Austin Bradford Hill’s nine viewpoints, set out in “The environment and disease: association or causation?” (Proceedings of the Royal Society of Medicine 1965;58:295–300), are the standard reference, and Hill presented them as considerations to weigh rather than as a checklist to satisfy. Experimental biology has the intervention Hill lacked, which is exactly why reviewers expect it to be used.
A phenotype that consists of less of something is the weakest place to make a causal claim, because almost any perturbation produces less of it. William Kaelin made the point directly in a review of preclinical target validation (Nature Reviews Cancer 2017;17:425–440): decreased proliferation, decreased viability and decreased tumour growth are all “down” readouts, and a down readout can reflect a nonspecific loss of cellular fitness rather than a specific requirement for the gene you perturbed.
Two practical consequences follow. First, a knockdown that reduces proliferation is a weaker specificity argument than a knockdown that changes a differentiation marker, a localisation, or a directional migration, because the second class of readouts is not produced by generic sickness. Second, the control that answers the objection is a perturbation known to be toxic but irrelevant to your pathway; if it reproduces your phenotype, the phenotype is fitness. Reviewers asking for more controls on a viability endpoint are usually asking this question in shorter words.
Do the perturbation. If a knockdown or knockout is tractable in your system, this is the direct answer. Include the rescue if you can — reviewers who have seen many knockdown papers know that off-target effects are common, and a rescue pre-empts the follow-up objection. Estimate the time honestly before you promise it in a response letter: a second independent siRNA or sgRNA plus its validation is a matter of weeks in a cell line, and a rescue construct with an immune target site adds further weeks on top of that.
Rewrite the claim. This is legitimate, frequently correct, and under-used. “X is associated with Y and is required for its induction” is a different sentence from “X drives Y”, and if you only have the association, the first is the true one. A precise correlative finding is publishable. An overstated causal one is not.
The mistake is treating rewriting as defeat. Reviewers read a considered narrowing as evidence the authors understand their own data.
What is not an honest route is a vocabulary swap. Changing “drives” to “plays a key role in” leaves the causal implication intact while removing the specificity, and it converts a checkable claim into an uncheckable one — which is the failure mode a hedge always has. If the title still says the same thing, the abstract still says the same thing and only the discussion has softened, the revision will come back. Narrow the title, the abstract’s final sentence and the discussion’s first paragraph together, or narrow nothing.
Consider an invented but entirely ordinary manuscript. Kinase K is elevated in 68 of 94 tumours relative to matched normal tissue; high K by immunohistochemistry associates with shorter survival; siRNA against K reduces proliferation in two cell lines by roughly 40%; the title reads “K drives tumour progression”.
Trace each claim to the result meant to support it. The tissue comparison supports “K is elevated in tumour tissue” and nothing about progression. The survival association supports “K expression marks a poorer-prognosis subgroup” and, because it is observational and unadjusted, not even that a reviewer will accept without the covariates. The siRNA experiment is a single reagent producing a down readout at a magnitude that overlaps generic fitness loss, so it supports “knockdown of K reduces proliferation in these lines” — a statement about the experiment, not about tumours.
Nothing in the set supports the word “drives”, and the shortest route to a defensible paper is not the mouse experiment. It is a second sgRNA, a rescue with a codon-altered K, a non-proliferation readout, and the title “K is required for proliferation in K-amplified lines and marks poorer survival”. That claim is smaller, entirely supported, and harder to argue with. Whether the 40% effect is worth reporting at all is a separate question, and a real one — see reviewer says the effect size is not meaningful.
A reviewer who has your perturbation in hand and still calls the work correlative is questioning the perturbation itself, not its absence. Usually one of four things:
Specificity. A single siRNA or one inhibitor at high concentration. The answer is a second independent reagent, a rescue, or a documented off-target profile. Off-target regulation by siRNA is a documented, sequence-driven phenomenon rather than a theoretical risk: Jackson and colleagues showed in Nature Biotechnology (2003;21:635–637) that siRNAs silence unintended transcripts through partial complementarity, and later work traced much of that silencing to the seed region, so two siRNAs with different seeds are a real control while two siRNAs tiling the same region are close to one. For CRISPR, two sgRNAs targeting different exons, and a small-molecule inhibitor paired with a resistant allele, carry the same logic.
Magnitude. A knockdown that removes 40% of the protein and produces a large phenotype invites the question of what the residual is doing. Report the knockdown efficiency, always — as a percentage of control protein, quantified against a dilution series rather than a single lane, with the number of independent preparations stated. A large phenotype at small depletion is not automatically wrong; it is a claim about a steep dose–response, and it should be written as one.
Timing. A chronic knockout can produce compensation, so the phenotype may reflect adaptation rather than the acute role. Acute degradation systems answer this; acknowledging it also answers it. The compensation is now mechanistically characterised, which makes it citable rather than hand-waved: Rossi and colleagues found zebrafish egfl7 mutants phenotypically normal while morphants showed severe vascular defects, and traced it to upregulation of extracellular-matrix genes in the mutants (Nature 2015;524:230–233). El-Brolosy and colleagues then reported a trigger in Nature (2019;568:193–197): degradation of the mutant transcript itself sets off the compensating programme, so alleles that transcribe no mutant mRNA at all show no such adaptation and give rise to more severe phenotypes than alleles whose mutant mRNA decays. For acute removal, the auxin-inducible degron (Nature Methods 2009;6:917–922) and the dTAG system (Nature Chemical Biology 2018;14:431–441) deplete a tagged protein on a timescale of minutes to hours, which is short enough to precede transcriptional adaptation. Where a mutant and a knockdown disagree in your own hands, that discrepancy is its own diagnosable result rather than a nuisance to be omitted.
Cell-autonomy. Whether the effect is in the cell where X was perturbed or downstream in another compartment. Often answerable with the data you already have: a mosaic in which perturbed and unperturbed cells sit side by side, a conditioned-medium transfer, a co-culture with the perturbation in one partner, or a conditional allele under a lineage-restricted driver. In vivo, the compartment question is usually the whole objection — the in vivo rigor expectations of randomisation, blinding and a stated unit of analysis exist so that a compartment claim can be read at all.
Some systems do not permit the experiment, and a reviewer who understands that will accept a bounded claim in place of it. State the obstruction explicitly rather than leaving a gap: the gene is essential and null cells are not viable; the material is human post-mortem tissue; no model recapitulates the phenotype; the experiment is a two-year mouse cross.
Four substitutions carry weight in that situation. A conditional or inducible allele converts an essential gene into a tractable one, and a degron converts it further into an acute one. A partial or hypomorphic perturbation with a stated residual can establish dose dependence where a null cannot. An orthogonal measurement of the same association — a different assay with a different failure mode, which is what orthogonal validation means — does not create causal evidence but removes the artefact account of the association, and reviewers frequently accept it as the realistic ceiling. And a naturally occurring perturbation, such as a loss-of-function human variant or a patient-derived line carrying the deletion, is an intervention nobody had to perform.
What does not work is substituting a mechanism cartoon, a pathway-enrichment analysis, or a citation to someone else’s knockout in a different tissue. If the argument rests on another system’s perturbation, name that system and say the claim is inferred from it — which is also the substance of the model relevance question. A separate objection, the mechanism is not shown, asks how X acts rather than whether it acts, and confusing the two produces revisions that answer neither.
Reviewers rarely write “correlative, not causal” as their opening words. They write the specific version, and these are the sentences that decide papers.
“The data establish an association between X and Y; the title and abstract assert causation.” “A single siRNA is insufficient to attribute the phenotype to loss of X; please provide a second independent reagent or a rescue.” “Knockdown efficiency is not reported for Figure 3.” “The rescue construct appears to be targetable by the siRNA used.” “The phenotype is a reduction in viability, which could reflect nonspecific loss of fitness rather than a specific requirement for X.” “The constitutive knockout may have compensated; can the authors deplete the protein acutely?” “It is not clear whether the effect is cell-autonomous.” “The authors write that X drives Y but demonstrate only that X is required for Y.” “The correlation is observed across conditions that differ in many respects.” “Please either provide the perturbation experiment or revise the claims throughout, including the title.”
The last one is worth reading closely, because it is an offer. A reviewer who names revision as an acceptable alternative has already decided the correlative paper is publishable, and the experiment is optional.
Run a verb audit before submission, on the title, the abstract’s last two sentences, the first paragraph of the discussion and every figure legend. Those four locations hold nearly all overclaims, because they are where results are summarised rather than reported.
For each causal verb, write the figure panel that supports it in the margin. A verb with no panel is an overclaim; a verb whose panel is a correlation is the objection this page is about. Then check the reverse direction: a panel showing a perturbation that the abstract never mentions is a claim you already earned and did not make.
Two additional checks catch most of the remainder. First, read the title as a stranger would and ask what experiment it promises; if the promised experiment is not in the paper, the title is the problem. Second, check that the limitations section names the missing evidentiary type by name — “we did not test sufficiency” is a specification, “further work is needed” is not, and reviewers read the second as an absence of analysis rather than an admission.
Name which evidence you have and which you do not, then say what you did about it. If you rewrote rather than experimented, say that explicitly — reviewers notice a claim quietly softened, and they read the silence as evasion rather than as compliance.
A workable shape: “We agree our data establish that X is required for Y but do not establish sufficiency. We have revised the title, abstract and discussion accordingly (changes listed below), and now describe X as required for rather than driving Y.”
When the experiment is genuinely out of reach, the response has a different shape, and it needs three parts rather than two: what the reviewer asked for, why the system does not permit it, and what you did instead. “The reviewer asks for a rescue. K is essential in these lines and stable re-expression is not viable, so we have instead added a second sgRNA targeting exon 2 (new Figure 3d), reported depletion efficiency for both reagents (Supplementary Figure 4), and narrowed the claim to a requirement for K in this context.” A refusal with a substitution reads as engagement. A refusal without one reads as a refusal, whatever its tone. The general structure is in how to write a response to reviewers.
The overclaim check is built for precisely this failure: it traces each claim in the abstract and discussion back to the specific result meant to support it and reports every one whose wording exceeds its evidence, supplying replacement sentences rather than advising caution. For epidemiological rather than experimental work, causal language discipline covers the same failure in observational designs.
PerfectPaper reads the title, abstract, discussion and figure legends against the results they cite, and reports every causal verb whose supporting panel shows an association rather than a perturbation. The check runs on the whole manuscript, not a sampled section, because an overclaim in a figure legend is as fatal at review as one in the abstract.
More in this family: reviewer comments on methods and design.
Usually a perturbation: loss of function, ideally with a rescue that restores the phenotype. Gain of function and dose dependence support the case. Correlation across conditions, however strong, does not establish it.
A knockdown alone is frequently accepted when the phenotype is large and a second independent reagent gives the same result. A rescue is what settles the specificity question, and reviewers ask for it most often when only one reagent was used. The rescue construct must be immune to the original reagent — a codon-altered cDNA or silent mutations in the target site — or it does not control for anything.
Yes, described as correlative. Association studies are a legitimate contribution when the claims match the design. What draws rejection is causal language attached to correlative evidence.
Say what the data establish precisely rather than vaguely. “Required for” is a real and interesting claim. “Associated with, and necessary for induction in this system” is more informative than “drives”, not less — it tells the reader what was actually tested. Changing “drives” to “plays a key role in” is not a rewrite; it removes the specificity and keeps the implication.
The objection has moved to specificity, magnitude, timing or cell-autonomy rather than the absence of a perturbation. Report knockdown efficiency, add a second reagent or a rescue, and address compensation if the system is a chronic knockout.
Say why, in the response and in the limitations, then substitute what the system does allow: a conditional or degron allele, a hypomorph with a stated residual, a naturally occurring loss-of-function variant, or an orthogonal measurement that removes the artefact account of the result. Then narrow the claim to what those support.
No. Dose dependence shows that Y tracks the level of X, which is consistent with causation and also with a shared upstream regulator that sets both levels. Temporal ordering after an acute stimulus is stronger, because a response that precedes the change in X rules out the direction your claim needs.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect