Author interview
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
SOLUTIONS
The ARRIVE guidelines are a 21-item checklist for reporting animal research, published by the NC3Rs in 2010 and revised as ARRIVE 2.0 in 2020 with an Essential 10.
Affiliations
Current platform
PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.
Prepare
PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.
Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.
Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.
Investigate
Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.
Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.
The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.
Revise
Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.
Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.
Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.
The ARRIVE guidelines are a reporting checklist for animal research — Animal Research: Reporting of In Vivo Experiments — published by the NC3Rs in 2010 and revised as ARRIVE 2.0 in 2020. ARRIVE 2.0 contains 21 items, split into an Essential 10 that every manuscript must report and a Recommended Set of 11 adding context to the study described.
ARRIVE is a reporting standard, not a design standard. The guidelines say what a reader needs to see in order to judge a study; they do not tell you how to run one. That distinction is the source of most misunderstanding about them, and it is worth holding onto through everything below.
ARRIVE is the acronym for Animal Research: Reporting of In Vivo Experiments. The guidelines are maintained by the NC3Rs — the UK National Centre for the Replacement, Refinement and Reduction of Animals in Research — and are hosted at arriveguidelines.org, where each item has its own explanatory page.
The original guidelines were published by Carol Kilkenny, William J. Browne, Innes C. Cuthill, Michael Emerson and Douglas G. Altman in PLOS Biology in 2010, under the title “Improving Bioscience Research Reporting: The ARRIVE Guidelines for Reporting Animal Research”. That version had 20 items. ARRIVE 2.0 was published in PLOS Biology in 2020 by Nathalie Percie du Sert and colleagues, with a companion Explanation and Elaboration article giving the rationale and worked examples for every item.
The motivating evidence was a survey of 271 published animal studies, co-funded by the NC3Rs and NIH/OLAW and reported by Kilkenny and colleagues in PLOS ONE in 2009. In that sample, only 12% reported random allocation of animals to experimental groups, and among the 35 papers that used qualitative outcome scores only 14% reported blinding. Only 59% stated the hypothesis or objective of the study together with the number and characteristics of the animals used. Those three numbers are the reason the checklist exists.
The Essential 10 are items 1 to 10 of ARRIVE 2.0: study design, sample size, inclusion and exclusion criteria, randomisation, blinding/masking, outcome measures, statistical methods, experimental animals, experimental procedures, and results. ARRIVE 2.0 describes this group as information that is “the basic minimum to include in a manuscript, as without this information, reviewers and readers cannot confidently assess the reliability of the findings presented.”
Several of the Essential 10 are more demanding than their one-word titles suggest, and the exact wording matters.
Item 1, study design, asks for the groups being compared including controls, and — sub-item 1b — for “the experimental unit (e.g. a single animal, litter, or cage of animals)”. Naming the experimental unit is a single clause that most manuscripts omit and that determines whether the statistics are defensible; see pseudoreplication.
Item 2, sample size, asks you to “specify the exact number of experimental units allocated to each group, and the total number in each experiment” and to “also indicate the total number of animals used”. Sub-item 2b asks how the sample size was decided, with details of any a priori calculation “if done” — the phrase permits an honest statement that no calculation was performed, which is a far better answer than a retrospective one. On why retrospective calculations do not help, see post hoc power.
Item 3, inclusion and exclusion criteria, is among the least often satisfied items in the set and has three sub-items. Sub-item 3a asks for the criteria and whether they were set a priori, adding: “If no criteria were set, state this explicitly.” Sub-item 3b asks you to report any animals or data points not included in the analysis and why, adding: “If there were no exclusions, state so.” Sub-item 3c asks you to “report the exact value of n in each experimental group” for each analysis. Two of the three sub-items require a sentence even when the answer is nothing, which is why silence is non-compliance rather than a neutral absence.
Item 4, randomisation, asks whether randomisation was used and, if so, “the method used to generate the randomisation sequence”. Sub-item 4b asks for the strategy used to minimise confounders such as order of treatment, order of measurement, and cage or rack position — with the same explicit fallback: “If confounders were not controlled, state this explicitly.” Cage position is not a decorative detail; temperature, light and handling gradients within a rack are real sources of confounding.
Item 5, blinding/masking, asks you to “describe who was aware of the group allocation at the different stages of the experiment (during the allocation, the conduct of the experiment, the outcome assessment, and the data analysis)”. The item is written as four stages precisely because “the experimenter was blinded” is uninterpretable — blinded at which stage, and to what.
Item 6, outcome measures, asks for a clear definition of every outcome assessed and, for hypothesis-testing studies, identification of the primary outcome — the one that drove the sample size.
Item 7, statistical methods, asks for the methods used for each analysis “including software used”, and separately for how the assumptions of the statistical approach were assessed “and what was done if the assumptions were not met”. The second half of item 7 is what turns a vague methods paragraph into something checkable; see wrong statistical test.
Item 8, experimental animals, asks for species, strain and substrain, sex, age or developmental stage, and weight where relevant, plus provenance, health and immune status, genetic modification status, genotype, and any previous procedures. Substrain is named deliberately: C57BL/6J and C57BL/6N differ at the Nnt locus and behave differently in metabolic work, so “C57BL/6” alone does not satisfy item 8a.
Item 9, experimental procedures, asks for what was done, how, when and how often, where, and why — in enough detail for replication.
Item 10, results, asks for summary or descriptive statistics for each group “with a measure of variability where applicable (e.g. mean and SD, or median and range)”, and, where applicable, the effect size with a confidence interval.
The Recommended Set is items 11 to 21 of ARRIVE 2.0, described in the guidelines as the group that “adds context to the study described”: abstract (11), background (12), objectives (13), ethical statement (14), housing and husbandry (15), animal care and monitoring (16), interpretation and scientific implications (17), generalisability and translation (18), protocol registration (19), data access (20), and declaration of interests (21).
The label “recommended” causes a predictable error. These eleven items are not optional courtesies; they are the items that are usually enforced by other machinery — journal ethics policies cover item 14, data availability policies cover item 20, and competing-interests forms cover item 21. What ARRIVE 2.0 did was rank the 21 items by what a reader cannot do without, not by what a journal will chase you for.
Three of the Recommended Set carry more weight in preclinical work than their placement implies. Item 15, housing and husbandry, includes cage type, bedding, light cycle, temperature and number of cage companions — the variables that most often explain why a phenotype does not travel between laboratories. Item 18, generalisability and translation, asks how the findings relate to human biology, which is where a single-sex, single-strain, single-age study must state its limits rather than imply none; see mouse model relevance. Item 19, protocol registration, asks whether a protocol with the research question, design and analysis plan was prepared before the study and where it was registered — a question most in vivo papers still answer with silence.
ARRIVE 2010 presented 20 items as a flat list, all nominally of equal weight. ARRIVE 2.0 kept substantially the same content but did three things to it.
It prioritised. Splitting 21 items into the Essential 10 and a Recommended Set was the central change, made after a structured consultation with researchers, journal editors and statisticians. A flat 20-item list gives an editor no basis for deciding what to insist on; a ranked list does.
It rewrote items to be checkable. ARRIVE 2.0 phrases items as instructions with explicit fallbacks — “if no criteria were set, state this explicitly”, “if there were no exclusions, state so”, “if confounders were not controlled, state this explicitly”. A reader can score those. “Describe inclusion criteria” cannot be scored, because an absent criterion and an unreported criterion look identical on the page.
It moved the experimental unit to item 1. In the 2010 version, the concept was diffused across design and statistics. In ARRIVE 2.0 it is sub-item 1b, before sample size, because every downstream number depends on it.
The ARRIVE 2.0 abstract is unusually candid about why a revision was needed at all: “Despite considerable levels of endorsement by funders and journals over the years, adherence to the guidelines has been inconsistent, and the anticipated improvements in the quality of reporting in animal research publications have not been achieved.”
The experimental unit is the smallest division of the experimental material that can be independently allocated to a treatment. ARRIVE 2.0 asks for it in sub-item 1b, and the answer propagates through items 2, 3c, 7 and 10.
If a treatment is applied to the drinking water of a cage of five mice, the cage is the experimental unit and n is the number of cages, not 25 mice across five cages. If a dam is dosed and pups are measured, the litter is the experimental unit for anything transmitted through the dam. If each animal receives its own injection, the animal is the unit — but if three tissue sections are quantified per animal, those sections are technical replicates and n remains the number of animals.
Getting this wrong inflates n by the number of measurements per animal, which shrinks standard errors and manufactures significance. It is a common statistical defect in in vivo manuscripts and the reason reviewers object that n is not independent. It is also invisible unless item 1b is stated, which is exactly why ARRIVE asks for it.
Detecting non-compliance is a set of specific arithmetic and cross-reference checks, not a subjective reading. The following seven checks find most of it, and each can be run against a submitted manuscript in a few minutes.
Reconcile the animal arithmetic. Sum the group ns given in the methods, compare with the total number of animals stated under item 2a, and compare both with the ethics-approval numbers if given. A discrepancy that is not explained by a stated exclusion is an item 3b failure. In practice this check finds attrition that nobody hid deliberately and nobody reported either.
Cross-check n between the figure legends and the methods. Legends reading “n = 3” against a methods section describing six animals per group indicates either unreported exclusions or a change in the experimental unit between the bench and the figure. Item 3c exists to close this gap: it asks for the exact n “for each analysis”, not once for the study.
Search for the words “randomly”, “randomised” and “randomized” and read what follows. “Animals were randomly divided into four groups” names no method and therefore fails item 4a. Compliant text names the mechanism: a random number generator, a randomisation sequence prepared by someone not involved in dosing, block randomisation stratified by cage or by baseline weight.
Search for “blind” and check whether a stage is named. “The experimenter was blinded” fails item 5, which is written as four stages. A compliant sentence identifies who was unaware of allocation during allocation, conduct, outcome assessment and analysis, and it is entirely acceptable to state that blinding was not possible at a given stage.
Check whether an exclusion statement exists at all. Absence is the failure mode. Under items 3a and 3b, a manuscript with no exclusions must say so; the checkable question is whether the sentence is present, not whether animals were excluded.
Check whether the primary outcome is identified and whether it matches the sample size justification. Item 6b and item 2b are a matched pair. A manuscript reporting eleven outcomes with no primary named, and a power calculation based on an outcome that does not appear in the results, is a common and consequential mismatch — see multiple comparisons.
Check whether the statistical unit in the analysis matches the stated experimental unit. A t-test across 30 sections from 6 animals is an item 1b and item 7a failure simultaneously. Image-based endpoints are where this most often hides; see image quantification reporting.
The NC3Rs publishes a compliance questionnaire built from operationalised questions on the Essential 10, intended for journal staff, editors and peer reviewers rather than for authors. PerfectPaper runs the equivalent checks on a submitted manuscript before a journal does, reporting which Essential 10 items have no supporting sentence and quoting the sentence that would need to change; see in vivo rigour and preclinical peer review.
Reviewer comments on ARRIVE failures are recognisable and repetitive. The phrasings below are the common ones in preclinical reports, and each maps to a numbered item.
“The manuscript does not state the experimental unit; if the cage was the unit of treatment, the reported n is inflated.” (item 1b)
“Please state how the sample size was determined. If no a priori power calculation was performed, say so.” (item 2b)
“No information is given on whether animals or data points were excluded. Please state exclusion criteria and report all exclusions with reasons, or confirm that there were none.” (items 3a, 3b)
“Animals are said to have been randomly assigned, but the method of sequence generation is not described.” (item 4a)
“It is unclear who was blinded and at which stage. Please specify whether outcome assessment and analysis were performed blind to allocation.” (item 5)
“A primary outcome is not identified, and it is therefore not possible to determine which comparisons were prespecified.” (item 6b)
“The statistical unit does not appear to match the experimental unit. Sections from the same animal are not independent observations.” (items 1b, 7a)
“The strain designation is incomplete. Please specify substrain and supplier.” (item 8a)
“Housing conditions, light cycle and group size are not reported, which limits assessment of reproducibility.” (item 15)
Two features of this list are worth noting. First, none of these comments is about the biology; they are all about whether the reader can evaluate the biology, which is why they arrive from statistical reviewers as often as from subject specialists. Second, most are unanswerable after the fact — an experiment run without randomisation cannot be retrospectively randomised, and the only honest response is to state the limitation. On writing that response, see how to write a response to reviewers.
The honest answer is: much less than expected, and the evidence for that is unusually good.
Baker, Lidster, Sottomayor and Amor examined journals that had endorsed ARRIVE and published “Two Years Later: Journals Are Not Yet Enforcing the ARRIVE Guidelines on Reporting Standards for Pre-Clinical Animal Studies” in PLOS Biology in 2014. Their finding is in the title: two years after endorsement, reporting of the core rigour items — randomisation, blinding, sample size justification — remained very low in the endorsing journals they sampled. Endorsement without enforcement moved almost nothing.
The strongest evidence is the IICARus trial — “A randomised controlled trial of an Intervention to Improve Compliance with the ARRIVE guidelines”, by Kaitlyn Hair, Malcolm R. Macleod and Emily S. Sena with the IICARus Collaboration, in Research Integrity and Peer Review in 2019. The trial randomised manuscripts submitted to PLOS ONE between an intervention arm, in which the journal asked authors to complete an ARRIVE checklist at submission, and standard editorial practice. Compliance with the checklist items was then scored in the manuscripts that went on to be published. Full compliance was essentially absent in both arms, overall compliance was low in both, and the intervention produced no broad improvement. The authors concluded that altering the editorial process to request a completed ARRIVE checklist is not enough to improve compliance, and that other approaches — such as more stringent editorial policies or a targeted approach on key quality items — may be required.
That result should change how an author uses ARRIVE. A checklist ticked at submission demonstrably does not produce compliant text, because ticking a box and writing the sentence are different acts and the trial separated them cleanly. What produces compliant text is checking the manuscript against the items — the seven checks in the section above — and adding the missing sentences.
The mechanism behind that failure is worth naming, because it tells you where to spend effort. A checklist question asks whether something is reported; the manuscript is what does or does not report it. An author who ran the study, knows the animals were allocated by a random number generator, and reads the question “did you randomise?” answers yes truthfully — and the methods section still says only “animals were randomly divided into four groups”, which fails item 4a because the method is not named. The checklist and the text drift apart in exactly the cases where the author’s memory is doing the work the sentence should do. Checking the text, item by item, is the only step that closes that gap, and it is a step that can be done by anyone reading the manuscript, including a coauthor who was not at the bench.
ARRIVE governs reporting of in vivo experiments in the life sciences. Several adjacent things fall outside it.
ARRIVE is not a design or ethics standard. The guidelines ask you to report whether randomisation was used; they do not require it, and they do not grant or substitute for ethical approval. Item 14 asks for an ethical statement, which is a reporting requirement about approvals you obtained elsewhere.
ARRIVE does not cover planning. The PREPARE guidelines, published by Norecopa in 2018, address planning animal experiments and are the upstream companion to ARRIVE’s downstream reporting focus.
ARRIVE does not cover in vitro work, clinical trials, or reviews. CONSORT covers randomised clinical trials, STROBE observational human studies, PRISMA systematic reviews, and ARRIVE animal experiments; MDAR and journal-specific reporting summaries sit across several of these. A study with a mouse cohort and a cell-culture arm needs ARRIVE for the former and something else for the latter.
ARRIVE does not settle whether a finding is interesting. Full compliance with the Essential 10 is compatible with a well-reported study nobody needs, and non-compliance is compatible with an important result. Reviewers who conflate the two produce a novelty objection when the actual objection was rigour, or the reverse.
ARRIVE does not make claims proportionate. An impeccably reported single-model experiment can still be written up as though it established a mechanism in humans; that is a separate defect, checked against the evidence actually presented. See overclaiming.
The items that cannot be satisfied retrospectively are items 1b, 2b, 4a, 4b, 5 and 6b — experimental unit, sample size justification, randomisation method, confounder control, blinding stages, and primary outcome. Every one of them is a decision made before the first animal is dosed, and no amount of careful writing recovers a decision that was never made.
The practical consequence is that ARRIVE is most useful as a study-planning document that happens to be published as a reporting checklist. Writing the six clauses above into a protocol before the study starts takes perhaps an hour, satisfies item 19 if the protocol is registered, and removes the category of reviewer comment that has no good answer. Reading the checklist for the first time during manuscript preparation, which is when most researchers meet it, leaves those six items as limitations to be confessed rather than methods to be described.
Pseudoreplication · In vivo rigour · Reviewer says n is not independent · Mouse model relevance · Preclinical peer review
Checked before submission by in vivo rigour, which reports which of the ARRIVE Essential 10 have no supporting sentence in the manuscript and names the sentence that would need to change.
ARRIVE stands for Animal Research: Reporting of In Vivo Experiments. The acronym names a reporting checklist maintained by the NC3Rs, first published in PLOS Biology in 2010 and revised as ARRIVE 2.0 in 2020. The current version contains 21 items describing what an animal-research manuscript must report.
The ARRIVE 2.0 checklist is the 2020 revision of the ARRIVE guidelines, containing 21 items divided into an Essential 10 and a Recommended Set of 11. The Essential 10 are described by the guidelines as the basic minimum without which reviewers and readers cannot confidently assess the reliability of the findings presented.
The ARRIVE Essential 10 are items 1 to 10 of ARRIVE 2.0: study design, sample size, inclusion and exclusion criteria, randomisation, blinding or masking, outcome measures, statistical methods, experimental animals, experimental procedures, and results. Every manuscript describing animal research is expected to report all ten.
The ARRIVE guidelines are the reporting standard for in vivo animal research, maintained by the NC3Rs — the UK National Centre for the Replacement, Refinement and Reduction of Animals in Research. Carol Kilkenny, William Browne, Innes Cuthill, Michael Emerson and Douglas Altman published the original in PLOS Biology in 2010; Nathalie Percie du Sert and colleagues published ARRIVE 2.0 there in 2020.
ARRIVE is endorsed rather than universally mandated, and endorsement has proved weak on its own. The IICARus randomised trial at PLOS ONE, reported in 2019, found that asking authors to complete an ARRIVE checklist at submission did not produce broadly compliant manuscripts, and concluded that other approaches, such as more stringent editorial policies, may be required.
ARRIVE 2010 listed 20 items of nominally equal weight. ARRIVE 2.0, published in 2020, lists 21 items ranked into an Essential 10 and a Recommended Set, rewrites items with explicit fallbacks such as “if there were no exclusions, state so”, and moves the experimental unit to sub-item 1b.
The ARRIVE guidelines require the experimental unit and group comparisons, exact group and total animal numbers, inclusion and exclusion criteria with all exclusions, the randomisation method, who was blinded at each stage, defined and primary outcomes, statistical methods and assumption checks, full animal characteristics, replicable procedures, and summary statistics with variability.
Last updated September 9, 2026
Upload your paper and receive structured, sourced feedback before you submit.
Connection lost
Reconnecting…
Something went wrong on our end
Attempting to reconnect