Skip to content

SOLUTIONS

The peer review capacity problem

Submissions grow, the invited reviewer pool does not, and decline rates rise. The levers editors actually control, and where computational checking legitimately fits.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

The peer review capacity problem

Peer review is a volunteer system carrying a workload that grows every year, allocated by editors whose main tool is asking people they can see. The result is predictable: a small group is over-asked and saturates, a large capable pool is never asked, decline rates climb, and time to first decision follows.

None of that is fixed by working harder. It is a structural problem with three levers — who gets asked, how much each reviewer is asked for, and how much reaches review at all — and only the third reduces total demand.

The pages

How much review work the system absorbs

Peer review absorbed over 100 million hours of researcher time worldwide in 2020, equivalent to more than 15 thousand years, on the estimate Aczel, Szaszi and Holcombe published in Research Integrity and Peer Review (2021;6:14). Their paper states the figure is very likely an under-estimate, because it reflects only a portion of the total number of journals worldwide.

That total is growing faster than the people available to absorb it. Hanson, Gómez Barreiro, Crosetto and Brockington reported in Quantitative Science Studies (2024;5(4):823–843) that articles indexed in Scopus and Web of Science in 2022 stood approximately 47% above the 2016 total, growth “which has outpaced the limited growth - if any - in the number of practising scientists.” Every one of those articles sits on top of two or three completed reviews and a larger number of invitations that produced nothing.

The supply side was already lopsided before that growth arrived. Kovanis, Porcher, Ravaud and Trinquart estimated in PLOS ONE (2016;11(11):e0166387) that 63.4 million hours went into biomedical journal peer review in 2015, of which 18.9 million hours — just under 30% — came from the top 5% of contributing reviewers. In the same analysis, 20% of researchers performed 69% to 94% of the reviews depending on the modelling scenario, and among researchers contributing to review at all, 70% dedicated 1% or less of their research work-time to it while 5% dedicated 13% or more.

Those three figures describe one system that is sustainable in aggregate and unsustainable per person. Kovanis and colleagues concluded exactly that: the volume is carryable, and the distribution of who carries it is not.

The three levers editors control

Editors control three levers over review capacity, and they act on different quantities. Targeting changes who is asked. Scoping changes how much each accepted invitation demands. Screening changes how many manuscripts generate an invitation at all, and it is the only one of the three that reduces total demand rather than redistributing it.

Targeting moves the acceptance rate without changing the workload. Invitation lists assembled from corresponding authors, previous reviewers and board members select for seniority and visibility, which is the mechanism that concentrates the effort Kovanis measured. Inviting from middle authorships and supporting credited co-review widens the frame the list is drawn from. Why it is so hard to find peer reviewers sets out what changes those numbers.

Scoping moves both acceptance and quality. A request to assess the statistics, or the clinical relevance, is answerable in a defined amount of work; a request to assess a paper is not. Two focused reviews frequently beat one exhaustive review, and the focused request is the easier one to accept.

Screening moves demand itself. A manuscript returned before invitation consumes editorial minutes rather than reviewer evenings, and it removes the entire invitation sequence that manuscript would have generated. Desk screening at scale covers what is determinate enough to screen on and what is not.

The order matters. Targeting and scoping compete for a fixed pool; screening changes the size of the demand the pool is being asked to meet. An editorial board that reaches for recruitment first is optimising the harder lever.

Why agreement rates fell even where reviewers were not asked more often

Reviewer agreement rates fell substantially over 13 years at four of six ecology and evolution journals studied by Fox, Albert and Vines in Research Integrity and Peer Review (2017;2:3), with an “average decline from 56% of review invitations generating a review in 2003 to just 37% in 2015.” A decline of that size is hard to dismiss as rhetoric.

The interesting part is the explanation that failed. Reviewer fatigue predicts that agreement fell because individuals were asked more often, and the first half of that prediction holds: the likelihood that an invitee agrees declines significantly with the number of invitations they receive in a year. The second half does not. Fox and colleagues found that “the average number of invitations being sent to prospective reviewers and the proportion of individuals being invited more than once per year has not changed much over these 13 years, despite substantial increases in the total number of review invitations being sent by these journals—the reviewer base has expanded concomitant with this growth in review requests.” Their conclusion is that “reviewer fatigue is not likely the primary explanation for this decline.”

Those journals had already done the recruitment work. The pool grew in step with the requests, per-reviewer load stayed roughly flat, and agreement fell anyway. An editor reading that should draw a narrow conclusion: widening the invitation list is necessary and is not sufficient, and a capacity plan that consists only of finding more names is planning for the lever that those six journals already pulled.

This does not contradict the concentration finding, because the two studies measure different quantities. Kovanis measures the share of total review effort carried by the most active reviewers across the biomedical literature. Fox measures invitations per invitee at six named journals through 2015. A system can widen its invited pool at individual journals and remain globally concentrated, because a reviewer active at eight journals is counted once by each of them.

The invitation arithmetic

An editor sends well over two invitations for every completed review that comes back, and each invitation that fails spends calendar time before the review clock starts. Albert, Gow, Cobra and Vines reported the per-manuscript numbers for Molecular Ecology in Research Integrity and Peer Review (2016;1:14): a mean of 5.68 reviewers invited per manuscript in 2009 producing 2.69 completed reviews, rising to 6.46 invited in 2015 producing 2.82 completed. The proportion of requests that led to a review moved from 0.47 (95% CI 0.43 to 0.52) in 2009 to 0.44 (95% CI 0.40 to 0.48) in 2015 — a change the authors report as not statistically significant for Molecular Ecology, though across all five journals they conclude that agreement rates have “probably declined slightly but not to the extent suggested by the anecdotal and rhetorical evidence.”

Convert those means into elapsed time and the capacity problem becomes a scheduling problem. Suppose a journal allows an invitee five days to answer before re-inviting. The 3.64 invitations per manuscript that produced no review in 2015 — the gap between 6.46 sent and 2.82 completed — are not simultaneous: they arrive in rounds, and three sequential rounds of unanswered invitations spend fifteen days before the first accepted reviewer has read a word. Nothing in that fifteen days is reviewer effort, editorial judgement, or author fault. It is queueing.

Two implications follow directly. First, shortening the invitation response window and pre-loading alternates does more for time to first decision than any request that reviewers work faster, because the review itself is the part with a deadline attached. Second, every manuscript removed before the invitation sequence begins removes the whole sequence — the six invitations, the three and a half that go nowhere, and the queueing days — which is why screening dominates recruitment arithmetically and not just rhetorically. A manuscript desk rejected without review within a week is a service to the author and a return of capacity to the pool.

Where computational checking legitimately fits

Computational checking covers the determinate layer of a manuscript and stops at the contested one. The honest boundary is between checking and judging.

Checkable without a specialist: whether reporting-guideline items are addressed, whether numbers reconcile across abstract, tables and figures, whether an ethics or registration statement exists, whether each abstract claim traces to a result, whether the statistical approach matches the stated design.

Not checkable, and should not be: whether the advance matters, whether the interpretation is right, whether a field should change its mind. This is the contested part, it is what reviewers are for, and automating it would hollow out the process while appearing to help.

Journals that get this wrong in one direction automate judgement and lose the thing peer review provides. Journals that get it wrong in the other direction keep asking unpaid specialists to spend evenings listing missing checklist items, and then wonder why invitations are declined.

The determinate list is longer than it looks, because each guideline supplies its own items. A randomised trial is checkable against the 30 items of CONSORT 2025; an observational study against STROBE; a systematic review against PRISMA; an animal study against ARRIVE; an interview study against COREQ. Whether an item is addressed is a determinate question. Whether the answer is good is not, and no screening step should pretend otherwise.

What a determinate check finds that a reviewer would otherwise write out

A determinate check finds omissions and inconsistencies, and those are the errors that actually dominate the published record. Nuijten, Hartgerink, van Assen, Epskamp and Wicherts ran the statcheck algorithm over psychology articles from 1985 to 2013 in Behavior Research Methods (2016;48:1205–1226) and reported that 49.6% of articles with null-hypothesis-significance-testing results contained at least one internally inconsistent p-value, and 12.9% contained at least one gross inconsistency — one large enough to change the statistical conclusion. Those are not sophisticated methodological failures. They are numbers that do not agree with each other, in half of a literature, found by an algorithm.

Denominators behave the same way. Take a trial manuscript whose abstract reports 248 participants, whose flow diagram randomises 240, whose Table 1 columns total 118 and 119, and whose primary model analyses 231. Four numbers, four different populations, no explanation for any of the gaps. CONSORT 2025 asks at item 26 for the numbers analysed and the numbers with available data at the outcome timepoint as separate quantities, and at item 21b for who is included in each analysis and in which group; STROBE item 13(a) asks for numbers at each stage and 13(b) for reasons for non-participation. A reviewer who opens with that list is doing arithmetic, not review.

The same holds for the formal apparatus: whether the data availability statement names a repository and an accession rather than promising availability on request, whether the trial registration identifier and the pre-specified primary endpoint match what the paper reports, whether a correction for multiplicity is named where many tests were run, and whether each conclusion sentence traces to a result rather than overstating what the design supports.

Every item on that list is one a reviewer can find. The question is whether an unpaid specialist should be the one who finds it, on an evening, in prose, after a three-week wait.

The confidentiality line

A reviewer uploading a manuscript to an external service is distributing material the author entrusted to them for one purpose, which most publishers prohibit. A journal running a check on a submission it holds under its own agreement with the author is a different act with a different consent basis.

The ICMJE states the reviewer side directly. Its Recommendations, section 3 on peer reviewers, hold that “Reviewers therefore should keep manuscripts and the information they contain strictly confidential”, that “Reviewers must maintain the confidentiality of the manuscript as outlined above, which may prohibit the uploading of the manuscript to software or other AI technologies where confidentiality cannot be assured”, and that “Reviewers must request permission from the journal prior to using AI technology to facilitate their review.” The operative condition is where confidentiality cannot be assured, and the assurance is a property of the arrangement, not of the reviewer’s intentions.

The same document declines to give journals a clean pass. Section 2, on journals, states that “Editors should be aware that using AI technology in the processing of manuscripts may violate confidentiality.” Custody is a better position than a reviewer’s borrowed custody; it is not an exemption. What converts it into a defensible one is specific: the journal states in its submission policy what checks it runs, the terms under which the manuscript is processed are contractual rather than assumed, and the manuscript does not leave that arrangement. Editors evaluating a supplier against those terms will find the questions worth asking in the institutional AI procurement checklist.

Editors need a policy that states both, because a journal with no policy is not preventing reviewer use — it is preventing disclosure of it.

What a journal AI policy has to state to be usable

A usable journal AI policy answers three questions in writing, and a policy that answers only the first fails in practice. Reviewers are already using these tools; a prohibition with no permitted alternative and no disclosure path suppresses reporting rather than use.

What is prohibited. Uploading the manuscript, or any substantial portion of it, to an external service. State it as an act, not as a technology, so the clause survives the next product launch.

What is permitted, and must be disclosed. Assistance with the reviewer’s own written text, without the manuscript present, is where most workable policies land, with disclosure required for anything that shaped the substance of the review. Can peer reviewers use AI on manuscripts? sets out why the line falls there.

What the journal itself runs, and on what basis. Name the checks performed on submissions, name the point in the workflow at which they run, and name the agreement under which the manuscript is held. This is the clause most policies omit, and it is the one that separates a journal exercising its own custody from a referee breaching someone else’s. Author-facing obligations are a separate question again, covered in journal AI policies for authors and referees.

The four asks worth designing for

Four asks recur, and none of them is a machine-written review.

They are: fewer manuscripts arriving that are not ready; reviewers who are not asked to do proofreading; a way to tell, before choosing reviewers, whether a paper has the reporting problems that will dominate the reviews; and a shorter time to first decision without asking anyone to work faster.

Each of those is a demand-side intervention. None requires replacing a reviewer.

They also decompose into different mechanisms. “Fewer manuscripts arriving that are not ready” is an author-facing publication of the screening criteria, and it works because most authors are not withholding anything — they do not know what is expected. “Reviewers not asked to do proofreading” is a scoping decision in the invitation letter. “A way to tell before choosing reviewers” is an editor-facing report rather than a decision rule, and its value is in reviewer selection: a manuscript whose problems are statistical wants a different second reviewer than one whose problems are in the model system, which is the triage question statistical review capacity addresses. “A shorter time to first decision” is the queueing arithmetic above, and it is won before the review clock starts.

Signals that a journal has a capacity problem rather than a recruitment problem

A journal can measure which problem it has, and the six numbers below are already in the submission system. Recruitment problems and capacity problems respond to different interventions, and the diagnosis is cheap.

Invitations per completed review, tracked yearly. A ratio drifting upward while the invited pool is stable is a targeting problem. A ratio stable while volume rises is a demand problem.

Agreement rate by invitation count within the year. Fox and colleagues found that the likelihood an invitee agrees declines significantly with the number of invitations they receive in a year, while the number sent per invitee barely moved. Plot the same curve locally: if it is flat, saturation is not what is driving your declines, and adding names will not lift the rate.

Share of completed reviews contributed by the top decile of reviewers. This is the local version of the Kovanis concentration figure, and it predicts how exposed the journal is to a handful of people retiring or stopping.

Days from submission to first invitation sent, separated from days from acceptance to review returned. These are two different queues with two different owners, and journals routinely report only their sum.

Proportion of first-round reviews whose opening paragraph is a list of missing reporting items. This is the clearest available measure of reviewer time spent on work no reviewer was needed for.

Return rate after desk screening, by failed item. A screen that returns the same item repeatedly is a screen the authors have not been told about.

What this does not fix

Computational checking does not add reviewers, does not shorten a review a reviewer has already agreed to write, and does not decide anything. It reduces how often a scarce specialist is spent on a determinate problem; it does not create specialists.

It also does not touch the contested layer at all. Whether a finding matters, whether the interpretation follows, whether a field should change its mind, whether an unusual result is fabrication or discovery — these are the questions peer review exists to answer, and a system that appears to answer them without a reviewer has removed the value while keeping the ceremony. The boundary stated above is not a modesty gesture; it is the design constraint.

Nor does it substitute for editorial policy. A journal that screens silently, invites the same twelve people, and publishes no expectations will have the same capacity problem with better instrumentation. The tooling changes what reaches a reviewer. The editor still decides who is asked, for what, and by when. Editors comparing options for the workflow layer itself will find that scope covered in peer review software.

PerfectPaper for journals

PerfectPaper checks a submitted manuscript against a fixed review contract — reporting-guideline items for the study design, numerical consistency across abstract, tables and figures, and whether each claim in the abstract traces to a result — and returns findings anchored to locations in the document, so the check is applied identically to every submission rather than varying with how a prompt was phrased.

Structured, anchored findings against a fixed output contract are what an editor receives; the specification and format are set by the application, not by whoever typed the request. The journal holds the manuscript under its own author agreement. Editorial arrangements, including what journal-level access looks like, are discussed directly rather than through a sales process.

Field-specific conventions that a generalist check misses are handled by custom reviewer agents, which encode a subfield’s own reporting expectations rather than a generic checklist. The findings are addressed to the editor and the author, not to the reviewer: the point is that the reviewer never receives the list.

Talk to us about journal use

Frequently asked questions

Why is peer review getting harder to staff?

Submission volume grows while editors continue inviting from a visible senior group who saturate and decline. The pool that is actually invited grows much more slowly than the pool that could be. Articles indexed in Scopus and Web of Science in 2022 were approximately 47% above the 2016 total, against limited growth in the number of practising scientists.

Is there really a shortage of peer reviewers?

The measured decline is real and is smaller than the rhetoric. Across four of six ecology and evolution journals, the proportion of invitations generating a review fell from 56% in 2003 to 37% in 2015 (Fox, Albert and Vines, 2017), while at Molecular Ecology the proportion completed moved only from 0.47 to 0.44 between 2009 and 2015. What is unambiguous is the concentration: 20% of researchers performed 69% to 94% of reviews in the Kovanis 2016 estimate.

Can AI reduce peer review workload?

Computational checking can remove the mechanical layer — reporting completeness, numerical consistency, claim-evidence alignment — which is the part reviewers most resent. It cannot make the judgements about significance and correctness that peer review exists to provide. Nearly half of psychology articles using significance tests contained at least one inconsistent p-value, which is the shape of problem a determinate check finds.

Is it acceptable for a journal to run AI checks on submissions to save reviewer time?

The journal holds the manuscript under its own agreement with the author, so it can state in its submission policy what checks it performs. That differs from a reviewer uploading a manuscript entrusted to them, which most publishers prohibit. The ICMJE nonetheless warns editors that “using AI technology in the processing of manuscripts may violate confidentiality”, so the policy statement and the processing terms both have to exist.

What reduces time to first decision most?

Screening out manuscripts that are not reviewable, and inviting reviewers who will accept. Both act before the review clock starts, which is where most of the elapsed time actually accumulates. A journal averaging 6.46 invitations per manuscript for 2.82 completed reviews is spending days in invitation rounds that no reviewer is working during.

Should journals recruit early-career reviewers?

Yes — they are the part of the pool that is not already saturated, which is where the unused capacity sits given how concentrated review effort is. The barrier is that invitation lists are built from corresponding authors, previous reviewers and board members, which structurally excludes them. Recruitment alone is not enough: the six journals studied by Fox and colleagues expanded their reviewer base in step with demand, and four of the six saw agreement rates fall regardless.

What can an editor actually do about peer review capacity?

Three things, in order of leverage: screen more manuscripts out before any invitation is sent, scope each invitation to a defined question rather than the whole paper, and widen the frame invitation lists are drawn from beyond corresponding authors. Only the first reduces total demand; the other two redistribute it.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.