Skip to content

SOLUTIONS

Desk screening at scale

The fastest way to protect reviewer capacity is to send fewer papers to review. What can be checked before an invitation goes out, and what cannot.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

Desk screening at scale

Desk screening is the set of checks a journal applies to a submission before any reviewer is invited. At scale it is the one intervention that changes the arithmetic of reviewer capacity, because it acts on how many manuscripts reach review rather than on how many people can be persuaded to accept an invitation.

Every manuscript that reaches a reviewer with a missing ethics statement, an unreported registration number, numbers that do not reconcile between the abstract and the table, or a conclusion the design cannot support consumes an unpaid evening on a problem that needed no reviewer to identify.

This is the part that is most reliably checkable before an invitation goes out, which makes desk screening the only lever that reduces demand rather than chasing supply.

This page covers what a screen can determine, what it must not, the formal and reporting-guideline passes item by item, how numerical consistency is actually checked, the sequence that works, what the screen returns and to whom, a worked screening week, how integrity screening differs, the four ways a screen at scale goes wrong, what authors should be told, what editors say when the screen is wrong, what to do when a journal has no staff to run one, and how to measure whether it worked.

The scale problem in numbers

Peer review labour is large, unpaid and measured. Aczel, Szaszi and Holcombe (Research Integrity and Peer Review 2021;6:14) estimated from publicly available data that reviewers worldwide contributed over 100 million hours to peer reviews in 2020, equivalent to more than 15,000 years of working time, and stated that their figures likely underestimate the true total because they capture only a portion of the world’s journals.

The relevant property of that number is not its size but its inelasticity. Reviewer hours do not expand when submissions do, because the supply is volunteers with day jobs, and the three levers an editor holds — who gets asked, how much each is asked for, and how much reaches review at all — act on different sides of the equation. Only the third reduces total demand, which is the argument set out in the peer review capacity problem.

A screen therefore has to be judged on demand, not on defect counts. A screen that finds twenty defects per manuscript and returns nothing has protected no reviewer hours. A screen that returns four manuscripts in a hundred before invitation, and hands the handling editor a one-page reconciliation note on nine more, has changed the workload of every reviewer downstream of it.

What can be checked before review

Desk screening can determine five things without a specialist reading the science, and each of the five is a presence-or-absence question rather than a judgement.

Formal completeness. Ethics approval naming an approving body, trial registration identifier, data availability statement naming a repository, competing interests, funding, author contributions, reporting checklist. Present or absent is a determinate question.

Reporting guideline items. Whether the manuscript addresses the items its study design requires — CONSORT, STROBE, ARRIVE, PRISMA — is checkable item by item without judging whether the answers are good.

Internal numerical consistency. Whether the n in the abstract matches the n in the flow diagram, whether percentages match their numerators, whether a table’s totals add, whether figures agree with the text describing them. Discrepancies here are common, they are mechanical, and they undermine reviewer confidence out of proportion to their importance.

Claim-evidence alignment. Whether each claim in the abstract traces to a specific result. Not whether the claim is right — whether it is supported by something in the paper.

Scope and fit. Whether the manuscript is the kind of paper the journal publishes, which is a common reason for desk rejection and the cheapest to decide.

Screening check What makes it determinate What it cannot tell the editor
Formal completeness Each required statement is present or absent, and the approving body, identifier or repository is either named or not Whether the approval was appropriate, or the registration prospective
Reporting-guideline items Each checklist item is addressed at a locatable place in the manuscript, or it is not Whether what is reported at that place is adequate
Internal numerical consistency Two numbers in the same manuscript either reconcile arithmetically or they do not Which of the two numbers is the correct one
Claim-evidence alignment Each abstract claim either points to a result in the paper or points to nothing Whether the result supports the claim strongly enough
Scope and fit The manuscript’s design and subject either resemble what the journal publishes or they do not Whether an unusual paper is the one worth making an exception for

The formal completeness pass, item by item

The formal completeness pass is the one screening category with published, quotable standards behind it, which is why it can be run identically on every submission and defended to any author who queries it.

The ICMJE Recommendations state that “The Methods section should include a statement indicating that the research was approved by an independent local, regional or national review body (e.g., ethics committee, institutional review board).” A statement that a study was “conducted ethically” without naming the approving body does not satisfy that sentence. For trials, the ICMJE recommends that all medical journal editors require “registration of clinical trials in a public trials registry at or before the time of first patient consent for enrollment as a condition of consideration for publication”, in a WHO ICTRP primary register or ClinicalTrials.gov. A screen can determine that an NCT number is present and findable in the manuscript body; it cannot determine that registration was prospective, which is why the finding goes to the editor as a question rather than out as a rejection — see trial registration and endpoints for the related failure the screen also cannot settle.

Data availability is the fastest-growing member of this category. The ICMJE asks that where “the data have been deposited in a public repository and/or are being used in a secondary analysis, authors should state at the end of the abstract the unique, persistent data set identifier, repository name and number.” In omics work an absent accession — GEO or ArrayExpress, SRA or ENA, dbGaP or EGA, PRIDE for proteomics — is a common reason a manuscript is returned before review, and “available on request” is a statement of intent rather than an accession. Ethics, consent and data sharing and data availability and accessions set out what each statement has to contain to pass.

Reporting-guideline screening without judging the answers

Reporting-guideline screening asks whether each item its design requires is addressed somewhere locatable, and it stops there. CONSORT 2025 comprises 30 items and a participant flow diagram, STROBE has 22 items, PRISMA 2020 has 27, and ARRIVE 2.0 has 21 arranged as an Essential 10 and a Recommended Set of 11. Those counts are the whole basis of the check: an item is addressed at a page and line, or it is not.

The distinction that keeps this honest is between reporting and quality. A trial with unconcealed allocation that describes its unconcealed allocation clearly satisfies the item and remains at high risk of bias. A screen that scores checklist compliance as a proxy for study quality has crossed into judgement it cannot make, and will reject careful work and pass careless work with equal confidence.

Two operational details decide whether the pass is worth running. The first is that a completed checklist uploaded with page numbers is not evidence the item is there — checklists are sometimes submitted with page numbers that do not contain the item, so the screen must resolve the page reference rather than accept it. The second is design-to-guideline matching: applying CONSORT to a single-arm phase I study produces a checklist full of not-applicable rows and no useful discipline, and applying the parallel-group standard to a cluster randomised trial misses the items the cluster extension adds. See the CONSORT checklist, the STROBE checklist, the PRISMA checklist and the ARRIVE guidelines for what each set actually asks for.

Internal numerical consistency, and what actually finds it

Internal numerical consistency is checkable by arithmetic, and the published prevalence of the failures it catches is high enough to justify running it on every submission rather than on a sample.

Nuijten and colleagues (Behavior Research Methods 2016;48:1205–1226) scanned 30,717 articles published between 1985 and 2013 in eight psychology journals with the statcheck procedure, recomputing each reported p-value from its own test statistic and degrees of freedom. Half of the papers using null-hypothesis significance testing contained at least one p-value inconsistent with its test statistic and degrees of freedom, and one in eight contained a grossly inconsistent p-value that may have affected the statistical conclusion. Brown and Heathers (Social Psychological and Personality Science 2017;8:363–369) applied the GRIM test — checking whether a reported mean is arithmetically attainable from an integer-scale measure at the reported sample size — to 260 recent articles, of which 71 were amenable to the technique; 36 of those 71, or 50.7%, contained at least one mean inconsistent with the reported sample size and scale characteristics, and 16 contained several.

Neither technique needs a statistician, and neither is field-specific in its logic. The manuscript-level version of the same discipline is denominator reconciliation: abstract N, Table 1 column totals, the flow diagram and the analysed N in the primary model are four numbers that should agree, and a paper reporting 1,204 enrolled, 1,180 in Table 1 and 1,061 in the adjusted model has 143 people whose exit is unexplained. A screen reports the discrepancy and the four locations; it does not decide which number is right, which is precisely the boundary that keeps it upstream of statistical review capacity.

Claim-evidence alignment in the abstract

Claim-evidence alignment asks, for each claim in the abstract, whether a specific result in the manuscript is the thing that claim rests on. Not whether the claim is right — whether it is supported by something in the paper.

The check runs as a mapping exercise. Every declarative sentence in the abstract’s results and conclusion gets a pointer to a table cell, a figure panel or a numbered result, and the screen reports the sentences with no pointer. Three shapes fail routinely: a conclusion naming an outcome the paper measured only as a secondary endpoint, an effect stated as an adjective (“significantly improved”) where the paper reports an estimate the abstract omits, and a causal verb attached to a design that supports an association. The last of these is what an overclaim check exists to catch, and it is the failure most likely to survive review if it is not caught before one.

The reason to run this pass at screening rather than leaving it to reviewers is that the abstract is the object on which the triaging editor decides fit and significance in the first place. An abstract that hides the design or buries the effect converts a strong study into a weak-looking one at the only moment the decision is being made, and the editor never sees the difference — how to write an abstract covers the six moves an abstract makes and the order to draft them in.

Scope and fit, decided in minutes

Scope and fit is a common reason for desk rejection and the one a screen can surface fastest, because the comparison is between the manuscript’s shape and the journal’s recent contents rather than between the manuscript and any standard.

Fit is decided on shape as well as topic — descriptive versus mechanistic, basic versus clinical, single-centre versus multi-centre, a methods paper versus an application of a method — so a manuscript can sit squarely inside a journal’s stated aims and scope and still be the wrong kind of paper for the people who read it. That is a judgement, and the screen’s contribution to it is narrow and useful: name the design, name the subject, name the closest comparable articles the journal published, and hand all three to the editor in one line. The decision stays with the editor; the ten minutes of looking does not.

Returning a fit rejection quickly is a service, and the ICMJE says so directly: editors “should endeavor to reject the manuscript as soon as possible to allow authors to submit to a different journal.” A journal that takes eleven weeks to send the same letter has done the author actual harm. The author-side view of this decision, including how to read the letter that carries it, is desk rejected without review; the same judgement raised later by a reviewer is an out-of-scope objection, and the decision that prevents both is journal selection made before the paper is written.

What cannot, and should not, be checked before review

Desk screening cannot determine three things, and a journal that asks it to has automated the part of review that peer review exists to provide.

Significance. Whether the advance matters to the field is an editorial judgement informed by reviewers, and automating it would be automating the part of review that is actually contested.

Novelty. Requires knowing the literature as a specialist reads it, including unpublished work and what a field currently believes.

Whether the interpretation is correct. A methods section can be complete and the conclusion still wrong. That is exactly what reviewers are for.

The line worth holding: screening establishes whether a manuscript is reviewable, not whether it is right. Confusing the two produces either a screen that rejects good papers or a review process that has quietly outsourced its judgement.

There is a fourth item that belongs on this list and is often mistaken for a determinate check: whether a violated statistical assumption invalidates a conclusion. Whether hazards are proportional, whether a covariate is a mediator rather than a confounder, whether the adjustment set is causally defensible — these look mechanical and are not. A screen can report that proportional hazards is never mentioned. It cannot report that the assumption fails, and a screen that claims to is making the same category error as one that scores novelty.

Sequencing that works

Desk screening at scale works in a fixed order, because the categories differ in how cheap they are to run and in how unambiguous their failures are, and running them out of order wastes editorial attention on manuscripts that were going to be returned anyway.

Formal completeness first, because it is cheapest and its failures are unambiguous — return for correction rather than rejecting, since these are usually fixable in a day and the paper may be excellent.

Scope second, because it is the largest category and returning quickly is a service to the author.

Consistency and claim-evidence third, as information for the handling editor rather than as an automatic decision. A manuscript with several unreconciled numbers is not thereby rejectable, but the editor should know before choosing reviewers, and the reviewers should not be the ones to discover it.

Two properties of the order matter more than the order itself. The first is that only the first stage produces an automatic outcome; stages two and three produce a recommendation and a note. The second is that a manuscript failing stage one is never assessed on stages two and three, because an incomplete submission is an incomplete object and reporting consistency findings against it generates author work that a corrected file would have made unnecessary.

What the screen returns, and to whom

A desk screen at scale produces three kinds of output with three different destinations, and collapsing them into a single decision is the most common design error.

A return to the author. Formal-requirement failures only, itemised, each naming the missing statement and where it belongs. “The trial registration identifier is not present in the abstract or Methods” is a complete item; “please check your submission against our requirements” is not, and generates a resubmission that fails a different item.

A note to the handling editor. Consistency discrepancies with all their locations, claim-evidence gaps, and the design-to-guideline match. The editor uses it to decide who to invite and what to ask them for — a paper whose numbers do not reconcile may need a statistical reader rather than a second subject specialist.

A decline, from the editor. Scope and significance decisions stay with a person, and the screen’s output is an input to them.

The property that makes the note usable rather than another reading task is anchoring: every finding names a location in the manuscript file, so the editor verifies it in seconds rather than re-deriving it. A finding that says “numbers are inconsistent” without saying which numbers, on which pages, is a second manuscript to read.

Worked example: a screening week at a mid-sized journal

Consider a journal receiving 100 submissions in a week, with one part-time managing editor and four handling editors, and a screen that runs the three stages above before anything reaches an editor’s queue.

Stage one returns the manuscripts missing a required statement. Suppose eleven are missing something itemisable — an ethics statement that names no approving body, three data availability statements reading “available on request”, a trial with its NCT number only in a supplementary file, several missing author contribution statements. All eleven go back the same day with the item named. Nine return corrected within a week, and the two that do not were going to be returned at some later, more expensive stage.

Stage two hands the four handling editors a design-and-subject line for the remaining 89, with the closest comparable articles the journal published. Fit declines are made by a person in minutes rather than in an afternoon of looking things up.

Stage three produces reconciliation notes. Suppose nineteen manuscripts carry at least one numerical discrepancy — the abstract N and the flow diagram disagreeing, a percentage that does not match its numerator, a table total that does not add. None of the nineteen is rejected. Each editor sees the discrepancy with its locations before choosing reviewers, and can ask the author to reconcile the numbers before the invitations go out.

The reviewer-facing result is the point. The reviewers invited for those nineteen manuscripts do not open their reports with a list of arithmetic, because the arithmetic was settled before they were asked. That is the demand reduction, and it is invisible in any count of manuscripts rejected.

Integrity screening is a separate lane

Integrity screening runs alongside desk screening and follows different rules, because its output is a query rather than a decision and its failure mode is accusation rather than delay.

The plagiarism arm is largely automated. Crossref Similarity Check, powered by iThenticate from Turnitin, compares a submission against a corpus Crossref describes as over 78 million full-text scholarly content items, and returns a similarity report to the editor. High overlap is not itself misconduct — methods sections repeat, and the report frequently flags the authors’ own earlier work — which is why the letter arrives as a question. At publisher scale the pattern-detection layer is shared: the STM Association runs an Integrity Hub through which participating publishers screen submitted manuscripts for research-integrity signals, including patterns associated with paper mills, using a growing suite of detection capabilities that publishers may enable at their discretion. The Hub does not publish a participant count or a screened-manuscript total, so treat any circulating figure as unsourced.

The operational rule that keeps this lane separate from the completeness lane: an integrity signal names something specific and asks the author a question, and it never travels in the same letter as a request to add a funding statement. Bundling the two teaches authors to treat a serious query as paperwork.

Where a screen at scale goes wrong

A desk screen fails in four identifiable ways, and three of them make the screen look like it is working.

Checklist theatre. The screen verifies that a checklist was uploaded rather than that the items are present at the pages the checklist claims. Compliance rises, reporting does not, and the reviewers still find the missing items.

Scope creep into judgement. A screen that begins by reporting that no multiplicity correction is named ends by scoring whether the correction was the right one. The first is determinate; the second is a reviewer’s judgement, and a screen that makes it will be wrong in ways nobody catches because its output arrives with the authority of a check.

Drift between submissions. A screen whose behaviour depends on how a question was phrased on a given day is not a screen, because two manuscripts with the same defect get different outcomes and the journal cannot defend either. Fixing the review specification and the output format in the application, rather than in a prompt written per submission, is what makes the check the same for every submission.

Rejecting on stage-three findings. Consistency and claim-evidence findings are information, not verdicts. A journal that starts declining manuscripts on unreconciled denominators will decline careful papers with a typo in a table and pass careless papers whose numbers happen to agree.

What authors should be told

Publish what you screen for. Authors who know the checklist meet it, which reduces the return rate and the work on both sides. A journal that screens silently and returns manuscripts without saying which item failed generates a resubmission that fails a different item.

The ICMJE is explicit that this is an editorial obligation rather than a courtesy: “Journals should publish a clear, transparent description of their peer-review process for all types of manuscripts.” A screening stage that returns manuscripts before review is part of that process, and describing it takes a page in the instructions to authors.

What that page should contain is narrow and specific: the list of statements checked for presence, the reporting guideline expected for each design and whether a completed checklist with page numbers is required at submission, the fact that numerical discrepancies are reported to the handling editor rather than treated as grounds for rejection, and the turnaround the journal aims for on a formal return. Naming the turnaround is the part most journals omit and the part authors most want, because a return that arrives in two days is a service and the identical return at six weeks is a loss.

What editors say when the screen is wrong

Editors do not complain that a screen is inaccurate. They complain about four specific things, and each maps to a design decision above.

“The report tells me the numbers are inconsistent but not which numbers.” Anchoring is missing. “It flagged the same non-issue on every paper in this specialty” — a design-to-guideline mismatch, usually CONSORT applied to a non-randomised design. “It told the author to add a data availability statement when the paper has an accession in the Methods” — the check resolved a required location rather than the manuscript. “The author replied asking what specifically was missing” — the return was a category, not an item.

Authors’ complaints are narrower and more consistent: a return that names no item, a second return for something the first return did not mention, and a screening standard that appears nowhere in the instructions to authors. All three are the same defect from the other side, and all three are fixed by publishing the list and itemising the return.

When a journal cannot run a screen at all

Most journals have no screening staff, which makes the arrangements described above unavailable rather than merely unfunded, and there are three partial substitutes that work at that size.

Move the check to submission. The submission system already asks for uploads. Making the ethics statement, registration identifier and data availability statement required text fields the author has to complete, rather than optional attachments, converts the most common return category into something the author cannot omit. This is one configuration change, made once.

Screen only the categories with published standards. A one-person editorial office cannot run five checks on every manuscript. It can run formal completeness on all of them, which is the category with the highest return rate and the shortest decision time, and leave consistency to the reviewers it invites.

State the standard even if you cannot enforce it. Publishing the list raises first-pass compliance whether or not anyone checks, because most authors are not withholding information; they do not know what is expected. A published expectation with no screen behind it outperforms a screen with no published expectation.

Where the field-specific conventions matter more than the general ones — flow cytometry gating, variant nomenclature, batch structure in an omics dataset — the checkable list is different for each, which is the case for custom reviewer agents written against a specialty rather than a general standard.

Measuring whether the screen works

A desk screen is measured on five numbers, and none of them is the number of manuscripts it rejected.

The first is the return rate for formal items, which should fall over the first year as the published checklist takes effect. The second is the repeat-return rate — manuscripts returned twice for different items — which measures whether returns are itemised. The third is invitations sent per completed review, the capacity number that the whole exercise exists to move. The fourth is time from submission to first decision, and specifically the share of that time spent before the review clock starts. The fifth is qualitative and the most informative: the proportion of reviewer reports whose first paragraph is a list of missing reporting items.

Set a baseline before the screen starts, because none of these is interpretable as a level. A 9% formal-return rate is neither good nor bad; a 9% rate that was 17% eighteen months ago is a working checklist.

Journals do not uniformly publish any of these figures, and the desk-rejection percentages in wide circulation are usually quoted without a stated method or denominator, so a journal cannot benchmark itself against a published distribution. It can only benchmark against its own previous year, which is a reason to record the baseline rather than a reason to skip the measurement.

What desk screening is not for

Desk screening is not a substitute for peer review, not a quality score, and not an integrity verdict.

Screening does not assess evidence, so a screened manuscript carries no assessment of its science and a returned manuscript carries no criticism of it. Screening does not rank submissions, so a manuscript that passes every check is not thereby a better paper than one returned for a missing funding statement. Screening does not adjudicate integrity: a similarity report is a question for the author, and treating it as a finding is how journals generate the complaints they later have to withdraw.

The comparison worth stating plainly: where a journal needs a judgement about whether a claim should change what a field believes, it needs a reviewer, and no amount of screening reduces that need. Where a journal needs to know whether the manuscript in front of it is complete, internally consistent and the kind of paper it publishes, it needs a screen, and asking a reviewer for that answer is what produces the declines described in why it is so hard to find peer reviewers.

What this looks like in practice

PerfectPaper produces structured, anchored findings against a fixed output contract — the review specification and format are set by the application rather than by a prompt, so results are consistent across submissions rather than varying with how a question was phrased. For a journal, the relevant properties are that the findings are anchored to locations in the manuscript, that the check is the same for every submission, and that the journal holds the manuscript under its own agreement with the author rather than a reviewer holding borrowed custody — see can reviewers use AI on manuscripts for why that distinction is the load-bearing one.

The boundary is the same one this page draws throughout. PerfectPaper reports whether a required statement is present, whether the reporting-guideline items for the stated design are addressed at locatable places, whether the denominators reconcile across abstract, tables, figures and model, and whether each abstract claim points to a result. PerfectPaper does not score significance, does not judge novelty, and does not decide whether an interpretation is correct, because those are the judgements a journal invites reviewers to make.

Editorial arrangements are discussed directly rather than through a sales process.

More in this cluster: the peer review capacity problem, why it is so hard to find peer reviewers, statistical review capacity.

Talk to us about journal use

Frequently asked questions

What should a journal check before sending a paper to review?

Formal completeness, reporting-guideline items for the study design, internal numerical consistency, whether abstract claims trace to results, and scope fit. All are determinate questions that do not require a specialist’s judgement.

How does desk screening work at scale?

In three stages: formal completeness first, returned to the author itemised; scope second, decided by an editor from a design-and-subject summary; consistency and claim-evidence third, reported to the handling editor as information before reviewers are chosen. Only the first stage produces an automatic outcome.

What can desk screening not check before review?

Significance, novelty and whether an interpretation is correct. Screening establishes whether a manuscript is reviewable, not whether it is right, and those three judgements remain editorial and peer judgements that should not be automated.

Does desk screening at scale reject good papers?

Not if failures return manuscripts for correction rather than rejecting them. Most formal failures are fixable within a day, and the paper may be excellent — the point is to fix them before a reviewer spends an evening listing them.

How much reviewer time does desk screening at scale save?

The saving is concentrated in the manuscripts that never reach review and in the reviews that no longer open with a list of missing reporting items. Aczel and colleagues estimated over 100 million reviewer hours worldwide in 2020, which is the pool the saving comes out of.

Should a journal publish what it screens for before review?

Yes. A published screening checklist raises first-pass compliance and reduces returns. Silent screening produces resubmissions that fail a different item, and the ICMJE asks journals to “publish a clear, transparent description of their peer-review process for all types of manuscripts.”

Can a small journal run desk screening without editorial staff?

Yes, in a reduced form. Make the ethics, registration and data availability statements required fields in the submission system, run formal completeness only, and publish the standard whether or not anyone checks it — a published expectation with no screen behind it outperforms a screen with no published expectation.

Last updated September 10, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.