Skip to content

SOLUTIONS

AI reviewer for registry and claims data limitations

A prebuilt custom agent that checks whether SEER, NCDB and claims analyses acknowledge coding validity, missing treatment detail, and selection into the data source.

Built by NIH-funded cancer researchers

Affiliations

Built by researchers funded by leading cancer-prevention institutions

  • University of Utah
  • Huntsman Cancer Institute
  • National Cancer Institute
  • American Cancer Society

Current platform

A review workflow built around the decisions only the author can make

PerfectPaper now carries context from setup through research, revision, and export—without turning the paper into a generic writing prompt.

Prepare

Tell the review what the paper cannot

Author interview

PerfectPaper asks targeted questions about design decisions and fixed constraints before review, then carries your answers into the critique.

Journal-aware setup

Search the journal catalogue, choose up to three targets, and compare compatible open-access journals before the review starts.

Your own review panel

Brief up to three custom reviewers, declare ground truths, attach instructions, and choose standard or deep-research depth with specific tools.

Investigate

Read the evidence as a connected whole

Methods, claims, citations, and visuals

Specialist reviewers inspect the full paper in context, including figures and tables—not isolated paragraphs.

Cited research

Deep-research reviewers can search the web and scholarly literature, inspect sources, and attach vetted citations to research-backed findings.

Visible review progress

The reading room shows which review areas are working, which findings have arrived, and when a research step could not complete.

Revise

Turn critique into a submission-ready draft

Anchored reading room

Move between each comment and its passage, read your paper as you wrote it in Word, filter feedback, and discuss any finding.

Apply, track, and undo

Preview suggested revisions, apply accepted changes, keep an edit history, and reverse a change without losing the review trail.

Submission exports

Export the revised paper and saved feedback as DOCX, annotated PDF, or print view, and prepare an anonymous copy for blinded review.

A custom AI reviewer for registry and claims data validity

PerfectPaper’s registry data agent reads a secondary analysis of cancer registry or administrative claims data for the gap between what the data source records and what the manuscript concludes. It checks coding validity, missing treatment and recurrence detail, how the cohort was selected into the source, and whether conclusions stay inside what the source can support. Copy the brief below into a custom agent slot.

Large registries produce large sample sizes, and large sample sizes produce narrow confidence intervals around estimates whose main source of error is not random. That mismatch is the recurring problem in this literature.

When to use this agent

  • Your manuscript analyses SEER, SEER-Medicare, NCDB, or an administrative claims database
  • Treatment exposure is identified from billing or registry codes
  • Your outcome is recurrence, which most registries do not capture
  • The cohort is very large and the effect size is small
  • The discussion draws conclusions about care quality or effectiveness

The agent brief

Paste this into a custom agent. Suggested settings: work type domain_review, skill level graduate, no tools required.

Name: Registry and claims data validity

You are reviewing a secondary analysis of cancer registry or administrative
claims data. Your task is to determine whether the conclusions stay within
what the data source can support.

Determine and report:

1. Source and version. Determine which database was used, which years, and
   which release. Report where this is unstated, since registry contents and
   coding change between releases.

2. Cohort selection. Determine how patients entered the database, not only how
   they entered the study. Report where the manuscript does not address who is
   represented: a hospital-based registry is not population-based, and a
   claims database covers only the insured population it draws from. State
   what the resulting cohort can and cannot generalise to.

3. Variable validity. For every key variable, determine whether it is directly
   recorded or inferred from codes. Report treatment exposure inferred from
   billing codes without acknowledging that codes capture billing rather than
   delivery. Report comorbidity indices computed from claims without
   acknowledging their dependence on coding intensity, which varies by
   institution and by insurance type.

4. Missing detail. Determine whether the analysis requires detail the source
   does not carry. Registries generally lack recurrence, dose, duration,
   performance status, and reasons for treatment choice. Report every
   conclusion that depends on a variable the source does not record, and name
   the variable.

5. Missing data. Determine how missing values were handled, particularly for
   stage and grade, which are frequently incomplete. Report complete case
   analysis presented without assessment of whether missingness relates to the
   outcome.

6. Confounding by indication. Where treatments are compared, determine whether
   the manuscript addresses that treatment was chosen for reasons recorded
   nowhere in the data. Report comparative effectiveness conclusions that do
   not address this, and state that statistical adjustment cannot resolve
   confounding by unrecorded indication.

7. Precision versus accuracy. Where the sample is very large, determine
   whether the manuscript distinguishes statistical significance from
   meaningful difference. Report narrow confidence intervals presented as
   precision when the dominant error is systematic rather than random.

8. Limitations. Determine whether the limitations section names the specific
   limitations of this source rather than generic ones. Report a limitations
   paragraph that would apply unchanged to any study.

For each finding, state what the source cannot support and propose a narrower
claim that it can.

What this agent catches

Failure What the source cannot support
Recurrence analysed in a registry lacking it The outcome is not recorded
Treatment from billing codes, unqualified Billing is not delivery
Comparative effectiveness without indication Choice reasons are unrecorded
Hospital registry called population-based Generalisation to the population
Narrow intervals presented as precision Systematic error dominates

How this differs from the built-in review

The standing team includes methodology, causal inference and conclusion validity specialists that assess a design on its own terms. What they do not carry is knowledge of what a specific data source contains — that most cancer registries do not record recurrence, that claims-derived comorbidity depends on coding intensity, or that a hospital-based registry is not population-based. Those facts are what turn a general methodological reading into a specific, actionable objection.

Pairs naturally with competing risks and time-related bias, since registry survival analyses commonly need both. Full set: AI peer review for epidemiology.

Review my manuscript

Frequently asked questions

Can I study cancer recurrence using SEER?

Generally not directly, because most cancer registries record vital status and cause of death rather than recurrence. Analyses needing recurrence usually require linked claims with a validated algorithm, and the algorithm’s performance should be stated.

Is NCDB population-based?

No. It is a hospital-based registry drawing from accredited cancer programmes, so it covers a large share of cases but is not a population sample. Conclusions about population incidence or population-level disparities need that limitation stated.

Does a large sample make an observational estimate reliable?

It makes it precise, not accurate. With hundreds of thousands of patients the confidence interval becomes narrow while confounding by indication and coding error remain unchanged, and those become the dominant sources of error.

What is confounding by indication?

Bias arising because treatment was chosen for reasons related to prognosis — reasons usually recorded nowhere in administrative data. Adjustment can only handle variables present in the data, so it cannot resolve confounding by an unrecorded indication.

Last updated September 9, 2026

A careful read when you need a second opinion.

Upload your paper and receive structured, sourced feedback before you submit.