aipaperdetect
APEAS v2.0 · 18 agents

What each agent actually checks

AIPaperDetect isn't one AI detector — it's eighteen specialized checks that run as a pipeline. This page explains each one in plain language: what it looks for, how it works, and — because honesty matters more in this product than most — what it can't do.

Every result is probabilistic evidence for your own revision process, never proof of misconduct. See our usage policy.

Core

Run on every plan, including Free. These seven produce the weighted overall score (0-10) shown on your report.

Writing Qualitystyle_tone

scoredAll plans

Scores the writing itself: academic register, clarity, coherence, and argument structure — the things a tired reviewer notices in the first two pages.

How: Reads representative samples from the beginning, middle, and end of your manuscript and evaluates them against discipline-standard academic writing conventions.

Limits: It judges writing quality, not correctness of content. A beautifully written paper with broken methodology will still score well here — that's what Scientific Quality is for.

Citation Consistencycitation_consistency

scoredAll plans

Cross-checks every in-text citation against the bibliography, both directions: citations pointing at entries that don't exist, and bibliography entries never cited in the text.

How: Extracts citation markers (numbered [12] and author-year styles) and the reference list, then matches them pairwise, reporting orphans on both sides plus naming inconsistencies.

Limits: Works best on cleanly extracted text. PDFs with unusual layouts can garble the reference section — if the count looks wrong, check the extraction, not just the paper.

AI Detectionai_detection

scoredAll plans

Estimates how likely the text is LLM-generated, as a calibrated 0-10 score — not a binary human/AI verdict.

How: Analyzes signals like lexical repetition, structural uniformity, hedging density, burstiness, and depth of specific detail, then reports each signal separately alongside the overall score.

Limits: Detection is probabilistic — mid-range scores mean 'inconclusive,' full stop. A score is never proof of misconduct, which is why our policy prohibits using results to publicly accuse others' work.

Reference Integrityfabrication

scoredAll plans

Verifies that cited references actually exist: real DOIs, real arXiv IDs, metadata that matches the paper being cited.

How: Resolves each DOI and arXiv ID against Crossref and arXiv's live APIs, compares returned titles/authors/years against your bibliography entry, and flags mismatches — including DOIs that resolve to a completely different paper.

Limits: References without any identifier (older books, some venues) can't be conclusively verified online and are reported as unverifiable, not fabricated.

Source Retrievabilityretrievability

scoredAll plans

Checks whether the sources you cite can actually be retrieved — dead URLs, non-resolving DOIs, paywalled-into-oblivion links.

How: Attempts live resolution of every DOI and URL in the bibliography via web requests and the Crossref API, and categorizes each entry as accessible, partially accessible, or inaccessible with the reason.

Limits: A source being unreachable today doesn't always mean it never existed — servers go down. It reports what it observed at scan time.

Reference Relevancerelevance

scoredAll plans

Judges whether each cited work is topically appropriate for this paper — catching padding, citation-stuffing, and copy-pasted bibliographies.

How: Reads the paper's abstract, introduction, and conclusion to establish what it's about, then rates every bibliography entry from essential to irrelevant.

Limits: Interdisciplinary work can legitimately cite far-afield sources; treat 'weak' flags as prompts to double-check the citation's purpose, not as errors.

Academic Integrityacademic_integrity

scoredAll plans

A methodology-transparency review: are methods described honestly, results reported completely, limitations acknowledged, conflicts disclosed?

How: Reads the full paper and scores plagiarism-adjacent signals, claim support, result plausibility, and disclosure practices — the checks an editor runs mentally before sending a paper to review.

Limits: It evaluates what the paper says about itself. It cannot detect fabricated lab data that is internally consistent — no automated tool can.

Qualitative

Deeper research-quality reviews, available from Starter. Qualitative by design — they inform you without moving the overall score.

Scientific Qualityscientific_quality

Starter +

Evaluates the science: methodology soundness, experimental rigor, reproducibility signals, and whether conclusions follow from the evidence.

How: Full-paper qualitative review calibrated to your field's norms (the pipeline auto-detects the field first), producing section-level findings rather than a single opaque number.

Limits: Qualitative by design — it doesn't feed the overall score. A domain expert will still catch things an automated reviewer can't.

Novelty Assessmentnovelty

Starter +

Assesses whether the claimed contribution is actually new, against the live literature — not against a stale training snapshot.

How: Searches Semantic Scholar and OpenAlex for closely related prior work, compares the paper's claims against what it finds, and excludes hits that are the paper itself (so public preprints aren't penalized as 'prior art').

Limits: Search coverage is partial: finding nothing similar supports novelty but can't prove it. Verdicts cite the specific prior works found so you can judge the overlap yourself.

Pro

Advanced audit modules on the Pro plan. Also qualitative — findings, not score changes.

Statistical Integritystatistical_integrity

Pro

Hunts for numbers that don't add up: percentages that sum past 100, deltas that don't match their operands, results implausibly clean for the stated sample size.

How: Cross-references every numeric claim it can find across abstract, tables, and prose, checks arithmetic consistency, and evaluates statistical-practice signals (p-values, intervals, multiple-comparison handling).

Limits: It verifies internal consistency, not truth — a paper can be arithmetically perfect and still wrong. Complex derived statistics beyond arithmetic checking are flagged for human review, not judged.

Ethics Complianceethics_compliance

Pro

Checks the compliance paperwork story: IRB/ethics approval, informed consent, data-privacy handling, dual-use discussion, and conflict-of-interest disclosure.

How: Determines what kind of research this is (human subjects? animal? sensitive data?) and then checks which statements are present, absent, or genuinely not applicable — a paper with no human subjects isn't penalized for lacking IRB approval.

Limits: Presence of a statement isn't proof the approval exists. Journals verify credentials; this agent verifies the manuscript is submission-complete.

Literature Coverageliterature_coverage

Pro

Asks: did the related-work section miss anything important? Seminal works, recent developments, competing approaches.

How: Reads the abstract and related-work sections, then searches Semantic Scholar and OpenAlex for the works a reviewer in this area would expect to see cited, reporting specific gaps.

Limits: The most search-intensive agent — its suggestions are leads, not mandates. Some 'missing' works may be legitimately out of scope for your framing.

Integrity screeners

Integrity screeners (Pro): the checks editors and research-integrity offices run. Qualitative, evidence-cited.

Tortured Phrasestortured_phrases

Pro

Detects the linguistic fingerprints of paper mills: mangled terminology ('counterfeit consciousness' for AI, 'irregular backwoods' for random forest), template writing, synonym-substituted boilerplate.

How: Pattern analysis over the full text against known papermill markers, reporting specific phrases with locations.

Limits: Deliberately narrow: it flags papermill patterns, not AI writing in general — a modern LLM's fluent prose correctly does not trigger it.

Retraction Checkretraction_check

Pro

Screens every cited reference against retraction records — so you don't build an argument on a retracted study (it happens more than you'd think).

How: Queries Crossref's retraction metadata (the infrastructure behind Retraction Watch data) for each DOI you cite, matching by exact DOI rather than fuzzy title search.

Limits: Covers retractions registered in Crossref's update records. Very recent retractions or venues with poor metadata hygiene may not appear yet.

Duplicate Submissionduplicate_submission

Pro

Searches for prior versions, duplicate publications, and salami-slicing in the authors' publication trail — the checks an editor runs before desk-rejecting for redundant publication.

How: Searches academic indexes for the title, distinctive phrases, and the authors' adjacent work; classifies findings (legitimate preprint, undisclosed prior version, duplicate publication) — and explicitly excludes matches that are this same paper, so posting a preprint never counts against you.

Limits: Search-based, so absence of evidence isn't evidence of absence: a clean result means 'nothing found,' never 'no duplicates exist.'

GenAI Disclosuregenai_disclosure

Pro

Audits the AI-use disclosure statement: is one present, and is it consistent with what the AI Detection agent actually measured?

How: Locates disclosure language, compares it against major venue policies (Nature, Science, IEEE, ACM norms), and cross-references the paper's own AI-detection score for consistency.

Limits: Venue policies change fast; treat its policy mapping as guidance and check your target venue's current author instructions.

Synthesis

Run automatically at the end of author/reviewer mode — they read every other agent's output, so they can't be selected standalone.

Author Feedbackauthor_feedback

Pro

The pre-submission coach: synthesizes every other agent's findings into prioritized, actionable revision guidance with a predicted review outcome.

How: Runs last, reading all prior agent results plus the full paper, and produces section-by-section recommendations ordered by how much each one moves your acceptance odds.

Limits: Its predictions are informed estimates, not promises — no tool knows your reviewer #2.

Reviewer Reportreviewer_assistant

Pro

A structured peer-review report — the strengths/weaknesses/verdict document a rigorous reviewer would write, from STRONG_ACCEPT to STRONG_REJECT.

How: Synthesizes all prior agent findings into standard peer-review format with cited evidence from the paper, flagging the red flags other agents raised.

Limits: Built to assist reviewers and authors, not replace review. Verdicts are calibrated to be conservative — a WEAK_ACCEPT here is a paper worth polishing, not shipping.

How the overall score works

Only the seven core agents contribute to the 0-10 overall score, each with a fixed weight (reference integrity weighs most at 0.25; citation consistency least at 0.10). Everything else is qualitative: findings and evidence, not score movements — so one experimental module can never silently tank a manuscript's number. Thresholds: ≥8 publication-ready, ≥6 needs revision, ≥4 significant issues, <4 not ready.

See them run on your manuscript

Free plan runs all seven core agents, three times a month. No card required.

Start freeSee plans