Thank you for a careful, transparent and potentially valuable audit. The reviewers agree that linking adjudicated reference labels to the organisers' automated flag, the LLM and human review scores and acceptance decisions is a distinctive contribution. The evaluation of the flag as 'a good screen but a poor measure' and the finding that LLM reviewers rarely identified fabrication explicitly are of clear practical interest. Your released code, caches and decision logs are exemplary. However, we cannot accept the paper in its current form. Its ground truth rests entirely on a single AI agent's adjudication, with no independent validation of either the manual decisions or the automated matches. In addition, Reviewer 3 found a corrupted author name in your own citation of Bianchi et al. 2025 ('Xian Sun' for 'Eric Sun'). This is precisely the kind of error your study classifies as fabricated, and it slipped past the identifier-based verification you describe. We invite a major revision. The central requirement is a blinded, independently adjudicated validation sample with agreement statistics and revised uncertainty bounds, together with a check of the automatically verified entries. Beyond that, please: - correct and fully re-verify your own reference list; - resolve the treatment of the 28 unmatched flagged examples; - fix the §5.6 accounting inconsistency and reconcile corpus counts with the organisers' report; - temper the cross-study prevalence comparison; - engage with the RefChecker and CiteAudit prior work; - document the timing of the analysis plan and protocol amendment; - frame the score and acceptance results as associations. With these changes, the paper could make a solid, well-evidenced contribution, and we look forward to a revised submission.
This submission audits every reference in the Agents4Science 2025 corpus. It combines a cached, reproducible bibliographic-lookup pipeline with manual adjudication of the 821 entries that pipeline could not verify. It then links the resulting labels to the organisers' automated flag, the three LLM reviewers' scores, human expert scores, acceptance decisions and self-reported AI autonomy. The panel agrees that the corpus-level linkage is the distinctive contribution. That means the evaluation of the organisers' flag against adjudicated labels (§5.3), the fabrication–score–acceptance associations (§5.5) and the reviewer-detection analysis (§5.6). Reviewer 1 places the novelty in 'adjudication across this venue's submission corpus linked to its reviews and decisions', not in the pipeline or taxonomy. Reviewer 3 reaches the same conclusion, calling the contribution 'real but incremental relative to an active literature'. Reviewer 2 judges it 'genuinely novel' and recommends acceptance. We accept that the linkage is new, but we weigh Reviewers 1 and 3 more heavily. Their concerns are specific and, in one case, backed by a cited source; Reviewer 2 acknowledges the same central weakness (single-agent adjudication) yet does not treat it as consequential. The decisive concern, raised by all three reviewers, is that the paper's ground truth has not been independently validated. Every one of the 821 manual decisions was made by a single AI agent, the author's own. No second coder, no blinded re-adjudication and no human-checked sample is reported. The paper's central methodological argument is that automated checks cannot substitute for careful adjudication, so leaving its own adjudication unvalidated undercuts the claim it is making. The weakness matters most where judgement is involved: the EXISTS vs EXISTS_CORRUPTED boundary accounts for 43% of all fabricated references. Reviewer 1 adds a second gap. None of the 5,864 automatically verified entries was sampled for correctness, and §7 concedes that identifier-based acceptance can miss wrong authors or years. Releasing the decision log is valuable, but it does not replace reported agreement statistics. Reviewer 3's fact-check gives this concern concrete force. Citing the arXiv version of the organisers' report, it shows that the manuscript's own reference list names 'Xian Sun' as an author of Bianchi et al. 2025, whereas the source lists 'Eric Sun'. The paper states that its own references 'were verified by resolving each identifier before submission'. That this author-list corruption survived is exactly the failure mode that §7 ('Identifier-based acceptance') acknowledges and that the paper classifies as EXISTS_CORRUPTED, which the paper counts as fabricated. It shows empirically that identifier resolution without author-level checks misses errors of the kind the paper counts. For a paper about citation integrity, this must be corrected, and every reference in the paper must be re-verified at the author, venue and year level, not just by identifier. Several further points require revision. - **Unmatched flagged examples (Reviewers 1 and 3).** 28 of 285 organiser-flagged examples could not be matched, 23 of them in accepted papers. They are explained only 'in the cases we could trace'. Yet the headline claim that no flagged example in an accepted paper was fabricated depends on how they are treated. Per-example evidence or explicit bounds are needed. - **Internal inconsistency in §5.6 (Reviewer 1).** The section reports 8 affected submissions with explicit LLM statements, then describes 5 independent findings plus 'the other four'. This must be reconciled. - **Cross-literature comparison (Reviewers 1 and 3).** §6 asserts that the gap with prior prevalence estimates is 'too large to be an artefact of method'. This is not supported, since the paper itself notes that those studies use different detection methods and outcome definitions. Either tone the claim down or support it with a sensitivity analysis on comparable definitions. - **Related work (Reviewer 3).** Two close prior verification pipelines identified by Reviewer 3 are neither cited nor compared: RefChecker/'Phantom References' (arXiv 2607.00738), which uses a conservative identity-level definition on ML venues, and CiteAudit (arXiv 2602.23452). - **Pre-registration timing (Reviewer 1).** The timing of the analysis plan and of the PLACEHOLDER amendment relative to data access should be documented verifiably. - **Reviewer-model identities (Reviewer 1).** The mapping of anonymised reviewer slots to named models is inferred from matching mean scores. It should be stated as an inference. - **Acceptance analysis (Reviewer 1).** The model adjusts only for mean LLM score, so it supports an association, not the claim that fabrication itself was penalised. - **Corpus counts.** The organisers' report, as quoted in Reviewer 1's fact-check, gives 62 desk rejections and 253 complete, reviewed submissions. The manuscript reports 61 desk-rejected submissions and 250 with three LLM reviews. The difference may be explained by withdrawals or incomplete reviews, but it should be reconciled explicitly. We also note, for transparency, that the author discloses operating this venue. This decision rests solely on the reviewers' grounded findings and the manuscript's content. The work is carefully executed in many respects. The categories are well defined, Wilson intervals are used throughout, limitations are candid and all materials are released. The flag-evaluation and reviewer-detection results would be a useful contribution once the ground truth is validated. At present, however, the core labels rest on one unvalidated AI adjudicator, and the manuscript contains a demonstrable citation error of the type it measures. Both are material evidential concerns that preclude acceptance or minor revision. A major revision is required.
This is a structurally complete empirical paper. It states six pre-specified research questions, gives a detailed and reproducible methodology (automated resolution against several registries followed by logged manual adjudication), reports quantitative results with confidence intervals, and includes a discussion, an extensive limitations section, a related-work section and verifiable references (19 of 20 identifiers resolve, and the one that does not is a quoted example of fabrication, not a cited source). On novelty, the evidence shows the Agents4Science organisers reported only an automated flag ([1] states "approximately 44% of submissions have no hallucinated references"). Other retrieved audits target NeurIPS and USENIX papers, not this corpus. Nothing in the evidence shows that an exhaustive, manually adjudicated audit of Agents4Science references, or its link to LLM and human review outcomes, is already published. The remaining issues are ordinary quality matters for reviewers: a single AI adjudicator, small internal count discrepancies, a missing related audit, and a disclosed conflict of interest.
The paper audits references extracted from Agents4Science 2025 submissions, combining automated bibliographic checks with AI-agent adjudication of unresolved entries (§§3–5). It reports lower fabrication prevalence than the organisers’ automated flag implied and examines associations with review scores, acceptance and whether reviewers explicitly noticed problematic references (§5).
The closest prior work is the organisers’ report, which describes the same conference and its review process (https://arxiv.org/pdf/2511.15534). GhostCite already studies citation validity at much larger scale (https://arxiv.org/html/2602.06718v2), while the NeurIPS audit examines fabricated citations that survived peer review (https://arxiv.org/html/2602.05930). What appears distinctive here is adjudication across this venue’s submission corpus linked to its reviews and decisions (§§3–5), not the verification pipeline or fabrication taxonomy alone. Its significance depends on the reliability of that adjudication, which the paper does not independently validate.
This paper presents an exhaustive empirical audit of all 6,685 references across 304 submissions to Agents4Science 2025, the first conference requiring AI systems as first authors. Using a multi-stage verification pipeline followed by protocol-driven manual adjudication of 821 unverified references, the authors find that 36.9% of reviewed submissions contained at least one fabricated reference (7.1% of all references), while accepted papers contained no wholly invented references. The study also evaluates the organizers' automated web-search flag, demonstrating high sensitivity (0.94) but modest specificity (0.66) and only 51.8% reference-level precision, and shows that while fabricated citations were associated with lower LLM and human scores and rejection, reviewers rarely explicitly identified the fabrications.
While the Agents4Science organizers previously reported an unverified 56% figure for papers with potential citation issues using an automated search check, this work provides the first exhaustive, manually adjudicated audit of the complete reference corpus. Unlike broader automated studies of the scholarly literature that rely on automated heuristics, this study inspects every unverified item, quantifies the exact false-positive profile of automated checking tools, and investigates the relationship between citation fabrication and multi-model LLM peer reviews in a live venue. The contribution is genuinely novel and provides concrete empirical benchmarks for citation integrity in agent-authored scientific writing.
The paper presents an exhaustive audit of all references in all Agents4Science 2025 submissions (6,685 references across 304 submissions), combining an automated verification pipeline with manual adjudication of 821 unresolved references. It reports that 36.9% of reviewed submissions contain at least one fabricated reference (vs the organisers' 56% automated-flag estimate), that accepted papers are nearly clean, that the organisers' flag is a good screen but a poor measure, and that LLM reviewers almost never explicitly noticed fabrication despite giving lower scores. The work is preregistered, with released logs and cached API responses.
The general problem of hallucinated-citation auditing at scale is well covered by prior work in the EVIDENCE: Zhao et al.'s 111-million-reference audit [6], the Lancet biomedical audit [9], RefChecker/Phantom References [2], and CiteAudit's multi-agent verification benchmark [13]. The organisers' own report [10] already describes Agents4Science's reference verification agents and provides per-paper verification statistics [16]. What is genuinely new here is (a) the exhaustive, per-reference manual adjudication of a complete single AI-first-author venue, (b) the evaluation of the organisers' automated flag against ground truth, and (c) the linkage of fabrication to LLM scores, human scores and acceptance — none of which the EVIDENCE shows in prior work. The contribution is therefore real but incremental relative to an active literature the paper only partially engages (RefChecker [2] and CiteAudit [13] are uncited close relatives of its pipeline).