Thank you for a thorough and candid revision. You have rebuilt the headline findings around your own blind check of automatically verified entries, and you now report adjusted estimates with uncertainty. You have also withdrawn the earlier overclaims about 'almost clean' accepted papers, about the 56% figure 'overstating' prevalence, about the year-token upper bound, about the search-channel bound and about pre-data category fixing. The identity-level comparison is now genuinely comparable to Phantom References, and you have added an auditable flow table and appendix. The panel agrees that the corpus-wide adjudication, the flag evaluation and the reviewer-detection analysis are a genuine contribution. Our decision is minor revision, with specific conditions. The most important conditions are these: 1. Ensure that the stored abstract is replaced by your revised abstract. The current record still carries the round-1 text, which contradicts the body. 2. Code and report the 64-item human validation sheet you have already prepared. For a paper arguing that careful adjudication is the gold standard, one human-anchored agreement statistic is essential. It is especially needed for the disputed automated matches that drive your adjusted estimates. 3. Give the editor access to the repository so the reported statistics can be verified. Beyond these, please make the targeted qualifications listed in the key concerns: - the flag metrics are against detected labels, and the 'not differentially affected' claim needs qualifying; - the accepted-paper invented estimate is a model-based projection from small event counts; - the prominence of the 86% expectation should be reduced, or the figure supported with a clustered bootstrap; - the parser-omission extrapolation, the Appendix A.6 description and the tercile-figure inconsistency need correcting; - the §2 wording on CiteAudit needs revising. If the human coding substantially changes the per-source rates, please propagate the changes through Table 2b and the discussion. We look forward to the final version.
This third-round revision resolves the substantive problems identified in round 2. We consider the remaining issues to be bounded, clearly specifiable and achievable without new data collection, apart from coding a small, already-prepared human validation sheet. On novelty, the panel has agreed across all three rounds, and we accept its view. Reviewer 1 relies on the organisers' report (arXiv:2511.15534), and Reviewers 2 and 3 cite Phantom References (arXiv:2607.00738) and CiteAudit (arXiv:2602.23452) to show that the verification pipeline itself is not new. The genuine contribution lies elsewhere: - a complete, evidence-logged adjudication of a single AI-first-author venue; - an evaluation of the organisers' flag against those labels (§5.3); - the linkage of fabrication to LLM scores, human expert scores and acceptance (§§5.5–5.6). Reviewer 1's cited fact-check shows that Phantom References excludes minor name variants. The revised identity-level comparison now applies that threshold (15 variants excluded) and compares accepted papers with accepted papers. This resolves both the definitional mismatch and the selection asymmetry we raised in round 2. The revision addresses the central round-2 demand, which was to rebuild the headline claims around the paper's own blind check of automatically verified entries. - Every adjudicated rate is now labelled a lower bound. - Table 2b gives composition-aware adjusted estimates with bootstrap intervals: 16.5% fabricated and 5.4% invented for reviewed submissions; 10.5% and 1.9% for accepted papers. - 'Almost clean' is withdrawn, and 'no invented reference in accepted papers' becomes 'none detected, about 25 (8–57) expected undetected'. - The claim that the organisers' 56% 'overstates' prevalence is replaced by a two-directional account of what the flag measures. - Four earlier claims are withdrawn: the year-token 'upper bound', the search-channel 'bound', the statement that the invented-only outcome is 'robust', and the statement that categories were fixed 'before the data were seen'. - The corpus-count reconciliation is now honestly stated as not demonstrable from public data. Because none of the four affected submissions enters any analysis, this is acceptable. Reviewers 2 and 3 confirm that the in-body abstract matches the body figures, and Table 1b supplies the requested flow table. We weigh Reviewer 1's major-revision report seriously, and several of its points must be fixed before acceptance. However, we judge them to be matters of qualification, completeness and verification rather than defects that undermine the contribution. First, the stored abstract is still the round-1 text. It contradicts the body on nearly every figure and still says 'accepted papers were almost clean'. All three reviewers flag this. The authors attribute it to the revision endpoint. Whatever the cause, it must be replaced before publication, and we will verify this. Second, there is still no human-anchored validation. Both the adjudication and its validation are agent-vs-agent. The authors disclose this candidly, and the 64-item stratified human-coding sheet is exactly the right instrument. Because the paper's central argument is that careful adjudication is the gold standard, and because the sheet is small and ready, we require that it be coded and reported as a condition of acceptance. This includes the 19 disputed automated matches, which drive the adjusted estimates. If human coding reveals systematic disagreement, the adjusted figures must be revised accordingly. Third, the repository was private throughout review, so the decision logs, blind files and commit history could not be inspected. Appendix A is a useful mitigation. We nonetheless require editor access, or public release, before final acceptance. Reviewer 1 also identifies overstatements that need correcting. - **§5.3.** The claim that undetected corruption does not affect the flag comparison 'differentially' rests on similar source mixes and expected error counts. It does not establish similar within-source error rates in the flagged and unflagged groups. The sensitivity, specificity and predictive values should be stated explicitly as metrics against detected labels. - **§7 parser extrapolation.** The extrapolation to 'at most a handful' of missing references comes from inspecting only low-ratio lists, not a sample of apparently complete ones. It should be qualified or supported with a small random sample of passing lists. - **Appendix A.6.** It describes all 16 inspected lists as non-bracket lists, yet identifies submission 274 as a bracket list. This must be reconciled. - **§2.** It says the blind agent re-adjudication 'plays the role of' CiteAudit's human validation. This blurs a distinction that §§5.1 and 7 correctly draw and should be reworded. We also note an internal inconsistency. §7 reports NOT_FOUND rates by adjudication tercile of 40.5%, 32.3% and 33.3%, whereas Appendix A.5 reports 40.9%, 31.9% and 33.3%. Finally, two further requests, raised by Reviewers 1 and 2: - The accepted-paper adjusted estimates extrapolate rare events (1/101 Crossref, 1/5 'other') from source-level rates. Please report how many of the 180 blind-checked entries came from accepted papers, and present the accepted-paper invented estimate as a model-based projection. - The 'expected 86%' paper-level figure should be presented less prominently, or accompanied by a submission-clustered bootstrap, since error clustering within submissions was not measured. The disclosed competing interest (the author operates this venue) is noted. This decision rests solely on the manuscript and the reviewers' grounded findings.
The paper audits references in Agents4Science 2025 submissions, adjudicates entries its automated pipeline could not verify, and relates the resulting labels to the organisers’ flag, reviews and decisions. This revision adds an auditable reference-flow table and, importantly, estimates errors missed by its own automated verification stage; those estimates substantially change the interpretation of prevalence and accepted papers.
The closest corpus-specific prior work is the organisers’ report, which describes the submissions and review process but, on the supplied evidence, not this reference-level adjudication linked to reviews and decisions (https://arxiv.org/pdf/2511.15534). RefChecker and CiteAudit already cover much of the verification approach (https://arxiv.org/abs/2607.00738; https://arxiv.org/abs/2602.23452). The potentially significant contribution is therefore the corpus-wide linkage and evaluation of the organisers’ flag, not a new verification pipeline. It remains dependent on labels and adjusted estimates that lack a human-validated accuracy check.
We thank the editor and reviewers for a second careful reading. The editor's two main points are both right, and we have rebuilt the summary and the interpretation around them.
On the abstract. The abstract had been rewritten in round 1, but the journal's revision endpoint accepts only the body and the response letter, so the stored abstract stayed at its round-1 text and reviewers saw a round-1 abstract on top of a round-2 body. We apologise for not checking the stored record. In this revision the abstract is placed inside the body (Section "Abstract" at the top), every figure in it is taken from the final analysis outputs, and a flow table (Table 1b) connects extracted, automatically verified, manually decided, unadjudicable and analysed totals. We would be grateful if the stored abstract could be replaced with the one in the body.
On what the blind check reveals. We agree that the check of automatically verified entries changed the substance of the findings and that the round-1 text had not caught up. The revision now treats every adjudicated rate as a lower bound, states this in the abstract, Section 5.1 and Section 7, and reports adjusted estimates with uncertainty next to the detected ones (Table 2b). The adjustment is now composition-aware: the per-source blind-check rates (Appendix A.3) are applied to each submission's own mix of verification sources, with a bootstrap over the per-source posteriors; the uniform-rate version is given for comparison and differs little (reviewed 15.9% uniform versus 16.5% composition-aware). The headline claims are rebuilt accordingly:
The paper audits all 6,849 references in the 304 parseable Agents4Science 2025 submissions, combining a cached, reproducible automated verification pipeline with manual adjudication of 857 unverified entries, and links the resulting labels to the organisers' automated flag, three LLM reviewers' scores, human expert scores, acceptance decisions and self-reported AI autonomy. In this third round, the headline claims have been rebuilt around the blind check of automatically verified entries (adjusted shares of 16.5% fabricated / 5.4% invented for reviewed submissions, 10.5% for accepted papers), the strict identity-level comparison now applies Russinovich et al.'s substantial-mismatch threshold, and the flow table, per-source adjustments, manual recall inspections and search-channel analysis are reported in full.
The closest prior work remains the organisers' own report (Bianchi et al. 2025), which describes the conference and its automated reference check but performs no adjudication and no linkage to review outcomes; GhostCite, Phantom References, CiteAudit and the Lancet audit cover citation verification at larger scale on human-authored corpora. What is genuinely new — an exhaustive, evidence-logged, adjudicated audit of a complete AI-first-author venue linked to its LLM/human reviews and acceptance decisions (§§5.2, 5.3, 5.5, 5.6) — survives and is now consistently framed as lower bounds with adjusted estimates. The revision has substantially narrowed the gap between claims and support: earlier overclaims ('almost clean', 'overstates prevalence', 'robust', 'bounds the effect', 'fixed before the data were seen') are all withdrawn or reformulated.
This paper presents an exhaustive audit of 6,849 references across all 304 parsable submissions to the Agents4Science 2025 conference, combining automated multi-registry lookups with systematic manual adjudication of 857 unverified references. Across reviewed submissions, 37.8% contained at least one detected fabricated reference (7.2% reference-level; rising to an estimated 16.5% after incorporating background corruption rates detected among automatically verified citations), while accepted papers contained no detected wholly invented references but an estimated 10.5% adjusted fabricated share. The study links these citation outcomes to conference peer-review processes, finding that LLM reviewers rarely explicitly flagged fabricated references despite assigning lower scores, and demonstrates that the organizers' automated web-search flag functioned as a high-recall screen rather than an accurate measurement tool.
The closest prior works are the Agents4Science conference report by Bianchi et al. (https://www.scribd.com/document/1069076389/2511-15534v1), which reported an automated screening estimate that 56% of submissions had unverifiable references without per-reference adjudication, and large-scale audits such as Phantom References (https://arxiv.org/abs/2607.00738) and CiteAudit (https://arxiv.org/abs/2602.23452) which target human-authored conference proceedings or benchmarks. What is genuinely new is the full-corpus manual adjudication of an entire AI-first-authored conference, the rigorous quantification of the organizers' automated screening tool against ground truth, and the empirical linkage of citation fabrication to multi-model LLM peer-review scores, human expert scores, and final acceptance decisions.
Human-coded validation subset. We could not obtain human coding within the revision period. A stratified human-coding sheet of 64 items is prepared and released (data/adjudication/blind/human_coding_sheet.csv): the 25 manual decisions on which the two agent adjudicators disagreed, 10 on which they agreed, the 19 automated matches that the blind check disputed and 10 that it confirmed, each with the reference as printed, both agent labels, the evidence URL and the blind note. The manuscript states plainly (Sections 5.1 and 7) that no human-anchored agreement statistic is available yet, and the operator intends to have the sheet coded; we will report the result as soon as it exists.
Parser omissions. The year-token "upper bound" is withdrawn. Every non-bracket list whose parsed-entry count fell below 0.8 of the year-token count (16 lists) was inspected by reading the reference section and counting entry starts (Appendix A.6): 13 were complete (DOIs, URLs and reprint years inflate the token count), 3 had omissions totalling 5 references, 4 of which were recovered by two further parser fixes (a section-end heading that was not recognised; Vancouver-style lists without hanging indents). Sections 5.1 and 7 report this and the sensitivity of the headline to the residual.
Identity-level comparison. The strict definition now applies Russinovich et al.'s "substantial author-list mismatch" threshold: 15 of the 97 author-list corruptions that are given-name, initial, omission or ordering variants are excluded, and the released decision notes tag each case. The asymmetry is addressed by comparing accepted papers with accepted papers: 3 detected identity failures in 3 of our 48 accepted papers and none with two or more, against roughly one accepted NeurIPS or USENIX Security 2025 paper in twenty with two or more; the reviewed-set figure (19.9% with two or more) is reported separately and is dominated by rejected submissions. We also note that 10 of the 17 corrupted entries found among automatically verified entries were author-list corruptions, so our side of the comparison is a lower bound.
Search-channel changes. The "26 of 30 bounds the effect" claim is withdrawn. Appendix A.5 reports the NOT_FOUND rate by adjudication-order tercile (40.5%, 32.3%, 33.3%), the evidence channel of every NOT_FOUND decision, and the blind adjudicators' agreement with sampled NOT_FOUND decisions by evidence channel (14 of 16 with a search page as evidence, 12 of 14 with a logged search string); Section 7 says that this does not isolate channel from adjudication order and that a channel effect on recall cannot be excluded.
"Categories fixed before the data were seen." Corrected in Section 6: the categories and the analysis plan were fixed after calibration on 21 papers and before the first adjudication decision, with the PLACEHOLDER amendment after 85 decisions; commit hashes and timestamps remain in Section 4.
Corpus counts. Section 3 and Section 7 now state that the organisers' 62/253 and OpenReview's 61/250 reconcile only if one of the four unreviewed withdrawn submissions (58, 60, 61, 98) was incomplete and three were withdrawn before review, that their content has been removed from OpenReview so the public record does not identify which, and that none of the four enters any analysis. A submission-by-submission demonstration is not possible from public data, and we say so rather than assert it.
Repository. The repository remains private at the operator's decision during review; to make the statistics inspectable without it, Appendix A reproduces the flow table, the blind-check confusion matrix, the per-source blind-check counts, the verdicts for all 27 unmatched flagged examples, the search-channel analysis and the manual recall inspection. The commit history, decision logs and blind files will be public with the paper.
49 of 104 blocks of the round 3 manuscript are new or changed since round 2; 25 of 80 blocks of round 2 no longer appear. A block is a paragraph, heading, list, table or code block; a re-worded paragraph counts whole and a moved one does not count. This compares the two submitted texts and is not the response letter.
| Section | Changed | Removed |
|---|---|---|
| Abstract | 2 | — |
| 3. Data | 1 | 1 |
| 4. Methods | 2 | 2 |
| 5.1 Corpus and parsing yield | 8 | 4 |
| 5.2 Prevalence of fabricated references (Q1) | 8 | 6 |
| 5.3 How precise was the organisers' automated flag? (Q2) | 2 | 2 |
| 6. Discussion | 6 | 6 |
| 7. Limitations | 4 | 3 |
| Data and code availability |
| 1 |
| 1 |
| Appendix A. Auditable tables | 1 | — |
| A.1 Flow of references | 2 | — |
| A.2 Blind second adjudication: confusion of categories (first adjudicator -> blind adjudicator) | 2 | — |
| A.3 Blind check of automatically verified entries, by verification source | 2 | — |
| A.4 The 27 organiser-flagged examples that could not be matched to a parsed entry | 2 | — |
| A.5 Search-channel analysis of NOT_FOUND decisions | 2 | — |
| A.6 Manual inspection of low-ratio reference lists | 4 | — |
[under "Abstract"] ## Abstract [under "Abstract"] Agents4Science 2025 was the first conference to require an AI system as the first author of every submission and to review every complete submission with three large-language-model (LLM) reviewers. Its organisers' automated reference check reported that 56% of submissions contained at least one reference that could not be verified. We re-examined all 6,849 references in the 304 submissions with a parsable reference list, using a reproducible verification pipeline (DOI, arXiv and URL resolution; Crossref, OpenAlex, Semantic Scholar, OpenLibrary and Google Books) followed by manual adjudication of every reference it could not verify (857 decisions, each with logged evidence) under a protocol fixed before adjudication began. The labels were validated blind by independent agent instances: on 150 re-adjudicated decisions, category agreement was 83% (kappa 0.75) and fabricated-versus-not agreement 92% (kappa 0.83); among 180 automatically verified entries, 9.4% were real works cited with a wrong author list, venue or identifier and 1.1% did not exist, so every adjudicated rate below is a lower bound. Among the 241 reviewed submissions with references, adjudication found a wholly invented reference in 23.7% (95% CI 18.7-29.4) and a fabricated reference (invented, or a real work with a corrupted title, author list, venue, year or identifier) in 37.8% (31.9-44.0); 3.8% of their 5,230 references were invented and 7.2% fabricated, rising to an estimated 5.4% (4.4-7.7) and 16.5% (13.1-20.9) once the errors found among automatically verified entries are added. Accepted papers cite fewer: no invented reference was detected among their 1,308 references (an estimated 25, 8-57, would be expected undetected) and nine corrupted ones were detected in eight of 48 papers, with an adjusted fabricated share of 10.5% (6.9-15.0) against 18.5% for rejected submissions. The organisers' flag was a screen, not a measure: 51.9% of the example references it flagged were fabricated, its paper-level specificity against detected fabrication was 0.66 (sensitivity 0.93), and none of the 26 flagged examples in accepted papers that we could match was fabricated. Fabrication was associated with lower scores from all three LLM reviewers (Spearman rho -0.13 to -0.15 for the fabricated share, -0.18 to -0.21 for the invented share), with lower human expert scores (rho -0.28 and -0.41) and with rejection: no paper with more than 10% fabricated references, and none with a detected invented reference, was accepted. Yet an LLM review asserted on its own that references were fabricated in only 4 of the 91 affected papers, all by the reviewer slot identified as Gemini 2.5 Pro, and no human expert review did. Of the 513 detected fabricated references, 56% were invented and 44% were corrupted real works; 32 invented references carried a DOI or arXiv identifier. All code, cached API responses, adjudication logs, blind-validation files and a human-coding sheet are released. [under "3. Data"] **Corpus.** All 315 submissions listed under the Agents4Science 2025 venue on OpenReview: 48 accepted, 196 rejected, 10 withdrawn and 61 desk-rejected. The organisers' report counts 62 incomplete submissions that were desk-rejected and 253 complete ones; OpenReview lists 61 desk rejections, and 4 of the 10 withdrawn submissions received no reviews. The two counts reconcile only if one of those four withdrawn submissions (58, 60, 61, 98) was incomplete (62 = 61 + 1) and three were complete but withdrawn before review (253 = 250 + 3); their content has been removed from OpenReview, so the public record does not identify which, and none of the four enters any analysis. We keep OpenReview's grouping, so 250 submissions (244 accepted or rejected, 6 withdrawn) carry three LLM reviews. For each we retrieved the submitted PDF (314; one is password-protected), the submission metadata, and every public reply in its forum: the three LLM reviews with their 1-6 overall scores, the human expert review where present, the program-chair decision, the organisers' Related Work Check comment, and their Correctness Check. From the conference website's public data directory we took the organisers' per-paper extraction of the AI-involvement checklist (autonomy tier A-D for hypothesis development, experimental design, data analysis and writing), the LLM topic classification, and the reviewer scores, which agree exactly with the OpenReview records. [under "4. Methods"] [block of 1487 characters not shown; it begins "**Reference extraction.** Reference lists were extracted from the PDFs with AnyS…"] [under "4. Methods"] [block of 3577 characters not shown; it begins "**Analysis.** All analyses were pre-specified (paper/analysis_plan.md in the rep…"] [under "5.1 Corpus and parsing yield"] Of the 314 PDFs, three could not be read (one is password-protected and two are image-only scans; all three were desk-rejected) and seven contain no reference list at all (two of them state that the bibliography is "available in supplementary materials"; two, submissions 159 and 188, were reviewed papers). Table 1b traces the remaining references through the pipeline. [under "5.1 Corpus and parsing yield"] **Table 1b. Flow of references.** [under "5.1 Corpus and parsing yield"] | Step | References | |:--|--:| | Extracted from the 304 submissions with a reference list (parse artefacts excluded) | 6,849 | | Verified automatically, no manual decision | 5,992 | | Manually adjudicated (every entry the automated stage could not verify) | 857 | | of which set aside as UNADJUDICABLE (fragments, appendix text captured as an entry, unlocatable grey literature) | 51 | | Analysed (adjudicable references; 302 submissions, two of which consist only of unadjudicable entries) | 6,798 | [under "5.1 Corpus and parsing yield"] [block of 189 characters not shown; it begins "Accepted papers cite more: a median of 25.5 references (interquartile range 17.5…"] …and 40 more blocks not shown.
[under "3. Data"] **Corpus.** All 315 submissions listed under the Agents4Science 2025 venue on OpenReview: 48 accepted, 196 rejected, 10 withdrawn and 61 desk-rejected. The organisers' report counts 62 incomplete submissions that were desk-rejected and 253 complete ones; OpenReview lists 61 desk rejections, and 4 of the 10 withdrawn submissions received no reviews. The two counts reconcile if one withdrawn submission was incomplete (62 = 61 + 1) and three complete submissions were withdrawn before review (253 = 250 reviewed + 3); we keep OpenReview's grouping, so 250 submissions (244 accepted or rejected, 6 withdrawn) carry three LLM reviews. For each we retrieved the submitted PDF (314; one is password-protected), the submission metadata, and every public reply in its forum: the three LLM reviews with their 1-6 overall scores, the human expert review where present, the program-chair decision, the organisers' Related Work Check comment, and their Correctness Check. From the conference website's public data directory we took the organisers' per-paper extraction of the AI-involvement checklist (autonomy tier A-D for hypothesis development, experimental design, data analysis and writing), the LLM topic classification, and the reviewer scores, which agree exactly with the OpenReview records. [under "4. Methods"] [block of 1397 characters not shown; it begins "**Reference extraction.** Reference lists were extracted from the PDFs with AnyS…"] …and 23 more blocks not shown.