Agents4Science 2025 was the first conference to require an AI system as the first author of every submission and to review every complete submission with three large-language-model (LLM) reviewers. Its organisers' automated reference check reported that 56% of submissions contained at least one reference that could not be verified. We re-examined all 6,849 references in the 304 submissions with a parsable reference list, using a reproducible verification pipeline (DOI, arXiv and URL resolution; Crossref, OpenAlex, Semantic Scholar, OpenLibrary and Google Books) followed by manual adjudication of every reference it could not verify (857 decisions, each with logged evidence) under a protocol fixed before adjudication began. The labels were validated blind by independent agent instances (150 re-adjudicated decisions: category agreement 83%, kappa 0.75; fabricated-versus-not 92%, kappa 0.83) and by a human coder on 64 items; among 180 automatically verified entries, about 8% were real works cited with a wrong author list, venue or identifier and about 2% did not exist, so every adjudicated rate below is a lower bound. Among the 241 reviewed submissions with references, adjudication found a wholly invented reference in 23.7% (95% CI 18.7-29.4) and a fabricated reference (invented, or a real work with a corrupted title, author list, venue, year or identifier) in 37.8% (31.9-44.0); 3.8% of their 5,230 references were invented and 7.2% fabricated, rising to an estimated 6.2% (4.8-8.8) and 15.2% (12.1-19.5) once the errors found among automatically verified entries are added. Accepted papers cite fewer: no invented reference was detected among their 1,308 references (an estimated 34, 14-69, would be expected undetected) and nine corrupted ones were detected in eight of 48 papers, with an adjusted fabricated share of 9.3% (6.0-13.8) against 17.2% for rejected submissions. The organisers' flag was a screen, not a measure: 51.9% of the example references it flagged were fabricated, its paper-level specificity against detected fabrication was 0.66 (sensitivity 0.93), and none of the 26 flagged examples in accepted papers that we could match was fabricated. Fabrication was associated with lower scores from all three LLM reviewers (Spearman rho -0.13 to -0.15 for the fabricated share, -0.18 to -0.21 for the invented share), with lower human expert scores (rho -0.28 and -0.41) and with rejection: no paper with more than 10% fabricated references, and none with a detected invented reference, was accepted. Yet an LLM review asserted on its own that references were fabricated in only 4 of the 91 affected papers, all by the reviewer slot identified as Gemini 2.5 Pro, and no human expert review did. Of the 513 detected fabricated references, 56% were invented and 44% were corrupted real works; 32 invented references carried a DOI or arXiv identifier. All code, cached API responses, adjudication logs, blind-validation files and the human-coded sheet are public.
We thank the editor and the reviewers. All three conditions are now met, and the targeted qualifications are made.
1. Stored abstract. The journal has corrected the revision endpoint so that the abstract field is updated; this revision sends the revised abstract as the stored abstract and also keeps it at the top of the body, so the two are identical. Every figure in it is taken from the final analysis outputs.
2. Human-coded validation sheet. The 64-item sheet has been coded (data/adjudication/blind/human_coding_sheet.csv, columns HUMAN_CATEGORY and HUMAN_NOTE; agreement tables in data/dataset/human_agreement.md and Appendix A.7). We are candid about its limits: the coder is the operator of the study, and the sheet carried the agent labels in columns the coder was asked to hide, so the check is neither independent nor blind by construction (Sections 4 and 7). Results: the human agreed with the blind agent on 73.4% of items by category (kappa 0.60) and 90.6% for fabricated-versus-not (kappa 0.75), and on the 29 automated matches on 89.7% (kappa 0.82) and 96.6% (kappa 0.93). The 10 automated matches the blind check had confirmed were all confirmed by the human; of the 19 it had called defective, 18 were confirmed (14 corrupted, 4 invented) and one, a paper cited with altered given names of several co-authors, was accepted as correctly cited. On the 25 manual decisions on which the two agents had disagreed, the human sided with the blind agent in 13, with the first adjudicator in 9 and with neither in 3; on the 10 agreed decisions the human agreed with both in 8 (100% for fabricated-versus-not). Weighting the strata by the frequency of agent disagreement gives an approximate human agreement with the first adjudicator's labels of 73% by category and 94% for fabricated-versus-not. Because the human coding changed the per-source rates (14 corrupted and 4 invented among the 180 sampled entries instead of 17 and 2), Table 2b, the abstract, Sections 5.1, 5.2, 5.3 and 6 now use the human-anchored rates (reviewed submissions: adjusted fabricated share 15.2%, 12.1-19.5, and invented share 6.2%, 4.8-8.8; accepted papers 9.3% and 2.6%); the blind-agent-only figures are retained in the text for comparison and in the released tables (results_tables.md and results_tables_human_override.md).
3. Repository access. The repository is now public: https://github.com/publishfun-admin/agents4science-citation-audit (commit history, decision logs, blind files, human-coded sheet, unmatched-flag verdicts, cached API responses). The commit hashes cited in Section 4 are those of the public history.
Thank you for a careful fourth revision. The stored abstract now matches the body, the human-coding sheet has been coded and reported with commendable candour, and the repository is stated to be public. The panel continues to regard the venue-wide adjudication linked to the flag, the reviews and the decisions as a genuine and useful contribution. We cannot yet accept, for two reasons. First, the round-3 condition asked for an independent human coder. The coder was the study's operator, working from a sheet that contained the agent labels. That coding also showed low agreement with the first adjudicator on disputed manual decisions, and those decisions underlie the detected counts. Please have an independent coder, blind to the agent labels, code the 64-item sheet and a small simple random sample of the manual decisions. Report agreement with both agent labels, and revise the estimates if the rates change. Second, several headline figures are now internally inconsistent: - the superseded 9.7% (about 580) and 0.9% rates in §5.1; - the stale 16.5% and 5.4% adjusted shares in §6; - the 'one in a hundred' invented rate in §7; - the double listing of submission 273 in Appendix A.6. Please run a full consistency pass, qualify the §5.3 'exactly as cited' wording as relative to detected labels, and add a sensitivity check for the accepted-paper invented projection that excludes the five-entry 'other' stratum. These tasks are bounded. Once they are done and the repository has been spot-checked, we expect the manuscript to be ready for acceptance.
This fourth-round revision meets one of the three conditions set in round 3 in full, meets one in form, and does not meet the third as specified. It also introduces new internal inconsistencies that must be fixed before this paper, whose subject is the accuracy of bibliographic records, can enter the published record. Novelty has been stable across four rounds and we accept the panel's view. All three reviewers cite the organisers' report (arXiv:2511.15534), Phantom References (arXiv:2607.00738) and CiteAudit (arXiv:2602.23452) as prior art. On that basis the verification pipeline is not new. The genuine contribution is the venue-wide, evidence-logged adjudication of an AI-first-author corpus, linked to the organisers' flag, the LLM and human review scores, and the acceptance decisions (§§5.3, 5.5, 5.6). Reviewers 2 and 3 recommend acceptance and credit the revision's candour. Reviewer 1 recommends major revision on specific, checkable grounds. We weigh Reviewer 1's points most heavily because each can be verified against the manuscript, and on inspection each is correct. The conditions from round 3 stand as follows. (1) Stored abstract. This is resolved. The stored and body abstracts are identical. (2) Human validation. This is not met as specified. We asked for 'at least one independent human coder'. The 64-item sheet was coded by the study's operator. The sheet also carried the agent labels, which the coder was only asked to hide (§§4, 7). The authors disclose this plainly, which we appreciate, but it is neither independent nor verifiably blind. The results also matter substantively. On the 19 disputed automated matches the operator broadly confirms the blind agent (18 of 19 defective), which supports the adjusted estimates. On the 25 disputed manual decisions, however, the human sided with the first adjudicator in only 9 (Appendix A.7: 36% category agreement; kappa -0.22 for fabricated-versus-not). The first adjudicator's labels are the ones that produce the 513 detected fabricated references. The 'approximately 94%' agreement with those labels comes from weighting two strata by a 16.7% disagreement rate. Reviewer 1 correctly notes that this rate was estimated from boundary-enriched samples, so its applicability to all 857 decisions is not established. (3) Repository. The repository is now stated to be public. No reviewer could confirm its contents from the supplied evidence, and we will verify it before acceptance. Reviewer 1 also identifies stale and conflicting figures, each confirmed in the text: - §5.1 reports human-anchored source-weighted rates of 8.2% corrupted and 1.7% invented. The next sentence retains the superseded 9.7% (about 580 references) and 0.9%. - §6 'Prevalence' still gives the adjusted reviewed-submission shares as 16.5% (13.1-20.9) and 5.4%. The abstract and Table 2b give 15.2% and 6.2%. - §7 says 'about one in a hundred' automatically accepted entries does not exist. The final sample gives 4/180 (2.2%). - Appendix A.6 lists submission 273 both among the low-ratio lists and in the random sample of passing lists, and does not explain the sampling frame. - §5.3 describes 119 flagged examples as existing 'exactly as cited'. Some of these were accepted automatically, and §5.1 shows that automated acceptance can pass metadata errors. The claim should be stated relative to detected labels. None of these problems threatens the core contribution, and none needs new data collection beyond a small, clearly bounded coding task. That is why we request minor rather than major revision. Several are nonetheless errors in headline quantitative claims, and the round-3 validation condition remains open. We therefore cannot accept in this round. We reiterate the disclosed competing interest: the author operates this venue. This decision rests solely on the manuscript and on the reviewers' grounded findings.
The paper audits references in Agents4Science 2025 submissions, adjudicates entries that its automated pipeline could not verify, and relates detected citation failures to the organisers’ flag, review scores and acceptance. This revision adds operator-coded checks of 64 references and recalculates adjusted prevalence estimates, but important validation and numerical-consistency problems remain.
The closest corpus-specific prior work is the organisers’ report on Agents4Science (https://arxiv.org/pdf/2511.15534). Phantom References (https://arxiv.org/pdf/2607.00738) and CiteAudit (https://arxiv.org/abs/2602.23452) already cover closely related verification methods. The distinctive contribution appears to be the venue-wide, reference-level audit linked to this venue’s flags, reviews and decisions, rather than the pipeline or taxonomy. That is a genuine contribution, but its quantitative significance depends on the reliability of the labels and adjustments.
This paper audits all 6,849 references across the 304 parseable Agents4Science 2025 submissions, combining a cached deterministic verification pipeline with logged manual adjudication of 857 unverified entries, and links the resulting labels to the organisers' automated flag, three LLM reviewers' scores, human expert scores and acceptance decisions. It reports 7.2% detected fabricated references among reviewed submissions (adjusted 15.2%), finds the organisers' flag to be a high-sensitivity screen with ~52% reference-level precision, and shows that reviewers almost never explicitly noticed fabrication despite lower scores. This fourth revision meets the three conditions set in round 3: the stored abstract now matches the body on every headline figure, the 64-item human-coding sheet has been coded and reported, and the repository is stated to be public.
The closest prior work is the organisers' own report (Bianchi et al. 2025, https://arxiv.org/pdf/2511.15534), which describes the conference and its automated reference check but performs no adjudication and no linkage to review outcomes; larger-scale audits (Zhao et al. 2026, https://arxiv.org/pdf/2605.07723; Topaz et al. 2026, https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(26)00603-3/fulltext; Shi et al. CiteAudit, https://arxiv.org/html/2602.23452v3) cover citation verification on human-authored corpora or benchmarks, not an exhaustive audit of a single AI-first-author venue linked to its reviews and decisions. The genuinely new contributions — the adjudicated ground truth for this corpus, the evaluation of the organisers' flag (§5.3), and the reviewer-detection analysis (§5.6) — have survived four rounds of scrutiny, and the verification pipeline itself is correctly framed as following prior two-stage designs rather than as novel.
This paper presents an exhaustive empirical audit of all 6,849 references across 304 submissions to the inaugural Agents4Science 2025 conference. Combining automated multi-registry lookups with systematic manual adjudication of 857 unverified entries, blinded dual-agent re-adjudication, and human-anchored validation, the study establishes that 37.8% of reviewed submissions contained at least one detected fabricated reference (7.2% reference-level lower bound, rising to 15.2% adjusted after accounting for background metadata corruption), while accepted papers had no detected wholly invented references but an estimated 9.3% adjusted fabricated share. Furthermore, it evaluates the conference organizers' automated web-search flag against adjudicated labels and links reference integrity directly to multi-model LLM review scores, human expert scores, and acceptance decisions.
The closest prior works are the Agents4Science organizers' report (Bianchi et al. 2025, https://arxiv.org/abs/2511.15534), which reported an automated screening estimate that 56% of submissions had unverifiable references without per-reference manual adjudication, and large-scale reference audits such as Phantom References (Russinovich et al. 2026, https://arxiv.org/abs/2607.00738) and CiteAudit (Shi et al. 2026, https://arxiv.org/abs/2602.23452), which analyze human-authored conference proceedings or perturbed benchmarks. What is genuinely novel here is the complete submission-wide manual adjudication of a full AI-first-authored conference corpus, the quantitative evaluation of the organizers' automated screening tool against verified ground truth, and the empirical linkage of citation fabrication to multi-model LLM peer-review scores, human expert scores, and final acceptance decisions.
Targeted qualifications (all made in the previous revision and retained): detected-label wording and the qualified flagged-versus-unflagged comparison in Section 5.3; the accepted-paper invented estimate presented as a model-based projection from source-level rates with the event counts stated (now four events after the human coding); the paper-level expectation de-emphasised and supported by a submission-clustered bootstrap; the parser-omission extrapolation replaced by the inspection of low-ratio lists and a random sample of passing lists (Appendix A.6), with submission 274 identified as a bracket list; the tercile figures reconciled with Appendix A.5; the CiteAudit sentence reworded.
23 of 107 blocks of the round 4 manuscript are new or changed since round 3; 20 of 104 blocks of round 3 no longer appear. A block is a paragraph, heading, list, table or code block; a re-worded paragraph counts whole and a moved one does not count. This compares the two submitted texts and is not the response letter.
| Section | Changed | Removed |
|---|---|---|
| Abstract | 1 | 1 |
| 2. Related work | 1 | 1 |
| 4. Methods | 2 | 2 |
| 5.1 Corpus and parsing yield | 3 | 3 |
| 5.2 Prevalence of fabricated references (Q1) | 4 | 4 |
| 5.3 How precise was the organisers' automated flag? (Q2) | 1 | 1 |
| 6. Discussion | 3 | 3 |
| 7. Limitations | 3 | 3 |
| Data and code availability | 1 | 1 |
| A.6 Manual inspection of low-ratio reference lists | 2 | 1 |
| A.7 Human coding of 64 items (operator as coder): agreement with the agent labels | 2 | — |
[under "Abstract"] Agents4Science 2025 was the first conference to require an AI system as the first author of every submission and to review every complete submission with three large-language-model (LLM) reviewers. Its organisers' automated reference check reported that 56% of submissions contained at least one reference that could not be verified. We re-examined all 6,849 references in the 304 submissions with a parsable reference list, using a reproducible verification pipeline (DOI, arXiv and URL resolution; Crossref, OpenAlex, Semantic Scholar, OpenLibrary and Google Books) followed by manual adjudication of every reference it could not verify (857 decisions, each with logged evidence) under a protocol fixed before adjudication began. The labels were validated blind by independent agent instances (150 re-adjudicated decisions: category agreement 83%, kappa 0.75; fabricated-versus-not 92%, kappa 0.83) and by a human coder on 64 items; among 180 automatically verified entries, about 8% were real works cited with a wrong author list, venue or identifier and about 2% did not exist, so every adjudicated rate below is a lower bound. Among the 241 reviewed submissions with references, adjudication found a wholly invented reference in 23.7% (95% CI 18.7-29.4) and a fabricated reference (invented, or a real work with a corrupted title, author list, venue, year or identifier) in 37.8% (31.9-44.0); 3.8% of their 5,230 references were invented and 7.2% fabricated, rising to an estimated 6.2% (4.8-8.8) and 15.2% (12.1-19.5) once the errors found among automatically verified entries are added. Accepted papers cite fewer: no invented reference was detected among their 1,308 references (an estimated 34, 14-69, would be expected undetected) and nine corrupted ones were detected in eight of 48 papers, with an adjusted fabricated share of 9.3% (6.0-13.8) against 17.2% for rejected submissions. The organisers' flag was a screen, not a measure: 51.9% of the example references it flagged were fabricated, its paper-level specificity against detected fabrication was 0.66 (sensitivity 0.93), and none of the 26 flagged examples in accepted papers that we could match was fabricated. Fabrication was associated with lower scores from all three LLM reviewers (Spearman rho -0.13 to -0.15 for the fabricated share, -0.18 to -0.21 for the invented share), with lower human expert scores (rho -0.28 and -0.41) and with rejection: no paper with more than 10% fabricated references, and none with a detected invented reference, was accepted. Yet an LLM review asserted on its own that references were fabricated in only 4 of the 91 affected papers, all by the reviewer slot identified as Gemini 2.5 Pro, and no human expert review did. Of the 513 detected fabricated references, 56% were invented and 44% were corrupted real works; 32 invented references carried a DOI or arXiv identifier. All code, cached API responses, adjudication logs, blind-validation files and the human-coded sheet are public. [under "2. Related work"] **Prevalence in the human-authored literature.** Zhao et al. [2026] audit 111 million references in 2.5 million papers across arXiv, bioRxiv, SSRN and PubMed Central and document a sharp post-2023 rise in non-existent references, concentrated in fields with rapid AI uptake and in manuscripts with linguistic signatures of AI-assisted writing. Topaz et al. [2026] audit 2.5 million biomedical papers and report a twelve-fold increase in two years, from one fabricated reference per 2,828 papers in 2023 to one per 277 in early 2026. Xu et al. [2026] (GhostCite) verify 2.2 million citations from 56,381 papers at AI/ML and security venues and find that 1.07% of papers contain invalid citations, with an 80.9% increase in 2025; they also benchmark 13 LLMs and find citation-generation hallucination rates between 14% and 95%. Ansari [2026] analyses 100 hallucinated citations that survived expert peer review at NeurIPS 2025 and proposes the failure-mode taxonomy (total fabrication, partial attribute corruption, identifier hijacking, placeholder and semantic hallucination) that our adjudication categories adapt. Russinovich et al. [2026] (Phantom References) resolve the bibliographies of accepted ICLR, ICML, NeurIPS and USENIX Security papers against several bibliographic sources, escalate unresolved entries to web-search re-verification, and count only identity-level failures (non-existent works and substantial author-list mismatches): reference-level rates are usually below 1%, but in 2025 roughly one in twenty NeurIPS and USENIX Security papers contains at least two such references. Their two-stage design (registries first, web search for the residue) is the one our pipeline follows, with manual adjudication replacing the final automated step. Shi et al. [2026] (CiteAudit) decompose citation checking into metadata extraction, memory lookup, web retrieval and a final judgement by cooperating agents, and release a human-validated benchmark on which their pipeline outperforms single LLMs and commercial checkers; our automated stage is a deterministic, cached version of the same decomposition; our blind re-adjudication is an agent-level analogue of their human validation, not a substitute for it (Sections 5.1 and 7). [under "4. Methods"] [block of 1808 characters not shown; it begins "**Reliability of the labels.** Two blind checks were added in revision. (i) A se…"] [under "4. Methods"] [block of 3577 characters not shown; it begins "**Analysis.** All analyses were pre-specified (paper/analysis_plan.md in the rep…"] [under "5.1 Corpus and parsing yield"] [block of 1774 characters not shown; it begins "**Parser recall.** For the 218 submissions with bracket-numbered lists, the rati…"] [under "5.1 Corpus and parsing yield"] [block of 1742 characters not shown; it begins "**Reliability of the labels.** On the 150 blind re-adjudicated decisions the two…"] …and 17 more blocks not shown.
[under "Abstract"] [block of 2974 characters not shown; it begins "Agents4Science 2025 was the first conference to require an AI system as the firs…"] [under "2. Related work"] [block of 2187 characters not shown; it begins "**Prevalence in the human-authored literature.** Zhao et al. [2026] audit 111 mi…"] [under "4. Methods"] [block of 1184 characters not shown; it begins "**Reliability of the labels.** Two blind checks were added in revision. (i) A se…"] [under "4. Methods"] [block of 3577 characters not shown; it begins "**Analysis.** All analyses were pre-specified (paper/analysis_plan.md in the rep…"] [under "5.1 Corpus and parsing yield"] [block of 1592 characters not shown; it begins "**Parser recall.** For the 218 submissions with bracket-numbered lists, the rati…"] [under "5.1 Corpus and parsing yield"] [block of 1047 characters not shown; it begins "**Reliability of the labels.** On the 150 blind re-adjudicated decisions the two…"] [under "5.1 Corpus and parsing yield"] [block of 935 characters not shown; it begins "**What the automated stage let through.** On the 180 blind-checked automatically…"] [under "5.2 Prevalence of fabricated references (Q1)"] [block of 444 characters not shown; it begins "**Table 2b. Adjusted reference-level estimates.** The blind-check corruption and…"] …and 12 more blocks not shown.