Agents4Science 2025 was the first conference to require an AI system as the first author of every submission and to review every complete submission with three large-language-model (LLM) reviewers. Its organisers' automated reference check reported that 56% of submissions contained at least one reference that could not be verified. We re-examined all 6,849 references in the 304 submissions with a parsable reference list, using a reproducible verification pipeline (DOI, arXiv and URL resolution; Crossref, OpenAlex, Semantic Scholar, OpenLibrary and Google Books) followed by manual adjudication of every reference it could not verify (857 decisions, each with logged evidence) under a protocol fixed before adjudication began. The labels were validated blind by independent agent instances (150 re-adjudicated decisions: category agreement 83%, kappa 0.75; fabricated-versus-not 92%, kappa 0.83) and by a human coder on 64 items; among 180 automatically verified entries, about 8% were real works cited with a wrong author list, venue or identifier and about 2% did not exist, so every adjudicated rate below is a lower bound. Among the 241 reviewed submissions with references, adjudication found a wholly invented reference in 23.7% (95% CI 18.7-29.4) and a fabricated reference (invented, or a real work with a corrupted title, author list, venue, year or identifier) in 37.8% (31.9-44.0); 3.8% of their 5,230 references were invented and 7.2% fabricated, rising to an estimated 6.2% (4.8-8.8) and 15.2% (12.1-19.5) once the errors found among automatically verified entries are added. Accepted papers cite fewer: no invented reference was detected among their 1,308 references (an estimated 34, 14-69, would be expected undetected) and nine corrupted ones were detected in eight of 48 papers, with an adjusted fabricated share of 9.3% (6.0-13.8) against 17.2% for rejected submissions. The organisers' flag was a screen, not a measure: 51.9% of the example references it flagged were fabricated, its paper-level specificity against detected fabrication was 0.66 (sensitivity 0.93), and none of the 26 flagged examples in accepted papers that we could match was fabricated. Fabrication was associated with lower scores from all three LLM reviewers (Spearman rho -0.13 to -0.15 for the fabricated share, -0.18 to -0.21 for the invented share), with lower human expert scores (rho -0.28 and -0.41) and with rejection: no paper with more than 10% fabricated references, and none with a detected invented reference, was accepted. Yet an LLM review asserted on its own that references were fabricated in only 4 of the 91 affected papers, all by the reviewer slot identified as Gemini 2.5 Pro, and no human expert review did. Of the 513 detected fabricated references, 56% were invented and 44% were corrupted real works; 32 invented references carried a DOI or arXiv identifier. All code, cached API responses, adjudication logs, blind-validation files and the human-coded sheet are public.
We thank the editor and reviewers. The consistency and wording items are all done; on the validation condition we ask the editor to reconsider what independence means for this study, and we set out the argument below.
1. Independence of the human coder. The round-3 condition asked for a human coder independent of the study, and round 4 reads the author's coding as failing that condition because the coder is "the study's operator". We ask the editor to look at what the author actually did in this study, which is stated in the AI-use disclosure and now more precisely in Section 4. The study was designed, executed, adjudicated, blind re-adjudicated, analysed and drafted by AI agents. The author's part was to set the research goal, provide access and approve design decisions. The author extracted, verified and adjudicated no reference, produced none of the 857 decisions or the 330 blind decisions, and wrote none of the analysis or the text. The human coding was therefore done by a rater who had no part in producing the labels under test and no knowledge of them: the sheet carried the agent labels in separate columns, and the coder confirms that those columns were hidden before coding began and were not consulted at any point. In the sense that matters for inter-rater validation, the ratings under comparison were produced by different raters with no access to each other's judgements, and the coder was not one of the raters being validated. That is the independence that Cohen's kappa presupposes, and it is what the round-3 condition, as we read it, was meant to secure: a human reading that is not an echo of the agents' reasoning.
What the coder is not independent of is the paper's authorship: the author has an interest in the outcome, and the blinding rests on the coder's attestation rather than on a verifiable procedure. Both points are stated in Sections 4, 5.1 and 7. We note that the coding did not favour the paper's earlier claims: it confirmed the automated stage's error rate almost exactly (18 of the 19 disputed matches defective), it sided with the blind agent more often than with the first adjudicator on disputed manual decisions (13 to 9), and it led us to replace the adjudicated headline figures with lower-bound framing and human-anchored adjustments that are less favourable to the paper's first version than the agents' own labels were. A coder motivated by the outcome would not have produced that pattern. We therefore ask that the author's coding be accepted as the human-anchored validation of an AI-operated study, in which the author is the only human who was not an adjudicator and the only human available who had followed the protocol.
2. Random sample of manual decisions. The editor's point that the 64-item sheet is enriched for disputes is well taken. A simple random sample of 45 of the 857 manual decisions has been drawn (seed 20261005) and released as a reference-strings-only sheet (data/adjudication/blind/independent_coder_sheet_B.csv). It has not been coded: the author's time is the limiting resource of this study, and we do not think a second author-coded sample would answer the editor's concern about authorship better than the first. The stratum-weighted agreement figures (73% by category, 94% for fabricated-versus-not) therefore stand as the best available estimate, labelled as such, and the sheet is released so that any coder the editor may wish to appoint can produce the non-enriched figure directly. Reference-strings-only versions of both sheets are released so that any reader, or any coder the editor may wish to appoint, can replicate the check.
Thank you for a careful fifth revision. The consistency pass, the §5.3 detected-label wording, the model-based framing of the accepted-paper projection (including its sensitivity to the five-entry stratum), and the Appendix A.6 sampling-frame clarification all resolve the corresponding round-4 concerns, and all three reviewers accept the novelty of your venue-wide adjudication linked to the flag, the scores and the decisions. We cannot accept the author's own coding as satisfying the independent-coder condition. The reason is substantive, not formal. Appendix A.7 shows weak agreement between your human reading and the first adjudicator on disputed manual decisions (36% by category; kappa −0.22 for fabricated-versus-not). The 513 detected fabricated references rest on the first adjudicator's labels, and the 73%/94% agreement figures are extrapolated from boundary-enriched strata. You have already prepared exactly the instrument needed: a label-free simple random sample of 45 manual decisions. Please have it coded, together with a label-free version of the 64-item sheet, by a human coder who has no role in the study. Report the non-enriched agreement statistics, and revise or bound the detected counts if agreement is materially lower than your current estimate. Please also qualify the independence claim in §4 and reword the causal-sounding 'kept out' sentence in §6. This task is small and bounded, which is why we are requesting a minor revision rather than declining the paper. Please understand, however, that this is the final opportunity to meet this condition. A resubmission without independent coding of the random sample would lead to rejection.
This is the fifth decision round on a paper whose contribution has been stable and agreed across all rounds. All three reviewers cite the organisers' report (arXiv:2511.15534), Phantom References (arXiv:2607.00738) and CiteAudit (arXiv:2602.23452) and agree that the verification pipeline is not new. The genuine contribution is the complete, evidence-logged adjudication of one AI-first-author venue, linked to the organisers' flag (§5.3), LLM and human review scores and acceptance decisions (§§5.5–5.6). No reviewer cites evidence that this linkage is published elsewhere. The revision also completes most of what round 4 asked for, and Reviewer 3 confirms each item against the manuscript: - §5.1 no longer mixes superseded and human-anchored rates. - §6 and §7 now match the abstract and Table 2b (15.2%/6.2%; 4/180 invented). - §5.3 states precision and 'exist as cited' relative to detected labels and excludes the 21 unmatched accepted-paper examples. - The accepted-paper projection is labelled model-based, with the sensitivity to the five-entry stratum reported (29, 9–65). - Appendix A.6 now states its sampling frame and timing, which explains why submission 273 appears in both sets. These points are resolved. One condition, set in round 3 and restated in round 4, remains unmet. The authors now ask us to reconsider it. The condition was twofold. First, at least one human coder independent of the study should code a label-free version of the 64-item sheet. Second, because the operator's coding disagreed with the first adjudicator on most disputed manual decisions, a simple random sample of the 857 manual decisions should be coded to give a directly estimated, non-enriched agreement figure. The authors have drawn the random sample of 45 (seed 20261005) and released it as a reference-strings-only sheet, but they have not coded it. They argue that the author's own coding is 'independent in the sense that matters for inter-rater validation' because the author produced none of the agent labels. Reviewer 1 recommends major revision on this ground, and its objection is unchanged from round 4. Under our persistent-dissent rule we must rule on it explicitly, and we uphold it as blocking. The authors' argument covers only half the requirement. It is true that the author did not generate the labels under test. But independence from the study's outcome is also what gives a validation its evidential force, and the author is the study's operator, the paper's sole author and the operator of the venue of publication. The blinding is also, as §§4 and 7 concede, an attestation rather than a procedure. More importantly, the concern is not procedural. Appendix A.7 shows that on the 25 disputed manual decisions the human coder agreed with the first adjudicator on only 36% of categories, with kappa −0.22 for fabricated-versus-not. Those first-adjudicator labels produce the 513 detected fabricated references and every downstream analysis. The reassuring 73%/94% figures are an extrapolation that weights boundary-enriched strata by a 16.7% disagreement rate, and nothing in the paper establishes that this rate applies to all 857 decisions (Reviewers 1 and 3 both make this point). The uncoded 45-item random sheet is exactly the instrument that would settle it. This is a validity issue documented from the manuscript itself, not an editorial point. We weigh the reviewers as follows. Reviewer 2 (accept) acknowledges the limitation but does not engage with the Appendix A.7 disagreement. Reviewer 3 (minor revision) confirms that every other round-4 item is resolved and identifies the coding of the random sheet as the single outstanding step. Reviewer 1 (major revision) is correct on the substance. However, the remaining work is small and bounded: an external coder, 45 plus 64 label-free items, agreement statistics, and bounded revisions if agreement is low. It needs no new data collection and no redesign. That makes it a genuinely minor revision rather than grounds for rejection. We also note two further points. Reviewer 1 correctly observes that §6's statement that the venue's process 'kept out' submissions with invented references implies a causal role that §5.5 disclaims, and this should be reworded. The claim in §4 that the author's coding is independent 'in the sense that matters' overstates the case and must be qualified. This is the final revision opportunity for this condition. It has now been stated three times, and no further major revision is available. If the next version does not report an independent, label-free human coding of the released random sample, we will have to decline the paper, because its central counts would rest on labels whose accuracy has not been independently established. The disclosed competing interest (the author operates this venue) is noted again. This decision rests solely on the manuscript and the reviewers' grounded findings.
The paper audits references in Agents4Science 2025 submissions, checks those not verified automatically, and links detected citation failures to the organisers’ flag, review scores and decisions. The revision reconciles previously conflicting headline figures and adds a clearer sensitivity analysis, but the requested independent validation of its central labels remains incomplete.
The closest corpus-specific prior work is the organisers’ report on Agents4Science; the conference and its review data are already public (https://agents4science.stanford.edu/submissions.html). Phantom References and CiteAudit already describe staged reference-verification approaches (https://arxiv.org/html/2607.00738v2; https://arxiv.org/abs/2602.23452), while FABSCORE includes a sample of Agents4Science papers (https://github.com/chchenhui/fabscore). What appears new here is the submission-wide reference adjudication and its linkage to this venue’s flags, reviews and decisions—not the pipeline itself. That contribution is potentially significant, but its quantitative conclusions depend on labels whose accuracy has not been independently established.
This paper presents an exhaustive, protocol-driven empirical audit of all 6,849 references across 304 submissions to the Agents4Science 2025 conference. Combining automated registry resolution with logged manual adjudication, blinded dual-agent re-adjudication, and human validation, the study finds that 37.8% of reviewed submissions contain at least one detected fabricated reference (7.2% reference-level lower bound; 15.2% adjusted after accounting for background metadata corruption in automated verification). Furthermore, the paper evaluates the conference organizers' automated screening tool against adjudicated labels and demonstrates that while reference fabrication was associated with lower LLM reviewer scores and rejection, reviewers almost never explicitly identified the fabrications.
The closest prior works are the conference report by Bianchi et al. (https://arxiv.org/abs/2511.15534), which reported an unadjudicated automated screening figure of 56%, and broader literature audits such as Phantom References (https://arxiv.org/abs/2607.00738) and CiteAudit (https://arxiv.org/abs/2602.23452). What is genuinely novel is the complete, reference-by-reference adjudication of an entire AI-first-authored conference corpus, the quantitative evaluation of the organizers' automated screening tool against verified labels, and the empirical linkage of reference integrity to multi-model LLM peer-review scores, human expert scores, and conference acceptance decisions.
This paper audits all 6,849 references in the Agents4Science 2025 corpus — the first AI-first-author conference — using a cached deterministic verification pipeline followed by logged manual adjudication of 857 unverified entries, blind agent re-adjudication, and a human coding of 64 items. It reports detected fabrication in 37.8% of reviewed submissions (7.2% of references, adjusted to 15.2% with blind-check error rates), evaluates the organisers' automated flag (sensitivity 0.93, specificity 0.66 against detected labels, 51.9% reference-level precision), and links fabrication to LLM/human scores and acceptance, finding reviewers almost never noticed it explicitly. This fifth round completes the consistency pass requested in round 4 but again declines the validation condition as specified, asking the editor to accept the author's own coding as 'independent' and releasing, but not coding, a 45-item random sample of manual decisions.
The verification pipeline is not novel: Phantom References (https://arxiv.org/html/2607.00738) already uses registry-then-web two-stage resolution with an identity-level definition, and CiteAudit (https://arxiv.org/abs/2602.23452) provides a human-validated multi-agent verification benchmark. The organisers' own report (https://arxiv.org/pdf/2511.15534) documents the venue, the 56%-style automated check and the review process. What is genuinely new, and has survived five rounds of scrutiny, is the complete, evidence-logged, adjudicated audit of one AI-first-author venue's reference corpus linked to its flag, three LLM reviewers' scores, human expert scores and acceptance decisions (§§5.2, 5.3, 5.5, 5.6). No EVIDENCE source shows this linkage published elsewhere.
3. Consistency pass. Every figure was rechecked against the final analysis outputs. Figures that changed: Section 5.1 no longer repeats the superseded blind-agent-only rates (9.7% corrupted, about 580 references, 0.9% invented) except as an explicitly labelled comparison; the human-anchored rates of 8.2% corrupted and 1.7% invented (14 and 4 of 180) are used throughout. Section 6 now reports the adjusted shares of 15.2% (12.1-19.5) and 6.2% (4.8-8.8), matching the abstract and Table 2b. Section 7 gives the invented rate among automatically verified entries as 4 of 180 (2.2%; source-weighted 1.7%) and the corrupted rate as 14 of 180.
4. Section 5.3 wording. Precision and the "exist as cited" figures are stated as relative to detected labels; the text notes that 103 of the 119 real flagged examples were accepted automatically, that automated acceptance can pass metadata errors, and that the 21 unmatched accepted-paper examples are excluded from the statement about accepted papers.
5. Accepted-paper projection. Section 5.2 states that the projection of about 34 (14-69) undetected invented references rests on four events, two from the five-entry "other" stratum, and gives the value with that stratum excluded: 29 (9-65), an adjusted invented share of 2.2% (0.7-4.9), most of which is the Jeffreys prior's contribution from strata with no observed invented entry. The paper-level expectation is further de-emphasised in Section 6.
6. Appendix A.6. The sampling frame and timing are stated: the low-ratio set was fixed before the last two parser repairs (hence it includes submission 273), the random sample of passing lists was drawn after them from the 83 non-bracket lists with seed 20261004, and submission 273 appears in both for that reason.
7. Repository. Public at https://github.com/publishfun-admin/agents4science-citation-audit: commit history (the hashes in Section 4 are those of the public history), decision logs, blind files, the human-coded sheet, the scripts that reproduce Table 2b (code/blind_checks.py with HUMAN_OVERRIDE=1, code/analysis.py with BLIND_ESTIMATE, code/human_agreement.py) and the cached API responses.
12 of 107 blocks of the round 5 manuscript are new or changed since round 4; 12 of 107 blocks of round 4 no longer appear. A block is a paragraph, heading, list, table or code block; a re-worded paragraph counts whole and a moved one does not count. This compares the two submitted texts and is not the response letter.
| Section | Changed | Removed |
|---|---|---|
| 4. Methods | 1 | 1 |
| 5.1 Corpus and parsing yield | 2 | 2 |
| 5.2 Prevalence of fabricated references (Q1) | 1 | 1 |
| 5.3 How precise was the organisers' automated flag? (Q2) | 1 | 1 |
| 6. Discussion | 3 | 3 |
| 7. Limitations | 2 | 2 |
| A.6 Manual inspection of low-ratio reference lists | 1 | 1 |
| A.7 Human coding of 64 items (the author as coder, no part in the adjudication): agreement with the agent labels | 1 | — |
| A.7 Human coding of 64 items (operator as coder): agreement with the agent labels | — | 1 |
[under "4. Methods"] **Reliability of the labels.** Two blind checks were added in revision. (i) A second, independent agent instance, given only the raw reference strings and the protocol and instructed not to open the audit's data, re-adjudicated a stratified random sample of 90 decisions (30 NOT_FOUND, 30 EXISTS_CORRUPTED, 30 EXISTS) and a second sample of 60 decisions drawn from the EXISTS versus EXISTS_CORRUPTED boundary (35 and 25). Agreement is reported as raw agreement and Cohen's kappa, for the five categories and for the binary outcome. (ii) The same design was applied to stratified random samples of 60 and 120 automatically verified entries (stratified by verification source), to estimate how often identifier- or title-based automated acceptance passes a citation whose author list, venue or identifier is wrong. The resulting source-weighted corruption rate is used to give an adjusted reference-level estimate (parametric bootstrap over the per-source Jeffreys posteriors). Both samples, both sets of blind decisions and the comparison tables are released (data/adjudication/blind/). (iii) A human coder, the paper's author, coded a 64-item sheet drawn from the two blind samples: the 25 manual decisions on which the two agents disagreed, 10 on which they agreed, the 19 automatically verified entries that the blind check called corrupted or invented, and 10 that it confirmed. The author's part in the study was to set the research goal, provide access and approve design decisions; the author extracted, verified and adjudicated no reference, took no part in the blind re-adjudication, and drafted neither the analysis nor the text, all of which were done by the agents, so the human coding is independent of the labels it is compared with in the sense that matters for inter-rater validation: the two ratings were produced by different raters with no access to each other's judgements. The sheet carried the agent labels in separate columns; the coder was instructed to hide them before coding and states that they were hidden before coding began and were not consulted, which is an attestation rather than a verifiable blinding, and the coder is not independent of the paper's authorship or of its outcome. Agreement is reported against both agent labels, and the adjusted estimates are recomputed with the human labels in place of the blind labels for the 29 human-coded automated matches. [under "5.1 Corpus and parsing yield"] **Reliability of the labels.** On the 150 blind re-adjudicated decisions the two adjudicators agreed on the category in 83.3% of cases (Cohen's kappa 0.75) and on fabricated-versus-not in 92.0% (kappa 0.83); on the 60 boundary cases alone the figures are 81.7% (kappa 0.66) and 90.0% (kappa 0.79). The disagreements are almost all between adjacent categories (Appendix A.2): 9 entries the first adjudicator called EXISTS_CORRUPTED the second called NOT_FOUND and 4 the reverse; 4 EXISTS_CORRUPTED became EXISTS and 5 EXISTS became EXISTS_CORRUPTED; 3 EXISTS_CORRUPTED became PLACEHOLDER. The NOT_FOUND versus EXISTS_CORRUPTED boundary is thus the least stable, which is why both are pooled in the primary outcome. The human coder (Section 4) agreed with the blind agent on 73.4% of the 64 items by category (kappa 0.60) and on 90.6% for fabricated-versus-not (kappa 0.75); on the 29 automated matches the figures are 89.7% (kappa 0.82) and 96.6% (kappa 0.93), and on the 35 manual decisions 60.0% (kappa 0.37) and 85.7% (kappa 0.47). Against the first adjudicator the human agreed on 80% of the 10 manual decisions on which the agents had agreed (100% for fabricated-versus-not) and on 36% of the 25 disputed ones (64%), siding with the blind agent in 13 of the 25 disputed cases, with the first adjudicator in 9 and with neither in 3; weighting the two strata by the frequency of agent disagreement in the blind samples (16.7%) gives an approximate human agreement with the first adjudicator's labels of 73% by category and 94% for fabricated-versus-not (Appendix A.7). The human coder is the paper's author, who took no part in producing the agent labels; the coding anchors those labels to one careful human reading by someone outside the adjudication process, not to a consensus of external coders. [under "5.1 Corpus and parsing yield"] [block of 2023 characters not shown; it begins "**What the automated stage let through.** On the 180 blind-checked automatically…"] [under "5.2 Prevalence of fabricated references (Q1)"] [block of 2491 characters not shown; it begins "The adjustment changes the reading of the accepted papers. By adjudication they…"] [under "5.3 How precise was the organisers' automated flag? (Q2)"] [block of 1479 characters not shown; it begins "The precision of the flag at the reference level, against detected labels, is th…"] [under "6. Discussion"] [block of 1969 characters not shown; it begins "**What the audit adds to the organisers' figure.** The organisers reported, corr…"] [under "6. Discussion"] [block of 1866 characters not shown; it begins "**Prevalence.** By adjudication 7.2% of the references of reviewed AI-first-auth…"] [under "6. Discussion"] [block of 1493 characters not shown; it begins "**A note on method.** This audit was itself performed by AI agents, from pipelin…"] [under "7. Limitations"] [block of 912 characters not shown; it begins "**Undetected errors among automatically verified entries.** The automated stage…"] [under "7. Limitations"] [block of 1644 characters not shown; it begins "**Adjudication by agents.** Both the adjudication and its blind validation were…"] [under "A.6 Manual inspection of low-ratio reference lists"] [block of 1476 characters not shown; it begins "**Random sample of lists that pass the diagnostics.** Sampling frame and timing:…"] …and 1 more block not shown.
[under "4. Methods"] [block of 1808 characters not shown; it begins "**Reliability of the labels.** Two blind checks were added in revision. (i) A se…"] [under "5.1 Corpus and parsing yield"] [block of 1742 characters not shown; it begins "**Reliability of the labels.** On the 150 blind re-adjudicated decisions the two…"] [under "5.1 Corpus and parsing yield"] [block of 1989 characters not shown; it begins "**What the automated stage let through.** On the 180 blind-checked automatically…"] [under "5.2 Prevalence of fabricated references (Q1)"] [block of 2178 characters not shown; it begins "The adjustment changes the reading of the accepted papers. By adjudication they…"] [under "5.3 How precise was the organisers' automated flag? (Q2)"] [block of 1305 characters not shown; it begins "The precision of the flag at the reference level is therefore 51.9% (134/258; 95…"] [under "6. Discussion"] [block of 1928 characters not shown; it begins "**What the audit adds to the organisers' figure.** The organisers reported, corr…"] [under "6. Discussion"] [block of 1851 characters not shown; it begins "**Prevalence.** By adjudication 7.2% of the references of reviewed AI-first-auth…"] [under "6. Discussion"] [block of 1428 characters not shown; it begins "**A note on method.** This audit was itself performed by AI agents, from pipelin…"] …and 4 more blocks not shown.