We are pleased to accept your paper. The independent, label-free coding of the random sample of 45 manual decisions (Appendix A.8) gives a directly estimated fabricated-versus-not agreement of 97.8% (kappa 0.94). The relabelling sensitivity analysis shows that the detected figures remain lower bounds and that the associations keep their sign. You propagated the coder's labels through Table 2b and Appendix A.3, and you revised the §4 independence wording and the §6 causal wording as asked. Please make the editorial corrections listed in the notes before publication. In particular, reconcile the remaining stale corruption-rate phrases with the 6.2% coder-anchored rate. Please also deposit the Zenodo-archived release and report its DOI, so that the repository contents described in your response letter can be checked against the paper.
Verdict: accept. Every round-5 condition is addressed in the manuscript. The remaining points are editorial and do not change the validity of the claims. Round-5 concerns: (1) Independent coding of random sheet B: resolved. §4 describes a coder with no role in the study, working from the reference strings alone. §5.1 and Appendix A.8 report 75.6% category agreement (kappa 0.63, 0.42–0.81) and 97.8% fabricated-versus-not agreement (kappa 0.94, 0.78–1.00), superseding the stratum-weighted 73%/94%. (2) Sheet A coded by the same coder, with agreement against the first adjudicator, the blind agent and the author: resolved (Appendix A.8; the coder's role is recorded in §4). (3) Bounds on detected counts and robustness: resolved. The A.8 relabelling table gives 549 (493–638) fabricated references and 45.2% of reviewed submissions affected; §5.3 gives flag sensitivity 0.83; §5.2 gives 4 (0–13) accepted papers with an invented reference. Table 2b and Appendix A.3 are recomputed with the coder's labels (13.5%, 6.2%, 7.8%). (4) Qualification of the §4 independence claim: resolved in §4, §5.1, §7 and the Appendix A.7 heading. (5) The §6 'kept out' sentence: resolved; it now states an association. (6) Repository contents: stated in the response letter but unverifiable from the supplied evidence (not_enough_evidence). This is reproducibility documentation, not a validity defect in the reported analysis. New issues: residual stale corruption figures remain in §§5.1, 5.2, 6 and 7, and the Data availability section omits the new coder files. All are editorial; the abstract and Table 2b are correct. The competing interest (the author operates the venue) remains disclosed.
Published with the paper; the authors were not required to address them before publication.
No new reviewer reports were commissioned for this round. The previous round asked for a minor revision, which the editor checks alone against the concerns it raised.
We thank the editor and the three reviewers. The reopened round-5 decision set one blocking condition, coding of the released reference-strings-only random sample by a human with no role in the study, and listed six items. This letter follows the six items and then the reviewers' remaining points.
Both sheets were coded by a person with no role in the study, recruited by the author from personal acquaintance. The coder's statement, data/adjudication/blind/CODER_STATEMENT.md, gives the coder's initials (JB), reports no relationship to the study, its author or Publish.fun, and confirms that only the sheets were opened, that no AI assistant was used, that no item was discussed with the author before the coding ended, and that the work took about 30 minutes for sheet B and 40 minutes for sheet A. The sheets contain only the reference strings; no agent label or author label was ever available to the coder. Sheet B carries a link or the searches run for 43 of its 45 items; sheet A carries a page, identifier or search link for 62 of 64 items and a note for 59.
Agreement with the first adjudicator on the 45 items (data/dataset/independent_agreement.md):
By first-adjudicator category: NOT_FOUND (n=20): coder NOT_FOUND 14, EXISTS_CORRUPTED 6. EXISTS_CORRUPTED (n=14): coder EXISTS_CORRUPTED 10, NOT_FOUND 4. EXISTS or web resource (n=9): coder EXISTS 8, EXISTS_CORRUPTED 1. PLACEHOLDER (n=1) and UNADJUDICABLE (n=1): agreed. The single fabricated-versus-not disagreement is therefore one entry the adjudicator accepted and the coder called corrupted; the ten category disagreements lie inside the fabricated class, between "invented" and "corrupted", and run in both directions. Against the blind agent, on the 11 items that were also blind re-adjudicated: 81.8% by category and 90.9% for fabricated-versus-not.
The directly estimated figure (97.8% for fabricated-versus-not; 75.6% by category) supersedes the stratum-weighted 94% and 73% estimates in Section 5.1, Section 7 and the appendix, where it is reported in a new Appendix A.8; the stratum-weighted figures remain only as the figures for the author's own coding in Appendix A.7. The abstract now reports the independent figure.
Materials. Sheet B is the simple random sample drawn with seed 20261005 from the 857 manual decisions; the generator, code/make_coder_sheets.py, is in the repository and regenerates both released sheets byte for byte. The scoring script, code/independent_agreement.py, reports raw agreement with Wilson intervals and Cohen's kappa with bootstrap intervals, and recomputes the detected headline values from the released data as a self-check (they match paper/results_tables.md exactly).
The same coder coded the label-free 64-item sheet (data/dataset/independent_agreement.md, Appendix A.8). Agreement on the 64 items, by category and for fabricated-versus-not, with 95% Wilson intervals for agreement and bootstrap intervals for kappa:
On the 25 manual decisions where the two agents disagreed, the coder sided with the first adjudicator in 6, with the blind agent in 14 and with neither in 5 (the author: 9, 13 and 3). On the 19 automated matches that the blind agent had called corrupted or invented, the coder confirmed a defect in 15 and accepted four as correctly cited: the entry the author had also accepted, and three real works that the blind agent and the author had called corrupted for a wrong identifier, author list or venue. The coder confirmed all 10 automated matches that the blind agent had confirmed. The sheet is boundary-enriched by design, so these figures are not an unbiased agreement rate; the random sample in item 1 is.
The coder never had access to any agent label or to the author's labels (item 1); the coder's role and relationship to the study are recorded in the statement and in Section 4. The author's coding remains presented as the author's own, non-independent check (Appendix A.7).
Fabricated-versus-not agreement on the random sample (97.8%) is not materially lower than the 94% previously extrapolated, so no corrected estimate replaces the detected figures. We nevertheless report the bound the editor asked for. The scoring script estimates from sheet B the coder's label given the first adjudicator's label (three states: invented, corrupted, other; Dirichlet posterior with a Jeffreys prior) and relabels every one of the 857 manual decisions accordingly in 4,000 simulations; automatically verified entries are unchanged, since their error rate is measured separately by the blind check. Median and 95% interval of the relabelled value against the detected value:
| Quantity | Detected | Under the coder's labels |
|---|---|---|
| Detected fabricated references (all submissions) | 513 | 549 (493-638) |
| Detected invented references | 286 | 271 (192-361) |
| Reviewed submissions with >=1 fabricated reference | 37.8% | 45.2% (38.6-56.0) |
| Reviewed submissions with >=1 invented reference | 23.7% | 29.0% (22.8-39.8) |
| Reviewed references fabricated | 7.2% | 7.7% (6.9-8.9) |
| Reviewed references invented | 3.8% | 3.7% (2.6-5.0) |
| Organiser flag, paper-level sensitivity | 0.93 | 0.83 (0.74-0.92) |
| Organiser flag, paper-level specificity | 0.66 | 0.66 (0.63-0.69) |
| Organiser flag, paper-level kappa | 0.54 | 0.49 (0.39-0.55) |
| Flagged examples fabricated (precision) | 51.9% | 51.9% (47.3-55.0) |
| Spearman rho, fabricated share vs LLM slots 1, 2, 3 | -0.15, -0.13, -0.13 | -0.14, -0.12, -0.11 (all intervals exclude zero) |
| Spearman rho, invented share vs LLM slots 1, 2, 3 | -0.20, -0.18, -0.18 | -0.15, -0.14, -0.13 (all intervals exclude zero) |
| Spearman rho, fabricated and invented share vs human expert score | -0.28, -0.41 | -0.23 (-0.32 to -0.12), -0.29 (-0.44 to -0.11) |
| Accepted papers with >=1 invented reference | 0 | 4 (0-13) |
| Accepted papers with >10% fabricated references | 0 | 1 (0-5) |
Where the coder and the adjudicator differ, the coder's reading is stricter, so the detected figures remain lower bounds under the coder's labels as well. Three consequences are now stated in the manuscript. Section 5.1 reports the relabelled headline values. Section 5.3 reports the flag metrics under the coder's labels: the paper-level sensitivity falls from 0.93 to 0.83 because the relabelling adds fabricated papers the flag missed, while specificity and precision are unchanged. Section 5.2 reports that a median of 4 (0-13) accepted papers would carry a reference classed as invented under the coder's labels, mostly through the invented-versus-corrupted boundary, so the absence of a detected invented reference among accepted papers is a statement about the adjudication's labels rather than a label-independent fact; the Section 5.5 acceptance result is stated as detected, and the score associations are attenuated but keep their sign, with all intervals excluding zero. Table 2b revised. Sheet A includes the 29 automatically verified entries whose labels anchor the blind-check error rates. The coder's labels differ from the author's on three of them (the three real works above, which the coder accepted as correctly cited), so, as the decision requested, the per-source rates and every downstream figure were recomputed with the independent coder's labels in place of the author's (data/adjudication/blind/coder_overrides.json; data/dataset/blind_checks_coder.md; data/dataset/results_tables_coder_override.md). The corrupted share among automatically verified entries falls from 8.2% (author-anchored) to 6.2% (source-weighted; 11 of 180 crude, 6.1%); the invented share is unchanged (1.7%; 4 of 180). Appendix A.3 now shows the per-source counts under all three label sets (blind agent, author, independent coder), and the Table 2b caption gives the author-anchored figures as a comparison. Every changed figure (old -> new):
Unchanged: all detected figures (513, 286; 37.8%, 23.7%; 7.2%, 3.8%), the flag metrics, the score and acceptance associations, the agent-agent agreement, the invented rates (1.7%; 6.2%, 2.6%) and the projection without the 'other' stratum (29, 9-65). The conclusions do not move: the adjusted fabricated shares fall by about two points and the accepted papers remain about half as affected as the rejected ones.
Revised in the abstract, Section 4, Section 5.1, Section 7 and the Appendix A.7 heading. The author's coding is presented as a disclosed, non-independent, attestation-blinded check, and the sentence claiming independence "in the sense that matters for inter-rater validation" has been removed.
In the human-coding context the word "independent" is now used only for the external coder. The blind agent instances are still called independent instances of the agent, which describes their separation from the first adjudicator, not human validation.
Reworded as an association: "Whatever the mechanism, no submission with a detected invented reference was accepted; the observational design of Section 5.5 shows an association with rejection, not that the venue's process (three LLM reviews, human expert review of the top-scoring papers, an automated reference flag and human program chairs) excluded those submissions because of their references. Corrupted citations of real works were accepted at a rate near that of our own automated stage."
21 of 111 blocks of the round 6 manuscript are new or changed since round 5; 17 of 107 blocks of round 5 no longer appear. A block is a paragraph, heading, list, table or code block; a re-worded paragraph counts whole and a moved one does not count. This compares the two submitted texts and is not the response letter.
| Section | Changed | Removed |
|---|---|---|
| Abstract | 1 | 1 |
| 4. Methods | 1 | 1 |
| 5.1 Corpus and parsing yield | 2 | 2 |
| 5.2 Prevalence of fabricated references (Q1) | 3 | 3 |
| 5.3 How precise was the organisers' automated flag? (Q2) | 2 | 2 |
| 5.7 What the fabrications look like (Q6) | 1 | 1 |
| 6. Discussion | 2 | 2 |
| 7. Limitations | 2 | 2 |
| A.3 Blind check of automatically verified entries, by verification source (corrupted / invented / sampled, under each label set) | 2 | — |
| A.7 Human coding of 64 items by the author (a disclosed, non-independent, attestation-blinded check): agreement with the agent labels | 1 | — |
| A.8 Independent human coding of the random sample of manual decisions (sheet B) and of the 64-item sheet (sheet A): agreement and sensitivity bound | 4 | — |
| A.3 Blind check of automatically verified entries, by verification source | — | 2 |
| A.7 Human coding of 64 items (the author as coder, no part in the adjudication): agreement with the agent labels | — | 1 |
[under "Abstract"] Agents4Science 2025 was the first conference to require an AI system as the first author of every submission and to review every complete submission with three large-language-model (LLM) reviewers. Its organisers' automated reference check reported that 56% of submissions contained at least one reference that could not be verified. We re-examined all 6,849 references in the 304 submissions with a parsable reference list, using a reproducible pipeline (DOI, arXiv and URL resolution; Crossref, OpenAlex, Semantic Scholar, OpenLibrary and Google Books) followed by manual adjudication of every reference it could not verify (857 decisions with logged evidence) under a protocol fixed in advance. The labels were re-adjudicated blind by independent agent instances (150 decisions: category agreement 83%, kappa 0.75; fabricated-versus-not 92%, kappa 0.83); an independent human coder agreed with the adjudication on 98% of a random sample of 45 manual decisions for fabricated-versus-not (kappa 0.94); among 180 automatically verified entries, about 6% were real works cited with a wrong author list, venue or identifier and about 2% did not exist, so every adjudicated rate below is a lower bound. Among the 241 reviewed submissions with references, adjudication found a wholly invented reference in 23.7% (95% CI 18.7-29.4) and a fabricated reference (invented, or a real work with a corrupted title, author list, venue, year or identifier) in 37.8% (31.9-44.0); 3.8% of their 5,230 references were invented and 7.2% fabricated, rising to an estimated 6.2% (4.7-8.8) and 13.5% (10.8-17.4) once the errors found among automatically verified entries are added. Accepted papers cite fewer: no invented reference was detected among their 1,308 references (an estimated 34, 14-70, expected undetected) and nine corrupted ones were detected in eight of 48 papers, with an adjusted fabricated share of 7.8% (4.7-12.0) against 15.4% for rejected submissions. The organisers' flag was a screen, not a measure: 51.9% of the example references it flagged were fabricated, its paper-level specificity against detected fabrication was 0.66 (sensitivity 0.93), and none of the 26 flagged examples in accepted papers that we could match was fabricated. Fabrication was associated with lower scores from all three LLM reviewers (Spearman rho -0.13 to -0.15 for the fabricated share, -0.18 to -0.21 for the invented share), with lower human expert scores (rho -0.28 and -0.41) and with rejection: no paper with more than 10% fabricated references, and none with a detected invented reference, was accepted. Yet an LLM review asserted on its own that references were fabricated in only 4 of the 91 affected papers, all by the reviewer slot identified as Gemini 2.5 Pro, and no human expert review did. Of the 513 detected fabricated references, 56% were invented; 32 invented references carried a DOI or arXiv identifier. All code, cached API responses, adjudication logs and validation files are public. [under "4. Methods"] [block of 3675 characters not shown; it begins "**Reliability of the labels.** Two blind checks were added in revision. (i) A se…"] [under "5.1 Corpus and parsing yield"] [block of 3520 characters not shown; it begins "**Reliability of the labels.** On the 150 blind re-adjudicated decisions the two…"] [under "5.1 Corpus and parsing yield"] [block of 2708 characters not shown; it begins "**What the automated stage let through.** On the 180 blind-checked automatically…"] [under "5.2 Prevalence of fabricated references (Q1)"] **Table 2b. Adjusted reference-level estimates.** The corruption and invention rates for each verification source (Appendix A.3, with the 29 coded entries at the independent coder's labels) are applied to each submission's own automatically verified entries and added to the adjudicated counts; intervals are bootstrap 95% intervals over the per-source rates. The uniform-rate version gives 12.7% (10.8-17.3), 6.6% (4.5-11.5) and 14.8% (12.9-19.3) for the fabricated share of the three groups; with the author's labels for the 29 coded entries the composition-aware figures are 15.2%, 9.3% and 17.2% for fabrication and 6.2%, 2.6% and 7.4% for invention, and with the blind-agent labels alone 16.5%, 10.5% and 18.5% and 5.4%, 1.9% and 6.6%. [under "5.2 Prevalence of fabricated references (Q1)"] | Group | Automatically verified entries (Crossref title matches) | Fabricated share, detected | Fabricated share, adjusted | Invented share, detected | Invented share, adjusted | Expected undetected invented references | |:--|--:|--:|:--|--:|:--|:--| | Reviewed | 4,638 (65%) | 7.2% | 13.5% (10.8-17.4) | 3.8% | 6.2% (4.7-8.8) | 122 (48-260) | | Accepted | 1,243 (55%) | 0.7% | 7.8% (4.7-12.0) | 0.0% | 2.6% (1.0-5.4) | 34 (14-70) | | Rejected | 3,395 (69%) | 9.4% | 15.4% (12.8-19.3) | 5.1% | 7.3% (6.0-9.9) | 88 (34-189) | [under "5.2 Prevalence of fabricated references (Q1)"] [block of 3286 characters not shown; it begins "The adjustment changes the reading of the accepted papers. By adjudication they…"] [under "5.3 How precise was the organisers' automated flag? (Q2)"] [block of 1873 characters not shown; it begins "The precision of the flag at the reference level, against detected labels, is th…"] [under "5.3 How precise was the organisers' automated flag? (Q2)"] [block of 1928 characters not shown; it begins "At the paper level, among the 237 reviewed submissions that received a check and…"] [under "5.7 What the fabrications look like (Q6)"] [block of 2446 characters not shown; it begins "Wholly invented references are usually plausible in form: real-sounding authors,…"] [under "6. Discussion"] [block of 2189 characters not shown; it begins "**Prevalence.** By adjudication 7.2% of the references of reviewed AI-first-auth…"] [under "6. Discussion"] [block of 1529 characters not shown; it begins "**A note on method.** This audit was itself performed by AI agents, from pipelin…"] …and 9 more blocks not shown.
[under "Abstract"] [block of 2993 characters not shown; it begins "Agents4Science 2025 was the first conference to require an AI system as the firs…"] [under "4. Methods"] [block of 2398 characters not shown; it begins "**Reliability of the labels.** Two blind checks were added in revision. (i) A se…"] [under "5.1 Corpus and parsing yield"] [block of 1802 characters not shown; it begins "**Reliability of the labels.** On the 150 blind re-adjudicated decisions the two…"] [under "5.1 Corpus and parsing yield"] [block of 2023 characters not shown; it begins "**What the automated stage let through.** On the 180 blind-checked automatically…"] [under "5.2 Prevalence of fabricated references (Q1)"] **Table 2b. Adjusted reference-level estimates.** The corruption and invention rates for each verification source (Appendix A.3, with the 29 human-coded entries at their human labels) are applied to each submission's own automatically verified entries and added to the adjudicated counts; intervals are bootstrap 95% intervals over the per-source rates. The uniform-rate version gives 14.5% (12.1-19.6), 8.5% (5.9-13.9) and 16.5% (14.1-21.5) for the fabricated share of the three groups; with the blind-agent labels alone the composition-aware figures are 16.5%, 10.5% and 18.5% for fabrication and 5.4%, 1.9% and 6.6% for invention. …and 12 more blocks not shown.