A controversial proposal for hallucinated references

 

Many prestigious scientific venues have suffered from papers containing hallucinated citations. GPTZero reports finding 100 hallucinations in NeurIPS 2025 accepted papers and 50 hallucinations in ICLR 2026 submissions. Hallucinated citations are references generated by an AI system that appear credible but point to sources that do not exist or do not support the claims attributed to them. They may contain plausible-looking but incorrect authors, titles, journals, dates, DOIs, or URLs, making them difficult to detect without thorough verification. Such citations, justifiably, undermine the confidence in both scientific publications and in the process that produces them. How should the scientific community address this issue?

Citation checks

One way is to ensure the correctness of published citations through technical means. Some advocate using tools to detect such citations when papers are submitted. Then, authors can be notified to fix them or the papers can be desk-rejected. Some even advocate disciplining the authors for violating scientific integrity, e.g. by placing their names in a list that might prevent them from submitting publications for a given time period. However, currently such checks are difficult and error-prone.

Rather than checking the text of submissions for errors, I think it’s more productive to adjust the document creation process to make it difficult for AI to add hallucinated citations and easy for authors, publishers, and PC chairs to check that the citations are correct. This can be done by specifying that each paper should embed in it machine-readable citation metadata obtained from the corresponding DOI registration agencies (RAs): the organizations authorized by the International DOI Foundation to register DOIs and maintain such data. If an AI system hallucinates such data, this can be trivially detected by comparing the submitted fields against those obtained from the RA via an API request. Tagging within the PDF each citation for which a DOI exists with that DOI makes it easy to verify the citation mechanically against authoritative registration metadata and to determine how many citations require manual verification (which is also a useful metric on its own). This check is quick and accurate, making it easy to integrate into the submission process. Of course, an AI system could obtain some correct metadata and then hallucinate the textual reference. Therefore, as a second line of defence, a tool can also provide a text similarity metric between the annotated citations and their metadata.

Such a mechanism can establish that a cited work exists and that its bibliographic description is accurate. It can’t, however, establish the much harder, and scientifically more important, property that the cited work supports the claim for which it is cited. So is this the way to go?

The suspect role of extensive citations

I’m skeptical regarding tools and processes aimed at pinpointing hallucinated references or even preventing their occurrence as I propose in the preceding section. What we’re dealing with in such cases is cargo cult science: sprinkling a study with references (correctly-cited or not) to give it a scientific appearance. Ensuring that the references’ metadata is correct is similar to a cult erecting more elaborate fake control towers and carving more convincing wooden headphones.

I believe the practice of having in each article tens of references, beyond what is strictly needed for attribution, evidence, and establishing novelty, reflects two needs that have become outdated. First, citations serve a gate-keeping role by having authors demonstrate scientific training and appropriate supervision, scholarship, attention to detail, and knowledge of existing work. Second, they serve the building of a citation network and the impact measures derived from it. (A cynic might argue that some conference venues provide two additional pages for references partly to boost citation-based impact metrics.)

The requirement for detailed references worked as a gate-keeping mechanism for excluding people who couldn’t produce them from publication venues, because it was easy to desk-reject such submissions. Extensive referencing was often mainly a costly signal that the author has studied and understood the relevant literature. A recent colleague’s statement that “hallucinated references are very much typos and do not substantively change the conclusions of the work” supports the view that many citations are perfunctory. Now with AI, this signal is gone: anyone can produce plausible references, and, increasingly, hallucinations will be dealt with through process adjustments, technical means, and better AI systems. But the ease and the trivial method for producing them (give me two citations to support claim XYZ) diminishes their value as a signal of scholarship. Once producing an impressive-looking bibliography becomes nearly costless, its size and apparent breadth cease to provide credible evidence of the effort or understanding that went into producing it.

Furthermore, if references are mainly obtained from AI, deriving impact measures from them is also a very roundabout and suspect way to judge impact. For example, currently committees for ten-year most influential paper awards start with a shortlist of highly-cited papers. In ten years, I’d feel more comfortable asking AI systems directly for a paper shortlist, rather than creating it via citation metrics derived from the number of references AI suggested to add to past papers. Similar arguments can be made regarding the (anyway fraught) practice of judging scientific worth through metrics such as citations and the h-index.

The scale and institutional importance of referencing increased substantially in the postwar period, driven by the rapid growth and specialization of the scientific literature, increasingly formal expectations concerning attribution and evidential support, the advent of citation indexes and bibliographic databases, and, later, citation-based research evaluation and digital tools that made references inexpensive to discover, manage, and follow.

However, excellent science can (and has) been produced without such references. While writing this text I went through Einstein’s seminal works on special and general relativity. The special relativity one contains absolutely no references in its 31 pages; the general relativity one (54 pages) contains just three footnote references to other work.

Based on the above, my proposal is to abolish the requirement for extensive and increasingly perfunctory — even when not hallucinated — references. Publication venues should de-emphasize in their review criteria the requirement for detailed examination of related work. Instead, a recorded, archived, and linked AI-assisted search of the scholarly literature, together with its queries, retrieved corpus, and analysis can likely provide more reliable evidence than many human-authored “related work” sections supported by AI-concocted references. In addition, venues that provide additional space just for references should do away with the pages allocated to them. Then, venues should judge submissions mainly on demonstrable substantial scientific significance, rejecting questionable marginal improvements over related work supported only through dense citations. Incidentally, such a policy should also reduce the number of published papers, which is also a desideratum when top AI systems can increasingly produce technically competent papers of little scientific significance.

In summary, venues should stop using extensive referencing as evidence of scholarship; require fewer, claim-relevant references; and replace ceremonial related-work sections with auditable evidence of literature search and novelty.

Comments   Post Toot! Share


Last modified: Monday, August 17, 2026 6:28 pm

Creative Commons Licence BY NC

Unless otherwise expressly stated, all original material on this page created by Diomidis Spinellis is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.