The short answer
Never put a reference in a paper unless you have opened the source and confirmed that it exists and says what you cite it for. Look up each reference by DOI or title in Crossref, OpenAlex, PubMed or Google Scholar, check that the authors, year and journal match, and read the passage that supports your claim.
This is not a reason to avoid AI. Tools that search real databases and link every answer to a source make far fewer of these errors. It is a reason to keep checking, which is good practice with any source.
How often do AI tools invent references?
The rate depends on the tool and has dropped as models improved, but it has not reached zero for general chatbots.
- Walters and Wilder (2023, Scientific Reports): 55% of citations from GPT-3.5 and 18% from GPT-4 were fabricated. Among real citations, many still had errors.
- Chelli and colleagues (2024, JMIR): when asked to reproduce the reference lists of systematic reviews, GPT-3.5, GPT-4 and Bard produced hallucinated references at rates of about 40%, 29% and 91%.
- Linardon and colleagues (2025, JMIR Mental Health): about one in five citations from GPT-4o were fabricated, and fabrication was much higher on less familiar topics.
- A 2026 preprint testing ten commercial models across four academic fields found hallucination rates between about 11% and 57%, depending on model, field and prompt.
- Tools built on real databases do better. In an informal 2026 check, librarian Aaron Tay (Singapore Management University) looked at 50 references each from Consensus and Undermind and found no invented references, only some metadata errors. He did not test whether the papers supported the claims, and that is the remaining risk: a real paper cited for something it does not say.
Why AI invents references
A language model writes the most plausible next words. A reference has a predictable shape (authors, year, title, journal), so the model can produce something that looks exactly like a citation without retrieving one. This is most likely when you ask a general chatbot for sources from memory, on a niche topic, with no search connected.
Ghost references are not new. Aaron Tay points out that wrong and non-existent references existed long before chatbots, through copied reference lists and broken database records. AI has made them more common, and interlibrary loan staff report receiving requests for papers that turn out to be AI hallucinations.
A five-minute verification routine
Do this for every reference that came from an AI tool, or that you have not personally read.
- 1. Resolve the DOI at doi.org. A fake DOI fails or points to a different paper.
- 2. No DOI? Search the exact title in Crossref, OpenAlex, PubMed or Google Scholar. Crossref’s free Simple Text Query lets you paste a whole reference list and returns the DOIs it can match.
- 3. Check that the authors, year, journal, volume and pages match the record.
- 4. Open the paper and find the passage that supports your sentence. A real paper cited for the wrong claim is the most common error left.
- 5. Check for retractions. The Retraction Watch database is now free through Crossref, and Zotero warns you if an item in your library has been retracted.
What is at stake
Fabricated citations are treated as a serious integrity problem. In 2023 a US court sanctioned two lawyers and their firm in Mata v. Avianca for filing fake cases produced by ChatGPT. In 2025 Springer Nature retracted a machine learning book after it could not verify 25 of its 46 references. arXiv has said it will ban authors for one year when there is incontrovertible evidence that they did not check AI-generated content.
The EU living guidelines put it simply: tools that give citations are useful, but the final responsibility for each citation remains with the researcher.
How Kahubi prevents invented references
Kahubi’s agent only cites papers that exist in your library or that it found during the conversation in OpenAlex and Semantic Scholar. Each citation links to the paper, so you can open the PDF and check the passage. When your sources do not support a claim, the agent writes [CITATION NEEDED] instead of inventing a reference. You still read and check, but you start from real papers.
Sources
- Walters WH, Wilder EI (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13:14045.
- Chelli M, et al. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews: comparative analysis. Journal of Medical Internet Research 26:e53164.
- Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023). Opinion and order on sanctions.
- Linardon J, et al. (2025). Influence of topic familiarity and prompt specificity on citation fabrication in mental health research using large language models: experimental study. JMIR Mental Health 12:e80371.
- Naser MZ (2026). How LLMs cite and why it matters: a cross-model audit of reference fabrication in AI-assisted academic writing and methods to detect phantom citations. arXiv:2603.03299 (preprint).
- Tay A (2026). arXiv tightens policy on hallucinated references. SMU Libraries.
- Tay A (2025). Why ghost references still haunt us in 2025. Musings about librarianship.
- University of Colorado Boulder Libraries (2024). AI chatbots offer possibilities and perils for researchers seeking library resources.
- Retraction Watch (2025). Springer Nature retracts book with fake citations.
- Crossref. Retraction Watch data.
- Crossref. Simple Text Query.
- Zotero (2019). Retracted item notifications.
- European Commission (2026). Living guidelines on the responsible use of generative AI in research, third version.
Last updated 2026-10-09.