CiteOwl
CiteOwl

How common are fake AI citations?

Fabricated citations are still rare in published papers, and they are getting less rare fast. In the largest audit of the published record, one paper in 2,828 carried a fabricated reference in 2023; by the first seven weeks of 2026 it was one in 277. A separate 2026 audit of four large repositories put a conservative floor of 146,932 hallucinated citations under work published during 2025 alone, running from 0.21% of references on bioRxiv to 1.91% on SSRN. But those are the numbers for citations that survived all the way to publication, and they are not the number that applies to you: when researchers ask a chatbot directly for references, between 11% and 57% come back fake. This page keeps those two questions apart, because most of what is written about this blends them.

The short answer, in three numbers

Most arguments about this go wrong in the first sentence, because "how common are fake AI citations" is three different questions with three different denominators. Keeping them apart is the whole job.

The first two numbers are small and the third is not, and there is no contradiction between them. The first two describe citations that made it past a supervisor, an editor, a copy editor and several reviewers. The third describes what arrives in your chat window before any human being has looked at it. You are standing at the third number, not the first.

Here is the part that is awkward for a company that sells citation software: the honest headline is that fabricated references remain rare in the published literature. Roughly 999 papers in every thousand do not have one. We would rather lead with a scarier figure, and there are plenty of scarier figures circulating, but almost none of them trace back to a study anyone can read.

What the two 2026 audits actually counted

Two large audits landed within a day of each other in May 2026. They are constantly reported as one study, partly because both happen to cover about 2.5 million papers. They are separate pieces of work with different corpora, different methods and different headline numbers, so it is worth being precise about which is which.

The four-repository audit

A team from Cornell, UCLA, Tsinghua and UC Berkeley's Haas School of Business audited 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN and PubMed Central. This is a preprint and has not been peer reviewed. Its headline figure is a conservative estimate of 146,932 hallucinated citations in 2025 alone, and phys.org and Nature both covered it.

That number is not a count of 146,932 individually confirmed fakes, and the authors are careful about this. They match every reference against a database built from Semantic Scholar and OpenAlex, then treat the excess of unmatchable references above the pre-2023 baseline as the signature of hallucination. In their words, they use the term to mean the estimated excess above the pre-LLM baseline "rather than classifications of individual references." It is a population measurement, not a list of names.

The biomedical audit

The other audit, by researchers at Columbia University's School of Nursing and Data Science Institute with colleagues in Israel and Finland, was published as correspondence in The Lancet. It screened the PubMed Central Open Access collection, not the whole of PubMed, from January 2023 to February 2026. Of 97.1 million verified references, 4,046 in 2,810 papers were fabricated. STAT and Nature reported it, and it is this study, not the four-repository one, that produced the year-by-year trend everybody quotes.

The trend: one in 2,828 to one in 277

This is the sequence worth memorising, and it belongs to the biomedical audit alone. It measures the share of published papers carrying at least one fabricated reference, in one corpus, over three years.

Period Papers with a fabricated reference Fabricated references per 10,000 papers
2023 1 in 2,828 about 4
2025 1 in 458 51.3 by the fourth quarter
First seven weeks of 2026 1 in 277 56.9

Read down the middle column and the rate is a tenth of what it was three years ago at getting caught. Read the right-hand column and the rise is more than twelvefold. The four-repository preprint puts a date on the inflection: the steepest climb begins in mid-2024, roughly eighteen months after ChatGPT was released, which its authors attribute to the arrival of AI search and research agents that assemble citations automatically rather than to chat assistance alone.

The forward-looking part is the part most coverage left out. Reasoning models and retrieval-augmented systems were supposed to end this, and on the audit's own data they have not. The authors note that citation hallucinations continued to rise through their late-2025 cutoff "with no signs of plateauing."

Where fabricated citations cluster

They are not spread evenly. Two patterns matter if you want to know your own exposure: which repository the work sits in, and which venue it was accepted at.

Repository What it mostly holds Share of references hallucinated, August 2025
bioRxiv Life sciences preprints 0.21%
PubMed Central Peer-reviewed journal articles 0.27%
arXiv Maths, physics, computing preprints 0.39%
SSRN Social sciences, law, humanities 1.91%

SSRN runs at roughly five times the rate of the other three. The audit also found that the pattern is not a few bad actors: papers with a small proportion of unmatchable references grew far faster than papers stuffed with them, which is what you would expect if ordinary authors are accepting a couple of suggested references without checking rather than a handful of people inventing bibliographies wholesale. Authors who cite hallucinated references had 62% fewer prior publications than a matched control group on arXiv, and 73% fewer on SSRN.

Peer review at the top of the field does not change the picture much. A 2026 preprint from Microsoft researchers scanned 48,095 accepted papers and 2.6 million references across four leading computer science conferences, counting only identity-level failures, meaning either no matching work exists or the author list diverges substantially from the citation.

Venue, 2025 Share of all references hallucinated Accepted papers with two or more
ICLR 0.38% 1.9%
ICML 0.54% 3.4%
USENIX Security 0.81% 4.8%
NeurIPS 0.68% 5.1%

Roughly one in twenty accepted NeurIPS papers from 2025 carries at least two references that cannot be matched to any real work. Separately, the detection company GPTZero screened 4,841 of the accepted NeurIPS 2025 papers by hand and confirmed 100 hallucinated citations across roughly 51 of them, each having passed three or more expert reviewers. The two numbers differ because the thresholds differ, and both point the same way.

Nothing downstream is catching them

The obvious hope is that these get filtered out on the way to publication. The four-repository audit measured that directly, at two separate gates, and the answer is no at both.

At the first gate, arXiv moderation, manuscripts that get rejected do carry hallucinated references at 4.5 times the rate of accepted ones, so screening is catching something. It is not catching enough: the authors estimate that 78.8% of non-existent citations pass moderation and appear on the platform anyway.

At the second gate, journal peer review, they traced 2,241 bioRxiv preprints containing unmatchable references through to their published versions.

Of the fabricated references already sitting in a preprint, 85.3% were still there in the published version. Peer review removed about one in seven.

Nor is this confined to weak journals. Stratifying biomedical journals into deciles by impact, the audit found the lowest and highest deciles behaved as you would expect but everything between them did not, with no consistent gradient and hallucinated citations appearing across the full impact spectrum. And once a fake reference is published, very little happens: as of February 2026, over 98% of the flagged articles had seen no publisher action at all.

That is the fact worth carrying away from this section. There is no institution downstream of you reliably removing these. If a fabricated reference is in your draft, the overwhelmingly likely outcome is that it is still there when the work is read.

Why every number here is a floor

Every audit on this page counts one specific failure: a reference whose title matches no real work. The four-repository team says so outright, focusing on "those with non-existent titles." The conference audit restricts itself to identity-level failures and explicitly excludes ordinary drift in venue, year or publication status. That is a defensible choice, because a non-existent title is the one thing you can check at scale without judgement. It also means all of them are blind to the same failure.

A citation can name a real paper and attach the wrong identifier to it. In our own August 2026 run, we asked current models for references on ten ordinary undergraduate topics and resolved every DOI against CrossRef. For GPT-5.6, 7% of DOIs were dead, and 12% resolved perfectly to a completely different paper. A bad DOI from that model was almost twice as likely to work as to fail. One reference gave "Road Pricing: Lessons from London" a live DOI belonging to an economics paper titled "Has the inflation process changed?"

None of the audits above would count that. The title exists, the paper is real, the identifier resolves, and a matching pipeline sees a valid reference. So treat every figure on this page as a floor rather than a measurement, and note that the standard advice to check whether a DOI resolves is aimed at the failure that is becoming the less common one. We cover the identifier version of this in why AI makes up citations, the harder version where a real source is cited for a claim it does not support in ChatGPT cited a paper that says something else, and the check itself in how to check if a citation is real.

The rate in the text you are handed

All of the above measures citations that survived to publication. The rate in raw model output is an order of magnitude higher, and it is the number that describes your situation when you paste a reference list into a draft.

Study Model or models Sample Fabrication rate
Walters & Wilder, 2023 GPT-3.5 and GPT-4 636 references, 42 topics 55% and 18%
Linardon et al., 2025 GPT-4o 176 citations 19.9%
Cabezas-Clavijo & Sidorenko-Bautista, 2026 Eight free-tier chatbots 400 references 39.8% wrong or fabricated
Naser, 2026 (preprint) Ten small and mid-tier models 69,557 citation instances 11.4% to 56.8%
CiteOwl, August 2026 GPT-5.6 100 references with a DOI 19% unusable

The Naser preprint is the largest of these and the most useful for one reason: it shows the spread is not random. Every model it tested is a small or cheap tier, so the range is not a frontier-model range, but within it the ordering is stable and the fivefold gap between best and worst is real. It also found the rate tracks how well a field is represented in training data, running at 26.6% for machine learning topics, 41.8% for climate and environmental science, and 50.1% for structural engineering. The narrower and less online your subject, the worse your personal rate runs above any published average.

So the answer to how common fake AI citations are depends entirely on where you stand. In the published record they are rare and rising sharply. In the draft on your screen they are common enough that a ten-item reference list from a general chatbot will usually contain at least one entry that does not survive checking. The gap between those two facts is made of people checking, and for your bibliography that person is you. If you want the mechanism behind all of this rather than the counts, that is why AI makes up citations; if you want the references to be real in the first place, making the model retrieve before it writes is the lever that actually moves.

Things worth knowing.

How many fake AI citations are there?
The largest audit to date estimated at least 146,932 hallucinated citations in papers and preprints published during 2025, across arXiv, bioRxiv, SSRN and PubMed Central. A separate audit of biomedical papers found 4,046 fabricated references across 2,810 papers between January 2023 and February 2026. Both figures are floors rather than totals, because both count only references whose titles match no real paper, and neither can see a citation that points at a real paper that is not the one being cited.
What percentage of published papers contain a fabricated reference?
In the first seven weeks of 2026, about one published paper in 277 contained at least one fabricated reference, up from one in 458 in 2025 and one in 2,828 in 2023. At the level of individual references rather than papers, the rate as of August 2025 ranged from 0.21% on bioRxiv to 1.91% on SSRN, the repository that mostly holds social sciences, law and humanities work.
Are fake citations becoming more or less common?
More common, and quickly. The rate per ten thousand papers rose more than twelvefold between 2023 and early 2026. The sharpest rise began in mid-2024, roughly eighteen months after ChatGPT was released, which researchers link to the arrival of AI search and research agents that generate citations automatically. The authors of the largest audit note that rates were still climbing through their late-2025 data cutoff with no sign of levelling off, despite newer reasoning models and retrieval systems.
Which fields and venues have the most fabricated citations?
Social sciences and computer science lead. SSRN runs at nearly five times the rate of the other major repositories, and at leading machine-learning conferences roughly one in twenty accepted 2025 papers carried at least two references that could not be matched to any real work. Rates also track how well represented a field is in a model's training data: in one 2026 audit of ten models, fabrication ran at 26.6% for NLP and AI topics, 41.8% for climate and environmental science, and 50.1% for structural engineering.
Read next.

CiteOwl doesn't invent citations

Free to start. No card needed.

Start writing