CiteOwl
CiteOwl

How common are fake AI citations?

It depends where you are standing. In the list a chatbot hands you, between 11 and 57 percent of the references are fake. In the published record, one paper in 277 carried a fabricated reference in early 2026, up from one in 2,828 in 2023. Those got past a supervisor, an editor and several reviewers. Yours have not been past anyone, so run the list through the citation checker before someone else does.

The short answer, in three numbers

Most arguments about this go wrong in the first sentence, because "how common are fake AI citations" is three different questions with three different denominators. Keeping them apart is the whole job.

The first two numbers are small and the third is not, and there is no contradiction between them. The first two describe citations that made it past a supervisor, an editor, a copy editor and several reviewers. The third describes what arrives in your chat window before any human being has looked at it. You are standing at the third number, not the first.

Here is the part that is awkward for a company that sells citation software: the honest headline is that fabricated references remain rare in the published literature. Roughly 999 papers in every thousand do not have one. We would rather lead with a scarier figure, and there are plenty of scarier figures circulating, but almost none of them trace back to a study anyone can read.

What the two 2026 audits actually counted

Two large audits landed within a day of each other in May 2026. They are constantly reported as one study, partly because both happen to cover about 2.5 million papers. They are separate pieces of work with different corpora, different methods and different headline numbers, so it is worth being precise about which is which.

The four-repository audit

A team from Cornell, UCLA, Tsinghua and UC Berkeley's Haas School of Business audited 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN and PubMed Central. This is a preprint and has not been peer reviewed. Its headline figure is a conservative estimate of 146,932 hallucinated citations in 2025 alone, and phys.org and Nature both covered it.

That number is not a count of 146,932 individually confirmed fakes, and the authors are careful about this. They match every reference against a database built from Semantic Scholar and OpenAlex, then treat the excess of unmatchable references above the pre-2023 baseline as the signature of hallucination. In their words, they use the term to mean the estimated excess above the pre-LLM baseline "rather than classifications of individual references." It is a population measurement, not a list of names.

The biomedical audit

The other audit, by researchers at Columbia University's School of Nursing and Data Science Institute with colleagues in Israel and Finland, was published as correspondence in The Lancet. It screened the PubMed Central Open Access collection, not the whole of PubMed, from January 2023 to February 2026. Of 97.1 million verified references, 4,046 in 2,810 papers were fabricated. STAT and Nature reported it, and it is this study, not the four-repository one, that produced the year-by-year trend everybody quotes.

The trend: one in 2,828 to one in 277

This is the sequence worth memorising, and it belongs to the biomedical audit alone. It measures the share of published papers carrying at least one fabricated reference, in one corpus, over three years.

Period Papers with a fabricated reference Fabricated references per 10,000 papers
2023 1 in 2,828 about 4
2025 1 in 458 51.3 by the fourth quarter
First seven weeks of 2026 1 in 277 56.9

Read down the middle column and the denominator is a tenth of what it was three years ago, so a paper is roughly ten times as likely to carry a fabricated reference as in 2023. Read the right-hand column and the rise is more than twelvefold. The four-repository preprint puts a date on the inflection: the steepest climb begins in mid-2024, roughly eighteen months after ChatGPT was released, which its authors attribute to the arrival of AI search and research agents that assemble citations automatically rather than to chat assistance alone.

The forward-looking part is the part most coverage left out. Reasoning models and retrieval-augmented systems were supposed to end this, and on the audit's own data they have not. The authors note that citation hallucinations continued to rise through their late-2025 cutoff "with no signs of plateauing."

Where fabricated citations cluster

They are not spread evenly. Two patterns matter if you want to know your own exposure: which repository the work sits in, and which venue it was accepted at.

Repository What it mostly holds Share of references hallucinated, August 2025
bioRxiv Life sciences preprints 0.21%
PubMed Central Peer-reviewed journal articles 0.27%
arXiv Maths, physics, computing preprints 0.39%
SSRN Social sciences, law, humanities 1.91%

SSRN runs at roughly five times the rate of the other three. The audit also found that the pattern is not a few bad actors: papers with a small proportion of unmatchable references grew far faster than papers stuffed with them, which is what you would expect if ordinary authors are accepting a couple of suggested references without checking rather than a handful of people inventing bibliographies wholesale. Authors who cite hallucinated references had 62% fewer prior publications than a matched control group on arXiv, and 73% fewer on SSRN.

Peer review at the top of the field does not change the picture much. A 2026 preprint from Microsoft researchers scanned 48,095 accepted papers and 2.6 million references across four leading computer science conferences, counting only identity-level failures, meaning either no matching work exists or the author list diverges substantially from the citation.

Venue, 2025 Share of all references hallucinated Accepted papers with two or more
ICLR 0.38% 1.9%
ICML 0.54% 3.4%
USENIX Security 0.81% 4.8%
NeurIPS 0.68% 5.1%

Roughly one in twenty accepted NeurIPS papers from 2025 carries at least two references that cannot be matched to any real work. Separately, the detection company GPTZero screened 4,841 of the accepted NeurIPS 2025 papers by hand and confirmed 100 hallucinated citations across roughly 51 of them, each having passed three or more expert reviewers. The two numbers differ because the thresholds differ, and both point the same way.

Nothing downstream is catching them

The obvious hope is that these get filtered out on the way to publication. The four-repository audit measured that directly, at two separate gates, and the answer is no at both.

At the first gate, arXiv moderation, manuscripts that get rejected do carry hallucinated references at 4.5 times the rate of accepted ones, so screening is catching something. It is not catching enough: the authors estimate that 78.8% of non-existent citations pass moderation and appear on the platform anyway.

At the second gate, journal peer review, they traced 2,241 bioRxiv preprints containing unmatchable references through to their published versions.

Of the fabricated references already sitting in a preprint, 85.3% were still there in the published version. Peer review removed about one in seven.

Nor is this confined to weak journals. Stratifying biomedical journals into deciles by impact, the audit found the lowest and highest deciles behaved as you would expect but everything between them did not, with no consistent gradient and hallucinated citations appearing across the full impact spectrum. And once a fake reference is published, very little happens: as of February 2026, over 98% of the flagged articles had seen no publisher action at all.

That is the fact worth carrying away from this section. There is no institution downstream of you reliably removing these. If a fabricated reference is in your draft, the overwhelmingly likely outcome is that it is still there when the work is read.

The one place the checking reliably happens is the most expensive one. Damien Charlotin, a legal data analyst, keeps a public database of court decisions in which a party relied on AI-fabricated material and a judge dealt with it. When Mashable wrote it up in late May 2025, a few weeks after it opened, it held 120 entries. A July 2026 write-up put it at roughly 1,490 decisions worldwide, more than 1,000 of them in the United States, as of that May. On 19 September 2026 the database itself listed 2,044 cases, 1,397 of them American: 16 from 2023, 61 from 2024, 851 from 2025 and 1,116 so far this year. Every one is a filing somebody was supposed to have checked, and the checking got done by the judge instead. A marked essay works the same way with a smaller audience.

Why every number here is a floor

Every audit on this page counts one specific failure: a reference whose title matches no real work. The four-repository team says so outright, focusing on "those with non-existent titles." The conference audit restricts itself to identity-level failures and explicitly excludes ordinary drift in venue, year or publication status. That is a defensible choice, because a non-existent title is the one thing you can check at scale without judgement. It also means all of them are blind to the same failure.

A citation can name a real paper and attach the wrong identifier to it. In our own August 2026 run, we asked current models for references on ten ordinary undergraduate topics and resolved every DOI against CrossRef. For GPT-5.6, 7% of DOIs were dead, and 12% resolved perfectly to a completely different paper. A bad DOI from that model was almost twice as likely to work as to fail. One reference gave "Road Pricing: Lessons from London" a live DOI belonging to an economics paper titled "Has the inflation process changed?"

None of the audits above would count that. The title exists, the paper is real, the identifier resolves, and a matching pipeline sees a valid reference. So treat every figure on this page as a floor rather than a measurement, and note that the standard advice to check whether a DOI resolves is aimed at the failure that is becoming the less common one. We cover the identifier version of this in why AI makes up citations, the harder version where a real source is cited for a claim it does not support in ChatGPT cited a paper that says something else, and the check itself in how to check if a citation is real.

The rate in the text you are handed

All of the above measures citations that survived to publication. The rate in raw model output is an order of magnitude higher, and it is the number that describes your situation when you paste a reference list into a draft.

Study Model or models Sample Fabrication rate
Walters & Wilder, 2023 GPT-3.5 and GPT-4 636 references, 42 topics 55% and 18%
Mugaanyi et al., 2024 GPT-3.5 102 references, 10 topics in two fields 27% natural sciences, 23% humanities
Linardon et al., 2025 GPT-4o 176 citations 19.9%
Cabezas-Clavijo & Sidorenko-Bautista, 2026 Eight free-tier chatbots 400 references 39.8% wrong or fabricated
Naser, 2026 (preprint) Ten small and mid-tier models 69,557 citation instances 11.4% to 56.8%
CiteOwl, August 2026 GPT-5.6 100 references with a DOI 19% unusable

The Naser preprint is the largest of these and the most useful for one reason: it shows the spread is not random. Every model it tested is a small or cheap tier, so the range is not a frontier-model range, but within it the ordering is stable and the fivefold gap between best and worst is real. It also found the rate tracks how well a field is represented in training data, running at 26.6% for machine learning topics, 41.8% for climate and environmental science, and 50.1% for structural engineering. The narrower and less online your subject, the worse your personal rate runs above any published average. Mugaanyi and colleagues are the one study in the table that splits its sample by discipline rather than by topic, and the split is narrow: about a quarter of GPT-3.5's references did not exist, 27% in the natural sciences against 23% in the humanities.

So the answer to how common fake AI citations are depends entirely on where you stand. In the published record they are rare and rising sharply. In the draft on your screen they are common enough that a ten-item reference list from a general chatbot will usually contain at least one entry that does not survive checking. The gap between those two facts is made of people checking, and for your bibliography that person is you. If you want the mechanism behind all of this rather than the counts, that is why AI makes up citations; if you want the references to be real in the first place, making the model retrieve before it writes is the lever that actually moves.

Read next.

Finish with a reference list you don't have to take on trust.

CiteOwl is an editor whose AI cites only papers it opened, with the quote under each claim.

Start writing