Citation fabrication rates in three chat models
We asked three models for ten peer-reviewed references on each of ten ordinary essay topics, in one plain prompt, with no tools and no web access. Then we resolved every DOI they supplied against CrossRef and sorted each reference into one of three buckets: the record is the work that was cited, the DOI resolves to nothing, or the DOI resolves to a different paper.
The failure everyone tests for is the smaller one. GPT-5.6 returned 100 references with DOIs, of which 7% were dead and 12% opened somebody else's paper. GPT-5.4-mini was worse on both counts, 25% dead and 16% wrong paper, leaving 41% of its references unusable. Claude Haiku 4.5 mostly declined to supply DOIs at all, offering 18 across the whole run. Checking that a DOI resolves therefore passes more bad references than it catches.
What it cannot tell you: one run on one date. Rates move between runs, so read these as an order of magnitude rather than a league table. Title-only references are reported but not scored.
Download the data (JSON, CC BY 4.0) Read the studyHow often a one-character DOI slip lands on a real paper
This is the study that explains the one above. We drew 600 journal articles at random from CrossRef's own sample endpoint, changed the last digit of each DOI up by one and down by one, and asked CrossRef whether either invented string was a registered work. All 1,186 lookups completed with no errors.
352 of the 600 had at least one live neighbour, and 350 of those sat in the same journal. DOI suffixes run in sequence rather than at random, so a typo does not usually produce a dead link. It produces a different real paper, published weeks apart, in the same volume. The file breaks the rate down across fifteen publishers and includes forty worked examples with both titles.
What it cannot tell you: this measures what a one-digit slip does, not how often students make one. It is a property of the identifier, not a rate of human error.
Download the data (JSON, CC BY 4.0) Read the studyHow many references a published paper carries, by field
Every guide answering "how many sources do I need" gives a number and cites nothing. So we measured the convention instead: OpenAlex journal articles from 2024 in core venues, in English, with the deposited reference-list length read from CrossRef, across 26 fields. The file carries the median, the quartiles, the tenth and ninetieth percentiles and the deposit rate for each field and document type, 52 rows in all.
Agricultural and Biological Sciences, to take one row, runs to a median of 51 references with a middle half between 39 and 68. The spread between fields is wide enough that any single number offered as a universal rule is wrong somewhere.
What it cannot tell you: these are published journal articles, not student papers. It describes professional convention, and a marker's expectation for an undergraduate essay is a different thing. Works with a zero deposited count are excluded as a metadata artefact.
Download the data (JSON, CC BY 4.0) Read the studyUniversities that turned off Turnitin's AI detection
Every institution named in any roundup of universities disabling Turnitin's AI writing indicator, checked against a page that institution publishes itself. An entry is verified only when the university's own page states it, and all 17 carry a link to that page. Where no such page could be found, the claim is not dropped and not repeated: it is recorded separately with the reason, which is why nine well-known names sit in an unverified list rather than in the table.
Roundups of this run on each other. This one runs on primary sources, and says out loud where it could not get one.
What it cannot tell you: a snapshot of what each institution published on the collection date. Policies change and pages move. Absence from the list is not evidence a university still runs the detector, only that we found no published statement either way.
Download the data (JSON, CC BY 4.0) Read the studyUsing any of this
All four files are CC BY 4.0. Reuse them, republish the figures, build on them, with attribution to CiteOwl and a link back to this page. If you re-run one of these and get a different answer, we would rather hear about it than not: every file names its method and its date precisely so that it can be checked.