ChatGPT cited a real paper that says something else
Nobody checks thirty references. A marker reads one sentence that sounds stronger than the literature should allow, clicks its citation and opens the paper, which is why this risk sits on your boldest claims instead of spreading evenly down the list. A model reaches for a reference that fits the topic, and fit is not support. Existence is cheap to check and support is not, so the tools are thickest exactly where the danger is thinnest.
This is the harder version of the fabricated-citation problem and almost nobody writes about it. A made-up reference dies the moment someone searches the title. A real reference attached to the wrong claim passes every automated check ever built, sits in your bibliography looking impeccable, and falls apart only when a human being opens the paper. Markers do open papers.
Our guide to verifying a citation names this as one of four ways a reference fails. This page is the long version of that one, because it is the failure that survives verification.
The failure has a name
Researchers call it a quotation error: a reference that fails to substantiate the claim it is cited for. It is a well-studied problem, and it existed long before chatbots. Baethge and Jergas's 2025 systematic review and meta-analysis in Research Integrity and Peer Review pooled 46 studies covering around 32,000 references in medical literature and found 16.9% of quotations were incorrect, with 8.0% classed as major errors, meaning the source did not support the claim at all. Their meta-regression found no improvement over time. A study of high-impact general science journals published in the Royal Society's Proceedings A put the total quotation error rate at around 25%.
Those are peer-reviewed papers by professional researchers. Roughly one citation in six in published medicine does not say what it was cited for. The point is not that everyone is careless; it is that this failure is invisible by design, so nothing catches it unless a person deliberately looks.
Why AI makes it worse
Three distinct mechanisms, all documented.
Summarising strips the qualifiers. Peters and Chin-Yee tested ten leading models across nearly 5,000 summaries and published the results in Royal Society Open Science. Models routinely stated findings more broadly than the original texts, with some overgeneralising in 26 to 73 percent of cases, even when explicitly prompted for accuracy. Compared directly against human-written summaries, LLM summaries were around five times more likely to contain broad generalisations. Newer models tended to do worse, not better. The mechanism is simple: the hedges are the boring part, and a summary drops the boring part. "In a sample of 43 undergraduates, sleep restriction was associated with lower recall scores" becomes "sleep deprivation impairs memory."
Attribution drifts even with retrieval. Turning on web search fixes fabrication, not accuracy. The Tow Center for Digital Journalism tested eight AI search tools on their ability to identify the source of excerpts they were shown, and found incorrect answers to more than 60% of queries, with tools frequently citing the wrong article and doing so with high confidence and no hedging. If a system given the text struggles to attribute it correctly, a system working from memory has no chance.
The nearest plausible paper wins. When a model needs a citation for a sentence, what it produces is a reference that fits: right field, right topic, right kind of journal. Fit is not support. This is the same prediction machinery that produces entirely fabricated citations, except that here it lands on something real, which makes it more dangerous rather than less.
The four shapes it takes
Recognising the pattern makes checking faster, because you learn where to look first.
Scope creep
The study found something in 60 undergraduates at one university. Your sentence says it about students. Or adults. Or people. This is the most common shape by far, and the fix is usually to narrow the sentence rather than to find a different paper.
Correlation promoted to cause
The paper reports an association, adjusts for confounders, and hedges carefully in the discussion. Your sentence says one thing causes the other. Check the study design before you check anything else: an observational study cannot support a causal claim no matter how strong its numbers are.
The claim belongs to a different paper
You find your sentence in the source, word for word almost, and it is in the introduction, with its own citation attached. You have cited a paper that was itself citing someone else. Go to that original and cite it instead, after checking that it says what this paper claims it says. Passing a claim along a chain without opening the first link is how errors propagate through a literature for decades.
Right topic, opposite finding
Less common, more embarrassing. The paper is about exactly your topic and reports the opposite of what you claimed, or reports a null result. This happens when a model recognises a famous paper's topic without recalling its conclusion. It is also worth checking whether a source has been contradicted since: scite classifies citing statements as supporting, contrasting or mentioning, which is a fast way to see whether a finding held up.
How a marker actually finds it
Not by scanning your bibliography. Nobody checks thirty references. What happens is narrower and much likelier than students assume.
A marker reads a sentence that surprises them. Either it is stronger than they expected the literature to support, or it is in their own subfield and they recognise the paper. They click one citation. That is the whole mechanism, and it is why the risk is not spread evenly across your reference list: it concentrates on your boldest claims, which are exactly the ones a model was most likely to overstate. The sentence that made your argument feel strong is the sentence that gets checked.
Two other routes are common. Supervisors reading a thesis chapter check the sources for anything they intend to build on, because they will have to defend it later. And once one reference has failed for any reason, including a broken DOI, the rest of the list gets read properly. That is why a fabrication accusation so often arrives as "several of your sources," even when only one was invented: someone started checking and kept going.
What they see when they open the paper is exactly what you would see. There is no ambiguity to argue about, no accuracy debate, no question of interpretation as there is with a detector score. The paper either says it or it does not.
The 90-second check
Per citation, in order:
- Read the abstract's results sentence, not the conclusion sentence. The results line carries the sample and the effect. The conclusion line is the authors' own generalisation and is already one step removed.
- Find the sample. Who, how many, where, when. Compare that against the population your sentence talks about. This alone catches most scope creep.
- Search the full text for your key term. Open the PDF or HTML and use Ctrl+F on the specific word your claim turns on. If the paper never uses it, that is your answer.
- Copy one supporting sentence, verbatim, with a page or section reference. Put it in a comment or footnote next to the citation. This is the whole method. Writing the quote down is what forces you to confirm it exists.
- Check whether that sentence is itself a citation. If it is, follow it to the original.
Keep the quotes. They cost nothing to store and they are the fastest possible answer if anyone ever questions your work, which is the situation an accusation of fabricated sources puts you in. They also make revision easier: when you rewrite a paragraph six weeks later, the quote tells you instantly whether the citation still fits the new sentence.
Three tests you can run mechanically
Part of the support question is not a reading question at all, and doing the mechanical part first tells you where to look. These three are the probes we run inside CiteOwl over the quote attached to every citation, and each one works just as well by hand.
The number test. If your sentence carries a figure, that figure has to appear in the sentence you are quoting, not merely somewhere in the paper. Check every number separately. A claim carrying two figures where the quote backs only one is still unsupported, and the number that did match is exactly what stops most people looking at the one that did not. Allow about a percent either way for rounding, and skip anything that looks like a year: a four-digit number between 1900 and 2100 sitting in a citation is almost always bibliographic rather than evidence.
The content-word test. Take your claim and your quote, throw away the words under four letters and the ordinary connectives, and count how many of the claim's remaining words survive in the quote. Near zero means the quote is about a different subject, whatever the paper's title suggested. One caveat, if you write in more than one language: the test measures nothing across languages, since a claim in Slovak backed by an English quote shares no words at all by construction. We learned that the expensive way, watching the check flag every single citation in a non-English draft.
The title test. If the only line you can find to support the claim turns out to be the paper's own title, you have nothing. A title names a work, it does not evidence a claim. It is also the one string about a paper that is available without opening it, which is why it is the most common thing an invented quote turns out to be.
What none of the three can settle is whether an observational finding supports a causal sentence, or whether an effect measured in one group carries to the group your sentence talks about. Those two need the reading, and they are the shapes this failure takes most often.
"The citation is real" is not "the citation is correct"
These are two separate tests and they need different evidence. Existence is a database question: a DOI resolver, an index, a publisher's record, all of which a machine can answer in milliseconds. Support is a reading question, and no database holds the answer, because it depends on a sentence you wrote. That asymmetry is why the tooling around citations is so lopsided: everything on the market checks the easy half.
Ours included, and we would rather say so than let a green tick imply otherwise. CiteOwl's free citation checker takes a pasted reference list and reports, per entry, whether CrossRef or OpenAlex hold the work and whether the record they hold agrees with the authors, title and year you wrote down. That is the existence half. It has no opinion on whether the paper supports your sentence, because it has not read your sentence. Any product that claims to verify support is claiming to have read the source, and that claim is worth testing on a paper you already know well before you trust it on one you do not.
So a reference can pass every check in the verification routine, existence, metadata, retraction status, and still be wrong in the only way that matters. The checks are worth running, and they are not the end of the job.
Academic integrity policies already treat the two as related failures. Virginia Tech's definitions of academic misconduct list citing a source for a proposition it does not support alongside inventing a source outright, and Iowa State's integrity tutorial defines fabrication to include attributing to a source ideas and information the source does not contain. From a marker's side of the desk, the distinction between a fake paper and a real paper cited for something it never said is thin: both mean the sentence has no evidence behind it.
The professional world has run into the same thing. When Deloitte Australia partially refunded a government contract in October 2025 over AI-generated errors in a report, the problems included both references to research that did not exist and a fabricated quote attributed to a real court judgment. The second kind is the one that is harder to spot and harder to explain. Quoted words fail differently from a paraphrase, and a quote the paper never contained can be checked against the page in a way a summary cannot.
Which leaves this market a strange shape, and we think it is worth naming. Existence is cheap to check and the answer is not arguable, so every tool checks it. Support is expensive to check and the answer can be argued about, so almost nothing does. The tooling is therefore thickest exactly where the danger is thinnest: a fabricated reference is the failure most likely to be caught by software and least likely to reach a marker, while a real paper cited for a claim it does not make passes every automated check on the market and lands in front of a person. So if you have time for one manual pass over a finished draft, do not re-run the DOIs. Machines are already good at that. Open your three boldest sentences and go and find the quote.
When nothing supports your sentence
Sometimes you check honestly and find that no source says what you wrote. That is useful information, not a research failure. Three legitimate moves, in order of preference:
Narrow the claim to what the evidence shows. "Sleep restriction reduced recall in a small sample of undergraduates" is a real sentence with real backing. It is also more interesting than the vague version, because it is specific.
Attribute instead of asserting. "One 2023 study of 43 undergraduates found…" puts the claim's weight where it belongs and lets the reader judge.
Cut it. A paragraph rarely misses the sentence you could not support. If the whole argument depends on it, that dependency was a problem whether or not you found a citation.
What not to do is keep hunting until you find a paper that can be made to fit. That is the process that produces quotation errors in the published literature, and it is easy to do accidentally when you are tired. If you want a structured way to search for real backing, we cover it in how to find a source for a claim you already wrote.
Writing so this doesn't happen
The structural fix is to reverse the order: find the source first, read it, and write the sentence the source supports. Most citation problems come from writing the sentence you wanted and then shopping for evidence. Starting from real papers you have actually read costs more time up front and removes an entire category of failure later.
If you do use a chatbot to gather sources, making it search rather than recall reduces fabrication. It does not touch this problem. Retrieval improves whether the paper exists; only reading tells you whether it supports your sentence.
A sensible order on a finished draft is to clear the mechanical half first so your attention is free for the half that needs you. Paste the reference list into the citation checker and let it tell you which entries the indexes cannot find and which ones carry a record that disagrees with what you wrote. Chase any single suspect identifier through the DOI lookup to see exactly which work it is registered to. Both are free and need no account. Then close them, open your three boldest sentences, and read.