CiteOwl

ChatGPT cited a real paper that says something else

Open the source and find the exact sentence, table or figure that supports your claim, then paste it verbatim into a note next to the citation. If you cannot find one after ninety seconds of searching the full text, the citation is wrong, even though the paper is real and the DOI resolves. That one habit, the quote test, catches the failure that reference checkers, DOI resolvers and plagiarism software all miss completely.

This is the harder version of the fabricated-citation problem and almost nobody writes about it. A made-up reference dies the moment someone searches the title. A real reference attached to the wrong claim passes every automated check ever built, sits in your bibliography looking impeccable, and falls apart only when a human being opens the paper. Markers do open papers.

Our guide to verifying a citation names this as one of four ways a reference fails. This page is the long version of that one, because it is the failure that survives verification.

The failure has a name

Researchers call it a quotation error: a reference that fails to substantiate the claim it is cited for. It is a well-studied problem, and it existed long before chatbots. Baethge and Jergas's 2025 systematic review and meta-analysis in Research Integrity and Peer Review pooled 46 studies covering around 32,000 references in medical literature and found 16.9% of quotations were incorrect, with 8.0% classed as major errors, meaning the source did not support the claim at all. Their meta-regression found no improvement over time. A study of high-impact general science journals published in the Royal Society's Proceedings A put the total quotation error rate at around 25%.

Those are peer-reviewed papers by professional researchers. Roughly one citation in six in published medicine does not say what it was cited for. The point is not that everyone is careless; it is that this failure is invisible by design, so nothing catches it unless a person deliberately looks.

Why AI makes it worse

Three distinct mechanisms, all documented.

Summarising strips the qualifiers. Peters and Chin-Yee tested ten leading models across nearly 5,000 summaries and published the results in Royal Society Open Science. Models routinely stated findings more broadly than the original texts, with some overgeneralising in 26 to 73 percent of cases, even when explicitly prompted for accuracy. Compared directly against human-written summaries, LLM summaries were around five times more likely to contain broad generalisations. Newer models tended to do worse, not better. The mechanism is simple: the hedges are the boring part, and a summary drops the boring part. "In a sample of 43 undergraduates, sleep restriction was associated with lower recall scores" becomes "sleep deprivation impairs memory."

Attribution drifts even with retrieval. Turning on web search fixes fabrication, not accuracy. The Tow Center for Digital Journalism tested eight AI search tools on their ability to identify the source of excerpts they were shown, and found incorrect answers to more than 60% of queries, with tools frequently citing the wrong article and doing so with high confidence and no hedging. If a system given the text struggles to attribute it correctly, a system working from memory has no chance.

The nearest plausible paper wins. When a model needs a citation for a sentence, what it produces is a reference that fits: right field, right topic, right kind of journal. Fit is not support. This is the same prediction machinery that produces entirely fabricated citations, except that here it lands on something real, which makes it more dangerous rather than less.

The four shapes it takes

Recognising the pattern makes checking faster, because you learn where to look first.

1. Scope creep

The study found something in 60 nursing students in one city. Your sentence says it about students. Or adults. Or people. This is the most common shape by far, and the fix is usually to narrow the sentence rather than to find a different paper.

2. Correlation promoted to cause

The paper reports an association, adjusts for confounders, and hedges carefully in the discussion. Your sentence says one thing causes the other. Check the study design before you check anything else: an observational study cannot support a causal claim no matter how strong its numbers are.

3. The claim belongs to a different paper

You find your sentence in the source, word for word almost, and it is in the introduction, with its own citation attached. You have cited a paper that was itself citing someone else. Go to that original and cite it instead, after checking that it says what this paper claims it says. Passing a claim along a chain without opening the first link is how errors propagate through a literature for decades.

4. Right topic, opposite finding

Less common, more embarrassing. The paper is about exactly your topic and reports the opposite of what you claimed, or reports a null result. This happens when a model recognises a famous paper's topic without recalling its conclusion. It is also worth checking whether a source has been contradicted since: scite classifies citing statements as supporting, contrasting or mentioning, which is a fast way to see whether a finding held up.

How a marker actually finds it

Not by scanning your bibliography. Nobody checks thirty references. What happens is narrower and much likelier than students assume.

A marker reads a sentence that surprises them. Either it is stronger than they expected the literature to support, or it is in their own subfield and they recognise the paper. They click one citation. That is the whole mechanism, and it is why the risk is not spread evenly across your reference list: it concentrates on your boldest claims, which are exactly the ones a model was most likely to overstate. The sentence that made your argument feel strong is the sentence that gets checked.

Two other routes are common. Supervisors reading a thesis chapter check the sources for anything they intend to build on, because they will have to defend it later. And once one reference has failed for any reason, including a broken DOI, the rest of the list gets read properly. That is why a fabrication accusation so often arrives as "several of your sources," even when only one was invented: someone started checking and kept going.

What they see when they open the paper is exactly what you would see. There is no ambiguity to argue about, no accuracy debate, no question of interpretation as there is with a detector score. The paper either says it or it does not.

The 90-second check

Per citation, in order:

  1. Read the abstract's results sentence, not the conclusion sentence. The results line carries the sample and the effect. The conclusion line is the authors' own generalisation and is already one step removed.
  2. Find the sample. Who, how many, where, when. Compare that against the population your sentence talks about. This alone catches most scope creep.
  3. Search the full text for your key term. Open the PDF or HTML and use Ctrl+F on the specific word your claim turns on. If the paper never uses it, that is your answer.
  4. Copy one supporting sentence, verbatim, with a page or section reference. Put it in a comment or footnote next to the citation. This is the whole method. Writing the quote down is what forces you to confirm it exists.
  5. Check whether that sentence is itself a citation. If it is, follow it to the original.

Keep the quotes. They cost nothing to store and they are the fastest possible answer if anyone ever questions your work, which is the situation an accusation of fabricated sources puts you in. They also make revision easier: when you rewrite a paragraph six weeks later, the quote tells you instantly whether the citation still fits the new sentence.

"The citation is real" is not "the citation is correct"

These are two separate tests and they need different evidence. Existence is a database question: a DOI resolver, an index, a publisher's record, all of which a machine can answer in milliseconds. Support is a reading question, and no database holds the answer, because it depends on a sentence you wrote. That asymmetry is why the tooling around citations is so lopsided: everything on the market checks the easy half.

So a reference can pass every check in the verification routine, existence, metadata, retraction status, and still be wrong in the only way that matters. The checks are worth running, and they are not the end of the job.

Academic integrity policies already treat the two as related failures. Virginia Tech's definitions of academic misconduct list citing a source for a proposition it does not support alongside inventing a source outright, and Iowa State's integrity tutorial defines fabrication to include attributing to a source ideas and information the source does not contain. From a marker's side of the desk, the distinction between a fake paper and a real paper cited for something it never said is thin: both mean the sentence has no evidence behind it.

The professional world has run into the same thing. When Deloitte Australia partially refunded a government contract in October 2025 over AI-generated errors in a report, the problems included both references to research that did not exist and a fabricated quote attributed to a real court judgment. The second kind is the one that is harder to spot and harder to explain.

When nothing supports your sentence

Sometimes you check honestly and find that no source says what you wrote. That is useful information, not a research failure. Three legitimate moves, in order of preference:

Narrow the claim to what the evidence shows. "Sleep restriction reduced recall in a small sample of undergraduates" is a real sentence with real backing. It is also more interesting than the vague version, because it is specific.

Attribute instead of asserting. "One 2023 study of 43 undergraduates found…" puts the claim's weight where it belongs and lets the reader judge.

Cut it. A paragraph rarely misses the sentence you could not support. If the whole argument depends on it, that dependency was a problem whether or not you found a citation.

What not to do is keep hunting until you find a paper that can be made to fit. That is the process that produces quotation errors in the published literature, and it is easy to do accidentally when you are tired. If you want a structured way to search for real backing, we cover it in how to find a source for a claim you already wrote.

Writing so this doesn't happen

The structural fix is to reverse the order: find the source first, read it, and write the sentence the source supports. Most citation problems come from writing the sentence you wanted and then shopping for evidence. Starting from real papers you have actually read costs more time up front and removes an entire category of failure later.

If you do use a chatbot to gather sources, making it search rather than recall reduces fabrication. It does not touch this problem. Retrieval improves whether the paper exists; only reading tells you whether it supports your sentence.

The quote lives with the claim

CiteOwl stores the exact passage behind every citation it writes, so you can see what a source actually says without opening the PDF.

Start writing

Things worth knowing.

Is citing a real paper for the wrong claim actually a problem?
Yes, and several university integrity codes name it directly. Virginia Tech's definitions of academic misconduct include citing a source for a proposition it does not support, alongside inventing a source outright. A marker who opens the paper sees the same thing you would see.
How common is this compared to fabricated references?
Much more common, and it predates chatbots. A 2025 meta-analysis of quotation accuracy in medicine, covering 46 studies and around 32,000 references, found 16.9% of quotations were incorrect and 8.0% were major errors where the source did not support the claim at all. AI adds volume to a failure human authors were already making.
Why does an AI do this even when it has read the paper?
Because summarising drops the qualifiers. Researchers testing ten leading models on nearly 5,000 summaries found they routinely stated results more broadly than the original texts, with some models overgeneralising in 26 to 73 percent of cases, and LLM summaries were around five times more likely than human-written ones to contain broad generalisations.
What is the fastest way to check?
The quote test. Open the source and find one sentence, table or figure that supports your claim, then paste it verbatim into a note beside the citation. If you cannot find one in ninety seconds of searching the full text, the citation is probably wrong even though the paper is real.
What if no source supports my sentence?
Change the sentence. Narrow it to what the evidence actually shows, attribute it to the single study that found it, or cut it. Attaching a citation that does not support the claim is worse than having no citation, because it looks like evidence and disappears the moment anyone checks.
Read next.