CiteOwl
CiteOwl

Does Gemini make up citations?

Yes, and the common failure is stranger than an invented paper. Gemini makes up the attribution. In the largest independent evaluation published so far, journalists at 22 public broadcasters graded 675 Gemini answers and found a significant sourcing problem in 72% of them, against 24% for ChatGPT and 15% each for Perplexity and Copilot. In the same study Gemini's plain factual accuracy improved faster than any other assistant's. Those two findings together are the thing worth understanding before you paste anything into a bibliography: Gemini's weak point is not mainly that it gets things wrong, it is that it cannot reliably tell you where an answer came from.

Most people do not choose Gemini. It arrives: in the side panel of a Google Doc, at the top of a search results page, in the Workspace account a university handed out. That makes the question less academic than it sounds, because the assistant you did not pick is the one you are least likely to have formed a view about.

The number, and what it is a number about

In October 2025 the BBC and the European Broadcasting Union published News Integrity in AI Assistants, the largest study of its kind so far. Journalists at 22 public service media organisations across 18 countries and 14 languages put the same questions to ChatGPT, Copilot, Gemini and Perplexity, using the free consumer version of each, and graded 3,113 answers on accuracy, sourcing, separation of fact from opinion, editorialisation and context. For the core question set there were 675 graded Gemini answers.

Across all four assistants, 45% of responses carried at least one significant issue and 81% had a problem of some kind. Sourcing was the single biggest cause, at 31% of all responses. Gemini recorded the highest rate of significant issues overall at 76%, double the next assistant. But the headline number conceals the shape of the problem, and the shape is the useful part.

Split a citation into its two halves. Does the source exist and load? And does it say what the sentence claims it says? Gemini's measured failure sits almost entirely in the second half, which is the half no automated check can see.

Sourcing is where the four assistants separate. The report gives Gemini 72% of responses with a significant sourcing issue, against 24% for ChatGPT and 15% each for Perplexity and Copilot. On accuracy they are close together. That asymmetry is the finding, and it is why "is Gemini accurate" and "can I cite what Gemini told me" are different questions with different answers.

A source that was named but not used

The report devotes a section to Gemini's sourcing specifically, and it breaks the 72% into behaviours you can recognise in your own transcripts.

Behaviour recorded by evaluators Gemini Other three assistants
Incorrect or unverifiable sourcing claim: an organisation named as the source, with a link to something else or no link at all 54% of responses None above 4%
No direct source of any kind: no URL pointing to a specific piece of content 42% of responses Reported as predominantly a Gemini problem
Any significant sourcing issue 72% of responses ChatGPT 24%, Perplexity 15%, Copilot 15%

That first row is the one to sit with. More than half of Gemini's answers credited a claim to a named organisation while linking elsewhere, or linking nowhere. No other assistant in the study did this in more than one answer in twenty-five. It is not a rounding difference, it is a different behaviour.

The worked examples in the report make it concrete. Asked by CBC about the origins of the Los Angeles fires, Gemini wrote that "CBC News reports highlight that climate change significantly contributed to the conditions" and that "CBC News emphasizes that human-caused climate change created the critical underlying conditions". Five sources were attached to the answer. None of them were CBC News, and CBC's own evaluators could not find any origin for those statements. Answering a question from NPR, Gemini framed its response with "Here's a breakdown of the key aspects, drawing from NPR sources" and then supplied 11 sources, none from NPR.

Sometimes it argues the point. Asked to use NOS sources on whether Türkiye is in the EU, Gemini replied that "while the NOS is a reliable news source, the status of EU membership is a fundamental fact that is widely known and doesn't need to be specifically linked to a recent NOS publication for this basic information". That is a coherent sentence and a terrible principle for anyone writing an essay, where the widely known things still need a citation because the marker is checking whether you can support a claim, not whether the claim is true.

An evaluator at the Finnish broadcaster Yle gave the pattern a name that deserves wider use: "ceremonial citations", meaning references added to create an impression of thorough research but which do not support the stated claims when checked. Every marker who has read a first draft knows the phenomenon. Now a machine produces it at scale.

Accuracy improved. Sourcing did not.

This is where the study earns its length. The BBC ran an earlier round of the same evaluation in December 2024, and although the two rounds are not directly comparable overall, the BBC's own answers can be compared with each other: 362 responses in the first round, 237 in the second. Read down the column for Gemini alone.

Gemini, BBC evaluations only December 2024 May and June 2025
Significant issues with accuracy 46% 25%
Significant issues with providing context 36% 5%
Significant issues with sourcing Not published separately 47%, described as broadly unchanged

Gemini was the biggest improver on accuracy of the four assistants, halving its error rate and joining the pack: all four now sit between 20% and 29%. On context it improved most of all, from 36% to 5%. On sourcing it did not move, while every other assistant dropped into the 10 to 15% range, Copilot falling from 27% to 10%.

One year of engineering effort moved the things that are easy to score and left the thing that is hard to score exactly where it was. You cannot wait this one out on the assumption that the next model release fixes it, and you should be sceptical of any advice that treats "AI accuracy is improving" as an answer to the citation question. Those are different tracks, and only one of them is moving.

Every figure above is about news

Now the caveat that most pages on this topic will not give you. The EBU study, and Columbia's Tow Center study alongside it, measured answers to news questions. Neither one asked an assistant for journal articles, and neither one graded whether a cited paper supports an academic claim. Nobody has published that study. If you find a page telling you that Gemini fabricates a specific percentage of scholarly references, ask which measurement it is quoting, because as far as we can establish there is not one.

The Tow Center's March 2025 comparison of eight AI search engines is still worth reading for one Gemini-specific result. Researchers fed each tool an excerpt from a known article and asked it to identify the source, 1,600 queries in total, and found the tools collectively answered more than 60% incorrectly. On links specifically: "More than half of responses from Gemini and Grok 3 cited fabricated or broken URLs" that led to error pages. Google's own Google-Extended crawler was permitted by ten of the twenty publishers tested, and Gemini still produced a completely correct response only once.

There is one measurement in the academic domain, and it complicates the story in a way worth taking seriously. In April 2026, Delip Rao, Eric Wong and Chris Callison-Burch at the University of Pennsylvania published an audit of citation URL validity covering 168,021 URLs across 32 academic fields, produced by three search-augmented models. They checked each URL against the live web and the Wayback Machine to separate a dead link from one that never existed.

Model, academic questions across 32 fields URLs per question Citation URLs that do not resolve
Gemini 2.5 Pro 10.7 4.2%
GPT-5.1 46.4 8.5%
Claude Sonnet 4.5 28.3 9.4%

Gemini came first. It cited the fewest URLs per question and the fewest of them were broken, and the paper notes it "consistently achieves the lowest rates across nearly all fields". That is not a contradiction of the EBU result, and reading it as one is the mistake. The Pennsylvania study asks whether a link loads. The EBU study asks whether the source supports the claim. Gemini is good at the first and, on the evidence we have, the worst of the four big assistants at the second. Both things can be true because they are the two halves of the callout above, and a reference list is only as good as the second half.

Which Gemini are you actually using?

"Gemini" names several products with different retrieval machinery, and the measurements do not transfer between them. Three distinctions are worth holding.

The free consumer assistant. This is what the EBU study tested, and the report's own version table names it: the default free tier, Gemini 2.5 Flash, at the time of the May and June 2025 test window. It is also, by construction, the tier most students meet, whether through the standalone app, the Ask Gemini button in the top right of a Google Doc, or an AI Overview in a search result. Google's help page states the ground truth without decoration: not every response includes sources, and "if you don't have the Sources button below a response, Gemini Apps didn't provide any links for that particular response".

Deep Research. A different product, and the Pennsylvania study found it fails differently too. Gemini's deep research agent produced 113 URLs per query, the most of any of the ten systems tested, and had the highest fabricated-URL rate of all of them at 13.3%, against 4.6% to 4.8% for the ordinary search-augmented Gemini models. More retrieval steps meant more invented addresses, not fewer. A long report with a hundred references is not more verified than a short answer with ten, it is a bigger surface to check.

The API and the paid tiers. Nobody has published an independent sourcing audit of these against the same rubric. Assume nothing from the free-tier numbers in either direction.

What this does to a reference list

Three of the failures above land directly in a bibliography.

A quotation you did not take from the source. Of the 290 Gemini responses in the EBU study that contained a direct quote, 20% had significant problems with the accuracy of that quote, against 12% across all 1,053 quoted responses from all four assistants. A quotation is the one element a marker can check in ten seconds, and the one they will assume you transcribed rather than paraphrased. If it is in quotation marks in your essay, you read it in the source or you do not use it.

An attribution you inherited. "According to the OECD" is not a citation, it is a sentence about a citation. If Gemini wrote that phrase and you carried it into your draft, you now have a claim with a named owner and no link, which reads to a marker as though you consulted the OECD. The version where the source is real and simply says something else is covered in when a chatbot cites a paper that says something else, and the underlying mechanism in why AI makes up citations.

A formatted reference list. Asking any assistant to output ten references in APA changes the task from retrieval to generation, because authors, volume, issue, pages and a DOI are usually not on the page the search returned, so the missing fields are written rather than read. We measured what that looks like in August 2026, asking three models for a hundred references each on ten ordinary essay topics with no web access and resolving every DOI against CrossRef. For the strongest model, 19% of DOI-bearing references were unusable, and the split matters more than the total: 7% were dead and 12% resolved cleanly to a real but different paper. The dataset is public. It did not test Gemini, and we are not presenting it as a Gemini number. It is a measurement of the operation, and the operation is the same whichever assistant performs it. The wrong-paper half of that result is its own quiet disaster, taken apart in the DOI links to the wrong paper.

The check that catches all three takes about a minute per source. Open the link. Find the sentence, not the page, by searching it for the number or phrase your claim rests on. Confirm the page is the study and not a news article about the study. Then take the author, year, journal and DOI from the article's own landing page rather than from the answer. Our free citation checker does the mechanical part of that for a whole list at once, resolving each entry against CrossRef and OpenAlex, and the full manual routine is in how to check if a citation is real.

The gap between a mention and a citation

What the evidence supports is narrower and more useful than "do not use Gemini". Gemini is competitive on accuracy, ahead of its rivals on whether a link loads, and last by a distance on whether the source it names is the source it used. If you treat it as a way to understand a topic quickly, that profile is fine. If you treat its output as the provenance of your claims, the one thing it is worst at is the one thing you are relying on.

The comparison holds across the category rather than singling out one vendor. Perplexity retrieves first and still attaches real pages to sentences they do not support, which we go through in does Perplexity make up citations. ChatGPT and Claude have their own division of labour, set out in ChatGPT vs Claude for research papers. None of them was built to produce a bibliography, because a bibliography is a record of what someone read, and nothing in a chat pipeline knows what was read.

That is the piece CiteOwl builds around. The agent searches OpenAlex, CrossRef, Unpaywall and the web, retrieves the papers it chose, reads them, and writes each claim from a specific passage that stays attached to the sentence, one hover away in the document. Every edit arrives as a diff you accept or reject, so the sentence and the source it came from are the same object rather than two things introduced to each other afterwards.

Things worth knowing.

Does Gemini make up sources?
Yes, and more often it makes up the attribution rather than the source. In the BBC and EBU evaluation of 675 Gemini answers, 72% had a significant sourcing problem, three times the rate of ChatGPT. The specific behaviour evaluators recorded in 54% of Gemini responses was a claim credited to a named organisation with a link to something else, or to nothing. No other assistant in that study was above 4% on the same failure, and 42% of Gemini's answers carried no link to specific content at all.
Is Gemini worse than ChatGPT at citations?
On the published measurements, yes, and by a wide margin on sourcing specifically. The BBC and EBU study recorded significant sourcing issues in 72% of Gemini's answers against 24% for ChatGPT, 15% for Perplexity and 15% for Copilot. On plain factual accuracy the four are much closer together, all in the 20 to 29% band on the same study's repeat measurements. So the gap is not that Gemini is more often wrong. It is that Gemini more often cannot show you where an answer came from.
Why does Gemini say "according to" a source it never links?
Nobody outside Google can say for certain. The BBC and EBU researchers noted that their questions carried a prefix asking each assistant to use the participating broadcaster's sources where possible, and that Gemini appeared to answer that instruction in words rather than in links. The behaviour is worth knowing because a sourcing phrase costs nothing to generate and looks exactly like evidence. Google's own help page confirms the underlying fact plainly: not every response includes sources, and if there is no Sources button, no links were provided.
Can I use Gemini to build a reference list?
Not as the last step. Asking any assistant for a formatted bibliography changes the job from retrieval to generation, because authors, volume, issue, pages and a DOI are usually not on the page a search returned, so the missing fields get written rather than read. Use it to find candidate reading, then open each paper, confirm it says what you are citing it for, and take the reference details from the article's own landing page rather than from the answer.
Read next.

The paper behind the sentence, every time

Free to start. No card needed.

Start writing