CiteOwl
CiteOwl

Does Perplexity make up citations?

Not often in the way a chatbot does. Perplexity searches before it writes, so the reference list of ghost papers that ChatGPT hands you is mostly not the failure here. The failure that has been measured, repeatedly, is quieter: a real page, correctly linked, sitting under a sentence it does not support. A Stanford audit found that roughly one in four of Perplexity's citations did not back the statement it was attached to. Journalists reviewing three thousand answers in 2025 found it the most careful of the four big assistants and still put a significant problem in 30% of its answers, including quotes that had been altered or made up outright. If you are pasting its links into a bibliography, that is the number that should worry you, because a citation that resolves feels safe.

The question behind this one is usually specific: you have an essay due, Perplexity gave you eight linked sources, and you want to know whether you can put them in the reference list without opening them. The answer is no, and the reason is more interesting than a blanket warning about AI.

The short answer, and why it is not "no"

Perplexity is a retrieval system with a language model on top. It runs searches, pulls back pages, and writes an answer from what came back, attaching numbered links to the finished sentences. That order of operations removes most of the fabrication problem that general chatbots have, and it is why "does Perplexity make things up" gets a different answer from the same question about ChatGPT. A model with no search has nothing to work from but its training data, which is why it produces a reference that looks exactly like a real one and is not, the mechanism we take apart in why AI makes up citations.

What retrieval does not fix is the join. The links are attached to the text after the text exists, and nothing in that step guarantees that the page under link three says what sentence three claims. So the errors move rather than disappear. Instead of a paper that was never written, you get a real paper that does not support the point, a quote that has been tidied up, or a list of nine sources where three were used. Every one of those survives the check most people actually perform, which is clicking the link and seeing that something loads.

The distinction that matters for your bibliography: a fabricated citation fails loudly the moment anyone looks for it. A misattributed one fails silently, and it fails at the worst possible time, which is when a marker who knows the field opens the source and finds it does not say that.

The Stanford audit: one citation in four

The most careful measurement of this is still Nelson Liu, Tianyi Zhang and Percy Liang's audit of generative search engines, published in Findings of EMNLP 2023. They put 1,450 queries through each of four systems, including perplexity.ai, and had 34 trained human annotators judge two things separately: whether each generated statement was supported by the citations attached to it (citation recall), and whether each citation supported the statement it sat under (citation precision).

System Statements fully supported by their citations Citations that support their statement
perplexity.ai 68.7% 72.7%
NeevaAI 67.6% 72.0%
Bing Chat 58.7% 89.5%
YouChat 11.1% 63.6%
Average across all four 51.5% 74.5%

Read the right-hand column as the one that decides whether you can trust a link. Perplexity scored 72.7%, which means about 27% of the citations it produced did not support the sentence they were attached to. The left column is the other half of the problem: 68.7% of its statements were fully backed by any citation at all, so roughly three sentences in ten carried a claim the cited pages did not cover.

The authors' own summary of all four systems is worth quoting because it names the trap rather than the score: responses "are fluent and appear informative, but frequently contain unsupported statements and inaccurate citations". Fluency is what you are judging when you skim, and fluency is uncorrelated with whether the link underneath is doing its job.

Two honest caveats. The study is from 2023, and the systems have changed. And two of the four engines it tested no longer exist in that form. What has not changed is the architecture: links are still matched to generated text rather than generated from read text, so the class of error the study measured is still available.

What 2025 measured, on news rather than papers

Two large studies landed in 2025, both about news reporting rather than academic sources. Read them as evidence about the mechanism rather than a measurement of what happens to journal articles, because nobody has published the equivalent audit for scholarly citations.

Columbia's Tow Center for Digital Journalism ran 1,600 queries across eight AI search tools in March 2025, feeding each one an excerpt from a known article and asking it to identify the source. Collectively the tools answered more than 60% of queries incorrectly. Perplexity was the best of the eight and still got 37% wrong. The behaviour around the errors is the part worth carrying away: most tools stated wrong answers confidently rather than declining. ChatGPT's search mode identified the wrong article 134 times out of 200 while signalling uncertainty in only 15 of those, and never once declined to answer.

The larger study came in October 2025, when the BBC and the European Broadcasting Union had journalists at 22 public service broadcasters, across 18 countries and 14 languages, assess answers from ChatGPT, Copilot, Gemini and Perplexity. Their News Integrity in AI Assistants report covers 3,113 questions in total, with 681 core answers per assistant graded on accuracy, sourcing, separation of fact from opinion, and context.

Assistant Answers with a significant issue of any kind Answers with a significant sourcing issue
Perplexity 30% 15%
ChatGPT 36% 24%
Copilot 37% 15%
Gemini 76% 72%
All four together 45% 31%

Perplexity comes out of that table well, and the report says so, noting that participating organisations "often felt Perplexity was strong on sourcing". It is also the table where you should notice that the winner still put a significant problem into three answers in ten. Across the whole dataset, sourcing was the single biggest cause of significant issues, which the report defines to include "information in the response not supported by the cited source". That is the academic failure mode in a news costume.

One more figure from the same report, because it is the one that transfers directly to essays. Of the 1,053 answers that contained a direct quote, 12% had significant problems with the accuracy of that quote. If you are quoting from an AI answer rather than from the source, that is your error rate.

Three failure modes you can see in the transcripts

Percentages are abstract. The EBU report publishes the individual failures, and Perplexity's are worth reading because each one maps onto something a student does.

Quotes that were altered or invented. Asked why Birmingham's bin collectors went on strike, Perplexity produced three problem quotes in one answer. Two, attributed to a trade union and to the city council, do not appear in the sources cited for them and could not be found anywhere. A third was altered from the wording in the cited BBC article. This is the failure with the worst consequences in a paper, because a quotation is the one thing a marker can check in ten seconds, and the one thing they will assume you copied rather than paraphrased.

Sources listed but never used. Answering a question about a Myanmar earthquake, Perplexity attached a block of 19 URLs and referred to three of them in the answer. Asked what NATO does, it supplied nine links and used three. One broadcaster's evaluator described the pattern as "providing long lists of URLs without actually referring to them in the answers". A different broadcaster found nine of its own articles cited on a question about the Gulf of America, including pieces about the abolition of first-class train seats and about power plants in the Netherlands. If you copy the source block into your bibliography, most of what you have added is decoration, and padding a reference list with sources you did not use is visible to a marker for the same reason it is visible here.

The wrong kind of source, presented as the right kind. Answering a question about why people dislike Tesla, Perplexity built part of its response on a satirical column and did not flag it as satire. A retrieval system ranks by relevance to your question, not by whether the page is the kind of thing you are allowed to cite, and satire ranks well because it is written about exactly the topic you asked about. Judging that is your job, and it is the same judgement covered in how to find credible sources.

The moment it turns back into a chatbot

Everything above assumes Perplexity is doing what it is built to do, which is answer a question from pages it just fetched. There is a specific request that quietly takes it out of that mode, and students make it constantly: asking for a formatted reference list.

"Give me ten APA references on urban heat islands" is not a retrieval task in the way "what causes urban heat islands" is. It asks for strings in a particular shape, and producing a citation string from a search result means filling in authors, volume, issue, pages and a DOI, fields that are often not on the page the system read. Whatever is missing gets generated, and generated citation metadata is exactly what fails.

We measured what that looks like when models do it with no search at all. In August 2026 we asked three current models for a hundred references each on ten ordinary essay topics and checked every DOI against CrossRef. For the strongest model, 19% of the DOI-bearing references were unusable: 7% dead, and 12% resolving cleanly to a completely different real paper. The smaller model was at 41%. The raw dataset is public. That is not a measurement of Perplexity, and we are not presenting it as one. It is a measurement of the operation you are invoking whenever you ask any of these systems for a bibliography rather than an answer.

The second version of the same problem is what happens when the search comes back thin. The Tow Center's finding that these tools rarely decline applies here: an answer gets written whether or not the retrieval supported one. In the EBU study, refusals ran at 0.5% across 3,113 questions.

The check that takes thirty seconds

You do not need to re-do Perplexity's work. You need to close the gap between the sentence and the page, and there is only one way to do it.

  1. Open the link and find the sentence. Not the page, the sentence. Search the page for the number, the name or the phrase your claim depends on. If you cannot find it in under a minute, the citation is not supporting your claim, whatever the link does.
  2. Check what kind of page you landed on. A study, a press release about the study, a news article about the press release and a blog post about the news article all rank for the same query, and each hop drops a qualifier. Cite the study.
  3. Take the reference from the paper, not from the answer. Authors, year, journal, volume and DOI come off the article's own landing page. If the DOI opens something other than the paper you just read, you have hit the failure in the DOI links to the wrong paper, which is more common than it sounds.
  4. Verify anything you are quoting, character by character. Given a 12% error rate on direct quotes, this is the highest-value minute you will spend.

Our free citation checker does the mechanical half of this: paste a reference list and it resolves each entry against CrossRef and OpenAlex and tells you which ones do not exist or do not match. It cannot tell you whether a real paper supports your sentence. Nothing can, except reading it. The full manual routine is in how to check if a citation is real, and the specific case where the source is real but says something else is in when a chatbot cites a paper that says something else.

A reading list is not a bibliography

The useful way to think about Perplexity for coursework is that it produces a reading list very fast. That is a real service, and on an unfamiliar topic it will save you an afternoon of the search work described in how to find peer-reviewed articles. What it does not produce is a bibliography, because a bibliography is a record of what you read, and nothing in the pipeline knows what you read.

The gap between those two things is where the marks are lost, and it is not a gap any answer engine is trying to close. If you want the tool comparison rather than the accuracy question, we set out both sides in CiteOwl vs Perplexity, including where Perplexity is the better twenty dollars.

CiteOwl is built in the other order. The agent searches OpenAlex, CrossRef, Unpaywall and the web, retrieves the papers it selected, reads them, and then writes the claim from a specific passage, which is stored beside the sentence and sits one hover away in the document. There is no step where finished prose goes looking for a source to sit next to. That does not make it a better conversationalist than Perplexity, and it is not trying to be. It means the reference list at the end is a record of reading rather than a list of links that were nearby.

Things worth knowing.

Does Perplexity hallucinate?
Yes, but not usually by inventing a paper that does not exist. Because it retrieves pages before it writes, the more common failure is a real source attached to a sentence it does not support, and quotes that have been altered or made up. In the BBC and EBU evaluation of more than three thousand questions, 30% of Perplexity's answers carried at least one significant issue, which was the best of the four assistants tested and still close to one answer in three.
Are Perplexity's citations real?
The links are usually real pages. Whether they support the sentence they sit under is a separate question, and it is the one that has been measured. A Stanford audit of generative search engines found that 72.7% of perplexity.ai's citations supported the statement they were attached to, so roughly one in four did not, and that only 68.7% of its statements were fully supported by any citation at all.
Is Perplexity better than ChatGPT for finding sources?
On the published measurements it is ahead on sourcing. In the 2025 BBC and EBU study, 15% of Perplexity's answers had a significant sourcing problem against 24% for ChatGPT and 72% for Gemini. Columbia's Tow Center recorded Perplexity as the strongest of eight AI search tools in March 2025, at 37% of queries answered incorrectly. Better than the alternatives is not the same as safe to cite unchecked.
Can I cite Perplexity in my paper?
Cite the source, not the tool. Perplexity is a way of finding pages, in the same category as a library search, and a search engine does not appear in a reference list. Open the page it pointed you at, read the passage that carries the claim, and cite that. If your assignment requires you to disclose the AI tools you used, that belongs in a disclosure statement or a methods note rather than in the bibliography.
Read next.

Papers read in full, not links from a search

Free to start. No card needed.

Start writing