CiteOwl
CiteOwl

Are there fake papers on Google Scholar?

Yes. In 2024 four researchers searched Google Scholar for two phrases only a chatbot writes and pulled out 139 papers with undeclared use of a language model. Nineteen sat in journals indexed by Scopus or Web of Science, eighty-nine in journals indexed nowhere, twelve were working papers, and nineteen were student papers in university repositories. That last number is the one that should change how you read a result list, because a coursework PDF on a university server meets Google Scholar's inclusion criteria in full. Scholar is a crawler, not a vetted database, and it publishes exactly what it checks before indexing a document. None of it is about whether the paper is any good.

The useful version of the question is not "is there junk in Google Scholar". Every index has junk. It is "what does appearing in Google Scholar prove", and Google answers that one itself, on a help page.

What was actually found, and where it was sitting

The measurement is Jutta Haider, Kristofer Rolf Söderström, Björn Ekström and Malte Rödl, in the Harvard Kennedy School Misinformation Review, September 2024. Their method is blunt. They scraped Google Scholar for documents containing one of two strings a chatbot emits and a researcher does not: "as of my last knowledge update" and "I don't have access to real-time data". That returned 227 papers, which all four authors read and coded together, discarding 88 where the use of a model was legitimate or declared and keeping 139 where it was neither.

Here is where the 139 were living.

Where the paper sat Papers
Journals not indexed by any citation database 89
Journals indexed by Scopus, Web of Science, DOAJ or the Norwegian register 19
Student papers in university databases 19
Working papers, mostly in preprint databases 12
Total 139

The two largest subject categories the authors named were computing, with 32 papers, and environment, with 27. The other 80 fall into further categories.

The 19 student papers are the part of this that belongs to you rather than to the publishing industry. Nobody smuggled them anywhere. A repository is a normal home for coursework and theses, Scholar indexes repositories as a matter of course, and in a result list a deposited essay written with an undeclared chatbot is indistinguishable from a refereed article. As the authors put it, Google Scholar "easily locates and lists these questionable papers alongside reputable, quality-controlled research".

They also traced where the papers spread. The 27 environment papers turned up across 26 different domains and 56 separate URLs. The five hosting the most of them:

Domain Environment papers hosted
researchgate.net13
orcid.org4
easychair.org3
ijope.com3
publikasiindonesia.id3

That is the durable part of the finding. Most of the papers, the authors write, "exist in multiple copies and have already spread to several archives, repositories, and social media. It would be difficult, or impossible, to remove them from the scientific record." Retracting the original does not retract the copies. Their word for the risk is evidence hacking, "the strategic and coordinated malicious manipulation of society's evidence base".

139 is a floor on a count, not a rate. The search finds only papers that left the chatbot's boilerplate in the text; anyone who deleted that sentence is invisible to it. The authors make no claim about what share of Scholar's index is machine-written, and neither will we. Google does not publish the size of that index, so the denominator does not exist.

What Google Scholar checks before it indexes a paper

This is the half most write-ups skip. Scholar publishes a list of what it checks. Start with the friendliest page in the inclusion guidelines, the part addressed to an individual author with no publisher behind them: upload the paper to your own website, link it from a publications page, and make sure that:

Then, verbatim: "That's it! Our search robots should normally find your paper and include it in Google Scholar within several weeks."

The rest of the documentation adds requirements of the same kind. A site must consist primarily of "scholarly articles", a category Google defines to include "journal papers, conference papers, technical reports, or their drafts, dissertations, pre-prints, post-prints, or abstracts", and the abstract must be visible without a login. Files must be HTML or PDF with searchable text, under 5MB, reachable from the homepage "by following at most ten simple HTML links". Three metadata tags are mandatory: a title, at least one author, and a publication date.

Line them up against what each one actually tests and the picture is unambiguous.

What Google Scholar requires What that is a test of
A PDF or HTML file with searchable text, under 5MBFile format
Title in a large font at the top of page oneLayout
Authors on a separate line below the titleLayout
A References or Bibliography section at the endLayout
Title, author and date meta tagsMetadata
Abstract or full text visible with no login or popupAccess
Reachable from the homepage within ten linksCrawlability
The site answers when the crawler visitsUptime

Not one of those tests whether the paper is correct, reviewed, or written by a person. They test whether a document is shaped like a paper and reachable by a robot. A well-formatted invention passes all eight. A brilliant article behind a login passes none.

Scholar's about page is worth reading for what it does not say. On ranking: "Google Scholar aims to rank documents the way researchers do, weighing the full text of each document, where it was published, who it was written by, as well as how often and how recently it has been cited in other scholarly literature." Where it was published is a ranking signal, not a gate, and peer review is not mentioned on the page at all.

Haider and colleagues state the consequence plainly: Scholar's "inclusion criteria are based on primarily technical standards", and "the search interface does not offer the possibility to filter the results meaningfully by material type, publication status, or form of quality control, such as limiting the search to peer-reviewed material". There is no peer-reviewed-only checkbox because there is no field behind it. The habits that make Scholar an excellent tool anyway are in how to find peer-reviewed articles.

The experiment that put six invented papers in the index

Someone tested this directly. In 2012 three bibliometrics researchers at the University of Granada, Emilio Delgado López-Cózar, Nicolás Robinson-García and Daniel Torres-Salinas, set out to find how hard it is to get fabricated documents into Scholar. They reported it in the Journal of the Association for Information Science and Technology as "The Google scholar experiment: How to index false papers and manipulate bibliometric indicators".

They invented an author, Marco Alberto Pantani-Contador, and gave him six documents. The contents were not research: text copied off their own group's website, run through Google Translate, with some graphs pasted in. What each document had was the shape, a title, a short abstract, an author line and a reference list, in this case 129 papers by members of their own group. On 17 April 2012 they put the six files on a page under their university's domain.

"Google indexed these documents nearly a month after they were uploaded, on May 12, 2012." Colleagues began receiving automated Scholar alerts saying someone named Pantani-Contador had cited them. For the youngest researchers, whose real counts were small, "their citation rates were multiplied by six".

The ending is the part worth remembering. They published what they had done on 29 May, and two days later Google erased Pantani-Contador and quarantined the authors' own profiles for weeks without telling them. But an earlier pilot document, unnamed in the announcement, stayed in Scholar's cache. Their reading, in the paper: the removal "was more of a reaction to our complaint than because GS had uncovered the deception".

That was fourteen years ago and the documents were crude. It still matters because the property they exploited, a crawler judging document shape, was never a bug to be patched. It is the published inclusion criterion, still on Google's help pages, quoted above.

Reading a Scholar result for what it is evidence of

Once you know what the index is, a result page reads differently. Four elements are worth separating.

The listing itself. A result means a crawler found a document shaped like a paper on a site that let it in, and parsed a title, an author and a date out of it. That is the whole claim, and it says nothing about the contents.

Results marked [citation]. The ones with no link to click. Scholar's help pages define them exactly: "These are articles which other scholarly articles have referred to, but which we haven't found online." That is an index entry built out of somebody else's reference list, for a document Scholar has never fetched. If a fabricated reference sits in a paper Scholar does index, the machinery making those entries cannot tell. A [citation] result is evidence that something cited a thing by that name, and nothing else.

The "cited by" number. It counts documents already in Scholar's index whose parsed reference lists its software matched to this record, so it mixes refereed articles with preprints, theses and repository deposits, unlabelled. Scholar is candid about what that means: when a count falls, it says, the likely cause is that citing papers "have either disappeared from the web entirely, or have become unavailable to our search robots". A number that moves when a website goes down measures the web, not the paper.

How fresh the record is. Scholar adds new papers several times a week, but on existing ones it states the lag: "updates to existing records take 6-9 months to a year or longer, because in order to update our records, we need to first recrawl them from the source website". The record in front of you is a snapshot taken at some point in the last year, so if the article was corrected or withdrawn since, the entry will not show it. That is why checking whether a paper has been retracted is a separate job.

The same mistake, one step later

The error underneath all of this is treating a machine-produced signal as though a person had checked something, and it returns the moment you have a reference list. In August 2026 we asked current chat models for reference lists on ten ordinary undergraduate essay topics, with no tools and no web access, then resolved every DOI against CrossRef. Of the 100 DOI-bearing references from the strongest model, 81 pointed at the paper cited, 7 were dead, and 12 resolved cleanly to a real but different article. A bad DOI from that model was almost twice as likely to open something as to open nothing. The dataset is published with its method and limits.

Put the two side by side. A Scholar result proves a document passed a crawler's format tests. A resolving DOI proves a string is registered to some record. Both are automatic, both are cheap to obtain, and neither is a person saying this is the paper you think it is. That gap is the subject of the DOI that links to the wrong paper, and the fix is the same either way.

A five-minute check before you cite a Scholar result

Not a checklist for everything. A routine for the sources you are actually going to cite.

  1. Read the venue line, not just the title. Scholar prints it in grey under each result and most people skip it. If it names a repository, a preprint server or a personal page, you have a document rather than a refereed article, which is often fine as long as you know it and say so. If it names a journal you have never heard of, spend two minutes checking the venue.
  2. Search the PDF for the giveaway phrases. This one comes free with the study above: "as of my last knowledge update" and "I don't have access to real-time data". Those are not stylistic tells you have to interpret. They are a chatbot's own boilerplate left in the manuscript, and they took four researchers to 139 papers.
  3. Treat a [citation] entry as unverified. No link means Scholar has never seen the document. Find it somewhere else before you cite it, or drop it.
  4. Read the passage you are about to cite. Not the abstract, not the snippet. The abstract is the part a fabricated paper gets right. Everything else in judging a source comes after actually opening it.
  5. Match the record to your reference before you submit. Authors, year, journal, DOI, all four. The full routine is here if you want to do it by hand, and our free citation checker runs a whole reference list against CrossRef and OpenAlex in one go.

What nobody has measured

Three honest gaps, because a page reporting only what is known is no use for judging how worried to be.

Nobody knows what proportion of Scholar's index is fabricated. The 139 papers came from two search phrases, so the figure is a floor on a count and cannot become a rate. Google publishes no index size, and its own help says the way to assess coverage from outside is to sample titles. Anyone quoting you a percentage for how much of Scholar is fake made it up.

Nobody has measured how many [citation] entries stand for papers that were never written. The mechanism allows it, since those entries come from reference lists rather than documents, and fabricated references are entering published reference lists. The question is answerable: sample the entries, resolve each one, hand-check the residue. As far as we can tell, nobody has.

And nobody outside Google can audit what happens after a paper is flagged: no public record of removals, no notice on a result whose source was withdrawn, and a documented lag of six to nine months to a year before an existing record is recrawled at all.

What you control is the one step none of this machinery performs. Every automatic signal here, indexing, resolution, citation counts, is a proxy for a person opening a document and reading it. The proxies are good for finding candidates and worthless as verdicts. Read the thing.

That order is what CiteOwl is built to preserve. It searches OpenAlex, CrossRef, Unpaywall and the web, reads the papers it retrieves rather than their titles, writes each claim from a passage it keeps beside the sentence, and delivers every edit as a diff you accept or reject. The reference in your bibliography is the record an index returned, and the sentence above it came from text that was read.

Things worth knowing.

Does Google Scholar check whether a paper is peer reviewed?
No. Scholar publishes its inclusion criteria and every one of them is technical: the file has to be a PDF or an HTML page under 5MB with searchable text, the title has to sit in a large font at the top of the first page, the authors go on a separate line below it, there has to be a section called References or Bibliography at the end, the abstract has to be readable without a login, and the URL has to be reachable by a crawler. Nothing on that list asks who reviewed the work, or whether anyone did. Scholar's own description of its ranking never mentions peer review either.
How do fake papers get into Google Scholar?
By meeting the technical criteria, which a document can do without being reviewed by anyone. Scholar's guidelines tell an individual author to upload a PDF to their own website, add a link to it from a publications page, and wait: "Our search robots should normally find your paper and include it in Google Scholar within several weeks." Researchers at the University of Granada tested this in 2012 by posting six invented documents under a false author name on their university domain. Scholar indexed them about a month later and the group's citation counts jumped.
Does a high cited by count mean a paper is trustworthy?
No. The number counts documents already in Scholar's index whose reference lists its parser matched to that record, and the index contains preprints, theses, repository deposits and papers from journals nobody vets alongside refereed work. Scholar's help says the count falls when citing documents go offline, because Scholar "generally reflects the state of the web as it is currently visible to our search robots". It is a measurement of the web, not a judgement by anybody about the paper.
Is Google Scholar safe to use for a university assignment?
Yes, as a way to find things. It is the best free search across academic literature there is. What it cannot do is vouch for anything it returns, and its interface offers no filter for publication status or peer review, so a refereed article and a student essay in a repository look the same in the result list. Use Scholar to find candidates, then open each one, check where it was published, and read the passage you plan to cite before you cite it.
Read next.

Sources that were opened before they were cited

Free to start. No card needed.

Start writing