Six search surfaces, and the hole in each one
"Search the literature" sounds like one activity. It is at least six, run over indexes that were built for different purposes and that miss different things. Picking the wrong one is not a small mistake: a surface's blind spot is invisible from inside it, because a search engine cannot show you what it never indexed.
Counts below marked as measured are what that index's own public API returned when we queried it on 11 August 2026. Counts marked as claimed are the vendor's own published figure, which is a different kind of fact and is treated as one.
| Surface | What it indexes | Size | What it misses |
|---|---|---|---|
| Your library | The discovery layer, usually Primo, Summon or EBSCO Discovery: everything your institution licenses, merged with open content, in one search box | Depends entirely on what your library buys | Anything your library does not license. It is also the only surface that knows what you can actually open, which is why it should usually be first |
| Google Scholar | Its own description: "articles, theses, books, abstracts and court opinions, from academic publishers, professional societies, online repositories, universities and other web sites" | Google publishes none, and offers no API | Unknowable, and that is the problem. No coverage list means nobody outside Google can audit what is absent. It also mixes preprints, predatory journals and several versions of one paper without labelling them |
| OpenAlex | An open index of works, authors, venues and citation links, built on Crossref plus repositories | 323,998,640 works, measured | Full text it has no open copy of. It can tell you a paper exists and cannot tell you what is on page 7 |
| Crossref | Metadata publishers deposited themselves, keyed by DOI | 185,346,103 records, measured | Anything nobody registered: many books, much of the humanities, a lot of non-English regional publishing |
| Semantic Scholar | Papers from publisher partnerships, data providers and web crawls, with extracted citation context | "over 200 million academic papers", claimed | The same deposit gaps, plus whatever its crawl did not reach |
| Scopus and Web of Science | Curated title lists. Clarivate says the Web of Science Core Collection indexes "22k+ peer-reviewed journals" and holds "99m+ records"; Elsevier says Scopus holds "over 217,000 book titles" from "more than 7,000 publishers" | Vendor figures, claimed | Everything off the list, on purpose. Both are subscriptions, so you have them only while you are enrolled |
Sources for the vendor figures: Google Scholar's about page, Semantic Scholar, Clarivate and Elsevier, whose Scopus page states its numbers are "current as of July 2025". The measured counts came from OpenAlex and Crossref.
Field databases are missing from the table on purpose. PubMed for the life sciences, ERIC for education, IEEE Xplore for engineering and their equivalents are not competing on size, they are competing on a boundary: everything in them belongs to one field, and that is the whole product. If your topic has one, it belongs in your search alongside a general surface, never instead of one.
Which one to open first, by what you are actually doing
- You need to read it this week. Start in your library's discovery layer. Every other surface will show you papers you cannot open, and a source you cannot read is a source you cannot cite honestly.
- You do not know the vocabulary yet. Start in Google Scholar. It is the best free discovery surface there is and the worst one to audit, so use it to learn what your field calls the thing, then take those words somewhere you can check.
- You are checking whether something exists. Use Crossref or OpenAlex. They are the only two here that answer an anonymous query, for free, with a number you can reproduce, which is exactly what verification needs.
- You need a defensible search for a systematic review or a methodology chapter. Use Scopus or Web of Science. A published title list is a limitation everywhere else and an asset here, because it is what makes a search reproducible by your examiner.
- Whatever you are doing, use two. Every surface above has a shaped hole, and the shape differs. Two surfaces will not cover everything, but the papers only one of them returns are the first evidence you have that a hole exists at all.
Where each guide fits
The twelve guides below run in the order the work does: search, judge, read, keep.
- Search. How to find sources for a research paper is the full routine; how to find peer-reviewed articles covers the surfaces above in practice and how to confirm a journal really is refereed.
- Judge. How to find and judge sources is the CRAAP pass, plus scholarly against popular and primary against secondary. Are there fake papers on Google Scholar is the same judgement applied to the index itself, and can I cite a preprint covers the source that has not been judged by anyone yet.
- Read. How to read a research paper is the three-pass method, which exists because nobody reads one top to bottom.
- Keep. The annotated bibliography and the worked literature review example turn a pile of PDFs into something you can write from, and the AI writing tools hub covers doing that with a tool without losing the traceability.
- Backwards. Finding a source for a claim you already wrote is the guide for the order everyone actually ends up in at least once, including what to do when nothing supports the sentence.