CiteOwl
CiteOwl

How to find sources for a research paper

Start with your library's databases, then Google Scholar, then free indexes like OpenAlex and CORE when you hit a paywall. Finding is the easy half. Two failures kill a source and only one is visible in a results list. A paper that does not exist announces itself the moment you open it. One that answers a neighbouring question survives every check except reading it. Open each source and read the part you will cite.

Turn your topic into search keywords

Before you open any database, turn your research question into search terms, because the words you type decide what comes back. Pull the two or three core concepts out of your question and, for each, list synonyms and related terms: "commuters" is also "travellers" and "passengers," and "congestion pricing" sits next to "road pricing" and "cordon charging." A database matches the word, not the idea, so the concept you only phrase one way is the one whose best papers you never see.

Then combine the terms with the operators every academic database understands. Quotation marks force an exact phrase, so "congestion pricing" stays together instead of matching any page with both words scattered around. AND narrows a search to results that contain both concepts; OR widens it to catch either synonym; and most databases have an advanced search where you stack these on separate lines. Two habits keep you out of the weeds: if a search returns thousands of hits, add a third concept or a date limit, and if it returns almost nothing, drop the narrowest term or swap in a broader synonym. A few minutes building a keyword list up front saves an hour of scrolling past results that only brushed your topic.

Start with your library's databases

Before you open Google, open your university library. It is the single most underused resource a student has, and it beats a plain web search for one concrete reason: access. Your tuition already pays for subscriptions to academic databases and full-text journals that the open web hides behind paywalls. JSTOR, Scopus, Web of Science, PsycINFO, PubMed, and dozens of subject-specific databases are sitting there, logged in and free to you, full of peer-reviewed work you simply cannot reach with a normal search.

The library also solves a problem Google creates. A web search returns whatever is popular and public; a subject database returns whatever is scholarly and relevant, filtered by discipline. If you are writing about psychology, a psychology database has already excluded the news articles, marketing pages, and forum posts that a general search drowns you in. You spend your time reading research instead of sifting noise.

Two practical moves. First, find your subject's database rather than the catalogue's generic search box; a librarian or the library's subject guides will point you to it, and asking a librarian for help with a real research question is the highest-leverage thing on this whole list. Second, look for the link resolver, usually a "Find it" or "Full text" button that takes a reference you found anywhere and checks whether your institution has access. International and European students should note that access follows your enrolment, not your country, so log in through your institution and most paywalls quietly disappear.

Use Google Scholar well

Google Scholar (scholar.google.com) is the best free starting point for almost any topic, and most people use maybe a tenth of it. A few habits turn it from a search box into a real research tool.

One caution. Google Scholar indexes broadly, which is its strength and its weakness. It includes preprints, predatory journals, and the occasional non-scholarly source alongside the good material, so it finds things but does not vouch for them. Treat it as a discovery engine, not a quality filter. The quality judgement is still yours, and we get to it below.

Free open databases that actually work

If your library access is thin, or you just want broader coverage, several free academic databases are genuinely good and ask for no login. These are also where a careful AI research tool searches, because they expose real metadata instead of guesswork.

A practical workflow that uses these well: discover broadly in Google Scholar or OpenAlex, then when you hit a paywall, paste the title into CORE, Semantic Scholar, or your library's full-text finder to track down a version you can actually open. A source you cannot read is a source you cannot cite honestly. If your brief specifically demands peer-reviewed work, our guide on how to find peer-reviewed articles shows which of these tools surface refereed papers and how to confirm a journal is genuinely peer-reviewed.

How many sources, and what kinds

There is no magic number, and chasing one is a mistake. You will find ranges quoted confidently all over the internet, and we went looking for what is behind them: they contradict each other by a factor of five, two of the pages asserting one contradict themselves within a paragraph, and almost none cites a study. We wrote up what university guides actually say in how many sources a research paper needs. The real rule is simpler: one solid source for every claim that needs backing. A tight paper built on a dozen relevant sources reads better than a padded one stuffed with forty weak ones. If your brief names a number, treat it as a floor and aim past it only with sources that earn their place.

Kind matters more than count. Two distinctions are worth keeping straight:

Judge a source fast

You cannot read every hit in full before deciding whether to keep it, so run a gate instead. Credibility is not a single yes or no: it sits on a scale, and the same source can be perfect for one claim and useless for another. A news report is a fine source for what happened and when; it is a weak source for the mechanism behind a scientific finding. What you are matching is the source to the weight you are asking it to carry.

Peer review, in one minute. Before a journal runs a study it sends the manuscript to other experts in the field, who probe the method, the analysis, and whether the conclusions follow, then recommend accepting, revising, or rejecting it. Work that survives has been pressure-tested by people with no stake in flattering the author, which is why "use peer-reviewed sources" really means build your argument on that checked layer of the record instead of whatever ranks well on the open web. It is strong evidence, not a guarantee: reviewers miss things and papers still get retracted, so judge the individual paper too. Confirming that a given article really is refereed, using database limiters, Ulrichsweb and the journal's own statement, is its own short procedure, laid out in how to find peer-reviewed articles.

Scholarly or popular. Most credibility judgements collapse into this one distinction, and the signals are easy to learn.

SignalScholarly sourcePopular source
AuthorNamed experts with credentials and affiliationsJournalists, staff writers, or no byline
AudienceOther researchers and studentsThe general public
ReferencesA full reference list you can followFew or none, links at best
ReviewPeer-reviewed before publicationEdited for style, not refereed
LanguageTechnical, precise, field-specificAccessible, simplified
PurposeReport original research or analysisInform, entertain, or sell

Popular sources are not banned. A well-reported news article can introduce a topic, point you toward the underlying study, or document a current event you genuinely need to cite. Use them for what they are good at, and chase the real research before you rest a factual claim on anything.

The CRAAP test, applied. When the scholarly-versus-popular line is not enough, run the source through the five-part checklist a Chico State librarian, Sarah Blakeslee, published in 2004 and writing centres everywhere now teach. Thirty seconds per source catches most of what should never make it into a paper.

Worked example: you find a page titled "Does insulation actually cut heating bills?" with no author, dated 2014, on an insulation installer's site, citing nothing. Currency: dated, and the field has moved on. Authority: no author, a commercial domain. Accuracy: no references to follow. Purpose: it exists to sell insulation. It fails four of five, so you set it aside and go looking for the peer-reviewed retrofit studies it was loosely paraphrasing.

One thing the checklist cannot see: a journal that fakes the badge. Predatory journals publish anything for a fee and dress the result in the furniture of real scholarship, and hijacked clones impersonate a legitimate title outright. The checks that separate them from the real record start with the ISSN, and they are in how to tell if a journal article is fake.

Verify it is real before you cite it

Here is the step that has quietly become essential. Finding a source that looks right is no longer the same as finding one that exists. AI assistants confidently produce references with a believable author, a real-sounding title, a journal, and a correctly formatted DOI, and a share of them point to nothing at all, for reasons baked into how these models work. The share is not small: a 2023 study in Scientific Reports found that 55% of GPT-3.5 citations and 18% of GPT-4 citations were entirely fabricated. The reference fits the shape of a real one without being one, and you only find out when you try to open it.

So before any source enters your paper, confirm it exists. The checks take seconds each:

A working link is not proof. Fabricated references sometimes carry a DOI that resolves to a real but unrelated paper, so the link loads and the citation is still wrong. The only check that cannot be faked is opening the source and confirming it says what you claim. Apply this to every AI-suggested source without exception.

This matters most for anything a chatbot suggested, but the habit is the same for a reference you found in another paper's bibliography or got from a classmate. We lay out the full method, including the edge cases, in how to check if a citation is real.

Track where every claim came from

The reason verifying later feels painful is almost always that you did not track sources as you went. You find a perfect quote on Tuesday, paste it into your notes, and by the time you write the paragraph on Friday you have no idea which paper it came from. Now you are reverse-engineering your own research, and that is when people give up and cite from memory, which is exactly how a half-remembered or invented reference slips in.

Fix it at the source. The moment you decide to use something, record enough to rebuild the citation without going back: title, authors, year, journal or publisher, the DOI or a stable link, and the page or section the claim sits on. A reference manager like Zotero (free and open source) automates most of this and exports a formatted bibliography at the end, but even a plain document where every note carries its source works. The rule is that no claim should ever exist in your draft without a trail back to where it came from. If you keep that trail, citing at the end is transcription. If you do not, it is detective work, and detective work is where mistakes live.

Let the workflow do this for you

Finding, judging, verifying, and tracking sources is the honest core of a research paper, and none of it can be skipped. It can be made faster. CiteOwl is built around exactly this workflow. It searches real academic databases, the same OpenAlex and web and DOI infrastructure you would use by hand, reads the papers it finds, and ties every claim in your draft to the source it actually retrieved, with the verbatim quote shown so you can confirm the source says what the sentence claims. The finding and the verifying are built in, not bolted on after. You still own the argument and the words. You just never paste in a reference you have not seen.

Read next.

Write the paper from sources you have actually read.

CiteOwl is an editor whose AI finds real papers, reads them, and writes from what it read.

Start writing