The chain from a claim to a paper you can open
Every problem the twelve guides below deal with is a break in one chain, and the useful question is never "is this citation good" but "which link snapped". There are five, and each one fails in its own way.
- The claim. A sentence in your draft asserts something. Either it needs outside backing or it does not. Your own reasoning, your own arithmetic and genuinely common knowledge need none, and a citation bolted onto them is noise.
- The reference string. Authors, year, title, journal, DOI. This is the link a chatbot can produce out of nothing, because a reference is a pattern and patterns are exactly what a language model learned.
- The metadata record. The string either matches a record some publisher deposited or it matches nothing. This is the first link that can be checked mechanically, in one query, for free.
- The resolver. A DOI is not a web address, it is a key.
doi.orglooks it up and forwards you to wherever the publisher currently hosts the work, which is why a DOI survives a journal moving its website and a raw URL does not. A DOI that resolves proves a record exists. It proves nothing about what is written in the paper. - The paper. You open it and read enough to confirm it supports your sentence. This is the link almost everyone skips and the one a marker actually tests, because it is the only one where reading is unavoidable.
Two things can still invalidate a citation after all five links hold: the paper was retracted after it was published, or the journal it appeared in is not doing the peer review it advertises. Both have their own register, and both are checkable in under a minute.
What each register holds, and what it cannot see
The registers behind links three, four and the two after-the-fact checks are public, and most of them answer an anonymous query with no account and no key. We ran those queries on 11 August 2026 and the counts below are what each register's own API returned that day, so treat them as a snapshot rather than a constant. Every one of them grows daily.
| Register | What it holds | Count on 11 August 2026 | What it cannot tell you |
|---|---|---|---|
| Crossref | Metadata deposited by publishers: journal articles, book chapters, conference papers, preprints | 185,346,103 records | Whether the paper supports your sentence. A record is a filing, not a reading |
| DataCite | DOIs for datasets, software, theses and repository deposits | 133,322,923 DOIs | The same. It also holds a lot of grey literature that was never peer reviewed |
| OpenAlex | An open index built on Crossref plus repositories, with authors, venues and citation links | 323,998,640 works | Full text it has no open copy of. It knows the work exists, not what page 7 says |
| Retraction Watch | Retractions, expressions of concern, corrections and reinstatements, hosted by Crossref | 71,686 records, of which 66,197 are retractions | Whether the paper is wrong. Retraction is an act by a publisher, and its reasons range from fraud to a duplicate submission |
| DOAJ | Open access journals that passed its own vetting | 23,303 journals | Whether any particular article in one of them is any good |
Do not read those four counts as a ranking. They describe overlapping and differently shaped universes: OpenAlex is larger than Crossref partly because it merges Crossref with repository and preprint records, so the same paper can arrive twice from different directions and be reconciled, and DataCite's DOIs are mostly not articles at all. What the table is for is the last column. Every register answers exactly one question, and a reference that clears all of them can still be a reference to a paper that does not say what you wrote.
One of these is newer than students realise. Retraction Watch spent years as a privately funded database; Crossref bought it in September 2023 and released the whole dataset publicly, which is the only reason checking a retraction is now a free lookup instead of a subscription. The counts above came from Crossref's REST API, DataCite's, OpenAlex's, the Retraction Watch dataset and DOAJ.
Where each guide plugs into the chain
Read in this order if you are working through a reference list you do not trust. Read the one that matches your link if you already know where it broke.
- Link 2, why the string is fake. Why AI makes up citations covers the mechanism and the measured rates, and how to get ChatGPT to cite real sources covers what actually reduces it.
- Link 3, does the record exist. How to check if a citation is real is the six-check routine against the registers above.
- Link 4, the resolver. A DOI that does not work separates the five different reasons a syntactically valid DOI resolves to nothing, and the DOI that links to the wrong paper covers the harder half, where it resolves perfectly to somebody else's article. My reference has no DOI is the case where there is nothing to resolve, which is a correct reference more often than students think.
- Link 5, does the paper say it. A real paper that says something else is the failure no register catches, and it is more common than outright fabrication. A made-up quote is the same failure in quotation marks, where the words are checkable against the page in a way a paraphrase is not.
- Where the reference came from. I cited a paper I never read is the habit under most of the failures above, and how to cite a source you found in another paper is the sentence that makes it honest in APA, MLA and Chicago.
- After the fact. Is the paper I cited retracted and how to tell if a journal article is fake cover the two checks that survive a clean chain.
- Habit and consequences. Cite as you write keeps link 1 honest, how to cite AI tools handles the formatting once you know the source is real, how to cite your own previous paper covers the reference students are most likely to leave out, and my professor says my sources do not exist is the guide for when the checking happens too late.