Can AI write my research paper? Yes, but not as an author
Yes, AI can draft a research paper, and a good one can do it well. What it can never be is the author, and that is what everything below follows from: a language model cannot be accountable for a claim, so the accountability is yours whether or not you read the sentence before you handed it in. A research paper is the document where that lands on the citations, because a research paper is the one that has to be built out of real literature. The famous audit of 636 model-produced references found 55% of the GPT-3.5 ones fabricated. The same audit found something worse buried underneath it, and this page is mostly about that second number, because it is the one your verification routine does not catch.
- The famous number, and the one underneath it
- What people found when they checked
- The check that passes while the paper fails
- Nobody will let the model take the blame
- Point it at the retrieval end
- A workflow built around the check
- The same question about a shorter and a longer document
- So, can AI write your research paper?
Three pages on this site answer the same question about three documents, and the answers differ because the documents do. An essay is short, argued from what you were handed, and marked on argument, which makes it the easy case, and how far a drafted essay gets takes the marking bands apart. A thesis is supervised for months and examined out loud, so the answer there is mostly no, which is can AI write my thesis. This page is the research paper, where the whole question collapses into whether the sources are real and whether they say what you claimed. Below: what the audits found, the check that passes while the paper still fails, why no one will let the model take the blame, and where we would point it instead.
The famous number, and the one underneath it
The number everyone quotes comes from a 2023 audit that took 636 references produced by ChatGPT and checked each one against bibliographic records. 55% of the GPT-3.5 references and 18% of the GPT-4 references referred to no real publication at all (Walters & Wilder, 2023, Scientific Reports). That is the statistic that made the rounds, and it is the one people design their defences around: check the reference exists, and you are safe.
The same audit reports a second number that almost nobody quotes. Among the references that were real, 43% of the GPT-3.5 ones and 24% of the GPT-4 ones carried substantive citation errors: wrong authors, wrong year, wrong journal, wrong pages. A quarter of the surviving references from the better model were still wrong about the paper they named, and every one of them passes an existence check, because the paper exists.
Hold those two numbers next to each other and the shape of the problem changes. Fabrication is loud and catchable. Misattribution is quiet, survives the check most students actually perform, and is what a marker finds when they do the thing a script cannot do, which is read the source.
What people found when they checked
Reference lists get audited more often than you would think, at very different scales, and the results are worth putting in one place. These are the checks that have been published, what each one looked at, and what came back.
| Audit | What was checked | What came back |
|---|---|---|
| 2023, Scientific Reports | 636 references generated by ChatGPT, matched against bibliographic records | 55% (GPT-3.5) and 18% (GPT-4) referred to nothing real. Of the ones that did exist, 43% and 24% carried substantive errors |
| April 2026, preprint | Over 221,000 citation URLs emitted by 10 commercial models and deep research agents | 3% to 13% of citations had no historical record; 5% to 18% did not resolve. Agents that search cite far more and hallucinate URLs at a higher rate. Handing the models a tool that checks whether each URL resolves cut non-resolving citations to under 1% |
| May 2026, preprint | 111 million references across 2.5 million papers on arXiv, bioRxiv, SSRN and PubMed Central | A conservative estimate of 146,932 non-existent references in 2025 alone, rising sharply after widespread model adoption |
Sources, in order: Walters & Wilder (2023); Rao, Wong & Callison-Burch, "Detecting and correcting reference hallucinations in commercial LLMs and deep research agents", posted 3 April 2026; Zhao et al., "LLM hallucinations in the wild", posted 8 May 2026. The last two are preprints and have not been through peer review, which is worth saying out loud on a page about citation integrity.
Three things fall out of the table. Newer models fabricate less but do not stop. Tools that go and search cite more sources than tools that do not, which means a higher raw count of bad links even where the rate improves. And the 2026 numbers are measured in the published literature rather than in a lab, which is the part that should bother anyone about to add references to a pile that is already carrying six figures of them.
The most useful line in that table is the last clause of the middle row. Once the models were given a tool that simply asks whether each URL resolves, non-resolving citations dropped to under 1%, from as much as 18%. The failure is cheap to catch the moment anything is actually looking. It survives because usually nothing is.
The check that passes while the paper fails
Most verification advice, ours included, tells you to confirm a reference exists: search the title, resolve the DOI, check the author publishes in the field. That method is correct, it takes a couple of minutes per reference, and we have written it up in full in how to check if a citation is real. What the 2023 audit's second number tells you is that clearing it is not the same as being right.
A real paper attached to a claim it does not make passes every check in that method. The DOI resolves. The authors publish in the field. The title is exactly what your reference list says. The only thing wrong is the sentence in front of it, and finding that requires opening the paper and reading the passage the claim is supposed to rest on. That is also, inconveniently, exactly what a marker does when a claim surprises them.
So the question to ask of each reference is not "does this exist" but "does this source say the thing my sentence says it says". In practice that means keeping the supporting line, the actual sentence or two from the paper, attached to your claim while you draft, so the check is a glance rather than a hunt. Do that and the mechanism behind fabrication stops mattering much, though it is worth understanding once, and that is what why AI makes up citations covers.
Nobody will let the model take the blame
The reason this lands on you rather than on the tool has been written down, repeatedly, by the people who run academic publishing. The Committee on Publication Ethics stated its position on 13 February 2023: AI tools cannot meet the requirements for authorship, because they cannot take responsibility for the submitted work. As non-legal entities they cannot declare a conflict of interest and cannot hold or assign copyright.
Springer Nature's guidance draws the operational line that follows from it, and the line is more precise than most course policies manage. A model cannot be listed as an author. Where one is used for substantive work, the use has to be documented in the Methods section or its equivalent. And AI-assisted copy editing, grammar, spelling, tone and formatting, requires no declaration at all.
That distinction is worth carrying into your own coursework even though no journal is involved. Polishing sentences you wrote is invisible and uncontroversial. Generating content is disclosable. Neither route moves the accountability, because accountability is the one thing a model structurally cannot hold, which is why every version of this rule ends in the same place: your name is on the citation. If you want the practical comparison of working this way against a chat window, it is in CiteOwl versus ChatGPT for research papers.
Point it at the retrieval end
Here is the part we would argue for, against a lot of the advice out there. The instinct after reading all this is to use AI less, and we think that is the wrong lesson. What actually matters is which end of the pipeline it touches. A model that drafts prose from what it already believes is dangerous exactly in proportion to how fluent it is, because fluency is the thing that stops you checking. A model that goes and fetches papers, then writes from what it just read, moves the risk somewhere you can inspect: a wrong source is a link you can open and a claim you can contradict, whereas an invented source is a sentence with nothing behind it at all. So we would rather see a student use AI heavily for retrieval and reading, where the failure mode is visible, than lightly for drafting, where it is not. The rule is not less AI. It is AI pointed at the part of the work that leaves evidence.
The middle row of that table is the honest complication, and we are not going to pretend it away. Retrieval does not make the number zero, and search agents produce more bad links in absolute terms precisely because they produce more links. What changes is that every one of those links is a thing you or a script can resolve in a second, which is how the same study got its models under 1%. That is the whole argument in one measurement: the point of retrieval is not that it never gets it wrong, it is that being wrong leaves a trace.
A workflow built around the check
What follows from all of the above is a way of working, not a rule about how much help is allowed.
You set the argument. Decide what the paper claims before anything drafts a word, because the claim is what the sources are being gathered for. If the paper does not have a position yet, how to write a research paper covers getting from a question to one, and a research paper outline tells each section what it has to do.
Retrieve before you write, section by section. Sources first, sentences second, in pieces small enough to actually read. A literature review, a methods section and a discussion are three different conversations and asking for all of them at once is how unchecked material gets in.
Keep the supporting line with the claim. Not the reference, the passage. This is the single habit that converts the misattribution problem from invisible to trivial, and it costs nothing at the moment the claim is written and a great deal at three in the morning three weeks later.
Review each change as a change. A block of finished text hides the decisions inside it. A diff you accept or reject does not, which is the difference between adopting an argument and reading one.
Open the sources the argument leans on. Not all of them. The three or four the paper would collapse without. Detection is not the thing to optimise for, and will AI detectors flag my writing explains why, but a claim you cannot trace is a real risk regardless of what any tool says about your prose. Where the institutional line falls is covered in whether using AI to write essays is cheating.
The same question about a shorter and a longer document
This page is about the research paper, because the research paper is where the citation layer carries the weight. The answer moves when the document does.
An essay is shorter, usually argued from material you were handed rather than material you had to find, and marked on the argument rather than the sourcing. That makes it the case where machine drafting comes closest to working, and the published marking bands show exactly how close and where it stops. We take that apart in how far a drafted essay actually gets.
A thesis runs for months under a supervisor, ends in an oral examination and comes with a declaration you personally sign. There the answer is mostly no, for reasons that have nothing to do with how good the text is, and can AI write my thesis walks through the paperwork and the room.
So, can AI write your research paper?
Yes, and it can write a good one. It cannot be the author of it, and on the evidence above it cannot be trusted with the reference list unaided, whether the failure is a source that does not exist or a real source that does not say what you claimed. Both are recoverable, and both are recoverable by the same move: make the tool fetch the paper before it writes the sentence, and keep the passage that backs the claim where you can see it.
That is what CiteOwl is built to do. It searches real literature, reads what it finds, and writes claims it can attach to a source it actually read, with the supporting quote shown, and every edit arrives as a plain diff you accept or reject. The paper that comes out is one you have read line by line, and every citation in it points at something you can open.