AI essay writer: how far a drafted essay gets
An AI essay writer drafts an essay for you, and on an essay it comes closer to working than on anything else you will be asked to write. An essay is short. It usually argues from material you were handed rather than material you had to go and find. It is marked by one person, against criteria your department has probably published. Nearly all of that plays to a language model's strengths, which is why the measured evidence on machine-written essays is less comforting than most guides let on. So this page takes the marking bands seriously instead: what they say the marks are for, how far a draft gets against them on its own, and the one line in the top band that describes something no tool can hand you.
Three pages on this site answer the same question about three different documents, and the answers are genuinely different, so it is worth knowing which one you are reading. A research paper lives or dies on whether its citations are real, and can AI write my research paper owns that evidence. A thesis is supervised for months and examined out loud, and can AI write my thesis owns that. This page is the essay: why it is the easy case, how far a draft gets against a published marking scale, what still gives it away, and what we would do instead.
Why an essay is the easy case
Start with the uncomfortable part, because pretending otherwise wastes your time. In 2023 a team ran a controlled comparison of argumentative student essays against essays written on the same 90 topics by GPT-3 and GPT-4, and had them scored on standard criteria by more than a hundred teachers. The model essays were rated higher for quality than the student ones, and the writing styles came apart in measurable ways rather than in mysterious ones (Herbold et al., 2023, Scientific Reports).
A year later a group at the University of Reading ran the harder version of that test inside their own live examinations system. They opened 33 fake student accounts and submitted entirely AI-written answers into five undergraduate modules across all years of a degree, without telling the markers. 94% of the AI submissions went undetected. The grades those submissions received were on average half a grade boundary higher than the ones real students got, and across modules there was an 83.4% chance that the AI answers would beat a random selection of the same number of real ones (Scarfe et al., 2024, PLOS ONE).
None of that is a reason to relax, and it is not an argument for handing your essay over. It is a reason to stop arguing about whether the prose is good enough, because on the published evidence it is, and to look instead at what the marks are actually being given for.
The band a drafted essay reaches
Most departments publish their marking criteria, and those documents tell you more than any tool's landing page will. The University of Essex's Department of Government publishes its undergraduate criteria as plain grade bands, which makes them a good ruler: the wording is typical of the British scale and unusually direct about what separates one class from the next.
| Band | What the published descriptor asks for | What a draft supplies on its own |
|---|---|---|
| 70 to 80%, first class | "a clear command of material, arguments and sources", a clear understanding of underlying principles "in answering the question", development of argument, and "independence of judgement" | All of it except the last clause |
| 60 to 69%, upper second | "a good knowledge of material, arguments and original and secondary sources", "some grasp of principles and development of argument", a clear point and "some critical acumen" | Usually all of it, in cleaner prose than the band expects |
| 50 to 59%, lower second | "a basic, clear and generally correct knowledge of material", correct summary of it, and "reasonably appropriate conclusions" | Reliably, on almost any prompt, first time |
| 40 to 49%, third | "some knowledge of basic material", used in a way that is "only just adequate", or "ill judged or even mistaken" | Below what a current model does unprompted |
Source: Marking criteria, undergraduate courses, Department of Government, University of Essex.
Read down the middle column and the shape of the problem appears. Every band below the top one is describing correct handling of material: knowing it, summarising it, drawing conclusions that follow from it. Correct handling of material is precisely what a language model is good at. The first-class band adds one clause that none of the others contain, and it is not about prose. It is independence of judgement.
That is not a claim that a machine draft cannot score well. Reading's data says the opposite, and says it loudly. It is a claim about which sentence the marks at the top are attached to, and that sentence is describing the writer rather than the writing.
Where the ladder stops
Three lines in the same document mark the edges of the easy case, and all three are about what the essay is anchored to rather than how it reads.
The title is the baseline. Essex puts it flatly: "The baseline for the assessment of the essay is the title. If this is provided to you by the teacher, you must follow that title as strictly and rigidly as you can." A model handed a title tends to answer the general version of it, because the general version is the one the training data is full of. The drift is small, it survives a proofread, and it is worth exactly the marks that the specificity of the title was carrying.
The assigned reading is assumed. Even at first-year level the minimum is to "refer to the basic literature in the area (assigned texts and other sources)". An essay set on a chapter everyone in the room read is marked by someone who knows what you were given. A draft that discusses the field in general instead of the text in front of you is visible on the first page, with no software involved.
By final year the easy case is over. The third-year minimum adds that students "should show evidence of independent literature searches", and at that point the essay has quietly turned into a research paper: a document whose sources have to be real, found by you, and checkable.
What actually gives it away
The tells that matter are not the ones people worry about. They are these.
The voice changes. The 2023 comparison found the model essays carried fewer discourse and epistemic markers, more nominalisations and greater lexical diversity than the student ones. In plain terms: fewer hedges and signposts, more abstract noun phrases, a wider vocabulary than the writer normally uses. A student who half believes something writes "this might suggest". A model writes "this suggests a tendency toward". A marker who has read your last three essays notices the change of register before they notice anything else, and no tool told them.
The reasoning thins out as the questions get harder. In the Reading study the AI submissions beat real students in every module but one, and the exception was the finalist module. The authors read that as consistent with current models struggling with more abstract reasoning, while noting fairly that only three AI submissions landed in that module. Whatever the size of the effect, the direction is the one you would expect: the further the question moves from summarising a body of material, the less a draft brings.
Quotations from a set text. This is the one place fabrication normally shows up in an essay. Ask for a quotation from a book the model was not given and you get something approximately right, with a page number attached to nothing. It is also the easiest failure in this entire article to prevent: open the text, find the line, type it yourself.
The detector is not what catches you
Most students shopping for an AI essay writer are really shopping for a detector score, which is aiming at the wrong target. A published case at the Office of the Independent Adjudicator, the body that reviews student complaints against providers in England and Wales, shows how the process really runs.
An international student's module assignment was flagged by Turnitin as containing a high proportion of AI-generated content. The provider called the student to a viva, then convened a misconduct panel, which found two offences and imposed a module retake, a 10% reduction in the overall degree mark, and a reflective essay on academic integrity. The OIA reviewed it in July 2025 and found the complaint partly justified. The provider had not explained what evidence led the panel to conclude that AI use amounted to misconduct, had misrepresented what the student said in the viva, and had not considered whether Turnitin's AI detection might be less reliable for non-native English speakers. The recommendation was to hear the allegations again from scratch.
Two things follow, and they point in opposite directions from the usual advice. The detector is a trigger for a conversation, not evidence in itself, and a provider that treats it as evidence can lose. But winning that argument is not an outcome anybody wants. The viva happened, the panel sat, the penalty was imposed, and the adjudication came later. What you are protecting yourself against is not a score. It is being unable to answer questions about your own essay in a room.
What we would do instead
Here is our position, and it is narrower than the usual advice to treat AI "as a tool". Do not let it draft until you have written down the argument. Not because drafting first is dishonest, and not as a rule about effort, but because of where the marks sit. The top band is paid for independence of judgement, and a draft whose position you did not choose deletes the only stage of the work where that happens. You spend the rest of the evening improving somebody else's argument, and what you hand in is fluent, correct and merely competent, which is the middle of the pile. Write the sentence you are arguing first, in your own words, badly if necessary. Then hand over paragraphs rather than the essay, and keep deciding what each one is for.
The rest is ordinary craft. Work in sections so you can actually read what comes back. Keep the assigned text open and quote it yourself. Check that the essay answers the exact title you were set rather than its more famous cousin. If you want the version of this with a clock attached, how to write an essay fast lays out the time-boxed plan, and where the line falls between support and misconduct is covered in whether using AI to write essays is cheating. If you are still choosing a tool, our round-up of AI tools for academic writing ranks the main names by whether they retrieve anything before they write.
The same question about a longer document
Everything above holds because an essay is short, argued from material you were given, and read by one person. Change any of those and the answer changes with it.
Add real sources and the failure moves to the citation layer, where it can be measured. That page carries the audits: what people found when they actually checked a machine's reference list, and the number that sits underneath the famous fabrication number. It is can AI write my research paper.
Stretch it over months with a supervisor watching and an oral examination at the end, and the answer turns to mostly no for reasons that have nothing to do with text quality. That is can AI write my thesis, which walks through the documents you sign and the room you sit in. If detectors are the part you are actually worried about, will AI detectors flag my writing covers what those tools do and do not measure.
CiteOwl is built for the version of this where the sources have to be real. It searches actual literature, reads what it finds, and links every claim it writes to a source it read, with the supporting quote shown, and each edit arrives as a plain diff you accept or reject. On a first-year essay from a set text that is more machinery than you need. On everything after it, it is the part that stops being optional.