Your Turnitin AI score: 4% of flagged sentences are human
Our own read, for what it is worth: the sentence-level 4% is the figure that should decide this, and it almost never gets quoted. The number is a prediction, not a measurement, and Turnitin publishes how often it misses: around 4% of the sentences it highlights as AI-written are human-written, roughly one in every twenty-five. Turnitin itself states that the indicator does not make a determination of misconduct. What decides your case is your own institution's policy, and a growing list of universities have switched the indicator off entirely.
If your score is high and you wrote the essay, do not start rewriting sentences. Editing after the fact destroys the one thing that helps you, which is the record of how the document was built. More on that below.
The two false positive rates Turnitin publishes
Turnitin publishes two false positive rates, and only the smaller one travels. The document-level rate, the one everybody quotes, is under 1%, with the caveat that the figure applies to documents where the tool reports 20% or more AI writing. In the same explanation Turnitin also gives a sentence-level false positive rate of around 4%: "there is a 4% likelihood that a specific sentence highlighted as AI-written might be human-written." That second number is the one you are actually looking at when someone opens the report and points at a highlighted paragraph. One sentence in twenty-five of that highlighting is wrong by the vendor's own account. It has also acknowledged publicly that real-world use produced different results from its lab testing, and it advises instructors to apply professional judgment rather than treat the score as a verdict.
The number that sounds small stops sounding small at scale, and that arithmetic is why some universities stopped showing the indicator at all. Vanderbilt disabled the detector in 2023 after working out that a 1% false positive rate, against the roughly 75,000 papers its students submitted in 2022, implies around 750 wrongly flagged. Which institutions have actually followed is a separate question with its own page here, and it is answered strictly: the universities that turned off Turnitin's AI detection lists only the ones we could check against the institution's own published announcement, seventeen of them, and lists the widely repeated names we could not verify separately. Northwestern and Johns Hopkins are in that second group, named in roundups and in one news report but not on any page either university publishes, so treat them as unconfirmed rather than settled. Plenty of institutions kept the feature and wrote rules around it instead, which is why your own institution's page is the fact that matters and a blog post cannot answer it for you. Ten minutes on your academic integrity pages tells you whether the number you are staring at is a trigger, a talking point or nothing.
A document-level rate answers a question no student is ever in the room for. The conversation students actually have is an instructor with a report open, pointing at highlighted sentences, and at that granularity the tool tells you itself that one in twenty-five of those highlights is a mistake. On a report carrying twenty-five highlights, the vendor's own arithmetic puts about one of them in the wrong, and nobody in the room knows which one. We would not run a product on a signal that noisy, and the universities switching it off are not being squeamish. They are reading the numbers the vendor published.
Detectors also do not err evenly. Liang and colleagues, publishing in Patterns, found that GPT detectors are biased against non-native English writers: a majority of TOEFL essays written by non-native speakers were misclassified as AI-generated, while essays by US school students were scored accurately. Their explanation is that these tools lean on statistical predictability, and writing that uses a smaller, plainer vocabulary reads as more predictable. Turnitin was not one of the seven detectors that study tested, so it is not a measurement of your own report, and what the study actually measured is taken apart line by line on the page that carries it in full. The mechanism it identifies is not product-specific, which is the part worth knowing if you write in your second language.
What the number counts
Turnitin splits your submission into segments of prose and scores each one, then reports the percentage of qualifying text predicted to be AI-generated. The result appears as an AI writing indicator alongside the similarity score, with a linked report highlighting the specific segments the model flagged.
Two details change how the number should be read. The tool works on long-form prose, so it is not designed for lists, tables, code, equations, heavily quoted text or very short submissions. And Turnitin marks low scores as less reliable: it displays an asterisk when the reported figure falls under 20%, signalling that at that end of the scale the prediction is noisier. A 6% score and a 60% score are not the same kind of claim.
The flagged segments matter more than the headline percentage. If you are shown the report, look at what got highlighted. Two flagged sentences in a methods paragraph and a flagged introduction are very different situations.
What it does not measure
This is where most of the confusion lives, so it is worth being blunt about each one.
It is not the similarity score. Similarity compares your text against a database of existing documents and reports overlap. The AI indicator compares nothing to anything. It predicts from style alone. The two numbers are independent, and what Turnitin can and cannot detect is a different question for each.
It cannot see whether you used AI. There is no log, no watermark, no trace of your session. The model is reading the finished text and guessing at its origin. That is why a student who used a chatbot heavily can score low and a student who wrote every word can score high.
It is not a probability that you cheated. Even taken at face value, the number is about text, not conduct. Plenty of legitimate AI use is disclosed, permitted, or limited to things nobody objects to. Whether it counts as misconduct depends on your course's rules, which is a policy question rather than a detection question.
It does not reliably survive rewriting. Paraphrasing tools and so-called humanizers exist specifically to move the number, which tells you how loosely the number is coupled to the underlying fact. We looked at whether humanizers actually work separately. The relevant point here is that a metric you can change without changing how the text was produced is not measuring how the text was produced.
What a score does and does not prove
A high score proves that a statistical model, reading your finished text, found it resembled AI-generated writing. That is a reason for a conversation. It is not, on its own, evidence of what you did, and Turnitin's own guidance is that the final judgement rests with the instructor and the institution's policy, not with the number.
A low score proves less than students hope. It says the model did not find the pattern. Text that has been paraphrased, edited or written with a lot of back-and-forth help can score low regardless of how it was produced. Nobody should treat a low score as a clearance certificate, and if your course requires you to disclose AI use, a low number does not remove that obligation.
What is left, in both directions, is process. How the document was written is a real fact with real traces: drafts, version history, notes, the sources you actually opened. That evidence answers the question the score only gestures at.
If your score is high and you wrote it
Two things tonight, and the rest belongs to a different guide. Do not edit the submitted work, because changing it after the fact reads badly and destroys your record. And preserve the process evidence while it still exists: Google Docs and Word version history, earlier drafts, outlines, handwritten notes, the PDFs you read, exported or screenshotted somewhere separate from your working folder.
Everything after that, how to answer the email, what to ask for in writing, what an academic integrity meeting is actually like, and what changes if you did use AI, is the whole subject of what to do when you are accused of using AI on an essay. If you are worried before you submit rather than after, will AI detectors flag my writing? covers what tends to raise scores on genuinely human work, while what actually gives a ChatGPT essay away covers what human readers notice, which is a different list.
The thing detectors will never see
Everything above is about style, and style is the weakest available signal. The things that get people into real trouble are checkable facts: a reference that does not exist, a quote nobody wrote, a source cited for a claim it never made. Those survive any amount of rewriting, and a marker who looks up one reference finds them in a minute. Guarding against fabricated citations protects you against a risk that is concrete, provable and entirely avoidable. Detector scores come and go.