What does my Turnitin AI score actually mean?
The number is a prediction, not a measurement. It reports the share of your document's qualifying prose that Turnitin's model predicts was generated by a language model, based on how the writing reads. It has no access to your drafts, your browser or your process. Turnitin itself states that the indicator does not make a determination of misconduct and should not be the sole basis for action. Before you do anything else, find out what your institution's policy says the number actually triggers, because that varies enormously and it is the thing that affects you.
If your score is high and you wrote the essay, do not start rewriting sentences. Editing after the fact destroys the one thing that helps you, which is the record of how the document was built. More on that below.
What the number counts
Turnitin splits your submission into segments of prose and scores each one, then reports the percentage of qualifying text predicted to be AI-generated. The result appears as an AI writing indicator alongside the similarity score, with a linked report highlighting the specific segments the model flagged.
Two details change how the number should be read. The tool works on long-form prose, so it is not designed for lists, tables, code, equations, heavily quoted text or very short submissions. And Turnitin marks low scores as less reliable: it displays an asterisk when the reported figure falls under 20%, signalling that at that end of the scale the prediction is noisier. A 6% score and a 60% score are not the same kind of claim.
The flagged segments matter more than the headline percentage. If you are shown the report, look at what got highlighted. Two flagged sentences in a methods paragraph and a flagged introduction are very different situations.
What it does not measure
This is where most of the confusion lives, so it is worth being blunt about each one.
It is not the similarity score. Similarity compares your text against a database of existing documents and reports overlap. The AI indicator compares nothing to anything. It predicts from style alone. The two numbers are independent, and what Turnitin can and cannot detect is a different question for each.
It cannot see whether you used AI. There is no log, no watermark, no trace of your session. The model is reading the finished text and guessing at its origin. That is why a student who used a chatbot heavily can score low and a student who wrote every word can score high.
It is not a probability that you cheated. Even taken at face value, the number is about text, not conduct. Plenty of legitimate AI use is disclosed, permitted, or limited to things nobody objects to. Whether it counts as misconduct depends on your course's rules, which is a policy question rather than a detection question.
It does not reliably survive rewriting. Paraphrasing tools and so-called humanizers exist specifically to move the number, which tells you how loosely the number is coupled to the underlying fact. We looked at whether humanizers actually work separately. The relevant point here is that a metric you can change without changing how the text was produced is not measuring how the text was produced.
How it behaves on human writing
Turnitin has published its own position: it says it aims to keep the document-level false positive rate under 1%, with the caveat that the figure applies to documents where the tool reports over 20% AI writing. It has also acknowledged publicly that real-world use produced different results from its lab testing, and it advises instructors to apply professional judgment rather than treat the score as a verdict.
The number that sounds small stops sounding small at scale. Vanderbilt University worked it through when it disabled Turnitin's AI detector in 2023: against the roughly 75,000 papers its students submitted in 2022, a 1% false positive rate implies around 750 papers wrongly flagged. Vanderbilt also objected to having no insight into how the tool worked. Other institutions kept the feature and set policies around it. Both responses are common, which is why your own institution's policy is the fact that matters to you.
Detectors also do not err evenly. Liang and colleagues, publishing in Patterns, found that GPT detectors are biased against non-native English writers: a majority of TOEFL essays written by non-native speakers were misclassified as AI-generated, while essays by US school students were scored accurately. Their explanation is that these tools lean on statistical predictability, and writing that uses a smaller, plainer vocabulary reads as more predictable. That study tested other detectors rather than Turnitin, so it is not a measurement of Turnitin's behaviour. The mechanism it identifies is not product-specific, and Stanford HAI's write-up is worth reading if you write in your second language.
What a score does and does not prove
A high score proves that a statistical model, reading your finished text, found it resembled AI-generated writing. That is a reason for a conversation. It is not, on its own, evidence of what you did, and Turnitin's own guidance is that the final judgement rests with the instructor and the institution's policy, not with the number.
A low score proves less than students hope. It says the model did not find the pattern. Text that has been paraphrased, edited or written with a lot of back-and-forth help can score low regardless of how it was produced. Nobody should treat a low score as a clearance certificate, and if your course requires you to disclose AI use, a low number does not remove that obligation.
What is left, in both directions, is process. How the document was written is a real fact with real traces: drafts, version history, notes, the sources you actually opened. That evidence answers the question the score only gestures at.
If your score is high and you wrote it
Four things, in order:
- Do not edit the submitted work. Changing it after the fact reads badly and destroys your record.
- Preserve the process evidence now. Google Docs and Word version history, earlier drafts, outlines, handwritten notes, the PDFs you read. Export or screenshot it while it still exists.
- Read the policy. Your institution's academic integrity page tells you whether a score triggers a conversation, an inquiry or nothing at all. Knowing the procedure before the meeting is worth more than any argument about detector accuracy.
- Be ready to talk about your own essay. Why you framed the argument that way, where a particular source came from, what you cut. This is usually the most convincing thing in the room, because it is the thing a model cannot fake for you.
If it has already moved beyond a score and someone has raised an allegation, we wrote a separate guide on what to do when you're accused of using AI on an essay, including what an integrity meeting is like. And if you are worried before you submit rather than after, will AI detectors flag my writing? covers what tends to raise scores on genuinely human work, while what actually gives a ChatGPT essay away covers what human readers notice, which is a different list.
The thing detectors will never see
Everything above is about style, and style is the weakest available signal. The things that get people into real trouble are checkable facts: a reference that does not exist, a quote nobody wrote, a source cited for a claim it never made. Those survive any amount of rewriting, and a marker who looks up one reference finds them in a minute. Guarding against fabricated citations protects you against a risk that is concrete, provable and entirely avoidable. Detector scores come and go.
A document with a paper trail
CiteOwl shows every change as a diff you accept or reject and keeps the history, so the record of how your draft was built exists by default.
Start writing