Table of Contents
- What OpenAI found about its own detector
- The Stanford study: detectors and non-native writers
- How AI detectors make their guess
- The maths problem: false alarms add up
- How students can protect themselves
- What to do if you are wrongly flagged
- The bigger lesson: understanding how AI works
- How we teach it
- Frequently asked questions
AI detectors promise to tell whether a piece of writing came from a person or from a tool like ChatGPT. Many schools and universities have used them, and some students have been accused of cheating because a detector said so. So it is fair to ask: do they actually work?
This guide looks at the published evidence, including figures from OpenAI itself and a Stanford study, explains in plain terms how detectors make their guesses, and shows two small calculations that make the problem obvious. It ends with practical advice for students who want to protect themselves, for parents, and for anyone who has been wrongly flagged.
What OpenAI found about its own detector
In January 2023, OpenAI, the company behind ChatGPT, released its own AI text classifier. In its announcement it was unusually frank: on its test set, the classifier correctly identified 26% of AI-written text as "likely AI-written", while incorrectly labelling human-written text as AI-written 9% of the time. OpenAI also noted that reliability improves as texts get longer, so short answers are even harder to judge.
On 20 July 2023, OpenAI withdrew the tool, stating that it was "no longer available due to its low rate of accuracy". If the company that makes ChatGPT could not build a reliable detector for its own model's writing, that tells you something about how hard the problem is.
The Stanford study: detectors and non-native writers
In 2023, researchers at Stanford University (Weixin Liang, James Zou and colleagues) tested seven widely used detectors on two sets of human writing: 91 essays written for the TOEFL English test by non-native speakers, and 88 essays by US 8th-grade students. The detectors were near-perfect on the US essays. But they misclassified over half of the TOEFL essays as AI-generated, with an average false positive rate of 61.22%. Of the 91 TOEFL essays, 89 were flagged by at least one detector.
Then the researchers ran a revealing experiment. When they used ChatGPT to make the non-native essays sound more like a native speaker, the false positive rate dropped from 61.22% to 11.77%. When they made the US essays use simpler words, the rate jumped from 5.19% to 56.65%. In other words, the detectors were largely reacting to how simple the vocabulary was, not to whether AI wrote it. The researchers cautioned against using these detectors in educational settings.
How AI detectors make their guess
Many detectors lean on a measure called perplexity: roughly, how surprised a language model is by each word in a passage. AI tools tend to choose likely, predictable words, so their text has low perplexity. The problem is that plenty of human writing is predictable too: short sentences, common words, a young writer, or someone writing in their second language. To show this, we built a tiny toy version. It learns how common each word is from the paragraphs of 150 posts on this blog, then scores two passages with the same meaning, both written for this demo:
import json, math, re, collections, pathlib
# 1. Learn how common each word is, from the paragraphs of 150 blog posts
words = []
for f in sorted(pathlib.Path("content/blog/data").glob("*.json"))[:150]:
post = json.loads(f.read_text(encoding="utf-8"))
for s in post.get("sections") or post.get("content", {}).get("sections", []):
if s.get("type") == "paragraph":
words += re.findall(r"[a-z']+", re.sub(r"<[^>]+>", " ", s["text"]).lower())
counts = collections.Counter(words)
total, vocab = len(words), len(counts)
# 2. Perplexity: how "surprised" the model is by a passage (lower = more predictable)
def perplexity(text):
ws = re.findall(r"[a-z']+", text.lower())
log_p = sum(math.log2((counts[w] + 1) / (total + vocab)) for w in ws)
return 2 ** (-log_p / len(ws))
simple = ("My school has a big library. I go there every day after class. "
"I like to read books about space. The books are very good and I learn a lot.")
varied = ("Our school's cavernous library has become my after-class refuge, where I devour "
"dog-eared volumes about nebulae, pulsars and the eerie physics of black holes.")
print(f"trained on {total:,} words")
print(f"simple wording: perplexity {perplexity(simple):,.0f}")
print(f"varied wording: perplexity {perplexity(varied):,.0f}")
trained on 176,424 words
simple wording: perplexity 971
varied wording: perplexity 7,239
The simple passage scored 971, the varied one 7,239. A detector that flags low perplexity would flag the simple passage first, even though the only difference is vocabulary. Real detectors use much larger models than our toy, but the Stanford researchers found the same pattern in theirs: the essays every detector flagged had significantly lower perplexity than the rest.
Why detectors are easy to fool
The same study had ChatGPT write 31 college admission essays, then asked it to rewrite each one using more literary language. Detection rates fell from up to 100% to up to 13%. So detectors tend to catch careless users and plain writers, while careful cheaters slip through.
The maths problem: false alarms add up
Even a detector that sounds accurate can produce many false accusations, because most students are honest. Here is the arithmetic for a class of 30, using OpenAI's published rates and supposing 5 students used AI:
# OpenAI's own published figures for its 2023 classifier
catch_rate = 0.26 # share of AI-written text it flagged
false_alarm = 0.09 # share of human-written text it wrongly flagged
class_size = 30
used_ai = 5 # suppose 5 of 30 students used AI
honest = class_size - used_ai
caught = used_ai * catch_rate
wrongly_flagged = honest * false_alarm
flagged = caught + wrongly_flagged
print(f"AI users caught: {caught:.2f}")
print(f"Honest students flagged: {wrongly_flagged:.2f}")
print(f"Share of flags that are innocent: {wrongly_flagged / flagged:.0%}")
# Even a 1% false positive rate adds up at scale
papers = 75_000
print(f"1% of {papers:,} papers = {papers * 0.01:,.0f} wrongly flagged")
AI users caught: 1.30
Honest students flagged: 2.25
Share of flags that are innocent: 63%
1% of 75,000 papers = 750 wrongly flagged
On average, the detector catches 1.3 of the 5 AI users but flags 2.25 honest students, so about 63% of the students it flags did nothing wrong. Better detectors shrink the problem but do not remove it. When Vanderbilt University disabled Turnitin's AI detection in August 2023, it noted that Turnitin had claimed a 1% false positive rate, and that of the 75,000 papers Vanderbilt submitted in 2022, around 750 could have been wrongly labelled. The last line of the output is that same calculation.
How students can protect themselves
- Write in a tool that keeps version history, such as Google Docs or Word with autosave. A document that grows over several sessions, with edits and changes of mind, is strong evidence of your own work.
- Keep your notes, outlines and rough drafts. Photos of handwritten plans count too.
- Know your own work. If you can explain your argument, why you chose each source, and how your draft changed, that is more convincing than any detector score.
- Follow your school's AI policy, and be open about any use that is allowed. Our guide to using ChatGPT to study without cheating shows how to use AI in ways that build your skills rather than replace them.
- Do not run your own writing through a paraphrasing tool to lower a detector score. It can make genuine work look less like yours.
What to do if you are wrongly flagged
Stay calm, and treat it as a conversation rather than a verdict. Ask exactly what the concern is based on. Share your version history, notes and drafts. Offer to talk through the work, or to write something similar under supervision. It is reasonable to point out, politely, that OpenAI withdrew its own detector for low accuracy, that some universities have disabled AI detection for the same reason, and that research shows detectors unfairly flag plain and non-native writing. If you are a parent, ask the school what its policy says about detector evidence, and whether a detector score alone can lead to a penalty.
For teachers
Evidence of the writing process is more reliable than any detector: drafts, in-class writing, short conversations about the work, and assignments that ask students to connect ideas to their own experience or to class discussions.
The bigger lesson: understanding how AI works
Detectors are a good example of why everyone, not only programmers, benefits from understanding how AI actually works. A tool that outputs a confident percentage can still be guessing. Knowing about training data, probabilities and error rates helps you question an AI result instead of simply trusting it. Our post on why machine learning models make mistakes explores this with small experiments, and our AI tools age guide covers which tools are meant for which ages.
A detector score is a guess about probability, not proof about a person.
How we teach it
Two principles on our how we teach page fit this topic closely. Learning by building: a small model like the toy detector above teaches more about how AI guesses than any definition. And mistakes are data, not failures, which is exactly the right way to think about a detector's errors. Our AI literacy classes for students focus on knowing when the machine is wrong, one to one or in small groups of 5 to 10.
Frequently asked questions
Not reliably enough to prove anything on their own. OpenAI's own classifier caught 26% of AI-written text while wrongly flagging 9% of human writing, and OpenAI withdrew it in July 2023 for low accuracy. Detectors can give a signal, but they should never be the only evidence.
Yes. A 2023 Stanford study found seven widely used detectors wrongly flagged non-native English speakers' essays as AI-generated with an average false positive rate of 61.22%, while being near-perfect on US 8th-grade essays.
Many detectors rely on how predictable the words are. Plain, simple or formulaic writing is predictable, so it can look like AI output even when a person wrote it. This affects younger writers and people writing in a second language most.
Show your process: version history in Google Docs or Word, notes, outlines and drafts. Offer to explain your work or write something similar under supervision. Evidence of how the work developed is far stronger than a detector score.
Turnitin has claimed a low false positive rate, but even small rates add up. Vanderbilt University disabled Turnitin's AI detection in August 2023, noting that a 1% false positive rate could have wrongly labelled around 750 of its 75,000 papers from 2022.
Yes. The Stanford researchers found that one extra prompt asking ChatGPT to use more literary language cut detection rates on its essays from up to 100% to up to 13%. This is one reason detector results should not be treated as proof.
Stay calm, ask what evidence the concern is based on, and help your child gather drafts, notes and version history. Ask whether school policy allows a penalty based on a detector score alone, and request a conversation about the work itself.