Back to blog

How Do AI Detectors Work? What They Measure and Where They Fail

Perplexity, burstiness, trained classifiers and watermarks explained in plain language, with sourced evidence on accuracy and false positives, for students and publishers.

October 2, 2026By Stefan Petrov
how AI detectors work

Key takeaways

  • AI detectors do not read meaning. They estimate how predictable a text looks, or they use a model trained on examples of human and AI writing, and they return a probability, not a fact.

  • A different approach, watermarking, hides a statistical signal in text as it is generated, but it only works for text from systems that add the mark.

  • False positives are a documented problem. A well-known Stanford study found that seven detectors misjudged non-native English writing at a high rate.

  • A detector score is a reason to look closer, not proof. Treat it that way whether you are the writer, the teacher, or the publisher.

You paste a paragraph into an AI detector and it says "87% likely AI." What did the tool actually do to arrive at that number, and how much should you trust it?

The honest answer is that detectors make an educated guess from patterns in the text. They can be useful, and they can be wrong. This guide explains how the main methods work, what the evidence says about accuracy, why people get falsely flagged, and what to do about it as a student, a teacher, or a publisher.

What an AI Detector Is Trying to Do

A detector looks at a piece of text and estimates whether a person or a language model wrote it. It cannot see who typed the words. It only sees the words, so it looks for statistical clues in how they are arranged.

Almost all detectors return a probability or a percentage. That is an important detail. A score of 87 percent is not "87 percent of this text is AI." In most tools it is closer to "the system's confidence that this text matches AI writing." The two are easy to mix up, and the difference matters when someone's grade or job is at stake.

Method 1: Perplexity and Burstiness

Early and widely discussed detectors, such as the first versions of GPTZero, were explained in terms of two ideas.

Perplexity measures how surprised a language model is by a text. A model predicts the next word from the words before it. If a passage keeps using the most likely next word, the model is not surprised, so perplexity is low. Language models write by choosing likely words, so their output tends to have low perplexity.

Burstiness is about variation in sentence length and structure. People tend to mix very short sentences with long ones. Model text is often more even.

A simple detector could therefore say: low perplexity and low burstiness suggest AI, and higher values suggest a person. This is easy to explain, but it has an obvious weakness. Plenty of human writing is plain and predictable, such as legal text, simple instructions, and writing by people who are using a language they are still learning. Many vendors have since moved to trained classifiers, so a tool you use today may work differently from this explanation.

Method 2: Trained Classifiers

Most current detectors are classifiers: models trained on large collections of text labeled as human or AI. The classifier learns which patterns tend to appear in each group, and gives a probability for new text.

This approach can be more accurate than a simple rule, but it has built-in limits:

  • It depends on its training data. If the detector was trained on older models, it may be less reliable on text from newer ones.
  • It can be fooled by editing. Heavily rewritten AI text, or AI text mixed with human text, is harder to classify.
  • It can misread human writing that happens to resemble the training examples of AI.
  • It does not explain itself. The score does not tell you which sentences caused it, so it is hard to check.

Many tools also give a per-sentence highlight. Those highlights are more guesses and can be wrong for the same reasons.

Method 3: Watermarking

A different idea is to mark AI text when it is created, instead of trying to spot it later.

In watermarking, the model's choice of words is nudged slightly in a pattern that a detector can later check for. Google DeepMind published a text watermarking method called SynthID-Text in the journal Nature in 2024. Reports at the time described that the watermark survived some changes, such as light editing or cropping, but was less reliable when text was heavily rewritten or translated, and less reliable for answers to factual questions, where there are fewer ways to word something without changing the facts.

Watermarking has two practical limits. First, it only works for text from systems that add a watermark and share a way to check it. Second, it says nothing about text from a model that does not use one. So a watermark can confirm that text came from a particular system, but the absence of a watermark proves nothing.

What the Evidence Says About Accuracy

Accuracy claims vary a lot, and many come from the companies selling the tools. Here are three data points from named sources.

OpenAI's own detector was withdrawn. OpenAI released an AI text classifier in January 2023. In its evaluations on a challenge set of English texts, it correctly flagged 26 percent of AI-written text as "likely AI-written," and wrongly labeled human-written text as AI-written 9 percent of the time. As of July 20, 2023, OpenAI said the tool was no longer available because of its low rate of accuracy. See OpenAI's announcement and update.

A Stanford study found bias against non-native writers. Researchers led by Weixin Liang tested seven GPT detectors on essays by US eighth graders and on TOEFL essays written by non-native English speakers. The detectors wrongly labeled more than half of the TOEFL essays as AI-generated, with an average false positive rate of 61.3 percent, and far fewer of the US student essays. The paper, "GPT detectors are biased against non-native English writers," appeared in Patterns in 2023. The authors' explanation is that these detectors are partly measuring how complex the language is, so simpler writing looks "machine-like."

Vendors publish their own error rates, and warn about low scores. Turnitin states that its AI writing detector's false positive rate is below 1 percent for documents with 20 percent or more AI writing, and that scores under 20 percent are less reliable, which is why it shows an asterisk instead of a percentage in that range. That is the vendor's own statement. Independent testing may find different results, and results depend on the kind of text. See Turnitin's guidance on using the AI Writing Report.

We are not aware of a single accuracy number that applies to all detectors and all kinds of text, and we will not give you one. Be skeptical of any tool that promises 99 percent.

Why Detectors Give False Positives

  • Plain, simple writing looks predictable, and predictable text looks "AI."
  • Non-native English writing often uses simpler vocabulary and more standard phrasing, as the Stanford study found.
  • Formulaic genres, such as lab reports, legal clauses, and product descriptions, are repetitive by design.
  • Heavy editing with grammar tools can smooth text into something more uniform.
  • Short texts give the detector less to go on. Scores on a paragraph are less reliable than on a long essay.
  • Writing that is in the training data. Well-known texts can be scored as AI because the model has seen them many times.

Why Detectors Give False Negatives

  • Light editing of AI text, or mixing AI and human writing.
  • Paraphrasing tools that rewrite the surface of the text.
  • Newer models that the detector has not been trained on.
  • Prompts that ask for a particular style, such as casual or quirky writing.

This is why an arms race exists between detection and tools that rewrite AI text. Neither side has a lasting advantage.

If You Are a Student

  • Know your school's policy on AI before you write. Rules differ by school, course, and assignment, and the policy is what counts, not a tool's score.
  • Keep your drafts. Version history in Google Docs or Word, notes, and outlines show how you worked.
  • Be ready to talk about your work. If you can explain your argument and your sources, that is the strongest evidence.
  • If you are flagged, stay calm. Ask what the score is based on, point to your drafts, and ask for a conversation. Mention that detectors can be wrong, citing sources like those above.
  • Do not rely on a "humanizer" to dodge a policy. If your school bans AI, rewriting the text does not make it allowed.

If You Are a Teacher or Publisher

  • Do not treat a score as proof. Use it, at most, as a reason to talk to the writer.
  • Compare with the writer's other work, drafts, and ability to discuss the content.
  • Be extra careful with non-native speakers and with short texts.
  • Set clear rules in advance about when AI is allowed and how to disclose it.
  • For publishing, focus on quality, accuracy, and originality. Our guide to what Google says about AI content covers the search side: Google's concern is helpfulness, not how a page was produced.

Can Humans Detect AI Writing?

Not reliably. Wikipedia's page on signs of AI writing summarizes research suggesting that people who rarely use language models do little better than chance at telling AI text from human text, while heavy users do better, and it cautions that the picture keeps changing as models and writers change. Our guide to the signs of AI writing lists the patterns people notice, with the same warning that none of them are proof.

Using Detectors Sensibly

If you want to use a detector, the free AI content detector from Writingful gives you a rough read on a draft. Treat the result as a prompt to review the text, not as a verdict. If a draft reads as machine-like, the better fix is editing for substance. Our guide to how to humanize AI content shows how, and the free text humanizer can help with stiff phrasing.

What We Do Not Know

We do not know how well any specific detector performs on your text today. Models change quickly, vendors update their tools, and published tests age fast. The studies above are snapshots from the dates given. Check recent independent evaluations before relying on a tool for a high-stakes decision, and consider that a false accusation harms a real person.

Frequently asked questions

They estimate whether text looks like AI writing, either by measuring how predictable it is (perplexity and burstiness) or by using a classifier trained on examples of human and AI text. They return a probability, not a certainty.
Not reliably. OpenAI withdrew its own detector for low accuracy, a Stanford study found high false positive rates for non-native English writing, and vendors' own numbers are not independently confirmed. Treat scores as a hint, not proof.
Turnitin offers an AI writing detection feature. Its own guidance says scores below 20 percent are less reliable and are shown with an asterisk, and that its false positive rate is below 1 percent for documents with 20 percent or more AI writing. Those are the vendor's statements.
Plain, formulaic, short, or simple writing can look predictable to a detector, and non-native English writing is flagged more often. Keep your drafts and be ready to explain how you wrote it.
Yes. Editing, paraphrasing, mixing human and AI text, and newer models can all reduce detection. That is why a score should never be the only evidence.
A method where a model's word choices are subtly shaped so a matching detector can later recognize the text. It only works for systems that add and share a watermark, and it is less reliable if the text is heavily rewritten.