How to spot AI-generated writing
The tells that actually hold up, the ones that get innocent people accused, and why running text through a detector is the worst available method.
Start with what does not work, because it is the thing everyone reaches for first.
Detectors are not the answer
A 2023 Stanford HAI study ran seven popular AI-text detectors against TOEFL essays written by non-native English speakers. The detectors flagged that genuinely human writing as AI-generated 61% of the time on average. One flagged 97% of the essays. On native-speaker essays, the same tools were under 10%.
The tools are not detecting machines. They are detecting predictable prose (low vocabulary variance, regular sentence structure) and then mislabelling it. Which means they systematically punish second-language writers, and anyone whose style is plain and disciplined.
OpenAI withdrew its own classifier in 2023 for low accuracy. If the company that built the model cannot detect its output, a $9/month browser extension cannot either.
Use them and you will accuse innocent people. That is not a hypothetical; it is the documented behaviour.
Structural tells that hold up
The reliable signals are about shape, not word choice.
Uniform paragraph length. Human writing is lumpy. Some paragraphs are one line because the point was short. Generated text tends toward evenly-sized blocks throughout.
Symmetrical lists. Every bullet the same length, every heading the same grammatical form, exactly three examples per point. Real thinking is asymmetric, one item deserves four sentences and another deserves five words.
Comprehensive but uncommitted. It covers every angle of the question and takes a position on none of them. The tell is not that it is wrong; it is that nothing is at stake.
The summary that adds nothing. A closing paragraph restating the piece, then a sentence about how this is an evolving space and it will be interesting to see what happens.
No specifics with edges. No names, no dates, no numbers, no anecdote that could be checked. Generic examples where a person would have used the real one they remembered.
Hedging on things that are not contested. Attribution to “many experts” and “recent studies” without ever naming one.
What is genuinely diagnostic
Verifiable claims that turn out to be false. A fabricated citation, a quotation nobody said, a book that does not exist. This is the strongest signal available, and unlike everything else on this list, checking it produces certainty. Ten of fifteen titles in a syndicated 2025 newspaper reading list were invented, real authors, imaginary books. That was not a style judgement. Someone looked them up.
Confident detail on obscure specifics. The more niche the subject, the less likely a precise-sounding claim about it is real.
The tells that are wrong
Em-dashes. Writers have always used them. Accusing someone over punctuation is how you end up apologising.
“Delve,” “tapestry,” “landscape.” These were markers a couple of years ago. They now appear in human writing partly because of AI, and are actively avoided by people using AI. The signal inverted.
Correct grammar. Some people are just good at writing.
Anything a single sentence made you feel. Style is not evidence.
The honest position
Well-edited AI writing is not reliably detectable, and it is getting less so. The realistic goal is not detection. It is deciding what you require.
Ask instead: is this checkable, is it specific, does it commit to anything, and is anyone accountable for it? Those questions work regardless of how the words got onto the page, which is the point, because the process was never really what mattered.