Can you really tell whether text is AI?
No one wants to be fooled by AI content. So, we’ve created a wide assortment of tools and assumptions that we can use to determine whether writing is AI. We identify anomalous grammar. We run content through AI detection suites. We’ve even embedded watermarking in text.
But actually identifying AI can be murkier than it seems. The tools that we have are largely flawed—as are the assumptions they’re founded on. Like many situations, few things are as dangerous as overconfidence. And if all you see is bad AI writing, you can easily talk yourself into thinking that you aren’t getting fooled by the good AI.
Trusting your ears and your eyes
Some humans can identify AI writing better than others. But, on average, we still aren’t great at it. It really seems to be dependent on how exposed we’ve been to AI content. When lecturers were tested, they were about 50% correct—a coin flip. But AI annotators were 99% correct.
Other results are likewise all over the place. And that makes sense, because AI is a constantly moving target. Just as AI is changing, people are becoming accustomed to how AI looks and feels. Just as people are becoming accustomed to how AI looks and feels, AI is changing. The hallmarks of AI yesterday are not necessarily the hallmarks of AI today.
Generally, humans are a bit behind on methods of AI detection—everyone started talking about the use of the word “load-bearing” well after that word had been mostly stripped from the model. People still look for telltale signs (such as the use of the infamous em-dash) that aren’t really signs of AI writing.
And while people have gotten quite good at picking up context clues (a Reddit post selling you a waffle iron is probably AI), there’s a problem: there’s usually no way to prove that most of the content we encounter is really AI. People can think they’re quite good at identifying AI just because they’re assuming they’ve always been right. Online, there’s really no way to ever know that your assumptions have been wrong—or how many times you’ve been fooled.
When the truth is never revealed, everyone has a hit rate of 100%.
AI detection suites
Since people aren’t the best at detecting AI, we’ve developed AI detection suites. But it turns out most of those are even worse than people.
There are many, many AI detection suites out there today. They are fairly hit or miss. There is no standard or rigorous methodology used to test for AI writing. Instead, they also look for a set of hallmarks. Mostly, they’re just doing the exact same thing as content generators: they’re using statistical patterns. In some respect, they’re not trying to identify AI writing; they’re trying to identify the telltale messiness of a human.
A human may suddenly pull a random word from thin air. I might use a word incorrectly or just creatively. I may start using longer or shorter sentence rhythms; shorter when excited, longer when contemplative. An AI system, by its nature, is going to be pulling statistically likely words. They’ll be building very structured content.
But you can see the problem with this immediately: What if I’m just a statistically likely person?
In fact, people who write with formality and structure tend to be disproportionately flagged by AI detection suites, because we do write like statistically average people. And just like people, the accuracy of AI detection is all over the place—they often give false positives for non-native speakers and can be easily tricked into giving false negatives.
Today, services like Pangram are all over the place, coaxing writers to check their own work to make sure that it isn’t flagged as AI—even if they haven’t used AI. But that introduces another question. If you change your writing just because an AI system thinks it looks like AI… haven’t you just been told by an AI what to write?
AI watermarking
Finally, we have AI watermarking for text. And the way that this works is actually pretty simple: the AI generator alters its own text so that there is an invisible, statistical watermark on the words generated.
Basically, there’s already a baseline probability for each word in a sentence; the AI generator just weights the probability of words slightly differently based on its own internal cryptographic guidance—and when any word will do, the AI generator will choose a word most heavily weighted. And because those weights are altering regular probability, the same AI generator can see its own invisible hand within the text.
This leaves some obvious vulnerabilities. Rephrasing the content, for example, should just about eliminate the watermarking. Long before LLMs were even popular, there were “content spinning” solutions that just algorithmically changed text. And, the text can only be checked by the system that generated it—the system with access to the cryptographic key. It is not visible to a human reader. In reality, text watermarking is mostly useful for the AI systems themselves—so they don’t ingest and train on their own content.
From a functional perspective…
It’s understandable that we want an easy way to detect AI. Our world and sense of reality is falling apart. And there are some mostly reliable ways to detect AI writing. Some AI texts are far more obvious than others.
But there’s also a very real danger of falling prey to unconscious bias—and these things are not trivial, and they do have consequences. Nowhere has this been more obvious than academia. A single serious accusation of AI use can derail a student’s entire college career. When these accusations are being made by finicky, outdated technological platforms that people think are infallible, the situation becomes even more volatile. When neurodivergent and non-native students are far more likely to be flagged, the situation becomes disastrous.
AI watermarking doesn’t really solve this problem; it’s a separate solution for a separate problem altogether. As AI systems continue to become more sophisticated, they are going to become progressively harder to detect.
What is true is that context still matters more than anything. What is a text trying to get you to do? What are the potential goals of this text? Is it a post that made you feel an emotion—like anger, or sadness, or hate? Is it a post that made you fire up an e-commerce site to make a purchase? Is the post you’re seeing by someone you know is real—or someone you’ll never meet?
Academia may start moving back to blue books. Creatives may need to show their process. Our world will adapt. But while it adapts, we do need to encounter each other with curiosity and grace before assumption.