Pangram Flagged My Own Writing as AI
Here's Why That's a Problem

Full disclosure: I used to work in Big Tech, before I left it all behind to live the writer’s life. (FWIW, Big Tech pays better). I worked with AI. A lot. And I used to be a Pangram champion. I touted it as a good way to tell if a writer here (or anywhere, for that matter) was trying to pull the wool over our eyes and pass off AI-generated work as their own (and charge for it).
I’m doing a complete one-eighty on that opinion, and here’s why.
Substack recently announced that it had partnered with Pangram to make their AI detection tools available on its platform. I was already a Pangram subscriber, so I took the announcement in stride and didn’t pay it much heed.
Then I started reading Notes from other writers who had decided to run some of their own work through the Substack-provided tool, and my alarm bells tinkled. One author reported running a chapter from a book-in-progress through the tool, and receiving validation that it had been 100% human-generated. During that review, they realized there had been a typo in their original manuscript. They fixed it, then ran it through the tool again. The response this time? 100% AI-generated. What. The actual. F**k.
This same writer conducted another test. They changed one word from the same manuscript, and actually created a broken sentence. The verdict: 100% human. Put the word back: 100% AI-generated. Now my alarm bells were clanging. How does this even make sense?
I decided to put my own writing to the test. I ran a short story I’d been working on through the tool. This is my own, original work – a story I’ve been working on for a few weeks now, and I had just started to finalize it. Pangram’s verdict? Thirty-one percent AI-generated. I used the standalone tool that I subscribed to a while ago, so it highlighted the passages for me where it detected AI text. The only “evidence” it could cite for its assessment was the phrase “moral clarity.”
Ironically, my story is about an AI in the future that has been tasked with running infrastructure and operations for New York City. It goes rogue, in a HAL sort of way, and a lone QA engineer is left to try and sort it out. So sure – I used some techno-speak in the piece. It’s relevant. But it’s my writing; it’s not an AI writing.
I was kind of freaking out. I combed through all the passages it had flagged and scrutinized the language I’d used. The phrasing I’d chosen. What about my work had said “AI” to this tool? I made some tweaks to the flagged passages. I changed the phrase “moral clarity” to “moral courage,” which I decided fit the story’s protagonist better anyway. Most of the other edits I made were designed to tighten the narrative, eliminate redundancies – you know, the usual stuff. The stuff we do all the time when we’re editing our work.
I ran it through Pangram again. The AI-generated score nearly doubled. And just like the writer I cite above, the only changes introduced were human-made. This isn’t measuring “AI-ness,” it’s measuring noise.
A tool that cannot produce consistent, stable feedback on trivial edits to human-authored text isn’t reliable for anything, let alone a result that could unravel an author’s reputation. A false positive like the ones I describe here could cost a writer a publication, a grade, a contract, or their credibility with an editor or reader base.
Meanwhile, the vendor of the tool faces no consequences for getting it wrong.
Yet institutions and businesses – like Substack – are already using tools like Pangram to make consequential decisions. This is genuinely dangerous, not just annoying.
Pangram’s own marketing claims a very low false positive rate: an astonishing 1 in 10,000.1 But independent reviews complicate that claim considerably. A more detailed self-reported figure puts it at 0.19% false positive and 1.4% false negative on standard datasets. And critically – this is important – that benchmark tests pure human text and pure AI text under controlled conditions. It doesn’t describe real-world use cases, where writers might run their text through multiple rounds of reviews and human editing before finally publishing (which is exactly what happened in the two cases I cite here). Additionally, one 2025 study found standard detectors frequently misclassify AI-polished text as fully human between 10% and 75% of the time, and these aren’t edge cases.2
Even the most favorable independent research on Pangram comes with caveats: researchers at Chicago Booth found its false-positive rate could be pushed close to zero – but only under specific, controlled circumstances. And performance still varied by document length, the underlying LLM used, and how conservatively the tool’s threshold was calibrated.3
Just by pointing one tool at a writer and accepting its outputs as the last word, feels reckless in the extreme. Even the limited experiences I describe here – which appear to be backed up by early research – are delivering false positives. The gap between how confidently these tools present a percentage, and how little that percentage may actually represent reality, is a serious flaw in the system.
Basically, we are asking these tools to perform beyond their capabilities, and risking writers’ reputations in the process.
The practical implications for us: don’t let the number anchor your sense of the work, and don’t keep re-running it hoping for a reassuring score. I, at least, have demonstrated to myself that the score isn’t a stable signal. Your version history and any editorial back-and-forth will be better evidence of authorship than anything Pangram outputs.
And hopefully, this too shall pass. To all the writers I ever casually and carelessly ran through Pangram, I apologize. Rest assured, I won’t do it again.
OR …
“Industry-leading 1 in 10,000” false positive claim — from Pangram's own blog post, All About False Positives in AI Detectors.
“0.19% false positive and 1.4% false negative on standard datasets” + the “10% to 75%” misclassification figure — from a Substack piece, Pangram and the All-Clear by Michael G Wagner, which cites the 2025 study, “Almost AI, Almost Human.”
“Do AI Detectors Work Well Enough to Trust?” — Chicago Booth Review
https://www.chicagobooth.edu/review/do-ai-detectors-work-well-enough-trust
Further reading:
Saha, Shoumik, and Soheil Feizi. “Almost AI, Almost Human: The Challenge of Detecting AI-Polished Writing.” Findings of the Association for Computational Linguistics: ACL 2025, July 2025, pp. 25414–25431.
ACL Anthology (official publication): https://aclanthology.org/2025.findings-acl.1303/
arXiv preprint (freely accessible PDF): https://arxiv.org/abs/2502.15666




I'm still looking for the world where Pangram makes any sense. To me, this is the most useless, deceptive, dishonest and hostile software in all of AI.
It's going to cause people to quit writing entirely. I refuse to use, as there is zero reason to. I know I write everything the old school way. Anyone who doubts that is not going to believe me, so to Hell with all of it.