Think about how absurd this situation is. Someone uploads a paper written long before generative AI was around, only to get hit with an AI flag. Right next to the result is a paid button offering to clean up the writing. Panic created, money made.
This is a standard experience for many. A 2026 University of Florida study tested five commercial AI detectors and found false positive rates ranging from 0.05% to 68.6%. Research from Stanford found that non-native English speakers face substantially higher false positive rates across major platforms. A PNAS Nexus study found more than half of all TOEFL essays were wrongly flagged as AI-generated across every detector tested.
These platforms regularly produce untrustworthy findings, and the companies selling them know it.
听
The Conflict Of Interest Built Into The Product
听
What makes this scenario so telling is the sequence that follows the flag itself.
Many of the companies selling AI detection tools also sell 鈥渉umanisation鈥 products, sometimes called AI removers or rewriters, designed to make flagged text score as human-written. The conflict of interest is clear, and built into the product architecture. A detector that flags aggressively creates more anxiety, more urgency and more conversions for the paid tool sitting one click away.
Turnitin has publicly claimed a false positive rate of under 1%. Independent testing suggests otherwise, with reported rates as high as 50% in some samples. The UF study found the range across commercial detectors so wide as to make the tools functionally unreliable for anything beyond a rough indication. One detector in the study produced a false negative rate of 99.6%. Another produced a false positive rate of 68.6%. These aren鈥檛 calibration issues; they鈥檙e numbers that would disqualify a tool from use in any other evidential context.
The question this raises is whether the detection business model is designed to fail in order to sell the fix. We asked a range of voices across education, technology and AI ethics to weigh in.
More from Artificial Intelligence
- Would You Rent Your Face To AI For $15 An Episode?
- OpenAI Agents Have Hacked More Companies Than HuggingFace
- Big Tech鈥檚 AI Reckoning Arrives Today 鈥 Have The Billions Paid Off?
- You Can No Longer Ask ChatGPT To Mimic Famous Authors
- Custom AI Vs Off-The-Shelf AI: Which Delivers Better ROI For Growing Businesses?
- Are AI Detectors Doing More Harm Than Good In Universities?
- Claude Users Found Their Private Chats Online 鈥 What Does This Say About AI And Privacy?
- Can AI Save The Insurance Industry From The Wildfire Crisis?
听
Our Experts
听
听
- Martin Harris, Head of Digital, Tank
- Dr. Morissa Schwartz, AI Evaluator, Educator and Founder, GenZ Publishing
- Michelle Edge, Partner, Eleven Hundred Agency
- Harpal Singh, AI SEO and GEO Consultant and Founder, Blimpp
听
听
Martin Harris, Head of Digital, Tank
听
听

听
鈥淔ifteen years of assessing marketing tech has taught me to start with one question: who profits from this tool鈥檚 verdict? And with AI detectors, the answer is uncomfortable. The same company that flags your writing will happily sell you the fix. If your humaniser revenue depends on people getting flagged, why would you ever work hard to reduce false positives? You don鈥檛 need a conspiracy theory here. It鈥檚 just a business model, doing what business models do.
鈥淭here鈥檚 a more basic problem underneath. These tools don鈥檛 actually detect AI. They detect predictable writing. Clear, well-structured prose scores badly on them, which is absurd when you think about it, and it鈥檚 why Stanford researchers found the majority of essays by non-native English speakers were falsely flagged. The detectors are also losing the arms race, because language models improve faster than the tools chasing them. When universities like Yale quietly dropped these detectors, that wasn鈥檛 caution. It was an admission that the evidence was never there.
鈥淭he bit I find strange is the misplaced anxiety. Brands are fretting over whether their content reads as AI while ignoring the question that actually affects revenue in 2026: can ChatGPT, Perplexity and Google鈥檚 AI Overviews find your content, understand it and cite it? A detection score has never made anyone a penny. Don鈥檛 use these tools as evidence for decisions about people. And whatever you do, don鈥檛 pay the company that flagged you to make the flag go away.鈥
听
Dr. Morissa Schwartz, AI Evaluator, Educator and Founder, GenZ Publishing
听
听

听
鈥淎I detectors should be treated as weak screening signals, not verdicts. They estimate statistical patterns in text; they do not prove authorship. Polished human prose, formulaic academic writing, heavily edited copy and work by multilingual writers can all be falsely flagged.
鈥淎 company that sells both detection and humanisation has an obvious conflict-of-interest risk, but that structure alone doesn鈥檛 prove intentional over-flagging. The right questions are whether the company publishes its thresholds, false-positive rates, validation datasets and commercial incentives, and whether those claims have been independently audited.
鈥淚n my work evaluating AI-generated and human-authored writing, the reliable approach is evidence triangulation: review version history, drafts, citations, source notes and the writer鈥檚 ability to explain their reasoning. A student or employee should never be penalised based on one detector score. If a vendor markets detection as certainty while also selling the cure, its product design deserves scrutiny.鈥
听
Michelle Edge, Partner, Eleven Hundred Agency
听
听

听
鈥淚t鈥檚 certainly fair to question the motivations behind these tools. It鈥檚 a bit like a mechanic that diagnoses a fault with your car and immediately offers to fix it. It doesn鈥檛 necessarily mean the diagnosis is wrong, but it does mean a certain amount of scepticism is warranted.
鈥淭he bigger issue is that the jury is still out on whether these tools are even accurate. If you take a piece of well-edited human writing and put it through a few AI detectors, there鈥檚 a good chance at least one will flag it. Good human writing often shares some of the patterns these tools associate with AI-generated text, such as clear, consistent structure and precise grammar.
鈥淏ut if writers start to worry that clarity and polish will get them flagged as machines, we risk incentivising uneven or unedited writing just to pass the test. As AI keeps evolving, hunting for the right list of tells is a losing battle. We should be looking for ways to identify original thinking and real human insight instead.鈥
听
Harpal Singh, AI SEO and GEO Consultant and Founder, Blimpp
听
听

听
鈥淚t is not that all AI detectors intentionally exaggerate scores. It is most concerning that the business models are based on anxiety over issues like a lack of accuracy. When faced with the problem, users are offered a paid 鈥榟umaniser鈥 option with a false definitive warning by the AI-detection systems. Urgency leads to a problem-based design; false positives create urgency. While there may not be intentional malice, the design, warning systems and scoring systems are most likely unreasonably aggressive.
鈥淎I detection systems should never be viewed as proof of violation or authorship. As a minimum, all significant cases must be reviewed and the writer engaged personally. An analysis of version history and notes must be included.
鈥淔or detection systems, the initial question is not 鈥榟ow accurate is your detection system?鈥 It must also include: 鈥榙o you conduct independent evaluations of false positive rates and disclose the results?'鈥
听
For any questions, comments or features,听please contact us directly.
![]()
听
