AI detectors: are Turnitin, GPTZero and Compilatio reliable?
Explainer11 min read · 27 September 2026
Your essay comes back with a high “AI” score, even though you wrote every word yourself. Or you just had an assistant check your spelling, and now you’re wondering whether that will light up in red. AI detectors arrived in universities in 2023, and their reliability has been debated ever since, including by the companies that sell them. Here’s what the research, the vendors and universities actually say, and what you can do if you’re wrongly accused.
In short: no, an AI detector proves nothing on its own. Vendors report false positive rates of around 1% or less in their own tests, but independent studies have measured far more frequent errors in some situations, particularly for non-native English writers and for text that was only lightly edited with AI. OpenAI withdrew its own detector in 2023 over its low accuracy, and several universities (Vanderbilt, Waterloo, Cape Town, Curtin) have switched off Turnitin’s AI detection. Turnitin, GPTZero and Compilatio all say their score is not proof. If you’re falsely accused, your best defense is the trail of your work: drafts, version history, notes, and your ability to explain how you got there.
How an AI detector works
An AI detector is itself a statistical model. It’s trained on two piles of text, one written by humans and one generated by AI, and learns what tells them apart. When it reads your essay, it doesn’t “know” anything: it calculates a probability.
Early detectors leaned heavily on perplexity, which Stanford researcher James Zou describes as “how surprising the word choice is”: very predictable text looks more like machine output. The tools have moved on since. GPTZero says it has not used perplexity and burstiness since autumn 2023, having switched to a deep-learning architecture that classifies text sentence by sentence. According to Journal du Net, Compilatio, a French vendor, fine-tuned a large language model on more than 7,000 texts produced by ChatGPT, Gemini, Mistral and Claude. Turnitin, for its part, didn’t explain its method in detail, a lack of transparency that Vanderbilt University criticized.
The difference from a plagiarism checker matters a lot. A similarity report shows you the source that was copied, so anyone can check it. An AI score shows no source at all. Debora Weber-Wulff’s team put it plainly in their 2023 study: with text-matching software, “at least it is possible to provide evidence of potential misconduct”, which isn’t the case with AI detectors.
What the vendors themselves say
The figures below are self-reported by each vendor, based on its own testing. They aren’t comparable with one another.
| Tool | What it claims | What it admits |
|---|---|---|
| Turnitin | Below 1% false positives per document, for documents where it detects more than 20% AI writing; around 4% at sentence level (May 2023) | Scores under 20% carry an asterisk because they’re less reliable. “We let probably 15% go by” to keep false positives down, its chief product officer told BestColleges |
| GPTZero | False positives kept at no more than 1%, 96.5% accuracy on mixed documents, 1.1% false positives on TOEFL essays (vendor site) | Use it “as a conversation starter, and not as the final verdict” |
| Compilatio | 94 to 99% accuracy, under 1% false positives, more than 35 languages (vendor site) | “This score is not proof,” but an indicator that must be interpreted; no detector is 100% reliable |
| OpenAI (withdrawn) | 26% of AI text caught, 9% of human text wrongly labeled as AI (January 2023) | Withdrawn on July 20, 2023 “due to its low rate of accuracy” |
Two things to keep in mind. First, the 98.5% accuracy figure Compilatio promotes comes with no external audit, as Journal du Net points out. Second, “1%” sounds small. Vanderbilt ran the numbers: with 75,000 papers submitted in 2022, a 1% rate could have meant around 750 papers wrongly flagged. Turnitin says it itself: behind every false positive there’s a real student, and the company cannot eliminate that risk completely. As for OpenAI, the maker of ChatGPT, it pulled its own detector six months after launch.
What independent research says
The bias against non-native writers (Stanford, 2023)
This is the most widely cited study. Weixin Liang, James Zou and their Stanford colleagues ran 91 TOEFL essays (the Test of English as a Foreign Language), all written by humans, and 88 essays by US eighth graders through 7 detectors. The result: an average false positive rate of 61.3% on the TOEFL essays. All 7 detectors unanimously labeled 19.8% of them as AI-written, and at least one detector flagged 97.8%. The US students’ essays were almost all classified correctly. The likely explanation: simpler vocabulary makes text more predictable. The authors “strongly caution against” using these tools in evaluative settings, especially for non-native speakers.
A test of 14 tools (2023)
Eight researchers from several countries, led by Debora Weber-Wulff, tested 14 tools, including Turnitin, Compilatio and GPTZero, between March and May 2023 on 54 documents. Their verdict: these tools are “neither accurate nor reliable.” All of them scored below 80% accuracy, and only 5 exceeded 70%. Turnitin came out on top, followed by Compilatio. Six of the fourteen tools produced false positives, with the risk rising sharply for machine-translated text. Their conclusion: the systems they tested should not be used in academic settings.
Tools are improving, but not everywhere (2025-2026)
More recent studies add nuance. Brian Jabarian and Alex Imas (University of Chicago) tested four detectors on 1,992 passages: news, blogs, consumer reviews, novels and résumés. Commercial tools performed far better than the open-source RoBERTa model, which misclassified 30 to 78% of human text depending on the scenario. Pangram reached near-zero error rates on medium and long passages. But accuracy dropped on passages under 50 words, and the authors describe a “technical arms race” between detectors and models. Their corpus contained no student essays.
A team from Charles University in Prague repeated the Stanford experiment in Czech in February 2026: across 450 learner essays, they found no systematic bias against non-native speakers. In their view, the bias depends on the language. Good news, though it tells us nothing about other languages.
Finally, a University of Notre Dame preprint from August 2026 highlights the problem most relevant to you. On scientific abstracts written before ChatGPT (2013-2015), Pangram and GPTZero produced zero false positives at the threshold used. But once an abstract was lightly rewritten by AI while keeping the author’s ideas, it was flagged 64 to 80% of the time by Pangram and 38 to 49% by GPTZero. In other words, the detector can’t tell an authorized polish from AI-written text. The authors conclude that detector scores shouldn’t serve as standalone evidence of misconduct.
The universities that switched detection off
Several institutions have drawn their own conclusions:
- Vanderbilt (US), August 2023: disabled Turnitin’s AI detection, saying it did not believe the software was “an effective tool that should be used.”
- Waterloo (Canada), September 2025: internal testing found the tool flagged human-written text as 100% AI-generated.
- University of Cape Town (South Africa), October 1, 2025: dropped Turnitin’s AI score in favor of assessing the process of learning (oral exams, observed work, students disclosing their AI use).
- Curtin (Australia), January 1, 2026: AI detection turned off, text-matching checks kept.
The counterexample is sobering. According to ABC News, Australian Catholic University recorded nearly 6,000 alleged misconduct cases in 2024, around 90% of them AI-related, often relying on the Turnitin report alone. Some students had their results withheld during investigations and had to hand over handwritten notes and search histories. The university dropped the tool in March 2025.
The picture in France
Compilatio is widely used in French universities. On the Lyon Saint-Étienne university site, 18 institutions have access to the tool, presented as able to identify copied or AI-generated passages; the same page reminds instructors that careful reading remains the primary detection method. By contrast, the Université de Lorraine decided not to adopt Magister+, the Compilatio package that includes AI detection, citing recommendations published in June 2025 under the national “L’IA en éducation” framework: frequent false positives, easy circumvention, scores with no justification.
As for the rules, there’s no single national standard: each institution has its own charter (we broke them down in our article on academic integrity). The Université d’Angers charter, for example, states that work may go through a plagiarism detection tool and receive “an in-depth analysis in case of doubt”, with possible penalties up to a disciplinary committee. One key point: failing to disclose your use of AI is itself treated as fraud.
Falsely accused: what to do
First, breathe. A score is not a verdict, and the vendors say so themselves. Then act methodically.
- Ask for the evidence in writing. Which tool, which version, which score, which passages were highlighted. Ask for the full report. You can’t defend yourself against a number you haven’t seen.
- Calmly point to what the tools themselves say. Compilatio writes that its score “is not proof.” Turnitin warns that its reports “should not be used as the sole basis for adverse actions against a student.” GPTZero calls its result “a conversation starter.” Quote them, without aggression.
- Gather the trail of your work. The version history in your word processor (Google Docs, Word on OneDrive) shows your text growing over days, crossings-out included. Add your drafts, outline, reading notes, Zotero library, data files and your exchanges with your supervisor.
- Offer to walk through your process out loud. Why this structure, where that idea came from, why this source over another. People who wrote their text can defend it. That’s exactly the logic of universities that now assess the process rather than just the product.
- If you used AI in an authorized way, say so right away. Spell-checking, comprehension questions, feedback on your outline: show your AI-use statement and, if you have them, your exchanges with the tool. The Notre Dame study shows that a simple AI polish can trigger a flag; that’s exactly why disclosing from the start protects you.
- Flag anything that can skew the score. You’re not a native speaker of the language you wrote in, your text is short or highly formulaic (methods, protocols): the studies above show these cases are more exposed.
- If it escalates, follow your institution’s procedure. Reread the charter and academic regulations, keep a copy of every exchange, and get support: student representatives, student associations, the academic office, or an ombudsperson if your university has one.
One thing never to do: run your text through a rewriting tool to lower the score. Charters explicitly target this kind of maneuver, and you’d turn a false positive into actual misconduct.
Our take
Detectors aren’t useless gadgets: the newest ones make far fewer mistakes on fully human text than in 2023, and an instructor is entitled to ask questions. But they remain probabilistic, opaque and, above all, unable to tell an authorized polish from AI-written text, which is precisely the line that university charters draw. A score should open a conversation, never close it. The lasting answer isn’t hunting for suspicious texts or playing cat and mouse: it’s transparent, disclosed use of AI, where the tool acts as a coach that questions and corrects, not as a ghostwriter.
What this means for you
In practice: work from day one in a tool that keeps version history, keep your dated notes and drafts, check your institution’s rules before opening an assistant, and disclose every use of AI, even proofreading. If you need to cite an AI response, do it properly (see how to cite ChatGPT or Claude in a thesis). The day a score lands on your desk, you won’t have to make anything up: you’ll just have to show your work.
FAQ
Can Compilatio detect ChatGPT?
Compilatio offers an AI detector (part of its Magister+ package) that, according to the vendor, flags text from models such as ChatGPT, Claude and Gemini. It claims 94 to 99% accuracy, but says itself that its score is not proof. Not every institution turns it on: the Université de Lorraine chose not to adopt it.
Is a high AI score enough to penalize me?
The vendors themselves say no: Turnitin, GPTZero and Compilatio all present their score as an indicator, not as proof. Any penalty goes through your institution’s procedure, which has to weigh all the evidence. Always ask to see the report and to explain yourself.
If I use AI to fix my spelling, could I get flagged?
It’s possible: a 2026 study shows that text lightly edited by AI is often flagged. First check that your institution allows this kind of help, then disclose it in your work. Authorized, disclosed proofreading is not misconduct.
I’m not a native speaker. Am I more at risk?
In English, the 2023 Stanford study measured an average 61.3% false positive rate on essays by non-native writers. A 2026 study in Czech didn’t find that bias, which suggests it depends on the language and the tools. If this applies to you, say so and lean on your drafts.
How can I prove I wrote my thesis myself?
Show the path, not just the result: version history, drafts, outline, reading notes, exchanges with your supervisor. Then offer to explain your process out loud. That’s the part detectors can’t fake.
Further reading
- AI and academic integrity: what universities actually say: what charters allow, ban, and how to disclose your use.
- How to cite ChatGPT or Claude in a thesis: APA, MLA, Chicago and ISO 690 formats, plus the AI-use statement.
- The citer-sans-plagier skill: learn to quote, paraphrase honestly and summarize without plagiarizing, with or without AI.
- Studying with AI without cheating: use AI as a coach that quizzes you, not as an answer machine.
Sources
- GPT detectors are biased against non-native English writers (Liang et al., 2023) — Patterns, Cell Press, via PubMed Central · accessed 27 September 2026
- GPT detectors can be biased against non-native English writers — ScienceDaily (Cell Press release) · accessed 27 September 2026
- Testing of detection tools for AI-generated text (Weber-Wulff et al., 2023) — International Journal for Educational Integrity · accessed 27 September 2026
- Artificial Writing and Automated Detection (Jabarian and Imas, 2025) — Becker Friedman Institute, University of Chicago · accessed 27 September 2026
- Do AI Detectors Work Well Enough to Trust? — Chicago Booth Review · accessed 27 September 2026
- Different Time, Different Language: Revisiting the Bias Against Non-Native Speakers in GPT Detectors (Al Ali et al., 2026) — arXiv · accessed 27 September 2026
- Why AI Detection Fails for Academic Integrity (Karr Jr et al., 2026) — arXiv · accessed 27 September 2026
- New AI classifier for indicating AI-written text — OpenAI · accessed 27 September 2026
- AI writing detection update from Turnitin’s Chief Product Officer — Turnitin · accessed 27 September 2026
- Testing Turnitin’s New AI Detector: How Accurate Is It? — BestColleges · accessed 27 September 2026
- GPTZero Technology — GPTZero · accessed 27 September 2026
- How do I interpret burstiness or perplexity? — GPTZero Support Center · accessed 27 September 2026
- Compilatio AI detector — Compilatio · accessed 27 September 2026
- How do AI detectors work? — Compilatio · accessed 27 September 2026
- 98,5 % de précision : Compilatio peut-il vraiment identifier les textes générés par IA ? — Journal du Net · accessed 27 September 2026
- Détection de plagiat (Compilatio) — Le numérique à l’Université de Lorraine · accessed 27 September 2026
- Lutter contre le plagiat et l’utilisation frauduleuse des intelligences artificielles — Université de Lyon · accessed 27 September 2026
- Charte d’utilisation de l’IA générative — Université d’Angers / Polytech Angers · accessed 27 September 2026
- Guidance on AI Detection and Why We’re Disabling Turnitin’s AI Detector — Vanderbilt University · accessed 27 September 2026
- Discontinuing use of AI detection functionality in Turnitin — University of Waterloo · accessed 27 September 2026
- UCT scraps flawed AI detectors — UCT News, University of Cape Town · accessed 27 September 2026
- Update on Turnitin AI-Detection Tool — Curtin University · accessed 27 September 2026
- University caught out using AI to wrongly accuse students of cheating with AI — ABC News (Australia) · accessed 27 September 2026






