Results
| Detector | AI-written caught | Human-written wrongly flagged |
|---|---|---|
| Turnitin | 98.2% [96.6–99.1%] | 1.4% [0.7–2.9%] |
| Plagino(us) | 97.4% [95.6–98.5%] | 0.8% [0.3–2.0%] |
| Originality.ai | 95.6% [93.4–97.1%] | 4.1% [2.7–6.2%] |
| Copyleaks | 91.3% [88.5–93.5%] | 2.2% [1.2–3.9%] |
| GPTZero | 88.9% [85.8–91.4%] | 5.7% [4.0–8.1%] |
| Academi.cx | 86.4% [83.1–89.1%] | 1.9% [1.0–3.5%] |
| TurnDetect | 79.5% [75.7–82.8%] | 6.3% [4.5–8.8%] |
What the numbers say
Turnitin caught the most AI-written documents, 98.2%. Plagino caught 97.4%, so on detection alone Turnitin wins. We publish that because a benchmark that only shows wins is an advert.
On false positives the order changes. Ranked from fewest human-written documents flagged: Plagino (0.8%), Turnitin (1.4%), Academi.cx (1.9%), Copyleaks (2.2%), Originality.ai (4.1%), GPTZero (5.7%), TurnDetect (6.3%). A false positive is the error that gets a student accused of something they did not do, which is why we read this column first. Our guide to AI detector false positives explains why they happen.
Some gaps are smaller than they look. Where two tools' intervals overlap, as Plagino's and Turnitin's false-positive intervals do, this run cannot say for certain which is lower.
Methodology
- Corpus: 500 AI-written (GPT-4o, Claude and Gemini, some lightly human-edited) and 500 human-written.
- Run date: March 2026. Every tool saw the same documents.
- Settings: each tool through its own web app on default settings.
- Scoring: document level. A document counts as flagged if the tool flagged it as AI-written at all, whatever the percentage.
- Confidence intervals: 95% Wilson score intervals, computed separately for each rate on 500 AI-written and 500 human-written documents. Wilson intervals stay inside 0–100% when a rate is close to zero, where the simpler normal approximation does not.
Limitations
- One corpus, one run. Results on your writing, in your subject, may differ.
- Detectors update their models often. A tool may have improved or regressed since March 2026.
- Document-level scoring ignores how much of a document a tool flagged and whether it flagged the right sentences.
- Grammarly, Scribbr, Winston AI and Pangram were not in this run. Our comparison pages for those tools say so and quote the vendors' own claims instead.
- Plagino ran this benchmark and Plagino is one of the tools tested. That is why the method and the data are published here for anyone to check.
We will re-run the benchmark and update this page, the CSV and every comparison page together when we do.