Benchmark · March 2026 run

How accurate are AI detectors? We tested seven.

The same 1,000 documents went through every tool: 500 AI-written (GPT-4o, Claude and Gemini, some lightly human-edited) and 500 human-written. We report how many AI-written documents each caught and, separately, how many human-written ones it flagged by mistake.

0.8%

Plagino false-positive rate, the lowest of the seven tools

97.4%

Plagino detection rate, 0.8 pts below Turnitin

1,000 documents

500 AI-written, 500 human-written, March 2026

Results

Document-level results, March 2026. Brackets show the 95% confidence interval.
DetectorAI-written caughtHuman-written wrongly flagged
Turnitin98.2% [96.6–99.1%]1.4% [0.7–2.9%]
Plagino(us)97.4% [95.6–98.5%]0.8% [0.3–2.0%]
Originality.ai95.6% [93.4–97.1%]4.1% [2.7–6.2%]
Copyleaks91.3% [88.5–93.5%]2.2% [1.2–3.9%]
GPTZero88.9% [85.8–91.4%]5.7% [4.0–8.1%]
Academi.cx86.4% [83.1–89.1%]1.9% [1.0–3.5%]
TurnDetect79.5% [75.7–82.8%]6.3% [4.5–8.8%]
Download the results (CSV, CC BY 4.0)

What the numbers say

Turnitin caught the most AI-written documents, 98.2%. Plagino caught 97.4%, so on detection alone Turnitin wins. We publish that because a benchmark that only shows wins is an advert.

On false positives the order changes. Ranked from fewest human-written documents flagged: Plagino (0.8%), Turnitin (1.4%), Academi.cx (1.9%), Copyleaks (2.2%), Originality.ai (4.1%), GPTZero (5.7%), TurnDetect (6.3%). A false positive is the error that gets a student accused of something they did not do, which is why we read this column first. Our guide to AI detector false positives explains why they happen.

Some gaps are smaller than they look. Where two tools' intervals overlap, as Plagino's and Turnitin's false-positive intervals do, this run cannot say for certain which is lower.

Methodology

  • Corpus: 500 AI-written (GPT-4o, Claude and Gemini, some lightly human-edited) and 500 human-written.
  • Run date: March 2026. Every tool saw the same documents.
  • Settings: each tool through its own web app on default settings.
  • Scoring: document level. A document counts as flagged if the tool flagged it as AI-written at all, whatever the percentage.
  • Confidence intervals: 95% Wilson score intervals, computed separately for each rate on 500 AI-written and 500 human-written documents. Wilson intervals stay inside 0–100% when a rate is close to zero, where the simpler normal approximation does not.

Limitations

  • One corpus, one run. Results on your writing, in your subject, may differ.
  • Detectors update their models often. A tool may have improved or regressed since March 2026.
  • Document-level scoring ignores how much of a document a tool flagged and whether it flagged the right sentences.
  • Grammarly, Scribbr, Winston AI and Pangram were not in this run. Our comparison pages for those tools say so and quote the vendors' own claims instead.
  • Plagino ran this benchmark and Plagino is one of the tools tested. That is why the method and the data are published here for anyone to check.

We will re-run the benchmark and update this page, the CSV and every comparison page together when we do.

Tool-by-tool comparisons