Clever AI Detector · Benchmark August 2026

Best AI Detectors: Which Ones Actually Catch AI Text

We ran 600 AI-involved essays — generated, machine-rewritten, AI-edited and humanized — plus 150 human essays through eight detectors and recorded every verdict. Detection rates ranged from 18.8% to 96.7%. No detector is safe on its average: the figure that separates these tools is the worst category, where the field runs from 0% to 92.0%.

600 AI texts· 150 human texts· 8 detectors· 6,000 verdicts· 4 levels of AI involvement· Tested August 2026
Try Clever AI Detector free See the full results No signup · 10,000 words · unlimited checks
BEST AI DETECTOR most accurate ai detector how accurate are ai detectors which ai detector is the most accurate DETECTOR AI TEXT CAUGHT · 600 TEXTS CLEVER AI DETECTOR 96.7% COPYLEAKS 95.0% ORIGINALITY 86.8% ZEROGPT 18.8% 96.7% CAUGHT · 20 MISSED 8 DETECTORS · 6,000 VERDICTS CAUGHT MISSED 0 FALSE FLAGS
92.0% the highest weakest-category score on this benchmark, held by Clever AI Detector. The next best worst case is 86.7%.
100% of raw, unedited AI essays were caught — by six of the eight detectors. This is the easy case.
0.7–93.3% detection range on the same 150 humanized AI essays. Only Copyleaks and Clever AI Detector clear 90%; two tools are effectively blind to it.
52.6% average agreement between any two detectors on humanized text. A coin flip does 50%.

Key findings

  1. The worst category is the number that matters, and Clever AI Detector holds the best one: 92.0%. It never dropped below that on any of the four AI categories. Copyleaks is next at 86.7%. Every other tool falls to 51.3% or lower somewhere.
  2. The overall ranking is close at the top. Clever AI Detector caught 96.7% of the 600 AI essays and Copyleaks 95.0%, but that gap of 10 texts is inside the margin of error, so treat the two as level overall. Then come Originality Lite (86.8%) and Winston AI (82.7%). GPTZero caught 43.7% and ZeroGPT 18.8%.
  3. Raw AI text is easy; edited AI text is not. Six of eight detectors caught 100% of unedited AI essays. On essays a model had edited, the same eight ranged from 0% to 96.0%.
  4. Humanized text defeats most detectors. On 150 essays run through a paraphrasing attack, detection ranged from 0.7% (ZeroGPT) to 93.3% (Copyleaks). Only two tools cleared 90%: Copyleaks at 93.3% and Clever AI Detector at 92.0%.
  5. Short text makes every verdict weaker. Under 200 words, QuillBot caught 12.5% and ZeroGPT 0%, against 78.3% and 27.0% at 250–299 words.
  6. The tools disagree with each other constantly. Any two of the eight reached the same verdict on 62.3% of the AI texts, and on humanized text just 52.6% — barely above chance.
  7. Not one tool wrongly flagged a human essay. All eight scored 0 false positives on the 150 human essays in the set. With a control set of that size, the true rate could still be as high as 2.5% for any of them.
  8. No detector output is evidence of authorship. These are probability estimates from statistical models, not a record of who wrote the text.

Tested August 2026 on the GEDE research corpus · Last updated 31 August 2026

Methodology

What Was Tested, and How

The texts come from GEDE — Generative Essay Detection in Education, a public research corpus by Lukas Gehring and Benjamin Paaßen (Bielefeld University), built specifically to model how students actually use language models — not just "human vs. ChatGPT", but the messy middle in between.

We took the first 150 texts from each of its four AI contribution levels, and the first 150 human-written essays as a control set. That gives 750 texts of roughly 288 words each, and 6,000 verdicts. Every text was submitted once to each detector in August 2026 and the provider's own verdict was recorded without modification.

This page measures two things. The first is how much AI-involved text each detector catches, over the 600 AI texts. The second is how often it wrongly flags human writing, over the 150 human essays. On this control set every tool scored zero false positives, but 150 texts is a small control set: the true rate could still be as high as 2.5% for any of them, and these essays are pre-ChatGPT student work on academic topics, not the full range of human writing.

One judgement call that changes the ranking: several tools return a third verdict — "mixed", "likely partially AI", "uncertain". We counted those as not AI, because that is how a teacher facing an ambiguous flag would have to treat it. This is harsh on ZeroGPT, which returned "mixed" for 175 of the 600 texts, every one of them genuinely AI. Counting "mixed" as AI would lift its recall from 18.8% to 48.0%, Pangram's from 67.5% to 73.2%, and GPTZero's from 43.7% to 46.3%. Both views are defensible; we show the strict one and disclose the other.

The five text types, 150 each

Direct AI Task
Essay generated by a model from the assignment description alone. Zero student contribution — the scenario every detector is built for.
AI-improved Improve-Human
The student's own essay, passed to a model to fix grammar and language only. The argument, structure and ideas remain the student's.
AI-rewritten Rewrite-Human
The student's essay rewritten by a model, not just corrected. Meaning survives; much of the original wording does not.
Humanized AI Humanize
Generated text pushed through DIPPER, an 11-billion-parameter paraphraser built specifically to defeat detectors. The adversarial case.
Human Control
Student essays written before ChatGPT existed, with no model involved at any stage. These measure the opposite error: how often a tool accuses a person who did the work.
A labelling choice you should know about: we score AI-improved and AI-rewritten essays as AI. The GEDE authors do the opposite in their own analysis — they place both on the human side, "assuming that most teachers would permit minor improvements through LLMs." So the two middle columns measure something contested: a student who used a grammar-and-style pass on their own argument. Whether that should be caught at all is a policy question, not a technical one — and it is worth deciding before you read the columns below.
The data

Which Dataset Is This? GEDE, and Where to Get It

Every text on this page comes from the GEDE corpus. It is public, free, and you can download it and repeat this benchmark yourself.

Full name
Generative Essay Detection in Education (GEDE)
Authors
Lukas Gehring and Benjamin Paaßen, Faculty of Technology, Bielefeld University
Published
August 2025 — "Assessing LLM Text Detection in Educational Contexts: Does Human Contribution Affect Detection?"
Size
916 human-written student essays and more than 12,500 LLM-generated essays across 826 unique assignment prompts
Source of the human essays
Three pre-ChatGPT corpora: Argument Annotated Essays, PERSUADE 2.0, and the British Academic Written English corpus
Generators used
GPT-4o-mini and Llama-3.3-70B-Instruct; humanized texts produced with DIPPER
Subset we tested
The first 150 texts from each of four AI contribution levels — Task, Improve-Human, Rewrite-Human and Humanize — plus the first 150 human-written essays as a control set, giving 750 essays of about 288 words each

Dataset and paper are the authors' work, not ours. We are not affiliated with them and their own conclusions differ from the framing on this page — they argue current detectors are not reliable enough for educational use.

The leaderboard

Which AI Detector Caught the Most AI Text?

Sorted by detection rate across all 600 AI-involved texts. The second column is the one that separates these tools: a detector's weakest category, because a tool is only as useful as its worst case. The top two are level on the average — 96.7% for Clever AI Detector against 95.0% for Copyleaks, a gap of ten texts that sits inside the margin of error. They are not level at their weakest: 92.0% against 86.7%. The third column is the control set of 150 human essays, where none of the eight raised a flag. Note that a detection rate on its own can always be gamed by flagging everything and be useless.

Eight detectors, 600 AI texts and 150 human texts each Click any column heading to sort
Benchmark results for eight AI detectors across 600 AI-involved texts and 150 human texts: overall detection rate, weakest category, human texts wrongly flagged, texts caught and missed, and 95% confidence interval.
Detector AI text caught Weakest category Human texts flagged Caught Missed 95% CI
Detection rate is measured over 600 AI-involved texts. "Weakest category" is the tool's lowest score across the four AI categories. "Human texts flagged" is measured over a separate control set of 150 human-written essays; every tool scored 0, and with 150 texts the true rate could still reach 2.5% for any of them. Verdicts are each provider's own label, with "mixed" counted as not-AI. Clever AI Detector was scored at its standard 0.5 threshold.
Clever AI Detector was the only tool above 90% everywhere

It is easy to score well on raw AI output — six tools did. Clever AI Detector is the only one that stayed above 90% on humanized text (92.0%), machine-rewritten text (100%) and AI-edited human drafts (94.7%). Copyleaks came closest at 86.7% on the hardest column.

Raw AI output tells you nothing

Six of the eight tools scored 100% on unedited AI essays. Any comparison built on that case alone ranks the field as a nine-way tie. The differences only appear once a human or a second model has touched the text.

The spread is enormous

Between the top and bottom of this table sits a 78-point gap in detection rate on identical inputs. "AI detector" is not a category with a shared standard of performance — it is a label eight very different products wear.

Run your own text through the tool with the best worst case

Clever AI Detector is free, needs no account, and handles up to 10,000 words per check with no limit on how many checks you run — including text that has been paraphrased or humanized.

Check my text
Where they break

Which AI Detectors Catch Edited and Humanized Text?

Short answer: only two of the eight. Clever AI Detector (92.0% on humanized text, 94.7% on AI-edited drafts) and Copyleaks (93.3% and 86.7%) were the only tools that stayed above 85% once a human or a second model had touched the text. This is the part the marketing pages leave out. Every tool here is near-perfect on raw AI output. Sort them by what happens when a human has touched the text — or when a machine has been asked to hide its fingerprints — and the ranking scrambles completely.

Show
Flagged as AI 0% – 100%
Swipe sideways for all four columns →
Percentage of texts correctly flagged as AI by each detector, split by how the AI was involved. Higher is better in every column.
Detector Direct AI
should flag
AI-rewritten
should flag
AI-improved
should flag
Humanized AI
should flag
150 texts per column, all of them AI-involved. Green means the detector caught them; sand means it missed them.
The AI-improved column is the real story

A student writes their own essay, then asks a model to tidy it up. GPTZero flags 1.3% of those. ZeroGPT flags 0%. Pangram flags 18%. Copyleaks reaches 86.7%, Clever AI Detector 94.7%, and Originality Lite 96%. Same essays, same argument, same student — and a seventy-fold difference in whether they get caught, decided entirely by which tool the institution bought.

Humanized text splits the field in half

Detection ranges from 0.7% to 93.3% on the identical 150 texts. Two tools — QuillBot at 22% and ZeroGPT at 0.7% — are effectively blind to it. Only two tools cleared 90%: Copyleaks at 93.3% and Clever AI Detector at 92.0%, and they differ by two texts out of 150, which is not a real gap. Anyone deliberately evading detection will find a tool that does not see them; the tools that catch this case are the exception, not the norm.

Sample size effects

Does Text Length Change an AI Detector's Accuracy?

Short answer: yes, sharply, for most tools. On texts under 200 words, five of the eight detectors lost more than 20 percentage points against their own best length band, and three of those lost more than 50. Detectors work statistically: they need enough text to measure how a passage's word choices deviate from what a model would typically produce. Below roughly 250 words, most of them lose the plot. This chart shows the share of AI-involved texts each tool caught, grouped by length.

AI texts caught, by word count Toggle a detector to show or hide its line
Buckets contain 24, 129, 267, 124 and 56 AI texts respectively. The under-200 bucket is small, so treat its exact values as indicative rather than precise.
Practical consequence: a verdict on a 150-word paragraph is close to worthless from most of these tools — QuillBot fell to 12.5% under 200 words, Winston AI to 41.7% and ZeroGPT to 0%. Clever AI Detector held up best at 91.7% and Copyleaks at 83.3%, but no tool held 100%, and the sample in that bucket is 24 texts, so treat even those figures as indicative rather than settled.
Consensus

How Often Do AI Detectors Agree With Each Other?

Short answer: about as often as a coin flip on hard cases. If AI detection were a solved measurement problem, any two tools looking at the same text would reach the same verdict nearly always. Across all 600 AI texts, any two of these eight reached the same verdict 62.3% of the time. On humanized text specifically, 52.6% — statistically indistinguishable from flipping a coin.

Of the 8 detectors, how many caught each AI text?
Each row is 150 AI texts, distributed by how many of the eight detectors caught them. A unanimous verdict is the exception everywhere except raw AI output.

Consensus voting works better than any single tool

We tested what happens if you combine the seven third-party detectors into a single vote, and how that stacks up against one tool on its own:

RuleAI caught
Any 1 of 7 says AI98.7%
At least 2 of 7 agree96.2%
At least 4 of 7 agree72.5%
Best single tool (Copyleaks)95.0%
Clever AI Detector alone96.7%

Stacking seven tools gets you to 98.7%, two points above what Clever AI Detector reached on its own, at seven times the cost and seven separate subscriptions to manage. We report that because it is what the data says, not because we recommend it: every extra tool you add is another chance for one of them to be wrong about a text, which is why loose "any tool flags it" screening should never be the basis of a decision about a person. On this run the loose rule cost nothing — none of the seven flagged a human essay — but that is one control set of 150 texts, not a guarantee.

What disagreement actually tells you

When eight tools split 4–4 on a text, that is not eight tools being individually unreliable. It is a signal that the text sits in genuinely ambiguous territory — prose that carries statistical traces of both a person and a model, because in the AI-improved and humanized categories that is literally what it is.

The measurement is not broken. The question is. "Was this written by AI?" assumes a binary that most real 2026 writing no longer fits — a student who drafts by hand, gets grammar suggestions from one model and structure feedback from another has produced something a yes/no detector cannot describe.

A more useful question: not "is this AI?" but "how much of this process was the student's own thinking, and can they walk me through it?" That question has an answer. The binary one frequently does not.
Plain-language explanation

Why Every AI Detector Advertises 99% Accuracy

Because on one specific task, it is true. Six of the eight tools we tested flagged 100% of unedited AI essays. If your test set is "human essays vs. text pasted straight out of a chatbot", 99% is an easy number to reach and an honest one to print.

It stops being honest the moment a reader assumes it describes their situation. Three things are usually true of an advertised accuracy figure, and none of them are stated on the page it appears on:

  • It was measured on raw AI output. Not paraphrased, not humanized, not a human draft that a model edited — the three cases that make up most real submissions.
  • It was measured on a balanced test set. Half AI, half human. Real inbound work is not half AI, and accuracy on a balanced set tells you very little about a skewed one.
  • It combines two different error types into one number. A tool that catches 99% of AI and wrongly accuses 3% of humans, and a tool that catches 96% and accuses nobody, can both print "99%" depending on which metric they picked.
Interpretation

How to Read an AI Detector Score

Detector outputs look like precise measurements. They are estimates with error bars that nobody prints. Move the slider to see what a given score is worth on the evidence from this benchmark.

Tool by tool

What Is Each AI Detector Actually Good At?

Short, specific verdicts based only on what we measured. The badge is balanced accuracy — AI recall and human specificity averaged. The strengths and weaknesses come from the per-category results above.

Clever AI Detector98.3 balanced

The only tool that stayed above 90% on all four AI categories — including humanized text, where most of the field falls below 65%, and AI-edited human drafts, where four tools scored under 40%. Its weakest category, 92.0%, sits five points above the next best tool's weakest. It is not first in every column: Copyleaks caught more humanized text (93.3% against 92.0%) and Originality Lite more AI-improved drafts (96% against 94.7%), and neither of those gaps is larger than the margin of error. Free, no signup, 10,000 words per check, unlimited checks.

92.0% on humanized94.7% on AI-improved100% on direct AI0 of 150 human essays flagged10,000 words per check
Copyleaks97.5 balanced

The strongest of the third-party tools and the most consistent of them: it never dropped below 86.7% on any AI category, which no other external tool managed, and it held 83.3% even on texts under 200 words. It beat Clever AI Detector on humanized text, 93.3% against 92.0%. Its one clear gap is AI-edited human drafts — 86.7% against 94.7%.

Most consistent93.3% on humanizedNo false positivesPaid, per-page pricing
Originality.ai (Lite model)93.4 balanced

Perfect on machine-written and machine-rewritten text — 100% on both — and the best in the field on AI-improved drafts at 96%. Its weak spot is humanized output, where it caught 51.3%, roughly a coin flip. Note that we tested the Lite model; Originality's heavier models are marketed as stronger on evasion and were not part of this run.

96% on AI-improved51.3% on humanizedWeak under 250 wordsLite model only
Winston AI91.3 balanced

A near-clone of Originality Lite's profile: 100% on generated and rewritten text, and only 44.7% on humanized output. It fell off harder at both length extremes than Originality did, dropping to 41.7% under 200 words and 76.8% over 350.

86% on AI-improved44.7% on humanizedLength-sensitive
Pangram83.8 balanced

An unusual profile: better on humanized text (64%) than on AI-improved drafts (18%), the reverse of Originality and Winston. It also returned "mixed" for 34 texts, all of them genuinely AI; counting those as flags lifts its recall from 67.5% to 73.2%. Reads as tuned to catch fully-generated prose rather than machine-assisted human prose.

100% on direct AI18% on AI-improved34 "mixed" verdicts
QuillBot AI Detector82.1 balanced

Free, no signup, and perfect on raw AI — which is roughly what you would expect from a free utility attached to a paraphrasing product. It caught 22% of humanized text and 12.5% of AI text under 200 words. Fine as a first look; not something to base a decision on.

Free, no signup96.7% on AI-rewritten22% on humanized12.5% under 200 words
GPTZero71.8 balanced

The most recognised name here and the second-weakest result on this corpus. It caught 92.7% of directly generated essays but 1.3% of AI-improved drafts and 7.3% of AI-rewritten ones — meaning almost anything that passed through a human's document first went undetected. It did, notably, catch 73.3% of humanized text, the third-best score in that column.

73.3% on humanized1.3% on AI-improved7.3% on AI-rewritten8.9% over 350 words
ZeroGPT59.4 balanced

Last by a wide margin under our strict rule, at 18.8% recall — but it returned "mixed" for 175 texts, more than any other tool, and all of them were genuinely AI. Counting those as detections lifts it to 48.0%, which would put it just above GPTZero at 46.3% and still far behind the other six. It flagged 0% of AI text under 200 words. Treat its confident-looking percentage bar with particular scepticism.

18.8% recall (strict)48.0% if "mixed" counts0% under 200 words0.7% on humanized
Practical use

How Should You Use an AI Detector Without Doing Harm?

These numbers support a fairly narrow set of uses. Here is what the evidence on this page will and will not carry.

Treat the output as a prompt for a conversation, never a conclusion

A flag means the text has statistical properties common in model output. It does not identify an author. The appropriate next step is asking someone about their process — not opening a case.

Check the length before you check the score

Under 200 words, five of the eight tools lost more than 20 points against their own best length band, and three of those lost more than 50. If the sample is short, no verdict from this list is worth acting on.

Know which category you are trying to catch

If your concern is students generating whole essays, almost any tool here works. If it is students running their own drafts through a model, the tools differ by a factor of seventy — and you need to have checked the AI-improved column before you buy.

Collect evidence that a detector cannot fake

Draft history, version timestamps, an in-class writing sample, a five-minute conversation about the argument. All of these establish things a probability score cannot, and all of them survive scrutiny that a score does not.

Tell people the rules in advance

Most disputes over AI use are disputes over an unstated policy. "You may use a model for grammar but not for structure" is enforceable and fair; "don't use AI" in 2026 is neither.

What This Benchmark Does Not Show

Every limitation we are aware of, listed rather than buried:

  • One text genre only. Argumentative student essays of about 288 words, largely by non-native English writers. Nothing here predicts performance on fiction, journalism, technical documentation, code comments or business writing.
  • One language. English. Detector accuracy is known to vary sharply by language and we did not test that.
  • The control set is small. False positives are measured, on 150 human essays, and every tool scored zero. With 150 texts that result is compatible with a true rate of up to 2.5%, so it is evidence that no tool is grossly over-flagging, not evidence that any tool is safe.
  • The control essays are one narrow kind of human writing. They are pre-ChatGPT student essays on academic prompts. A tool that never flags those may still flag a non-native writer, a formulaic report or a heavily edited draft.
  • Not a random sample. We took the first 150 texts of each type from the corpus rather than sampling randomly, so topic clustering may affect results.
  • The top two are not separated by this test. Clever AI Detector caught 580 of 600 AI texts and Copyleaks 570. On a paired test that difference gives p = 0.15, so it is inside the margin of error. The one difference that does hold up is the weakest category, 92.0% against 86.7%.
  • One snapshot, August 2026. These are live commercial products that update without notice. A result from three months ago may already be wrong; treat the date as part of the finding.
  • One threshold per tool. Several detectors let you move the sensitivity dial. We used each tool's default verdict, and Clever AI Detector's standard 0.5 threshold — the numbers a typical user gets, not the best numbers any tool can produce.
  • "Mixed" verdicts counted as not-AI. A defensible choice that materially penalises ZeroGPT and Pangram. The alternative figures are published in the methodology section.
  • Single run per text. We did not test whether a detector returns the same verdict for the same text twice. Some do not.
  • The AI labels are contestable. The corpus counts a human essay improved by a model as "AI", while the GEDE authors themselves classify it as human. If you side with them, the AI-improved and AI-rewritten columns become false positives rather than catches, and the ranking changes substantially. Decide your own policy before reading the table.

Contradict it. These are live products that change, and a benchmark is only as good as its last run. If you test these tools and get different numbers, tell us — we will re-run and publish the correction on this page.

Questions people actually ask

Best AI Detectors: FAQ

What is the best AI detector in 2026?

There is no single winner on the averages. Clever AI Detector caught 96.7% of the 600 AI texts and Copyleaks 95.0%, a gap of ten texts that is inside the margin of error. Then come Originality Lite (86.8%) and Winston AI (82.7%), both with a significant blind spot for humanized text.

The one separation that holds up is the worst case. Clever AI Detector never fell below 92.0% on any category; Copyleaks never fell below 86.7%; every other tool drops to 51.3% or lower somewhere. None of the eight flagged a human essay in the 150-text control set.

The more useful answer: "best" depends on what you need to catch. If your concern is students editing their own drafts with a model, Originality Lite (96%), Clever AI Detector (94.7%) and Copyleaks (86.7%) are far ahead of GPTZero (1.3%) and ZeroGPT (0%). If it is deliberate evasion, Copyleaks (93.3%) and Clever AI Detector (92.0%) lead while QuillBot (22%) and ZeroGPT (0.7%) do not compete. Check the column that matches your problem, not the headline.

Does this page tell me how often detectors falsely accuse human writers?

Partly. The run includes a control set of 150 human-written student essays, and every one of the eight tools scored zero false positives on it. That is the best result the test can return.

It is not enough to call any tool safe. With 150 texts, a rate of zero is still compatible with a true rate of up to 2.5% — one wrong accusation in forty. The control essays are also one narrow kind of writing: pre-ChatGPT student work on academic prompts. A tool that leaves those alone may still flag a non-native writer or a formulaic report. Do not read a high detection rate on this page as a safety claim about any tool, ours included.

Can AI detectors detect humanized or paraphrased AI text?

Some can; most partly can; a couple cannot at all. On the same 150 humanized essays the detection rate ranged from 0.7% (ZeroGPT) to 93.3% (Copyleaks), with Clever AI Detector at 92%, GPTZero at 73.3%, Pangram at 64%, Originality Lite at 51.3%, Winston at 44.7% and QuillBot at 22%.

This is the widest spread anywhere in the benchmark, and it is the case where a tool's marketing tells you the least about its behaviour.

Why do detectors claim 99% accuracy when independent tests show less?

Because both figures are usually measured on different tasks. On unedited AI text the 99% claim is real — six of eight tools here scored 100% on it.

Independent tests tend to include edited, paraphrased and humanized text, and there the same tools spread out: 0% to 96% on AI-edited drafts, and 0.7% to 93.3% on humanized text. Neither number is a lie; the vendor's number just answers an easier question than the one you have.

Does text length change the result?

Sharply, for most tools. Comparing texts under 200 words with texts of 250–299 words: QuillBot went from 12.5% to 78.3%, ZeroGPT from 0% to 27%, Originality Lite from 45.8% to 92.9%, Winston from 41.7% to 89.5%. Copyleaks was the most stable third-party tool, moving only from 83.3% to 96.6%.

Detectors need a statistical sample. Short text does not provide one, so short-text verdicts should not be acted on.

Is a detector result enough to accuse a student of cheating?

No, and no vendor in this test claims otherwise in their own documentation. A detector estimates the probability that text resembles model output. It does not observe who wrote it, and there is no way for an accused person to prove authorship after the fact — which makes a false positive uniquely difficult to recover from.

Use the score as one input alongside draft history, version timestamps, an in-class sample, and a conversation about the work. Several universities have disabled detector integrations rather than manage the false-accusation risk.

Is it worth running several AI detectors on the same text?

For catching AI, yes. Flagging a text when any one of the seven third-party detectors called it AI caught 98.7% of AI-involved texts, with no false positives on the 150-text human control set — two points better than the best single tool, including ours at 96.7%.

The catch is symmetrical: every extra tool is another chance at a wrong flag. That trade-off is acceptable for screening a pile of submissions and unacceptable as the basis for a decision about one person.

Are free AI detectors as good as paid ones?

Not on detection. The two free-tier third-party tools we ran — QuillBot (64.2%) and ZeroGPT (18.8%) — finished sixth and eighth. Clever AI Detector is also free and finished at the top of the table, so price is not the dividing line; what the tool was built to catch is.

That said, all of them were perfect or near-perfect on raw AI output. If all you need is a quick check on whether something came straight out of a chatbot, a free tool answers that. It is the harder cases you are paying for.

Why does the table show a "weakest category" column?

Because an average hides the failure that will actually affect you. Pangram catches 67.5% overall, which sounds workable — until you see that the figure is an average of 100% on generated text and 18% on AI-edited drafts. If your students edit their own work with a model, that tool will miss four in five of them.

The weakest-category column is the number to compare if you do not know in advance which kind of AI use you are facing. Most people do not.

Which dataset was used, and where can I download it?

GEDE — Generative Essay Detection in Education — a public research corpus by Lukas Gehring and Benjamin Paaßen of Bielefeld University, released in August 2025. It contains 916 human-written student essays and over 12,500 LLM-generated essays across eight levels of student contribution.

The dataset and code are on GitHub and the paper is on arXiv. We tested the first 150 texts from each of four AI contribution levels, plus the first 150 human-written essays. Anyone can download it and repeat the run.

How was this benchmark run, and can I reproduce it?

750 texts from the public GEDE research corpus — 150 each of direct AI, AI-improved, AI-rewritten and humanized AI, plus 150 human-written essays as a control set — submitted once to each of eight detectors in August 2026. That is 6,000 verdicts. Each provider's own label was recorded unchanged; "mixed" verdicts were counted as not-AI, and Clever AI Detector was scored at its standard 0.5 threshold.

The corpus is public, so anyone can repeat the run. If your numbers differ from ours, send them to us and we will publish the correction here.

What makes Clever AI Detector different from the others?

The gap shows up on edited text, not generated text. Almost every detector handles raw chatbot output — six of the eight here scored 100% on it. Clever AI Detector was built for the cases that come after: text that has been paraphrased, humanized, or produced by a model editing a human draft. On those three categories it scored 100%, 92.0% and 94.7%. The rest of the field ranged from 0% to 100% across the same three, and no other tool stayed above 90% on all of them — though on each category taken alone, another tool matched or beat it.

It is free, needs no account, accepts up to 10,000 words per check with unlimited checks, and returns a score with the passages that drove it rather than a single verdict.

Check a text with the detector that topped this benchmark

Clever AI Detector catches paraphrased and humanized text that most tools miss. Free, no account, 10,000 words per check, unlimited checks.

Open AI Detector