Can AI detectors reliably distinguish human-written content from AI-generated text? What happens when several detectors analyze the same document but return completely different results?
To investigate these questions, we tested five AI detection tools using six writing samples with known origins. Our experiment included academic writing, blog and marketing content, and writing produced by a non-native English speaker.
Each sample was submitted to CopyChecker, Copyleaks, GPTZero, Originality.ai, and QuillBot. We recorded each tool's AI percentage and evaluated the results using a standardized classification threshold.
Across 30 individual detector evaluations, 22 classifications were correct, producing a combined descriptive accuracy of 73.3%.
However, overall accuracy revealed only part of the picture.
Some detectors correctly classified both AI-generated samples but incorrectly flagged human writing. Another classified three of the four human samples correctly while failing to identify either AI-generated sample.
Our non-native English samples also produced notable disagreements. The same human-written document received a 0% AI score from one detector and a 100% AI score from another.
These observations show why accuracy percentages must be examined alongside false positives, false negatives, sample characteristics, and individual detector disagreements.
Study scope: This is a small exploratory comparison involving six documents, not a representative benchmark of the AI detection industry. All reported percentages describe performance on our specific samples under a standardized 50% classification rule.
Key Findings
Our AI detector accuracy study examined six writing samples across three content categories. Four samples were human-written, while two were generated using ChatGPT.
We tested every document with the same five detectors, producing 30 individual evaluations. The results revealed differences in overall accuracy and, more importantly, in the types of classification errors each detector made.

Source: CopyChecker's September 28–30, 2026 experiment. We calculated classifications using a standardized threshold of 50% or higher for AI-generated content.
Three detectors, CopyChecker, Copyleaks, and GPTZero, correctly classified five of the six samples under our testing rule. Originality.ai correctly classified four samples, while QuillBot correctly classified three.
However, their error patterns differed substantially. Originality.ai identified both AI-generated samples but classified two human-written samples as AI. QuillBot correctly identified three human samples but classified both AI-generated samples as human.
These results illustrate why two detectors with similar overall accuracy figures may behave differently when examining human and AI-generated content.
Five Important Findings
Several observations emerged from the experiment. All are worth noting before moving on.
Overall Accuracy Varied Across Tools
Recorded accuracy ranged from 50% to 83.3%. Because the experiment contained only six samples, a single additional error would change a detector's overall accuracy by approximately 16.7 percentage points. So, overall accuracy varied across the tested tools.
False Positives Present for All Detectors Tested
Every tool incorrectly classified at least one of the four human-written samples under our standardized threshold. This means all five detectors produced false positives.
Inconsistent False Negative Behavior
Third, false-negative behavior differed. Four detectors identified both AI-generated samples, while QuillBot classified both as human.
Substantial Disagreement for Non-Native Samples
Fourth, non-native English writing produced substantial disagreement. Across ten evaluations involving two genuine human-written samples from the same non-native English writer, four classifications incorrectly identified the content as AI-generated.
Detector Scores Were Not Interchangeable
Finally, individual detector scores were not interchangeable. A percentage displayed by one service does not necessarily have the same statistical meaning as a percentage displayed by another.
Our findings therefore reflect a common classification rule imposed for this experiment, not necessarily each provider's intended interpretation of its scores.
These findings should be understood as observations from a limited test, not estimates of general detector performance.
Why We Ran This AI Detector Accuracy Study
AI detectors are increasingly relevant to education, publishing, content production, and other activities where people need to assess whether writing may have been generated by artificial intelligence.
However, an incorrect detection result can have very different consequences depending on how it is used.
For a content editor, a suspicious result may lead to additional editorial checks. For a student whose original work is classified as AI-generated, an inaccurate result could raise questions about academic integrity.
The challenge is knowing how much confidence to place in a detector's output.
Different AI detectors use different models, scoring systems, thresholds, and methods of presenting their conclusions. One tool may display a numerical AI score, while another may assess individual passages or distinguish between human, AI-generated, and mixed writing.
These differences make direct comparisons difficult.
Published accuracy claims also depend heavily on testing methodology. A detector evaluated against thousands of AI-generated documents may produce different results when tested on academic writing, short marketing articles, or text written by people with different linguistic backgrounds.
Independent academic research has raised similar concerns. A 2023 study published in the International Journal for Educational Integrity evaluated multiple AI detection systems and found substantial reliability problems, including differences in false-positive and false-negative behavior and reduced effectiveness when AI-generated content was modified. DOI
Non-native English writing presents another important research question. In a separate 2023 study, researchers from Stanford University examined seven AI detectors using 91 TOEFL essays written by non-native English speakers and 88 essays written by US eighth-grade students.
Across the tested detectors, an average of approximately 61.3% of the TOEFL essays were incorrectly classified as AI-generated. The researchers reported substantially different results for the US student essays. Their findings raised concerns about how certain linguistic characteristics may affect AI detection. DOI
This research provided an important reason to examine genuine non-native English writing in our experiment. Our objective was not to reproduce those studies or establish that one detector is universally reliable or unreliable.
Instead, we wanted to investigate how five accessible AI detection services responded to the same small collection of known-origin texts.
We focused on three questions: whether the detectors agreed with the actual origin of each document, what kinds of mistakes they made, and how they responded to writing produced by a non-native English speaker.
Our Research Methodology
We conducted our experiment between September 28 and September 30, 2026, using the web interfaces of five AI detection services.
We tested six English-language writing samples across three content categories. The dataset contained four human-written documents and two ChatGPT-generated documents. We submitted each document to all five detectors.
We recorded the percentage each service displayed and converted those scores into binary Human or AI classifications using a standardized 50% threshold.
We then compared those classifications against the known origin of each sample.

The three writing categories were Academic, Blog/Marketing, and Non-Native English.
The study was not balanced equally between human and AI writing. Academic and Blog/Marketing each contained one human and one AI sample. The Non-Native English category contained two human-written samples and no AI-generated samples.
Consequently, we analyzed the non-native category separately for human-writing classification and false positives.
Writing Categories Tested
We selected three writing categories to examine how the detectors responded to different forms of English-language content.
The final sample distribution was determined partly by the availability of source material and the practical limits of conducting the experiment within a short testing period.
Academic Writing
The academic category contained one archived human-written sample and one ChatGPT-generated sample. The human sample was archived, while the AI sample was generated using ChatGPT.
The academic material used in this project included science-oriented explanatory writing. This distinction matters because research-based science journalism and conventional university essays can have different linguistic characteristics.
Accordingly, our results should not be interpreted as an assessment of AI detector performance across every form of academic writing. We tested both academic samples using the same five detectors.
Blog and Marketing Writing
The Blog/Marketing category contained one human-written sample sourced from material published online before ChatGPT's public release and one ChatGPT-generated sample.
The human document was selected because its publication history provided a basis for establishing its human origin. The AI sample was generated using ChatGPT and recorded separately. These two documents formed the blog and marketing category for the experiment.
However, with only one document from each origin, differences in subject matter, style, and other writing characteristics could influence the results. The dataset does not support conclusions about blog and marketing writing as an entire category.
Non-Native English Writing
Our third category focused on human-written content produced by a non-native English speaker. We included two original samples from the same writer on our team.
We recorded both documents as Human because we knew their authorship independently of the AI detection results.
Unlike the previous categories, we did not include an AI-generated comparison sample in this group.
This category therefore examined false positives rather than balanced Human-versus-AI classification accuracy. We did not treat AI-generated text written in simplified English as genuine non-native human writing.
Although prompting an AI model to use simple vocabulary and sentence structures may provide a useful supplementary experiment, it cannot reproduce the full linguistic characteristics of an actual person writing in an additional language.
How Human Writing Was Selected
We selected human samples from archived content, previously published online material, and original writing produced by a known team member.
We established their origin through available source information rather than by relying on another AI detector.
This distinction is essential. If a document is labeled Human simply because another detector believes it is human-written, the resulting experiment cannot provide an independent assessment of detector accuracy.
For our analysis, the recorded Actual Origin served as the ground truth. The study's published dataset identifies each sample's source category. Original source URLs and the complete source texts are not included in this article.
This limits how far an external reader can independently verify sample provenance, which we acknowledge in the study's limitations.
How AI Writing Was Generated
We recorded both AI-generated samples as ChatGPT outputs. One belonged to the academic category, while the other belonged to the Blog/Marketing category.
Their known AI origin allowed us to measure whether the five detectors classified them as AI-generated under our standardized threshold. The experiment did not include outputs from Claude, Gemini, or other generative language models.
The recorded dataset also does not preserve complete generation prompts, exact ChatGPT model identifiers, or a complete account of any reference material supplied during generation.
We therefore do not claim that these samples represent all ChatGPT models, all AI-generated writing, or generation without reference material.
The resulting false-negative measurements apply specifically to the two AI documents included in this experiment.
Sample Length and Matching
The study was designed around comparable writing forms within the selected categories.
However, the final spreadsheet does not record individual word counts or enough prompt-level information to demonstrate exact matching by length, subject, or writing complexity.
For that reason, we have not treated the samples as tightly controlled experimental pairs. Document length and writing characteristics may still influence the results.
Future testing could improve this aspect of the methodology by recording each document's word count and matching human and AI writing using the same assignment or content brief.
How the Five AI Detectors Were Tested
Every sample was submitted to CopyChecker, Copyleaks, GPTZero, Originality.ai, and QuillBot through their web interfaces. For each submission, we recorded the tool's numerical AI score.
We entered these scores into a spreadsheet alongside the sample ID, writing category, source information, and actual origin. The spreadsheet then calculated a binary classification and compared it against the actual origin. We used the same standardized classification rule across all five services.
We classified scores below 50% as Human and 50% or higher as AI.
This provides a consistent calculation method, but it introduces an important limitation.
AI detection services do not necessarily use percentages in the same way. Some systems may display probability-like confidence scores, while others estimate the proportion of AI-generated content or apply additional classification rules.
Our standardized results must therefore be distinguished from evaluations using each detector's native classification method. We retained the original numerical scores so readers could examine the results independently of our binary classification rule.
How We Defined a Correct Result
Each sample had a known origin recorded as either Human or AI. We compared detector classifications against that origin using four possible outcomes.

These definitions formed the basis of our calculations. For example, a human-written document receiving a 78% AI score would be classified as AI under our threshold. Because its actual origin was Human, that result would count as a false positive.
Similarly, an AI-generated sample receiving an 8% AI score would be classified as Human, producing a false negative.
How Accuracy Was Calculated
We evaluated three principal performance measures.
- Overall accuracy measures the percentage of all tested documents classified correctly. Accuracy = Correct classifications ÷ Total classifications × 100
- False-positive rate measures the percentage of known human-written documents incorrectly classified as AI. False-positive rate = Human samples classified as AI ÷ Total human samples × 100
- False-negative rate measures the percentage of known AI-generated documents incorrectly classified as Human. False-negative rate = AI samples classified as Human ÷ Total AI samples × 100
Each detector was evaluated on six samples: four human documents and two AI documents. The denominators are important because the dataset is small.
For example, one false positive produces a 25% false-positive rate when only four human samples are tested. Similarly, one false negative produces a 50% false-negative rate when only two AI-generated samples are tested.
These percentages accurately describe the recorded results but should not be interpreted as precise estimates of real-world performance.
AI Detectors Tested
We selected five AI detection services and evaluated them using their publicly accessible web interfaces during September 28–30, 2026.

We did not record the exact underlying detector model versions. Results should therefore be interpreted as a snapshot of the web services accessed during the stated testing period.
We also did not verify that the five services offered equivalent sensitivity settings or classification modes.
Our standardized scoring method provides a common numerical comparison, but it does not eliminate differences in the tools' underlying detection systems.
Research Disclosure
CopyChecker conducted and published this research and is one of the five products included in the experiment.
The same six samples and standardized scoring rule were used for all five detectors. CopyChecker's incorrect classifications have been retained and reported alongside errors produced by the other services.
Readers should consider the publisher's commercial interest when interpreting this comparison. The numerical results and recorded sample-level scores are included to make the calculations transparent.
Overall AI Detector Accuracy Results
Across the six documents, the five detectors produced 30 individual classifications. Of those classifications, 22 matched the known origin of the submitted material, while eight were incorrect.
That produces a combined descriptive accuracy of 73.3% across the 30 evaluations.
However, this combined figure should not be interpreted as an industry-wide accuracy rate. The same six documents were tested repeatedly, meaning the 30 evaluations were not independent observations. Individual detector accuracy ranged from 50% to 83.3%.
CopyChecker, Copyleaks, and GPTZero each correctly classified five of six documents. Originality.ai correctly classified four, while QuillBot correctly classified three.
The differences become clearer when we examine the underlying mistakes. CopyChecker, Copyleaks, and GPTZero each made one error involving human-written non-native English content.
Originality.ai incorrectly classified two human documents. QuillBot incorrectly classified one human document and both AI-generated documents.
These patterns show why reporting overall accuracy without the underlying error types provides an incomplete picture.
How Often Did the Detectors Agree?
Only one document, SA01, all five detectors classified correctly. This was the archived human sample in the academic category.
Every detector returned an AI score below 50%, producing five correct Human classifications. The other five documents produced at least one incorrect classification.
Some disagreements were particularly substantial. For example, the first non-native English sample received scores ranging from 0% to 100%. Three detectors classified it as AI, while two classified it as Human.
The second non-native sample also produced disagreement, although four detectors classified it correctly.
The AI-generated samples demonstrated another pattern. Four detectors classified both correctly, while QuillBot classified both as Human. No sample was classified incorrectly by every detector.
These observations illustrate how strongly the interpretation of an individual document can depend on the detection service used.
False Positive Analysis
False positives were a central focus of this experiment because they involve human-written work being incorrectly classified as AI-generated. Our dataset contained four verified human-written samples.
Every detector incorrectly classified at least one of those documents under the standardized 50% threshold.
Across all five services, six of the 20 evaluations involving human-written documents produced false positives.
This represents a combined observed false-positive proportion of 30% within our experiment. Because the evaluations involve repeated testing of the same four documents, that figure is not a population-level estimate.
The results also show that false-positive behavior was not consistent across individual samples. All five detectors correctly identified the archived academic human sample.
The human Blog/Marketing sample was incorrectly classified by Originality.ai and QuillBot. The two non-native samples produced different patterns of disagreement.
What Happened With Non-Native English Writing?
Our non-native English analysis included two documents written by the same human author. We independently confirmed both documents were human-written before testing.
The five detectors produced the following results.
Four incorrect classifications occurred across ten evaluations of these two documents.
CopyChecker, Copyleaks, GPTZero, and Originality.ai each produced one false positive. QuillBot classified both samples as Human under our threshold.
However, the individual scores tell a more interesting story. For SN01, GPTZero assigned a 0% AI score, while Originality.ai assigned 100%. CopyChecker and Copyleaks also classified this sample as AI.
For SN02, Copyleaks returned 0%, while GPTZero returned 54%. The same writer therefore received substantially different assessments across the two documents and five services.
Across the ten evaluations of our two other human-written samples, there were two false positives. The observed false-positive proportions were consequently 40% for the non-native samples and 20% for the other human samples.
These figures do not establish that non-native English writing generally produces higher false-positive rates.
Both non-native samples came from one person. The other human samples also differed in genre and subject matter, and their writers' language backgrounds were not established as a controlled comparison group.
The observed differences could reflect multiple document characteristics rather than language background alone.
Nevertheless, the results provide concrete examples of how differently AI detectors can assess genuine writing produced by someone who uses English as an additional language. Readers interested in the broader issue can consult our related article on why human writing can be detected as AI.
Accuracy by Writing Type
Our experiment included two documents in each of three content categories. We calculated category-level results by comparing detector classifications with each document's known origin.
The Academic and Blog/Marketing categories each contained one human and one AI sample.
In contrast, both Non-Native English documents were human-written. Its category-level percentages therefore measure only correct identification of human writing.
These results should not be compared as though all three categories had identical sample compositions.
Academic Writing
The academic category contained two documents. CopyChecker, Copyleaks, GPTZero, and Originality.ai classified both correctly.
QuillBot correctly identified the human sample but classified the AI-generated sample as Human.
However, the category contained only one example of each origin. These findings cannot establish general performance on academic essays, research papers, dissertations, or other scholarly material.
Blog and Marketing Writing
The blog and marketing category also contained one human and one AI document. CopyChecker, Copyleaks, and GPTZero correctly classified both samples.
Originality.ai identified the AI-generated document correctly but incorrectly classified the human document as AI. QuillBot misclassified both samples.
The human marketing sample was particularly interesting because it received AI scores of 0% from Copyleaks and GPTZero, compared with 97% from Originality.ai and 66% from QuillBot.
This demonstrates substantial disagreement over a document whose actual origin was known.
Non-Native English Writing
The non-native category contained only human-written documents, allowing us to examine false-positive behavior.
CopyChecker, Copyleaks, GPTZero, and Originality.ai each correctly classified one of the two samples. QuillBot correctly classified both.
The individual errors were not identical across tools, showing that detectors did not consistently agree about which document appeared AI-generated.
Because both texts came from one writer, these observations cannot establish how the tools would perform across the wider population of non-native English speakers.
Detector-by-Detector Results
The following breakdown examines each detector using the same three primary measurements.
These findings describe performance under our standardized 50% classification rule, not necessarily each service's official interpretation of its output.
CopyChecker
CopyChecker correctly classified five of the six documents. Its only incorrect result involved SN01, the first non-native English sample.
That document was human-written but received a 78% AI score, exceeding the study's classification threshold. CopyChecker correctly classified the remaining three human documents and both AI-generated samples.
Its recorded results were 83.3% overall accuracy, a 25% false-positive rate, and a 0% false-negative rate.
The non-native false positive is important because CopyChecker conducted the study and is one of the evaluated products. The incorrect result remains part of the published analysis.
Copyleaks
Copyleaks also correctly classified five of six documents. Like CopyChecker, its incorrect classification involved SN01. Copyleaks assigned that human-written document an 80% AI score.
It correctly classified the remaining human documents and both AI-generated samples. Its overall accuracy was 83.3%, with a 25% false-positive rate and a 0% false-negative rate.
The result shows that correctly identifying the two AI documents did not prevent the detector from incorrectly flagging genuine human writing.
GPTZero
GPTZero correctly classified five of the six samples. Its incorrect result involved SN02, the second non-native English sample.
The human-written document received a 54% AI score and was therefore classified as AI under our standardized threshold. GPTZero correctly identified the other three human documents and both AI-generated samples.
Its recorded overall accuracy was 83.3%, with a 25% false-positive rate and a 0% false-negative rate. The contrast between its two non-native results was notable. It assigned SN01 a 0% AI score but assigned SN02 a score above our classification threshold.
Originality.ai
Originality.ai correctly classified four of the six samples. It correctly identified both AI-generated documents and two human documents. However, it incorrectly classified SM01, the human Blog/Marketing sample, and SN01, the first non-native English sample.
Those documents received AI scores of 97% and 100%, respectively. Its recorded overall accuracy was 66.7%, with a 50% false-positive rate and a 0% false-negative rate.
The results show that a detector can identify AI-generated examples in a small dataset while still misclassifying human-written content.
The provider's own scoring and allowance settings may differ from the standardized threshold imposed in our experiment.
QuillBot
QuillBot correctly classified three of the six documents. It correctly identified the archived human academic sample and both non-native human samples.
However, it classified the human Blog/Marketing sample as AI. More notably, it classified both ChatGPT-generated samples as Human.
The AI academic sample received an 8% AI score, while the AI marketing sample received 0%. Its recorded overall accuracy was 50%, with a 25% false-positive rate and a 100% false-negative rate.
That 100% false-negative figure represents two incorrect classifications out of two AI-generated samples.
It should not be interpreted as evidence that QuillBot fails to detect AI-generated content generally.
Where AI Detectors Made Mistakes
Individual sample results provide more information than overall percentages alone. Three examples illustrate the different types of classification disagreement observed during testing.
Example 1: Human Marketing Content Classified as AI
Sample SM01 was a human-written Blog/Marketing document sourced from material published online before ChatGPT's public release. The five detectors returned very different scores.
| Detector | AI Score | Classification |
|---|---|---|
| CopyChecker | 30% | Human |
| Copyleaks | 0% | Human |
| GPTZero | 0% | Human |
| Originality.ai | 97% | AI |
| QuillBot | 66% | AI |
Three detectors correctly classified the document as Human. Originality.ai and QuillBot incorrectly classified it as AI under the standardized threshold.
The difference between the lowest and highest reported AI scores was 97 percentage points.
Because the tools do not necessarily use equivalent scoring systems, this spread should not be interpreted as a direct comparison of calibrated probabilities.
Nevertheless, the conflicting binary classifications demonstrate how different outcomes can emerge from the same text.
Example 2: Human Non-Native Writing Produced Major Disagreement
SN01 was written by a non-native English-speaking member of our team. CopyChecker returned 78%, Copyleaks returned 80%, and Originality.ai returned 100%.
All three scores exceeded the classification threshold. GPTZero returned 0%, while QuillBot returned 49%. The resulting classifications were split between three AI decisions and two Human decisions.
Because the document was known to be human-written, three of those classifications were false positives. This example illustrates the importance of evaluating a document's actual origin independently of detector output.
Example 3: AI-Generated Writing Classified as Human
The two ChatGPT-generated samples also revealed disagreements. For the academic AI sample, four detectors returned scores of at least 68%, while QuillBot returned 8%.
For the marketing AI sample, CopyChecker returned 54%, while Copyleaks, GPTZero, and Originality.ai each returned 100%. QuillBot returned 0%. Under our threshold, QuillBot classified both AI documents as Human.
This example demonstrates why false negatives need to be reported alongside false positives. A detector that frequently classifies documents as Human may avoid certain false accusations while also missing AI-generated writing.
How Our Results Compare With Published AI Detector Accuracy Claims
Our results should be considered alongside existing research, but they cannot be treated as directly equivalent to larger published benchmarks.
Differences in sample size, document origin, detector version, scoring methodology, and AI-generation models can substantially affect reported performance.
Published Vendor Research
GPTZero published a benchmarking methodology in February 2026 that evaluated its detector across multiple writing domains and AI models. Its reported combined benchmark results included 99.76% accuracy and a 0.08% false-positive rate for the specified detector version and evaluation conditions. That evaluation used considerably larger datasets and a different classification methodology from our six-sample experiment.
Copyleaks published updated testing information on September 28, 2026, describing evaluations of its V11 model. Its published methodology included large collections of human-written and AI-generated material, separate testing by data science and quality-assurance teams, and multiple sensitivity settings.
The company reported a 0.30% false-positive rate for its balanced sensitivity setting. Those tests used its API, whereas our experiment used its web interface and a researcher-imposed 50% classification threshold.
Originality.ai also published updated accuracy information in September 2026. Its documentation described an AI Allowance approach that applies different permitted-AI thresholds rather than relying exclusively on a traditional binary Human-or-AI classification.
The company reported 99.4% accuracy on a large internal binary benchmark using its 15% AI Allowance configuration.
This distinction matters to our study because we did not evaluate Originality.ai using that published benchmark methodology or independently establish an equivalent allowance configuration.
These published figures demonstrate why accuracy comparisons require more than placing percentages side by side. Our experiment examined how the five accessible web services responded to six selected documents under one standardized rule.
It was not designed to reproduce or invalidate the vendors' larger benchmark studies.
What Independent Research Says About False Positives
The Stanford research discussed earlier found substantial misclassification of non-native English writing among the seven detectors evaluated in its 2023 experiment.
That study involved different tools, older detector versions, and a substantially larger collection of texts than our experiment. Its results cannot be applied directly to the five current services we tested. DOI
Similarly, the 2023 study published in the International Journal for Educational Integrity demonstrated that AI detection reliability can vary substantially depending on the detector, the type of error measured, and modifications to the submitted text. DOI
Our experiment contributes a smaller, more recent set of observations rather than independently confirming either study.
The relevant connection is methodological: all three investigations demonstrate why detector results require careful interpretation and why false positives and false negatives should be analyzed separately.
Limitations of This Study
Several limitations are essential when interpreting our findings. Sample size is the most significant limitation. We tested six documents, with only two samples in each writing category.
This means individual classification errors strongly affect the resulting percentages.
One additional error would change a detector's overall accuracy by approximately 16.7 percentage points. Within a two-sample category, one error changes the category-level result by 50 percentage points.
The reported results therefore cannot establish general detector accuracy or provide reliable population-level false-positive estimates.
Unbalanced Dataset for Category
The dataset was not balanced across every writing category. Academic and Blog/Marketing each contained one human and one AI sample, while the Non-Native English category contained two human samples.
Its results measure human-writing classification only and are not directly comparable with the mixed-origin categories as a general measure of accuracy.
Single Author For Non-Native Sample
Both non-native samples came from the same writer. The results do not represent the linguistic diversity of people who use English as an additional language.
We also did not establish a matched comparison group of native English speakers. Consequently, we cannot attribute differences between the non-native and other human samples specifically to language background.
AI Samples Only From ChatGPT
We did not test writing generated by Claude, Gemini, or other models. All our samples were from ChatGPT.
We did not preserve the exact ChatGPT model identifiers or complete generation prompts in the final dataset. This limits reproducibility and prevents detailed analysis of whether particular prompt characteristics influenced detector behavior.
Broad Writing Categories
Our writing categories were broad. The academic material included science-oriented explanatory writing, while the other categories contained different forms of online writing.
Genre, vocabulary, structure, and subject matter may affect detector behavior, but the dataset was too small to isolate those factors.
Standardized 50% Threshold
Different providers may assign different meanings to their displayed percentages. However, we maintained a standardized 50% threshold.
Our calculated classifications should not be confused with a formal evaluation of every detector's native classification system or recommended settings.
This distinction becomes especially clear when examining scores close to the threshold.
For example, QuillBot assigned SN01 a score of 49%, while GPTZero assigned SN02 a score of 54%. Changing the classification threshold could alter how those results are counted without changing any underlying detector output.
Exact Detector Model Not Recorded
Exact detector model versions were not recorded. We documented the September 28–30, 2026 testing period and used the web interfaces available then.
However, detector services may update their models, scoring systems, and interfaces without providing users with detailed version information. Future versions of the same tools may produce different results.
For these reasons, interpret the study as an exploratory comparison conducted during a specific three-day period.
Its value lies in documenting the results and disagreements observed in that experiment, not in claiming definitive accuracy figures for the wider AI detection industry.
What These Results Mean for Students, Teachers, and Writers
The practical implications of our findings depend on the context in which AI detection is used.
A false positive and a false negative may have very different consequences, particularly when detector output influences academic or professional decisions.
For Teachers and Educational Institutions
Our experiment included examples of genuine human writing receiving AI scores above the study's classification threshold.
The non-native English samples also produced conflicting results across detectors. These observations illustrate why a detector classification should not automatically be treated as proof of unauthorized AI use.
The University of Sydney's published guidance similarly states that AI detection scores should not constitute the only evidence in academic-integrity cases. It recommends considering them alongside other relevant information. The University of Sydney
When investigating potential unauthorized AI use, teachers can consider drafts, document revision histories, assignment requirements, sources, and the student's explanation of their writing process.
The appropriate process should follow the institution's established academic-integrity policies. Our related resource, How Teachers Should Interpret AI Detector Results, discusses these considerations in greater detail.
For Students
Students should understand that a detector's AI percentage is not necessarily a verified measurement of how much of their document was generated by artificial intelligence.
The interpretation depends on the specific tool and its scoring methodology. Our study demonstrates that known human-written documents can receive substantially different scores across services.
Students who write independently may benefit from retaining drafts, research notes, revision histories, and other evidence of their writing process. These records may provide useful context if questions arise about a submitted assignment.
Students should also understand their institution's rules regarding generative AI, including whether brainstorming, editing, translation, or other forms of AI assistance are permitted. For additional context, see What AI Percentage Is Acceptable?
For Writers, Editors, and Publishers
AI detectors can provide an additional signal during editorial review, but you should consider their output alongside other evidence. Our experiment found examples of human writing being classified as AI and AI-generated writing being classified as Human.
For content teams, this means detector results are more informative when combined with editorial review, source verification, revision history, and knowledge of the writer's production process.
A high AI score may warrant further investigation, but the score itself does not establish authorship. Similarly, a low AI score does not guarantee that a person wrote the content entirely.
Full AI Detector Accuracy Dataset
The following tables present the data recorded in our experiment. Each sample has a unique identifier, writing category, source classification, and known origin.
The five numerical columns preserve the original AI percentages recorded during testing.
Sample Information
| Sample ID | Category | Source | Actual Origin |
|---|---|---|---|
| SA01 | Academic | Archived writing | Human |
| SM01 | Blog/Marketing | Online material published before ChatGPT | Human |
| SN01 | Non-Native English | Original team writing | Human |
| SA02 | Academic | ChatGPT-generated | AI |
| SM02 | Blog/Marketing | ChatGPT-generated | AI |
| SN02 | Non-Native English | Original team writing | Human |
Both non-native samples were written by the same person.
Complete Recorded AI Scores
| Sample | CopyChecker | Copyleaks | GPTZero | Originality.ai | QuillBot |
|---|---|---|---|---|---|
| SA01 | 5% | 0% | 0% | 15% | 0% |
| SM01 | 30% | 0% | 0% | 97% | 66% |
| SN01 | 78% | 80% | 0% | 100% | 49% |
| SA02 | 68% | 100% | 100% | 100% | 8% |
| SM02 | 54% | 100% | 100% | 100% | 0% |
| SN02 | 22% | 0% | 54% | 15% | 13% |
Each percentage represents the numerical score recorded from the corresponding detector's web interface.
We then applied the standardized 50% threshold to calculate the classifications used throughout this article.
Correct and Incorrect Classifications
In the following table, 1 indicates a correct classification and 0 indicates an incorrect classification.
| Sample | CopyChecker | Copyleaks | GPTZero | Originality.ai | QuillBot |
|---|---|---|---|---|---|
| SA01 | 1 | 1 | 1 | 1 | 1 |
| SM01 | 1 | 1 | 1 | 0 | 0 |
| SN01 | 0 | 0 | 1 | 0 | 1 |
| SA02 | 1 | 1 | 1 | 1 | 0 |
| SM02 | 1 | 1 | 1 | 1 | 0 |
| SN02 | 1 | 1 | 0 | 1 | 1 |
We calculated the correctness values by comparing each detector's threshold-based classification against the Actual Origin column.
Reproducibility Notes
Readers can reproduce the reported accuracy, false-positive, and false-negative calculations using the sample origins, recorded percentages, and 50% threshold provided above.
The complete spreadsheet additionally separates the numerical scores, calculated Human/AI classifications, and binary correctness results.
The study's conclusions are limited to the recorded data and testing conditions.
Original source texts and their URLs have not been included in this publication.
Conclusion
Our September 2026 experiment produced 30 detector evaluations across six human and AI-generated writing samples.
Twenty-two classifications matched the recorded origins, for a combined descriptive accuracy of 73.3%. However, overall accuracy alone concealed important differences in the results.
Some detectors identified both AI-generated samples but incorrectly classified genuine human writing. Another correctly identified several human samples while missing both AI-generated documents.
The two non-native English samples also revealed substantial disagreement, including cases where the same known human-written text received sharply different assessments.
These observations reinforce the importance of examining false positives, false negatives, and sample-level results rather than relying exclusively on a single accuracy percentage.
Our study's small sample size, limited writing categories, single AI-generation source, and standardized threshold prevent broad conclusions about general detector performance.
The results are therefore best understood as a documented exploratory experiment conducted between September 28 and September 30, 2026.
They demonstrate what happened when five AI detectors examined the same six documents, while identifying questions that require more extensive testing.
Frequently Asked Questions
How Accurate Are AI Detectors?
Accuracy depends on the detector, its underlying model, the writing being tested, and the evaluation methodology. In our experiment, individual accuracy figures ranged from 50% to 83.3%. Because we tested only six documents, these figures should not be generalized to other datasets or writing categories.
How Often Do AI Detectors Produce False Positives?
In our experiment, individual false-positive rates ranged from 25% to 50%. We tested each detector against only four human-written documents. Consequently, a single incorrect classification represented 25 percentage points. The recorded rates describe our selected samples, not the frequency of false positives across all writing.
Can Human Writing Be Detected as AI?
Yes. Our experiment included several examples of genuine human-written documents that received scores above the standardized AI classification threshold.
Are AI Detectors Less Accurate for Non-Native English Writers?
Our two genuine non-native English samples produced four false positives across ten detector evaluations. The two other human-written samples produced two false positives across ten evaluations.
Are AI Detectors More Accurate on Academic Writing?
The results varied by detector. Four services correctly classified both documents in our academic category, while one correctly classified only the human sample.
Can AI Detector Scores Prove That Someone Used AI?
An AI detection score should not be treated as definitive proof of authorship. Different detectors may use different scoring systems, and incorrect classifications can occur.
How Was This AI Detector Accuracy Study Conducted?
We tested six known-origin writing samples across five AI detection services between September 28 and September 30, 2026. We recorded each detector's numerical AI score. We then applied a standardized 50% threshold and classified the results as Human or AI.
Which AI Detectors Were Included?
Our experiment included CopyChecker, Copyleaks, GPTZero, Originality.ai, and QuillBot. We tested all five through their web interfaces during the same three-day testing period.
