InDetect

Turnitin AI Detection Test 2026: GPT-5.5, Claude Opus 4.8 & Gemini 3.1 Pro Results

Author image
Written by  Eleanor Rose Sterling
2026-07-05 11:26:28 7 min read

Bottom line: Turnitin's AI detection is remarkably effective. We tested 9 AI-generated texts across three frontier models — including texts that went through dedicated humanization — and Turnitin flagged every single one. Meanwhile, real human text scored 0%. Here's the full breakdown.


Why We Ran This Test

AI writing tools are everywhere. ChatGPT, Claude, and Gemini are used daily by students, researchers, and content creators. Turnitin — deployed across 16,000+ institutions — is the de facto gatekeeper for academic integrity, and its AI detection feature increasingly determines whether a paper gets flagged.

Three questions keep coming up:

  1. Can Turnitin actually catch AI-generated text from the latest models?

  2. Do "humanization" techniques — rewriting, prompt engineering — actually lower the AI score?

  3. Does Turnitin produce false positives on real human writing?

We ran a controlled test to answer all three.


How We Tested

The Models

Three frontier models as of mid-2026:

  • ChatGPT (OpenAI GPT-5.5)

  • Claude Opus 4.8

  • Gemini 3.1 Pro

The Three Writing Modes

For each model, we generated text three ways:

Mode

What We Did

What It Tests

Direct Generation

Prompt the model to write about marine litter degradation. No extra instructions.

Baseline AI detection rate

Humanized Rewriting

Generate first, then ask the model to rewrite "to sound more natural and human."

Whether post-hoc editing evades detection

Humanized Requirements

Include humanization instructions in the initial prompt.

Whether prompt engineering evades detection

The exact prompts used are documented in the appendix.

The Human Control Group

To test for false positives, we submitted real human writing that predates large language models:

  • 3 excerpts from real legal documents (U.S. Constitution, Bill of Rights — formal, highly structured language)

  • 9+ paragraphs from pre-2015 academic papers (dissertation acknowledgments, economics papers — impossible to be AI-generated)

Topic and Length

All AI texts covered marine litter degradation — a neutral, factual topic that keeps output consistent across models. Every text was at least 3 paragraphs to ensure Turnitin had enough material to analyze.

Due to Turnitin's per-check cost, we merged multiple texts into grouped submissions (no more than 4 per submission).


The Results

AI Detection Scores

Model

Writing Mode

AI Score

GPT-5.5

Direct Generation

100%

GPT-5.5

Humanized Rewriting

100%

GPT-5.5

Humanized Requirements

100%

Claude Opus 4.8

Direct Generation

100%

Claude Opus 4.8

Humanized Rewriting

100%

Claude Opus 4.8

Humanized Requirements

100%

Gemini 3.1 Pro

Direct Generation

100%

Gemini 3.1 Pro

Humanized Rewriting

100%

Gemini 3.1 Pro

Humanized Requirements

95%

Human Text Results

Text Type

AI Score

Legal documents (U.S. Constitution, Bill of Rights)

0%

Pre-2015 academic papers

0%


Screenshots: Turnitin AI Detection Reports

GPT-5.5

Direct Generation — AI: 100%

GPT-5.5 Direct Generation Turnitin AI Detection Report

Humanized Rewriting — AI: 100%

GPT-5.5 Humanized Rewriting Turnitin AI Detection Report

With Humanized Requirements — AI: 100%

GPT-5.5 with Humanized Requirements Turnitin AI Detection Report


Claude Opus 4.8

Direct Generation — AI: 100%

Claude Opus 4.8 Direct Generation Turnitin AI Detection Report

Humanized Rewriting — AI: 100%

Claude Opus 4.8 Humanized Rewriting Turnitin AI Detection Report

With Humanized Requirements — AI: 100%

Claude Opus 4.8 with Humanized Requirements Turnitin AI Detection Report


Gemini 3.1 Pro

Direct Generation — AI: 100%

Gemini 3.1 Pro Direct Generation Turnitin AI Detection Report

Humanized Rewriting — AI: 100%

Gemini 3.1 Pro Humanized Rewriting Turnitin AI Detection Report

With Humanized Requirements — AI: 95%

Gemini 3.1 Pro with Humanized Requirements Turnitin AI Detection Report


Human Text

Real-world Legal Documents — AI: 0%

Legal Documents Turnitin AI Detection Report

Pre-2015 Academic Papers — AI: 0%

Pre-2015 Academic Papers Turnitin AI Detection Report


Analysis

1. Turnitin Caught Everything

Across three different frontier models, across three different writing strategies, Turnitin flagged every single AI-generated text. Not one slipped through.

The lowest score was 95% — and that is still an unambiguous "this is AI" flag. No instructor would look at a 95% AI detection score and think "maybe this is human-written."

2. Humanization Does Not Work

We tested the two most common humanization strategies:

  • Post-generation rewriting: "Make this sound more human."

  • Pre-generation prompt engineering: "Write like a thoughtful human, not a template."

Neither worked. GPT-5.5 and Claude stayed at 100% regardless. Gemini budged from 100% to 95% — a negligible change that still screams "AI."

This tells us something important: Turnitin is not looking at surface-level features like transition words or sentence uniformity. The patterns it detects — likely statistical regularities in word choice, syntax, and structure — survive paraphrasing, restructuring, and stylistic adjustments.

Looking at the texts side by side, the humanized versions did read differently. Sentence lengths varied more. Formulaic transitions were reduced. Some passages adopted a more observational tone. But Turnitin saw through all of it. The underlying signal was too strong.

3. Zero False Positives on Human Text

This is equally important. When we submitted formal legal text and pre-2015 academic writing:

  • U.S. Constitution and Bill of Rights: 0% AI. These are highly structured, formal texts — exactly the kind critics worry might be flagged.

  • Pre-2015 dissertation and economics papers: 0% AI. These were written before GPT existed, so Turnitin correctly identified them as human.

The fear that Turnitin might flag formal, structured human writing as AI-generated is one of the most common objections to AI detection tools. In our test, that fear did not materialize.

4. Why Gemini 3.1 Pro Dropped to 95%

The only blip in the data was Gemini 3.1 Pro with humanized requirements, which scored 95% instead of 100%. What was different about that text?

  • It was longer (5 paragraphs vs. 4 for the direct generation)

  • It opened with a narrative scene ("Standing on a beach...")

  • It used more rhetorical questions and reflective language

  • Sentences varied more in length and rhythm

Despite all of that, 95% is still a definitive AI flag. The difference is academically interesting but practically meaningless — no one would mistake a 95% score for human writing.


Key Takeaways

Turnitin's AI Detection Is Extremely Accurate

100% detection on 8 of 9 AI texts, 95% on the ninth. Turnitin is not making probabilistic guesses — it is reliably identifying AI-generated content from three different frontier models with near-perfect consistency.

Humanization Is a Dead End (Against Turnitin)

The "AI humanization" industry promises to make generated text undetectable. Our test shows that, at least against Turnitin, these claims do not hold up:

  • Rewriting after generation: zero effect

  • Humanizing during generation: negligible effect (5% drop on one model, zero on the other two)

The patterns Turnitin detects are too fundamental — they survive paraphrasing, restructuring, and stylistic adjustments.

Turnitin Does Not False-Posit on Human Writing

Our control group scored 0% AI. Formal legal prose and academic writing — styles that critics worry might be misclassified — were correctly identified as human. The false-positive fear, at least for formal and academic writing, appears unfounded based on our results.


Limitations

Be transparent about what this test does and does not cover:

  1. Sample size: 9 AI texts and 2 human text groups. Meaningful but not exhaustive.

  2. Single topic: All AI texts were about marine litter. A personal narrative or creative topic might produce different patterns.

  3. Formal writing only: We tested informative prose, not creative writing, poetry, or technical content.

  4. Snapshot in time: Turnitin updates its algorithms. These results reflect July 2026.

  5. Merged submissions: Texts were grouped to manage costs. This could theoretically affect scoring, though Turnitin operates at the sentence level.

  6. No adversarial techniques: We tested straightforward humanization. We did not test aggressive obfuscation (character-level perturbations, intentional errors, back-translation) — methods that typically degrade text quality to the point of being useless.


What This Means

For Students

If you are generating text with AI and hoping Turnitin will not notice — it will. The detection rate in our test was essentially 100%.

The better approach: use AI as a tool, not a replacement. Brainstorm ideas, outline arguments, refine your own prose. Keep your voice and ideas at the center. If you do use AI assistance, be transparent per your institution's policies.

For Educators

Turnitin's AI detection is a reliable signal — but it is still a signal, not a verdict. A 100% AI score should trigger a conversation, not an automatic penalty. Review the student's writing history, discuss the assignment, consider the full context.

Our test also confirms that formal, well-structured human writing does not trigger false positives. You can trust a 0% score with reasonable confidence.

For Content Creators

If you need content that passes AI detection, generating text with any major model and then "humanizing" it is unlikely to work against Turnitin. The most reliable strategy is to write the content yourself and use AI as a supplementary tool.


Test Methodology Reference

Parameter

Detail

Test Date

July 2026

Models Tested

ChatGPT (GPT-5.5), Claude Opus 4.8, Gemini 3.1 Pro

AI Texts

9 (3 per model, 3 writing modes each)

Human Texts

Legal documents (3 excerpts), Pre-2015 academic papers (9+ paragraphs)

Topic

Marine litter degradation

Minimum Length

3 paragraphs per text

Detection Tool

Turnitin AI Detection

Submission Method

Merged documents (≤4 texts per submission)


Appendix: Prompts Used

For full reproducibility, here are the exact prompts used for each writing mode. These were applied identically across all three models.

Prompt 1: Direct Generation

Please write a short article about "the degradation of marine litter" in English.

Requirements:
1. The article should contain at least three paragraphs.
2. Discuss the sources of marine litter, its environmental impacts, the difficulty of degradation, and the importance of reducing marine pollution.
3. Use a clear and formal style suitable for a general science or environmental awareness article.
4. Do not use headings or bullet points.
5. Do not include fabricated statistics, studies, organizations, or citations.
6. Keep the writing coherent and logically structured.

Prompt 2: Humanized Rewriting

This prompt was applied to the output of Prompt 1 — the model was asked to revise its own text.

Please revise the following article to make it read more naturally, as if it were written by a real person with a thoughtful writing style.

Revision requirements:
1. Preserve the original meaning, main ideas, and overall topic.
2. Do not add fabricated statistics, studies, organizations, cases, or citations.
3. Vary the sentence length and rhythm so the writing does not feel overly uniform.
4. Reduce mechanical transitions such as "firstly," "secondly," "moreover," "in conclusion," or similar formulaic phrases.
5. Make the tone more natural while still keeping it professional and clear.
6. Avoid making the article too casual or conversational.
7. Keep at least three paragraphs.

Original article:
[Paste the original article here]

Prompt 3: Generation with Humanized Requirements

This prompt was used for a fresh generation — the humanization instructions were embedded from the start, rather than applied as a second pass.

Please write a short article about "the degradation of marine litter" in English.

Content requirements:
1. The article should contain at least three paragraphs.
2. Discuss the main sources of marine litter.
3. Explain why marine litter, especially plastic waste, is difficult to degrade in ocean environments.
4. Describe its impacts on marine life, ecosystems, and human society.
5. Mention the importance of reducing marine litter and improving waste management.

Writing style requirements:
1. Write in a natural and thoughtful style, as if the article were written by a real person rather than generated from a fixed template.
2. Avoid making every paragraph follow the same structure.
3. Vary sentence length and rhythm.
4. Include a few natural observations or reflections where appropriate.
5. Avoid overusing formulaic transitions such as "firstly," "secondly," "finally," or "in conclusion."
6. Do not use headings or bullet points.
7. Do not include fabricated statistics, studies, organizations, cases, or citations.
8. Keep the writing clear, accurate, and suitable for general readers.

Key Design Decisions

The three prompts were designed to test a specific hypothesis: can simple prompt rewriting reduce Turnitin AI detection scores?

  • Prompt 1 represents the most common use case — a student pasting an assignment prompt into ChatGPT and submitting the output.

  • Prompt 2 simulates the popular "humanize my text" workflow — generate first, then ask the AI to rewrite it to sound less artificial.

  • Prompt 3 tests whether baking humanization instructions into the initial prompt is more effective than applying them after the fact.

The prompts deliberately avoid fabricated statistics and citations to ensure the test isolates the model's writing style rather than its factual accuracy.


Final Verdict

Turnitin's AI detection passed this test convincingly. It caught AI-generated text from all three frontier models — GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro — at 100% (or 95% in one case). It correctly identified human-written text at 0%. Humanization techniques made no meaningful difference.

Simple prompt rewriting cannot bypass Turnitin's AI detection. This is the most important takeaway from our test. We tried two distinct approaches — post-generation rewriting and pre-generation style instructions — and both failed. Across GPT-5.5, Claude Opus 4.8, and Gemini 3.1 Pro, the results were unanimous: Turnitin detected the AI signal regardless of how the prompt was phrased. The statistical patterns that Turnitin identifies are embedded at a level too deep for surface-level stylistic adjustments to mask.

The message is clear: if you submit AI-generated text through Turnitin, it will be detected. The best way to avoid an AI flag is to write your own content.