Epoch AI Tests Three Leading AI Text Detectors: Up to Nearly 30% of Human-Like AI Content Goes Undetected
Epoch AI has found that leading AI text detectors are highly effective at identifying standard AI-generated content, but their accuracy drops significantly when large language models are deliberately prompted to mimic human writing styles. The study found that scientific writing is the most challenging scenario, with some tools failing to detect nearly 30% of AI-generated text written in an academic style.

New research from Epoch AI suggests that today’s leading AI writing detectors still perform extremely well on standard machine-generated text, but their accuracy drops once large language models are instructed to imitate a particular author’s style. The decline is especially noticeable in scientific and academic writing, which emerged as the most difficult category for detection.
Three major detectors were evaluated
The study examined three widely used AI text detection systems: Pangram (3.3.2), GPTZero (2026-05-11-base), and Originality.ai (Turbo 3.0.2).
To create a benchmark for comparison, the researchers assembled a test set of 495 human-written texts spanning three types of writing:
Blog content
Fiction
Scientific writing
All of these samples were written before the public launch of ChatGPT in November 2022, a step intended to reduce the possibility that the texts had already been absorbed into later model training data.
Ordinary AI-generated text was usually detected
When the detectors were tested on conventional AI-generated writing, all three systems performed strongly. The highest miss rate recorded in this setting was just 0.7%, indicating that routine AI output remains relatively easy for current tools to flag.
Performance on genuine human writing was also mostly solid, though not identical across tools. Pangram and GPTZero produced no false positives in the human sample set. Originality.ai, however, incorrectly labeled 19 human-written pieces as AI-generated, which corresponds to a 3.8% false positive rate.
Style imitation made detection substantially harder
The picture changed once the researchers moved beyond generic AI writing. In the next phase of the study, Claude Opus 4.8, GPT-5.5, and Gemini 3.1 Pro were each given five authentic texts from a specific author and asked to generate new material in that same style.
Across 297 style-mimicked samples, the detectors failed to identify roughly 13% on average. Broken down by tool, the miss rates were:
Pangram: 10%
GPTZero: 11%
Originality.ai: 18%
In other words, once language models were guided to reproduce the tone, rhythm, and phrasing patterns of a known human writer, the reliability of current detection systems weakened noticeably.
Scientific writing proved the toughest category
Among all tested domains, scientific writing created the biggest challenge for every detector in the study. For AI-generated text written in an academic style, the miss rates climbed sharply:
Pangram: 25%
GPTZero: 24%
Originality.ai: 29%
Some model-detector pairings were even more problematic. Pangram failed to catch 48% of the academic-style texts generated by Gemini. Originality.ai missed 39% of the scientific samples produced by GPT-5.5.
These results suggest that academic prose may be particularly difficult to distinguish from human work when AI systems are trained or prompted to mimic established writing conventions.
Different detection methods showed similar weaknesses
Although the three detectors rely on different technical approaches—including neural network methods, predictability-based analysis, and statistical pattern recognition—the study found that they struggled in broadly similar ways.
The shared limitation is clear: as language models become better at reproducing human stylistic signals, current AI text detectors become less dependable. According to Epoch AI, this creates an ongoing challenge for settings where the authenticity of written work matters most, particularly in education and research.
Overall, the findings point to a growing gap between what AI systems can imitate and what existing detection tools can confidently identify. Standard AI writing may still be easy to spot, but author-style imitation—especially in scientific contexts—is making that task far more uncertain.