A humanizer tool can score close to perfect on passing a detector and still fail at the one job that actually matters more: keeping the original meaning intact. A benchmark that only reports a detector pass rate misses this entirely, which is why a serious comparison needs a second measure sitting right alongside the first one, counting what the rewrite actually changed, not just whether it reads as human.
This second measure rarely gets the same attention as a headline pass rate, but it answers a question that matters just as much to anyone actually relying on a rewrite to represent their own work accurately, which is exactly why it deserves equal billing rather than a footnote.
A Pass Rate Alone Answers the Wrong Question
A detector score measures statistical texture, whether a passage reads as predictable and uniform in the way AI-generated text tends to, or varied and irregular in the way human writing tends to. It says nothing about whether the rewritten version still claims the same facts, cites the same figures, or represents the same argument as the original text handed to the tool.
A rewriting engine that achieves a high pass rate by substantially altering content, swapping a statistic, changing a date, inventing a supporting detail that was not in the source, has solved a narrower problem than it appears to, passing a detector while quietly introducing inaccuracies a reader would have no way to notice without comparing the rewrite against the original line by line.
This is precisely the failure mode a detector-only benchmark cannot catch, since nothing about a passing score indicates whether the underlying content still matches what was originally said.
How Phrasly’s Benchmark Defines and Counts This
Phrasly’s September 2026 benchmark defines a hallucination specifically as an unsupported factual addition or change in an output relative to its input, reported two ways: the average count per output across each tool’s completed rewrites, and the share of outputs with no hallucinations at all among those with a recorded count.
Across the four tools tested on the same 72 texts, Phrasly Ultra averaged 0.85 hallucinations per output with 48.6 percent of its outputs entirely hallucination-free. StealthGPT averaged 3.54 per output with 5.6 percent hallucination-free. WriteHuman averaged 2.64 per output with 14.5 percent hallucination-free. Undetectable AI averaged 6.16 per output with 1.4 percent hallucination-free.
The spread between tools here is considerably wider than the spread in detector pass rates, suggesting meaning preservation and detection evasion are not just conceptually separate measures, but genuinely different capabilities that a rewriting model can develop independently of one another. Undetectable AI, for instance, recorded the lowest detector pass rate among the lenient-measure tools tested yet still produced more than seven times as many hallucinations per output as Phrasly Ultra, a combination that would not be obvious from a pass rate alone.
Why This Measure Deserves Equal Weight to a Pass Rate
A detector pass rate and a hallucination count answer two genuinely independent questions, and a tool can score well on one while scoring poorly on the other. A rewrite that changes very little of the original wording may preserve meaning almost perfectly while still reading as statistically predictable enough to fail a detector. A rewrite that changes a great deal may pass a detector easily while drifting substantially from what the original text actually said.
This independence is why a benchmark reporting only one of the two measures gives an incomplete picture no matter how rigorously that single measure was tested, since a reader has no way to know which of these two very different tradeoffs a high-scoring tool actually made to get there.
Reading both measures together, rather than either one alone, reveals different kinds of tools:
- High pass rate, low hallucination count: a rewrite that changes style effectively while preserving meaning
- High pass rate, high hallucination count: a rewrite that defeats detection partly by substantively altering content
- Low pass rate, low hallucination count: a cautious rewrite that preserves meaning but does not change enough to pass
- Low pass rate, high hallucination count: a rewrite that changes content without the stylistic shift needed to actually pass
- Checking where a specific tool actually falls in this matrix, rather than assuming, is only possible when both measures are reported
The full methodology and per-tool figures behind these categories are published in the Phrasly Benchmark, including the downloadable data needed to check how each tool’s specific balance of these two measures was calculated.
Why This Matters Beyond a Single Benchmark
Any comparison of rewriting tools that reports only a detector pass rate is, by omission, treating meaning preservation as unimportant. For a student protecting their own genuine argument from a false-positive flag, a professional preserving a client’s exact claims, or a researcher protecting a precise technical description, a rewrite that passes a detector while quietly changing what was actually said solves the wrong problem entirely.
A detector score and a hallucination count, read together, give a much more complete picture of what a rewriting tool actually does to a piece of writing, and a reader evaluating any such tool is better served asking for both numbers than settling for whichever one a vendor chose to lead with, since the one left out is often the one that would have changed the decision.
For more on how AI detection and meaning preservation are measured together, further reading on the Phrasly blog covers the underlying research for anyone evaluating a rewriting tool on more than one dimension, including how meaning preservation is measured alongside detection scores.