A score can hear static and miss the sentence
A synthetic voice can sound clean and still say a word badly. It can stress the wrong syllable, flatten a question or pause in a place that changes the sentence.
A new ServiceNow study asks whether automatic voice evaluators hear those mistakes the way people do. The researchers tested four traditional mean-opinion-score predictors and four audio-capable language models on 860 speech samples labelled by trained linguists.
The short answer is no, at least not consistently. The score predictors were good at obvious acoustic damage. The audio language models noticed a different, prompt-dependent mix of problems. None covered the full set of ten dimensions.
That is a problem for teams trying to improve a voice system without asking people to listen to every version.
Ten ways speech can go wrong
The benchmark breaks the vague idea of naturalness into ten parts. At word level, it checks phonetic accuracy and lexical stress. At sentence level, it checks intonation, emphasis, boundaries and speaking rate. The final group covers emotion, expressiveness, speaker identity and whether the voice sounds plausibly human.
Researchers started with Harvard Sentences, material from another TTS benchmark and generated text. They created targeted errors by changing text or phonetic input, or by manipulating audio directly. Cartesia Sonic-3 produced the synthetic speech.
Three first-language English linguists labelled every sample on every dimension. Majority votes became the ground truth, then the team balanced positive and negative examples for each test.
This design makes it possible to ask what an evaluator is actually hearing. It also makes the task more controlled than everyday voice use.
The judges listened differently
The four MOS predictors were NISQA, DNSMOS-Pro, UTMOSv2 and four variants of Meta's Audiobox Aesthetics model. Their strongest results clustered around speaking rate, expressiveness and human plausibility — changes tied closely to the acoustic signal.
None of those predictors significantly detected phonetic accuracy, lexical stress or prosodic boundary placement. Some significant relationships even ran in the wrong direction, meaning a degraded sample could receive the more favourable score.
The audio language models were Gemini 3 Flash, Gemini 3.5 Flash, Qwen3-Omni and Step-Audio-2-mini. Under a generic naturalness prompt, their sensitivity was sparse. Gemini covered the broadest set in some conditions; the other systems usually found only a few dimensions, if any.
All eight evaluators were relatively good at the most visible kind of failure: a glitch inserted to make speech sound less human. That does not mean they understood why a sentence sounded wrong.
Better prompting helped, but only in places
The team then changed how the audio models were asked to judge. A detailed ten-part rubric recovered some word-level sensitivity that the generic prompt missed. Asking about one dimension at a time usually aligned better with the linguists than asking for everything together.
But the gains did not travel evenly. Some model-and-prompt combinations returned the same score for every sample, leaving no useful correlation to measure. Other improvements appeared in one dimension and disappeared in another.
This is a practical warning for automated evaluation. A capable audio model is not automatically a reliable judge, and the prompt is part of the measuring instrument.
If a product team cares about pronunciation, timing and emotion, it may need separate checks for each. One overall naturalness number hides too much.
What is confirmed, found and still open
Confirmed: the nine-author ServiceNow paper was submitted on 10 August 2026. It evaluates four MOS model families and four audio language models against majority labels from three trained linguists on 860 English samples.
The research finding: current automated evaluators had different blind spots. MOS predictors concentrated on signal-level changes, while audio language models were selective and sensitive to prompt design. No system matched human labels across all ten dimensions.
The authors' claim: the dataset, annotation schema and evaluation code are publicly released. The paper does not provide a direct repository or dataset link, so Model Current could not independently inspect those artefacts at publication time.
Still open: whether the result holds for longer conversation, other languages and naturally occurring failures. The test used short sentences, controlled manipulations and one TTS architecture for phoneme-level errors. Two dimensions also had smaller balanced samples: 46 for emotional appropriateness and 90 for lexical stress.
The benchmark is not a ranking of voice providers. It is a reminder that sounding natural is not one thing — and measuring it probably should not be either.
Sources
- Bamgbose et al. — Beyond NaturalnessPrimary paper record submitted 10 August 2026. Source for authorship, sample size, evaluated model classes and headline findings.
- Bamgbose et al. — Full experimental manuscriptFull primary manuscript. Source for the ten dimensions, prompts, model list, statistical results, sample construction and limitations.
- Cartesia — Sonic text-to-speech documentationPrimary vendor documentation for the speech model interface used to generate the study's synthetic samples. It does not validate the paper's evaluation claims.



