Paper deep dive
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct "naturalness" into a linguistically grounded annotation schema spanning 10 distinct perceptual dimensions, and use it to construct the first dimension-level meta-evaluation benchmark for TTS, comprising 860 utterances annotated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLM judges reveal that MOS predictors collapse onto acoustic signal quality, while Audio-LLM judges show selective, prompt-dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation.
Tags
Links
- Source: https://arxiv.org/abs/2608.09930v1
- Canonical: https://arxiv.org/abs/2608.09930v1
Trouble viewing inline? Open PDF directly â
Full Text
101,505 characters extracted from source content.
Expand or collapse full text
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions Oluwanifemi Bamgbose * , Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols ServiceNow nifemi.bamgbose@servicenow.com Abstract Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predic- tors and Audio Large Language Models (Audio- LLMs) judges) are expected to reflect human perception, yet it is unclear how well they cap- ture the distinct aspects of speech that listen- ers actually perceive. We deconstruct "nat- uralness" into a linguistically grounded an- notation schema spanning 10 distinct percep- tual dimensions, and use it to construct the first dimension-level meta-evaluation bench- mark for TTS, comprising 860 utterances an- notated by trained linguist raters. Results from benchmarking four MOS predictors and four Audio-LLMs judges reveal that MOS predic- tors collapse onto acoustic signal quality, while Audio-LLMs judges show selective, prompt- dependent detection that does not generalise across all dimensions. Neither class reliably captures a breadth of linguistically structured speech errors. Our dataset, annotation schema, and evaluation code are publicly released to support more targeted and interpretable TTS evaluation. 1 Introduction Text-to-speech (TTS) systems are increasingly de- ployed in production contexts where the natural- ness of synthesized speech directly affects user experience and comprehension, with potential con- sequences for user engagement (Tan et al., 2021; Xie et al., 2025). Listeners are sensitive to a mul- tidimensional set of failures; a TTS system may stress the wrong syllable in a word, break a sen- tence at an unnatural point, or deliver a question with the flat intonation of a statement, any of which would be immediately human-perceptible (Geneva et al., 2023; Gutierrez et al., 2021; Lee et al., 2026). The gold standard for measuring TTS quality across these dimensions remains human perceptual judgment (Baughan et al., 2023; Ulgen et al., 2026). * Corresponding author. Table 1: Capability comparison with recent meta- evaluation resources. ETE: EmergentTTS-Eval (Manku et al., 2025); ITE: InstructTTSEval (Huang et al., 2025); SJ: SpeechJudge (Zhang et al., 2025); S: Speaker- Sleuth (Lee et al., 2026).âpresent, ~ partial,âabsent; partial=human labels only validate a model judge (ETE/ITE), or controlled difficulty without injected lin- guistic errors (S). Only ours pairs expert, aspect-level human labels with a linguistically-grounded taxonomy and controlled error injection, auditing both MOS pre- dictors and Audio-LLM judges. CapabilityOursETEITESJSS Human-annotated ground truthâ~~â Dimension-level annotation per sampleââ Linguistically-grounded aspect taxonomyââ Covers MOS predictors & Audio-LLM judgesâââ As deployment scales, however, relying solely on human judgment becomes expensive and time con- suming. Two main paradigms have emerged to automatically evaluate speech quality at scale: neu- ral Mean Opinion Score (MOS) predictors (Mittag et al., 2021; Saeki et al., 2022), trained to output in- dividual, holistic quality scores that reflect human MOS; and Audio Large Language Model (Audio- LLM) Judges (Zhang et al., 2025; Manakul et al., 2026), which prompt audio-capable LLMs to rate or compare audio samples. A good speech eval- uator should respond to all aspects of perceptual quality to which humans attend, yet it is not clear whether models of either type do so (i.e whether in- dividual output scores align with quality along any or all human-perceptible dimensions). For Audio- LLMs in particular, it additionally remains an open question whether they even have the underlying ca- pability, and/or can be prompted, to attend to each dimension. To better understand and identify the aspects of speech to which MOS models and Audio-LLMs ac- tually attend, we construct the first meta-evaluation benchmark for automated TTS evaluators that has annotations based on linguistically-grounded, human-perceptible speech dimensions. Our contributions are as follows: arXiv:2608.09930v1 [cs.SD] 10 Aug 2026 Figure 1: Dataset Construction Pipeline. Sentences from three sources (Harvard Sentences, EmergentTTS-Eval, synthetic text) are fed through three error-generation pathways: LLM-modified plain text, LLM-modified IPA transcriptions, and direct acoustic manipulation. All pathways pass through the TTS model to produce clean (Score 1) and error-targeted samples. Samples are then annotated by three linguists on all ten dimensions. After linguist annotation, majority filtering and downsampling produce a balanced final dataset of 860 samples. âą Annotation Schema. A linguistically grounded evaluation schema decomposing naturalness into 10 phonologically and phonetically grounded di- mensions across word, prosodic, and paralinguis- tic levels. This enables future evaluation to be aligned with human perception of speech quality. âąMeta-Evaluation Dataset. A per-dimension annotated dataset of 860 utterances, labeled by trained linguist raters across the 10 dimensions. âąDimension-Level Audit. An evaluation of four neural MOS models as well as four Audio- LLMs judges across four prompting conditions, with and without reference transcripts. This reveals uneven phonetic, prosodic and paralin- guistic sensitivity, with neither MOS predictors nor Audio-LLMs judges reliably detecting er- rors across these dimensions under any condition tested. Our benchmark enables researchers and practition- ers to attribute TTS system failures to specific in- terpretable dimensions rather than opaque holistic scores. This provides a diagnostic foundation that prior evaluation frameworks have lacked. 2 Related Works Automated TTS Evaluation. The dominant paradigm for scalable TTS evaluation has been MOS predictors: learning to approximate human naturalness scores from raw waveforms without a reference signal. Through the VoiceMOS Chal- lenge series (2022â2024) (Huang et al., 2022, 2024), Semi-supervised Learning (SSL)-based MOS predictors (Patton et al., 2016; Reddy et al., 2021; Saeki et al., 2022; Baba et al., 2024) have demonstrated strong system-level correlation with human judgments, approaching the practical ceil- ing of aggregate naturalness scoring. However, these methods produce holistic quality signals without grounding them in independently inter- pretable linguistic attributes, leaving no diagnostic path when a system performs poorly. More re- cently, researchers have explored leveraging Audio- LLMs as richer judges for TTS evaluation, evalu- ating speaking style (Chiang et al., 2025), speech quality (Zhang et al., 2025; Manakul et al., 2026; Monjur and Nirjon, 2025), and speaker consis- tency (Lee et al., 2026). Yet, LLM-based judges inherit the same limitation: scores remain holistic, and the perceptual dimensions driving any given rating are opaque. TTS Evaluation Benchmarks. Several bench- marks have moved beyond scalar MOS towards richer TTS evaluation (Manku et al., 2025; Huang et al., 2025), yet share a fundamental limitation: evaluation dimensions are organized around pre- defined scenario categories rather than a prin- cipled speech quality schema.EmergentTTS- Eval (Manku et al., 2025) introduces a model- as-a-judge benchmark spanning six pre-defined scenario types (emotions, paralinguistics, foreign words, syntactic complexity, complex pronuncia- tion, and questions) and demonstrates that Audio- LLMs judges can distinguish broad scenario-level variation. InstructTTS-Eval (Huang et al., 2025) takes a complementary direction, centralizing eval- uations on expressive and style-following capa- bility of TTS systems across pre-defined dimen- sions such as speaking rate, pitch, and emotional expressiveness. However, neither provides human- annotated ground truth at the attribute level, making it impossible to audit whether automated evalua- tors track the aspects listeners actually perceive. We bridge this gap with a human-annotated dataset grounded in 10 interpretable aspects across three linguistically-motivated quality dimensions, paired with a calibrated annotation framework, enabling direct attribution of TTS system failures to specific and interpretable aspects. More information related to other benchmarks can be found at Table 1 3 Benchmark Design 3.1 Evaluation Schema Current TTS evaluation conflates perceptually dis- tinct failure types under a single quality judgment. Terms such as ânaturalness,â âfluency,â and âclarityâ are used interchangeably in MOS elicitation, yet listeners explain their ratings by invoking factors ranging from pronunciation accuracy to affective delivery (Kirkland et al., 2023), aspects that are not only perceptually distinct but defined at dif- ferent levels of linguistic structure. Drawing on Crystalâs distinction between linguistic and paralin- guistic features (Crystal, 1969), we organize evalu- ation around three levels corresponding to where each property is linguistically specified: the lexi- con, the utterance, and the speaker. We propose a ten-dimension schema in which each attribute is operationally defined and perceptually separable (see Table 5). 3.1.1 Word level. Quality at the word level is determined by whether each word contains the correct sounds and carries stress on the correct syllable. We organize this level around fixed properties of lexical entries, including both phoneme identity and stress position. Errors at this level (whether phonetic or stress-related) con- stitute word identity errors regardless of sentence context (Jesse et al., 2017). Phonetic accuracy: Whether the sounds in each word fall within the acceptable range of realization for their lexical targets. Lexical stress: Whether primary stress falls on the correct syllable of each word, including appro- priate vowel reduction in unstressed syllables. 3.1.2 Prosodic level. Quality at the prosodic level is determined by whether suprasegmental properties are appropri- ate for the syntactic and semantic properties of the utterance. Prosody is a perceptually separa- ble dimension that listeners isolate from overall speech quality (Gutierrez et al., 2021) and one where current systems remain measurably below natural speech (VallĂ©s-PĂ©rez et al., 2021). We define the dimensions below based on lin- guistic function. Intonation: The pitch contour over phrases and sentences, encoding utterance type (e.g. ques- tion, statement) and discourse structure (e.g. topic boundaries, turn-taking signals). Prosodic stress: Prominence on specific words or phrases relative to the rest of the utterance, de- termined by information structure. Prosodic boundary appropriateness: The chunking of the utterance into prosodic phrases aligned with syntactic structure, realized through boundary tones, lengthening, and pause. Speech rate appropriateness: The overall pace of delivery, which can fail without any correspond- ing failure in pitch or prosodic structure. 3.1.3 Paralinguistic level. Quality at the paralinguistic level is determined by whether the speech conveys speaker characteristics appropriate to the intended speaker and content. Al- though paralinguistic dimensions share audio prop- erties with prosody, such as pitch and timing, they convey information about the speaker rather than the linguistic message. Following Crystal (1969), we distinguish speaker-stable voice quality features from affective features that modulate in response to content and context (see Appendix 6.4). Where no external reference speaker exists, dimensions are assessed relative to the systemâs own output. Emotional appropriateness: The degree to which the emotional tone matches the affective content of the utterance. Expressiveness: Natural variation in delivery style and vocal energy, assessed relative to the sys- tem itself. Human plausibility: The degree to which vocal qualities fall within the range of physically realiz- able human production. Speaker identity consistency: Stability of per- ceived vocal characteristics across the utterance, assessed relative to the system itself. 3.2 Dataset Construction Our dataset consists of 860 samples of annotated and majority-voted audios. We generated synthetic speech using Cartesia Sonic-3 and introduced con- trolled errors through text modification and acous- tic manipulation. Below we describe our source corpora, annotation scale, error generation meth- ods, and human annotation procedure. Table 2: Error generation method, sample counts, and inter-annotator agreement per dimension.nis the balanced evaluation set after majority-class downsampling to a 50/50 positive/negative split (positives=negatives= n/2for every dimension). Krippendorffâs α is reported on this set with 95% bootstrap CIs. Total: n = 860. LevelDimensionMethodText SourceManipulationnKrippendorffâs α Word Phonetic AccuracyLLM-altered IPAHarvard SentencesPhoneme substitution1100.767 [0.660, 0.852] Lexical StressLLM-altered IPAHarvard SentencesStress mark transposition900.821 [0.728, 0.907] Prosodic IntonationAcoustic manipulationEmergentTTS-EvalF0 contour manipulation800.469 [0.332, 0.611] Prosodic StressLLM-altered plain textHarvard SentencesFunction word quoting1020.658 [0.542, 0.766] Prosodic Boundary PlacementLLM-altered plain textSynthetically generatedStructurally targeted sentences880.727 [0.604, 0.834] Speech Rate AppropriatenessAcoustic manipulationHarvard SentencesDuration scaling1280.707 [0.605, 0.801] Paralinguistic Emotional AppropriatenessAPI-enforced emotionSynthetically generatedTextâemotion mismatch via Cartesia tag 460.596 [0.400, 0.767] ExpressivenessAcoustic manipulationHarvard SentencesF0 manipulation, loudness compres- sion, vowel shortening 940.460 [0.332, 0.597] Speaker Identity ConsistencyAcoustic manipulationHarvard SentencesCross-speaker fading660.713 [0.577, 0.836] Human PlausibilityAcoustic manipulationHarvard Sentences Insertedaudiodistortionsand glitches 560.710 [0.565, 0.849] Total860 Source corpora. To systematically cover the annotation schema dimensions, we selected ut- terances from two existing corpora and synthet- ically generated additional texts where needed (Table 2). Harvard Sentences (IEEE Subcom- mittee on Subjective Measurements, 1969) pro- vide 720 phonetically balanced sentences cover- ing the English phoneme inventory broadly, well- suited to word-level and some prosodic dimensions. EmergentTTS-Eval (Manku et al., 2025) provides sentences with complex syntactic structures and varied question types (useful for prosodic bound- ary and intonation dimensions) alongside sentences with explicitly defined affective content, used for emotional appropriateness. For dimensions lacking appropriate utterances in these sources, we synthet- ically generated sample texts (Appendix 7.4). Scale design.All dimensions are rated as binary (0/1), where 0 means any error is present and 1 means the audio is perfect, following the annotation schema. Error generation pipeline.For each source sen- tence, we generated utterances at the intended qual- ity level by introducing targeted errors via one of three pathways (see Figure 1): âąLLM-altered IPA. Used for Phonetic Accuracy and Lexical Stress. An LLM (GPT-5) produces a modified IPA transcription instantiating the in- tended error, passed to Cartesia Sonic-3 (Carte- sia AI, 2025) for IPA-driven synthesis (e.g., phoneme substitution: seed/sid/â /sib/; stress transposition: harvest /"hAr.vIst/â /hAr."vIst/). âąLLM-altered plain text. Used for Prosodic Stress and Prosodic Boundary Placement. For stress, quotation marks around LLM-selected function words induce atypical prominence (e.g., Type out three lists âofâ orders); for bound- aries, commas split LLM-selected syntactic con- stituents (e.g., I dropped my, keys). âąAPI-enforced emotion. Used for Emotional Ap- propriateness. Sentences with a predetermined emotion are synthesized via Cartesia with deliber- ately mismatched emotion tags (e.g., My beloved dog passed away quietly in his sleep tonight. with tag euphoric). âąAcoustic Manipulation. Used for Intonation, Speech Rate, Expressiveness, Speaker Identity Consistency, and Human Plausibility. Praat- based manipulation (Boersma, 2001) via Parsel- mouth (Jadoul et al., 2018) introduces targeted perturbations to a clean synthesis baseline: F0 contour manipulation for intonation; duration scaling for speech rate; combined F0, loudness, and vowel duration modification for expressive- ness; cross-speaker fading for speaker identity; and inserted distortions for human plausibility. We treat these as controlled upper-bound cases: a judge that cannot detect a deliberately introduced perturbation cannot detect subtler naturally oc- curring deviations of the same type. The prompts and details of the acosutic modi- fication are provided in Appendix 7.5 and Ap- pendix 7.2 3.3 Human Annotation Raters.Each sample was annotated by three pro- fessional linguists with experience in analytical work with natural language data. Three raters is the practical minimum for stable Krippendorffâsαesti- mation; sample counts per dimension are reported in Table 2. Pilot and guideline refinement.To reduce guide- line ambiguity (Uma et al., 2022), we conducted a pilot annotation (n = 50samples) with all raters and iteratively refined the guidelines after collec- tive review of disagreements. Procedure.Raters are initially presented with the audio only, but have the option to toggle the tran- script. Every sample is rated on all applicable di- mensions, not only the dimension it was generated to target. This yields cross-dimensional contam- ination data and allows us to quantify perceptual spillover: when a sample was generated to elicit a failure on dimension X, what do raters observe on dimension Y? (See Appendix 2) Inter-rater reliability. Table 2 reports Krippen- dorffâsα(Krippendorff, 2011) on the construct- aligned subset (n = 46-128per dimension), after class balancing, with 95% bootstrap confidence in- tervals (1117 resamples, percentile method, seed 42). 3.4 Dataset Finalization Aggregation and downsampling. Ground truth annotations for each sample, for each dimension, were determined by majority vote. Each intended (analysis) dimension was downsampled to a bal- anced dataset of positive and negative samples, ran- domly sampling from the larger class with prior- ity for consensus annotations. Final dataset statis- tics, including per-dimension sample counts, are reported in Table 2. 4 Experiments 4.1 Judges Evaluated We evaluate two classes of automated judge that differ in output type and evaluation mechanism. MOS predictors.We include UTMOSv2 (Baba et al., 2024), DNSMOS-Pro (Cumlin et al., 2024), NISQA (Mittag et al., 2021), and Audiobox- Aesthetics (Tjandra et al., 2025). Each produces either a single continuous scalar per audio or dimension-wise scores. Audio-LLMs judges.We evaluate a set of open- source and closed audio large language mod- els: Gemini 3.5 Flash , Gemini 3 Flash, Qwen3- Omni-30B-A3B and Step-Audio-R2-Mini, Audio- LLMs judges are prompt-sensitive, motivating the systematic prompting conditions described in Sec- tion 4.2. 4.2 Prompting Strategy Conditions We define four prompting conditions varying in the amount of schema guidance provided and the granularity of the expected output. In all conditions, models produce binary labels matching the ground- truth annotation format. Each judge is queried three times per sample per condition, with the majority label taken as the final prediction. All conditions are run with and without the reference transcript. Condition 1: Underspecified MOS-style prompt. The model is prompted to rate the "naturalness" of the sample without a detailed schema, producing a single holistic score analogous to a neural MOS predictor. This condition establishes what the judge attends to in the absence of detailed task guidance, and serves as the primary baseline. Condition 2a: Schema-guided prompt, single score. The model is prompted with the full 10- dimension schema but instructed to produce a sin- gle overall score. Since this yields one score per sample, we correlate it against human labels inde- pendently for each dimension, treating the same predicted score as the candidate signal for each. This condition reveals whether schema exposure alone is sufficient to direct the judge toward the right perceptual dimensions. Condition 2b: Schema-guided prompt, per-di- mension scores.The model is prompted with the full schema and instructed to produce a score for each dimension, enabling direct per-dimension cor- relation against ground-truth labels. This condition establishes whether models can score all dimen- sions simultaneously when given the a detailed schema. Condition 3: Isolated per-dimension prompts. The model is prompted for each dimension indi- vidually. This condition reveals which dimensions the judge can reliably detect without interference from the full schema. It tests dimension-specific sensitivity in its purest form. Full prompt text for all conditions is provided in Appendix 6.7. 4.3 Metrics For each (rating dimensionĂmodel) pair, we eval- uate how well model scores track binary human ground-truth labels. Per-dimension effect size and significance. KendallâsÏ b serves as a unified comparison metric for per-dimension effect sizes across both model types: for continuous models (MOS predictors) it measures rank-order correlation against binary labels, and for binary models (Audio-LLMs) it is algebraically equivalent to PearsonâsÏ. This equiv- alence enables consistent comparison of continu- ous and binary models within the same evaluation framework. Table 3: Per-dimension KendallâsÏagainst human ground truth per evaluation dimension. AUDIO-LLMS judge rows use the C1 naturalness prompt (single holistic score, closest MOS analogue). ModelWordProsodicParalinguistic Phon.Lex.Inton.Pros.SPros.B RateEmo. Expr.Spk.Hum. MOS Models AudioBox CE -0.0510.1610.371 â -0.241 â -0.0360.550 â 0.050 0.187 â -0.270 â 0.606 â AudioBox CU -0.0520.0150.0980.0680.0480.485 â -0.025 0.418 â -0.1770.511 â AudioBox PC -0.058-0.019-0.0990.337 â 0.1520.020-0.050 0.326 â -0.142-0.615 â AudioBox PQ -0.0880.0380.272 â -0.169 â -0.0790.447 â 0.015 0.304 â -0.379 â 0.606 â DNSMOS-Pro BVCC -0.112-0.0550.0920.1210.179-0.147 â -0.061 -0.072-0.1960.450 â DNSMOS-Pro VCC 0.005-0.1210.026-0.048-0.0380.0710.128 0.406 â -0.0920.397 â NISQA 0.027-0.0470.0680.1420.0990.236 â -0.015 0.347 â -0.0360.377 â UTMOSv20.0000.0870.289 â 0.200 â -0.0240.390 â -0.062 0.0990.0500.313 â AudioLLM Models Gemini 3 Flash â 0.323 â 0.251 â 0.1250.125-0.0810.1230.089 0.274 â 0.0360.202 Gemini 3 Flash âł -0.0220.089-0.225 â 0.099-0.1040.0910.000 -0.0500.0360.121 Gemini 3.5 Flash â 0.412 â 0.312 â 0.0830.227 â 0.276 â 0.1160.255 0.1700.254 â 0.238 Gemini 3.5 Flash âł -0.0910.082-0.0260.1110.2300.0000.045 0.1350.0680.219 Qwen3 Omni â 0.000-0.124-0.1130.0390.0640.214 â 0.105 0.0610.293 â 0.000 Qwen3 Omni âł 0.1730.1920.026-0.1030.2620.300 â 0.088 0.1380.2280.237 Step Audio 2 Mini â -0.0180.022-0.128-0.0800.116-0.096-0.132 0.0650.1290.073 Step Audio 2 Mini âł 0.0190.0230.1000.040-0.111-0.078-0.039 -0.0210.0950.036 Model Subscripts:: CE = Content Enjoyment; CU = Content Usefulness; PC = Production Complexity; PQ = Production Quality. DNSMOS-Pro subscripts: BVCC = trained on BVCC corpus; VCC = trained on VCC2018 corpus. AUDIOLLM rows use the C1 naturalness prompt. Columns: Phon. = Phonetic Accuracy; Lex. = Lexical Stress; Inton. = Intonation; Pros.S = Prosodic Stress; Pros.B = Prosodic Boundary; Rate = Speech Rate; Emo. = Emotional Appropriateness; Expr. = Expressiveness; Spk. = Speaker Identity; Hum. = Human Plausibility.ç Transcript: â = with transcript; âł = without transcript. Bold = significant (p < .05). â p < .05, â p < .01, â p < .001 (two-sided). To test significance, for MOS predictors (contin- uous output), we report Mann-WhitneyU, which tests for distributional separation between score distributions of positive and negative labels. For Audio-LLMs(binary output), we also report Mc- Nemarâs test, which tests for marginal symmetry between model predictions and ground truth. For MOS predictors, we additionally report AU- ROC, which is more naturally interpretable for con- tinuous output and is equivalent to a normalized Mann-Whitney U statistic. AUROC rankings are consistent with effect sizes reported by Kendallâs Ï b across predictors (Spearmanâs Ï = 0.995). 4.4 Main Findings MOS Predictors and AudioLLM Judges show sensitivity to different dimensions.Under a ba- sic naturalness prompt (C1), MOS predictors and Audio-LLMs judges attend to fundamentally differ- ent aspects of speech quality. Neither class aligns reliably with human judgments across the full set of perceptual dimensions. We observe that quality degradations in many dimensions go unnoticed by these models. Every MOS predictor reaches significance on at least one dimension, with their attention con- centrating on paralinguistic dimensions generated via acoustic manipulation, while remaining insen- sitive to word-level and prosodic dimensions (see Table 3). Human Plausibility, constructed via in- serted glitches, is the most consistently detected dimension, reaching significance across all tested models. Speech Rate Appropriateness, constructed via duration scaling, is also significant across 6 of the 8 models. Both patterns are consistent with MOS predictors tracking signal-level degradation, reflecting their origins in telephony-era voice qual- ity assessment where such artifacts were the pri- mary target. Expressiveness also reaches signifi- cance for six of the eight models, though whether this reflects genuine sensitivity to expressive de- livery or a response to shared low-level acoustic properties across the Expressiveness, Speech Rate, and Human Plausibility stimuli remains unclear. In contrast, no MOS models show significant de- tection of Phonetic Accuracy, Lexical Stress, and Prosodic Boundary Placement, dimensions where TTS systems produce linguistically meaningful er- rors that human listeners reliably detect, but MOS predictors do not. Some significant correlations are negative.We observe a misalignment between some predictors and human perceptual judgments. AudioBox PC shows a strong negative relationship with Hu- man Plausibility; negative effects also appear across AudioBox CE , AudioBox PQ and DNS- MOS Pro BVCC on Speaker Identity Consistency, Table 4: Per-dimension KendallâsÏfor AUDIO-LLMS judges using C2a, C2b, and C3 prompts (with and without a provided schema, single or per-dimension outputs). UndefinedÏ(model predicted same value for all samples) indicated by (.) ModelWordProsodicParalinguistic Phon.Lex.Inton. Pros.S Pros.B RateEmo.Expr.Spk.Hum. C2a â Multi-dimension schema; single holistic score Gemini 3 Flash â 0.292 â 0.0370.1970.1410.0640.0000.2640.1420.1770.135 Gemini 3 Flash âł 0.206 â 0.327 â -0.1600.0000.283 â 0.1720.1290.211 â 0.1240.000 Gemini 3.5 Flash â 0.486 â 0.294 â 0.1190.216 â 0.0640.0000.396 â -0.0310.1220.146 Gemini 3.5 Flash âł 0.396 â 0.438 â -0.0520.064-0.1040.1020.0990.1700.296 â 0.115 Qwen3 Omni â 0.0970.1460.052 -0.1010.192..0.1820.435 â 0.000 Qwen3 Omni âł 0.194 â 0.212 â .-0.1410.000...-0.1240.135 Step Audio 2 Mini â 0.0000.186.-0.1010.0000.0000.213-0.1470.413 â 0.135 Step Audio 2 Mini âł 0.0410.0000.160 -0.039-0.111-0.221 â 0.000-0.1650.061-0.153 C2b â Multi-dimension schema; one score per dimension Gemini 3 Flash â 0.397 â 0.106..0.196.0.1490.104.. Gemini 3 Flash âł 0.194 â 0.186.....0.104.. Gemini 3.5 Flash â 0.514 â 0.290 â ..0.1370.1260.309 â 0.211 â 0.1770.135 Gemini 3.5 Flash âł 0.210 â 0.1060.113 .0.1370.0890.1490.1470.2180.135 Qwen3 Omni â .......... Qwen3 Omni âł .......-0.104.. Step Audio 2 Mini â .......... Step Audio 2 Mini âł ..0.095 ...0.149... C3 â Single-dimension focus; one dimension at a time Gemini 3 Flash â 0.366 â 0.0000.083 .0.319 â 0.0890.2130.061.. Gemini 3 Flash âł 0.261 â 0.267 â 0.000 .0.0550.0210.213-0.054.0.135 Gemini 3.5 Flash â 0.355 â 0.178..0.2430.181 â 0.459 â -0.0600.286 â 0.135 Gemini 3.5 Flash âł 0.194 â 0.170-0.1550.1000.276 â 0.232 â 0.303 â 0.0430.218-0.135 Qwen3 Omni â ....0.137-0.184 â .0.026.0.165 Qwen3 Omni âł 0.1360.243 â ..0.298 â 0.050-0.149-0.214 â .0.139 Step Audio 2 Mini â .......0.104.0.000 Step Audio 2 Mini âł .-0.108.0.100...0.069.0.038 Columns: Phon. = Phonetic Accuracy; Lex. = Lexical Stress; Inton. = Intonation; Pros.S = Prosodic Stress; Pros.B = Prosodic Boundary; Rate = Speech Rate; Emo. = Emotional Appropriateness; Expr. = Expressiveness; Spk. = Speaker Identity; Hum. = Human Plausibility. Transcript: â = with transcript; âł = without transcript. Bold = significant (p < .05). â p < .05 , â p < .01, â p < .001 (two-sided). Prosodic Stress and Speech Rate. We hypothesize that these predictors assign higher scores to the controlled acoustic distortions applied to the par- alinguistic dimensions (see Appendix 7.3), rating degraded samples more favorably than human lis- teners do. Generic naturalness prompts yield sparse, in- consistent sensitivity. In contrast to MOS mod- els, Audio-LLMs judges prompted for a generic naturalness rating (C1) show sparse and scattered attention to human-perceptible dimensions (see Ta- ble 3). Gemini 3.5 Flash with transcript shows the broadest coverage of any Audio-LLMs condition, reaching significance on dimensions spanning the word-level, prosodic, and paralinguistic tiers, while every other model and condition reaches at most two or three significant dimensions. Beyond the Gemini family, sensitivity under C1 is sparse to absent, with no dimension reliably detected across more than one model, and the few significant ef- fects are directionally inconsistent. Taken together, the C1 baseline shows selective sensitivity across the Gemini models, while the other models show little evidence of reliable sensitivity to any dimen- sion under an underspecified naturalness prompt. Schema guidance helps selectively.We find that schema guidance recovers word-level sensitivity absent under the unguided naturalness prompt (C1): providing the full 10-dimension schema while maintaining a single aggregated score (C2a) in- creases sensitivity to word-level dimensions for models that showed none under C1, suggesting that this capacity was present but not elicited without schema guidance (see Table 4). The clearest gain appears in the no-transcript condition: both Gemini variants show no signifi- cant word-level correlations under the basic natu- ralness prompt (C1) but achieve statistically sig- nificant correlations when provided the detailed schema (C2a). This supports schema guidance improving model sensitivity on word-level dimen- sions, and underscores how under-specified a bare naturalness prompt is for surfacing sensitivity to which the model evidently has access. However, this benefit does not extend uniformly across all dimensions: some correlations signifi- cant under C1 lose significance under C2a, while new ones emerge. Paralinguistic dimensions re- main largely resistant across all conditions, with significant correlations sparse and inconsistent. Output collapse (where a model predicts the same score for every sample, leaving correlation undefined) is also observed under C2a; for exam- ple, Step Audio 2 Mini assigns the maximum score to every sample regardless of true label. When models are provided the schema and also required to score each dimension independently (C2b), output collapse becomes the dominant fail- ure mode. Two of the four models assign the maxi- mum score across nearly all dimensions and both transcript conditions, leaving correlation undefined in almost all cells. Transcript access does not consistently increase sensitivity to any dimension. Transcript avail- ability produces inconsistent effects across condi- tions and models, with no clean directional pattern. The models that show sensitivity to word-level di- mensions under the unguided naturalness prompt C1 (Gemini 3 Flash and Gemini 3.5 Flash) only do so when the transcript is provided. However, once the schema is provided (C2a), this dependency does not persist; both Gemini variants become signifi- cant on the two word-level dimensions without the transcript, although transcript provision is associ- ated with stronger effect sizes. When models are prompted for each dimension independently (C3), both Gemini variants remain significant on Phonetic Accuracy with stronger ef- fect sizes when the transcript is provided. For paralinguistic dimensions, Gemini 3.5 Flash also reaches significance on Emotional Appropriateness in both transcript conditions. However, the pattern runs the opposite way for paralinguistic dimensions for two other models; Qwen3 Omni and Step Audio 2 Mini show significant Speaker Identity sensitivity under C2a only with the transcript, and show no sensitivity to the other paralinguistic dimensions in either condition. The absence of a consistent transcript benefit suggests that current Audio-LLMs judges are not consistent in their ability to leverage textual ground- ing to improve dimension-specific detection, even for dimensions such as Phonetic Accuracy where the reference transcript is directly informative. Audio-LLMs sensitivity remains inconsistent un- der per-dimension prompting. Under C3 (per- dimension prompting), sensitivity remains sparse and inconsistent, with a pattern that largely dif- fers from C1 and C2a. An exception is Geminiâs increased sensitivity to Phonetic Accuracy under transcript provision, which holds across prompt variations. Negative correlations still exist for some models on isolated dimensions. Step Audio 2 Mini shows no significant correlation under any tran- script condition in C3, consistent with its perfor- mance across all other conditions and suggesting a fundamental sensitivity limitation rather than a prompting artifact. 5 Conclusion We present the first dimension-level meta- evaluation benchmark for TTS, decomposing natu- ralness into 10 linguistically grounded perceptual attributes annotated by trained linguists. Across four MOS predictors and four Audio-LLMs judges, we find that current automated evaluators ex- hibit systematic but distinct blind spots: MOS predictors align strongly with signal-level arti- facts but fail on word-level and prosodic dimen- sions, whereas Audio-LLMs judges show selective, prompt-dependent detection that never generalises across the full dimension set. We further show that how Audio-LLMs judges are prompted in practice substantially affects their measured performance: querying dimensions in isolation consistently improves alignment with hu- man judgments, while asking models to score all dimensions jointly degrades output reliability through prediction collapse. Taken together, our findings suggest that natural- ness cannot be treated as a single scalar construct for modern TTS evaluation and that robust evalua- tion will require dimension-aware benchmarks. Limitations Ecological validity. Due to the difficulty of sourc- ing naturally occurring TTS failures at scale across specific perceptual dimensions, errors were elicited through TTS prompting and Praat manipulation. Praat-manipulated stimuli are thus best understood as controlled upper-bound conditions rather than representative deployment samples. Stimulus complexity ceiling. Harvard Sentences average 7â8 words with SVO structure, limited dis- course context, and few idioms or proper nouns. This is ideal for eliciting clear word-level failures, but limits generalisability to more complex natural- istic speech where prosodic and syntactic interac- tions are richer. Single architecture for IPA-controlled genera- tion. Phoneme-level input is currently only sup- ported by Cartesia Sonic among the evaluated providers; all IPA-driven samples consequently reflect one TTS architecture. We identify multi- system evaluation as a priority for future work. Annotation scale granularity. Binary scales do not capture full fine-grained quality gradients. In contrast, ordinal scales provide granularity but also the potential for response biases and incon- sistent boundary interpretation between levels. In this context, binary scales were a deliberate scope choice; the benchmark targets error detection, not full perceptual grading. Annotation reliability. Majority-class downsam- pling to achieve class balance reduces effective sample sizes, most consequentially for Emotional Appropriateness (n=46) and Lexical Stress (n=90), where agreement estimates carry wider confidence intervals. Ethics Statement Our proposed benchmark is intended solely to sup- port more targeted and interpretable TTS evalua- tion research and is not intended for discrimina- tory applications or commercial ranking of TTS providers. All annotators are first-language English speakers and trained linguists who participated vol- untarily at standard rate compensation. No sensi- tive or personally identifiable data was collected. All of our speech samples are synthetically gen- erated or acoustically manipulated; no real human speech recordings were collected. Source corpora and APIs are used in accordance with their re- spective licenses and terms of service (see Ap- pendix 7.5). Regarding Large Language Model (LLM) usage in manuscript preparation, we utilize them solely to refine the language used in paper to improve clarity and correctness, without generating any substantial content or claims. References Kaito Baba, Wataru Nakata, Yuki Saito, and Hiroshi Saruwatari. 2024. The t05 system for the VoiceMOS Challenge 2024: Transfer learning from deep im- age classifier to naturalness MOS prediction of high- quality synthetic speech. In IEEE Spoken Language Technology Workshop (SLT). Amanda Baughan, Xuezhi Wang, Ariel Liu, Allison Mercurio, Jilin Chen, and Xiao Ma. 2023. A mixed- methods approach to understanding user trust after voice assistant failures. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1â16. Paul Boersma. 2001. Praat, a system for doing pho- netics by computer. Glot International, 5(9/10):341â 345. Cartesia AI. 2025.Sonic-3:Streaming text- to-speech model.https://docs.cartesia.ai/ build-with-cartesia/tts-models/latest. Ac- cessed: 2025. Cheng-Han Chiang, Xiaofei Wang, Chung-Ching Lin, Kevin Lin, Linjie Li, Radu Kopetz, Yao Qian, Zhen- dong Wang, Zhengyuan Yang, Hung-yi Lee, and Li- juan Wang. 2025. Audio-aware large language mod- els as judges for speaking styles. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 467â480, Suzhou, China. Association for Computational Linguistics. David Crystal. 1969. Prosodic Systems and Intonation in English. Cambridge University Press, Cambridge. Fredrik Cumlin, Xinyu Liang, Victor Ungureanu, Chan- dan K. A. Reddy, Christian SchĂŒldt, and Saikat Chat- terjee. 2024. DNSMOS Pro: A Reduced-Size DNN for Probabilistic MOS of Speech. In Interspeech 2024, pages 4818â4822. Diana Geneva, Georgi Shopov, Kostadin Garov, Maria Todorova, Stefan Gerdjikov, and Stoyan Mihov. 2023. Accentor: An Explicit Lexical Stress Model for TTS Systems. In Interspeech 2023, pages 4848â4852. Elijah Gutierrez, Pilar Oplustil-Gallegos, and Cather- ine Lai. 2021. Location, location: Enhancing the evaluation of text-to-speech synthesis using the rapid prosody transcription paradigm. In Proc. SSW 2021, pages 25â30. Kexin Huang, Qian Tu, Liwei Fan, Chenchen Yang, Dong Zhang, Shimin Li, Zhaoye Fei, Qinyuan Cheng, and Xipeng Qiu. 2025.Instructttseval: Bench- marking complex natural-language instruction fol- lowing in text-to-speech systems. arXiv preprint arXiv:2506.16381. Wen Chin Huang, Erica Cooper, Yu Tsao, Hsin-Min Wang, Tomoki Toda, and Junichi Yamagishi. 2022. The voicemos challenge 2022. In Proc. Interspeech 2022, pages 4536â4540. Wen-Chin Huang, Szu-Wei Fu, Erica Cooper, Ryand- himas E Zezario, Tomoki Toda, Hsin-Min Wang, Ju- nichi Yamagishi, and Yu Tsao. 2024. The voicemos challenge 2024: Beyond speech quality prediction. In 2024 IEEE Spoken Language Technology Work- shop (SLT), pages 803â810. IEEE. IEEE Subcommittee on Subjective Measurements. 1969. IEEE recommended practice for speech quality mea- surements. IEEE Std 297-1969. Appendix C: 1965 Revised List of Phonetically Balanced Sentences (Harvard Sentences). Yannick Jadoul, Bill Thompson, and Bart de Boer. 2018. Introducing Parselmouth: A Python interface to Praat. Journal of Phonetics, 71:1â15. Alexandra Jesse, Katja Poellmann, and Ying-Yee Kong. 2017. English listeners use suprasegmental cues to lexical stress early during spoken-word recognition. Journal of Speech, Language, and Hearing Research, 60(1):190â198. Ambika Kirkland, Shivam Mehta, Harm Lameris, Gus- tav Eje Henter, Eva Szekely, and Joakim Gustafson. 2023. Stuck in the MOS pit: A critical analysis of MOS test methodology in TTS evaluation. In 12th ISCA Speech Synthesis Workshop (SSW2023), pages 41â47. Klaus Krippendorff. 2011. Computing Krippendorffâs alpha-reliability. Departmental papers, University of Pennsylvania, Annenberg School for Communica- tion, Philadelphia, PA. Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, and Yohan Jo. 2026.Speakersleuth:Evaluat- ing large audio-language models as judges for multi-turn speaker consistency. arXiv preprint arXiv:2601.04029. Potsawee Manakul, Woody Haosheng Gan, Michael J Ryan, Ali Sartaz Khan, Warit Sirichotedumrong, Ku- nat Pipatanakul, William Barr Held, and Diyi Yang. 2026. Audiojudge: Understanding what works in large audio model based speech evaluation. In Pro- ceedings of the 19th Conference of the European Chapter of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 3644â3663. Ruskin Raj Manku, Yuzhi Tang, Xingjian Shi, Mu Li, and Alexander Smola. 2025. Emergenttts-eval: Eval- uating tts models on complex prosodic, expressive- ness, and linguistic challenges using model-as-a- judge. 38. Gabriel Mittag, Babak Naderi, Assmaa Chehadi, and Sebastian Möller. 2021. Nisqa: A deep cnn-self- attention model for multidimensional speech qual- ity prediction with crowdsourced datasets. arXiv preprint arXiv:2104.09494. MahathirMonjurandShahriarNirjon.2025. Speechqualityllm:Llm-based multimodal as- sessment of speech quality. arXiv preprint arXiv:2512.08238. Brian Patton, Yannis Agiomyrgiannakis, Michael Terry, Kevin Wilson, Rif A. Saurous, and D. Sculley. 2016. Automos: Learning a non-intrusive assessor of naturalness-of-speech. In NIPS 2016 End-to-end Learning for Speech and Audio Processing Work- shop. Chandan KA Reddy, Vishak Gopal, and Ross Cutler. 2021. Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors. In ICASSP 2021-2021 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP), pages 6493â6497. IEEE. Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. 2022. Utmos: Utokyo-sarulab system for voicemos challenge 2022. In Proc. Interspeech 2022, pages 4521â4525. Xu Tan, Tao Qin, Frank Soong, and Tie-Yan Liu. 2021. A survey on neural speech synthesis. arXiv preprint arXiv:2106.15561. Andros Tjandra, Yi-Chiao Wu, Baishan Guo, John Hoff- man, Brian Ellis, Apoorv Vyas, Bowen Shi, Sanyuan Chen, Matt Le, Nick Zacharov, and 1 others. 2025. Meta audiobox aesthetics: Unified automatic qual- ity assessment for speech, music, and sound. arXiv preprint arXiv:2502.05139. Ismail Rasim Ulgen, Zongyang Du, Junchen Lu, Philipp Koehn, and Berrak Sisman. 2026. Objective evalua- tion of prosody and intelligibility in speech synthesis via conditional prediction of discrete tokens. IEEE Open Journal of Signal Processing. Alexandra N. Uma, Tommaso Fornaciari, Dirk Hovy, Silviu Paun, Barbara Plank, and Massimo Poesio. 2022. Learning from disagreement: A survey. J. Artif. Int. Res., 72:1385â1470. IvĂĄn VallĂ©s-PĂ©rez, Julian Roth, Grzegorz Beringer, Roberto Barra-Chicote, and Jasha Droppo. 2021. Im- proving Multi-Speaker TTS Prosody Variance with a Residual Encoder and Normalizing Flows. In Inter- speech 2021, pages 3131â3135. Tianxin Xie, Yan Rong, Pengfei Zhang, Wenwu Wang, and Li Liu. 2025. Towards controllable speech syn- thesis in the era of large language models: A system- atic survey. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Process- ing, pages 764â791. Xueyao Zhang, Chaoren Wang, Huan Liao, Ziniu Li, Yuancheng Wang, Li Wang, Dongya Jia, Yuanzhe Chen, Xiulin Li, Zhuo Chen, and 1 oth- ers. 2025.Speechjudge: Towards human-level judgment for speech naturalness. arXiv preprint arXiv:2511.07931. 6 Appendix Glossary of Terms Linguistic Theory Acoustic.Of or relating to the physical properties of sound as a mechanical wave. In speech sci- ence, acoustic analysis involves measurements of frequency, amplitude, duration, and spectral char- acteristics derived from the audio signal. Lexical.Of or relating to the words or vocabulary of a language, as distinct from its grammatical structure. In speech, lexical properties include word identity, lexical stress assignment, and phono- logical form. Linguistic. Of or relating to language or linguistics. In speech evaluation, pertaining to features con- ventionally encoded in the grammar of a language, including segmental and prosodic structure. Phoneme.The smallest contrastive unit of sound capable of distinguishing meaning. Phonemes are abstract categories realized by phonetically simi- lar variants (allophones) whose distribution is gov- erned by phonological context. Phonetics.The branch of linguistics concerned with the physical and perceptual properties of speech sounds. Phonology.The branch of linguistics concerned with the abstract, rule-governed organization of sound systems. Phonology describes the system- atic patterns governing the distribution and combi- nation of phonemes. Segmental.Pertaining to individual consonant and vowel segments comprising the phonemic inven- tory of a language. Segmental quality encompasses place and manner of articulation, voicing, vowel quality, and allophonic processes such as aspiration and flapping. Semantic.Of or relating to meaning in language. In TTS evaluation, semantic considerations include whether prosodic realizationâparticularly intona- tion and focusâaccurately reflects information structure, such as given/new status and contrastive emphasis. Suprasegmental.Pertaining to phonological fea- tures spanning units larger than a single segment: stress, pitch accent, intonation, rhythm, and bound- ary placement. Also termed prosodic features. Syntactic. Of or relating to the grammatical struc- ture and word order of sentences. Syntactic con- stituent boundaries frequently align with prosodic boundaries, and syntactic relationships influence prominence placement and intonational contour shape. Paralinguistic.Pertaining to communicative cues that accompany linguistic expression without con- stituting part of its lexical, grammatical, or phono- logical structure. Such cues may convey affect, stance, or speaker identity. Although paralinguistic features may be vocal, visual, gestural, or other- wise embodied, this glossary uses the term in the narrower context of TTS evaluation to refer only to vocal cues. Notational f 0 (Fundamental Frequency)The lowest fre- quency component of a voiced speech waveform, corresponding to the rate of vocal fold vibration. f 0 is the primary acoustic correlate of perceived pitch and a key parameter of intonation and stress. IPA (International Phonetic Alphabet) A stan- dardized phonetic notation system maintained by the International Phonetic Association, providing a unique symbol for every sound in any human lan- guage. Used for both broad (phonemic) and narrow (phonetic) transcription. Speech & Language Technology LALM (Large Audio Language Model)A mul- timodal extension of large language models incor- porating audio as an input and/or output modality. LALMs can process and generate spoken language, in many architectures subsuming TTS and ASR within a unified model. LLM (Large Language Model)A neural lan- guage model trained on large-scale text corpora via self-supervised objectives, capable of genera- tion, classification, and reasoning across a wide range of text-based tasks. TTS (Text-to-Speech) A class of speech synthesis technology converting written text to spoken audio. Modern neural TTS systems leverage sequence-to- sequence architectures and neural vocoders trained at scale to produce high-fidelity synthetic speech. MOS (Mean Opinion Score) A perceptual evalu- ation metric in which listeners rate stimulus quality on a scale of 1â5; the MOS is the arithmetic mean of collected ratings. Originally standardized for telephony (ITU-T P.800), MOS and its variants are standard in TTS evaluation. Table 5: The ten-aspect evaluation schema organized by dimension. Full annotation criteria and worked examples are provided in Table 6. LevelDimensionFailure condition Word Phonetic accuracySound outside acceptable range of lexical target Lexical stressStress assigned to wrong syllable Prosodic IntonationPitch contour inappropriate to utterance type or pragmatic mean- ing Prosodic stressPost-lexical prominence misplaced relative to information struc- ture Prosodic boundary placementPhrase boundaries inconsistent with syntactic structure Speech rate appropriatenessTempo inappropriate to context Paralinguistic ExpressivenessAffective range insufficient or excessive for content Emotional appropriatenessEmotional tone inconsistent with content Speaker identity consistencyVoice characteristics vary within the utterance Human plausibilityVocal qualities outside physically realizable range 6.1 Data Scale 6.1.1 Scale Collapse For most ternary dimensions, collapsedαis com- parable to raw ordinalα; for expressiveness, col- lapsedαexceeds rawα, indicating that the mid- dle level adds noise rather than signal. Three-way splits on ternary dimensions (each rater assigning a distinct level) occur only on the 1-versus-2 bound- ary and are irresolvable by majority vote; these samples are excluded from the final dataset for that dimension. Table 6: Annotation rubric for all ten speech quality dimensions. Each dimension is described in full in §3.3. DimensionScoreMeaning Phonetic Accuracy (§6.2.1) 1One or more consonants or vowels are wrong and the error(s) are very noticeable. You might mishear the word entirely. 2One or more consonants or vowels are slightly off, but you can still tell what word was intended without much effort. 3All consonants and vowels sound correct. No perceivable phonetic errors. Lexical Stress (§6.2.2) 0At least one word has stress on the wrong syllable. Even one error earns a 0. 1Every multi-syllable word has stress on the correct syllable with appropriate vowel quality. No errors detected. Intonation (§6.2.3) 1One or more pitch contour errors cause high negative impact. The intonation may confuse the intended sentence type or signal an unintended attitude. 2 One or more errors cause low negative impact. Something about the melody sounds slightly unnatural but does not cause confusion. 3No perceivable errors. The intonation sounds natural and appropri- ate for the sentence and its context. Prosodic Stress (§6.2.4) 0Prominence falls on words that do not warrant it, or words that should be prominent are not. 1The choice of prominent words sounds predictable and appropriate. New or important information is highlighted; given material is de-emphasized. Prosodic Boundary Placement (§6.2.5) 1A boundary is missing where it is needed or present where it should not be, with severe impact. You cannot recover the intended sen- tence structure while listening. 2A boundary is in the wrong position and noticeably hurts the expe- rience, but you can still follow the sentence structure. 3Words are grouped into phrases that align naturally with the sen- tenceâs syntactic and semantic structure. Speech Rate Appropriateness (§6.2.6) 1 Rate is clearly inappropriate: so fast words blur together, or so slow the pace feels labored. 2Rate is slightly too fast, too slow, or too uniform, but you can follow the content without real difficulty. 3 Tempo feels natural and varies appropriately. Complex material gets more time; parentheticals move briskly. Emotional Appropriateness (§6.2.7) 0The emotional tone is wrong for the content. The mismatch is clearly perceivable. 1The emotional tone is a plausible match for the text and context. Expressiveness (§6.2.8) 1Delivery is extremely flat and robotic, or excessively exaggerated, to the point of seriously detracting from the experience. 2 Delivery is slightly too flat or slightly too dramatic for the content, but not very distracting. 3Delivery has natural, appropriate variation; sounds engaged and dynamic without drawing attention to itself. (Continued on next page) (Continued from previous page) DimensionScoreMeaning Speaker Identity Consistency (§6.2.9) 0The voice changes in a way that makes it sound like a different person, either within or across sentences. 1The voice is recognizably consistent throughout. It sounds like a single individual, even with natural variation in expressiveness or loudness. Human Plausibility (§6.2.10) 0The audio contains at least one moment that could not have been pro- duced by a human speaker: robotic or metallic artifacts, impossible pitch jumps, audio glitches, or unnaturally cut-off syllables. 1The entire audio stream sounds like it could have been produced by a human speaker. 6.2 Human Annotations Annotator BackgroundAll annotators are first- language English speakers and trained linguists. Annotators participated voluntarily at the standard rate compensation. The annotation task involved listening to synthesized speech samples and mark- ing perceptual quality dimensions based on the linguistic-verified guidelines; no sensitive or per- sonal data was collected. Annotator Iteration For dimensions that con- tain inherent human subjectivity, we initially tested capturing a severity gradient with a ternary scale (1-3). However, annotators could not reliably dis- tinguish between non-perfect rating levels, so we collapsed these to binary, trading nuance for dataset reliability. Since not all generations reliably exhibit the in- tended failure, we generated an excess of samples and relied on human annotation to establish ground- truth labels (see Table 9). Annotation is conducted through a purpose-built interface presenting the audio, the dimension defi- nition, the revealable transcript, and the response options in a single screen, minimizing cognitive load and reducing interface-induced variance be- tween raters. Annotation Guide The annotation guide pro- vides detailed explanations and rubric ratings on how to rate synthetic (TTS) speech on each metric in the evaluation schema. For every metric, you will find: a plain-language explanation of what the metric captures, what to listen for, a rating scale with clear descriptions of each score, and examples to anchor annotatorsâ judgments. The metrics are grouped into three dimensions: Word-level (are the right sounds and stress patterns produced for each word?), Prosodic (is the speech organized into phrases with appropriate melody, emphasis, and timing?), and Paralinguistic (does the speech convey appropriate emotion, expressive- ness, and speaker characteristics?). The details are provided in Table 6. General Labeling Guideline Listen to each ut- terance at least twice before scoring. On the first pass, get a general impression. On the second pass, focus on whichever metric you are currently rating. Try not to let a strong error in one metric bias your rating in another: a word might be mispronounced (a Phonetic Accuracy issue) while the prosody is perfectly fine, or vice versa. Rate each metric inde- pendently. Word-Level Metrics Word-level metrics ask: does each individual word sound right? This covers whether the TTS system picked the correct sounds (phonemes) and placed stress on the correct sylla- ble within each word. 6.2.1 Phonetic Accuracy What you are listening for. Does every word con- tain the right speech sounds? This includes conso- nants, vowels, and coarticulation. You are checking whether the system produced the correct phonemes for each word, especially for tricky cases like het- eronyms, rare words, proper nouns, and loanwords. Example. âShe will lead the group to the lead mine.â The first lead should rhyme with feed and the second with red. If the system pronounces both the same way, that is a phonetic accuracy error. How disruptive it is determines whether you score 1 or 2. 6.2.2 Lexical Stress What you are listening for. Is the stress on the right syllable within each word? For example, the noun record is stressed on the first syllable (RE-cord) while the verb is stressed on the second (re-CORD). This metric is strictly about within-word stress; do not consider sentence-level emphasis here (see Prosodic Stress, Metric 4). Why binary? Lexical stress errors are relatively unambiguous, and even a single error can substan- tially impair word recognition. Lexical vs. prosodic stress. Lexical stress is a fixed dictionary property (which syllable within a word). Prosodic stress (Metric 4) is a speaker-level choice about which words in a sentence receive prominence. Prosodic Metrics Prosodic metrics ask: is the speech organized and delivered in a way that com- municates the structure and meaning of the sen- tence? Prosodic errors can change what a lis- tener understands even when every individual word sounds correct. Because pitch, duration, loudness, and pauses serve multiple prosodic functions, a single acoustic event can be relevant to more than one metric. Rate each metric for its own question, not the acous- tic signal in isolation. Table 7 summarizes which channels each metric draws on. 6.2.3 Intonation What you are listening for. Intonation is the pitch melody across phrases and sentences. It signals sentence type (questions typically rise; statements Table 7: Acoustic channels for each prosodic metric. MetricPitchDur.Loud.Pause Intonationâ Prosodic Stressâ Boundary Placementâ Speech Rateâ fall), speaker attitude, and discourse structure. This metric concerns pitch only; do not factor in rate or pauses. 6.2.4 Prosodic Stress What you are listening for. Which words in the sentence receive extra prominence? Speakers sig- nal focus via higher pitch, greater loudness, and longer duration. Do not confuse with lexical stress (Metric 2), which concerns syllables within a word. 6.2.5 Prosodic Boundary Placement What you are listening for. Speakers group words into phrases using pitch movements, final-syllable lengthening, and pauses. Misplaced boundaries disrupt parsing even when no alternative meaning is at stake. Focus on whether the grouping aligns with plausible syntactic structure and produces a pleasant listening experience. Example. Boundary errors go in both directions. A spurious boundary (âIt was a|beautiful dayâ) sounds choppy; a missing boundary (âWhen the dog bites the man screamsâ) creates a parsing fail- ure. In both cases the question is whether the boundary made the speech harder or less pleasant to follow. 6.2.6 Speech Rate Appropriateness What you are listening for. Is the overall tempo appropriate, and does it vary naturally? Human speakers slow for complex or emotionally weighty material and speed up through predictable or par- enthetical content. This metric is distinct from Boundary Placement (where phrase breaks occur). Paralinguistic Metrics Paralinguistic metrics ask: beyond the linguistic message, does the speech convey appropriate speaker characteristics? This covers emotion, expressiveness, voice consistency, and whether the audio could have been produced by a human. 6.2.7 Emotional Appropriateness What you are listening for. Does the emotional tone of the voice match the content? A sentence about exciting news should not sound bored, and a condolence message should not sound cheerful. This metric does not ask whether the emotion is strong enough (see Expressiveness, Metric 9); it asks only whether the type of emotion is correct. 6.2.8 Expressiveness What you are listening for. Expressiveness captures the degree of natural variation in delivery. TTS speech can be flat (monotone, robotic) or overdone (theatrical, exaggerated). Both extremes are errors. Rate expressiveness relative to what is appropri- ate for the content: informational material calls for more measured delivery; emotionally charged content warrants greater vocal variation. 6.2.9 Speaker Identity Consistency What you are listening for. Does the voice maintain a consistent perceived identity (stable pitch range, vocal quality, perceived age, perceived gender, and accent) both within a single sentence and across sentences in a set? Inconsistencies can occur across sentences or partway through a single sentence. 6.2.10 Human Plausibility What you are listening for. This metric is inde- pendent of linguistic correctness. The question is whether a human being could have physically produced this audio. It targets digital artifacts: robotic buzzing, clicks, metallic resonance, im- possible pitch jumps, abruptly cut-off syllables, or anything outside the range of human vocal pro- duction. Do not penalize linguistic errors here; a mispronounced word is a Phonetic Accuracy issue but may still sound physically human. Quick Reference Table 8: All metrics at a glance. Dim.MetricScaleSummary Word Phonetic Accuracy1â3Right sounds produced? Lexical Stress0/1Right syllable stressed? Pros. Intonation1â3Pitch melody fits sen- tence? Prosodic Stress0/1Right words empha- sized? Boundary Placement1â3Natural phrase group- ing? Speech Rate1â3Tempo appropriate? Para. Emotional Appr.0/1 Emotion matches con- tent? Expressiveness1â3Deliverydynamic/fit- ting? Speaker Consistency0/1Same person through- out? Human Plausibility0/1Could a human produce this? Remember: Rate each metric independently. A single utterance can score 3 on Phonetic Accuracy but 1 on Intonation, or 0 on Human Plausibility but 1 on Lexical Stress. Do not let a strong impression from one dimension bleed into another. 6.3 Acoustic Overlap in the Prosodic and Paralinguistic Tiers The prosodic dimensions share acoustic channels: f 0 , for instance, carries intonation, prosodic stress, and boundary tones simultaneously. Crucially, however, sharing an acoustic channel does not en- tail these dimensions cannot fail independently. A misplaced prosodic boundary may disrupt the pitch contour without any corresponding failure in ac- cent placement, and a stress error may occur with an otherwise well-formed intonation contour. This independence motivates rating each dimension sep- arately rather than holistically: each stimulus is rated on all applicable dimensions independently, so cases where a single acoustic anomaly produces failures across multiple categories are captured ex- plicitly rather than collapsed into a single quality judgment. The paralinguistic tier presents an analogous case. These four dimensions share acoustic chan- nels with prosody but convey information about the speaker rather than the linguistic message. A syn- thesis artifact may simultaneously affect perceived speaker identity consistency and human plausibil- ity, yet the two dimensions target distinct speaker properties and can fail independently of one an- other. Figure 2 further demonstrates the idea of correlations. 6.4 Paralinguistic Framework and Neural TTS Crystalâs original framework distinguishes voice quality features, which are speaker-stable proper- ties and remain constant across an utterance, from affective paralinguistic features, which modulate in response to content and context. The four par- alinguistic dimensions in our schema map directly onto this distinction: speaker identity consistency and human plausibility reflect speaker-stable prop- erties, while expressiveness and emotional appro- priateness reflect content-driven modulation. Human plausibility warrants particular note. Modern neural TTS can produce voices that are internally consistent and expressively adequate but physically implausible (a failure mode not antici- pated by Crystalâs framework but characteristic of neural synthesis). We treat it as a distinct paralin- guistic dimension rather than subsuming it under overall quality, since it represents a novel category of synthesis failure with no natural analog in human speech perception research. 6.5 Intended versus observed labels. Table 9 reports the agreement between intended quality levels (design variable) and majority-vote annotator labels (ground truth). Of 741 samples intended to contain the targeted failure, annotators majority-labeled 239 (32.3%) as having no fail- ure: approximately 1 in 3 designed errors was not perceived as such. Misses are highest for emo- tional appropriateness (46.2%), intonation (40.0%), and lexical stress (38.4%), and lowest for prosodic boundary placement (22.4%) and phonetic accu- racy (21.2%). The false-alarm rate on intended- clean items is only 3.5%, indicating that raters sys- tematically under-call failures rather than over-call them which is a pattern consistent with the dif- ficulty of eliciting perceptually salient errors on higher-level dimensions. 6.6 AUROC Analysis for MOS Models As demonstrated in Table 10, MOS models ex- hibit striking performance gap across dimensions. Most presented models achieve strong alignment with human judgments on human plausibility (AUC 0.72â 0.92) and speech rate (AUC0.40â 0.89), yet systematically fail on word-level and lower prosodic dimensions. Phonetic accuracy and lexi- cal stress yield near-chance performance across all models (AUC0.41â 0.61), while intonation and prosodic stress show modest, inconsistent correla- tion (AUC 0.33â 0.74). 6.7 Prompt Details Condition 1: Underspecified MOS-style prompt [ROLE] You are a speech quality evaluator assessing text-to-speech audio. [TASK] % if show_text %A synthesized speech sample is attached above, along with the text it was generated from. Evaluate the overall naturalness of the speech. Intended text: "original_text"% else %A synthesized speech sample is attached above. Evaluate its overall naturalness based on the audio alone.% endif % [CRITERIA] Rate the naturalness on a binary scale (0 = failure present, 1 = no failure): 0: One or more issues affecting naturalness are perceivable, whether minor or severe. 1: Completely natural, indistinguishable from human speech. [PROCEDURE] Follow these steps: 1. Attend to the attached audio sample. % if show_text %2. Compare the spoken output to the intended text. 3. Identify any issues that affect naturalness. 4. Decide the most appropriate score.% else %2. Identify any issues that Table 9: Intended versus observed label distributions on targeted samples. Caught = majority vote agrees with intended failure. Missed = majority vote labels intended failure as clean. Tie = no majority among the first 3 valid binary votes (only possible when fewer than 3 raters scored the dimension). Miss rate = missed / intended-fail. Bottom panel: of 376 intended-clean items, 363 (96.5%) were observed clean; false-alarm rate = 3.5%. TierDimensionNInt. failCaughtMissedTieMiss rate Word-level Phonetic accuracy120806317021.2% Lexical stress113734528038.4% Prosodic Intonation120804432440.0% Prosodic stress124795326032.9% Prosodic boundary placement95584513022.4% Speech rate1681289335027.3% Paralinguistic Emotional appropriateness58392118046.2% Expressiveness100705020028.6% Speaker identity consistency104643231148.4% Human plausibility115704919227.1% Total1117741495239732.3% Intended-clean items (n = 376): observed clean 363 (96.5%); false-alarm rate 3.5%; ties 0. Table 10: Construct-aligned AUC for MOS models per evaluation dimension. ModelWordProsodicParalinguistic Phon. Lex.Inton.Pros.SPros.B RateEmo. Expr.Spk.Hum. AudioBox CE 0.460.610.76 â 0.33 â 0.470.89 â 0.530.63 â 0.31 â 0.92 â AudioBox CU 0.460.510.570.550.530.84 â 0.480.79 â 0.380.86 â AudioBox PC 0.460.490.430.74 â 0.610.510.470.73 â 0.400.07 â AudioBox PQ 0.440.530.69 â 0.38 â 0.440.81 â 0.510.71 â 0.23 â 0.92 â DNSMOS-Pro BVCC 0.420.460.560.590.630.40 â 0.460.450.360.82 â DNSMOS-Pro VCC 0.500.410.520.470.470.550.590.79 â 0.440.78 â NISQA0.520.470.550.600.570.67 â 0.490.74 â 0.470.76 â UTMOSv20.500.560.70 â 0.64 â 0.480.77 â 0.460.570.530.72 â AudioBox subscripts: CE = Content Enjoyment; CU = Content Usefulness; PC = Production Complexity; PQ = Production Quality. DNSMOS-Pro subscripts: BVCC = trained on BVCC corpus; VCC = trained on VCC2018 corpus. Column key: Phon. = Phonetic Accuracy; Lex. = Lexical Stress; Inton. = Intonation; Pros.S = Prosodic Stress; Pros.B = Prosodic Boundary; Rate = Speech Rate; Emo. = Emotional Appropriateness; Expr. = Expressiveness; Spk. = Speaker Identity; Hum. = Human Plausibility. Bold = significant (p < .05). â p < .05, â p < .01, â p < .001 (two-sided MannâWhitney U ). affect naturalness. 3. Decide the most appropriate score.% endif % [OUTPUT] Respond in this exact JSON format: "evaluation_steps": "<your step-by-step reasoning>", "score": <0 or 1> Condition 2a: Schema-guided prompt, single [ROLE] You are an expert speech quality evaluator with training in phonetics and prosodic phonology. [TASK] % if show_text %A synthesized speech sample is attached above, along with the text it was generated from. Evaluate the overall quality of the speech, considering all dimensions in the evaluation schema below. Intended text: "original_text"% else %A synthesized speech sample is attached above. Evaluate the overall quality of the speech, considering all dimensions in the evaluation schema below.% endif % [CRITERIA] Consider the following ten dimensions when forming your overall judgment: WORD-LEVEL: - Phonetic Accuracy: Does every word contain the right speech sounds? This includes consonants, vowels, and the way adjacent sounds blend together (coarticulation). Check whether the system produced the correct phonemes for each word, especially for heteronyms, rare words, proper nouns, and loanwords. - Lexical Stress: Is the stress on the right syllable within each word? Every multi-syllable English word has a predictable stress pattern. Incorrect lexical stress often changes vowel quality too. This is strictly about within-word stress. PROSODIC: - Intonation: The melody of speech: the way pitch rises and falls across phrases and sentences. It signals sentence type, speaker attitude, and discourse structure. Concerns pitch only. - Prosodic Stress: Which words in a sentence receive extra prominence via higher pitch, greater loudness, and longer duration, signaling which information is new, important, or contrastive. This is about entire words within a sentence, not syllables within a word. - Prosodic Boundary Placement: Speakers group words into phrases using pitch movements, slight lengthening of the final syllable, and pauses. Misplaced boundaries disrupt the listening experience. Focus on whether grouping aligns with plausible syntactic structure. - Speech Rate Appropriateness: Is the overall speed appropriate, and does it vary in natural ways? Human speakers slow down for complex material and speed up through predictable material. PARALINGUISTIC: - Emotional Appropriateness: Does the emotional tone of the voice match the content of the text? This does not ask whether the emotion is strong enough (that is Expressiveness). It asks only whether the type of emotion is correct. - Expressiveness: Natural variation in delivery style, pace, and vocal energy. Human speech is dynamic. TTS speech can be flat or overdone. Rate relative to what would be appropriate for the content. - Speaker Identity Consistency: Does the voice maintain a consistent perceived identity in terms of pitch range, vocal quality, perceived age, gender, and accent? Check for shifts both within and across sentences. - Human Plausibility: Could a human have physically produced this audio? Targets digital artifacts and glitches: robotic buzzing, clicks, unnatural metallic resonance, impossible pitch jumps, abrupt syllable cutoffs. Independent of linguistic correctness. Considering ALL of the above, rate overall quality on a binary scale (0 = failure present, 1 = no failure): 0: One or more issues detected across the ten dimensions, whether minor or severe. 1: No perceivable issues detected across all ten dimensions. [PROCEDURE] Follow these steps: 1. Attend to the attached audio sample. 2. For each of the three dimension groups (word-level, prosodic, paralinguistic), note whether any issues are present. 3. Identify which specific dimensions (if any) show failures. 4. Decide the most appropriate score. [OUTPUT] Respond in this exact JSON format: "word_level_issues": "<any word-level issues noted>", "prosodic_issues": "<any prosodic issues noted>", "paralinguistic_issues": "<any paralinguistic issues noted>", "score": <0 or 1> Condition 2b: Schema-guided prompt, per-dimension scores [ROLE] You are an expert speech quality evaluator with training in phonetics and prosodic phonology. [TASK] % if show_text %A synthesized speech sample is attached above, along with the text it was generated from. Evaluate the speech independently on each of the ten dimensions below, using the specified scale for each. Intended text: "original_text"% else %A synthesized speech sample is attached above. Evaluate the speech independently on each of the ten dimensions below, using the specified scale for each.% endif % [CRITERIA] Rate each metric independently. Do not let a strong impression from one dimension bleed into another. ALL DIMENSIONS use a binary scale (0 = failure present, 1 = no failure): Lexical Stress: Is the stress on the right syllable within each word? This is strictly about within-word stress, not sentence-level emphasis. 0: At least one word has stress on the wrong syllable. 1: Every multi-syllable word has stress on the correct syllable. Prosodic Stress: Which words in the sentence receive extra prominence via higher pitch, greater loudness, and longer duration? 0: Prominence falls on words that do not warrant it, or words that should be prominent are not. 1: The choice of which words receive prominence sounds predictable and appropriate. Emotional Appropriateness: Does the emotion of the voice match the content of the text? If the text is neutral, the emotion should be neutral as well. 0: The emotion is wrong for the content. 1: The emotion is a plausible match for the text. Speaker Identity Consistency: Does the voice maintain a consistent perceived identity? 0: The voice changes in a way that makes it sound like a different person. 1: The voice is recognizably consistent throughout. Human Plausibility: Could a human have physically produced this audio? 0: The audio contains at least one moment that could not have been produced by a human speaker. 1: The entire audio stream sounds like it could have been produced by a human speaker. Phonetic Accuracy: 0: One or more consonants or vowels are outside the range of possible expected sounds for the utterance spoken in general american english. 1: All consonants and vowels fall within the range of possible expected sounds for the utterance spoken in general american english. Intonation: 0: Perceivable errors in the pitch contour, regardless of severity. 1: No perceivable errors in the pitch contour. Prosodic Boundary Placement: 0: Misplaced boundary disrupts sentence structure, noticeably or severely. 1: Phrase grouping aligns naturally with syntactic structure. Speech Rate Appropriateness: 0: Speech rate is too slow or too fast during parts or the whole of the utterance. 1: Speech rate feels natural and varies appropriately. Expressiveness: 0: Delivery is too flat or too dramatic, regardless of severity. 1: Delivery has natural, appropriate variation. [PROCEDURE] For each dimension: 1. Attend to the attached audio sample focusing on that dimension. 2. Compare what you hear to the rubric anchors above. 3. Record your reasoning and score. Evaluate each dimension independently. [OUTPUT] Respond in this exact JSON format: "phonetic_accuracy":"reasoning": "...", "score": <0 or 1>, "lexical_stress":"reasoning": "...", "score": <0 or 1>, "intonation": "reasoning": "...", "score": <0 or 1>, "prosodic_stress":"reasoning": "...", "score": <0 or 1>, "prosodic_boundary":"reasoning": "...", "score": <0 or 1>, "speech_rate": "reasoning": "...", "score": <0 or 1>, "emotional_appropriateness": "reasoning": "...", " score": <0 or 1>, "expressiveness":"reasoning": "...", "score": <0 or 1>, "speaker_identity":"reasoning": "...", "score": <0 or 1>, "human_plausibility":"reasoning": "...", "score": <0 or 1> Condition 3: Isolated per-dimension prompts [ROLE] You are an expert speech quality evaluator with training in phonetics and prosodic phonology. [TASK] % if show_text %A synthesized speech sample is attached above, along with the text it was generated from. Evaluate the speech on ONE specific dimension. Intended text: "original_text"% else %A synthesized speech sample is attached above. Evaluate the speech on ONE specific dimension based on the audio alone.% endif % [CRITERIA] dimension_block [PROCEDURE] Follow these steps: 1. Attend to the attached audio sample. 2. Focus exclusively on dimension_name. 3. Compare what you hear to the score anchors above. 4. Select your score. Do not consider any other aspect of speech quality. Evaluate ONLY dimension_name. [OUTPUT] Respond in this exact JSON format: "reasoning": "<your reasoning about dimension_name>", "score": <score_range> 7 Generation Pipeline Details This appendix documents the model configura- tions and hyperparameters for the error genera- tion pipeline described in Section 3.2. All audio: 44,100 Hz, 16-bit PCM, mono. All dimensions use binary labels: 1 (acceptable) and 0 (error). For dimensions originally generated with three severity levels, scores 1 and 2 were collapsed to label 0; score 3 maps to label 1. 7.1 Shared Defaults LLM. Azure OpenAIgpt-5;api_version 2025-03-01-preview;max_completion_tokens = 16384; temperature: API default. TTS (default).Cartesiasonic-3(Blake: a167e0f3-df7e-4d52-a9c3-f949145efdab, Brooke: e07c00bc-4134-4eae-9ea4-1a55fb45746b) TTS(intonation).ElevenLabs eleven_multilingual_v2. 7.2 Pathway 1: IPA-Controlled TTS gpt-5produces a modified transcript for each in- tended error level, passed to Cartesia Sonic-3. Ta- ble 11 lists voices and prompts are listed in 7.5; Table 12 lists intended errors. Table 11: IPA pathway: voice per dimension. All use Cartesia Sonic-3. Dim.Voice Phon. Acc.Brooke Lex. Str.Blake Pros. Str.Blake Table 12: IPA pathway: intended error per label. Dim.LabelIntended error Phon.1Canonical IPA Acc.0 Phonemesubstitution (within- or cross-class) Lex.1Canonical stress Str.0Shifted stress + vowel reduc- tion Pros.1Emphasis on content word Str.0Emphasis on function word (Cartesia quote-markup) 7.3 Pathway 2: Acoustic Manipulation A clean TTS baseline is synthesised first; Parsel- mouth/Praat PSOLA or local DSP then introduces a targeted perturbation. Intonation. Source: EmergentTTS-Questions split (declaratives and questions; 4â12 words; no commas/dashes).ElevenLabs voice Brian (nPczCjzI2devNBz1zQrb) serves as the label-1 an- chor; label-0 samples are derived via F0 manipula- tion with the Praat parameters in Table 13. Label 1: verbatim ElevenLabs synthesis. Label 0: F0 contour manipulation ; either terminal-rise compression (lastâŒ30% of contour compressed to 15â30% of original excursion) or one of 10 high- impact strategies (Table 14), randomly selected. Table 13: Praat PSOLA parameters (Intonation and Speaker Identity). ParameterValue Analysis floor75 Hz Manipulation floor50 Hz Pitch ceiling600 Hz Output clip range60â320 Hz Control points45 Time step0.01 s Two strategies may be stacked with probability âstack-rate P (default 0.0). Table 14: F0 strategies for Intonation label 0. All oper- ate within mean± 70â130 Hz. StrategyParametersKept terminal-rise comp.LastâŒ30% of contour com- pressed to 15â30% of orig- inal excursion â reverseReverse contour; scale 1.2â 1.5Ă;±25 Hz noise â invertInvert around mean; scale 1.3â1.6Ă;±25 Hz noise â questionLinear rise: meanâ40â mean+130 Hz â exaggerateRangeĂ3.5â5.0; clip at ceiling â staircase3â5 levels; mean±70 Hzâ random jumps âŒ50% frames random; mean±90 Hz â flatten+spike 0.7Ămeanbase;3â5 bumps 0.5â1.6Ă â sine wave3â5 Hz wobble;±40â 65 Hz â monotone drift130â180 Hz drift;±30 Hz wobble â random walk Cumulative; Ï â35 Hz/step â Speech Rate. Voice: default English. Direction (fast/slow) randomly assigned. Label 1: Cartesia speed=1.0. Label 0: either Cartesiaspeed=1.4 (fast) or0.6(slow), optionally further stretched vialibrosatime-stretch â factor 0.7 (fast; net â2.0Ă) or 1.3 (slow; netâ0.78Ă). Speaker Identity Consistency.Each sample is a pair: label 0 (inconsistent) and label 1 (consistent). Label 0 uses either (a) equal-power crossfade of two renderings by different voices across a target vowel (â„50 ms), or (b) stitch of two utterances with a 400 ms gap. Voice B undergoes PSOLA pitch alignment when F0 mismatch exceeds 0.5 semitones; pairs with>3 semitones are rejected. Voice pairs: BrookeâEmma (feminine), Ronald â Henry (masculine), all Cartesia Sonic-3 SSE. Human Plausibility.gpt-5selects 1â2 artifact types from Table 15 (always from distinct cate- gories); local DSP applies the recipe to clean Carte- sia synthesis (default English voice). Label 1: un- modified; label 0: post-processed. Table 15: DSP artifact taxonomy (24 types, 8 cate- gories). CategoryTypes (â= used in kept dataset) Compressionmp3 compressionâ, quantiza- tion noiseâ Clippinghard clippingâ, soft clipping â, bit crushingâ Glitchbuffer dropout, stutter repeat , granular fragmentationâ Noisewhite noiseâ, pink noise, vinyl crackleâ, EM hum 50 Hzâ, EM hum 60 Hzâ Time/Pitchtime-stretch artifactâ, pitch- correction artifactâ, phase cancellationâ Digital bit errorsâ, sample-rate mis- match, buffer underrunâ Delayecho feedbackâ, resonant comb filterâ, feedback squeal â SpectralFFT smearingâ, phase distor- tion, aliasingâ 7.4 Pathway 3: Synthetically Generated Text Prosodic Boundary Placement. We created 60 manually curated sentences from Claude Opus 4.8, stratified by error severity and type. The set con- tains 20 sentences with severe errors (Score: 0), 20 with moderate errors (Score: 0), and 20 with correct boundaries (Score: 1). Within each tier, 10 sentences contain boundary insertion errors and 10 contain boundary deletion errors. The severe-deletion subset uses garden-path sen- tences (e.g., âthe old man the boatâ) where a nec- essary boundary deletion forces readers to reparse syntactic structure. This is a failure mode where both human readers and modern TTS systems con- sistently underperform. Severe-insertion sentences feature mid-constituent commas that create infelici- tous boundaries (e.g., âThe quick brown, fox jumps over the lazy dogâ). Moderate-deletion errors are produced by delet- ing spaces around coordinating conjunctions, cre- ating rushed but recoverable runs (e.g., âThe rain stoppedso we went for a walkâ). Moderate- insertion errors place commas between major con- stituents that listeners can still regroup (e.g., âHe bought, a new car last weekendâ). TTS: Cartesia Sonic-3; no LLM phase. Emotional Appropriateness.38 items use Carte- sia Sonic-3 with a mismatched inline emotion tag (e.g.euphoricon sad text) for label 0. 20 addi- tional items use ElevenLabs voice Gerald (Emotion- less, Flat and Clear) for flat delivery of emotion- ally loaded sentences (label 0). Label 1: matched emotion or natural delivery. 7.5 LLM Prompts Prompts are reproduced exactly as submitted to gpt-5. Per-utterance slots shown as placeholders. Phonetic Accuracy Prompt â Phonetic Accuracy Task. You are a phonetic transcription assistant. Given a sentence and a target word, produce three versions of the sentence in which the target word is replaced by three distinct IPA transcriptions in Cartesia sonic-3 phoneme notation, representing decreasing levels of segmental accuracy (score 1 = most erroneous; score 3 = canonical). Phoneme Format. Strings are enclosed in << . . . >> with pipe-separated symbols; stress markers precede the stressed syllable. Example: Cartesia â <<k|ĂĆ|ĂĆș|ĂÄč|t|i|ĂĆ |ĂĆč>> Ph. Ex. Ph. Ex. aI myA cot aU nowO law b bedOI boy d do@ sofa dZ jeansE bed eI day3r bird f fish g go h hatI bit i seer red j yesS ship k keyU book l lipZ vision m man" primary stress n no, secondary stress oU goT thin p penD this s sip"ĂŠ cat t topN sing tS church u two v van w we z zoo Input. Sentence: INPUT_UTTERANCE Target word: TARGET_WORD Output Format. Return only valid JSON â no explanation, markdown, or code fences. "target_word": "<original word>", "versions": [ "score": 1, "error_description": "<substitution + why highly salient>", "transcription": "<<IPA-with-error>>", "sentence": "<sentence with inline transcription>" , "score": 2, "error_description": "<substitution + why moderate >", "transcription": "<<IPA-with-error>>", "sentence": "<sentence with inline transcription>" , "score": 3, "error_description": "Canonical American English pronunciation.", "transcription": "<<canonical-IPA>>", "sentence": "<sentence with inline transcription>" ] Constraints. 1. Scores 1 and 2 must not produce a real English word. 2. Score 1: target the primary stressed syllable; use large vowel-class or manner shifts. 3. Score 2: clearly audible but less dramatic â consonant voicing/manner change, or a one-step vowel shift in a secondary-stressed syllable. Do not swap schwas in fully reduced syllables (inaudible in synthesis). 4. Do not use word-final /z/â /s/ swaps â the contrast is undetectable in synthesis. 5. Use only phonemes from the inventory above. Do not introduce ĂĆœ, ĂĆ, ĂÂż, ĂĆ€, length marks, or diacritics absent from the table; write ĂĆœ as Ăİ (stressed) or ĂĆč|ĂĆș (unstressed). 6. Score 3 is the unmodified canonical American English transcription. Lexical Stress Prompt â Lexical Stress Task. You are a phonetic transcription assistant. Given a sentence and a target word, produce two versions of the sentence in which the target word is replaced by a Cartesia sonic-3 phoneme string. The two versions differ only in lexical stress placement. When stress shifts to a different syllable, apply the phonological changes that naturally follow in American English: vowel reduction (newly unstressed full vowels reduce, e.g. /ĂĆ»/â /ĂĆč/); vowel restoration (reduced /ĂĆč/ in newly stressed syllables becomes its full counterpart); and any other stress-conditioned segmental alternations. Phoneme format and inventory: identical to Phonetic Accuracy above (see §7.2). Input. Sentence: INPUT_UTTERANCE Target word: TARGET_WORD Output Format. Return only valid JSON â no explanation, markdown, or code fences. "target_word": "<original word>", "versions": [ "score": 0, "error_description": "<which syllable incorrectly stressed, where it should be, and phonological changes applied>", "transcription": "<<IPA-with-wrong-stress>>", "sentence": "<sentence with inline transcription>" , "score": 1, "error_description": "Canonical American English pronunciation with correct lexical stress placement.", "transcription": "<<canonical-IPA>>", "sentence": "<sentence with inline transcription>" ] Constraints. 1. Target word must be polysyllabic. If monosyllabic, return an empty versions array. 2. Score 0: move primary stress to a different syllable. If a secondary stress exists, swapping primary and secondary is preferred. Apply all resulting phonological changes. 3. Phoneme changes in score 0 must be a direct consequence of the stress shift â no arbitrary substitutions. 4. Use only phonemes from the inventory (same constraint as Phonetic Accuracy). Prosodic Stress Prompt â Prosodic Stress Task. You are a linguistic annotation tool. Given an English sentence, identify function words (words that should not carry prosodic stress under a neutral broad-focus reading) and output two versions: (1) score 1 â original sentence, unchanged (natural prosody baseline); (2) score 0 â every contiguous run of function words is wrapped in typographic double quotes " . . . ", forcing a pitch accent on the quoted span and producing audibly misplaced prosodic stress. Function-word categories. âą Determiners: a, an, the, this, that, these, those âą Prepositions: in, on, at, to, from, of, for, by, with, about, into, through, over, under, between, among, during, before, after, against, without âą Auxiliaries: is, am, are, was, were, be, been, being, has, have, had, do, does, did, will, would, shall, should, can, could, may, might, must âą Pronouns: I, me, you, he, him, she, her, it, we, us, they, them, my, your, his, its, our, their âą Conjunctions: and, but, or, nor, so, yet, because, although, if, when, while, that, which, who âą Complementizers / relativizers: that, which, who, whom âą Particles: to (infinitival) Ambiguous cases: âthatâ as demonstrative â function word; âthatâ in contrastive focus â content word. Wh-words in questions â content words; as relativizers â function words. Contractions â leave unquoted. Rules. 1. Preserve original spelling, capitalization, and punctuation exactly. Sentence-final punctuation stays outside any quoted span. 2. A single function word gets its own quote pair; two or more adjacent function words share one pair. 3. Do not add, remove, or reorder words. Input. Sentence: INPUT_UTTERANCE Output Format. Return only valid JSON â no explanation, markdown, or code fences. target_word is always the literal string "stress" (used for folder naming). "target_word": "stress", "versions": [ "score": 1, "error_description": "Natural prosody -- content words stressed, function words reduced.", "transcription": "", "sentence": "<original sentence unchanged>" , "score": 0, "error_description": "Stress placed on \"<region>\" (<pos>), \"<region>\" (<pos>).", "transcription": "<region1>; <region2>; ...", "sentence": "<sentence with function-word runs in \"...\">" ] error_description for score 0: begin with Stress placed on, then list each quoted region as "<region>" (<pos>), with multi-word POS joined by â+â, e.g. "on the" (preposition + determiner). Example. Input: The cat sat on the mat. "target_word": "stress", "versions": [ "score": 1, "error_description": "Natural prosody -- content words stressed, function words reduced.", "transcription": "", "sentence": "The cat sat on the mat." , "score": 0, "error_description": "Stress placed on \"The\" ( determiner), \"on the\" (preposition + determiner).", "transcription": "The; on the", "sentence": " 201CThe 201D cat sat 201Con the\ u201D mat." ] 01020 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 02040 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 0102030 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 010203040 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 010203040 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 0204060 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 010203040 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 02040 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 0102030 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility 0204060 Phonetic Accuracy Lexical Stress Intonation Prosodic Stress Prosodic Boundary Placement Speech Rate Appropriateness Emotional Appropriateness Expressiveness Speaker Identity Consistency Human Plausibility Intended dimension score score = 0 score = 1 Failures on each rating dimension, split by intended-dimension ground truth score emotional_appropriatenessexpressiveness human_plausibilityintonation lexical_stressphonetic_accuracy prosodic_boundaryprosodic_stress speaker_identity_consistencyspeech_rate Figure 2: Co-occurrence of perceptual failures across speech quality dimensions. Per-dimension failure counts stratified by intended-dimension ground-truth label. Each subplot corresponds to one intended dimension (column panels); bars show the number of score-0 (failure) observations on each rating dimension (y-axis), split by whether the sampleâs ground-truth label on its own intended dimension is a failure (orange, score = 0) or a pass (blue, score = 1). Tall blue bars on dimensions other than the intended one reveal that samples annotated as passing on their intended construct nonetheless exhibit failures on unrelated dimensions. Human Plausibility Prompt â Human Plausibility Task. You are an audio signal processing expert and evaluation dataset designer. Given a sentence and a target word, produce two versions of the same utterance at controlled quality levels. Both versions use the same clean spoken sentence â the sentence field is identical across both. The difference is a post-processing artifact recipe applied after synthesis. Score 1 is an unmodified clean reference; score 0 has 1â2 digital artifacts applied, making it clearly sound artificial or degraded. Artifact Types. Use these exact identifiers in the type field. Compression: mp3_compression (lossy codec smearing, muffled quality); quantization_noise (low-bit-depth noise/distortion). Clipping & Distortion: hard_clipping (peaks chopped flat; harsh); soft_clipping (warmer harmonic distortion); bit_crushing (lo-fi crunch). Glitch / Stutter: buffer_dropout (sudden silence or frozen chunk); stutter_repeat (rapid looping of a tiny slice); granular_fragmentation (micro-segments scrambled). Noise: white_noise (broadband injection); pink_noise (spectrally weighted); vinyl_crackle (crackle and hiss); em_hum_50hz (50 Hz EM hum); em_hum_60hz (60 Hz EM hum). Time & Pitch: time_stretch_artifact (metallic smearing); pitch_correction_artifact (robotic, over-quantised tuning); phase_cancellation (comb-filtering from offset copies). Digital Errors: bit_errors (random bit flips; sharp pops); sample_rate_mismatch (aliasing or chirping); buffer_underrun (choppy, skipping playback). Delay-Based: echo_feedback (feedback runaway); resonant_comb_filter (resonant comb); feedback_squeal (feedback loop squeal). Spectral: fft_smearing (FFT smearing from over-processed EQ); phase_distortion (heavy-filter phase distortion); aliasing (sample-rate-reduction aliasing). Location Values. target_word : target wordâs duration only; sentence_start : first âŒ20%; sentence_middle : middle âŒ60%; sentence_end : last âŒ20%; throughout : full signal. Severity Values. mild : subtle, attentive listening needed; moderate : clearly audible, intelligibility intact; severe : highly noticeable, degrades intelligibility. Input. Sentence: INPUT_UTTERANCE Target word: TARGET_WORD Artifact assignment for score 0: ARTIFACT_SELECTION Output Format. Return only valid JSON â no explanation, markdown, or code fences. "target_word": "<original target word>", "versions": [ "score": 1, "sentence": "<original sentence, unchanged>", "error_description": "Clean reference -- no artifacts applied.", "artifacts": [] , "score": 0, "sentence": "<original sentence, unchanged>", "error_description": "<listener-perspective description>", "artifacts": [ "type": "<identifier from list above>", "location": "<value from list above>", "severity": "<mild | moderate | severe>" ] ] Constraints. 1. sentence must be identical across both versions. 2. Score 1 must have "artifacts": []. 3. Score 0 must use exactly the specified artifact type(s) â no substitutions or additions. 4. At least one score 0 artifact must be moderate or severe; mild-only is insufficient. 5. Use location: target_word for localised glitch artifacts (buffer_dropout, stutter_repeat, bit_errors, granular_fragmentation). Use throughout for diffuse artifacts (noise, compression, hum). 6. Write error_description from a listenerâs perspective, e.g. âA sharp digital stutter interrupts the stressed vowel of âeclipseâ, looping a tiny slice twice before the sentence continues.â 7. Match location and severity to the target wordâs phonetics: stutter/dropout are most salient on vowel-heavy or sonorant words; clipping/bit-crushing on fricatives and plosives; noise/compression are phoneme-agnostic. 7.6 Artifacts Licenses We provide documentations of the artifact licenses that were used to assist with our data generation and annotation pipeline together with detailed analyses in Table 16, 17, 18. Table 16: TTS APIs and source datasets used in this work. ArtifactVersion / ModelLicense TTS APIs / Services Cartesia TTS APIcartesia SDKCommercial ElevenLabs TTS APIeleven_multilingual_v2Commercial Azure OpenAI / GPT-5openai SDKCommercial Source Datasets Harvard SentencesIEEE Std 269IEEE EmergentTTS-Eval bosonai/EmergentTTS- Eval Apache 2.0 Table 17: Signal processing and audio tools used in this work. ArtifactVersion / NotesLicense Signal Processing / Audio Tools Praat (praat-parselmouth)parselmouthGPL v3 Montreal Forced Alignerenglish_us_arpaMIT SciPysignal processingBSD3- Clause SoundFile / libsndfileaudio I/OBSD3- Clause 7.7 Model Details MOS Predictors NISQA, DNSMOS-Pro, UT- MOSv2, and Audiobox-Aesthetics were each served via a custom FastAPI inference endpoint on a single NVIDIA Tesla P100-PCIE-12GB GPU paired with an Intel Xeon Gold 6126 CPU @ 2.60 GHz. Audio-LLMs We evaluate four Audio-LLMs as automatic judges. Provider, endpoint identifier, in- put modality, and parameter count are summarized in Table 19. Where the vendor has not published an official parameter count, we report it as not disclosed. Table 19 summarises hardware require- ments for each LALM judge. The two Gemini mod- els are accessed via the Google API. Qwen3-Omni (30B total /â3B active, MoE) requiresâ60 GB VRAM atbfloat16and is served on 2Ă80 GB GPUs. Step-Audio-2-mini (8B) requiresâ16 GB VRAM but is served on 1Ă80 GB GPU to accom- modate its custom vLLM backend and audio deto- kenizer (tensor-parallel-size = 2). All open- weight models were served via vLLM on NVIDIA A100/H100 GPUs. Each judge was queried three times per sample per condition; with860samples Ă4 conditionsĂ2 transcript settings, this yields up to 20,640 inference calls per model. Table 18: Python packages used in the pipeline and analysis stages of this work. PackagePurposeLicense Pipeline Packages openaiAzure OpenAI / GPT-5 API clientMIT cartesiaCartesia TTS API clientMIT elevenlabsElevenLabs TTS API clientMIT numpyNumerical computationBSD 3-Clause pandasTabular data manipulationBSD 3-Clause datasetsHuggingFace dataset loadingApache 2.0 huggingface_hubHuggingFace model/data accessApache 2.0 openpyxlExcel file I/OMIT python-dotenvEnvironment variable managementBSD 3-Clause ruamel-yamlYAML configuration parsingMIT soundfileAudio file I/OBSD 3-Clause Analysis Packages krippendorffInter-annotator agreementMIT scikit-learnClassification and evaluation metricsBSD 3-Clause statsmodelsStatistical modelingBSD 3-Clause pingouinStatistical testsBSD 3-Clause crowd-kitCrowdsourcing aggregationApache 2.0 dominance-analysisRelative importance analysisMIT plotlyInteractive visualizationMIT matplotlibStatic visualizationPSF / BSD seabornStatistical visualizationBSD 3-Clause streamlitInteractive data exploration UIApache 2.0 kaleidoStatic image export for plotlyMIT Table 19: Audio-LLM judges and compute requirements. T = text, I = image, A = audio, V = video. Full model name for row 3: Qwen3-Omni-30B-A3B-Instruct. GPU counts reflect our serving configuration; open-weight models served via vLLM on NVIDIA A100/H100 GPUs. # JudgeProvider ModalityParametersVRAM GPUsAccess 1 Gemini 3 FlashGoogleT+I+A+V not disclosed- -API 2 Gemini 3.5 FlashGoogleT+I+A+V not disclosed- -API 3 Qwen3-OmniAlibaba T+I+A+V 30B/â3B (MoE) â60 GB 2Ă80 GB open weights 4 Step-Audio-2-mini StepFun T+A8Bâ16 GB 1Ă80 GB open weights