Paper deep dive
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses
Guangzheng Hu, Ziyue Jiang, Weixu Qiao, Lixin Zhang, Jianye Kang, Yuru Wu, Rong Bao, Niantong Li, Wei Wang, Ziyi Cheng, Xinfa Zhu, HangRui Hu, Ting He, Bing Zhao, Lin Qu, Hu Wei, Jin Xu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/29/2026, 4:06:27 AM
Summary
The paper introduces D3-Omni, a balanced and decoupled benchmark for diagnosing multimodal 'OmniJudge' models across text-to-image (T2I), text-to-video (T2V), and text-to-speech (TTS) generation. It addresses biases in existing benchmarks by creating a dataset with 10,671 samples and 53 orthogonal binary dimensions, ensuring dual balance (score-level and dimension-level) and decoupling of error modes. The study reveals that strong OmniJudges often suffer from Yes-bias, struggle with modality-specific dimensions, and conflate distinct attributes, highlighting the need for balanced evaluation to expose systematic blind spots.
Entities (10)
Relation Signals (7)
D3-Omni â coverstask â Text-to-Video (T2V)
confidence 98% · covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks.
D3-Omni â coverstask â Text-to-Speech (TTS)
confidence 98% · covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks.
D3-Omni â coverstask â Text-to-Image (T2I)
confidence 98% · covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks.
D3-Omni â usesmethodology â D3 Framework
confidence 95% · We propose the D3 frameworkâDual-balanced, Decoupled and Dynamic
Qwen3.5-Omni-Plus â evaluatedon â D3-Omni
confidence 92% · Figure 1: Top: per-segment accuracy over the full ground-truth score range... Qwen3.5-Omni-Plus (74.6%)
Gemini 3.1 Pro â evaluatedon â D3-Omni
confidence 92% · Figure 1: Top: per-segment accuracy over the full ground-truth score range... Gemini-3.1-Pro (78.6%)
OmniJudge â exhibitsbias â Yes-bias
confidence 90% · even strong OmniJudges tend to ... confirm satisfied requirements far more reliably than they detect violated ones
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multimodal understanding models that can jointly judge text-to-image (T2I), text-to-video (T2V) and text-to-speech (TTS) generation are increasingly used as "OmniJudges" for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D3-Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering 53 orthogonal binary dimensions (17/22/14) and 10,671 samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak information across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models this http URL suite reaches near 1:1 per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide systematic blind spots that a balanced and decoupled lens can help expose and, in turn, address.
Tags
Links
- Source: https://arxiv.org/abs/2608.24160v1
- Canonical: https://arxiv.org/abs/2608.24160v1
Trouble viewing inline? Open PDF directly â
Full Text
133,608 characters extracted from source content.
Expand or collapse full text
OmniJudge or OmniBias? Diagnosing Multimodal Judges through Balanced, Decoupled Lenses Guangzheng Hu 1,4 Ziyue Jiang 2,5 Weixu Qiao 1 Lixin Zhang 1 Jianye Kang 1 Yuru Wu 1 Rong Bao 1 Niantong Li 1 Wei Wang 1 Ziyi Cheng 1 Xinfa Zhu 2 HangRui Hu 1 Ting He 1 Bing Zhao 3 Lin Qu 1 Hu Wei 1,â Jin Xu 2,⥠1 Alibaba Group 2 Qwen Team 3 Alibaba DAMO Academy 4 University of Melbourne 5 Zhejiang University guangzhengh@student.unimelb.edu.au â kongwang@alibaba-inc.com ⥠renjun.xj@alibaba-inc.com https://github.com/SKYLENAGE-AI/D3OmniFramework https://huggingface.co/datasets/skylenage-ai/D3OmniBench Abstract Multimodal understanding models that can jointly judge text-to-image (T2I), text- to-video (T2V) and text-to-speech (TTS) generation are increasingly used as âOm- niJudgesâ for evaluation and automatic annotation. How reliably they understand what they score remains unclear, since existing benchmarks and training data tend to overemphasize positive examples and to conflate distinct failure modes, so a judge may score well without recognizing failures while its capability gaps stay hidden. Motivated by this, we introduce D 3 -Omni, a balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding, covering53 orthogonal binary dimensions (17/22/14) and10,671samples (3,526/1,998/5,147) across the three tasks. Rather than re-generating outputs, which may leak informa- tion across dimensions, we fix verified fully positive seeds and derive negatives through controlled prompt rewriting and atomic, dimension-isolating perturbations. The resulting D 3 design is Dual-balanced, which helps alleviate negative-sample scarcity and per-dimension label imbalance; Decoupled, so that each error is attributable to a single capability; and Dynamic, steering construction toward under-represented regions of the label distribution as generative models improve. The suite reaches near1:1per-dimension parity and a uniform distribution over all total-score levels. Under this balanced view, even strong OmniJudges tend to struggle on modality-related dimensions, to confirm satisfied requirements far more reliably than they detect violated ones, and to treat nominally distinct attributes as largely a single decision, suggesting that aggregate accuracy may hide system- atic blind spots that a balanced and decoupled lens can help expose and, in turn, address. 1 Introduction Multimodal large language models (MLLMs) are rapidly evolving toward unified systems capable of understanding, reasoning, and generating across text, image, audio, video, and their combinations. Recent omni-modal benchmarks systematically assess these modelsâ ability to understand and reason across visual, auditory, acoustic, and textual inputs [1,2]. Beyond serving as response generators, MLLMs are increasingly deployed as automatic annotators, preference judges, reward models, and distillation teachers; when tasked with cross-task scoring across T2I, T2V, and TTS, they are arXiv:2608.24160v1 [cs.AI] 25 Aug 2026 020406080100 Expected Score (%) 60 70 80 90 100 Accuracy (%) 86.5 82.3 86.9 80.8 82.3 76.2 79.8 72.3 79.8 71.9 77.1 67.2 77.4 66.1 73.2 63.3 74.0 62.6 74.5 64.1 72.7 63.8 74.2 63.9 74.9 63.7 77.3 64.6 79.1 68.2 82.9 70.7 86.1 73.5 94.6 78.3 T2I Segment Accuracy Gemini-3.1-Pro (78.6%) Gemini-3.5-Flash (77.6%) Gemini-3-Flash (77.5%) Grok-4-Fast-Reasoning (75.9%) GPT-5.4 (75.5%) Qwen3.5-Omni-Plus (74.6%) GPT-5.2 (74.4%) Claude-Opus-4.6 (73.8%) Claude-Opus-4.7 (73.3%) Qwen3.5-Omni-Flash (72.4%) 020406080100 Expected Score (%) 50 60 70 80 90 100 98.0 89.6 94.9 85.4 89.2 80.4 89.4 78.1 84.2 76.4 81.9 72.2 81.0 72.8 77.8 67.0 75.4 65.1 74.0 58.9 73.4 56.6 73.2 60.6 70.8 58.7 69.8 58.4 70.0 56.9 69.4 52.6 68.0 53.6 70.4 61.9 71.8 61.6 72.6 62.2 76.2 66.8 79.3 67.0 82.3 68.9 T2V Segment Accuracy Gemini-3.5-Flash (76.7%) Gemini-3.1-Pro (76.0%) Gemini-3-Flash (75.2%) Qwen3.5-Omni-Plus (72.6%) Qwen3-Omni-Flash (71.3%) Qwen3.5-Omni-Flash (70.9%) Qwen3.6-27B (70.0%) Qwen3-Omni-30B (69.5%) 020406080100 Expected Score (%) 40 50 60 70 80 90 100 99.4 86.1 90.9 80.0 82.9 73.0 74.2 66.1 67.8 60.5 61.4 55.2 55.7 52.1 50.2 48.7 49.7 45.6 50.1 42.6 54.9 41.3 59.3 41.0 66.3 43.4 74.9 43.6 80.6 46.3 TTS Segment Accuracy Qwen3.5-Omni-Plus (65.7%) Gemini-3.1-Pro (60.7%) Qwen3.5-Omni-Flash (60.4%) Gemini-3-Flash (60.2%) Qwen3-Omni-30B (60.1%) GPT-Audio-1.5 (58.7%) 020406080100 Expected Score (%) 20 10 0 10 20 Predicted Expected (%) T2I Total Score Deviation 020406080100 Expected Score (%) 10 0 10 20 30 T2V Total Score Deviation 020406080100 Expected Score (%) 10 0 10 20 30 40 50 TTS Total Score Deviation Figure 1: Top: per-segment accuracy over the full ground-truth score range; every judge dips in the mid-score regime, forming a U. Bottom: total-score deviation (predictedâexpected): positive= Yes-bias, negative = No-bias, zero = calibrated. commonly referred to as âOmniJudgesâ. However, the evaluation of these models in such roles, particularly as judge models (JMs) or reward models (RMs), remains considerably less developed than the assessment of their generation capabilities. Given their expanding real-world applications, addressing this evaluation gap has become increasingly critical. Recent efforts evaluate LLM-as-a-judge systems, reward models, multimodal reward/preference models, and omni-modal reward models through dedicated benchmarks and preference datasets [3â7]. Many multimodal generation and evaluation benchmarks also rely on closed-source frontier models such as GPT-4o and Gemini for automated scoring or preference annotation [6,8]. Although these protocols streamline evaluation, they reveal an important limitation: existing judge and reward benchmarks are often not explicitly designed around distributional balance. Samples are typically skewed across modalities, tasks, score levels, and evaluation criteria, with certain score ranges or quality categories represented disproportionately. Such distributional biases obscure whether high judging performance reflects genuine fine-grained multimodal understanding or simply memorization of dataset priors and majority-class patterns. The need for balance is sharpened by a fundamental asymmetry between generation and judging. Real- world generations are predominantly acceptable outputsâcatastrophic failures are, by construction, the minorityâyet a judge is valuable precisely in the opposite regime, where it needs to catch the occasional error, the fine-grained mismatch, and the subtle factual or perceptual flaw inside an otherwise plausible sample. On a high-quality-dominated, imbalanced suite a judge can thus obtain a high score while remaining blind to the very failure modes that motivate its deployment: when90%of a dimensionâs samples are positive, a constant âYesâ already reaches90%accuracy without any genuine discrimination. A diagnostic omni-modal judge benchmark therefore needs two kinds of balanceâscore-level balance across quality intervals and dimension-level positive/negative parityâtogether with an evaluation taxonomy that is as decoupled as possible, so that failures can be attributed to specific abilities rather than conflated [9]. Another challenge is the cost of benchmark construction. Although human annotation is important for reliability, constructing and updating a balanced and decoupled omni-modal benchmark entirely through manual labeling is expensive and inefficient. The cost becomes even higher when balance is required simultaneously across modalities, score levels, and fine-grained evaluation dimensions. This is particularly limiting in the Omni-LLM setting, where model capabilities evolve quickly and static 2 benchmarks may soon lose their discriminative power. Therefore, instead of relying on exhaustive human annotation, benchmark construction should shift human effort toward targeted verification and quality control, while using automatic sample construction, model-assisted filtering, and consistency checking to support scalable data generation and dynamic benchmark updates [6, 8]. Balanced test sets are, in fact, standard practice elsewhere: in long-tailed and imbalanced learning, models trained on skewed data are routinely evaluated on class-balanced test sets, precisely because accuracy is only meaningful once the label distribution is controlled. Why, then, does no explicitly balanced omni-modal judge benchmark yet exist? The obstacle is not a lack of motivation but the difficulty of construction: the quality of generated media is not directly controllable, the natural prevalence of each fine-grained dimension is intrinsically imbalanced, and building decoupled negatives that isolate a single dimension demands precise, controllable sample construction that ordinary data collection cannot provide. Breaking through this barrier is the core enabler of our work: rather than chasing controllable generations, we reverse-construct the prompt from a verified positive seed and apply controllable, dimension-isolating rewriting, coupled with a dynamic dual- balancing loop, which finally makes a balanced and decoupled evaluation suite feasible and exposes the capability blind spots that skewed benchmarks had kept hidden. Figure 1 previews one such blind spot: on our balanced suite every judgeâs per-segment accuracy collapses on the mixed-quality middle of the score rangeâa U-shaped dip that any aggregate number conceals. To address these challenges, we propose D 3 -Omni, a balanced and decoupled benchmark for diag- nosing the fine-grained multimodal understanding of Omni-LLMs as JMs and RMs. Unlike existing multimodal benchmarks that mainly emphasize broad modality coverage or aggregate preference accuracy (e.g., [4,7]), D 3 -Omni focuses on distributional balance, decoupled assessment, and scal- able construction. Specifically, it is designed to reduce bias induced by skewed score distributions, disentangle different judging dimensions for fine-grained diagnosis, and support low-human-cost benchmark construction and dynamic updates as Omni-LLM capabilities continue to evolve. In summary, D 3 -Omni represents a paradigm shift from asking âhow well does the judge scoreâ to asking what the judge truly understands, and, equally importantly, what it fails to understand and evaluate. Contributions. Benchmark. We introduce D 3 -Omni, to the best of our knowledge the first balanced and decoupled omni-modal judge/reward benchmark for diagnosing the fine-grained multi- modal understanding of Omni-LLMs, covering T2I, T2V and TTS with10,671samples. Framework. We propose the D 3 frameworkâDual-balanced, Decoupled and Dynamicâwhich controls the total- score distribution and the per-dimension positive/negative ratio jointly, reaching near1 : 1parity on every dimension and a uniform distribution over all score levels. Taxonomy. We design a decoupled taxonomy of53orthogonal binary dimensions (17for T2I,22for T2V,14for TTS) that separates prompt-related compliance (instruction following, attribute binding, spatial and temporal reasoning, text typesetting, audio-text alignment) from modality-related perceptual fidelity (visual realism, anatomical coherence, temporal stability, audio quality, speaker characteristics), so that each error is attributable to one capability. Pipeline. We develop a low-human-cost, dynamically updatable construction pipeline combining automatic sample synthesis, model-assisted filtering, consistency checking and targeted human verification, allowing the benchmark to be extended as Omni-LLM capabilities evolve. Diagnosis. Using D 3 -Omni we surface blind spots that aggregate accuracy hides and that recur across model families: a U-shaped competence collapse on mixed-quality samples, a pervasive Yes-bias that widens from vision to speech, and pseudo-decoupling in which nominally orthogonal decisions collapse onto a single latent factor. Each shortcoming maps back to a concrete D 3 operator, turning diagnosis into actionable data-side interventions; and all of it is legible only on a score- and dimension-balanced suite, since reading a single high-score segment alone would already reorder the leaderboard. 2 Related Work 2.1 Benchmarks for omni-modal generative tasks. A large body of work evaluates the generation quality of individual modalities: for text-to-image, compositional benchmarks such as T2I-CompBench [10] and GenAI-Bench [11] probe object pres- ence, attribute binding, and spatial relations; for text-to-video, VBench [12], EvalCrafter [13], and T2VScore [14] cover text-video alignment, visual and motion quality, and temporal consistency; and 3 for text-to-speech, automatic evaluators predict perceptual quality, naturalness, intelligibility, and speaker similarity [15â17]. These benchmarks measure how well generators behave, but they treat the underlying evaluator as a trusted black box, leaving untested whether the evaluator itself reliably understands the fine-grained perceptual, alignment, and instruction-following properties it is asked to score. This gap motivates a closer look at the judge and reward models used for generative evaluation. 2.2 Judge and reward models for generative content. LLM-as-a-judge has become a scalable alternative to human evaluation [3], with text judges such as PandaLM, JudgeLM, Prometheus, and CompassJudger performing pairwise comparison, scalar scoring, and rubric-based critique [18â22]. In the generative-media domain, judges and reward models remain largely modality-specific: image reward models such as ImageReward, PickScore, and HPSv2 learn human preferences over generated images [23â26], while speech and video rely on learned quality and alignment metrics [12â17]. In parallel, generic omni-modal MLLMs are increasingly used as default judges across all three tasksâclosed frontier systems (GPT-4o, GPT-5, Gemini, Claude Opus, Grok) [27â32] and open omni-LLMs (Qwen3-Omni, MiniCPM-o, Mini- Omni2, Baichuan-Omni) [33â36]. Yet none has been independently verified as a unified judge applying the same fine-grained criteria across T2I, T2V, and TTS: modality-specific evaluators approximate human preference scores and generic MLLMs are optimized for general task-solving, so in neither case is the judge itself tested for the perceptual, alignment, and instruction-following understanding that reliable evaluation requires. This motivates benchmarking the judges directly. 2.3 Benchmarks for evaluating judge models. A growing line of work probes judge reliability. Text-only judge benchmarks test whether LLM judges can compare, score, or rank outputs across dialogue, instruction following, reasoning, coding, and safety [3,4,37,38], revealing biases such as position and verbosity bias and weak sensitivity to factual correctness [3,4,37]. Recent multimodal and omni-modal benchmarks extend this to vision-language and any-modality judging [1,2,5â7,9,39,40]. Three limitations persist. First, they concentrate on text-only or vision-language settings, leaving speech and video judgment under-studied. Second, they score the final judgment directly without disentangling the underlying sub-abilities, so a failure cannot be attributed cleanly to perceptual misreading, prompt-alignment failure, or criterion misinterpretation. Third, their label distributions are inherited from naturally collected outputs and rarely controlled, so high accuracy may reflect label priors rather than understandingâwhich turns attention to how the labels themselves are distributed. 2.4 Imbalanced label distribution undermines diagnostic capacity. Because most judge benchmarks are built from naturally collected outputs or human preferences [3, 4,39,40], the positive and negative labels within a fine-grained dimension (e.g., object presence, speaker emotion, temporal consistency, prompt faithfulness) are typically skewed. A judge can then post high accuracy by exploiting label priorsâindeed a trivial majority-class baseline already doesâ so aggregate metrics fail to surface real capability gaps and cross-dimension comparisons become unreliable. The generation-judging asymmetry compounds this: real generations are themselves skewed toward acceptable samples, so a benchmark that mirrors this distribution leaves precisely the regime where a judge is most neededâcatching the occasional error and the subtle perceptual or factual flawâessentially untested. A diagnostic omni-modal judge benchmark therefore needs explicitly balanced per-dimension labels across text-to-speech, text-to-image, and text-to-video, so that it measures whether judges recognize both the presence and the absence of each fine-grained property rather than dataset bias. 3 Benchmark Dimension Design We aim to evaluate whether multimodal understanding modelsâso-called âOmniJudgesââcan make accurate, fine-grained, and decoupled judgments about generated content. To this end, each benchmark instance is formalized as a triplet(p,x,y), wherepâPis an input prompt,xâXis a generated modality (image, video, or speech), andyâ0, 1 D is a ground-truth binary label vector. A judge model f predicts Ë y = f (p,x), and its performance is measured by comparing Ë y to y. 4 3.1 Motivation A judge that produces only an aggregate quality or preference score reveals nothing about which sub-ability is failingâexactly the diagnostic gap motivated in our introduction and related-work discussion. Yet judging generated content fundamentally requires two mechanistically distinct abilities: interpreting natural-language intent and perceiving low-level signal integrity. The first is a language-grounding problem operating jointly over(p,x); the second is a perceptual-fidelity problem operating purely overx. Folding them into a single rubric makes failure modes inseparableâa model that perfectly understands intent but cannot detect visual artifacts is, under aggregate scoring, indistinguishable from one with the opposite weakness. We therefore decompose every judgment into two disjoint categories, Prompt-related (D p ) and Modality-related (D m ), and inside each category we further refine the requirement into the smallest atomic dimension on which a controlled perturbation can flip the label without affecting any other dimension. This recursive decoupling is the structural prerequisite for the targeted, dimension-isolating negative-sample construction in Section 4. 3.2 Formalization and Two-Category Decomposition Each dimensiond â 1,...,Dcorresponds to a binary requirement of the form: âDoes the generated output satisfy this specific requirement?â with answery d = 1(Yes) if satisfied andy d = 0 (No) otherwise. This formulation ensures judgments are unambiguous and directly attributable to specific capabilities. The two categories are formally defined as follows. Prompt-related dimensions (D p ). These assess semantic alignment between the prompt and the output. They encompass requirements concerning subject identity, spatial and relational structure, compositional logic, stylistic intent, and adherence to explicit or implicit constraints in the prompt. For these dimensions, the ground-truth judgment depends on the joint interpretation of p and x. Modality-related dimensions (D m ). These evaluate intrinsic properties of the generated signal that are independent of the prompt. They cover perceptual and structural qualities such as visual or audio fidelity, temporal coherence, physical plausibility, and low-level artifact presence. Here, the judgment is determined solely byx, as it pertains to the modalityâs internal consistency and realism. This categorization enables disentanglement of two fundamental judgment capabilities: understanding user intent versus perceiving signal integrity. Diagnosing which category exhibits systematic errors reveals whether a modelâs limitations lie in language grounding or perceptual robustness; for example, a judge that scores high onD p but consistently fails onD m is one that follows the rubric semantically yet remains blind to subtle perceptual flawsâan extremely common pattern that we document empirically in Section 5. To support meaningful fine-grained analysis, all dimensions are designed to be orthogonal: the satisfaction of any requirement should not depend on the state of others. For instance, whether an image correctly depicts the requested color (attribute binding) should be independent of whether the objects are in the correct spatial arrangement (spatial reasoning); a sample is allowed to fail on color while passing on layout, enabling precise localization of the deficit. The same orthogonality principle is applied recursively to the fine-grained sub-dimensions inside each category: each requirement is the smallest atomic unit whose label can be flipped through a single controlled perturbation without altering the label of any other dimension. This ensures that a failure on one dimension reflects a localized judgment deficit rather than a side effect of correlated attributes, and it is precisely what makes targeted negative-sample construction tractable downstream. The total score is defined as the integer sum: s(y) = D X d=1 y d â0, 1,...,D, which counts the number of satisfied requirements. Due to orthogonality, deviations from the maximum score can be precisely attributed to specific violated dimensions. 5 3.3 Granularity and Coverage Most existing automatic evaluators for generative media collapse âqualityâ or âalignmentâ into one or a handful of coarse metricsâFID, CLIP-similarity, MOS, single-axis preference scores, or a small set of generic Likert categoriesâwhich neither pinpoint specific failure types nor cover the breadth of attributes a modern OmniJudge is expected to inspect. Our taxonomy goes substantially finer and substantially broader: 17, 22, and 14 atomic binary requirements for T2I, T2V, and TTS respectively, totallingD = 53decoupled dimensions. Critically, every fine-grained sub-dimension is selected to target a specific failure mode that frontier generators still occasionally exhibit despite producing visibly high-quality outputs overallâe.g., subtle text-rendering glitches and finger-anatomy errors in T2I, audioâvideo desynchronization and inter-frame flicker in T2V, or mismatched speaker personality and unstable volume in TTS. These rare-but-critical errors are exactly the cases a deployed judge is expected to detect reliably, yet they are also precisely the cases for which naturally collected data offers very few negative examples. Designing each requirement as small, atomic, and mutually orthogonal is therefore not a stylistic preference but a functional necessity: only such a taxonomy can isolate, balance, and diagnose each rare failure mode in turn, and only such a taxonomy admits dimension-specific negative samples whose construction does not leak into other dimensions. 3.4 Instantiation Across Three Tasks We instantiate this design across three generation scenarios; the complete dimension lists, definitions, and category assignments are provided in Tables 3, 4, and 5 in the appendix. Text-to-Image (T2I).D = 17: 14 prompt-related dimensions covering composition, attribute binding, spatial reasoning, typography, and negative-instruction handling, plus 3 modality-related dimensions evaluating material realism, edge clarity, and anatomical coherence; the latter three being precisely the residual perceptual flaws that even state-of-the-art diffusion models still leak through on otherwise impressive renderings. Text-to-Video (T2V).D = 22: 16 prompt-related dimensions spanning subject, scene, lighting, style, and audio alignment, plus 6 modality-related dimensions targeting temporal stability, focal sharpness, motion rhythm, audio quality, and audioâvideo synchronization, where even strong T2V models still produce occasional glitches that aggregate quality scores routinely overlook. Text-to-Speech (TTS).D = 14: 10 prompt-related dimensions covering textual and punctuation faithfulness together with eight separately-judged speaker characteristics (age, gender, personality, timbre, speed, pitch, tone, emotion), plus 4 modality-related dimensions on voice clarity, background noise, volume stability, and spectral integrity. This is a fine-grained decoupling of speaker attributes that current TTS systems often blur into a single âlooks-goodâ voice. However, real-world data lacks the balance and isolation required to evaluate this taxonomy reli- ably: outputs are skewed toward high scores, single-dimension negative examples are scarce, and conventional negative synthesis often violates orthogonality through unintended cross-dimensional effects. To address this, we introduce a dynamic dual-balanced construction framework (Section 4) that generates triplets(p,x,y)with dimension-isolated negatives and near-exact balance across both per-dimension labels and total score levels. 4 Balanced Benchmark Construction As established in Section 3, reliable fine-grained diagnosis of an omni-modal judge requires per- dimension label balance and score-level uniformity over a taxonomy whose sub-dimensions are mutually orthogonal. Achieving this in practice, however, is far from trivial: directly resampling new outputs from a perturbed prompt unavoidably introduces cross-dimensional leakage: even minor edits (e.g., ârunningâââwalkingâ) drift the pose, motion, background, or rendering quality, simulta- neously corrupting several supposedly independent dimensions and destroying the orthogonality on which fine-grained attribution depends. We therefore propose D 3 -Construction, a benchmark-construction pipeline organized around three pillars whose names (Decoupling, Dual-Balancing, and Dynamic) all begin with the letter D and 6 together operationalise the design objectives of Section 3. Decoupling (§4.1) makes every generated sample atomically attributable to a single dimension by fixing a fully positive seed and altering only one factor at a time; Dual-Balancing (§4.2) gives the precise mathematical objective that the resulting benchmark must satisfy and proves it is realizable; Dynamic (§4.3) provides the iterative procedure that actually drives any partial benchmark towards that objective. Notation.LetÏ âT2I, T2V, TTSdenote the generation task,D (Ï) p andD (Ï) m the disjoint sets of prompt-related and modality-related sub-dimensions, andD Ï =|D (Ï) p | +|D (Ï) m | their total size. The catalogues defined in Section 3 yieldD T2I = 17,D T2V = 22, andD TTS = 14. Whenever the task is unambiguous we drop the superscript (Ï ). Each benchmark example is a triple p, x, y â P ĂX Ï Ă0, 1 D Ï ,(1) wherepis a natural-language prompt,xa generated modality artifact (image, video, or speech waveform), andy = (y 1 ,...,y D Ï )a binary label vector aligned positionally with the rubric in Section 3:y d = 1(YES) iff(p,x)satisfies thed-th rubric requirement, andy d = 0(NO) otherwise. The total score of a sample iss(p,x,y) = P D Ï d=1 y d â 0, 1,...,D Ï , and we writes(y)when (p,x) are clear from context. 4.1 Decoupling: Atomic, Single-Dimension Negative Construction Why decouple at the construction step. Aggregate quality on its own conveys nothing about which sub-ability is failing; pinpointing the failing dimension requires negatives whose violation is, by construction, attributable to one and only oned â . We obtain such atomically-attributable negatives by decoupling modality generation from semantic mismatch creation: rather than re-synthesizing a new modality, we fix a fully positive seed(p + ,x + )and inject a single, dimension-isolated edit on either the prompt side (for d â âD p ) or the modality side (for d â âD m ). Definition 1 (Fully positive seed). A pair(p + ,x + )is a fully positive seed for taskÏiff its label vector under the rubric of Section 3 isy + = 1 D Ï , i.e. every prompt-related and every modality-related dimension is satisfied. We denote the set of such seeds by S (Ï) full = (p + ,x + ) : y(p + ,x + ) = 1 D Ï . Definition 2 (Single-dimension flip operator). For a target dimensiond â â D p âȘD m , a single- dimension flip operatorF d â :PĂXâPĂX is any map satisfying y F d â (p + ,x + ) d = 0, d = d â , 1, dâ1,...,D Ï \d â , â(p + ,x + )âS full .(2) F d â is prompt-side if it modifies only p + and modality-side if it modifies only x + . Step 1: Seed prompt construction with a tri-LLM modality-aware expert pipeline. A correct negative is meaningful only on top of a verifiably correct seed. We therefore buildS full through a modality-aware inverse-prompting procedure executed independently on three large language models (the TTS / T2V / T2I captioning experts implemented respectively as Qwen-3.5 Plus, Gemini-3.1 Flash, and Gemini-3.1 Pro), each conditioned on a modality-specific inverse template that explicitly enumerates every prompt-related sub-dimensiond â D (Ï) p from Section 3. The three captions are reconciled by a human-in-the-loop calibration pass that resolves disagreements and rewrites the seed promptp + until every sub-dimensiondinD p is grounded by a literal, modality-faithful requirement (e.g., for TTS the prompt explicitly specifies rate, volume, timbre, emotion, etc., matching what the audio actually exhibits). Concretely, given the raw modality x, p + (x) = CALIBRATE RECONCILE(c 1 ,c 2 ,c 3 ), D (Ï) p , x , c i = LLM (Ï) i x; T (Ï) inv ,(3) whereT (Ï) inv is the modality-specific inverse template. Modality-related labels for the samexare then verified by the same LLM ensemble (with conflicts again adjudicated by a human auditor); we deliberately keep this verification ofD m manual because a mis-certified seed would silently propagate as a false negative once perturbed. Only samples whose certified labels equal1 D Ï enter S full . 7 Step 2: Prompt-side flips with graded counter-semantic rewriting.Ford â âD p we instantiate F d â as a controlled prompt rewrite executed by Gemini-3.1 Pro, governed by: (i) minimal modification (alter only tokens directly tied tod â ); (i) dimensional isolation (assert all other tokens that ground dimensionsdÌž= d â remain unchanged); and (i) structural coherence (preserve fluency). To stress- test judges across the full counter-semantic spectrum, the rewriter does not flip in a single fixed direction; it samples a semantic-distance levelâ â L d â and emits ap â at that level. For example, when the seed TTS prompt requires âvery fastâ speaking rate (d â =rate), the rewriter randomly draws amongââmoderate-opposite: medium, strong-opposite: slow, maximal-opposite: very slow, all of which still violate the rate dimension but at different degrees of semantic divergence, ensuring that the resulting negatives span a wide range of promptâmodality semantic distance and prevent judges from over-fitting to one particular flip template. Formally, the prompt-side flip is realized at a levelâ drawn uniformly fromL d â , giving (p â ,x + ) =F (â) d â (p + ,x + ), and the resulting negative example is (p â ,x + ,y (d â ) ) with y (d â ) obtained from 1 D Ï by setting the d â -th coordinate to 0. Step 3: Modality-side flips via human-verified atomic operators. Ford â âD m a prompt edit cannot induce the failure; we instead apply a modality-specific atomic operatorO d â :X Ï âX Ï to the seed modalityx + (the promptp + is held fixed). Eachx + entering this stage has been manually inspected, ensuring that the post-perturbation defect is the only reason a modality-related dimension would flip. At a high level, the operators target three families of perceptual defects: spatial / textural distortions (e.g. frequency-domain blurring, anatomical deformation, segmentation-driven warping), temporal / synchronization perturbations (e.g. frame-order shuffling,2Ăspeed masking, audioâ video offset), and audio-fidelity manipulations (e.g. spectrogram-domain bilateral filtering, additive reverberation, gain modulation, band-targeted frequency manipulation). The full list of atomic operators used in this workâ6 for T2V, 3 for T2I, and 4 for TTSâtogether with their implementation details is provided in Appendix A.3. For every(d â ,Ï )the corresponding operatorO (Ï) d â has been validated to leave alldÌž= d â unchanged on a held-out audit set, so that Ìx =O (Ï) d â (x + )paired with the unaltered p + produces the desired triple (p + , Ìx,y (d â ) ) in Eq. (1). Diagnostic metric for decoupling: the dimension-coupling matrix.To verify that the decoupling holds end-to-end (both in the rubric itself and in any judgeâs predictions), we report the empirical dimension-coupling matrixÎŁ â R D Ï ĂD Ï whose(d,d âČ )entry is the Pearson correlation of the YES/NO indicators across the benchmark: ÎŁ d,d âČ = P N i=1 (y i,d â Ìy d )(y i,d âČ â Ìy d âČ ) q P N i=1 (y i,d â Ìy d ) 2 q P N i=1 (y i,d âČ â Ìy d âČ ) 2 , Ìy d = 1 N N X i=1 y i,d .(4) We computeÎŁon (a) the ground-truth labels (the reference matrix) and on (b) each judgeâs predicted labels; substantial off-diagonal mass in (b) that is absent from (a) is direct evidence of spurious coupling, i.e., the judge cannot truly separate the underlying capabilities, which is one of the four diagnostic axes evaluated in Section 5. 4.2 Dual-Balancing: Definition and Existence Theorem The two balance constraints. A benchmarkB = (p i ,x i ,y i ) N i=1 is dually balanced iff it is uniform along the score-segment axis and along the per-dimension label axis. WritingN = M (D Ï + 1) with M â N + samples per score level, we require i : s(y i ) = s = M,âsâ0, 1,...,D Ï ,(Score uniformity) N X i=1 y i,d = N 2 = M(D Ï +1) 2 , âdâ1,...,D Ï .(Dimension parity) Eq.(Score uniformity)enforces segment balance so that each total-score bin is equally represented; Eq.(Dimension parity)enforces per-dimension balance so that no single sub-ability can be solved by predicting the majority class. Importantly, neither constraint alone is sufficient: a benchmark can be score-uniform yet have one dimension always equal to1, and conversely a benchmark can be per-dimension-balanced yet entirely concentrated near sâD Ï /2 (cf. §5.2 of the experiments). Lemma 1 (Attainability of dual balance). Fix a task Ï and suppose 8 (A1) S (Ï) full Ìž= â (fully positive seeds exist). (A2) For everyd â 1,...,D Ï , a single-dimension flip operatorF d satisfying Eq.(2)is realizable. (A3)For everyS â1,...,D Ï the compositionF S =â dâS F d acts as the characteristic flip y(F S (p + ,x + )) d = 1[d /â S] for every seed. Then for every tolerance(Ï,ÎŽ)withÏâ [0, 1)andÎŽ â„ 0there is an achievable benchmark meeting the stopping criterion(7); the target can be tightened to exact dual balance (Eqs.(Score uniformity)â (Dimension parity)), which is realizable with|B| = M (D Ï + 1)wheneverM (D Ï + 1)is even. Consequently the three monitored quantities of §4.3 can be driven below any prescribed threshold, however strict. Proof sketch.Exact dual balance is the tightest instance (Ïâ 1,ÎŽ = 0) and dominates every looser tolerance, so it suffices to construct it. We do so explicitly in Appendix A.1: single-dimension flips are composed via (A2)+(A3) into multi-dimension flips whoseM (D Ï + 1)label vectors realise each score level exactlyMtimes and load every dimension exactlyN/2times, establishing Eqs.(Score uniformity)â(Dimension parity)whenM (D Ï + 1)is even; any looser(Ï,ÎŽ)is then met a fortiori. 4.3 Dynamic: Iterative Construction Algorithm Theorem 1 guarantees that a dually balanced benchmark exists, but in practice the seed pool, the LLM- rewriter, and the modality operators all incur finite costs, so we cannot enumerate the combinatorial template once and stop. Instead, we growBdynamically: at each step we identify the least-populated non-full score, the most positive-dominated dimensions, and the least-loaded seed, and append one controlled multi-dimension negative for that target. Monitored quantities. WriteN t =|B (t) |and letÏ(i)âS full be the seed of samplei. OverB (t) we track exactly the three quantities that define the target: the per-dimension positive and negative proportions, the score histogram, and the per-seed negative load, Ï YES d = 1 N t X i 1[y i,d = 1], Ï NO d = 1â Ï YES d , N (t) s = i : s(y i ) = s , u (t) Ï = i : Ï(i) = Ï, s(y i ) < D Ï . (5) Per-step target. Each step addresses the least-populated non-full score level, the most positive- dominated dimensions, and the least-loaded seed: s â = arg min 0â€s<D Ï N (t) s , T â â arg min |T|=D Ï âs â X dâT Ï NO d Ï YES d , Ï â = arg min ÏâS full u (t) Ï ,(6) with ties broken uniformly at random, so thatT â gathers theD Ï â s â dimensions of smallest negative/positive ratio. Applying the composite flipF T â of §4.1 toÏ â = (p + ,x + )âits prompt- related coordinatesT â â©D p by graded rewriting and its modality-related coordinatesT â â©D m by the atomic operatorsâproduces a negative of scores â with labely (T â ) , wherey (T â ) d = 1[d /â T â ], which is appended toB (t+1) . Stopping criterion. The loop halts once the two balance axes are within tolerance: min d Ï NO d Ï YES d > Ïandmax s N (t) s â min s N (t) s < ÎŽ.(7) The first condition forces every dimension to a negative/positive ratio of at leastÏ(perfect balance is ratio1); the second flattens the score histogram; the seed loadu Ï enters only through the least- loaded-seed choice in Eq.(6), keeping negatives spread across seeds rather than being a stopping target. For all three tasks of D 3 -Omni we useÏ = 0.95andÎŽ = 6. Each step raises the negative count of the currently most positive-dominated dimensions and of the least-filled score, so both deficits 9 Algorithm 1 D 3 -Construction: dynamic dual-balanced benchmark assembly Require:Raw modalitiesX;D p ,D m ; tri-LLM ensembleLLM 1 , LLM 2 , LLM 3 ; reconciler / rewriter Gemini-3.1 Pro; modality operatorsO d dâD m ; tolerances Ï,ÎŽ. Ensure: BenchmarkB meeting the stopping criterion Eq. (7). Stage A: Seed curation (Decoupling, Step 1). 1:Run inverse-prompting onXwith eachLLM i , reconcile + human-calibrate to obtainp + (x)for every x. 2: Run modality-integrity check on every x viaLLM i + human audit. 3: S full â(p + (x),x) : y(p + (x),x) = 1 D Ï ; B âS full . Stage B: Monitorâgenerateâupdate loop. 4: Compute Ï NO d /Ï YES d , N s , u Ï overB.â· Eq. (5) 5: while min d Ï NO d /Ï YES d â€ Ï or max s N s â min s N s â„ ÎŽ doâ· Eq. (7) 6: s â â arg min s<D Ï N s ; T â â arg min |T|=D Ï âs â P dâT Ï NO d /Ï YES d ; Ï â â arg min Ï u Ï . â· Eq. (6) 7:(p â ,x â ) â F T â (Ï â ): prompt coords ofT â by rewriting, modality coords by operators; y â â y (T â ) . 8: B âBâȘ(p â ,x â ,y â ); update Ï NO d /Ï YES d ,N s ,u Ï . 9: end while 10: returnB. decrease monotonically; by Lemma 1 the tolerance region is attainable for any(Ï,ÎŽ)âindeed down to exact balanceâso the loop terminates. The output of Algorithm 1 is a benchmark meeting the stopping criterion Eq.(7)âthe attainable target guaranteed by Lemma 1âin the canonical triplet format of Eq.(1), equipped with the diagnostic metrics of §4.2; this is exactly the benchmark evaluated by the experiments in Section 5. 4.4 Diagnostic Metrics Derived from the Dual Balance The dual balance is precisely what makes the metrics below readable as capability signals rather than as artifacts of an imbalanced label distribution. Let Ë y i â0, 1 D Ï be a judgeâs predicted label vector for the i-th benchmark sample, and letB s =i : s(y i ) = s. Metric 1: Per-dimension accuracy. This metric quantifies the judgeâs resolution of thed-th sub-ability: Acc d = 1 N N X i=1 1 Ëy i,d = y i,d .(8) Metric 2: Per-segment accuracy. This metric is the mean per-sample dimension match inside score bin s: Acc s = 1 |B s | X iâB s 1 D Ï D Ï X d=1 1 Ëy i,d = y i,d .(9) Under Eq.(Score uniformity)every|B s |equalsM, so the curves7â Acc s is undistorted by segment mass. Metric 3: Per-segment perfect-match rate.This metric reports the fraction of samples in segment s whose predicted label vector exactly matches the ground truth: PM s = 1 |B s | X iâB s 1 h Ë y i = y i i .(10) Metric 4: Dimension-restricted Yes/No exact-match rates.These metrics report respectively the fraction of samples on which a judge gets every ground-truth-YES dimension right, and analogously 10 for NO: PM YES = 1 |I YES | X iâI YES 1 h Ëy i,d = 1, âdâY + i i ,Y + i =d : y i,d = 1,(11) PM NO = 1 |I NO | X iâI NO 1 h Ëy i,d = 0, âdâY â i i ,Y â i =d : y i,d = 0,(12) whereI YES =i :Y + i Ìž= â andI NO =i :Y â i Ìž= â . Metric 5: Dimension-coupling matrix.ÎŁ(Eq.(4)) is evaluated separately ony i (reference) and on Ë y i (predicted), with off-diagonal excess ÎŁ pred d,d âČ â ÎŁ ref d,d âČ flagging spurious coupling. These five quantities, taken jointly, instantiate the four diagnostic axes listed in Section 5: segment robustness (Eq.(9)), dimension-wise resolution (Eq.(8)), dimensional decoupling (Eq.(4)), and Yes/No commit behavior (Eq. (12)). 5 Experiments We evaluate a suite of state-of-the-art multimodal judges on D 3 -Omni and organise the diagnosis around five questions that an unbalanced benchmark cannot answer: (i) at the macro level, do todayâs OmniJudges differ enough to be distinguished, and does the ranking transfer across modalities (Sec. 5.1)? (i) as the ground-truth total score varies, in which score regime does each judge break down, and in which direction does it miscalibrate (Sec. 5.2)? (i) when forced to commit on all-Yes or all-No segments, does a judge expose a structured Yes/No class bias, and does its aggregate accuracy actually reflect an ability to localize the few defects that a segment contains (Sec. 5.3)? (iv) when nominally orthogonal dimensions are to be judged independently, does a judge keep its per-dimension decisions decoupled or let one decision drag the others (Sec. 5.4)? (v) at the per-dimension level, which fine-grained requirements floor every judge near chance, and which expose the largest inter- model gaps (Sec. 5.5)? Because D 3 -Omni is balanced over both dimensions and total scores, every drop in accuracy is attributable to a localized capability gap rather than to a skewed label prior, turning aggregate numbers into a transparent diagnostic map. The diagnosis is built around four figures (Figures 1, 2, 3 and 5); Sec. 5.7 then maps the four diagnosed shortcomings back to the three operators of the D 3 construction framework (Sec. 4). The full judging prompts are reproduced in Appendix A.4. 5.1 Setup and Overall Accuracy Judges, tasks and metrics. We benchmark ten judges on T2I, eight on T2V and six on TTS, drawn from the Gemini, GPT, Claude, Grok and Qwen families; we treat them under a single pool, since the goal is to characterize judging behavior rather than to compare access types. Every judge of a given task is queried with the same prompt template (Appendix A.4) and answers in a JSON array ofDYes/No decisions. Unless stated otherwise, the observations below refer to the judges evaluated here. All judges are queried at temperature0with every other decoding parameter left at its provider default, so the reported behavior reflects each modelâs deterministic mode rather than sampling variance. Because every dimension is a binary judgment, we adopt four complementary metrics: (i) per-dimension binary accuracy, the primary metric; (i) segment-wise accuracy on the total-score slices(y) = P d y d â0,...,D, with one curve per judge per task, together with the total-score calibration curve; (i) the off-diagonal Pearson correlation matrix of each judgeâs predicted per-dimension decisions, which measures whether the judge decides each near-orthogonal dimension independently or allows one latent factor to drive several outputs (a pair is flagged at Ï> 0.6and treated as strongly entangled atÏ> 0.8); and (iv) per-segment Yes/No perfect-match rates, which expose class-prior shortcuts and, more strictly, whether a judge can localize every minority defect within a segment. Overall accuracy hides the diagnosis. Averaging per-dimension binary accuracy over the entire test set gives a tight band on the visual tasks: the ten T2I judges sit between72.4%and78.6%and the eight T2V judges between69.5%and76.7%, a6â7point spread that one would dismiss as a saturating leaderboard. TTS does not so much widen the band as reorder it: accuracy compresses to 11 T2IT2VTTS JudgeD p D m LowMid HighD p D m LowMid HighD p D m LowMid High Gemini-3.1-Pro 83.21 56.93 81.274.2 80.385.20 51.41 71.1 71.784.5 64.9350.16 58.450.5 73.2 Gemini-3-Flash 81.9157.1781.5 73.078.1 84.8149.52 69.372.3 83.563.09 52.9846.8 50.783.0 Qwen-3.5-Omni-Plus 78.27 57.65 78.8 67.2 77.9 80.30 52.04 68.1 66.3 82.6 66.33 64.20 67.2 51.4 78.5 Qwen-3.5-Omni-Flash 76.64 52.74 69.9 66.9 80.5 77.85 52.27 66.2 63.1 82.3 64.02 51.54 55.5 50.4 75.4 Gemini-3.5-Flash 82.85 52.92 78.0 73.4 81.3 85.57 53.3071.371.3 87.0â GPT-5.4 80.2053.70 78.8 70.077.7â Grok-4-Fast 80.00 56.88 79.468.5 79.9â GPT-5.2 78.96 52.94 77.3 68.4 77.4â Claude-Opus-4.7 77.18 55.57 78.7 66.1 75.2â Claude-Opus-4.6 77.10 58.60 80.8 64.7 76.0â Qwen-3.6-27Bâ 77.76 49.16 65.0 65.678.8â Qwen-3-Omni-Flashâ 76.96 56.18 71.8 60.4 80.3â Qwen-3-Omni-30Bâ 76.92 49.53 63.0 61.6 83.0 64.35 49.6251.8 50.378.3 GPT-Audio-1.5â 62.1850.25 43.251.0 81.9 Table 1: Per-judge accuracy with the three tasks side by side. Top four rows: cross-task judges present in all tasks; below: task-specific judges (â: not evaluated). For each task,D p /D m are the prompt-/modality-related macro accuracies and Low/Mid/High the accuracies in the0â33/33â66/66â 100%score bands. Per column, bold marks the best andunderlinethe second-best value, computed separately within the cross-task and task-specific blocks. 58.7%â65.7%and the family that leads the visual tasks no longer leads, with Qwen-3.5-Omni-Plus on top (65.70%) and the two Gemini judges, dominant on T2I and T2V, sliding into the lower half. The four diagnostic axes that follow show that the visual-task tightness is itself an artifact: every T2I/T2V judge loses substantial accuracy on specific fine-grained regimes, but each one loses it in a different regime, so the macro average conceals these losses. 5.2 Score-Segment Behavior and Calibration Reading the two rows. We slice the test set by the ground-truth total scores(y) = P d y d , which counts how many of theDrequirements a sample satisfies; by construction every score segment holds the same number of samples. The top row of Figure 1 plots, for each judge, its per-dimension binary accuracy within each segment (one curve per judge, the legend value being the overall mean), and thus reveals the difficulty regime on which a judge decides accurately or breaks down. The bottom row plots the total-score deviation, predicted minus expected total, so a point above the zero baseline is over-scoring and one below is under-scoring, revealing the direction in which the judgeâs scores are biased. Read together, the two rows turn a single aggregate accuracy into a difficulty-resolved capability curve and a bias direction. Top row: a U-shaped collapse on the mixed-quality middle.Every curve is markedly U-shaped: a judge decides easily on the all-violated and all-satisfied extremes but loses substantial per-dimension accuracy in the middle, where roughly half the requirements hold and the other half fail, i.e. the ambiguous mixed-quality samples on which no all-Yes or all-No prior helps and each dimension has to be decided on its own. On T2I, GPT-5.2 reads84.9%ats = 0and83.4%ats =Dyet drops to about 66%in the middle; at the aggregate level Gemini-3.1-Pro leads T2I (78.6%), Gemini-3.5-Flash leads T2V (76.7%), and TTS is hardest, where the best judge, Qwen-3.5-Omni-Plus, reaches only65.7%. The dip marks a judgeâs real weakness and sits precisely where ordinary benchmarks are sparsest, so an unbalanced set yields a high mean from the two easy tails while the collapse remains hidden. Table 1 gives the banded counterpart (Low/Mid/High=the0â33/33â66/66â100%score bands): for almost every judge the Mid band is the trough, e.g. Gemini-3.1-Pro on T2I reads81.2/74.2/80.3and Qwen-3.5-Omni-Plus 78.8/67.2/77.9. The shape of the dip is itself a quality profile.Judges differ in the depth, position and symmetry of the U, and each shape reads as a distinct competence profile. Gemini-3.1-Pro traces the shallowest, flattest U on T2I (trough72.7%, endpoints89.2%and85.2%), a judge equally at ease on clearly- bad, mixed and clearly-good samples, which is why it also tops the macro table; GPT-5.2 and Gemini-3.5-Flash are near-symmetric, concentrating their error squarely on the half-satisfied middle. Claude-Opus-4.6 is the opposite extreme, a deep narrow valley (trough62.6%) whose all-violated 12 tail is the highest of any judge (94.6%) but whose recovery is weak, the profile of a judge that only reliably flags blatant defects, while Grok-4-Fast-Reasoning shows high shoulders over a deep trough, so reading only its extremes would overstate it. The lone T2I exception is Qwen-3.5-Omni-Flash: instead of a symmetric U it traces a rising slope whose trough sits early (near29%) and whose all-satisfied end (83.9%) exceeds its all-violated end (78.3%), the only judge more reliable on good samples than on blatantly bad ones. Endpoint asymmetry: confirming quality versus detecting defects. Beyond depth, the relative height of the two arms is diagnostic: the differencehiâlobetween the all-satisfied and all-violated ends says whether a judge is better at confirming good content or at catching bad content. Averaged over judges the sign flips with the modality: T2I is mildly negative (â3.67; Claude-Opus-4.6 reachesâ10.6), so judges are stricter on poor images, whereas T2V (+19.05) and TTS (+32.45) are strongly positive, so on the temporal tasks judges confirm all-good content well but cannot flag every defect on all-bad content; the asymmetry is same-signed and slightly larger in Chinese (T2V +21.22, TTS+36.81). The trough migrates accordingly, deepening and shifting toward the low-score, mostly-violated side on the temporal tasks (the three T2V Gemini judges bottom out near27%, and Qwen-3-Omni-30B reaches52.57%, one of the deepest collapses in the study), so the hardest samples change from half-right-half-wrong on T2I to mostly-wrong-but-not-all on T2V and TTS. The extreme is GPT-Audio-1.5 on TTS, closer to a monotonic ramp than a U, rising from46.3%on all-violated speech to 98.5% on all-satisfied speech and essentially saying Yes until the audio is clearly clean. Bottom row: a regression-to-the-mean bias. The deviation curve falls monotonically with the ground-truth score, giving a consistent pattern of over-scoring the bad and under-scoring the good: every judge sits above the zero baseline ats = 0and below it ats =D, crossing zero somewhere in the low-to-mid range. On T2I, GPT-5.2 over-scores by+15points on all-violated samples and under-scores byâ17on all-satisfied ones, so the error is directional and predictable rather than random. The bias is most severe on the all-violated tail of the temporal tasks: even the best-calibrated TTS judge, Qwen-3.5-Omni-Plus, still mislabels about one in five No dimensions as Yes ats = 0, and GPT-Audio-1.5 mislabels more than half (46.3%accuracy there). This is the macro-level footprint of the Yes-biased class prior that Sec. 5.3 isolates dimension by dimension, where no TTS judge produces all-No on more than 12.24% of the samples for which all-No is correct. Why only a balanced set reveals this, and a counterfactual. Both the U trough and the sign- changing deviation line can be read only when every score segment is populated. A direct coun- terfactual makes the point: reading a single high-score segment would rewrite the leaderboard, since the all-satisfied segment alone would rank Gemini-3-Flash first on TTS (99.4%), a judge third from last under balanced macro accuracy (60.2%), while the all-violated segment would instead promote Qwen-3.5-Omni-Plus; on T2I the all-violated segment would elevate Claude-Opus-4.6 (94.6%), eighth of ten on the macro average. Score uniformity is therefore a necessary condition for a trustworthy leaderboard, not a stylistic preference. Cross-modal synthesis. Every judge converges on the same strategy: exploit a majority-class shortcut wherever the labels concentrate, and pay for it in the mixed-quality middle. T2I expresses this as a mildly low-biased U, T2V and TTS as a high-biased U whose trough migrates toward the low-score side and deepens with the length of the temporal signal, so the amplitude grows from T2I through T2V to TTS. Importantly, the rebound of these curves on high-score segments is a per-dimension average and should not be read as evidence that a judge has localized the few remaining defects; Sec. 5.3 shows that this rebound is largely carried by confirming the many satisfied dimensions. 5.3 Yes/No Perfect Match: Does the Headline Number Localise Defects? Reading the two rows. Per-dimension accuracy averages over all dimensions of a sample, so a judge can post a high mid-to-high score by confirming the many satisfied dimensions while missing the few violated ones. Figure 2 applies a stricter test per score segment: the top row is the Yes-rate, the fraction of samples on which the judge is correct on all ground-truth-Yes dimensions, and the bottom row is the symmetric No-rate on all ground-truth-No dimensions. Because getting every dimension of one polarity right is demanding, the absolute rates sit far below the segment accuracy of Figure 1, so the figure is read for relative contrasts, between judges and between the Yes and No rows, 13 020406080100 Expected Score (%) 0 10 20 30 40 50 60 70 Yes Perfect Match Rate (%) 66.5 30.5 61.1 20.7 31.3 15.4 30.3 11.3 32.3 6.2 27.7 3.6 27.7 2.6 27.7 3.1 18.9 2.5 17.4 0.5 20.4 0.0 13.8 0.5 14.7 0.0 16.8 0.5 12.8 1.5 20.8 1.5 38.8 14.8 T2I Yes Perfect Match Qwen3.5-Omni-Flash (24.3%) Gemini-3.5-Flash (16.3%) Grok-4-Fast-Reasoning (14.6%) Gemini-3.1-Pro (14.3%) Qwen3.5-Omni-Plus (11.3%) GPT-5.2 (10.4%) Gemini-3-Flash (10.3%) GPT-5.4 (10.1%) Claude-Opus-4.7 (9.7%) Claude-Opus-4.6 (9.0%) 020406080100 Expected Score (%) 0 10 20 30 40 50 60 70 80 76.9 17.6 65.1 18.6 68.6 11.6 58.6 12.6 49.4 3.5 60.5 0.0 64.8 8.1 45.4 4.7 41.9 0.0 40.7 2.2 41.4 0.0 41.1 1.2 26.5 0.0 29.8 0.0 27.3 1.2 34.8 1.1 32.0 0.0 16.1 1.1 24.4 0.0 28.1 1.1 43.0 2.3 41.9 10.5 T2V Yes Perfect Match Qwen3-Omni-30B (40.2%) Qwen3-Omni-Flash (17.6%) Gemini-3.5-Flash (17.3%) Gemini-3-Flash (15.4%) Qwen3.6-27B (12.3%) Qwen3.5-Omni-Flash (11.7%) Gemini-3.1-Pro (11.5%) Qwen3.5-Omni-Plus (7.6%) 020406080100 Expected Score (%) 0 20 40 60 80 100 94.5 9.2 86.0 6.4 81.3 9.1 66.6 7.0 65.8 6.7 59.2 8.1 59.6 8.8 41.3 4.4 49.0 4.1 38.2 2.0 35.3 4.1 41.9 6.7 40.6 7.8 57.1 22.2 TTS Yes Perfect Match GPT-Audio-1.5 (58.3%) Gemini-3-Flash (46.6%) Qwen3-Omni-30B (29.7%) Qwen3.5-Omni-Plus (18.3%) Qwen3.5-Omni-Flash (13.3%) Gemini-3.1-Pro (10.9%) 020406080100 Expected Score (%) 0 10 20 30 40 50 60 No Perfect Match Rate (%) 51.8 35.2 60.5 27.2 58.0 16.4 48.2 9.2 48.2 7.7 38.5 1.5 31.3 3.1 30.6 4.6 27.7 2.6 27.0 1.0 29.1 3.1 23.9 3.0 21.4 1.0 32.3 3.1 35.0 1.0 32.1 2.5 58.9 3.5 T2I No Perfect Match Claude-Opus-4.6 (34.3%) Gemini-3-Flash (34.2%) Gemini-3.1-Pro (33.4%) Claude-Opus-4.7 (28.8%) Qwen3.5-Omni-Plus (26.1%) Grok-4-Fast-Reasoning (25.3%) GPT-5.4 (23.8%) GPT-5.2 (22.2%) Gemini-3.5-Flash (21.1%) Qwen3.5-Omni-Flash (7.4%) 020406080100 Expected Score (%) 0 10 20 30 40 50 60 70 69.8 23.3 34.9 5.8 34.5 5.8 26.4 4.6 12.8 0.0 15.1 2.3 15.1 0.0 4.7 0.0 8.0 0.0 4.5 0.0 9.3 0.0 12.8 0.0 11.1 0.0 8.1 1.1 1.1 0.0 7.0 0.0 9.2 0.0 10.5 0.0 4.5 0.0 10.5 0.0 9.3 0.0 10.5 0.0 T2V No Perfect Match Gemini-3.1-Pro (11.1%) Qwen3-Omni-Flash (9.3%) Qwen3.5-Omni-Plus (8.8%) Qwen3.6-27B (8.3%) Gemini-3-Flash (8.0%) Gemini-3.5-Flash (7.1%) Qwen3.5-Omni-Flash (3.7%) Qwen3-Omni-30B (3.1%) 020406080100 Expected Score (%) 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 12.2 0.3 2.6 0.0 1.2 0.0 0.9 0.0 2.0 0.0 1.2 0.0 1.2 0.0 1.8 0.0 2.0 0.0 2.0 0.0 1.8 0.0 1.7 0.0 1.5 0.0 3.2 0.0 TTS No Perfect Match Qwen3.5-Omni-Plus (2.1%) Gemini-3.1-Pro (1.1%) Qwen3.5-Omni-Flash (1.1%) GPT-Audio-1.5 (0.4%) Qwen3-Omni-30B (0.4%) Gemini-3-Flash (0.1%) Figure 2: Per-segment Yes/No perfect-match rates. Top row: Yes perfect-match, correct on all ground-truth-Yes dimensions. Bottom row: No perfect-match, the symmetric quantity on ground- truth-No dimensions. A calibrated judge would show roughly symmetric rows; the visible asymmetry is class bias, and the low No-row values quantify how poorly the few defects are localized. rather than for absolute height. The No row is the strict measure of defect attribution: a high-score sample keeps only one or two No dimensions, and No-rate asks whether the judge catches every one of them. Judges confirm far better than they detect. On each task the peak Yes-rate far exceeds the peak No-rate and no judge is symmetric, and the gap widens sharply as the modality moves from vision to audio: the peak No perfect-match is60.51%on T2I (Gemini-3.1-Pro),69.77%on T2V (Gemini-3.1-Pro) and only12.24%on TTS (Qwen-3.5-Omni-Flash), while the Yes side stays high and reaches94.52%on TTS (Gemini-3-Flash), whose overall TTS Yes rate (46.7%) is more than four times that of its sibling Gemini-3.1-Pro (10.9%), a within-family split that echoes the TTS family collapse of Sec. 5.1. The polarity of the imbalance is itself a personality: on T2I a lenient judge such as Qwen-3.5-Omni-Flash carries the highest Yes but the lowest No perfect-match, whereas strict judges such as Claude-Opus-4.6 and Gemini-3-Flash reach a high No (⌠34%) but a low Yes, so class bias does not even track family membership (on T2I Qwen-3.5-Omni-Flash is Yes-biased, with Yes-rate minus No-rate equal to+16.9while its sibling Qwen-3.5-Omni-Plus is No-biased at â14.8). On TTS the asymmetry is extreme: GPT-Audio-1.5 confirms58.3%of all-satisfied speech yet catches all-violated dimensions on only0.44%of samples, and across the six TTS judges the No perfect-match averages between0.06%and2.15%(the re-run Qwen-3-Omni-30B no exception at 0.42%). The high-score rebound is confirmation, not localization.This is the key complement to Figure 1: the high-score rebound in segment accuracy does not mean a judge has found the defects. Restricting to high-score segments (at least76%of the requirements satisfied, so only one or two No dimensions remain), the mean No perfect-match is44.92%on T2I,23.31%on T2V and merely2.67%on TTS (a peak of12.24%, an all-segment mean below1%). In other words, a TTS judge that posts a strong mid-to-high segment average is, on the very samples where only one or two defects remain, almost never able to name them: its high accuracy is carried entirely by confirming satisfied dimensions. The ordering44.9%on the visual task,23.3%on video,2.7%on speech, is a modality statement: when a judge can no longer rely on visual evidence and has to attribute a defect from temporal or audio cues alone, its localization ability collapses. The capability center of todayâs OmniJudges is still visual. 14 At the all-violated end the gap is combinatorial, and total on temporal tasks. The low-score end makes the mechanism explicit. When the ground-truth total is zero every dimension is a defect, so No perfect-match asks the judge to catch all of them at once. Per-dimension accuracy can still look reassuring here, yet the all-or-nothing rate is far lower: GPT-5.2 reads84.9%per dimension on all-violated T2I samples but scores only4.06%of them fully correct, the mark of an all-or-nothing score that compounds many per-dimension decisions. On T2I the stricter judges still retain a usable rate (Claude-Opus-4.658.9%, Grok-4-Fast35.5%), but on the temporal tasks it collapses to the floor for every judge: on T2V seven of the eight stay at or below1.9%(four at exactly0%, the maximum being Qwen-3-Omni-Flash at10.5%), and on TTS five of the six stay at or below0.6% (peak Qwen-3.5-Omni-Plus3.2%). Two forces compound. The metric runs over the maximum number of No dimensions, so the Yes-bias every judge carries on low-score segments (Sec. 5.2) is almost certain to leak at least one false Yes and void the sample; and a temporal defect cannot be read from a single frame but must be attributed from evidence spread over time or audio, exactly the per-dimension detection that is weakest there and that the coupling map shows is frequently decided by overall impression rather than dimension by dimension (Sec. 5.4). On T2V and TTS a judge therefore almost never labels a wholly defective sample as defective throughout. The asymmetry traces back to the generation-judging pipeline.The Yes-bias reproduces across every judge, both modalities and exactly the segments where the test set forces a No commitment. Real-world generations skew toward acceptable samples; if pre-training and benchmark suites inherit that skew, an all-Yes default is locally optimal almost everywhere. Without low-score samples in a dimension-and-score-balanced benchmark, this entire diagnosis would be unreachable: there would be no No perfect-match to fall toward zero in the first place. Cross-modal synthesis. On T2I and T2V the two rows roughly mirror each other across judges, so a Yes-biased judge is paired with a No-biased counterpart and a balanced evaluator can be assembled from the pool. On TTS the rows are incommensurable: the Yes row approaches perfect prediction while the No row hugs the floor for every judge. The Yes/No gap thus identifies TTS as the most extreme manifestation of the generation-judging asymmetry, and pinpoints class-prior correction, rather than fine-grained perception alone, as the primary data-side intervention the TTS sub-leaderboard demands. 5.4 Judgment Decoupling: Pseudo-Decoupling Diagnosis Reading the matrix. A faithful OmniJudge should decide each near-orthogonal requirement on its own, so we compute the off-diagonal Pearson correlation matrix of each judgeâs per-dimension predictions (Figure 3). The reading is a statement about the judge, not about the dimensions: a high correlation means the judge lets its decision on one dimension drag its decision on another (a halo effect), whereas a low correlation means it decides the two independently. In the figure a coloured pie marks a pair on which one or more judges exceedÏ> 0.6(each slice is one model) and a white cell reports the cross-judge mean correlation when none does; counts below are over unordered dimension pairs (136on T2I,231on T2V,91on TTS). One caveat applies to the white cells: a white, low-correlation cell is only evidence of independent judgment when the dimension is judged well above chance, since a near-chance dimension such as T2V Video Reality (0.42â0.48, Sec. 5.5) produces low correlations by near-constant prediction rather than by genuine independence. Pseudo-decoupling, defined.We term a judge pseudo-decoupled when its macro accuracy remains competitive while its predicted dimension-vs-dimension matrix is dominated by a few high-Ïblocks, indicating that nominally distinct decisions are internally driven by a single latent decision; we treat Ï> 0.8 as strong entanglement. The entangled pairs are semantically adjacent, and shared.Whatever the model, the high-Ïpairs land on semantically adjacent dimensions rather than at random: text-accuracy with text-typesetting on T2I, the audio-content trio on T2V, and the speaker-identity attributes (pitch, timbre, age, gender) on TTS. This is a shared capability limit rather than an idiosyncrasy of one judge. Which pairs a judge fails to separate, however, is model-specific: T2I colour-with-lighting is coloured for Claude, Grok and Qwen but stays white for GPT and Gemini, so the same pair is entangled or not depending on the judge. 15 Composition Color Lighting Emotion Basic Attr. Appearance Behavior Neg. Prompt Spatial Rel. Scene Type Style Text Acc. Text Font Shot Req. Material Edge/Clarity Anatomy Anatomy Edge/Clarity Material Shot Req. Text Font Text Acc. Style Scene Type Spatial Rel. Neg. Prompt Behavior Appearance Basic Attr. Emotion Lighting Color Composition 4851525327502536201311 4853434530474825344520128 51455250305154253648231410 52534650265052283747221310 434546265243243246211413 52505129515049263546251512 53455051275045273547201413 273030262629272924251926251194 4751505129482635221512 4854525250502448223050221310 50434945254848223246261511 25252528242627192622222211128 36343637323535263530323013127 4548474646472550462230231412 202023222125201122222611132335 131214131415149151315121214 11810101312134121011871235 T2I Dimension Correlation (threshold=0.6) Subject Type Subject Count Completeness Appearance Action BG Scene Subj-BG Spatial Inter-Subj Pos. Subj-Camera Shot/Angle Lighting Color Tone Style/Atmos. Audio Content Audio Style Audio Params Video Reality Video Focal Video Contin. Video Rhythm Audio Quality AV Sync AV Sync Audio Quality Video Rhythm Video Contin. Video Focal Video Reality Audio Params Audio Style Audio Content Style/Atmos. Color Tone Lighting Shot/Angle Subj-Camera Inter-Subj Pos. Subj-BG Spatial BG Scene Action Appearance Completeness Subject Count Subject Type 46374040-63341012 3637333535-3325911 45343837-3355911 44353939-35571112 40413539390468810 463645444045313433-33551011 333737-33561112 4545333737-4435910 313535-35581012 293435-2558811 323737-4558911 374145303536-5547910 36-53571111 37333435353133333129323036-423539 403538393934373735343735-533537 403537393933373735353736-424638 -6-3-3-30-3-3-4-3-2-4-5-5-4-5-4710 33354334555532321110 32556553555453341112 45578565888775561720 10991181011910899113937387111117 1211111210111210121111101110101220 T2V Dimension Correlation (threshold=0.6) Text Consist. Punctuation Age Gender Personality Timbre Speed Pitch Tone Emotion Clarity Background Volume Stab. Frequency Frequency Volume Stab. Background Clarity Emotion Tone Pitch Speed Timbre Personality Gender Age Punctuation Text Consist. 25262629232827285525 20192525282527275775 2520404659 2619473345444546 262547477768 29254788711 2328403347475149465677 2825516669 272745496767 282744466767 55447856661927 57657866771916 275467766616 5596811797727 TTS Dimension Correlation (threshold=0.6) Claude-Opus-4.6 Claude-Opus-4.7 GPT-5.2 GPT-5.4 GPT-Audio-1.5 Gemini-3-Flash Gemini-3.1-Pro Gemini-3.5-Flash Grok-4-Fast-Reasoning Qwen3-Omni-30B Qwen3-Omni-Flash Qwen3.5-Omni-Flash Qwen3.5-Omni-Plus Qwen3.6-27B Figure 3: Predicted dimensional correlation. Each cell is a dimension pair: a coloured pie indicates one or more judges exceedÏ = 0.6on that pair (each slice is one model); a white cell shows the mean correlation (Ă100) when no judge exceeds the threshold. A coloured cell reflects the judge coupling its own decisions, not a property of the dimension design. Entanglement density varies sharply across judges.On T2V the entangled-pair count spans more than an order of magnitude within the same task, from5â8for the most decoupled judges to67and74 for the two most entangled (a more-than-tenfold gap,16â18of the latter strong). Notably, this range runs within a family and tracks version iteration: entanglement falls steeply across Qwen generations, from67â74pairs for Qwen-3-Omni to11â23for Qwen-3.5-Omni and only7for Qwen-3.6-27B (on par with any Gemini judge), so decoupling is a judge-level capability that newer releases visibly acquire rather than a property of the vendor. The most entangled judges compress much of the prompt-related half of T2V into a handful of meta-judgments while still reporting⌠70%macro accuracy, the pseudo-decoupling signature. On TTS the most entangled judges are GPT-Audio-1.5 (22pairs,7strong) and the re-run Qwen-3-Omni-30B (19pairs,5strong), and GPT-Audio-1.5 is also the macro-lowest TTS judge; the two Gemini TTS judges, by contrast, stay near-orthogonal (3and7 pairs) yet only reach the middle of the TTS table, so low coupling is necessary but not sufficient for a high rank, and the diagnosis should be read together with the per-dimension competence of Sec. 5.5. On T2I the coupling is mild for all judges (2â11pairs), with the Gemini family the most decoupled (2 each) and Claude, Grok and Qwen-3.5-Omni-Plus the most entangled (11 each). TTS judgeEntangled pairs (Ï>0.6)Macro acc.Macro rank Qwen-3.5-Omni-Plus1065.70%1 Gemini-3.1-Pro360.70%2 Qwen-3.5-Omni-Flash960.44%3 Gemini-3-Flash760.18%4 Qwen-3-Omni-30B1960.12%5 GPT-Audio-1.52258.71%6 Table 2: Entangled-pair count vs. macro accuracy on TTS. The most entangled judge (GPT-Audio-1.5) is also the weakest, but the reverse does not hold: Gemini-3.1-Pro is the most decoupled yet only mid-pack, because it is separately floored on the audio-quality dimensions of Sec. 5.5. Cross-modal synthesis. Each entangled cluster collapses around a single perceptual cue that the judge has substituted for the decoupled decisions: legible text on T2I, âaudio matches the promptâ on T2V, and voice identity on TTS. Whenever a macro number moves without a matching change in the coupling map, the improvement is structural rather than capability-driven. Which polarity carries the coupling. Splitting this matrix by ground-truth polarity (Figure 4) shows that the coupling is not symmetric across the Yes and No decisions. On T2I it is almost entirely a confirmation-side halo: nearly every coupled cell sits on the Yes side, and the only defect-side coupling is confined to the fine-grained modality block (Material, Edge & Clarity, Object Anatomy). On T2V both sides couple heavily, and on TTS the two are balanced. The split also exposes couplings 16 Composition Color Lighting Emotion Basic Attr. Appearance Behavior Neg. Prompt Spatial Rel. Scene Type Style Text Acc. Text Font Shot Req. Material Edge/Clarity Anatomy Anatomy Edge/Clarity Material Shot Req. Text Font Text Acc. Style Scene Type Spatial Rel. Neg. Prompt Behavior Appearance Basic Attr. Emotion Lighting Color Composition 4651515014826191111 464047441447477224120810 51414647165264424911 465117727221113 40414648134512254020811 5147465148154992740231214 5044474848165148122543221117 14141617131516141312121712784 4714724211313 47511341725811 45494812422301416 8757129121274461108 262226272527251724172218687 414440404312618231112 19202422202322721253016233331 11891181211813814108113336 11101113111417413111687123136 T2I Yes-side Coupling (GT=Yes) Subject Type Subject Count Completeness Appearance Action BG Scene Subj-BG Spatial Inter-Subj Pos. Subj-Camera Shot/Angle Lighting Color Tone Style/Atmos. Audio Content Audio Style Audio Params Video Reality Video Focal Video Contin. Video Rhythm Audio Quality AV Sync AV Sync Audio Quality Video Rhythm Video Contin. Video Focal Video Reality Audio Params Audio Style Audio Content Style/Atmos. Color Tone Lighting Shot/Angle Subj-Camera Inter-Subj Pos. Subj-BG Spatial BG Scene Action Appearance Completeness Subject Count Subject Type 233432152019-221728 313828252614212044810513 23362026302026250356710 303225241619187559812 192929271621196561116 23312330192132293341352323212316103 34363243162123-340627 3238252921433828233130182626043559 32381318191659211 29281520229831239 28202429332315232269410411 252627413115202568512510 263035302030301458711 15142016162316181315151520-3245 20212619212321261820232030-2146 19202518192123261922222530-2137 -240762-3019661-3-2-254 243553446898421133 185561035345544336 7106911665912101285671215 2578110252345753312 8131012637911911101143615 T2V Yes-side Coupling (GT=Yes) Text Consist. Punctuation Age Gender Personality Timbre Speed Pitch Tone Emotion Clarity Background Volume Stab. Frequency Frequency Volume Stab. Background Clarity Emotion Tone Pitch Speed Timbre Personality Gender Age Punctuation Text Consist. 1517191913158185142 521213161315146171 155225-233 1723917342-21-1 191239244-241 1913286065 1316221724283431304-132 1513344-433 81534314-351 1814305-240 565246444557 11-2-2-20-1-4-3-254 473146335448 213-115231078 TTS Yes-side Coupling (GT=Yes) Composition Color Lighting Emotion Basic Attr. Appearance Behavior Neg. Prompt Spatial Rel. Scene Type Style Text Acc. Text Font Shot Req. Material Edge/Clarity Anatomy Anatomy Edge/Clarity Material Shot Req. Text Font Text Acc. Style Scene Type Spatial Rel. Neg. Prompt Behavior Appearance Basic Attr. Emotion Lighting Color Composition 621614414-328117-1326835 626261132214-225011 2126364211-3121710-149654 626360711-113163106402 14140139-112623111334 41271316-2834016423 1431111916-3211151714558 -32-3-1-1-2-3-2-11103-132 282121312821-271011112757 11117166311-17180112303 7410324511018149421 -1-2-113011101-1-155 3240117011143835 2659611614312129-13535 8064345-1734-185 31503253502533 51424382731555 T2I No-side Coupling (GT=No) Subject Type Subject Count Completeness Appearance Action BG Scene Subj-BG Spatial Inter-Subj Pos. Subj-Camera Shot/Angle Lighting Color Tone Style/Atmos. Audio Content Audio Style Audio Params Video Reality Video Focal Video Contin. Video Rhythm Audio Quality AV Sync AV Sync Audio Quality Video Rhythm Video Contin. Video Focal Video Reality Audio Params Audio Style Audio Content Style/Atmos. Color Tone Lighting Shot/Angle Subj-Camera Inter-Subj Pos. Subj-BG Spatial BG Scene Action Appearance Completeness Subject Count Subject Type 232533302019-20-2127 1823231520202314 252321172019123333 262426181918223265 1619112264 23182127171921302336 25262128121517-1245108 242121182715181-11356 23262024141923023468 2523212726211822122466 3321181917334434 3021182022123366 232827242616-121345 151718161712151421182216-21-141520 2020201919191518191819-20031518 191918211723221700141520 -221213-110131-1-2-2012 0022102-12232210015 -2233224132431-10115 133223534443343420 2136631056636415151512151520 7435468686465201820 T2V No-side Coupling (GT=No) Text Consist. Punctuation Age Gender Personality Timbre Speed Pitch Tone Emotion Clarity Background Volume Stab. Frequency Frequency Volume Stab. Background Clarity Emotion Tone Pitch Speed Timbre Personality Gender Age Punctuation Text Consist. 1621172116191119-1202 1081617231621171654 1610340248 21840263536-1025 171640382647 2117413669 1623342638414638360487 191646400467 112135380556 19173636402445 -110-12300021628 262066445416 0542468654 248579776528 TTS No-side Coupling (GT=No) Claude-Opus-4.6 Claude-Opus-4.7 GPT-5.2 GPT-5.4 GPT-Audio-1.5 Gemini-3-Flash Gemini-3.1-Pro Gemini-3.5-Flash Grok-4-Fast-Reasoning Qwen3-Omni-30B Qwen3-Omni-Flash Qwen3.5-Omni-Flash Qwen3.5-Omni-Plus Qwen3.6-27B Figure 4: Polarity-conditioned dimensional coupling. Top row: correlation computed only on ground-truth-Yes dimension pairs (confirmation-side halo); bottom row: only on ground-truth-No pairs (defect-side halo). Glyph and threshold follow Figure 3; cells with insufficient same-polarity samples or a near-constant prediction are left undetermined. that the mixed matrix averages away, and judges that look independent overall yet couple on a single polarity (for example Gemini-3-Flash on TTS Audio Clarity with Volume Stability,0.25in the mixed matrix but0.71on the Yes side). The most instructive case is Gemini-3.1-Pro on T2V: the strongest and most balanced video judge overall, it nonetheless couples almost exclusively on the defect side (Completeness with Subject Count rising from0.54in the mixed matrix to0.73on the No side), fusing dimensions specifically when attributing a defect. The mixed matrix of Figure 3 therefore understates the coupling. 5.5 Per-Dimension Competence: the Chance Floor and Capability Gaps Reading the fan.Figure 5 is a53-axis tri-sector radar that reads like a three-bladed fan: each blade is one modality (T2I17dims, T2V22, TTS14), and within a blade the dimensions are sorted by best-model accuracy clockwise from 12 oâclock, so travelling inward along a blade traces each judgeâs capability gradient from its strongest to its weakest dimension. The radius is per-dimension judging accuracy (inner30%to outer100%), so the outer arc collects the dimensions a judge decides reliably and the collapsed inner region the shared blind spots, while the spread between the four contours on one axis is a per-dimension capability gap between judges. We overlay only the four judges common to all three tasks (Gemini-3-Flash, Gemini-3.1-Pro, Qwen-3.5-Omni-Flash, Qwen-3.5-Omni-Plus) so the comparison is cross-modal, and the English and Chinese dials are drawn side by side so a judgeâs cross-lingual stability can be read off directly. Strengths are semantic; blind spots are fine-grained and physical. Every sector is strongest on a semantic or categorical dimension (T2I Scene Type0.899, T2V Subject Count0.955, TTS Audio Background0.765) and weakest on a perceptual, fine-grained one, with T2V showing the widest internal range of the three sectors, from Subject Count (0.955) at the top down to its physical dimensions at the floor. All four judges trace nearly the same blade profile, full at the rim and pinched 17 40% 50% 60% 70% 80% 90% 100% Scene Type Negative Prompt Basic Attributes Color Performance Style & Aesthetic Emotional Expression Text Accuracy Appearance Behavior Shot Requirement Spatial Relationship Composition & Focus Lighting & Atmosphere Text Font/Typeset Edge & Clarity Material & Detail Object Anatomy Subject Count Subject Type Color Tone Lighting Background Scene Style/Atmosphere Shot/Depth/Angle Appearance Subj-Camera Frame Action/Behavior Audio Content Type Subj-BG Spatial Audio Style Audio Params Inter-Subj Position Completeness Audio Quality Audio-Video Sync Video Continuity Video Focal Video Rhythm Video Reality Audio Background Subject Pitch Subject Timbre Subject Emotion Subject Personality Subject Speed Subject Tone Subject Age Audio Frequency Subject Gender Text Consistency Punctuation Audio Clarity Volume Stability EN 40% 50% 60% 70% 80% 90% 100% Basic Attributes Scene Type Negative Prompt Style & Aesthetic Text Accuracy Shot Requirement Appearance Emotional Expression Behavior Color Performance Spatial Relationship Composition & Focus Lighting & Atmosphere Text Font/Typeset Edge & Clarity Material & Detail Object Anatomy Subject Count Color Tone Subject Type Lighting Shot/Depth/Angle Background Scene Appearance Style/Atmosphere Action/Behavior Audio Content Type Subj-Camera Frame Subj-BG Spatial Audio Style Audio Params Inter-Subj Position Completeness Audio Quality Audio-Video Sync Video Continuity Video Reality Video Focal Video Rhythm Audio Background Subject Pitch Subject Speed Subject Timbre Subject Emotion Subject Personality Subject Tone Subject Age Text Consistency Subject Gender Punctuation Audio Clarity Volume Stability Audio Frequency CN Gemini-3-Flash Gemini-3.1-Pro Qwen3.5-Omni-Flash Qwen3.5-Omni-Plus T2I (17 dims) T2V (22 dims) TTS (14 dims) Figure 5:53-dimension tri-sector radar for T2I (17dims), T2V (22dims) and TTS (14dims), EN on the left and CN on the right. Four cross-task judges are overlaid; dimensions are sorted by per-dim best accuracy (strongest at 12 oâclock). The inner region collapsing toward the50%ring marks shared blind spots; the spread between contours on an axis marks a capability gap. at the hub, so the strong-to-weak dimension ordering is shared rather than model-specific; and the whole TTS blade retracts inward (radii0.60â0.66) relative to the comparable T2I and T2V blades (0.71â0.79), marking speech as the hardest modality to judge. The single deepest blind spot is T2V Video Reality, on which the best of the four judges reaches only0.484, below chance, with Video Rhythm (0.498) and Video Focal (0.505) close behind; the T2I floor is Object Anatomy (0.550) and Material & Detail (0.577), and the TTS floor is Volume Stability (0.545) and Audio Clarity (0.607). Several of these lock onto values indistinguishable from a constant prediction, for example Qwen-3.5-Omni-Flash at0.5001on TTS Audio Clarity, the textbook signature of a dimension the judge has not learned to evaluate. The chance floor has two compatible readings, and a caveat.A near-50%score on a dimension can arise either because the output is locked to an entangled companion dimension (Sec. 5.4) or because the judge emits a constant token; both converge on the same conclusion, namely that the fine-grained capability is not activated. This also qualifies the decoupling diagnosis: a dimension can appear decoupled simply because the judge emits a near-constant prediction on it, so a low correlation combined with a near-chance accuracy is spurious independence rather than genuine independent judgment, and the two figures should be read together. A per-polarity accuracy split (Figure 6) further shows the floor to be specifically a near-constant Yes prediction rather than unbiased guessing: on essentially every near-chance dimension the Yes accuracy is near1while the No accuracy is near0 (T2V Video Focal0.99/0.02, TTS Volume Stability0.99/0.03), so on these dimensions the judge is closer to a rubber stamp than to a detector. A judge that lacks an input modality altogether falls onto the same floor: Qwen-3.6-27B, which has no audio channel, sits near chance on all five of its T2V audio dimensions (0.46â0.58), splitting into a constant-Yes default on the intrinsic quality and sync dimensions and near-chance responses on the prompt-alignment ones (Appendix 6). The English and Chinese dials nearly coincide. Because the fan is drawn once per language, a judgeâs cross-lingual stability is legible as the near-coincidence of its two dials: the English and Chinese blades overlap almost exactly, the strong-to-weak ordering along each blade is preserved (Video Reality stays pinned at the T2V hub in both), and the residual difference is a mild Chinese- favoring shift confined to a few audio dimensions rather than a global rotation. Qwen-3.5-Omni-Flash is the most language-stable of the four judges and Gemini-3.1-Pro the most sensitive; Sec. 5.6 quantifies both across all four diagnoses. Aggregating overD p versusD m .Averaging the per-dimension accuracies over the prompt-related and modality-related halves defined in Sec. 3 (T2ID m =D15âD17, T2VD m =D17âD22, TTS D m =D11âD14) is reported in theD p /D m columns of Table 1. On T2I and T2V every judge scores18â35points lower onD m than onD p , with the largest drop reached by Gemini-3-Flash on 18 Composition Color Lighting Emotion Basic Attr. Appearance Behavior Neg. Prompt Spatial Rel. Scene Type Style Text Acc. Text Font Shot Req. Material Edge/Clarity Anatomy 0 20 40 60 80 100 Yes-dim Accuracy (%) T2I Yes-dim Accuracy Qwen3.5-Omni-Flash (77.2%) Gemini-3.5-Flash (73.2%) Gemini-3.1-Pro (70.7%) Gemini-3-Flash (68.2%) Grok-4-Fast-Reasoning (66.8%) GPT-5.4 (66.5%) GPT-5.2 (65.6%) Qwen3.5-Omni-Plus (64.9%) Claude-Opus-4.7 (61.9%) Claude-Opus-4.6 (60.6%) Subject Type Subject Count Completeness Appearance Action BG Scene Subj-BG Spatial Inter-Subj Pos. Subj-Camera Shot/Angle Lighting Color Tone Style/Atmos. Audio Content Audio Style Audio Params Video Reality Video Focal Video Contin. Video Rhythm Audio Quality AV Sync 0 20 40 60 80 100 T2V Yes-dim Accuracy Qwen3-Omni-30B (85.0%) Gemini-3.5-Flash (81.6%) Gemini-3-Flash (80.6%) Gemini-3.1-Pro (77.5%) Qwen3.5-Omni-Flash (77.1%) Qwen3.6-27B (73.2%) Qwen3.5-Omni-Plus (72.3%) Qwen3-Omni-Flash (69.2%) Text Consist. Punctuation Age Gender Personality Timbre Speed Pitch Tone Emotion Clarity Background Volume Stab. Frequency 0 20 40 60 80 100 TTS Yes-dim Accuracy Gemini-3-Flash (88.4%) GPT-Audio-1.5 (86.9%) Qwen3-Omni-30B (79.1%) Qwen3.5-Omni-Flash (74.6%) Qwen3.5-Omni-Plus (73.5%) Gemini-3.1-Pro (73.1%) Composition Color Lighting Emotion Basic Attr. Appearance Behavior Neg. Prompt Spatial Rel. Scene Type Style Text Acc. Text Font Shot Req. Material Edge/Clarity Anatomy 0 20 40 60 80 100 No-dim Accuracy (%) T2I No-dim Accuracy Claude-Opus-4.6 (87.1%) Gemini-3-Flash (86.9%) Gemini-3.1-Pro (86.5%) Grok-4-Fast-Reasoning (85.0%) Claude-Opus-4.7 (84.8%) GPT-5.4 (84.5%) Qwen3.5-Omni-Plus (84.4%) GPT-5.2 (83.2%) Gemini-3.5-Flash (81.9%) Qwen3.5-Omni-Flash (67.6%) Subject Type Subject Count Completeness Appearance Action BG Scene Subj-BG Spatial Inter-Subj Pos. Subj-Camera Shot/Angle Lighting Color Tone Style/Atmos. Audio Content Audio Style Audio Params Video Reality Video Focal Video Contin. Video Rhythm Audio Quality AV Sync 0 20 40 60 80 100 T2V No-dim Accuracy Gemini-3.1-Pro (74.5%) Qwen3-Omni-Flash (73.6%) Qwen3.5-Omni-Plus (72.9%) Gemini-3.5-Flash (72.3%) Gemini-3-Flash (69.8%) Qwen3.6-27B (66.9%) Qwen3.5-Omni-Flash (64.7%) Qwen3-Omni-30B (53.8%) Text Consist. Punctuation Age Gender Personality Timbre Speed Pitch Tone Emotion Clarity Background Volume Stab. Frequency 0 20 40 60 80 100 TTS No-dim Accuracy Qwen3.5-Omni-Plus (57.9%) Gemini-3.1-Pro (48.3%) Qwen3.5-Omni-Flash (46.3%) Qwen3-Omni-30B (41.2%) Gemini-3-Flash (32.0%) GPT-Audio-1.5 (30.5%) Figure 6: Per-dimension Yes/No accuracy for every judge on T2I/T2V/TTS. Top row: Yes accuracy (recall on ground-truth-Yes dimensions); bottom row: No accuracy (specificity on ground-truth-No dimensions). A dimension whose Yes row is high while its No row is near zero is a near-constant âYesâ predictor, which averages to the near-chance radius seen in Figure 5. T2V (84.81%vs.49.52%,â35.29). On TTS the picture inverts the old expectation: the modality gap is smallest not for the Gemini judges but for Qwen-3.5-Omni-Plus (66.33%vs.64.20%, only â2.12), the one judge that has genuinely learned the audio-quality block, while both Gemini judges keep a normalâ10toâ15point gap. The Yes/No split (Figure 6) shows this reflects a genuine ability rather than a fortunate label prior: on the audio-quality dimensions the other judges collapse to a near-constant âcleanâ verdict (No-accuracyâ 0) and rarely flag degraded audio, whereas Qwen- 3.5-Omni-Plus alone stays two-sided where the signal permits (No-accuracy0.85and0.67on two of the four), genuinely separating clean from degraded speech. The macro average collapses these qualitatively different failure modes into one number: without dimension parity, the long tail of easy prompt dimensions would dilute theD m floor and a20â35point cliff would look like a few-point dip. Cross-modal synthesis, and a reading caveat. The radar reduces the per-dimension diagnosis to one rule: T2I and T2V share a bottom-region floor confined toD m and large top-region spreads on D p , so improving these judges is mostly a matter of activating dormant prompt-side capabilities. A radius on this chart, however, is a Yes/No mixed average and cannot by itself tell whether a high value reflects genuine detection or mere confirmation of the many satisfied dimensions; that distinction requires the perfect-match view of Sec. 5.3. Read this way, the moderate TTS radii (0.60â0.66) are especially deceptive. 5.6 Cross-Lingual Robustness (English vs. Chinese) All four diagnoses were run in English and Chinese; only the English figures are shown in the main text, and the Chinese-language counterparts of Figures 1â5 are deferred to the appendix. The picture is language-invariant. The macro leader is the same in both languages on every task (Gemini-3.1-Pro on T2I, Gemini-3.5-Flash on T2V, Qwen-3.5-Omni-Plus on TTS), and per-model accuracy barely moves across languages: on the53-dimension radar the English and Chinese means of any judge differ by at most0.022(the largest being Gemini-3.1-Pro on TTS,0.607vs.0.629), with Chinese marginally higher on the visual tasks. Every structural conclusion survives the switch: Video Reality remains the deepest blind spot in both languages (0.484in English,0.513in Chinese); the No-class collapse on TTS persists (high-score No perfect-match2.67%in English vs.1.64%in Chinese, 19 against44.9%and45.3%on T2I); and the most entangled video judges remain the most entangled in both languages (74 then 65 pairs), with only marginal count changes. The residual language sensitivity is concentrated on a handful of fine-grained audio and temporal dimensions rather than spread across the taxonomy. The largest per-dimension swings are all Chinese- favoring and sit on TTS Audio Background (Gemini-3.1-Pro,0.494to0.639), TTS Audio Frequency (Gemini-3.1-Pro,0.462to0.564) and T2V Completeness (Qwen-3.5-Omni-Plus,0.609to0.741); the Chinese matrices are also marginally more entangled on the same semantically-adjacent pairs, without changing the decoupling structure. D 3 -Omni thus yields the same capability diagnosis in both languages, and the few cross-lingual gaps are themselves diagnostic, isolating a small set of speaker- timbre and audio-frequency dimensions on which a judgeâs competence is language-dependent. 5.7 Synthesis: From Diagnosis to Construction The four diagnostic axes converge on a consistent picture. We summarize four shortcomings and, for each, indicate how it should steer the construction of training and evaluation data (Sec. 4), so that any lab can re-aim the same pipeline at the exact gap diagnosed in its own model. The failure modes are industry-wide, not vendor-specific. Before mapping the shortcomings, we stress that these behaviors recur across every judge and, crucially, across vendors. The U-shaped competence collapse on the mixed-quality middle, the regression-to-mean calibration bias that over- scores low-quality content and under-scores high-quality content (Sec. 5.2), and the temporal-endpoint asymmetry by which a judge confirms all-good content far more readily than it catches all-bad content on T2V and TTS, all hold for the full panel rather than for a single family; the temporal-endpoint asymmetry in particular is positive without exception, for all eight T2V and all five TTS judges in both languages. The clearest sign that these are properties of current OmniJudges as a class is cross-family convergence: on the T2V endpoint asymmetry the Gemini and Qwen families land on nearly identical means (+19.30and+18.91), so the confirm-heavy temporal profile is a shared regularity of todayâs judges rather than one vendorâs implementation artifact. This is what makes the four shortcomings below worth closing generically rather than model by model. One thread across the four figures. The four diagnostics share a single reading protocol. The segment-average accuracy of Figure 1 and the per-dimension radii of Figure 5 are both means over dimensions, so they are inflated whenever a judge merely confirms the many satisfied dimensions of a high-score sample. Whether the few remaining defects are actually caught, and whether each dimension is decided on its own, is revealed only by the No perfect-match of Figure 2 and the decoupling map of Figure 3. Read together, the four figures separate a genuine capability from a confirmation artifact, and the gap between the two widens from vision to speech: on high-score samples the No perfect-match falls from44.9%on T2I to23.3%on T2V and2.7%on TTS, so an aggregate score can stay high precisely where defect attribution has collapsed. A per-polarity decomposition (Figures 4 and 6) sharpens both halves of this thread: the mean-inflation is specifically a near-constant Yes prediction rather than random guessing, and the entanglement splits into a confirmation-side halo (dominant on T2I) and a defect-side halo (dominant on the temporal tasks), so even a judge that looks decoupled overall can fuse dimensions on one polarity alone. (S1) Modality-internal perception is the floor of every judge.The modality-related dimensions cluster near chance on the inner ring of Figure 5, with several locking onto constant-prediction values just above 50% (Sec. 5.5), the deepest being T2V Video Reality at 0.484. (S2) Decisions collapse along modality-specific clusters. Despite competitive macro accuracy, judges fuse nominally orthogonal dimensions into a single bit, and the cluster is task-specific: a text pair on T2I, an audio trio on T2V, a speaker block on TTS. Wherever a strong macro number coexists with a dense coloured block in Figure 3, it is only pseudo-decoupled (Sec. 5.4); on the temporal tasks the most entangled judges accumulate dozens of such pairs. (S3) Class priors dominate, and aggregate accuracy masks poor attribution. The U-shaped curves and the TTS Yes-bias both stem from majority-class shortcuts; more sharply, the high-score rebound in segment accuracy is not defect localization, since the No perfect-match falls from44.9% on T2I to 2.7% on TTS on the same high-score segments (Sec. 5.2, Sec. 5.3). 20 (S4) No OmniJudge is universal.The leaderboard reorders between vision and speech: the Gemini family leads T2I and T2V but Qwen-3.5-Omni-Plus leads TTS and is the only judge to close the audio-quality gap. Cross-modal generalization of judging ability remains open (Sec. 5.1). How the diagnoses steer data construction. Read as requirements on the data rather than as a scoreboard, the four shortcomings point to concrete adjustments in how training and evaluation triplets are built. (S1) The modality floor calls for over-representing negatives that perturb only the modality-related dimensionsD m while holding the prompt fixed, so the data rewards genuine perceptual discrimination instead of prompt-level shortcuts. (S2) The entangled clusters call for mining triplets that violate exactly one member of a highly correlated pair, forcing each dimension to be decided on its own and breaking the halo that inflates macro accuracy. (S3) The class-prior shortcuts call for score-uniform samplingâequal mass at every total scoresâ0,...,Dand a balanced Yes/No count per dimensionâwith extra all-violated, low-score negatives on the temporal tasks, where defect attribution collapses. (S4) The cross-task reordering calls for a joint, task-agnostic balanced schema spanning T2I/T2V/TTS, so that cross-modal judging is trained for rather than assumed. These are precisely the decoupled, dual-balanced, and dynamic construction moves of Sec. 4, now aimed at the specific cells each model fails. Closing the loop.We intend these diagnoses as a constructive starting point for improvement rather than a final assessment of any model. The same dynamic, dual-balanced, decoupled construction can be re-aimed at whichever cells of Figures 1â5 a given judge is weakest on, turning a diagnostic benchmark into a broad recipe for supplementary training data. We also state a limitation plainly: the seed corpus and the pool of negative-construction operators used here are our own and far from exhaustive, so the coverage of any single effortâincluding oursâis necessarily partial. We therefore encourage the community to plug in its own seed pools and its own libraries of negative-construction operators, broadening and enriching the balanced benchmark so that the same method surfaces a more complete picture of each modelâs blind spots. Pooled this way, the individual diagnoses can grow into a genuinely collective multimodal diagnostic lens that keeps pace with Omni-LLMs as they evolve. 6 Conclusion We introduced D 3 -Omni, the first balanced and decoupled benchmark for diagnosing fine-grained multimodal understanding in present-day OmniJudges. Built through the D 3 construction framework (Sec. 4), the benchmark combines Dual-balanced sampling that closes the negative-sample gap and enforces near1 : 1Yes/No parity on every dimension, Decoupled orthogonal-flip operators that fix verified positive seeds and perturb one dimension at a time, and a Dynamic construction loop that steers generation toward the sparsest regions of the label space. The released benchmark spans T2I, T2V and TTS with53near-orthogonal binary dimensions (17/22/14) over10,671samples and a uniform distribution across all total-score levels, while remaining a stable evaluation set. Applied to ten T2I, eight T2V and six TTS judges, the same balanced lens converts an appar- ently saturated leaderboard into four mutually reinforcing diagnoses (Sec. 5.2â5.3, synthesized in Sec. 5.7): (S1) modality-internal perception is the universal floor, withD m macro accuracies floored near50%on every task and an exact0.500xconstant-prediction signature on TTS Audio Clarity; (S2) decisions collapse along modality-specific clusters, with the most entangled video judges compressing much of the prompt half into a few meta-judgmentsâdozens of entangled T2V dimension pairsâwhile still reporting⌠70%macro accuracy; (S3) class priors dominate when the sample regime changes, producing U-shaped or monotonic accuracy curves under which every judge confirms satisfied content far more reliably than it detects violated content, leaving the No class almost unlabeled on TTS (No perfect-match†12.24%across all six judges, and only2.7% even on high-score segments, versus44.9%on T2I); and (S4) no OmniJudge generalizes across modalities, with the TTS podium led by a different family than the visual tasks (Gemini-3.1-Pro 78.56%on T2I but60.70%on TTS, while Qwen-3.5-Omni-Plus leads TTS at65.70%). Crucially, all four recur across model families rather than isolating a single vendor, so they characterize current OmniJudges as a class. Together, these results identify the gap between OmniJudge and OmniBias: aggregate accuracy can be high while the underlying capability distribution is shallow, entangled and class-prior-driven, and only a benchmark balanced on both dimensions and total scores can surface that gap. 21 The same construction pipeline is the natural prescription for the diagnosis it surfaces. The orthogonal- flip operator targets S1 by generating one-hot-D m negatives that share the same prompt as a positive seed, the attainability guarantee (Lemma 1) targets S2 by mining triplets that violate exactly one member of an entangled pair, the score-uniformity constraint targets S3 by removing the empirical reward for all-Yes / all-No shortcuts, and the task-agnostic schema targets S4 by enabling joint balanced fine-tuning across T2I, T2V and TTS without bespoke per-modality pre-processing. D 3 - Omni therefore delivers more than a leaderboard: it is a closed-loop instrument that turns each diagnosed weakness into an actionable training-data prescription, and, since our own seed corpus and pool of negative-construction operators are inevitably partial, we invite the community to plug in its own seed pools and operator libraries, generate dimension- and score-balanced triplets aimed at each modelâs weakest cells of Figures 1â2, and grow these individual diagnoses into a genuinely collective multimodal diagnostic lens. Limitations and future work. Three directions complement the present study. First, the current benchmark covers T2I, T2V and TTS; extending the same balanced decoupled construction to image-to-X, video-to-X and any-to-text judging would test whether the four shortcomings persist in the reverse direction. Second, sample-level explainability that pairs each diagnostic curve with a representative example would make the diagnoses inspectable without re-running every judge; scaling this to a public gallery is a natural next step. Third, the dynamic loop is currently triggered by sparsity in the label distribution; coupling it directly to a target judgeâs most recent error profile would turn D 3 -Omni from a stable evaluation set into an adaptive training-data generator, closing the diagnose-to-construct loop in real time. References [1] Yizhi Li, Ge Zhang, Yinghao Ma, Ruibin Yuan, Kang Zhu, Hangyu Guo, Yiming Liang, Jiaheng Liu, Zekun Wang, Jian Yang, Siwei Wu, et al. Omnibench: Towards the future of universal omni-language models. Advances in Neural Information Processing Systems, 38, 2025. [2] Yiman Zhang, Ziheng Luo, Qiangyu Yan, Wei He, Borui Jiang, Xinghao Chen, and Kai Han. Omnieval: A benchmark for evaluating omni-modal models with visual, auditory, and textual inputs. arXiv preprint arXiv:2506.20960, 2025. [3]Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems, 36:46595â46623, 2023. [4]Nathan Lambert, Valentina Pyatkin, Jacob Morrison, LJ Miranda, Bill Yuchen Lin, Khyathi Chandu, Nouha Dziri, Sachin Kumar, Tom Zick, Yejin Choi, et al. Rewardbench: Evaluating reward models for language modeling. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 1755â1797, 2025. [5] Michihiro Yasunaga, Luke Zettlemoyer, and Marjan Ghazvininejad. Multimodal reward- bench: Holistic evaluation of reward models for vision language models. arXiv preprint arXiv:2502.14191, 2025. [6]Tianyi Xiong, Xiyao Wang, Dong Guo, Qinghao Ye, Haoqi Fan, Quanquan Gu, Heng Huang, and Chunyuan Li. Llava-critic: Learning to evaluate multimodal models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 13618â13628, 2025. [7] Zhuoran Jin, Hongbang Yuan, Kejian Zhu, Jiachun Li, Pengfei Cao, Yubo Chen, Kang Liu, and Jun Zhao. Omni-reward: Towards generalist omni-modal reward modeling with free-form preferences. arXiv preprint arXiv:2510.23451, 2025. [8]Robert Wijaya, Ngoc-Bao Nguyen, and Ngai-Man Cheung. Multimodal preference data synthetic alignment with reward model. arXiv preprint arXiv:2412.17417, 2024. [9]Zeyu Chen, Huanjin Yao, Ziwang Zhao, and Min Yang. Advancing multimodal judge models through a capability-oriented benchmark and mcts-driven data generation. arXiv preprint arXiv:2603.00546, 2026. 22 [10]Kaiyi Huang, Kaiyue Sun, Enze Xie, Zhenguo Li, and Xihui Liu. T2i-compbench: A compre- hensive benchmark for open-world compositional text-to-image generation. In Advances in Neural Information Processing Systems, 2023. [11]Baiqi Li, Zhiqiu Lin, Deepak Pathak, Jiayao Li, Yixin Fei, Kewen Wu, Tiffany Ling, Xide Xia, Pengchuan Zhang, Graham Neubig, and Deva Ramanan. Genai-bench: Evaluating and improving compositional text-to-visual generation. In Synthetic Data for Computer Vision Workshop at CVPR, 2024. [12]Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, et al. Vbench: Comprehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21807â21818, 2024. [13]Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and evaluating large video generation models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 22139â22149, 2024. [14]Jay Zhangjie Wu, Guian Fang, Haoning Wu, Xintao Wang, Yixiao Ge, Xiaodong Cun, David Jun- hao Zhang, Jia-Wei Liu, Yuchao Gu, Rui Liu, et al. Towards a better metric for text-to-video generation. arXiv preprint arXiv:2401.07781, 2024. [15]Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang, Junichi Yamagishi, Yu Tsao, and Hsin-Min Wang. Mosnet: Deep learning-based objective assessment for voice conversion. In Interspeech 2019, pages 1541â1545. ISCA, 2019. [16]Takaaki Saeki, Detai Xin, Wataru Nakata, Tomoki Koriyama, Shinnosuke Takamichi, and Hiroshi Saruwatari. Utmos: Utokyo-sarulab system for voicemos challenge 2022. In Interspeech 2022, pages 4521â4525. ISCA, 2022. [17] Soumi Maiti, Yifan Peng, Takaaki Saeki, and Shinji Watanabe. Speechlmscore: Evaluating speech generation using speech language model. In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1â5. IEEE, 2023. [18]Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. In The Twelfth International Conference on Learning Representations, 2024. [19]Lianghui Zhu, Xinggang Wang, and Xinlong Wang. Judgelm: Fine-tuned large language models are scalable judges. In The Thirteenth International Conference on Learning Representations, 2025. [20]Junlong Li, Shichao Sun, Weizhe Yuan, Run-Ze Fan, Hai Zhao, and Pengfei Liu. Genera- tive judge for evaluating alignment. In The Twelfth International Conference on Learning Representations, 2024. [21]Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Representations, 2024. [22] Maosong Cao, Alexander Lam, Haodong Duan, Hongwei Liu, Songyang Zhang, and Kai Chen. Compassjudger-1: All-in-one judge model helps model evaluation and evolution. arXiv preprint arXiv:2410.16256, 2024. [23] Jiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong, Qinkai Li, Ming Ding, Jie Tang, and Yuxiao Dong. Imagereward: Learning and evaluating human preferences for text-to-image generation. Advances in Neural Information Processing Systems, 36:15903â15935, 2023. [24] Yuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana, Joe Penna, and Omer Levy. Pick-a-pic: An open dataset of user preferences for text-to-image generation. Advances in Neural Information Processing Systems, 36:36652â36663, 2023. 23 [25]Xiaoshi Wu, Yiming Hao, Keqiang Sun, Yixiong Chen, Feng Zhu, Rui Zhao, and Hongsheng Li. Human preference score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv preprint arXiv:2306.09341, 2023. [26]Sixian Zhang, Bohan Wang, Junqiang Wu, Yan Li, Tingting Gao, Di Zhang, and Zhongyuan Wang. Learning multi-dimensional human preference for text-to-image generation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8018â8027, 2024. [27]Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card. arXiv preprint arXiv:2410.21276, 2024. [28] OpenAI. Gpt-5, 2025. URL https://openai.com/gpt-5/. Accessed: 2026-05-21. [29] Gemini Team, Petko Georgiev, Ving Ian Lei, Ryan Burnell, Libin Bai, Anmol Gulati, Garrett Tanzer, Damien Vincent, Zhufeng Pan, Shibo Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024. [30] Google. A new era of intelligence with gemini 3, 2025. URLhttps://blog.google/ products-and-platforms/products/gemini/gemini-3/. Accessed: 2026-05-21. [31]Anthropic.Claude 3.5 sonnet, 2024.URLhttps://w.anthropic.com/news/ claude-3-5-sonnet. Accessed: 2026-05-21. [32] Anthropic. Introducing claude opus 4.6, 2026. URLhttps://w.anthropic.com/news/ claude-opus-4-6. Accessed: 2026-05-21. [33]Jin Xu, Zhifang Guo, Hangrui Hu, Yunfei Chu, Xiong Wang, Jinzheng He, Yuxuan Wang, Xian Shi, Ting He, Xinfa Zhu, et al. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765, 2025. [34]Junbo Cui, Bokai Xu, Chongyi Wang, Tianyu Yu, Weiyue Sun, Yingjing Xu, Tianran Wang, Zhihui He, Wenshuo Ma, Tianchi Cai, et al. Minicpm-o 4.5: Towards real-time full-duplex omni-modal interaction. arXiv preprint arXiv:2604.27393, 2026. [35]Zhifei Xie and Changqiao Wu. Mini-omni2: Towards open-source gpt-4o with vision, speech and duplex capabilities. arXiv preprint arXiv:2410.11190, 2024. [36] Yadong Li, Haoze Sun, Mingan Lin, Tianpeng Li, Guosheng Dong, Tao Zhang, Bowen Ding, Wei Song, Zhenglin Cheng, Yuqi Huo, et al. Baichuan-omni technical report. arXiv preprint arXiv:2410.08565, 2024. [37]Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, and Ion Stoica. Judgebench: A benchmark for evaluating llm-based judges. In The Thirteenth International Conference on Learning Representations, 2025. [38]Guijin Son, Dongkeun Yoon, Juyoung Suk, Javier Aula-Blasco, Mano Aslan, Vu Trong Kim, Shayekh Bin Islam, Jaume Prats-CristiĂ , LucĂa Tormo-Bañuelos, and Seungone Kim. Mm-eval: A multilingual meta-evaluation benchmark for llm-as-a-judge and reward models. arXiv preprint arXiv:2410.17578, 2024. [39]Dongping Chen, Ruoxi Chen, Shilin Zhang, Yaochen Wang, Yinuo Liu, Huichi Zhou, Qihui Zhang, Yao Wan, Pan Zhou, and Lichao Sun. Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark. In Forty-first International Conference on Machine Learning, 2024. [40]Shu Pu, Yaochen Wang, Dongping Chen, Yuhang Chen, Guohao Wang, Qi Qin, Zhongyi Zhang, Zhiyuan Zhang, Zetong Zhou, Shuang Gong, et al. Judge anything: Mllm as a judge across any modality. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V. 2, pages 5742â5753, 2025. 24 A Additional Details Additional details can be placed in the appendix. A.1 Proof of Lemma 1 Proof of Lemma 1.WriteD ⥠D Ï . We give an explicit construction; throughout, a sample is identified with its flip setT â1,...,Dof zeroed dimensions, so its label vector isy d = 1[d /â T ] and its score is s = Dâ|T|. Step 1 (target multiset of flip sizes). Score uniformity(Score uniformity)requires exactlyMsamples at each scoresâ0,...,D, i.e.Mflip sets of sizeDâ sfor eachs. Enumerate theseM (D + 1) required sizes in any order asâ 1 ,...,â M(D+1) , where the valueDâ soccursMtimes for eachs. Their total is T tot = X r â r = M D X s=0 (Dâ s) = M D(D + 1) 2 . Step 2 (cyclic-arc placement). Identify the dimensions withZ D . Assign to ther-th sample the length-â r cyclic arc T r = (a r + k) mod D : k = 0,...,â r â 1 + 1, a r = P q<r â q mod D, so that each arc starts exactly where the previous one ended and the arcs tile the cycle end-to-end. By (A1) pick any seed(p + 0 ,x + 0 ) â S (Ï) full ; by (A2)+(A3) the composite operatorF T r = â dâT r F d realizes the flip set exactly, yielding a sample withy d = 1[d /â T r ]and scoreDâ â r . Collect all M (D + 1) samples intoB. Step 3 (both constraints hold).By Step 1 each scoresis realized exactlyMtimes, so(Score uniformity)holds. For(Dimension parity), note that laying consecutive arcs of total length T tot around a cycle of sizeDcovers each of theDpositions exactlyT tot /Dtimes wheneverD | T tot . HereT tot /D = M (D + 1)/2, which is an integer precisely becauseM (D + 1)is even (the lemmaâs hypothesis). Hence every dimension is flipped (labeled0) exactlyM (D + 1)/2times, so it is labeled 1 exactly M (D + 1)â M (D + 1)/2 = M (D + 1)/2 = N/2 times, which is (Dimension parity). Step 4 (consistency). The two constraints imply the same total positive count, P d N/2 = MD(D+1) 2 = P D s=0 Ms , confirming that they are mutually compatible and jointly satisfied by the explicitB above. A.2 Full Evaluation Dimension Tables The complete lists of evaluation dimensions, with category assignments and per-dimension descrip- tions, are given below for the three generation scenarios used in this work. These tables enumerate the atomic binary requirements that constitute the decoupled taxonomy described in Section 3. 25 CategoryDimensionDescription Prompt-related Composition & FocusDoes the composition of the image match the prompt? Color PerformanceDoes the color palette of the image match the prompt? Lighting & AtmosphereDoes the lighting effectively create the specific atmosphere described in the prompt? Emotional ExpressionDoes the overall emotion conveyed by the image precisely match the emotional requirement specified in the prompt? Basic AttributesDoes the core identity of the subject (including count, category, gender, and other basic attributes) match the prompt? Appearance Do the visual appearance details of the subject (such as color, material, shape, clothing, and hairstyle) match the prompt? BehaviorDoes the subjectâs state and behavior (including facial expression, pos- ture, and action) match the prompt? Negative PromptDoes the image successfully avoid elements explicitly excluded in the prompt (e.g., âno redâ, âno carsâ)? Spatial RelationshipDo the spatial positions, layout, and occlusion relationships among objects or subjects conform to the prompt? Scene TypeDoes the scene type in the image match the one specified in the prompt? Style & Aesthetic AdherenceDoes the image accurately reflect the specific artistic style, design move- ment, or aesthetic system specified in the prompt? Text AccuracyIs any text in the image clear, legible, and free of garbled characters or typos? Text Font/TypesettingDo the font choice and typesetting (e.g., centering, line spacing) of the text match the prompt requirements? Shot RequirementDoes the image accurately reflect the shot size, camera angle, or specific photographic effect (e.g., bokeh, fisheye) specified in the prompt? Modality-related Material & DetailIs the surface material (such as human skin, fabric, or metal) rendered with realistic texture and rich detail? Edge & ClarityAre object contours and edges sharp, clean, and free from unnatural blurring or blending? Object Anatomy Are the structures of subjects (whether objects or people) natural and plausible? For humans, are facial proportions, body anatomy, and hands accurate and anatomically correct? Table 3: Evaluation dimensions for Text-to-Image (T2I) generation. The 17 dimensions are divided into prompt-related dimensions that assess semantic alignment with the input prompt, and modality- related dimensions that evaluate intrinsic image quality. 26 CategoryDimensionDescription Prompt-related Subject TypeDoes the type of main subject in the video match the one specified in the prompt? Subject CountDoes the number of main subjects in the video match the count specified in the prompt? Completeness: Extra/MissingDoes the video contain all required elements and no extra or missing subjects as described in the prompt? Appearance: Color/Shape/ClothingDo the visual attributes of the subject (such as color, shape, size, and clothing) match those specified in the prompt? Action/Behavior/StateDoes the subject perform the correct action, behavior, or state described in the prompt? Background Scene TypeDoes the type of background scene in the video match the one specified in the prompt? Subject-Background Spatial RelationIs the spatial relationship between the subject and the background con- sistent with the prompt? Inter-Subject Relative PositionAre the relative positions among multiple subjects (or key visual ele- ments) consistent with the prompt? Subject-Camera Position/FramingIs the subjectâs position relative to the camera (e.g., framing, distance, orientation) consistent with the prompt? Shot/Depth/Angle/MotionDo the shot parameters (including shot size, depth of field, camera angle, and motion) match those specified in the prompt? Lighting: Direction/Source/QualityDoes the lighting direction, source, and quality (e.g., soft/hard, natu- ral/artificial) match the prompt? Color Tone: Hue/Sat/BrightAre the overall hue, saturation, and brightness of the video appropriate and consistent with the prompt? Video Style/AtmosphereDoes the visual style and overall atmosphere of the video match the one described in the prompt? Audio Content TypeDoes the type of audio content in the video (e.g., speech, music, ambient sound) match the prompt? Audio StyleDoes the stylistic character of the audio (e.g., cheerful, somber, tense) align with the prompt? Audio ParamsDo the audio parameters (such as volume, speaking rate, or instrument choice) match those specified in the prompt? Modality-related Video RealityWhether the video has no unreasonable errors in subject or background, such as distortion, deformation, drift, or jitter? Video FocalWhether the video is accurately focused with clear texture and no unrea- sonable blur? Video ContinuityWhether the video has no frame-to-frame flickering? Video RhythmWhether the action rhythm is reasonable (no sudden speed changes)? Audio Quality Whether the audio in the video is clear with background noise controlled reasonably? Audio-Video SyncWhether the audio and video are synchronized? Table 4: Evaluation dimensions for Text-to-Video (T2V) generation. The 22 dimensions are divided into prompt-related dimensions that assess semantic alignment with the input prompt, and modality- related dimensions that evaluate intrinsic video and audio quality. 27 CategoryDimensionDescription Prompt-related Text ConsistencyDoes the spoken content in the audio match the text specified in the prompt? Punctuation ConsistencyDoes the phrasing and pausing in the audio reflect the punctuation and sentence structure in the prompt? Subject AgeDoes the speakerâs perceived age in the audio match the age specified in the prompt instruction? Subject GenderDoes the speakerâs perceived gender in the audio match the gender specified in the prompt instruction? Subject PersonalityDoes the speakerâs vocal expression convey the personality described in the prompt instruction? Subject TimbreDoes the speakerâs voice timbre align with the timbre described or implied in the prompt instruction? Subject SpeedDoes the speaking rate in the audio match the speed specified or implied in the prompt instruction? Subject Pitch Does the speakerâs vocal pitch match the pitch described or implied in the prompt instruction? Subject ToneDoes the intonation and vocal attitude align with the tone specified in the prompt instruction? Subject EmotionDoes the emotional state expressed in the voice match the emotion specified in the prompt instruction? Modality-related Audio ClarityIs the human voice in the audio clear? Audio BackgroundIs there reverberation or background noise in the audio? Audio Volume StabilityIs the audio volume stable? Audio FrequencyAre there any anomalies in the audio frequency? Table 5: Evaluation dimensions for Text-to-Speech (TTS) generation. The 14 dimensions are divided into prompt-related dimensions that assess alignment with the speaker instruction and text content, and modality-related dimensions that evaluate intrinsic audio quality. A.3 Modality-Side Atomic Operators This subsection lists, by task, the modality-side atomic operatorsO (Ï) d â used in Step 3 of §4.1 to flip a single modality-related dimension while leaving every other dimension intact. Each operator was validated on a held-out audit set to confirm dimensional isolation. T2V atomic operators (6). 1.Subject / background warping, jitter and drift via SAM3-based instance segmentation with adaptive in-painting; 2.Texture-focus dynamic blur via bilateral filtering driven by a dynamic high-frequency-region detector; 3. Frame-order shuffling and inter-frame flicker via randomized frame whitening and re- ordering; 4. Implausible re-timing via 2Ă speed masking on the audio-video stream; 5. Audio-track Gaussian noise injection; 6. Audioâvisual desynchronization via audio-track separation and temporal offset. T2I atomic operators (3). 1. Greasy-texture distortion via frequency-domain filtering; 2. Image blurring via Gaussian blur; 3. Anatomical (limb / joint) deformation via keypoint detection followed by localized smearing. TTS atomic operators (4). 1. Muffled voice via bilateral filtering on the spectrogram; 2. Background reverberation via audio replay plus additive noise; 3. Volume instability via random smoothed gain modulation; 4. High-frequency squeal / low-frequency mud via targeted frequency-domain manipulation. 28 A.4 Evaluation Prompt Templates This subsection reproduces the verbatim English prompt templates used to elicit every per-dimension Yes/No score reported in Section 5. Each prompt is sent to every judge of the corresponding task without modification; the only field substituted at evaluation time is the user-side prompt (and, for TTS, the speaker instruction and the spoken text). A.4.1 Text-to-Image Judging Prompt (17 dimensions) You are a Text-to-Image quality evaluation expert. Please carefully examine the provided image and objectively evaluate the quality of the generated image based on the corresponding image generation prompt. ## Evaluation Task Please evaluate based on the following inputs: - **Image Generation Prompt**: image_generation_prompt - **Image to Evaluate**: The provided image After carefully examining the image, assess each of the following 17 dimensions to determine whether the image content meets the requirements of the generation prompt and quality standards: - Dimension 1 Composition & Focus: Does the composition of the image match the prompt? - Dimension 2 Color Performance: Does the color palette of the image match the prompt? - Dimension 3 Lighting & Atmosphere: Does the lighting effectively create the specific atmosphere described in the prompt? - Dimension 4 Emotional Expression: Does the overall emotion conveyed by the image precisely match the emotional requirement specified in the prompt? - Dimension 5 Basic Attributes: Does the core identity of the subject--including count, category, gender, and other basic attributes--match the prompt? - Dimension 6 Appearance: Do the visual appearance details of the subject--such as color, material, shape, clothing, and hairstyle--match the prompt? - Dimension 7 Behavior: Does the subjectâs state and behavior--including facial expression, posture, and action--match the prompt? - Dimension 8 Negative Prompt: Does the image successfully avoid elements explicitly excluded in the prompt (e.g., "no red", "no cars")? - Dimension 9 Spatial Relationship: Do the spatial positions, layout, and occlusion relationships among objects or subjects conform to the prompt? - Dimension 10 Scene Type: Does the scene type in the image match the one specified in the prompt? - Dimension 11 Style & Aesthetic Adherence: Does the image accurately reflect the specific artistic style, design movement, or aesthetic system specified in the prompt? - Dimension 12 Text Accuracy: Is any text in the image clear, legible, and free of garbled characters or typos? - Dimension 13 Text Font/Typesetting: Do the font choice and typesetting (e.g., centering, line spacing) of the text match the prompt requirements? - Dimension 14 Shot Requirement: Does the image accurately reflect the shot size, camera angle, or specific photographic effect (e.g., bokeh, fisheye) specified in the prompt? - Dimension 15 Material & Detail: Is the surface material--such as human skin, fabric, or metal--rendered with realistic texture and rich detail? - Dimension 16 Edge & Clarity: Are object contours and edges sharp, clean, and free from unnatural blurring or blending? - Dimension 17 Object Anatomy: Are the structures of subjects--whether objects or people--natural and plausible? For humans, are facial proportions, body anatomy, and hands accurate and anatomically correct? ## Output Format Provide a Yes or No answer for each dimension in order, formatted as a JSON array. Ensure the results are arranged in the order of the dimensions: ["Yes", "No", "Yes", ...] (17 elements in total) 29 ## Notes 1. You must strictly return the result in JSON array format, without any additional explanatory text 2. The array must contain exactly 17 elements, each being either "Yes" or "No" 3. Please first examine the image thoroughly, then evaluate each item against the prompt and quality requirements 4. The evaluation must be objective, based on the degree of match between the actual image content and the prompt, as well as the actual image quality A.4.2 Text-to-Video Judging Prompt (22 dimensions) You are a video quality assessment expert. Please carefully watch the provided video and objectively evaluate the quality of the generated video based on the corresponding video generation prompt. ## Evaluation Task Please evaluate based on the following two inputs: - **Video Generation Prompt**: video_generation_prompt - **Video to Evaluate**: The provided video After carefully watching the video, please evaluate the video content against the generation prompt requirements and quality requirements based on the following 22 dimensions: - Dimension 1 Subject Type: Does the type of main subject in the video match the one specified in the prompt? - Dimension 2 Subject Count: Does the number of main subjects in the video match the count specified in the prompt? - Dimension 3 Completeness (Extra/Missing): Does the video contain all required elements and no extra or missing subjects as described in the prompt? - Dimension 4 Appearance (Color/Shape/Clothing): Do the visual attributes of the subject--such as color, shape, size, and clothing--match those specified in the prompt? - Dimension 5 Action/Behavior/State: Does the subject perform the correct action, behavior, or state described in the prompt? - Dimension 6 Background Scene Type: Does the type of background scene in the video match the one specified in the prompt? - Dimension 7 Subject-Background Spatial Relation: Is the spatial relationship between the subject and the background consistent with the prompt? - Dimension 8 Inter-Subject Relative Position: Are the relative positions among multiple subjects (or key visual elements) consistent with the prompt? - Dimension 9 Subject-Camera Position/Framing: Is the subjectâs position relative to the camera (e.g., framing, distance, orientation) consistent with the prompt? - Dimension 10 Shot Parameters (Shot/Depth/Angle/Motion): Do the shot parameters-- including shot size (e.g., close-up, wide), depth of field, camera angle, and motion--match those specified in the prompt? - Dimension 11 Lighting (Direction/Source/Quality): Does the lighting direction, source, and quality (e.g., soft/hard, natural/artificial) match the prompt? - Dimension 12 Color Tone (Hue/Sat/Bright): Are the overall hue, saturation, and brightness of the video appropriate and consistent with the prompt? - Dimension 13 Video Style/Atmosphere: Does the visual style and overall atmosphere of the video match the one described in the prompt? - Dimension 14 Audio Content Type: Does the type of audio content in the video (e.g ., speech, music, ambient sound) match the prompt? - Dimension 15 Audio Style: Does the stylistic character of the audio (e.g., cheerful, somber, tense) align with the prompt? - Dimension 16 Audio Parameters: Do the audio parameters--such as volume, speaking rate, or instrument choice--match those specified in the prompt? - Dimension 17 Video Reality: Whether video has no unreasonable errors in subject or background, such as distortion, deformation, drift, or jitter? - Dimension 18 Video Focal: Whether video is accurately focused with clear texture and no unreasonable blur? - Dimension 19 Video Continuity: Whether video has no frame-to-frame flickering? 30 - Dimension 20 Video Rhythm: Whether action rhythm is reasonable (no sudden speed changes)? - Dimension 21 Audio Quality: Whether audio in video is clear with background noise controlled reasonably? - Dimension 22 Audio-Video Sync: Whether audio and video are synchronized? ## Output Format Please provide Yes or No answers for each dimension in order (22 answers total, only use Yes or No, answer Yes if not violated or not applicable), in JSON array format, ensuring results are arranged in dimension order: ["Yes", "No", "Yes", ...] (22 elements total) ## Notes 1. Must return strictly in JSON array format, do not include any other explanatory text 2. Array must contain exactly 22 elements, each element can only be "Yes" or "No" 3. Please watch the video completely first, then evaluate each item independently against the prompt and quality requirements 4. Evaluation must be objective, based on the match between actual video content and prompt, and actual video quality A.4.3 Text-to-Speech Judging Prompt (14 dimensions) You are a TTS generated voice quality evaluation expert. Please carefully listen to the provided audio, and objectively evaluate the quality of the generated voice audio based on the corresponding audio generation prompts (speaker setting instruction <prompt_instruction> and speaking content <prompt_text>). ## Evaluation Task Please evaluate based on the following inputs: - **Audio Generation Prompt: Speaker Setting Instruction**: <prompt_instruction> - **Audio Generation Prompt: Speaking Content**: <prompt_text> - **Audio to Evaluate**: The provided audio Please carefully listen to the audio, and evaluate whether the audio content and quality meet the requirements of the generation prompts and quality standards across the following 14 dimensions: - Dimension 1 Text Consistency: Does the spoken content in the audio match the text specified in the Audio Generation Prompt: Speaking Content? - Dimension 2 Punctuation Consistency: Does the phrasing and pausing in the audio reflect the punctuation and sentence structure in the Audio Generation Prompt: Speaking Content? - Dimension 3 Subject Age: Does the speakerâs perceived age in the audio match the age specified in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 4 Subject Gender: Does the speakerâs perceived gender in the audio match the gender specified in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 5 Subject Personality: Does the speakerâs vocal expression convey the personality described in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 6 Subject Timbre: Does the speakerâs voice timbre align with the timbre described or implied in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 7 Subject Speed: Does the speaking rate in the audio match the speed specified or implied in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 8 Subject Pitch: Does the speakerâs vocal pitch match the pitch described or implied in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 9 Subject Tone: Does the intonation and vocal attitude align with the tone specified in the Audio Generation Prompt: Speaker Setting Instruction? 31 - Dimension 10 Subject Emotion: Does the emotional state expressed in the voice match the emotion specified in the Audio Generation Prompt: Speaker Setting Instruction? - Dimension 11 Audio Clarity: Is the human voice in the audio clear? - Dimension 12 Audio Background: Is there reverberation or background noise in the audio? - Dimension 13 Audio Volume Stability: Is the audio volume stable? - Dimension 14 Audio Frequency: Are there any anomalies in the audio frequency? ## Output Format Please provide a Yes or No answer for each dimension in order, formatted as a JSON array. Ensure the results are arranged in the order of the dimensions: ["Yes", "No", "Yes", ...] (14 elements in total) ## Notes 1. Must strictly return in JSON array format, without any other explanatory text 2. The array must contain exactly 14 elements, each of which can only be "Yes" or " No". If the audio content does not contain the situation described by the dimension at all, answer "Yes" 3. Please listen to the audio completely first, then evaluate each item against the prompts and the audioâs own quality 4. The evaluation must be objective, based on the degree of matching between the actual audio content and the prompts, and the actual quality of the audio # User Input to Process <prompt_instruction>prompt_instruction</prompt_instruction> <prompt_text>prompt_text</prompt_text> A.5 Polarity-Resolved Judging Diagnostics The two diagnostics below refine the main-text analysis by splitting every per-dimension quantity by the ground-truth polarity, separating a judgeâs ability to confirm satisfied requirements (the Yes side) from its ability to detect violated ones (the No side). Both use the English runs; the Chinese counterparts are qualitatively identical. Per-dimension Yes/No accuracy. Figure 6 reports, for every judge and dimension, the Yes accu- racy (recall on ground-truth-Yes dimensions) and the No accuracy (specificity on ground-truth-No dimensions), a per-dimension decomposition of the perfect-match view of Figure 2 and of the radii of Figure 5. It shows that the near-chance radii of the radar are not unbiased guessing: on essentially every near-50%dimension the two rows are far apart (for example T2V Video Focal at Yes0.99/No 0.02and TTS Volume Stability at0.99/0.03), the signature of a near-constant Yes prediction rather than a random 50/50 split. T2V audio dimensionTypeYes acc.No acc. Audio Content Typeprompt-alignment0.510.61 Audio Styleprompt-alignment0.520.63 Audio Paramsprompt-alignment0.510.62 Audio Qualitymodality-quality0.660.26 Audio-Video Syncmodality-quality0.660.27 Table 6: Qwen-3.6-27B, a judge without an audio channel, on the five T2V audio dimensions, split by ground-truth polarity. On the prompt-alignment dimensions it stays near chance on both rows (it cannot confirm whether the audio matches the prompt), whereas on the intrinsic quality and sync dimensions it defaults to âYesâ (Yes accuracy0.66, No accuracy⌠0.26), rarely flagging a defect it cannot hear. The two behaviors give the same near-chance radius in Figure 5 for opposite reasons. Polarity-conditioned coupling.Figure 4 recomputes the off-diagonal correlation matrix of Figure 3 separately on the dimension pairs whose ground truth is Yes on both axes (the confirmation-side halo) and No on both axes (the defect-side halo), attributing each coupling in the mixed matrix to a polarity. On T2I the coupling is almost entirely a confirmation-side halo (31coupled dimension pairs are 32 Yes-driven against4on the No side), and the only defect-side coupling is confined to the fine-grained modality block (Material, Edge & Clarity, Object Anatomy); on T2V both sides couple heavily (61 Yes-side against77No-side pairs) and on TTS they are balanced (22each). The split also surfaces couplings that the mixed matrix averages away, and judges that look independent overall yet couple on a single polarity (for example Gemini-3-Flash on TTS Audio-Clarity-with-Volume-Stability,0.25 in the mixed matrix but 0.71 on the Yes side). 33