Paper deep dive
A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation
Jianlin Chen, Wenhui Chen, Ziyao Lin, Chi Man Vong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/29/2026, 4:37:48 AM
Summary
This paper formalizes construct validity for LLM-as-a-judge evaluation using a two-dimensional profile: invariance (S), the probability that a verdict remains unchanged under construct-preserving edits, and construct sensitivity (R), the probability that it changes under minimal construct-changing edits. The authors demonstrate that S and R are independent coordinates, meaning high reliability (invariance) does not guarantee validity (sensitivity). Empirical results across 7 judges and 4 domains show that while judges achieve high invariance (S >= 0.90), they exhibit low construct sensitivity (R = 0.319). A significant asymmetry is found where judges are more sensitive to 'scope' edits than 'strength' edits. Additionally, the study reveals that public label sets often leak surface information, allowing simple predictors to reproduce human votes, undermining the validity of standard benchmarks.
Entities (8)
Relation Signals (5)
Invariance (S) â isindependentof â Construct Sensitivity (R)
confidence 95% · We show that S and R are independent and that no scalar summary preserves all relevant comparisons.
Scope Edit â hashighersensitivitythan â Strength Edit
confidence 93% · Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262
LLM-as-a-judge â exhibitshighinvariance â Invariance (S)
confidence 92% · At matched invariance S >= 0.90, judges average S = 0.945
LLM-as-a-judge â exhibitslowsensitivity â Construct Sensitivity (R)
confidence 92% · judges average S = 0.945 but R = 0.319
Public Label Sets â leaksurfaceinformation â Surface Form
confidence 90% · surface-only predictors reproduce 55%-67% of labels in paired mode
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.
Tags
Links
- Source: https://arxiv.org/abs/2608.24419v1
- Canonical: https://arxiv.org/abs/2608.24419v1
Trouble viewing inline? Open PDF directly â
Full Text
125,086 characters extracted from source content.
Expand or collapse full text
A Judge Should Know What Changed: Construct Validity for LLM-as-a-Judge Evaluation Jianlin Chen â thanks: 202330450231@mail.scut.edu.cn Affiliation: South China University of Technology Wenhui Chen â thanks: mc35092@um.edu.mo Affiliation: University of Macau Ziyao Lin â thanks: mc35081@um.edu.mo Affiliation: University of Macau Chi Man Vong â thanks: Corresponding author. cmvong@um.edu.mo Affiliation: University of Macau Abstract An evaluator is a measurement instrument, and an instrument can be reliable without being valid. The judge literature has established the first failure thoroughly: judges are inconsistent, and they move on surface properties (length, format, register, injected metadata) that should not change a verdict. The more basic property is untested. A valid judge must satisfy two conditions, not one: it must not change its verdict under an edit that preserves the construct, and it must change its verdict under a minimal edit that changes the construct. We formalise the pair as a two-dimensional profile (invariance S, construct sensitivity R) and prove that the coordinates are independent (Proposition 1), that R is uninterpretable without S (Proposition 2), and that the profile admits no total order, so a scalar summary necessarily discards a comparison a reader may need (Proposition 3). We then measure it, over 7 judges and 4 domains, against minimal interventions along 7 construct slots and 5 register-only controls, with direction established by human annotators rather than by any model at a pre-registered agreement bar of 0.75, and generation, verification and judging drawn from three disjoint model families. Every profile is read off a frontier we obtain by eliciting a graded verdict and cutting it ourselves, so all judges are compared on identical terms whatever their vendor exposes. At matched invariance Sâ„0.90Sâ„ 0.90 judges sit at S=0.945S=0.945 against R=0.319R=0.319: reliable, and markedly less responsive to construct change than their reliability suggests. The failure is structured, and the structure is the result. Overclaiming decomposes into scope (how many situations a claim covers) and strength (how hard it commits to each), a distinction the nearest works reach and then deliberately unify, theoretically [27] and empirically [25]. We treat the unification as a hypothesis with a stated refutation condition, and it does not survive. Judge sensitivity to the two axes differs by +0.121 at matched invariance, with the same sign on all 7 judges across vendors, parameter scales and reasoning modes; a bootstrap over base items excludes zero, an axis-label permutation null gives p<0.002p<0.002, and the gap is larger on the subset where all 3 annotators agreed. The one length bias we can measure runs against the effect, so +0.121 is a lower bound. Judges register that a claim now covers more and largely miss that it now commits harder. The two axes further respond with opposite sign to accuracy pressure, which supplies a mechanism for the reported paradox that prompting for accuracy increases overgeneralisation [41]: pooling the axes hides it. Why is a gap this basic not in the record? Because the rulers leak. Across 5 public label sets, a frozen family reading only surface form, given the same paired input the judge gets, reproduces 55%â67% of their labels, including 67.4% of MT-Benchâs human votes. RewardBenchâs length control is the sharpest case: constraining chosen responses to be no longer than rejected ones did not remove the confound but inverted it, leaving prefer the shorter response correct on 59.7%59.7\% of pairs. And whether a set leaks turns out to be a property of the set and the input mode together, which is why validation power is defined relative to a mode (Definition 6). High judge agreement is therefore not evidence of valid evaluation. We release the interventions, the human direction labels, the elicitation pilot, the diagnostic, and a checklist. 1 Introduction A judge is an instrument, and the field has spent two years establishing that this instrument is noisy. Judges disagree with humans more than raw agreement suggests, they disagree with themselves across reorderings, and they move on properties that carry no information about quality: response length [60], list and emphasis formatting [59], the order the candidates are shown in [54], injected metadata about the author [33]. Reward models show the same pattern, well enough that benchmarks are now designed against it [29, 30, 22]. Every one of those results tests the same half of the same property. Each asks: the construct did not change, so did the verdict stay put? That is invariance, and failing it makes an instrument unreliable. Passing it does not make an instrument valid. A thermometer welded to a fixed reading is perfectly invariant to everything. The untested half. A valid judge must satisfy two conditions. Under an edit that preserves the construct, the verdict must not change. Under a minimal edit that changes the construct, it must. Prior judge evaluation measures the first and, to our knowledge, never the second. We measure both, as a profile âĄ(J)=(S,R)V(J)=(S,R), and report that judges are far more invariant than they are sensitive. This reframes what a high agreement number licenses. Reliability and validity come apart, and a judge can be stable, unbiased on every surface property yet tested, and still fail to notice when the thing it is supposed to measure has changed. That is not noise. It is the instrument measuring something other than the construct. The failure has structure, and the structure is the finding. We develop the pair on a construct where minimal edits are well posed: whether a scientific claim overreaches its evidence. Overreaching decomposes into two operations that are not the same and have no reason to be handled alike. scope edits enlarge the set of situations the claim is about: widen the population, extend to an adjacent domain, restate one observation as a standing property, strengthen a quantifier. strength edits leave that set alone and raise the commitment: delete the conditions the effect was observed under, remove a hedge, add an intensifier the evidence does not support. The method may improve performance in this basin fails differently from the method improves performance in all basins and from the method substantially and consistently improves performance, and a judge can be blind to one while catching the other. We find that it is. That asymmetry pays a debt. Peters and Chin-Yee [41] report that prompting a model for accuracy makes its overgeneralisation worse. On one axis that is incoherent and reads as an instruction-following quirk. On two it is expected: a demand for accuracy is a demand to sound certain, which narrows scope while amplifying strength, because hedges and stated conditions are precisely what reads as uncertainty. We rerun the manipulation with the axes scored separately. Why a gap this basic is not already in the record. Because the rulers cannot see it. Consider two 2026 findings that look incompatible. A large judge audit reports verbosity bias as negligible, under 0.0110.011 across all 2121 judges, over roughly 541,000541,000 judgments [38]. And a released epistemic-calibration resource reports that on its own judge-calibration set, a rule reading nothing but response length reproduces the labels better than the judge does [6]. Both hold. The leak is not in the judge but in the label set: when a label is inherited from the generative condition that produced an item, and that condition also fixes surface form, the set is solvable without the construct, and agreement with it rewards provenance rather than the construct. Such a set has no headroom in which a sensitivity failure could show up. We measure that headroom. 1.1 What is claimed, and on what evidence âą C1 (§2). Construct validity for an evaluator, formalised as the pair (S,R)(S,R) over two intervention classes, with an argument for why it must stay two-dimensional and what a validity frontier means. Conceptual; the vocabulary is not ours (§3), the application to evaluators is. âą C2 (§4). The measured profile: 7 judges Ă 4 domains, against interventions whose direction is set by human annotators. The headline. âą C3 (§4). The scope/strength asymmetry, and its opposite response to accuracy pressure, which resolves Peters and Chin-Yee [41]âs paradox. Stated as a test of a published unification rather than as a new distinction (§3.1); a null here is evidence for that unification and is reported as such. âą C4 (§4.4, §5.4). Validation power, defined relative to an input mode (Definition 6, Corollary 2), and measured on 5 public sets in both modes. The mode is load-bearing: the same sets read as sound per item and as 55%â67% recoverable paired. Needs no annotation of ours. The per-axis form of the prediction is not testable on these sets and remains open. âą C5 (§5.5). Against a blinded per-sentence replacement set, the inflation Î in judge accuracy from validating on the inherited set. âą C6 (§6.1). Released interventions, human direction labels, diagnostic, and checklist. What is not claimed. The invariance / directional-expectation pair is Ribeiro et al. [45]âs, named for task models; we claim its systematic application to evaluators. Shortcut exploitation is a mature literature on the model side (§3), and our control arm is a replication that exists to make the sensitivity arm interpretable. scope and strength do not exhaust overreach; they are the two that can be edited minimally and verified by humans. The two axes have been noticed before, twice, and both times deliberately collapsed: we claim that the collapse costs something measurable, which is narrower and more checkable than novelty (§3.1). A judge insensitive on strength may be perfectly serviceable where strength is not the construct, and human labels are not ground truth in general [24]. Companion submission, disclosed. A second submission by an overlapping author set characterises the corpus that supplies our probe items and C4âs case study [4]. It adopts this paperâs decomposition and its construct-blind predictor family, and reports the corpus-side measurements; this paper reports the judge-side ones. Neither depends on the otherâs results, and the one quantity that would cross between them, a bound on judge accuracy from a per-sentence gold set, is unrun in both and stated as such in both (§5.5). Anti-circularity. An intervention belongs to the sensitivity arm only if it changes the correct verdict, and if a model decides that, then the experiment tests a model against itself. Direction is therefore set by human annotators, at agreement â„0.75â„ 0.75, and generation, verification, and judging are three disjoint model families. §4.2 gives the protocol and reports what it discarded. 2 Construct validity for an evaluator Figure 1: The paperâs claim structure, in two tiers. The left group is the claim that stands on its own: no external label set enters C1, C2 or C3, and direction is assigned by humans under Protocol 1. The right group is support, and each of its stages carries the consequence of its own null result in the dashed box beneath it. The independence is the figureâs shape rather than an assertion in this caption, and it is why the premise cannot evaporate on one experiment (§1.1). 2.1 The object, and the two conditions The vocabulary is borrowed and we use it in its original sense. Construct validity is the question of whether an instrument measures the theoretical construct it is claimed to measure, as opposed to correlating with it [10]; the modern treatment makes validity a property of an interpretation of scores rather than of the instrument in isolation [35], and the measurement-theoretic reading insists that an instrumentâs validity is a claim about a causal relation between the attribute and the score [2]. Jacobs and Wallach [24] bring this vocabulary into machine learning. What has been missing on the evaluator side is not the vocabulary but an operational test, and the rest of this section is one. Write an evaluator as a function J:Ăâ¶,JâĄ(x,c)=y,J:XĂC , J(x,c)=y, (1) mapping an item x and an evaluation criterion c to a verdict y in a finite, ordered verdict space Y. Fix c throughout: the question is not whether a judge understands the criterion it was given, but how it responds to changes in x under that criterion. Let yââ(x,c)y (x,c) denote the correct verdict, which for the constructs we study is a fact about the relation between the claim and its evidence rather than about any annotatorâs opinion, though establishing it in practice requires annotators, and §4.2 is about that gap. An intervention is a map T:âT:X . Partition the interventions of interest by their effect on the correct verdict, not on the surface of x. Definition 1 (Intervention classes). T is construct-preserving for c, written TâTITâ T_I, if yââ(Tâx,c)=yââ(x,c)y (Tx,c)=y (x,c) for all x in the domain of interest. It is construct-changing, TâTCTâ T_C, if yââ(Tâx,c)â yââ(x,c)y (Tx,c)â y (x,c). Three consequences of defining the classes this way, all of which matter downstream. Membership depends on c: register normalisation is in TIT_I when the criterion is claim calibration and in TCT_C when the criterion is tone. Membership is unrelated to edit magnitude: a two-word quantifier change is in TCT_C while a rewritten paragraph can be in TIT_I, so the classes cannot be approximated by an edit-distance threshold. And membership is not observable from x alone; it is a claim about yây , which is why Protocol 1 exists and why a pipeline that assigns it by model output is circular. Definition 2 (Construct validity of an evaluator). J is construct-valid for c on (TI,TC)(T_I,T_C) if JâĄ(x,c) J(x,c) =JâĄ(Tâx,c) =J(Tx,c) for all âTâTI, all Tâ T_I, (2) JâĄ(x,c) J(x,c) â JâĄ(Tâx,c) â J(Tx,c) for all âTâTC. all Tâ T_C. (3) Figure 2: The two conditions of Definition 2 on one claim. Left: the property the judge literature measures: an edit that changes register, restatement, clause order, length or formatting must not move the verdict. Hedging is deliberately absent from that list, for the reason in Appendix B. Right: the property we could not find measured: an edit that changes what the claim covers or how hard it commits must move it. The right columnâs frame is drawn open because the correct verdict there is set by human annotation (§4.2), not read off the sentence. Equation (2) is the invariance property that judge-robustness work measures under many names: consistency, bias, shortcut sensitivity. Equation (3) is its complement. Definition 2 is not new as a form: it is the invariance / directional-expectation pair of Ribeiro et al. [45], stated for an evaluator rather than for a task model. What is new is measuring the second half for evaluators and finding that it does not follow from the first. Definition 3 (Validity profile). For a judge J, an item distribution D over X, and intervention families TI,TCT_I,T_C, S(J)=PrxâŒ,TâŒTI[J(x)=J(Tx)],R(J)=PrxâŒ,TâŒTC[J(x)â J(Tx)],S(J)\;=\; _x ,\,T T_I\! [J(x)=J(Tx) ], R(J)\;=\; _x ,\,T T_C\! [J(x)â J(Tx) ], (4) and the validity profile is âĄ(J)=(S,R)â[0,1]2V(J)=(S,R)â[0,1]^2. Both are properties of the triple (judge, distribution, family): a profile is reported with its families named or it is not interpretable. 2.2 Why the profile is not a scalar A reviewer will ask for one number. The answer is that no scalar summary of Eq. (4) can distinguish a valid judge from a broken one, and the reason is structural rather than a matter of losing nuance. Proposition 1 (The coordinates are independent). For every (s,r)â[0,1]2(s,r)â[0,1]^2 there exists a judge attaining =(s,r)V=(s,r). In particular neither coordinate constrains the other, and no inequality between them holds in general. Proof. Fix â0,1Y \0,1\. Construct J by independent randomisation on the two families: on an input reached through TâTITâ T_I, return the verdict assigned to the base item with probability s and the other verdict otherwise; on an input reached through TâTCTâ T_C, return the other verdict with probability r and the base verdict otherwise. The two rules act on disjoint sets of (base, intervention) pairs, so the two probabilities in Eq. (4) are set independently and equal s and r respectively. Since (s,r)(s,r) was arbitrary, the whole unit square is attained. â Proposition 1 is elementary, and that is the point: it says the literatureâs decade of invariance results places no bound whatever on R. A field could drive S to 11 across every known surface property and learn nothing about whether its instruments track the construct. Proposition 2 (Both coordinates are individually gameable, in opposite directions). Let Jconstâ(x,c)=y0J_const(x,c)=y_0 for a fixed y0y_0, and let JflipJ_flip return a verdict drawn to differ from the base verdict whenever the input was produced by any intervention. Then âĄ(Jconst)=(1,0)V(J_const)=(1,0) and âĄ(Jflip)=(0,1)V(J_flip)=(0,1). Neither reads x in a way that depends on c. Proof. Immediate from Eq. (4): JconstJ_const satisfies JâĄ(x)=JâĄ(Tâx)J(x)=J(Tx) identically, giving S=1S=1, and never satisfies JâĄ(x)â JâĄ(Tâx)J(x)â J(Tx), giving R=0R=0. Symmetrically for JflipJ_flip. â Corollary 1 (No monotone scalar separates valid from degenerate). Let g:[0,1]2ââg:[0,1]^2 be any function non-decreasing in each argument, and suppose g assigns a passing score to some judge with =(sâ,râ)V=(s ,r ), sâ,râ<1s ,r <1. If g is symmetric (as a weighted mean with λ=12λ= 12, or a harmonic mean, is) then gâĄ(1,0)=gâĄ(0,1)g(1,0)=g(0,1), and any threshold that admits one degenerate judge admits the other. More generally, for λâ(0,1)λâ(0,1) the weighted mean satisfies λâ 1+(1âλ)â 0=λ· 1+(1-λ)· 0=λ, so a threshold below maxâĄ(λ,1âλ) (λ,1-λ) admits a judge that never reads its input. The operational rule we adopt follows directly, and it is a commitment rather than a preference: R is never reported without S measured on the same judge over matched interventions, and a judge whose S falls below a pre-registered floor is reported as uninterpretable rather than as sensitive. Figure 3: The (S,R)(S,R) plane. The two degenerate instruments of Proposition 2 are the marked corners: a judge that never moves sits at (1,0)(1,0), and one that flips on anything sits at (0,1)(0,1). The dashed line joining them is the level set on which a symmetric scalar summary scores them identically (Corollary 1), which is why this paper reports the pair. â±âĄ(J)F(J) traces one judge as Ï varies, and its default Ï, JbJ_b, is the only point usually reported. JaJ_a sits at the same invariance as JbJ_b and in a different quadrant, with the projection to the axis drawn: what invariance testing alone resolves is that one coordinate, which does not separate valid from insensitive. The corners are glossed in Table 1. Table 1: The four corners of the profile, and what each is. The two cells in the top row are the ones that matter: invariance testing alone cannot separate them, because it does not measure the coordinate on which they differ. R high R low S high valid insensitive: reliable, and measuring something else S low unstable: may be flipping indiscriminately invalid 2.3 The frontier, and what a fair comparison is A judge exposing a decision threshold can trade one coordinate for the other, so a single profile is one point on a curve and two judges compared at their defaults are two arbitrary points on two different curves. Definition 4 (Validity frontier). For a judge family JÏ\J_Ï\ indexed by a decision threshold Ï, â±âĄ(J)=(SâĄ(JÏ),RâĄ(JÏ)):Ïâ[0,1]2.F(J)\;=\; \\, (S(J_Ï),\,R(J_Ï) )\;:\;Ï\, \\;â\;[0,1]^2. (5) Proposition 3 (Threshold comparisons are not validity comparisons). There exist judges J(1),J(2)J^(1),J^(2) and thresholds Ï1,Ï2 _1, _2 such that RâĄ(JÏ1(1))>RâĄ(JÏ2(2))R(J^(1)_ _1)>R(J^(2)_ _2) while â±âĄ(J(2))F(J^(2)) dominates â±âĄ(J(1))F(J^(1)) pointwise; that is, the judge with the higher reported sensitivity is the strictly worse instrument. Proof. Take â±âĄ(J(1))=(1âr,r)F(J^(1))=\(1-r,\,r)\ and â±âĄ(J(2))=(1âr+Ï”,r)F(J^(2))=\(1-r+Δ,\,r)\ for râ[0,1âÏ”]râ[0,1-Δ], so J(2)J^(2) attains strictly greater S at every R. Choosing Ï1 _1 with r=0.9r=0.9 and Ï2 _2 with r=0.1r=0.1 gives RâĄ(JÏ1(1))=0.9>0.1R(J^(1)_ _1)=0.9>0.1, while J(2)J^(2) dominates. â Hence: judges are compared by frontier, or at a matched S, and never by raw R at default thresholds. This is the same discipline as reporting an ROC curve rather than one accuracy, for the same reason. Remark 1 (Getting a frontier out of a judge that has no visible parameters). A judge behind an API exposes no decision threshold, no logits one can rely on across vendors, and no internal state. It would seem to follow that most judges admit only a single point, and that Definition 4 is unavailable for them. It does not follow, and the fix is a design choice rather than an approximation: elicit a graded verdict and threshold it ourselves. A judge asked for a calibration score on a fixed scale, together with a rule mapping scores to verdicts, is a family JÏ\J_Ï\, where Ï is the cut we apply to its own output and control exactly. Every judge then has a frontier on identical terms, whatever is or is not visible inside it. Two things this buys beyond convenience. Comparisons become uniform, so no row of Table 5 is a weaker claim than another for reasons of vendor access. And the threshold is ours, so it cannot be tuned per judge to flatter a result: one rule, applied to all, stated in Appendix D. What it costs is that the graded scale is part of the instrument now: a judge that grades coarsely has a coarser frontier, and Appendix H S3 reports how much of the measured insensitivity is attributable to scale granularity rather than to the judge. We therefore do not manufacture a frontier by varying a prompt and present it as one; we obtain one by asking for a score and cutting it. A judge that refuses to grade, or grades degenerately, is reported as a single point with that fact stated. Remark 2 (Relation to reliability, and why fixing one does not fix the other). Testâretest consistency and human agreement are properties of J at fixed input; S extends them to perturbed input. By Proposition 1 none of the three constrains R. So the finding that judges are unreliable [38] and the finding that they are insensitive are independent facts about the same instrument, and an intervention that raises agreement (ensembling, self-consistency, rubric refinement [20]) has no predicted effect on R at all. §4 reports whether that prediction holds. 3 Related work We read the nearby literature for what each result had to assume rather than for what it found, because a result can be correct and replicated while resting on an assumption that fails at the limit, and the assumption is what a new paper can move. Three recur, and this paper tests all three. That reliability is the binding constraint on judge quality, assumed by the invariance literature and false by Proposition 1, which shows the two coordinates place no bound on each other. That overclaiming is one thing, assumed by the works closest to ours, which reach the scope/strength distinction and then unify it deliberately, theoretically [27] and empirically [25]; we treat the unification as a hypothesis with a refutation condition. That agreement with a validated label set is evidence of validity, assumed nearly everywhere and the subject of §5.5, where a construct-blind family reproduces 55%â67% of five public setsâ labels. Four lines of work assume nothing this paper tests, and are cited as ancestry rather than as contrast. Behavioural testing supplies the invariance versus directional expectation distinction we adopt [45]; judge evaluation took up the first and not the second, so what we claim is the application to evaluators, not the distinction. Validity position papers make the call we answer [12]. NLI artifacts and contrast sets are the methodological ancestor, benchmarks solvable without doing the task [21, 42, 16]; the contribution here is the instrument, the decomposition and the audit rather than the idea. Measurement modelling supplies the vocabulary [24], to which we add an operational test. Two more are contrasts that fit in a clause. Automatic reliability stress-testing assumes the benchmarkâs labels are sound [11], which is exactly what our harness tests, so the two compose rather than compete. And positional bias of faithfulness locates its defect at predictable positions in the output [52], where mis-scoping has no position at all: it is a relation between a sentence and its evidence. Table 2 states each remaining assumption against the result that carries it; the right-hand column is where we agree with a result and disagree with what it is taken to license. Table 2: What the nearby results establish, and the assumption each one needs that we test. work establishes assumes, and we test Judge reliability audits [38] judges are unreliable, and exact-match agreement overstates discrimination that the label sets are ground truth. Judges are measured against fixed sets, never audited Mechanistic accounts of judge bias [55] surface manipulations move judge scores, visibly in activations only specificity: edits that should not flip a verdict. Sensitivity is untested, and no bias type there is scope or quantifier Noise-corrected evaluation [14] TPR/FPR estimated on a calibration set can be corrected for that the calibration set is sound. The guarantees inherit whatever it leaks Judge input-shortcut probes [33] judges follow injected metadata cues and never acknowledge them that the shortcut is in reading the input. Perturbs inputs, not the construct, and audits no label set Automatic factual-consistency metrics [57] claim support can be scored by one alignment function that the axis of interest is support. A mis-scoped claim is not unsupported, so such a metric is silent by construction, not by weakness Surface and style confounds, in judges and in reward models [60, 54, 59, 22, 29, 32, 30, 7, 56, 53] verbosity, order, formatting and length shift verdicts, and preference labels are recoverable from them, well enough that benchmarks are now designed against it that the property should not have mattered, and that the confound is model-side. Every one is an invariance test; the label-set side has no headroom metric, no cross-set audit, no reporting convention Preference optimisation [43, 47] a preference signal distils into a policy, and a proxy reward is gameable in the limit that the signal means what it says. What a judge cannot distinguish is what the policy may exploit, so R is a property of the training signal, not only of the report 3.1 The decomposition, seen twice and collapsed twice The claim that nobody separates scope from strength would be false. Two published works reach both mechanisms and unify them deliberately, which is a stronger position than a gap claim: a collapse has authors who chose it and a cost that can be measured. Collapsed theoretically. Jiang et al. [27] reinterpret linguistic calibration as answer-set prediction, under which widening what a claim covers and hedging how hard it commits are the same operation: both enlarge the set of possible worlds in which the claim holds. The unification is what buys the conformal guarantee, since a single nested family of answer sets is what a calibrated set predictor needs, and for generation under a coverage guarantee it is the right modelling choice. Our claim is about evaluation, where §2.2âs argument applies to their collapse as much as to a scalar V: a judge that misses a hedge deletion and catches a quantifier widening has a property one answer-set coordinate cannot express. Collapsed empirically, in the nearest workâs own appendix. James et al. [25] score overstatement as one continuous value, but their appendix decomposes linguistic certainty into extent, number, framing and probability, and reports that increased overstatement is driven by greater use of probability and extent certainty. Those are our two axes under other names, found in their data and discarded by the headline metric. It is the strongest external support the decomposition has: two mechanisms were seen driving one score, and nobody asked whether an evaluator responds to them differently. Observed together outside our domain. Studying whether a chatbot should assert generics about social groups, Zhu [61] reports that models inconsistently hedge and decline generalisations reflecting documented patterns. Generic formation is the scope operation and hedging the strength one, and a model turning both inconsistently is not what a single unstructured dial produces. The evidence is convergent from a domain with no stake in our construct. This makes C3 a test with a named opposing prediction rather than a gap-filling exercise: had scope and strength sensitivity not differed at matched S, that would have been evidence for the collapse. On the sensitivity arm, the same correction. Applying edits intended to change a verdict is not unprecedented. Park et al. [40] build minimally edited counterfactual responses isolating perceptual errors, catching a judge that fails to downgrade them. Three things separate it from §4: it is multimodal and about visual-versus-textual conflict; the intended direction is established by construction rather than by human adjudication, which is the circularity Protocol 1 closes; and its object is mitigation via reward modelling rather than measurement, so it reports no invariance arm and therefore no profile. The claim we keep is the pair, not the direction: S and R together, for text constructs, with direction set by people. The framing is converging. Roy et al. [46] treat judge rubrics as measurement specifications and audit them for structural adequacy, reliability, preference fit and adversarial robustness, concluding that no source is simultaneously reliable, preference-predictive and robust, and that high inter-rater agreement does not prevent exploitability. That is this paperâs premise reached from the rubric side by different authors. Where it stops is where §4.4 starts: PReMISE audits the rubric and the judge, not the label set the judge is scored against, and Definition 6 is the quantity for that set. Table 2 is the positioning; three remarks it cannot carry follow. The wider ancestry. That a benchmark can be solvable without doing the task is a general phenomenon: shortcut learning in deep networks [17], syntactic heuristics passing inference benchmarks [34], annotator identity recoverable from labels [18], counterfactually augmented data proposed against such correlations [28], the benchmarking critique that a leaderboard number can be uninformative about the capability it names [3, 44], and a formal treatment of reward hacking [47]. Our contribution is not to observe that this can happen to evaluators but to give the test that says whether it has. The nearest work. Jiang et al. [27] formalise a factuality/specificity trade-off and rewrite claims to trade one against the other. The distinction is the direction of the question: they ask how a generator should set the specificity of what it writes, we ask whether an evaluator notices when specificity has been changed for it. A calibration result about generation places no constraint on evaluator sensitivity. Ou et al. [39] is adjacent in material rather than in question, asking how well grounded LLM critiques of papers are. Which constructs are at risk, and an adjacent negative result. The constructs most exposed are those whose labels are expensive and whose surface correlates are cheap. Claim specificity is the case we develop, because a resource exists whose labels are known to be length-separable [6] and because the construct is well posed [27, 41]. Probing work on a neighbouring construct, the strength of evidence supporting a clinical claim, recovered a graded label linearly in every model tested and then found the recoverable signal was largely lexical and did not transfer across topics [48]. That is our shape of finding from the representation side, and it is why §5.4 reports transfer rather than only within-set recovery. 4 Method: the interventions, the direction protocol, and the diagnostic This section fixes the instrument before any result is read off it: how the two intervention classes of Definition 1 are realised as concrete edits, who decides that an edit belongs in the sensitivity arm, which judges and domains the profile is measured over, and (§4.4) the second instrument the paper needs, which measures not a judge but the label set a judge is validated against. Every choice here is frozen in the repository before the first judge call, and §4.2 is the one a reader should push hardest on. 4.1 Instantiating the two intervention classes The construct is whether a claim overreaches its evidence. TCT_C splits into two axes, fixed in probes/edit_taxonomy.py before any judge is run: scope (referent set grows, commitment fixed): population widening, domain extension, tense/modality shift (one observation restated as a standing property) and quantifier strengthening. strength (referent set fixed, commitment rises): condition removal (deleting the circumstances the effect was observed under) hedge removal, and intensifier addition. Both axes are linguistic distinctions before they are machine-learning ones. Hedging is a studied feature of scientific writing with a documented rhetorical function [23], and uncertainty together with its scope has been annotated at corpus scale in biomedical text [50], which is direct evidence that the two things we separate are separately annotatable by humans. The adjacent verification literature asks whether a claim is supported [49, 51]; our construct begins where that one ends, since every item in the TCT_C arm is supported on both sides of the pair and differs only in what it covers or how hard it commits. And strength is a manipulable quantity on the model side, not only a rhetorical one, which matters because an axis no model represents is an axis no judge could plausibly be asked to track. Linguistic calibration (making an agentâs expressed confidence match its competence) has been trained for directly [36]; expressed uncertainty has since been located as a linear feature in representation space and steered along it [26]; and Tayebi Arasteh [48] recover a graded evidence-strength label linearly in every model tested. Our claim is not that models cannot represent commitment. It is the sharper and more awkward one: commitment is representable, steerable and linearly decodable, and judges still do not act on it. The scope axis has the parallel grounding on the generation side, where decontextualising a scientific snippet (lifting it off the source without widening what it covers) is a defined task with gold data [37], and over-generalisation under that pressure is measured at scale [41]. TIT_I is the control arm, and it deliberately covers the surface properties the judge literature has already shown to move verdicts, because a low flip rate on properties nobody has implicated proves nothing: hedge padding, same-scope elaboration, register shift, clause reordering, verbosity padding [60], and format shift [59]. Figure 4: The decomposition on one worked claim. From a single licensed sentence, scope moves right (the referent set grows while commitment is untouched) and strength moves up: the referent set is fixed and commitment rises. The seven TCT_C edit types sit beside the axis each moves, and in each cell the changed span is coloured with the axis that moved it; every cell is one edit from the origin. The diagonal is the case the verifier rejects: an item that moved both axes cannot separate them, and admitting it would contaminate the very contrast §5.2 rests on. Remark 3 (Why these two axes and not one). The axes are separable in the sense measurement requires: an edit on one can be applied while holding the otherâs operators fixed, and a verifier can check which moved. They are not claimed to exhaust overreach. Splitting them is a modelling choice that earns its place only if judges behave differently across it, which is a result, reported below, not an assumption. 4.2 Who decides that the verdict should change An intervention is in TCT_C only if it changes the correct verdict, which makes this the paperâs load-bearing methodology. If a language model adjudicates that, we are testing models against models and any asymmetry could be a shared blind spot rather than a property of judges. Protocol 1 (Direction assignment). Three disjoint model families, and humans hold the deciding vote. 1. Generate the candidate edit (family A). 2. Verify mechanically (family B, never A): facts unchanged; the intended axis moved and the other axisâs operators untouched; no provenance anchor introduced; register comparable; length within tolerance. 3. Assign direction by human annotation. Annotators see the pair without knowing which member is the edit, which axis it belongs to, or what any model said, and answer only whether the second sentence claims more than the evidence licenses. An item enters the sensitivity arm only at agreement â„0.75â„ 0.75 (pre-set; HUMAN_DIRECTION_AGREEMENT_MIN). 4. Judge (family C, never A or B). Two reporting rules follow. Per-slot yield appears in the same table as the result, because if strength edits are harder to construct cleanly than scope ones then the surviving strength items passed a stricter filter and part of any asymmetry is the pipeline rather than the judges. And humanâverifier disagreement is reported rather than resolved silently, since it bounds how far the mechanical gate can be trusted where humans were not asked. The pool and the pre-set bar of 0.75 are specified in Appendix C, frozen before any data was collected, together with the rule attached to it: a slot or axis below the bar is reported as not measurable and excluded from R rather than reported with a wider interval, because a wide interval around an ill-defined construct still asserts the construct exists. What the annotators found. 3 annotators independently labelled all 478 items; the third worked the full batch rather than arbitrating the first twoâs disagreements, so the three label sets are independent. Raw pairwise agreement is 0.852 (Fleissâ Îș 0.776, the three pairwise rates spanning 0.814â0.877), and every axis clears the bar: scope 0.898 (n=144n=144), strength 0.889 (n=144n=144), invariance 0.789 (n=190n=190). Only 2 items (0.4%) drew three different answers. The two construct axes are equally well defined for human readers. scope is nominally higher, but by less than the spread between annotator pairs, so we claim no ordering; what the bar was set to rule out, that one of the two distinctions is not one people reliably make, happened on neither. Of the annotated items, 246 of 288 construct items carry a majority direction matching the intended one (116 scope, 130 strength) and 96 of 190 invariance items survive; Table 3 gives the per-slot breakdown, with agreement and yield in separate columns because they fail differently. tense_modality_shift is the lowest-yield slot at 57.8% on ordinary agreement: the readers agree, and what they agree on is that the edit often left the claim where it was. That is a generator yield problem rather than a construct problem, so the slot stays in scope with its yield stated. Table 3: Human direction annotation, per slot. agree is mean pairwise agreement over three annotators; usable counts items whose majority label matches the interventionâs intended direction. The two columns answer different questions and a slot can fail either one alone: tense_modality_shift has ordinary agreement and low yield â readers agreeing that the edit changed nothing â while the two length-manipulating invariance slots have ordinary yield and agreement below the pre-set bar of 0.75. slot axis n agree usable note domain_extension scope 25 0.973 25 population_widening scope 42 0.857 34 quantifier_strengthening scope 32 0.958 31 tense_modality_shift scope 45 0.852 26 low yield condition_removal strength 35 0.867 33 hedging_removal strength 58 0.851 48 intensifier_addition strength 51 0.948 49 elaboration invariance 33 0.404 0 below bar; excluded format_shift invariance 20 1.000 20 register_shift invariance 55 0.976 54 reordering invariance 23 0.971 22 verbosity_padding invariance 59 0.689 0 below bar; excluded Two invariance slots did not survive. The two that manipulate length, verbosity_padding and elaboration, fall below the bar at 0.689 and 0.404, while the other 3 invariance slots agree at â„0.971â„ 0.971. They are excluded under the pre-registered rule. Two separate problems were stacked there, and only the annotators could have separated them. The first was ours: the generator padded sentences with as is often observed, in practice and as is well known, which read as filler but each assert something the evidence never stated. On the first batch 67.4% of 46 items drew a majority directional verdict. Three revisions of the generation prompt failed to stop it, and what worked was a fixed marker list checked programmatically before the verifier sees the pair (epistemic_marker_change): a prompt states an intention, a filter enforces a property. Regenerating the same slots under the gate, with the same annotators, took the rate to 10.9% of 46. The second is untouched by the gate. Agreement barely moved (0.601 to 0.572) as the contamination cleared, because the disagreement changed hands: afterwards one annotator still called an appositive gloss directional (32 of 46) while the other two almost never did (3), and that annotator agreed with the others at 0.877 on the main batch. Whether naming an entity more precisely changes what a sentence claims is a question our codebook never answered, and one competent reader in three answers it the other way. We report it as a boundary of the invariance construct: surface-preserving is well defined for formatting, register and word order, and not for appositive elaboration (§6.2). 4.3 Judges and domains The design floor is six judge families spanning vendors, parameter scales, and reasoning-enabled versus not, over at least four evaluation domains. That floor is not negotiable downward for a reason of kind rather than of power: a single-domain result is a result about that domain, and this paperâs claim is about the instrument. Restricting the study to one construct in one domain would make the finding a fact about geoscience claim-scope rather than about evaluators, so the profile is measured across domains whose criteria differ in kind, with the scientific-claim domain as the one where minimal edits are best posed and human verification is cheapest. Table 4: Validity profiles. S from the TIT_I control arm, R from TCT_C split by axis. Per-slot n and human confirmation are part of the result, not appendix material: §4.2 explains why. n counts items whose majority human label matched the interventionâs intent and whose slot met the agreement bar: the items that actually enter the arm. human dir. is that count as a fraction of items annotated, so a low value marks a slot the generator often failed to move rather than one humans could not read: tense modality shift is the clearest case, at 57.8% with ordinary agreement. â below the pre-set bar of 0.75 and therefore excluded (§4.2). Pipeline discard before annotation was 24.2% overall, dominated by the length tolerance; it is reported as an aggregate because build_pairs.py keys its rejections by reason rather than by slot, and a per-slot figure we did not record is not one we will estimate. intervention axis n human dir. agreement flip rate TCT_C â should flip: scope domain extension scope 25 100% 0.973 0.445 population widening scope 34 81% 0.857 0.299 quantifier strengthening scope 31 97% 0.958 0.445 tense modality shift scope 26 58% 0.852 0.358 TCT_C â should flip: strength condition removal strength 33 94% 0.867 0.379 hedging removal strength 48 83% 0.851 0.210 intensifier addition strength 49 96% 0.948 0.235 TIT_I â should not flip: surface controls elaboration control 0 â 0.404â â format shift control 20 100% 1.000 0.029 register shift control 54 98% 0.976 0.053 reordering control 22 96% 0.971 0.084 verbosity padding control 0 â 0.689â â =(S,R)V=(S,R), mean over judges (0.945, 0.319)(0.945,\ 0.319) 4.4 The diagnostic: what a label set could ever have detected Definition 5 (Provenance-inherited label). Let each item xix_i be produced under a generative condition CiC_i, which side of a preference pair it occupied, which model emitted it, which prompt variant produced it. A label set â=(xi,yi)L=\(x_i,y_i)\ is provenance-inherited if yi=gâĄ(Ci)y_i=g(C_i) for some g, rather than the result of a reading of xix_i that is blind to CiC_i. Definition 5 is about how the label was obtained, not about whether it is correct. A provenance-inherited label can be perfectly correct and still be useless for validation, and that is the point: correctness and informativeness come apart here. The condition is not exotic. It is what happens by default whenever a label set is assembled from a generation pipeline rather than from independent reading, which is now the common case: surveys of synthetic-data curation treat the generating condition as a first-class part of the artifact [31], and the metrics proposed for certifying generated data are about quality and trustworthiness of the content rather than about whether the label is reachable without the construct [58]. Validation power is that missing quantity. It is also why the diagnostic is cheap: CiC_i is usually recorded somewhere in the pipeline, so a setâs authors can compute VPVP on their own artifact before releasing it, and the sixth item of §6.1 asks only that they release the field so others can. Definition 6 (Validation power, relative to an input mode). Let âłM be an input mode: the structure in which an instrument is shown an item. Two modes matter here. In per-item mode the instrument sees one response and returns a label. In paired mode it sees two responses to the same prompt and returns which one the label set prefers. For a label set âL and a mode âłM, let aimpâ(â,âł)a_imp(L;M) be the agreement with âL achieved by the best predictor in a fixed family of impoverished predictors operating in mode âłM, and let ahumâ(â,âł)a_hum(L;M) be the human ceiling in that mode. Then VPâĄ(â,âł)=ahumâ(â,âł)âaimpâ(â,âł).VP(L;M)\;=\;a_hum(L;M)-a_imp(L;M). VPVP is undefined until the mode is fixed. A per-item predictor can exploit only an absolute signature (responses over this length are preferred); a paired predictor exploits a relative one (whichever of these two is shorter is preferred). Where a setâs surface confound is relative, the per-item estimate understates recoverability by an amount nothing bounds in advance: MT-Bench gives aimp=0.30a_imp=0.30 per item and 67.4% paired. The mode reported must be the mode the instrument is used in. Corollary 2 (The two modes are not orderable in general). Neither mode dominates. A set whose label is a function of an absolute property is recoverable per item and may be near chance paired, if both sides of every pair share that property. A set whose label is a function of a relative property is the reverse. Hence a single reported aimpa_imp without its mode is uninterpretable, and reporting the lower of the two is not conservative; it is simply the answer to whichever question was not asked. Validation power is the discriminative room a judge has to earn. Where VPâ0VPâ 0, a judgeâs agreement with âL is uninformative at any magnitude, because an instrument that reads only surface form already reaches the ceiling. Where VPVP is large, agreement is evidence. Reporting VPVP rather than agreement is the single change we ask for. The impoverished family, frozen. Three predictors, each denied the construct by construction: 1. Length only. Token count of the item and nothing else. 2. Surface only. Character n-grams, punctuation and formatting statistics, and hedge-marker counts, no content words carrying the construct. 3. Provenance-recoverable. A predictor of CiC_i from xix_i. If C is recoverable and y=gâĄ(C)y=g(C), then y is reachable without the construct, and this predictor bounds how much of âL any surface-reading instrument can obtain. What makes it a test rather than a statistic. Three requirements, all of which the released harness enforces. Agreement is chance-corrected, so VPVP is not inflated by class imbalance. Each aimpa_imp is reported against a label-permutation null, so a small positive VPVP is distinguishable from noise. And ahuma_hum carries its own interval, because a ceiling estimated from few double-annotated items can be the larger source of uncertainty in Definition 6. Remark 4 (Why not simply audit the judge harder). A judge audit answers âdoes this instrument behave consistentlyâ. It cannot answer âdoes agreement with this label set mean anythingâ, because that question is about the label set. The two are independent: §1âs apparent contradiction is exactly a well-behaved judge meeting a leaky ruler. 5 Experiments: the profile, the asymmetry, and what the rulers could have seen Five measurements, in the order they depend on each other. The profile (§5.1) is the headline and needs only §4âs instrument. The axis decomposition (§5.2) splits it, and the accuracy-prompting manipulation (§5.3) is the splitâs out-of-sample test, since it predicts a published resultâs sign structure rather than fitting our own. The validation-power sweep (§5.4) then asks why the first three findings are not already in the record, and §5.5 prices the answer. Every measured number below is generated from results/ rather than typed into the manuscript, and the two arms that remain unrun are named as such where they would appear. 5.1 Judges are far more invariant than sensitive 000.10.10.20.20.30.30.40.40.50.50.60.60.70.70.80.80.90.911000.20.20.40.40.60.60.80.811S (invariance)R (sensitivity)haikunemotrongeminimistraldeepseekglmllama4 Figure 5: Validity profiles for 7 judges. Frontiers are drawn for the judges that expose a decision threshold, which under Remark 1 is all of them that grade; any judge that will not grade is a single labelled point, per Proposition 3. The dashed diagonal is S+R=1S+R=1, where a judge is trading one property for the other rather than possessing both. Curves are the monotone hull over our threshold sweep: at each invariance we keep the best sensitivity achieved at that invariance or stricter, so a curve is a frontier rather than a trace of every threshold. Every vertex is an attained threshold, so a long straight segment is interpolation between two of them and not a measured path; the ringed vertex is the matched-invariance operating point Table 5 reads off. Table 5: Per-judge profiles. R is reported per axis and never without S (Corollary 1); the final column records whether the row was read off a frontier or a default threshold, because the two are not comparable. â predicted, not measured judge Ï S R RscopeR scope RstrengthR strength pts regrade anthropic/claude-haiku-4.5 2.75 0.917 0.561 0.629 0.500 9 0.41 nvidia/nemotron-3-nano-30b-a3b 0.25 0.906 0.447 0.474 0.423 7 0.83 google/gemini-2.5-flash-lite 0.25 0.968 0.408 0.474 0.349 6 1.00 mistralai/mistral-small-3.2-24b-instruct 2.25 0.958 0.309 0.405 0.223 4 0.96 deepseek/deepseek-v4-flash 3.25 0.906 0.245 0.362 0.140 9 0.68 z-ai/glm-4.7-flash 0.25 0.979 0.168 0.239 0.107 2 0.98 meta-llama/llama-4-scout 2.25 0.979 0.093 0.095 0.092 6 0.97 mean over the 7 judges compared 0.383 0.262 scopeâstrength scope- strength gap: positive on 7/7 judges, +0.003 to +0.222 Figure 5 shows the vertical gap the study was built to look for, and it is large. Held at matched invariance Sâ„0.90Sâ„ 0.90, the 7 judges average S=0.945S=0.945 against R=0.319R=0.319: the insensitive cell of Table 1, not the valid one. This is not an artifact of a weak field. The strongest instrument in the sweep, anthropic/claude-haiku-4.5, reaches R=0.561R=0.561 and so still misses more than two construct changes in five while holding its verdict on 0.945 of edits that should not move it. The comparison that decides Remark 2. If ensembling and self-consistency raise S without moving R, then the fieldâs standard reliability remedies do not touch validity, and Proposition 1 has an empirical instance rather than only a proof. We run those two remedies as additional judge rows for this reason. 5.2 The scope/strength asymmetry Table 6: Sensitivity per intervention, at each judgeâs matched-invariance threshold, averaged over the 7 judges. The magnitude column is the median |Î|| | length the operation mechanically carries: it is a property of the edit type, not a free variable, which is why magnitude cannot be stratified independently of slot. intervention axis n R median |Î|| | length quantifier_strengthening scope 31 0.445 3.2% domain_extension scope 25 0.445 8.8% tense_modality_shift scope 26 0.358 4.8% population_widening scope 34 0.299 7.0% condition_removal strength 33 0.379 13.0% intensifier_addition strength 49 0.235 4.3% hedging_removal strength 48 0.210 4.2% haikunemotrongeminimistraldeepseekglmllama4mean000.20.20.40.40.60.6flip rate at matched SSTCT_C scopeTCT_C strengthTIT_I control (must be low) Figure 6: Sensitivity per intervention axis, per judge, at matched S, with the control-arm flip rate on the same scale. The control bar is the reference that makes the other two interpretable: a construct-edit bar near the control bar is a judge responding to a construct change no more than to an edit that should not move it at all. Judge labels are shortened vendor names; per-judge thresholds and instrument quality are in Table 5. The pattern holds. Compared at matched invariance, Rscope=0.383R scope=0.383 against Rstrength=0.262R strength=0.262: a gap of +0.121 at a control flip rate of 0.055. The gap has the same sign on 7 of the 7 judges, spanning vendors, parameter scales and reasoning-enabled or not, with per-judge values in Table 5. Judges detect that a claim now covers more and largely miss that it now commits harder. The sign, the magnitude and the condition that would have refuted them were all registered before measurement (Appendix F): the gap was to be reported as refuted if it came in below 0.100.10 or failed a paired bootstrap over base items. It clears both. Three further checks, all specified before the numbers existed, hold. A bootstrap resampling base items rather than pairs (several pairs share a base sentence across slots, so resampling pairs would treat those as independent draws) puts the gap at +0.052+0.052 to +0.189+0.189 over 243 items, excluding zero. Permuting the axis label within the construct arm while holding every judgeâs grades fixed gives p<0.002p<0.002, with a largest null gap of +0.123. And on the 217 items where all 3 annotators agreed on direction, the pre-committed robustness split, the gap is +0.125: larger than on the full set, which is the direction that makes it a fact about judges rather than about annotation difficulty. The length controls, and why the gap is a lower bound. Edit magnitude is the alternative explanation with the most force: strength edits are shorter on average than scope ones, so a judge merely less sensitive to small edits would reproduce the asymmetry with no axis structure. Three checks close it, and the third turns the confound into support. The magnitude match survives human filtering. Items enter the arm only on human confirmation, and confirmation rates differ by slot, so the match built at construction had to be re-verified on what actually entered: median |Î|| | length is 5.0% on scope against 5.3% on strength. The signed change does differ, with scope edits lengthening (+4.1% median, 64% longer) and strength edits shortening (-2.5%, 39%). So the second check asks whether grades track length at all: across every graded sentence, the correlation between a sentenceâs length and its grade is +0.025 to +0.114 over the 7 judges. It is positive, meaning judges grade longer sentences as slightly better scoped. Since the scope edit is the longer member, that bias pushes judges toward naming the original on scope items and the edit on strength ones. It works against the observed gap, which makes +0.121 a lower bound rather than an inflated estimate. Third, the gap survives stratifying on the sign: +0.149 where the edit lengthened (6 of 7 judges) and +0.073 where it shortened (6 of 7). Per slot, which is where the result is most useful. Table 6 gives R for each intervention. The ordering is more informative than the axis means: quantifier_strengthening and domain_extension are the most detected at 0.445, and hedging_removal the least at 0.210, a factor of more than two between two edits that a reader would call equally clear cases of overreach. The magnitude column shows why magnitude cannot be stratified independently: operation fixes it, since a deleted condition clause cannot be a small edit and an added intensifier cannot be a large one. The slot distributions overlap across the axes: condition_removal is a strength slot detected at 0.379, above 2 of the four scope slots. The axis gap is therefore a difference between the central tendencies of two heterogeneous groups rather than a categorical separation, and the registered decomposition accounts for η2=0.40η^2=0.40 of the variance in slot-level R; §6.2 records an alternative grouping that fits comparably. 5.3 The accuracy-prompting paradox If overreach were one quantity, an instruction to be accurate could not increase it. Under Definition 1 it can: a demand for accuracy is a demand to sound certain, which suppresses scope while inducing strength, because hedges and stated conditions are exactly what reads as uncertainty. We rerun the manipulation with the axes scored separately, on the generator side, and then ask whether judgesâ per-axis sensitivity predicts which of the two a model will exploit under that pressure. Not measured here. The two per-axis effects. The registered prediction is that they have opposite signs, with the refutation condition in Appendix F: if both axes move the same way, the decomposition does not explain the paradox. The manipulation is generator-side, so it needs fresh generation and a per-axis readout that does not yet exist. Declared overlap. The control arm replicates established results: Xu et al. [55] on surface manipulation, Marioriyad et al. [33] on injected metadata cues, Zhang et al. [59] on format, Zheng et al. [60] on verbosity, and Huang et al. [22] on length in reward models. We claim novelty for none of it. It is here because R is uninterpretable without S (Proposition 2), which makes the replication a measurement requirement rather than a contribution. What is new is the TCT_C arm, its decomposition, the human-set direction, and the profile. We note that none of the seven bias types in Xu et al. [55] is scope or quantifier structure. 5.4 How widespread it is: validation power of the public label sets Table 7 answers the objection raised in §1: if judges really miss strength change, why has no one reported it. It requires no annotation of ours and runs on public artifacts at pinned revisions. The per-item sweep returned aimpa_imp between 0.07 and 0.30, which reads as healthy headroom. In the mode these sets are actually used in, a predictor reading nothing but surface form reproduces 55%â67% of their labels. Both columns appear in Table 7 because the difference between them is the finding. What the two modes measure. Definition 6 carries the mode explicitly and Corollary 2 says why neither dominates. RewardBench pairs are built so the two sides differ within a pair, but across the set no absolute length band marks the chosen side, so a per-item predictor is near chance and a paired one is not. Our corpus is the opposite, with its preferred side in a narrow absolute band, so both modes succeed. Reporting the per-item number for a pairwise benchmark answers a question nobody asks of that benchmark. Table 7: Construct-blind recovery of the label sets the literature validates judges against, at the revisions pinned in results/labelsets/PROVENANCE.json, in both input modes of Definition 6. Per-item aimpa_imp is chance-corrected agreement; paired accuracy has chance at 50%50\% with side order randomised. Group-aware folds throughout; every paired row clears its label-permutation null, which sits at or below 0.026. HelpSteer2 has no paired column because it ships per-response ratings rather than comparisons. The final row is a provenance-inherited set audited under the identical protocol. per-item mode paired mode label set label origin n aimpa_imp pairs Îș accuracy RewardBench pair role 5,970 0.303 2,985 0.249 62.4% RewardBench 2 pair role 8,977 0.200 5,926 0.248 62.4% RM-Bench pair role 7,962 0.221 3,981 0.095 54.7% HelpSteer2 human rating 8,008 0.068 â â â MT-Bench (human) human vote 6,710 0.300 2,575 0.348 67.4% provenance-inherited case pair role 10,156 0.993 5,078 0.993 99.6% The design assumption that fails. RewardBench constrains chosen responses to be no longer than rejected ones [29], on the implicit assumption that removing the naive direction of the length signal removes the length confound. Measured on the released set, the chosen response is the longer one in 40.3%40.3\% of pairs, so the content-free rule prefer the shorter response is correct on 59.7%59.7\%. The confound was not removed; its sign was flipped and its magnitude preserved, and a reward model that learned to prefer brevity scores above chance on the benchmark built to detect it. A control defined on a property of the items is not a control on what a predictor can do with that property. The one set that holds up, and the one that does not. RM-Bench is closest to chance in paired mode at 54.7%. It is built by crossing every pair with three deliberate style levels, putting style variation inside the benchmark rather than leaving it outside as a confound; this row is what makes the others interpretable, and the design is the one we would ask others to copy. MT-Bench is the opposite. Its labels are human pairwise votes, and a surface-only paired predictor reproduces 67.4% of them, the highest of the public sets and higher than either pair-role benchmark. Definition 5 predicts the reverse ordering, since labels further from the generative condition should be harder to reach without the construct. Either the definition is incomplete or human preference on this task is itself substantially a length preference; we cannot separate the readings here and do not report the ordering as support for the definition (§6.2). The contrast that sets the scale. Run the identical family, folds and nulls against a set whose labels are a function of the generative condition, and the paired predictor is correct on 99.6% of pairs against 55%â67% for the public sets [5]. So Definition 5 does describe a real and severe failure mode. What has changed is that the public benchmarks are no longer the clean comparison they appeared to be in the per-item sweep: they sit well above chance, not at it. What is still untested, stated plainly because neither mode settles it. C4âs prediction was per axis: that strength-axis validation power would be near zero even where aggregate power is healthy. That prediction is untested in either mode and cannot be tested on these sets, because splitting VPVP by axis requires the pairs in a row set to differ by an identifiable scope or strength edit and none of the five is labelled that way. The released intervention set is, so the per-axis form is answerable on it and we leave it to future work. The reversal. Under per-item scoring the human-labelled sets look cleanest, which is what Definition 5 predicts; under paired scoring, the mode judges are actually used in, MT-Bench moves from the cleanest set to the most recoverable. That reversal is Corollary 2 with data attached: the two modes are not orderable, and reporting whichever is lower is not the conservative choice it appears to be. Transfer. Within-set recovery is necessary but not sufficient [48]: a predictor that recovers a set may be exploiting a set-specific idiom rather than a general surface correlate. We therefore fit each impoverished predictor on one set and evaluate it on the others. Of 20 ordered set pairs, 2 admit the question at all: the remaining 18 have label spaces that do not share two categories, so a predictor fitted on one cannot be scored on the other even in principle. That is a finding about the state of public preference labels rather than about our predictors, since these sets are routinely discussed as though they measured the same thing. On the 2 pairs where transfer is defined, the surface predictor carries Îș=0.157Îș=0.157 and 0.1090.109, far below its within-set recovery (Table 7). The recovery is therefore substantially a set-specific idiom, which is what makes within-set recovery necessary but not sufficient. 5.5 What it costs: a blinded replacement, and Î The construct is claim specificity: whether a sentence asserts more than its evidence licenses. We choose it because a provenance-inherited label set for it already exists and is known to leak [6], so the comparison is against a real, deployed ruler rather than a straw one. Figure 7: Why a provenance-inherited label set is separable without reasoning about the construct, and what the blinded design withholds. Left: labels attach to whole records by their role in the pair, so one threshold on response length reproduces them, at the agreement reported by Chen et al. [6]. Right: the unit is one sentence with its evidence context, every field that identifies its side withheld, sides interleaved so pair membership is unrecoverable, and the label the majority of three independent annotators with no arbitration pass. Both panels read in two bands: above the rule, what the annotator sees and how its label is formed; below it, the surface route the design leaves open or closes. A class chip is drawn identically wherever it appears, so the two classes a pair role can supply compare directly against the four. Design. Four per-sentence classes (clean, anchored, over-scoped, unsupported) rather than the two a pair role can supply. Every annotation unit is one sentence with its evidence context, presented with no indication of which side of a pair it came from, no indication of the producing system, sentences from both sides interleaved and globally shuffled so pair membership is unrecoverable, and the length cue broken by presenting sentences individually rather than whole responses. Three independent annotators, no arbitration pass and no annotator seeing anotherâs labels, the label taken as the majority of the three, with Fleissâ Îș and raw pairwise agreement reported per class [8, 15]. Figure 7 contrasts the two designs. The full protocol, codebook, and sampling frame are in Appendix C; the frozen specification is docs/TMLR_EXPERIMENT_PLAN.md C3. The headline would be the difference. A per-class precision/recall table against the blinded set alone would be a fact about one corpusâs labels. Î (how much validating on the inherited set overstates accuracy) is the quantity that transfers, and it is the number this design exists to produce. Not measured here. Î itself, and VPVP of both sets on identical terms. What blocks it is neither compute nor access but a second human annotation round the size of the one reported above, over a four-class codebook. The design, the protocol and the sampling frame are given in full (Appendix C) so that a reader can run it against their own label set: the design is the part that transfers. 6 Discussion and limitations 6.1 What a judge author can do on Monday 1. State how each label was obtained. If it is a function of a generative condition, say so. 2. Report the best impoverished predictorâs agreement, not only the judgeâs. 3. Report a human ceiling with an interval, and report validation power against it. 4. Report a label-permutation null. 5. Probe in both directions: construct edits and surface edits. 6. If the label set is released, release the provenance field, so others can run 1â5. 6.2 Limitations R is a lower bound and S an upper bound. Everything rests on a TCT_C edit really changing the correct verdict; Protocol 1 puts that decision with humans, and an item a majority of the 3 misread enters the arm as a false positive that lowers measured R. A judge may also register that a claim changed and still return the same verdict because the verdict space is too coarse to express it. In the other direction, S over 3 surviving control types bounds invariance from above, since a richer control family can only find more failures. Both bounds point the same way as the headline, which is why §5.2 reads +0.121 as conservative. A third construct axis is visible at the edge of the control arm. 2 of the 5 specified control types fell below the agreement bar, and the reason is substantive rather than noise: one competent reader in three holds that making a referent more precise changes what a sentence claims of it. If that reader is right, adding detail is a third axis alongside extension and precision, and it is a well-posed target for the next study. The disagreement is localised, since the remaining control types agree at â„0.971â„ 0.971, so it is elaboration specifically and not surface-preservation in general that is contested. The decomposition is a modelling choice, and not the only one the data supports. It accounts for η2=0.40η^2=0.40 of the variance in slot-level R; a post-hoc regrouping by whether an edit changes an explicit scope element or a modality-and-degree marker accounts for 0.50. We keep the registered split and disclose the alternative, since distinguishing them needs a construct designed to separate the two. The frontier is elicited, and the elicitation is part of the instrument. Cutting a judgeâs own grade makes every row of Table 5 comparable on identical terms, and it also makes the frontier a property of the judge under our elicitation: a judge clustering its grades on two values yields fewer distinguishable operating points than a finer readout would show, so Table 5 carries the realised operating-point count and regrade agreement beside each profile. The diagnosticâs family and mode are choices. VPVP is defined against a fixed predictor family, so every value is an upper bound on soundness: a richer family could only raise aimpa_imp. Definition 6 covers two modes because those are the two we needed, and Corollary 2 says a third modeâs aimpa_imp is not deducible from ours. One prediction of Definition 5 does not hold on our own data: labels further from the generative condition should be harder to reach without the construct, yet MT-Benchâs human votes are the most recoverable set in paired mode. Either the definition is incomplete or human preference on open-ended answers is itself substantially a length preference, and we cannot separate the two here. Scope, and provenance. The profile is measured on the domains of §4 and generalises no further; constructs whose minimal edits are not well posed lie outside what this method can measure. The paradox result is correlational, making the decomposition a sufficient explanation of Peters and Chin-Yee [41]âs finding rather than a mechanism. The resource whose label set we audit was constructed by an overlapping author set, so we lead with a result about judges and use no field absent from the public release, in particular no internally retained pair-type label. Every number is against judge versions pinned in Appendix D. Artifact availability. The diagnostic harness, the blinded label set, the edit taxonomy with its acceptance gates, and the frozen judge versions are released. The corpus that supplies probe items and C4âs case study is public [5]; it is material here, cited, and this paper would stand with it replaced by any resource carrying per-sentence scope annotations. 7 Conclusion The shortest true summary. An evaluator is a measurement instrument, and validity needs two properties rather than one: a verdict must hold still under an edit that preserves the construct, and move under a minimal edit that changes it. The judge literature has measured the first thoroughly and the second, to our knowledge, not at all. Measured as a profile over 7 judges and 4 domains with direction set by 3 humans at agreement 0.852, judges sit at S=0.945S=0.945 against R=0.319R=0.319: reliable, and markedly less responsive to construct change than their reliability suggests. The failure is structured rather than random. Sensitivity differs by +0.121 between scope and strength at matched S, in that direction on every one of the 7 judges and against the only length bias we can measure, so the gap is a lower bound; and the two axes respond with opposite sign to accuracy pressure, which is why prompting a generator to be accurate can make its overgeneralisation worse [41] without anything being incoherent. The reason none of this is in the record is that the rulers leak: across 5 public label sets, a predictor reading only surface form, given the same paired input the judge gets, reproduces 55%â67% of their labels, including 67.4% of MT-Benchâs human votes. Whether a set leaks is a property not of the set alone but of the set and the input mode together, which is why Definition 6 carries the mode. Three statements are worth carrying away separately from the result, because each is a claim about practice. A high agreement number is not evidence of valid evaluation; it is evidence of invariance, which Proposition 2 shows a constant function attains perfectly. An intervention that raises agreement has no predicted effect on sensitivity; Proposition 1 says the coordinates are independent, and §5.1 runs ensembling and self-consistency as judge rows against exactly that prediction. And a threshold comparison is not a validity comparison: Proposition 3 shows a judge dominated everywhere on its frontier can still win at default thresholds, which is why every row of Table 5 is read off a frontier we cut ourselves rather than off whatever operating point a vendor shipped. What we ask for is small, and it is not a new benchmark. Report (S,R)(S,R) where a judge paper today reports agreement; report the best impoverished predictorâs agreement beside the judgeâs; and say how each label in the validation set was obtained, since Definition 5 is a fact about provenance that costs one sentence to disclose and cannot be recovered afterwards. §6.1 is the six-line version. The open question we would most like answered is whether strength blindness is a property of this construct or of judges in general. That needs a second construct with well-posed minimal edits and cheap human verification, and the method transfers unchanged: the interventions, the direction protocol and the frozen predictor family are released for it. The per-axis form of the validation-power claim is a second piece of future work: it cannot be asked of any public row set, since none is labelled by edit axis, and the intervention set we release is. References [1] Y. Benjamini and Y. Hochberg (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. Journal of the Royal Statistical Society: Series B 57 (1), p. 289â300. Cited by: §A.2. [2] D. Borsboom (2005) Measuring the mind: conceptual issues in contemporary psychometrics. Cambridge University Press. Cited by: §2.1. [3] S. R. Bowman and G. E. Dahl (2021) What will it take to fix benchmarking in natural language understanding?. In Proceedings of NAACL-HLT, p. 4843â4855. Cited by: §3.1. [4] J. Chen, W. Chen, Z. Lin, and C. M. Vong (2026) TriDeHall-ecp: what auditing a million-pair scope-calibration preference corpus reveals. Note: Companion submission. Identifier added on posting. Cited by: §1.1. [5] J. Chen, W. Chen, and W. Liu (2026) TriDeHall-ecp. Note: https://huggingface.co/datasets/Airjiannan/TriDeHall-ECPDataset Cited by: §5.4, §6.2. [6] J. Chen, W. Chen, and W. Liu (2026) TriDeHall: a unified tri-level hallucination mitigation framework for multimodal cross-lingual scientific knowledge agents. In Proceedings of the 15th CCF International Conference on Natural Language Processing and Chinese Computing (NLPCC), Lecture Notes in Computer Science. Note: Oral presentation External Links: Link Cited by: §1, §3.1, Figure 7, §5.5. [7] L. Chen, C. Zhu, D. Soselia, J. Chen, T. Zhou, T. Goldstein, H. Huang, M. Shoeybi, and B. Catanzaro (2024) ODIN: disentangled reward mitigates hacking in RLHF. External Links: 2402.07319 Cited by: Table 2. [8] J. Cohen (1960) A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), p. 37â46. Cited by: §A.3, §5.5. [9] J. Cohen (1968) Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit. Psychological Bulletin 70 (4), p. 213â220. Cited by: §A.3. [10] L. J. Cronbach and P. E. Meehl (1955) Construct validity in psychological tests. Psychological Bulletin 52 (4), p. 281â302. Cited by: §2.1. [11] S. Dev, A. Sloan, J. Kavner, N. Kong, and M. Sandler (2026) Judge reliability harness: stress testing the reliability of LLM judges. External Links: 2603.05399 Cited by: §3. [12] L. Dietz, O. Zendel, P. Bailey, C. Clarke, E. Cotterill, J. Dalton, F. Hasibi, M. Sanderson, and N. Craswell (2025) LLM-evaluation tropes: perspectives on the validity of LLM-evaluations. In Proceedings of the ACM SIGIR International Conference on the Theory of Information Retrieval (ICTIR), External Links: Document Cited by: §3. [13] B. Efron (1979) Bootstrap methods: another look at the jackknife. The Annals of Statistics 7 (1), p. 1â26. Cited by: §A.2. [14] C. Feng, M. Shen, A. Balashankar, C. Gerner-Beuerle, and M. R. D. Rodrigues (2026) Noisy but valid: robust statistical evaluation of LLMs with imperfect judges. External Links: 2601.20913 Cited by: Table 2. [15] J. L. Fleiss (1971) Measuring nominal scale agreement among many raters. Psychological Bulletin 76 (5), p. 378â382. Cited by: §5.5. [16] M. Gardner, Y. Artzi, V. Basmov, J. Berant, B. Bogin, S. Chen, P. Dasigi, D. Dua, Y. Elazar, A. Gottumukkala, N. Gupta, H. Hajishirzi, G. Ilharco, D. Khashabi, K. Lin, J. Liu, N. F. Liu, P. Mulcaire, Q. Ning, S. Singh, N. A. Smith, S. Subramanian, R. Tsarfaty, E. Wallace, A. Zhang, and B. Zhou (2020) Evaluating modelsâ local decision boundaries via contrast sets. In Findings of the Association for Computational Linguistics: EMNLP 2020, p. 1307â1323. Cited by: §3. [17] R. Geirhos, J. Jacobsen, C. Michaelis, R. Zemel, W. Brendel, M. Bethge, and F. A. Wichmann (2020) Shortcut learning in deep neural networks. Nature Machine Intelligence 2 (11), p. 665â673. Cited by: §3.1. [18] M. Geva, Y. Goldberg, and J. Berant (2019) Are we modeling the task or the annotator? an investigation of annotator bias in natural language understanding datasets. In Proceedings of EMNLP-IJCNLP, p. 1161â1166. Cited by: §3.1. [19] P. I. Good (2000) Permutation tests: a practical guide to resampling methods for testing hypotheses. Springer Series in Statistics. Cited by: §A.2. [20] A. Gunjal, A. Wang, E. Lau, V. Nath, B. Liu, and S. Hendryx (2025) Rubrics as rewards: reinforcement learning beyond verifiable domains. arXiv preprint arXiv:2507.17746. Cited by: Remark 2. [21] S. Gururangan, S. Swayamdipta, O. Levy, R. Schwartz, S. R. Bowman, and N. A. Smith (2018) Annotation artifacts in natural language inference data. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL), p. 107â112. Cited by: §3. [22] Z. Huang, Z. Qiu, Z. Wang, E. M. Ponti, and I. Titov (2025) Post-hoc reward calibration: a case study on length bias. External Links: 2409.17407 Cited by: §1, Table 2, §5.3. [23] K. Hyland (1998) Hedging in scientific research articles. John Benjamins. Cited by: §4.1. [24] A. Z. Jacobs and H. Wallach (2021) Measurement and fairness. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT), p. 375â385. Cited by: §1.1, §2.1, §3. [25] J. James, C. Xiao, Y. Li, N. S. Moosavi, and C. Lin (2026) RIGOURATE: quantifying scientific exaggeration with evidence-aligned claim evaluation. External Links: 2601.04350 Cited by: §3.1, §3, Abstract. [26] Z. Ji, L. Yu, Y. Koishekenov, Y. Bang, A. Hartshorn, A. Schelten, C. Zhang, P. Fung, and N. Cancedda (2025) Calibrating verbal uncertainty as a linear feature to reduce hallucinations. arXiv preprint arXiv:2503.14477. Cited by: §4.1. [27] Z. Jiang, A. Liu, and B. Van Durme (2025) Conformal linguistic calibration: trading-off between factuality and specificity. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2502.19110 Cited by: §3.1, §3.1, §3.1, §3, Abstract. [28] D. Kaushik, E. Hovy, and Z. C. Lipton (2020) Learning the difference that makes a difference with counterfactually-augmented data. In International Conference on Learning Representations (ICLR), Cited by: §3.1. [29] N. Lambert, V. Pyatkin, J. Morrison, L. Miranda, B. Y. Lin, K. Chandu, N. Dziri, S. Kumar, T. Zick, Y. Choi, N. A. Smith, and H. Hajishirzi (2024) RewardBench: evaluating reward models for language modeling. arXiv preprint arXiv:2403.13787. Cited by: §1, Table 2, §5.4. [30] Y. Liu, Z. Yao, R. Min, Y. Cao, L. Hou, and J. Li (2025) RM-Bench: benchmarking reward models of language models with subtlety and style. In International Conference on Learning Representations (ICLR), Note: arXiv:2410.16184 Cited by: §1, Table 2. [31] L. Long, R. Wang, R. Xiao, J. Zhao, X. Ding, G. Chen, and H. Wang (2024) On LLMs-driven synthetic data generation, curation, and evaluation: a survey. External Links: 2406.15126 Cited by: §4.4. [32] S. Malik, V. Pyatkin, S. Land, J. Morrison, N. A. Smith, H. Hajishirzi, and N. Lambert (2025) RewardBench 2: advancing reward model evaluation. External Links: 2506.01937 Cited by: Table 2. [33] A. Marioriyad, M. H. Rohban, and M. Soleymani Baghshah (2025) The silent judge: unacknowledged shortcut bias in LLM-as-a-judge. Note: NeurIPS 2025 Workshop on Reliable ML from Unreliable Data External Links: 2509.26072 Cited by: §1, Table 2, §5.3. [34] R. T. McCoy, E. Pavlick, and T. Linzen (2019) Right for the wrong reasons: diagnosing syntactic heuristics in natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (ACL), p. 3428â3448. Cited by: §3.1. [35] S. Messick (1989) Validity. In Educational Measurement, p. 13â103. Cited by: §2.1. [36] S. J. Mielke, A. Szlam, E. Dinan, and Y. Boureau (2022) Reducing conversational agentsâ overconfidence through linguistic calibration. Transactions of the Association for Computational Linguistics 10, p. 857â872. Cited by: §4.1. [37] B. Newman, L. Soldaini, R. Fok, A. Cohan, and K. Lo (2023) A question answering framework for decontextualizing user-facing snippets from scientific documents. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 3194â3212. Note: arXiv:2305.14772 Cited by: §4.1. [38] J. D. Norman, M. U. Rivera, and D. A. Hughes (2026) Reliability without validity: a systematic, large-scale evaluation of LLM-as-a-judge models across agreement, consistency, and bias. External Links: 2606.19544 Cited by: §1, Table 2, Remark 2. [39] J. Ou, W. Walden, K. Sanders, Z. Jiang, K. Sun, J. Cheng, W. Jurayj, M. Wanner, S. Liang, C. Morgan, S. Han, W. Wang, C. May, H. Recknor, D. Khashabi, and B. Van Durme (2025) CLAIMCHECK: how grounded are LLM critiques of scientific papers?. In Findings of the Association for Computational Linguistics: EMNLP 2025, p. 21712â21735. External Links: Document Cited by: §3.1. [40] S. Park, J. Choi, J. Kang, S. Lee, J. Shin, and H. Shim (2026) Mitigating perceptual judgment bias in multimodal LLM-as-a-judge via perceptual perturbation and reward modeling. External Links: 2606.02578 Cited by: §3.1. [41] U. Peters and B. Chin-Yee (2025) Generalization bias in large language model summarization of scientific research. Royal Society Open Science 12 (4), p. 241776. Note: Ten models; odds ratio 4.85, 95% CI [3.06, 7.70], p<0.001p<0.001 External Links: Document Cited by: 3rd item, §1, §3.1, §4.1, §6.2, §7, Abstract. [42] A. Poliak, J. Naradowsky, A. Haldar, R. Rudinger, and B. Van Durme (2018) Hypothesis only baselines in natural language inference. In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics (*SEM), p. 180â191. Cited by: §3. [43] R. Rafailov, A. Sharma, E. Mitchell, S. Ermon, C. D. Manning, and C. Finn (2023) Direct preference optimization: your language model is secretly a reward model. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. Note: arXiv:2305.18290 Cited by: Table 2. [44] I. D. Raji, E. M. Bender, A. Paullada, E. Denton, and A. Hanna (2021) AI and the everything in the whole wide world benchmark. In NeurIPS Datasets and Benchmarks Track, Cited by: §3.1. [45] M. T. Ribeiro, T. Wu, C. Guestrin, and S. Singh (2020) Beyond accuracy: behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, p. 4902â4912. Cited by: §1.1, §2.1, §3. [46] S. Roy, R. Pujari, T. Kumarage, C. Peris, R. Gupta, A. Rumshisky, P. Natarajan, and V. Saligrama (2026) PReMISE: policy rubrics as measurement specifications for LLM judges. External Links: 2605.30803 Cited by: §3.1. [47] J. Skalse, N. H. R. Howe, D. Krasheninnikov, and D. Krueger (2022) Defining and characterizing reward hacking. Advances in Neural Information Processing Systems (NeurIPS). Cited by: §3.1, Table 2. [48] S. Tayebi Arasteh (2026) The strength of clinical evidence is recoverable from language model representations but not from their stated grades. External Links: 2606.29034 Cited by: Appendix E, §3.1, §4.1, §5.4. [49] J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal (2018) FEVER: a large-scale dataset for fact extraction and VERification. In Proceedings of NAACL-HLT, p. 809â819. Cited by: §4.1. [50] V. Vincze, G. Szarvas, R. Farkas, G. MĂłra, and J. Csirik (2008) The BioScope corpus: biomedical texts annotated for uncertainty, negation and their scopes. In BMC Bioinformatics, Cited by: §4.1. [51] D. Wadden, S. Lin, K. Lo, L. L. Wang, M. van Zuylen, A. Cohan, and H. Hajishirzi (2020) Fact or fiction: verifying scientific claims. In Proceedings of EMNLP, p. 7534â7550. Cited by: §4.1. [52] D. Wan, J. Vig, M. Bansal, and S. Joty (2024) On positional bias of faithfulness for long-form summarization. External Links: 2410.23609 Cited by: §3. [53] A. Wang, I. Arcuschin, and A. Conmy (2026) Automatically finding reward model biases. External Links: 2602.15222 Cited by: Table 2. [54] P. Wang, L. Li, L. Chen, Z. Cai, D. Zhu, B. Lin, Y. Cao, Q. Liu, T. Liu, and Z. Sui (2023) Large language models are not fair evaluators. External Links: 2305.17926 Cited by: §1, Table 2. [55] Z. Xu, S. Li, H. Liu, X. Wang, S. Li, Z. Song, and X. Chen (2026) Inside the unfair judge: a mechanistic interpretability account of LLM-as-judge bias. External Links: 2607.11871 Cited by: Table 2, §5.3. [56] W. Ye, G. Zheng, and A. Zhang (2025) Rectifying shortcut behaviors in preference-based reward learning. In Advances in Neural Information Processing Systems (NeurIPS), Note: Introduces PRISM. arXiv:2510.19050 Cited by: Table 2. [57] Y. Zha, Y. Yang, R. Li, and Z. Hu (2023) AlignScore: evaluating factual consistency with a unified alignment function. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Cited by: Table 2. [58] K. Zhang, M. Hu, H. A. D. Le, F. K. Torsha, Z. Jiang, M. K. Bui, C. Chang, Y. Chuang, Z. Xiong, Y. Lin, G. Wang, and N. Zou (2026) A survey on evaluating quality and trustworthiness in LLM-generated data. External Links: 2601.17717 Cited by: §4.4. [59] X. Zhang, W. Xiong, L. Chen, T. Zhou, H. Huang, and T. Zhang (2025) From lists to emojis: how format bias affects model alignment. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), p. 26940â26961. External Links: Document Cited by: Appendix B, §1, Table 2, §4.1, §5.3. [60] L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica (2023) Judging llm-as-a-judge with mt-bench and chatbot arena. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: Appendix B, item length_only., §1, Table 2, §4.1, §5.3. [61] T. Zhu (2026) How should AI talk about us? LLMs and social generics. AI & Society. External Links: Document Cited by: §3.1. The appendices carry what the argument does not need but a replication does: the full intervention specification with its prompts, the annotation instruments, the pinned judge configurations, the predictor family, the register of predicted values, and the secondary experiments that inform the design without supporting a headline claim. Appendix A Extended definitions and proofs A.1 Measurability and the item distribution Definition 3 writes S and R as probabilities over xâŒx and T drawn from an intervention family. Two details the main text suppresses. First, the two probabilities are over different product spaces (ĂTIDĂ T_I and ĂTCDĂ T_C), so they are not two coordinates of one joint distribution and no covariance between them is defined; this is what makes Proposition 1 available. Second, D in practice is the empirical distribution over base items that survived Protocol 1, which is not the distribution over items a judge meets in deployment. Every profile in this paper is therefore conditional on the surviving item set, and Appendix H reports how much the profile moves when the survival filter is relaxed. A.2 The paired estimator, and why the bootstrap resamples base items For base items x1,âŠ,xnx_1,âŠ,x_n with mim_i interventions applied to xix_i, the natural estimator of R is R^=1nâi=1n1miâj=1mi[J(xi)â J(Tiâjxi)], R\;=\; 1n _i=1^n 1m_i _j=1^m_i1\! [J(x_i)â J(T_ijx_i) ], (6) which weights base items equally rather than interventions equally. The alternative (pooling all âimi _im_i interventions) over-weights base items that happened to admit more edits, and admissibility is not random: a sentence with several hedges admits several strength edits, so pooling would let hedge-dense sentences dominate the axis they are most relevant to. Confidence intervals resample base items with replacement, never interventions [13]: interventions on one base item share that itemâs content and are not independent draws. Nulls are label-permutation nulls over the same grouping [19], and where a table reports many rows we control the false discovery rate across them [1]. Figure 8: The unit of analysis. (6) weights base items equally and, within an item, its interventions equally; pooling all interventions instead lets a hedge-dense sentence dominate the axis it is most relevant to, since a sentence carrying several hedges admits every strength slot and a bare one admits none. Intervals resample whole base items with replacement and never interventions, and folds split on the source document. The per-item counts drawn here are illustrative; measured per-slot yields are in Table 4. A.3 Agreement statistics Judge-versus-human agreement is chance-corrected [8]; where the verdict space is ordered we also report the linearly weighted variant [9], because an ordered space punishes an adjacent-class error and an opposite-class error identically under the unweighted statistic, and the two are not equally bad for a scope judgement. Degenerate cases (both raters constant, so expected agreement is 11 and Îș is 0/00/0) are reported as 00 rather than as undefined, matching the harness implementation. Appendix B The intervention families, in full Table 8: Every intervention type, its class, and the axis it moves. Frozen in probes/edit_taxonomy.py before any judge was run; the assertion in that file fails if a type is added without an axis, so the table and the code cannot drift. type class / axis what it does domain extension TCT_C / scope extend the claim to an adjacent domain population widening TCT_C / scope widen who or what the claim covers quantifier strengthening TCT_C / scope strengthen a quantifier (some â all, often â always) tense modality shift TCT_C / scope restate a one-time observation as a standing property condition removal TCT_C / strength drop the circumstances under which it was observed hedging removal TCT_C / strength remove a hedge the evidence relied on intensifier addition TCT_C / strength add an intensifier the evidence does not support (substantially, consistently, markedly) elaboration TIT_I / control add a clause that renames or restates something the sentence already names, using no epistemic language format shift TIT_I / control add or remove list/bold/heading formatting, same claim, same scope register shift TIT_I / control rewrite in a more formal textbook register, same claim, same scope reordering TIT_I / control reorder clauses without changing content or scope verbosity padding TIT_I / control lengthen by syntactic expansion only â nominalise, unpack a compound, spell out an abbreviation â with no epistemic language Why the control arm covers exactly these 5. Three are register manipulations. The other two exist because the judge literature has already shown verbosity [60] and formatting [59] to move verdicts: a control arm that omitted them would report a flattering S on properties nobody had implicated, which is the failure mode Proposition 2 warns about in its milder form. A sixth slot, hedge padding, was specified and then removed: adding a hedge moves commitment, so it is a strength edit and a control arm containing it is not construct-preserving. The taxonomy was internally contradictory and still looked fine as a list of names, which is why Table 8 is generated from the code rather than maintained beside it. Generation and verification prompts. The three templates below are reproduced verbatim from probes/edit_taxonomy.py, which is the single source of truth; braces in the JSON schema are doubled there because the templates are Python format strings, and are shown singly here. edit_type is filled from Table 8 and edit_gloss from the gloss column, so a template is never edited to accommodate a slot. tol is the length tolerance, fixed at 1010 percent before generation. Construct edit (treatment arm, TCT_C) You rewrite scientific sentences to make them claim MORE than their evidence supports, changing nothing else. Input sentence (well-calibrated: it states exactly what the evidence licenses): sentence Evidence context from the source article: evidence Apply EXACTLY ONE edit of type edit_type (edit_gloss). Hard constraints. Violating any one of them makes the output useless: 1. Do NOT add, remove, or alter any factual content: no numbers changed, no entities swapped, no new findings introduced. 2. Do NOT introduce provenance language. Never write âthis studyâ, âthe authorsâ, âFigure Nâ, âin our dataâ, or any equivalent. Neither sentence may point at a source. 3. Keep the register identical. If the input is plain declarative prose, so is the output. 4. Keep the length within tol percent of the inputâs token count. If your edit would shorten the sentence, compensate elsewhere WITHOUT adding content. 5. Change exactly one scope slot (the one named above) and leave the others alone. Return JSON only: "edited": ..., "slot_changed": "edit_type", "what_widened": ..., "what_it_became": ..., "facts_preserved": true, "anchor_free": true Register edit (control arm, TIT_I) You rewrite scientific sentences to change their SURFACE FORM ONLY, leaving the claim and its scope exactly as they were. Input sentence: sentence Apply EXACTLY ONE edit of type edit_type (edit_gloss). Hard constraints: 1. The claim must remain true of exactly the same population, conditions, and modality. Do not widen or narrow anything. 2. Do NOT add or remove factual content. 3. Do NOT introduce provenance language. 4. Keep the length within tol percent of the inputâs token count. Return JSON only: "edited": ..., "surface_change": "edit_type", "scope_unchanged": true, "facts_preserved": true, "anchor_free": true Verifier (family B), five gates Two sentences are given. Judge ONLY the following, independently of style or fluency. A: original B: edited Evidence context: evidence Answer each question separately: 1. facts_identical: do A and B assert the same factual content (same numbers, same entities, same findings) with nothing added or removed? 2. scope_relation: is Bâs scope broader, narrower, or same relative to A? 3. slot_changed: which scope slot differs, if any? One of population_widening, condition_removal, tense_modality_shift, quantifier_strengthening, hedging_removal, domain_extension, or none. 4. anchor_free_both: is neither sentence pointing at a source (no âthis studyâ, no âFigure Nâ, no âthe authorsâ)? 5. register_comparable: are they in the same register, so that neither reads as more formal or more hedged than the other beyond the slot change? Return JSON only: "facts_identical": <bool>, "scope_relation": ..., "slot_changed": ..., "anchor_free_both": <bool>, "register_comparable": <bool>, "notes": ... What the verifier gates, and what it does not. A generated pair enters the sensitivity arm only if the verifier returns facts_identical, anchor_free_both, register_comparable, and a slot_changed equal to the requested slot; the control arm additionally requires scope_relation == same. The verifier does not decide which member of a pair is the over-scoped one. That is the direction assignment, it is made by human annotators (Appendix C), and it is separated from generation and verification on purpose: a direction supplied by any model would make R a measurement of agreement between two models rather than of a judge against a construct. Generation, verification, and judging draw on three disjoint model families, so no family scores its own output. Appendix C Annotation protocol and codebook This paperâs human labels do one job: given a pair (x,TCâx)(x,T_Cx), the annotator decides which member a correctly calibrated reader should prefer. Delegating that to a model would make R an agreement statistic between two models, on which a judge sharing the direction-assignerâs blind spot scores as sensitive exactly where both are blind. Unit and blinding. The unit is one pair, shown as sentences A and B in randomised order with the evidence context and nothing else. Withheld: which member is the original, which edit slot was requested, which axis it belongs to, the generating and verifying model identities, and any batch-level composition cue. Presentation order is recorded, so a systematic first-position preference is detectable rather than absorbed into the labels. Treatment and control pairs are interleaved in one stream, since an annotator able to tell the arms apart could infer the control armâs answer from the arm itself. Figure 9: One pairâs route through Protocol 1, and the five places it can stop. Four outcomes are decided per item. The fifth is decided per slot and per axis after aggregation, against a bar fixed in probes/edit_taxonomy.py before collection, and it excludes the slot whole rather than widening an interval. Of 478 annotated pairs, 246 construct and 96 control items survive, 2 carry no label and 2 slots were excluded; per-slot yields are in Table 4. The basis field is a mandatory quotation of the span that exceeds the evidence, which is what makes a disagreement inspectable rather than a number, and presentation order is recorded so a systematic first-position preference is detectable rather than absorbed into the labels. The judgement. Three fields per pair. direction â A claims more than the evidence licenses, B does, neither, cannot tell from this context. basis: a mandatory one-sentence quotation of the exact span that exceeds what the evidence supports, plus what the evidence does support. same_facts: whether the two sentences assert the same factual content, a check on the generator rather than a judgement about scope. A pair marked cannot tell leaves the primary analysis and is reported as a separate rate; a pair where the annotators say the facts differ is discarded outright and counted as a generator failure, because a probe built on it would measure factuality. Class definitions and the boundary cases that matter. The four scope slots and three strength slots of Table 8 are defined for annotators through worked examples rather than through the slot names, which are ours and not theirs. Three boundaries carry most of the disagreement. Removal is expansion: turning âresults suggest M may reduce noiseâ into âM reduces noiseâ widens the claim although no word was added and nothing became false, and an annotator reading only for added content will call such a pair a tie. Generality is not always over-claiming: a typicality quantifier the evidence supports, physical constants, unit conversions and definitions are not over-scoped, and without this boundary an annotator drifts toward marking every general-sounding sentence as the over-scoped one, inflating R for exactly the judges that share the heuristic. An appositive gloss is not a change of scope: rewriting âaerosolsâ as âaerosols, the artificially designed particlesâ names the same referent more precisely without altering what is claimed of it. The third boundary was added after collection, because the codebook as run did not answer the question and one annotator read it the other way consistently (§4.2); the affected slots are reported as not measurable rather than ruled on retrospectively. Annotators, the majority rule, and the bar. 3 independent annotators label every pair, none seeing anotherâs labels, with domain competence in the pairâs field and a training round of shared items that is discussed and then discarded. The registered plan was two annotators with a third arbitrating disagreements; all three labelled the full batch instead, which is what makes a systematic disagreement distinguishable from a noisy one. The label is the majority of three with no automatic tie-break, so an item on which all three differ carries no label and the rate is reported (2 of 478). The pre-registered bar is a raw pairwise agreement of â„0.75â„ 0.75, fixed in probes/edit_taxonomy.py before any data was collected and applied per axis and per slot. Raw agreement rather than Îș, because Îș is degenerate wherever one answer legitimately dominates: register_shift has raw agreement 0.976 with Fleissâ Îș 0.324, since almost every label is same, and a Îș bar would have discarded the cleanest slot in the arm. Generated tables flag every cell where one category holds â„90%â„ 90\% of the labels. A slot or axis below the bar is reported as not measurable and excluded from R and from the axis comparison rather than reported with a wider interval: the question is whether the construct is well defined for humans at all, and a wide interval around an ill-defined construct still asserts it exists. Recruitment and consent. Annotators are paid contributors, not volunteers or students of the authors, compensated at or above the local hourly minimum for their jurisdiction. They are told what the data will be used for, that their labels are released per item without their identity attached, and that they may withdraw. No task presents personal data or content selected to be distressing. Two requirements on any span-quoting field. Both cost us a batch of correct labels before they were understood. Normalise typography before comparing: source sentences carry non-breaking hyphens, en dashes and curly quotes, an annotator who retypes produces the ASCII forms, and an unnormalised substring test rejects a correct label over one character. Accept more than one span: an edit can have two loci that are contiguous nowhere, and a single-span field cannot express it. A validation rule stricter than the property it checks rejects correct work silently. Appendix D Judge versions and decision protocols Every judge is reached through a single aggregating API, which is what makes a multi-family study tractable and the anti-circularity requirement of Protocol 1 easy to satisfy: disjoint vendors are strings in a config rather than procurement problems. An aggregator makes âpinned model versionâ harder, not easier. A model identifier at an aggregating API is not a model instance: the same identifier can be served by several upstream providers at different quantisations and with different sampling implementations, and the default is to route to whichever is available and fall back silently. A run that pins the judge version while accepting default routing has pinned a name, not an instrument, and two rows of Table 5 could differ by provider rather than by judge. We therefore pin the provider as well as the model, disable fallback, and record the resolved upstream per request; requests resolving to a different provider are discarded and re-issued. The count of discards is reported, because it is a fact about the measurement apparatus and hiding it would misrepresent how reproducible these numbers are. Quantisation is the specific hazard worth naming, because it is invisible and directional: a more aggressively quantised serving of the same weights is a different instrument for our purposes, and there is no reason its (S,R)(S,R) should match. Where an upstream does not disclose its quantisation we say so in the table rather than assuming parity. Every judge is run at temperature 00 on the same frozen prompt, with the provider pinned through the routing layer and the resolved provider recorded per call. What the elicitation pilot did to Remark 1. The remark asserted that a graded scale buys resolution over a binary verdict, and a pilot showed that it does not follow automatically. The two failure modes are opposite: a judge can be perfectly repeatable and still emit effectively one decision, and another can offer many usable cuts while reproducing its own grade on the same item a small fraction of the time at temperature 00. Both appear in the measured roster, where usable operating points span 2 to 9 and regrade agreement 0.412 to 1.000 (Table 5). A study reporting only distinct grades would have called every one of them fine. The consequence is that the score-to-verdict rule is fixed once for all judges and the regrade agreement is reported beside every profile rather than folded into its interval. 002244668810101212glmmistralllama4gemininemotronhaikudeepseekgrade valuesusable operating pointsdistinct grades emitted000.50.5110.980.960.971.000.830.410.68regrade agreement Figure 10: Instrument quality per judge, and the two independent ways a graded judge fails. Left: the grade values a judge emits, against those surviving the one-percent usability cut, so the segment is resolution the judge appears to offer and does not deliver. Right: whether it reproduces its own grade on the same item at temperature 00. Rows are sorted by usable points and the regrade column does not follow that order, so neither quantity predicts the other. Every judge here emits at least seven distinct grades, which is why a study reporting distinct grades would have called all of them fine. Appendix E The impoverished predictor family, in full The family is deliberately weak and deliberately frozen: weak so that every reported VPVP is an upper bound and therefore generous to the label set, frozen so that VPVP does not change meaning between papers. The harness asserts the member set, and a test fails if a predictor is added. Table 9: The three members. Feature counts and fitting procedure are fixed in analysis/validation_power.py; group-aware cross-validation splits on source document so no document appears in both folds. predictor sees does not see length_only response length in characters any token identity numeric_surface counts of digits, punctuation, and casing patterns any word surface_ngram character n-grams, nâ€4n†4 sentence structure, evidence context Why these three. The members are ordered by how much of the surface they see, so a label setâs failure locates itself. If length_only suffices, the set is separable by a quantity nobody would defend as the construct. If numeric_surface is needed, the leak is in formatting or numeric density, the signature of a set whose two conditions differed in how they render evidence. If only surface_ngram succeeds, the leak is lexical: a generative condition left an idiom behind. The three answers imply different repairs, which is why the per-member breakdown is reported rather than only aimp=maxa_imp= . Every member is a function of the response text alone, and none is shown the evidence context, so none can compute the relation between claim and evidence. That is what makes aimpa_imp interpretable: a member reaching the human ceiling has not solved the task, it has shown the task was not being asked. The invariant is structural, since the extractors take a single string and no code path delivers the evidence to them. Fitting. Cross-validation is group-aware on the source document, because two items from one document share topic, register and often exact phrasing, so a random split lets a predictor recognise the document and recover the label through it. Each member is refit on labels permuted within group under identical folds; refitting rather than assuming 12 12 also catches a broken harness, since a permutation null above chance means the pipeline leaks. No hyperparameter search is performed: a tuned impoverished predictor would make aimpa_imp a function of how hard we tried. Every member ends in LogisticRegression(max_iter=2000, class_weight="balanced"), weighted because several audited sets are imbalanced and an unweighted classifier can reach high accuracy at Îș near zero. length_only. One feature, the whitespace token count, standardised. It isolates the single surface property the preference literature repeatedly finds correlated with human choice [60]. numeric_surface. Ten standardised scalars: character count; whitespace token count; count of .,;:; count of (; newline count; hedge-marker rate; universal-quantifier rate; digit rate; uppercase rate; mean token length. The two rate features use closed hand-written lists, hedges may, might, could, appears, suggests, possibly, likely, generally, typically, often, tends to, in some cases and universals all, always, never, every, any, invariably, universally, counted case-insensitively as substrings over the token count. The lists are closed so that this member cannot drift into a content classifier as vocabulary grows. surface_ngram. Character n-grams, word-boundary aware, nâ[3,5]nâ[3,5], minimum document frequency 3, at most 20,000 features, sublinear term frequency, no scaler. Character n-grams carry some content, which is why this member is reported separately: recovering a set that length_only does not means separable by style, a weaker and different claim. Folds and scoring. Five folds, GroupKFold on the group key where the set has at least five distinct groups and KFold(shuffle=True) with a fixed seed otherwise; the output records which was used, since a aimpa_imp computed without grouping is not comparable to one computed with it. A predictorâs score is the Îș of the pooled out-of-fold predictions, not the mean of per-fold Îș : Îș is not linear in the confusion matrix, and the pooled quantity is what Definition 6 is written against. aimpa_imp is chance-corrected, ahuma_hum carries its own interval, and where the two overlap we report VPVP as not resolvable. Each member is also fit on one set and evaluated on the others, because within-set recovery can be a set-specific idiom rather than a general surface correlate [48]. Appendix F Register of predicted values Every value in this paper marked â predicted, not measured is listed here with the reasoning that produced it. This appendix exists so that the predictions are falsifiable as a set rather than adjustable one at a time after the fact: a prediction recorded before measurement is a commitment, and a prediction recorded afterwards is a description. Table 10 is generated from docs/PREDICTED_VALUES.md by analysis/make_predicted_register.py, which also fails the build if the paper renders a â predicted, not measured value the register does not explain. Transcribing it by hand would defeat its purpose: the appendix is only a commitment device if it cannot drift from the file the commitments were written in. Two properties of the register are worth stating because they constrain how the rest of the paper may argue. First, no sentence in this paper argues from a predicted value. Where a sectionâs design depends on one, the dependence is named in the prose. Second, predictions exist here for a mechanical reason rather than a rhetorical one: a plot needs coordinates to render, and there is no hole-marker equivalent for a data point, so the choice for a figure is between a marked prediction and no figure at all. Given that choice, the useful move is to record the prediction with its refutation condition attached and let the experiment kill it. Every such figure carries a predicted watermark drawn inside the axis rather than in the caption, because a caption disclaimer does not survive the figure being screenshotted into a slide. This appendix is a defect that shrinks. Each line disappears when its measurement lands, and the appendix disappears with the last one. Table 10: Register of predicted values. Every â predicted, not measured value in this paper, with the reasoning that produced it and the measurement that would refute it. A dash in the final column marks a design parameter or a value with no single refuting observation. key value basis refuted if C2 â the validity profile Sinv 0.88 Judges are already known to be fairly invariant to individual surface properties, and the strongest published verbosity-bias estimate is small (reliability2026, under 0.011 correlation across 21 judges). But our control arm is six properties at once including two with published effects, so invariance should sit below what any single-property study reports. High, not near-1. S<0.75S<0.75 â meaning the control arm is finding failures the single-property literature misses, which would be a finding in its own right and would make every R comparison harder Rsens 0.41 Pure judgement, and the weakest-founded number here. Anchored on the intuition that a judge does respond to some construct edits â quantifier changes are lexically salient â but not most. If it were near 0.8 the gap would have been noticed already; if near 0.1 judges would fail obviously broken cases. R>0.75R>0.75 at Sâ0.88Sâ 0.88 â judges are construct-sensitive and the paperâs headline is wrong nJudges 6 The scale the external review set as the floor for a claim about instruments rather than about one vendor. â (design parameter, not a prediction) nDomains 4 Same. See PEER_REVIEW_VWV.md §5 for the minimum-viable fallback of two. â C3 â the scope/strength asymmetry RscopeAx 0.57 Scope edits change quantifiers, populations and domains â lexically explicit, and the closest thing to them in the literature (over-generalisation in summaries, peters2025generalization) is detectable enough to have been measured at scale. â RstrengthAx 0.24 Hedge and condition deletion leave a fluent, confident sentence with no lexical marker of what was removed. A judge sees an absence, and absences are the classic blind spot. This is the paperâs central bet. the gap is not significant at matched S, or reverses axisGap 0.33 Difference of the two above. <0.10<0.10, or not significant under a paired bootstrap over base items ctrlFPR 0.12 1âS1-S. Reported alongside every sensitivity figure by rule. â C3b â the accuracy-prompting paradox dScopeAcc â0.09-0.09 Accuracy pressure should make a model narrow its reach â the obvious, intended effect of the instruction. â dStrengthAcc +0.14+0.14 The counter-effect: sounding certain means dropping hedges and stated conditions. Signed to make the pooled effect small and positive, which is what makes peters2025generalizationâs result look paradoxical. both axes move the same way. Then the decomposition does not explain the paradox and §paradox must say so in those words C4 â validation power of public label sets vpMedian 0.11 Low but non-zero across sets, reflecting that some are human-labelled (MT-Bench, HelpSteer2) and should retain headroom, while pair-role-derived sets should not. median >0.30>0.30 â the sets are sound, C4 becomes a negative result, C2âC3 unaffected vpStrength 0.03 The prediction that closes the paperâs loop: no public set has headroom on the strength axis, so none could have detected C3âs asymmetry. >0.15>0.15 â the sets could have found it, and the paper owes a different explanation for why nobody did C5 â the inflation Delta 0.19 The size that would make the result matter without straining credibility: large enough to change a paperâs conclusions, small enough to be plausible for one construct. Frankly a placeholder. it is not the point â Î is an existence proof, and any significant positive value supports the claim as stated humanAgree 0.81 Above the pre-set bar of 0.75, by design of the protocol rather than by prediction. below 0.75, in which case the affected axis is reported as not measurable rather than measured with weak labels Appendix G What is not measured here Table 11: What this paper specifies and does not measure, so that a deliberate omission is distinguishable from an oversight. Two omissions are argued in place instead, where the omission bears on how a nearby number should be read: the accuracy-prompting arm (§5.3) and the blinded replacement set (§5.5). quantity why not measured here Per-row commentary on each public setâs recoverability needs construction details the sets do not publish Which audited set each checklist item would have flagged needs construction details the sets do not publish Quantisation and API-version instrument columns most providers do not disclose them Self-consistency and ensemble remedy rows two further judges in the same harness; not yet run Sensitivity under a finer verdict space needs a second elicitation with its own validation The profile per language probe items are English only The profile at relaxed agreement bars a trajectory, not a point estimate; awaits a larger pool Appendix H Secondary experiments and background work None of the following supports a headline claim. Each rules out an alternative explanation, bounds a design choice, or records a path considered and rejected. S1, S3, S6: robustness checks specified and not run. Three checks are named in Table 11 with the reason each awaits. Reliability remedies (S1): Proposition 1 predicts that self-consistency voting and a judge ensemble raise S without moving R; both coordinates are reported for every remedy row, never S alone. Edit magnitude (S2) has run and is reported in §5.2; the one thing it cannot do is stratify, because operation fixes magnitude. Verdict granularity (S3) asks how much insensitivity survives a finer verdict space. The elicitation pilot already answers half of it, and not in the direction the check assumed: a finer space does not automatically buy resolution, since one judge yields 2 usable operating points and another reproduces its own grade on the same item 0.412 of the time at temperature 00. âRepeat with a graded scoreâ is therefore not a strictly more informative measurement. Item-survival sensitivity (S6) recomputes the profile at relaxed agreement bars; if the asymmetry grows as the bar relaxes it is partly an artefact of which items survive. S4. A free replication of the asymmetry, from the pair-construction pipeline. Building the intervention set requires a verifier model to check that each edit preserved the facts, the anchor-free property and the register (Appendix B). That verifier is also asked, for its own record, whether the edited sentence claims more than the original: the same question put to the annotators, put to a model. It notices 24.1% of scope edits (957 pairs) and 8.8% of strength edits (794 pairs), a factor of 2.7. Where it does register a change it assigns the same axis we do 82% of the time (301 pairs), so the gap is not an artefact of disagreeing about what the axes mean. This is not R. The verifier is not a judge, the task is not the judgeâs task, and no human assigned direction to these pairs, so there is no ground truth here: only a modelâs agreement with our construction. It is reported as an independent instrument, built for another purpose, showing the predicted asymmetry at no additional cost. The verifierâs own insensitivity is why direction is a human decision. An earlier build gated the construct arm on the verifier agreeing that an edit widened the claim, and that gate rejected 12 of 12 verified strength edits, each with a note conceding the change and denying it mattered. Such a gate keeps only the strength edits a model can already see, so every judge would then be scored on a set pre-filtered for detectability. The verifierâs opinion is recorded rather than enforced. S5. Alternative decompositions considered and rejected. Before settling on scope/strength we considered a single graded overreach magnitude, rejected because it cannot express the opposite-sign prediction of §5.3; a three-way split separating quantifier from temporal generalisation, rejected because almost every temporal generalisation in the domainâs writing conventions also strengthens a quantifier, so the two are not separable by a minimal edit; and a split by evidence type rather than claim operation, rejected because it makes membership depend on the source document and so cannot be verified from the pair alone. The chosen decomposition is a modelling choice, and a reader should know what it was chosen against.