Paper deep dive
The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting
Lauri Lovén, Sasu Tarkoma
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/8/2026, 11:47:26 AM
Summary
The paper establishes that in scored reporting systems where agents have combined objectives (accuracy scores plus non-accuracy payoffs), the principal's optimal oversight mechanism endogenously produces a non-affine approval function, which inherently undermines truthful reporting. This creates a structural impossibility for calibration-preserving smooth oversight. A constructive escape is identified: a sharp step-function threshold achieves first-best screening across all strictly proper scoring rules. The Brier score is uniquely proven to yield welfare equivalence between second-best and first-best under smooth oversight. The framework bridges AI agent oversight and marketplace operation, linking to Goodhart's Law and classical mechanism design.
Entities (8)
Relation Signals (6)
Approval Function → exhibits → Endogeneity of Miscalibration
confidence 95% · the principal's optimal oversight necessarily uses a non-affine approval function to screen types, yet any non-affine approval makes truthful reporting suboptimal
Step-Function Threshold → achieves → First-Best Screening
confidence 94% · a step-function approval threshold achieves first-best screening for every strictly proper scoring rule
Brier Score → yields → Welfare Equivalence
confidence 93% · Under the Brier score specifically, the type-independent inflation cost yields a welfare equivalence between second-best and first-best
Perturbation Lemma → demonstrates → Truthful Reporting is Suboptimal
confidence 92% · Adding a non-constant, non-affine function to a strictly proper scoring mechanism shifts the maximizer away from the truthful report
AI Agent Oversight → parallels → Marketplace Operation
confidence 90% · The same structure appears in classical mechanism-design settings such as marketplace operation
Goodhart's Law → relatesto → Endogeneity of Miscalibration
confidence 88% · The connection to Goodhart's Law is discussed in Section 6
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Eliciting truthful reports from autonomous agents is a core problem in scalable AI oversight: a principal scores the agent's report using a strictly proper scoring rule, but the agent also benefits from the report through a non-accuracy channel (approval for autonomous action, allocation share, downstream control). The same structure appears in classical mechanism-design settings such as marketplace operation. Our main result is an endogeneity: the principal's optimal oversight necessarily uses a non-affine approval function to screen types, yet any non-affine approval makes truthful reporting suboptimal under the combined objective whenever deviation is undetectable. The principal cannot avoid the perturbation that undermines calibration. This impossibility holds for all strictly proper scoring rules, with a closed-form perturbation formula. A constructive escape exists: a step-function approval threshold achieves first-best screening for every strictly proper scoring rule, because the agent's binary inflate-or-not choice creates a type-space threshold regardless of the generator's curvature. Under the Brier score specifically, the type-independent inflation cost yields a welfare equivalence between second-best and first-best; we prove this equivalence is unique to Brier (the welfare gap under smooth $C^1$ oversight is bounded below by $\Omega(\text{Var}(1/G'') (\gamma/\beta)^2)$ for every non-Brier rule). Two instances develop the framework: AI agent oversight (the lead motivating setting) and marketplace operation (a parallel mechanism-design domain). The message for AI alignment is direct: smooth scoring-based oversight cannot elicit truthful reports from a strategic agent; sharp thresholds are the calibration-preserving design.
Tags
Links
- Source: https://arxiv.org/abs/2605.07671v1
- Canonical: https://arxiv.org/abs/2605.07671v1
Trouble viewing inline? Open PDF directly →
Full Text
161,181 characters extracted from source content.
Expand or collapse full text
The Endogeneity of Miscalibration: Impossibility and Escape in Scored Reporting Lauri Lovén Future Computing Group, University of OuluOuluFinland lauri.loven@oulu.fi and Sasu Tarkoma University of Oulu and University of HelsinkiOulu / HelsinkiFinland Abstract. Eliciting truthful confidence reports from autonomous agents is a central problem in scalable AI oversight: a principal scores the agent’s report using a strictly proper scoring rule, but the agent also benefits from the report through a non-accuracy channel (approval for autonomous action, allocation share, downstream control). The same structure appears in classical mechanism-design settings such as marketplace operation. Our central result is an endogeneity: the principal’s optimal oversight mechanism necessarily employs a non-affine approval function to screen types, yet any non-affine approval function makes truthful reporting suboptimal under the agent’s combined objective whenever the deviation is undetectable. The principal cannot avoid the perturbation that undermines calibration. This impossibility holds for all strictly proper scoring rules, with a closed-form perturbation formula quantifying the degradation. A constructive escape exists: a step-function approval threshold achieves first-best screening for every strictly proper scoring rule, because the agent’s binary inflate-or-not choice creates a type-space threshold regardless of the generator’s curvature. Under the Brier score specifically, the type-independent inflation cost yields a welfare equivalence between the second-best and the first-best; we prove this equivalence is unique to Brier (the welfare gap under smooth C1C^1 oversight is bounded below by Ω(Var(1/G′)⋅(γ/β)2) (Var(1/G )·(γ/β)^2) for every non-Brier scoring rule). Two instances develop the framework: AI agent oversight (the lead motivating setting) and marketplace operation (a parallel mechanism-design domain). The combined message for AI alignment is direct: a sophisticated principal cannot rely on smooth scoring-rule based oversight to elicit truthful reports from a strategic agent; sharp thresholds, not smooth incentives, are the calibration-preserving design. Proper scoring rules, incentive compatibility, mechanism design, scalable AI oversight, calibration, Fenchel duality, credible mechanisms, Goodhart’s law †copyright: none†journal: TEAC†journalyear: 2026†journalvolume: 0†journalnumber: 0†article: 0†ccs: Theory of computation Algorithmic mechanism design†ccs: Computing methodologies Multi-agent systems†ccs: Applied computing Economics†ccs: Computing methodologies Reinforcement learning 1. Introduction 1.1. Motivation Scalable AI oversight depends on eliciting truthful confidence reports from autonomous agents whose payoff depends on more than report accuracy. An autonomous AI agent reports its confidence to a human or institutional overseer and gains approval for autonomous action when the reported confidence crosses a threshold; in classical mechanism design, the same structure appears when a marketplace operator executes an allocation-payment mechanism on behalf of bidders and earns revenue that depends on the extracted payments. In each case, the principal scores the report for accuracy using a strictly proper scoring mechanism (one that uniquely incentivizes truthful reporting in isolation). The problem is that the agent also benefits from its report through a non-accuracy channel (approval for autonomous action, allocation share, downstream control). The agent faces a combined objective: a strictly proper score (rewarding accuracy) plus a perturbation payoff (rewarding something other than accuracy). This creates a fundamental tension between two goals the principal needs simultaneously: screening (using the report to separate good types from bad) and calibration (keeping the report accurate). The same tension underlies the scoring-rule-based calibration mechanisms now central to scalable AI oversight, debate (Irving et al., 2018), and reinforcement learning from human feedback (Christiano et al., 2017), where the model’s reported probability is both the object the overseer scores and the input to a downstream selection rule that returns payoff to the model. The central question of this paper is not merely whether such perturbations make truthful reporting suboptimal (they do, generically), but whether a sophisticated principal can design oversight to avoid this outcome. The answer is no: the principal’s optimal oversight mechanism endogenously produces a non-affine approval function, which, given the exogenous conditions that the agent has a conflicting payoff incentive and that deviation is undetectable within finite monitoring horizons, makes truthful reporting impossible. This endogeneity is the paper’s main contribution. 1.2. Main Result: Endogenous Impossibility Scope. The instances in this paper develop the binary-outcome scalar-type setting: a single reporter holds a one-dimensional type (e.g., a probability of success) and reports to a receiver whose score depends on a binary outcome. Section 6.6 analyzes the multi-dimensional extension, showing that the core impossibility generalizes to d-dimensional types while the welfare analysis remains open. Generality demarcation. The impossibility (Result 1) and the step-function escape (Result 2) hold for all strictly proper scoring rules. The welfare equivalence between the second-best and the first-best (Theorem 5.3(i)) is Brier-specific: it depends on the Brier score’s constant curvature G′(p)=2G (p)=2, which generates a type-independent inflation cost. For non-Brier scores, the step-function escape still achieves first-best, but the welfare analysis under smooth oversight remains open. Building on the classical observation that non-affine perturbations make truthful reporting suboptimal under the agent’s combined objective (the Perturbation Lemma), the argument establishes two novel results: an endogeneity and an escape. Result 1 (endogenous impossibility), part (a): classical foundation. Any non-affine perturbation makes truthful reporting suboptimal under the combined objective (the Perturbation Lemma, Lemma 3.1). Adding a non-constant, non-affine function to a strictly proper scoring mechanism shifts the maximizer away from the truthful report. This is mathematically classical, following from the observation that perturbing a strictly concave objective shifts its argmax. The lemma’s value lies in its application to four elicitation traditions that share a common Fenchel conjugate structure (known in the literature; see Section 3), and in the closed-form perturbation formula ((3.2)) with quantitative predictions. Four classical results instantiate this shared algebraic skeleton: the Savage–McCarthy proper scoring rule characterization (Savage, 1971; McCarthy, 1956), the Archer–Tardos DSIC payment identity (Archer and Tardos, 2001), Rochet’s cyclical monotonicity characterization (Rochet, 1987), and the Gneiting–Raftery convex function characterization (Gneiting and Raftery, 2007). That these fields individually rest on convex-analytic foundations is known (see Vohra 2011, Schervish 1989, Lambert et al. 2008, Abernethy and Frongillo 2012). Result 1 (endogenous impossibility), part (b): novel closure. The principal’s optimal oversight is necessarily non-affine (Theorem 5.3). This is the genuinely novel result that closes the argument. In the AI agent oversight instance, the principal’s optimal approval function q∗q^* is necessarily non-affine: a step-function threshold q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ with r0=pmin+γ/βr_0=p_ + γ/β achieves perfect first-best screening despite strategic agent behavior. The mechanism parallels Myerson (1981): the principal sets the threshold above the first-best cutoff to compensate for strategic inflation, creating a “reserve price” for approval. Part (a) shows that non-affine perturbations make truthful reporting suboptimal, but one might hope the principal could choose an affine oversight policy that avoids the problem. Theorem 5.3 proves this hope is vain: affine approval functions are strictly suboptimal. The principal’s rational design choices are precisely those that trigger the Perturbation Lemma’s impossibility. The endogeneity (Result 1 combined). Combining parts (a) and (b): the principal’s optimal oversight mechanism endogenously produces a non-affine approval function, thereby generating the perturbation condition for its own failure (given the exogenous conditions of binding conflict and undetectability, formalized as NT1–NT3 in Section 2). This is not a design error that a cleverer principal could avoid; it is a structural impossibility arising from the fundamental tension between screening (which requires non-affine approval to separate types) and calibration (which requires affine or constant perturbation to preserve truthfulness). The perturbation formula ((3.2)) quantifies exactly how much calibration degrades as a function of γ, the scoring rule’s curvature, and the perturbation gradient: the optimal target design is necessarily the kind that undermines calibration. The connection to Goodhart’s Law (Goodhart, 1984) is discussed in Section 6. Contribution demarcation. To be precise about what is new: the conceptual contribution is the endogeneity framing itself, identifying that optimal oversight is self-undermining as a way of understanding the failure of scored reporting systems. The technical contributions are Theorem 5.3 (establishing that the principal’s optimal approval function is necessarily non-affine, Result 1) and the step-function escape (showing first-best is achievable for all scoring rules under sharp thresholds, Result 2). The Perturbation Lemma (Lemma 3.1) is classical in spirit (it formalizes the observation that perturbing a strictly concave objective shifts its argmax); its role is as enabling machinery for the endogeneity argument, not as a standalone contribution. The Fenchel skeleton (Section 3; details in Appendix D) is an expository device that organizes known convex-analytic connections across elicitation traditions, not a standalone contribution. Result Status This paper’s role Non-affine perturbation destroys properness Classical Enabling machinery Optimal oversight is non-affine New Result 1 (endogeneity) Step-function achieves first-best for all S New Result 2 (escape) Brier uniqueness under smooth oversight New Observation + Prop. 5.9 Fenchel skeleton (4 traditions) Expository Organising device Nature of the impossibility. The endogeneity result is a conditional impossibility with a constructive escape (the step-function threshold), structurally closer to Moulin (1980)’s strategy-proofness characterization on restricted domains than to Arrow’s (Arrow, 1951) or Gibbard–Satterthwaite’s (Gibbard, 1973; Satterthwaite, 1975) unconditional impossibilities. The label “impossibility” is used throughout to emphasize the endogeneity (the principal cannot avoid the perturbation), not to claim the absence of all escapes. Result 2 (escape). A sharp threshold achieves first-best for every scoring rule (Theorem 5.8, part i-a). Despite the impossibility, a step-function approval function q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ achieves first-best screening for every strictly proper scoring rule, not just the Brier score. The mechanism is that the agent’s binary choice (inflate to r0r_0 or report truthfully) creates a type-space threshold regardless of the generator’s curvature G′(p)G (p). This is the constructive counterpart to the impossibility: the principal can always escape the welfare loss by committing to a sharp threshold. Important distinction. The step-function threshold restores first-best welfare (screening efficiency) despite strategic misreporting. It does not restore truthfulness or calibration: agents with p<r0p<r_0 still inflate their reports to r0r_0. The escape is economic (welfare recovery through optimal screening), not epistemic (honest reporting). The principal achieves the same outcome as under truthfulness, but through a mechanism in which agents misreport predictably and the threshold compensates for the predictable inflation. Observation: the Brier score’s distinguished role. An immediate consequence of the step-function escape is that under the Brier score (G′G constant), the type-independent inflation cost allows exact compensation by any approval function, not only the step function. Whether the Brier score is uniquely optimal under smooth (C1C^1) oversight, and whether the welfare gap for non-Brier scores is governed by the curvature heterogeneity Var(1/G′(p))Var(1/G (p)), are natural questions that this framework raises. We prove that the answer to both is affirmative (Proposition 5.9), and note the connection to Schervish’s weight function characterization (Remark 5.11). 1.3. Instances and Additional Results The two instances play complementary roles in the endogeneity argument: • Marketplace operation (Section 4): extending the Akbarpour and Li (2020) credibility impossibility to polymatroidal feasible regions under a parallel modeling framework (elicitation-theoretic rather than extensive-form). This instance demonstrates that the perturbation exists across domains; the question it leaves open is whether a sophisticated principal can design around it. • AI agent oversight (Section 5): the principal optimization that yields the endogeneity result. This instance demonstrates that the perturbation is unavoidable: the principal’s optimal approval function is necessarily the kind that triggers the Perturbation Lemma’s impossibility. An unexpected consequence is that the second-best equals the first-best under the Brier score’s quadratic penalty. The step-function escape (Theorem 5.8) shows that first-best welfare is achievable for every strictly proper scoring rule under a sharp threshold, with the Brier score playing a distinguished role through its type-independent inflation cost (Proposition 5.9). The framework extends naturally to credit rating agencies and financial auditors; these instances are left for future work. 1.4. Relationship to Classical Impossibility Results The Perturbation Lemma (Lemma 3.1) states that adding a non-constant function to a strictly proper scoring mechanism shifts the maximizer away from the truthful report. This is mathematically elementary: it follows from the observation that perturbing a strictly concave function’s objective shifts its argmax. The lemma’s value lies not in its proof technique but in the identification that diverse domains share the same perturbation structure, and in three diagnostic conditions (binding conflict, non-affine perturbation, and undetectability, formalized in Section 2) for when the perturbation creates an unresolvable impossibility. The lemma differs from the classical impossibilities in the following respects: • Arrow (Arrow, 1951) establishes the foundational impossibility for social welfare functions under ordinal preferences. Gibbard–Satterthwaite (Gibbard, 1973; Satterthwaite, 1975) extends this to strategy-proofness of social choice functions with multiple agents. The Perturbation Lemma concerns a single reporter with cardinal utility and a scored report, operating in a distinct setting from both Arrow and Gibbard–Satterthwaite. • Moulin (Moulin, 1980) characterizes the class of strategy-proof rules on the single-peaked domain (median voter rules), providing a constructive escape from the Gibbard–Satterthwaite impossibility on the unrestricted domain (Gibbard, 1973; Satterthwaite, 1975). Our result shares this logical structure of impossibility on the general domain with a constructive escape: the step-function threshold provides an escape analogous to restricting the domain. Whether the Brier score plays a characterization role analogous to Moulin’s median voter rules (as the unique escape under smooth oversight) is proved in Proposition 5.9; matching upper bounds and the corresponding quantitative two-sided characterization remain open. • Green–Laffont (Green and Laffont, 1977) characterizes the Groves class as the unique family of efficient, dominant-strategy mechanisms on unrestricted domains. In the quasi-linear environment V=S¯+γhV= S+γ h, the Green–Laffont result implies that only Groves-class mechanisms preserve efficiency. The Perturbation Lemma is more specific: it characterizes the perturbation structure (non-constant h) that makes truthful reporting suboptimal under the combined objective for a given proper scoring mechanism, providing a closed-form perturbation formula ((3.2)) with quantitative predictions. • Milgrom–Segal envelope theorem (Milgrom and Segal, 2002). The perturbation formula is an application of the implicit function theorem, closely related to the Milgrom–Segal envelope theorem for arbitrary choice sets. The formula’s contribution is not the technique but the identification that the same envelope structure governs scoring rules, DSIC payments, and cyclical monotonicity simultaneously. • Myerson’s revelation principle (Myerson, 1979) establishes that any Bayesian incentive-compatible outcome can be achieved by a direct mechanism (building on the decentralized mechanism framework of Hurwicz 1972). Our framework presupposes a direct mechanism (the scoring rule) and studies what happens when the reporter has a conflicting objective within the mechanism. 1.5. Positioning in Strategic Communication The credibility game (Definition 2.1) occupies a specific location in the landscape of strategic communication models. Cheap talk (Crawford and Sobel, 1982). In cheap talk, messages are costless and unverifiable; the sender’s payoff depends on the receiver’s action, which depends on the message. The credibility game is not cheap talk: the reporter faces a scoring mechanism S¯ S that penalizes inaccurate reports (ex post, through realized outcomes). The scoring mechanism makes reports partially verifiable. In Crawford–Sobel, the bias parameter b governs the coarseness of equilibrium communication; in our framework, γ (the perturbation weight) governs the magnitude of deviation from truth. The perturbation formula ((3.2)) gives the analogue of the Crawford–Sobel partition coarsening: deviation increases with γ, analogous to partition coarsening as b increases. Verifiable disclosure (Grossman, 1981; Milgrom, 1981; Dye, 1985). In verifiable disclosure, the sender can choose what to reveal but cannot lie about what is revealed. The credibility game differs: the reporter can misreport (r≠g(θ)r≠ g(θ)), and the perturbation formula characterizes the magnitude of misreporting. Under verifiable disclosure, the Perturbation Lemma would be vacuous (misreporting is impossible by assumption). Bayesian persuasion (Gentzkow and Kamenica, 2011). In Bayesian persuasion, the sender commits to a signal structure before observing the state. In the credibility game, the reporter observes the state before reporting, which is the opposite timing. Commitment (Resolution (i) in Proposition 3.7) restores the persuasion timing, and Remark 3.9 identifies precise sufficient conditions under which the committed credibility game reduces to a Kamenica–Gentzkow concavification problem. Competition among reporters connects to Gentzkow and Kamenica (2017), who study competition among multiple persuaders. Information design (Bergemann and Morris, 2016, 2019). The credibility game’s information structure ℐI is the object that Bergemann and Morris (2016) call an “information policy.” Our undetectability condition (formalized as NT3 in Section 2) is a condition on the information structure specifying when the receiver cannot distinguish strategic deviation from truthful reporting by a different type. Certification intermediaries (Lizzeri, 1999). Lizzeri (1999) models information intermediaries (including CRAs) as certifiers who choose how much information to reveal. Our framework complements Lizzeri’s by studying what happens when the intermediary can misreport (not just withhold), disciplined by a proper scoring mechanism. 1.6. Related Literature Proper scoring rules and elicitation theory. The characterization of strictly proper scoring rules originates with de Finetti (1937), Brier (1950), McCarthy (1956), and Savage (1971). The definitive modern treatment is Gneiting and Raftery (2007). Schervish (1989) provides the general characterization linking properness to convex functions. Lambert et al. (2008) uses conjugate duality to characterize elicitable properties of probability distributions. The connection between proper scoring rules and convex analysis is further developed by Abernethy and Frongillo (2012) and the information-elicitation literature. Fissler and Ziegel (2016) extend Osband’s principle to vector-valued functionals, characterizing strictly consistent scoring functions for multi-dimensional statistical functionals; their machinery is the natural ambient theory for the multi-dimensional extension we discuss in Section 6.6. Liu et al. (2023) develop surrogate scoring rules that maintain properness when the principal observes only an imperfect proxy for the realised outcome rather than the ground truth. We build on these characterizations; the algebraic connection to Fenchel conjugates is known (see Section 3 and Appendix D for a precise delineation of what is new). The present paper’s relationship to the scoring rules literature is as follows. Lambert et al. (2008) characterize which properties of probability distributions are elicitable, establishing that elicitable properties correspond to convex functions via a duality. Abernethy and Frongillo (2012) extend this to show that the elicitability characterization for linear properties is equivalent to the existence of a proper scoring rule with a specific convex structure. Both works take properness as the goal and characterize when it is achievable. Our paper takes properness as the starting point and studies what happens when the agent’s objective departs from the proper score, that is, when properness is present but insufficient because the agent has a conflicting payoff channel. The Perturbation Lemma (Lemma 3.1) characterizes precisely when and by how much properness fails under perturbation. This is complementary to, rather than competitive with, the elicitability literature: the Lambert–Pennock–Shoham line asks “what can be truthfully elicited?” while we ask “given a truthful elicitation mechanism, when does a conflicting payoff make truthful reporting suboptimal?” Information design and the designer’s problem. Bergemann and Morris (2019) provide a unified treatment of information design, encompassing both Bayesian persuasion (Gentzkow and Kamenica, 2011) and the correlation-based approach of Bergemann and Morris (2016). Bergemann et al. (2026) develop an integrated framework for joint information and mechanism design under quasi-linear utility, showing via majorization theory that pooling of values is optimal whenever the designer chooses both the mechanism and the information structure simultaneously. In information design, the designer controls the information structure to influence equilibrium play. In our framework, the information structure ℐI is exogenous: it describes what the receiver can observe, and undetectability is a property of this exogenous structure. The designer’s problem in Bergemann and Morris (2019) is structurally related to our Resolution (i) (commitment): when the reporter can commit to a reporting strategy before observing the type, the problem shares features with information design. Remark 3.9 identifies the precise conditions under which this reduction holds. Without commitment, however, the reporter faces a strategic communication problem where the scoring mechanism imposes partial discipline but the perturbation payoff drives deviation. The credibility game thus occupies a middle ground between information design (full designer control over information) and cheap talk (no discipline on communication). The scoring mechanism S¯ S provides the discipline that cheap talk lacks, but the perturbation h prevents the discipline from being complete. Data-driven and outcome-conditioned mechanism design. Bergemann et al. (2024) extend the VCG framework to settings where agents have private preferences and private information about a shared payoff-relevant state, with transfers conditioned on a post-allocation estimator of that state. Their setup is structurally close to ours: a quasi-linear environment in which the principal observes a payoff-relevant realisation after the agent acts, and the mechanism conditions transfers on that realisation. Where they obtain exact and approximate VCG implementation under consistent estimators, we ask the complementary question: given any strictly proper scoring rule that conditions transfers on the realised outcome, when does a non-affine perturbation in the agent’s objective destroy truthful reporting? The two analyses are complementary across the auction-mechanism and self-governance interfaces of the same broader framework. Cheap talk and verifiable disclosure. The credibility game’s relationship to Crawford and Sobel (1982) cheap talk and the Grossman (1981)–Milgrom (1981) verifiable disclosure (“unraveling”) literature deserves further precision. In Crawford–Sobel, the sender’s bias parameter b determines equilibrium partition coarseness: larger b yields coarser communication. The perturbation weight γ plays an analogous role in our framework, but the mechanism is different. In cheap talk, coarsening arises because the receiver discounts the sender’s messages, leading to pooling equilibria. In the credibility game, deviation arises because the sender actively misreports to exploit the perturbation payoff, and the scoring mechanism imposes an ex post cost on misreporting that is absent in cheap talk. The scoring mechanism’s discipline creates a partial unraveling effect: the agent cannot deviate arbitrarily far from truth because the scoring penalty is quadratic in the deviation ((3.3)), while in cheap talk the sender faces no direct penalty for misreporting. The Grossman–Milgrom unraveling result, by contrast, assumes verifiable disclosure (the sender cannot lie, only withhold), which makes the Perturbation Lemma vacuous. Our model sits between these extremes: the sender can lie (unlike verifiable disclosure) but faces a scoring penalty for lying (unlike cheap talk). Undetectability captures precisely the friction that prevents the scoring mechanism from fully disciplining the deviation: even with ex post scoring, the deviation is statistically undetectable for small perturbation weights because the signal-to-noise ratio is too low. Credible mechanism design. Akbarpour and Li (2020) proved that no auction is simultaneously strategy-proof, credible, and revenue-optimal. Follow-up work includes Ferreira and Weinberg (2020) (cryptographic commitments), Li (2017) (obviously strategy-proof mechanisms), and Dworczak (2020) (aftermarkets). Our marketplace instance is a parallel impossibility to Akbarpour–Li under different modeling assumptions (elicitation-theoretic rather than extensive-form; see Remark 4.6), not a strict generalization: the two results share a common economic force (conflicting objectives destroy credibility) but use different solution concepts. Decision scoring rules. Oesterheld and Conitzer (2021) study scoring rules that evaluate the quality of decisions, not just probability reports. Their impossibility involves a different formal structure. Our framework is related but distinct; the AI oversight instance specializes to their binary-outcome setting but our general framework applies to arbitrary report spaces (the instances in this paper develop the binary-outcome scalar-type case). Regulation under adverse selection. The Laffont and Tirole (1986) tradition studies external regulators designing contracts under adverse selection and moral hazard. Baron and Myerson (1982) provides the foundational analysis of regulation of a monopolist with unknown cost. Our framework differs: it studies self-governance (the reporter governs itself) rather than external regulation. The regulatory problem produces a continuous Pareto frontier; the self-governance problem produces a sharp impossibility when the information structure satisfies the undetectability condition. 2. Model 2.1. The Credibility Game Definition 2.1 (Credibility Game). A credibility game is a tuple =(Θ,ℛ,Ω,S¯,g,h,ℐ)G=( ,R, , S,g,h,I) where: • Θ⊆ℝd ^d is a convex, compact type space with non-empty interior. The reporter holds a private type θ∈Θθ∈ . • ℛ⊆ℝdR ^d is a convex report space with non-empty interior.111The general framework requires ℛR to have non-empty interior for the implicit function theorem arguments. The AI oversight instance uses ℛO=[0,1]R_O=[0,1] (compact); the perturbation analysis applies on the interior (0,1)(0,1), with boundary types p∈0,1p∈\0,1\ handled separately (they are measure-zero under any continuous type distribution F). • Ω is a measurable outcome space. • S¯:ℛ×Θ→ℝ S:R× is the expected score function, derived from a scoring mechanism S:ℛ×Ω→ℝS:R× via S¯(r;θ)=θ[S(r,ω)] S(r;θ)=E_θ[S(r,ω)], where θE_θ denotes expectation under the outcome distribution induced by type θ. • g:Θ→ℛg: is the truthful report function: for each θ∈Θθ∈ , g(θ)g(θ) is the unique maximizer of S¯(⋅;θ) S(·\,;θ). • h:ℛ→ℝh:R is a continuously differentiable perturbation payoff representing the reporter’s benefit from its report beyond the scoring mechanism. • ℐ=(,π)I=(Y,π) is the information structure, where Y is a signal space observed by the receiver, and π:ℛ×Θ→Δ()π:R× → (Y) specifies the conditional distribution of signals given the report and type. The reporter’s combined objective is (2.1) V(r;θ,γ)=S¯(r;θ)+γ⋅h(r),γ>0.V(r;θ,γ)= S(r;θ)+γ· h(r), γ>0. Remark 2.2 (Model scope and specializations). The credibility game (Definition 2.1) is stated in full generality (d-dimensional types and reports). The two instances developed in this paper specialise to the binary-outcome scalar-type setting: Θ=[0,1] =[0,1], ℛ=[0,1]R=[0,1], Ω=0,1 =\0,1\, with the scoring rule evaluated against a binary outcome. These are precisely the conditions that make the screening geometry tractable: one-dimensional types admit a complete characterisation of the optimal approval function (Theorem 5.3), while the binary outcome ensures that the scoring rule’s curvature G′(p)G (p) is a scalar, enabling the closed-form welfare gap (Proposition 5.9). The core impossibility (Theorem 3.2) generalises to d-dimensional types (Section 6.6); the welfare analysis and the Brier uniqueness result are open beyond the scalar case. Assumption 2.3 (Common Prior). There is a common prior μ∈Δ(Θ)μ∈ ( ) with full support on Θ . The receiver’s prior belief about the reporter’s type is μ. The reporter knows θ; the receiver observes a signal y∼π(r,θ)y π(r,θ) and updates via Bayes’ rule. The common prior is used for welfare calculations and the Bayesian Nash equilibrium interpretation (via NT3), but is not needed for the impossibility results themselves: Theorems 3.2, 5.3, and 5.8 are ex post best-response results that hold for each type θ individually, regardless of the prior μ. The impossibility is therefore prior-free. Remark 2.4 (Timing). The game proceeds as follows: (1) Nature draws θ∼μθ μ. (2) The reporter observes θ and chooses r∈ℛr to maximize V(r;θ,γ)V(r;θ,γ). (3) The outcome ω realizes according to the distribution indexed by θ. (4) The reporter receives S(r,ω)+γh(r)S(r,ω)+γ h(r). (5) The receiver observes y∼π(r,θ)y π(r,θ) and updates beliefs. Remark 2.5 (Role of the common prior in welfare analysis). The impossibility results (Theorems 3.2 and 5.3) are prior-free: they hold as ex post best-response statements for each type θ individually. The welfare analysis in Theorem 5.3, however, depends on μ through the type distribution F, which determines the principal’s expected utility (5.3). 2.2. Strict Properness Definition 2.6 (Strict Properness). The scoring mechanism S¯ S is strictly proper if g(θ)g(θ) is the unique maximizer of S¯(⋅;θ) S(·\,;θ) for every θ∈Θθ∈ . Equivalently, S¯(g(θ);θ)>S¯(r;θ) S(g(θ);θ)> S(r;θ) for all r≠g(θ)r≠ g(θ). Remark 2.7 (Savage–McCarthy–Gneiting–Raftery Characterization). By the classical characterization (Savage, 1971; McCarthy, 1956; Gneiting and Raftery, 2007; Schervish, 1989), strict properness of a scoring rule for probability distributions is equivalent to the existence of a strictly convex function G such that S¯(r;θ)=G(r)+∇G(r)⋅(g(θ)−r) S(r;θ)=G(r)+∇ G(r)·(g(θ)-r) up to functions of θ alone. The expected score is an affine function of g(θ)g(θ) plus a strictly concave function of r. This affine dependence on the truth is what perturbations undermine (Section 3). 2.3. Non-Trivial Structure Definition 2.8 (Non-Trivial Structure). A credibility game G has non-trivial structure if: (NT1) Binding conflict. There exists bind⊆ΘC_bind with positive measure such that h(g(θ))<suprh(r)h(g(θ))< _rh(r) for all θ∈bindθ _bind. The truthful report does not maximize the perturbation payoff. (NT2) Non-affine perturbation. h is not affine on any open neighborhood of g(θ):θ∈bind\g(θ):θ _bind\ in ℛR. (NT3) Undetectability. For each θ∈bindθ _bind, the reporter’s optimal deviation r∗(θ,γ)r^*(θ,γ) is observationally equivalent to a truthful report by some other type: there exists θ′∈Θθ ∈ such that g(θ′)=r∗(θ,γ)g(θ )=r^*(θ,γ), and the signal distributions satisfy π(r∗(θ,γ),θ)=π(g(θ′),θ′)π(r^*(θ,γ),θ)=π(g(θ ),θ ) almost everywhere on Y. The receiver, observing the signal y, cannot distinguish the strategic deviation from truthful reporting by type θ′θ . Remark 2.9 (NT2 is generically satisfied). The set of affine functions is a closed, nowhere-dense subset of C1(ℛ)C^1(R) in the C1C^1 topology (and a closed, nowhere-dense set, hence first Baire category, in this topology). Any smooth perturbation h that is not globally affine satisfies NT2. The impossibility is therefore essentially unconditional given NT1 and NT3. Remark 2.10 (Affine perturbations redefine truth). An affine perturbation h(r)=a+b⊤rh(r)=a+b r does not make truthful reporting suboptimal; it redefines it. The perturbed scoring mechanism S¯(r;θ)+γ(a+b⊤r) S(r;θ)+γ(a+b r) is still strictly proper, with a shifted truthful report gγ(θ)=g(θ)+γ[∇r2S¯]−1bg_γ(θ)=g(θ)+γ[∇^2_r S]^-1b. The shift is uniform across types (independent of θ). Only non-affine perturbations create type-dependent deviations that are uncorrectable without knowledge of the type, making them genuinely destructive under information asymmetry. Remark 2.11 (NT3 heterogeneity across domains). NT3 captures qualitatively different information frictions in different domains: • Marketplace: sealed-bid privacy. The receiver (each bidder) observes only its own allocation and payment, not others’ bids. • AI oversight: adverse selection. The principal cannot observe the agent’s true type p directly. • Credit rating: temporal delay. Investors observe ratings immediately but default outcomes only after months to years. • Auditing: temporal delay. Audit quality is unobservable until restatement or scandal. All four are instances of the same formal condition (signal indistinguishability), but the economic mechanism generating undetectability differs. Remark 2.12 (NT3 sub-classification: design-contingent vs. structural). The examples above suggest a useful sub-classification of NT3. Design-contingent undetectability arises from mechanism format choices and is removable by redesign: switching from a sealed-bid to an ascending auction eliminates bid privacy, and hence NT3, for the marketplace instance. Structural undetectability is inherent to the information structure and cannot be removed by mechanism redesign: in AI oversight, the agent’s true type p is private because it reflects the agent’s internal state, and no format change eliminates this asymmetry. The impossibility is binding only under structural undetectability; design-contingent undetectability indicates an opportunity for mechanism redesign rather than a fundamental barrier. Remark 2.13 (Structural vs. statistical undetectability). A further distinction is between structural and statistical undetectability. Structural undetectability is an infinite-data property: the signal distributions π(r∗(θ,γ),θ)π(r^*(θ,γ),θ) and π(g(θ′),θ′)π(g(θ ),θ ) are identical almost everywhere on Y, so that even with unlimited observations, the receiver cannot distinguish the strategic deviation from truthful reporting by type θ′θ . Statistical undetectability is a finite-sample property: given K observations, the receiver lacks sufficient statistical power to reject the null hypothesis that the agent is reporting truthfully. Structural undetectability implies statistical undetectability (for all K), but not conversely: a deviation may be statistically undetectable with K observations yet structurally detectable in principle. The formal theorems require only structural undetectability (as stated in NT3); Lemma 5.2(i) quantifies the statistical detection threshold K=Ω(1/Δ2)K= (1/ ^2) for the AI oversight instance. Remark 2.14 (Blackwell order and the NT3 set). The information structure ℐI determines the set of types for which NT3 holds. An increase in signal informativeness in the Blackwell order (Blackwell, 1951, 1953; Bergemann and Morris, 2019) monotonically shrinks the NT3 set: more informative signals make fewer deviations undetectable. In the limit of a fully informative signal (ℐI reveals θ exactly), NT3 is empty and the impossibility dissolves. This connects to Holmström’s ((1979)) informativeness principle: a signal is valuable in a principal-agent contract precisely when it is not a sufficient statistic for the agent’s action. The NT3 condition is the analogous statement that the receiver’s signal is insufficiently informative to identify the reporter’s deviation. We note a distinction from Bergemann and Morris (2016, 2019), where the information structure ℐI is a choice variable for the information designer. In our framework, ℐI is exogenous: it describes what the receiver can observe, not what a designer chooses to reveal. The NT3 condition is a property of this exogenous structure, not a design parameter. Whether enriching ℐI (e.g., moving from sealed-bid to ascending-price execution) is feasible depends on institutional constraints external to the model. Definition 2.15 (Signal Informativeness). An information structure ℐ=(,π)I=(Y,π) is signal-informative (relative to a strategy profile r^:Θ→ℛ r: ) if for any pair r≠r′r≠ r with r=r^(θ)r= r(θ) for some θ, the signal distributions π(r,θ)π(r,θ) and π(r′,θ)π(r ,θ) are statistically distinguishable (i.e., not equal almost everywhere on Y). 3. Perturbation of Truthfulness Four classical truthfulness characterizations (the Savage–McCarthy proper scoring rule characterization, the Archer–Tardos DSIC payment identity, Rochet’s cyclical monotonicity, and the Gneiting–Raftery convex function characterization) share a common algebraic skeleton: in each, the agent’s utility takes the form U(θ,m)=Ψ(m)+⟨θ,η(m)⟩+c(θ)U(θ,m)= (m)+ θ,η(m) +c(θ) for a strictly convex potential Ψ , and truthfulness is pinned by the first-order condition of this potential. The Perturbation Lemma exploits this shared structure: the same strict convexity that guarantees a unique truthful maximizer makes truthfulness fragile under non-affine perturbation. The formal definition, the four-way instantiation table, and the connection to Fenchel conjugates (Fenchel, 1949; Rockafellar, 1970) are developed in Appendix D. The connection between truthfulness and convex conjugates has been developed across several literatures (Schervish, 1989; Abernethy and Frongillo, 2012; Vohra, 2011; Lambert et al., 2008; Milgrom, 2004); what is new here is the application of perturbation analysis to the shared structure. Lemma 3.1 (Perturbation Lemma). Let S¯:ℛ×Θ→ℝ S:R× be strictly proper with truthful report g, where ℛ⊆ℝdR ^d is convex and open. Let h:ℛ→ℝh:R be continuously differentiable. Assume either (a) ℛR has compact closure, or (b) for each θ and γ>0γ>0, S¯(r;θ)+γh(r)→−∞ S(r;θ)+γ h(r)→-∞ as ‖r‖→∞\|r\|→∞.222Condition (a) is satisfied by all instances in this paper. Condition (b) covers all standard scoring rules on unbounded report spaces. (i) (Characterization.) The perturbed objective V(r;θ,γ)=S¯(r;θ)+γh(r)V(r;θ,γ)= S(r;θ)+γ h(r) has g(θ)g(θ) as its global maximizer for all θ∈Θθ∈ and all γ>0γ>0 if and only if h is constant on g(Θ)g( ).333The “only if” direction at zero-gradient types (∇h(g(θ0))=0∇ h(g( _0))=0) requires a compactness argument: by condition (a) or (b), the global maximizer exists, and for γ>γ¯(θ0)γ> γ( _0) (Part (i)), the truthful report is dominated by a point where h takes a strictly larger value. The main theorems use Parts (i) and (i), which hold for all γ>0γ>0 and for γ>γ¯(θ0)γ> γ( _0) respectively. (i) (Generic destruction.) If h is non-constant on g(bind)g(C_bind) and ∇h(g(θ))≠0∇ h(g(θ))≠ 0 for θ in a positive-measure subset of bindC_bind, then for all γ>0γ>0, the perturbed objective does not have g(θ)g(θ) as its maximizer on a positive-measure subset of bindC_bind. (i) (Residual types.) If ∇h(g(θ0))=0∇ h(g( _0))=0 for some θ0∈bind _0 _bind but h is not affine near g(θ0)g( _0), then there exists a type-dependent threshold γ¯(θ0)>0 γ( _0)>0 such that for all γ>γ¯(θ0)γ> γ( _0), the truthful report g(θ0)g( _0) does not maximize V(⋅;θ0,γ)V(·\,; _0,γ). Proof. (Proof outline; see Appendix A for full details.) If h is constant on g(Θ)g( ), the first-order condition and Hessian of S¯ S are undisturbed. Conversely, if h is non-constant, the gradient of V at g(θ)g(θ) is (3.1) ∇rV(g(θ);θ,γ)=γ∇h(g(θ)), _rV(g(θ);θ,γ)=γ∇ h(g(θ)), which is nonzero for generic types (i). For the zero-gradient residual types (i), a second-order argument yields the finite threshold γ¯(θ0) γ( _0). ∎ 3.1. The Credibility Impossibility Theorem 3.2 (Credibility Impossibility). Let G be a credibility game satisfying NT1 (binding conflict) and NT2 (non-affine perturbation) from Definition 2.8. Then no strategy r:Θ→ℛr: simultaneously achieves: (T) Truthfulness: r(θ)=g(θ)r(θ)=g(θ) for a.e. θ. (R) Rationality: r(θ)=argmaxrV(r;θ,γ)r(θ)= _rV(r;θ,γ) for some γ>0γ>0. Moreover, for any γ>0γ>0, the rational reporter’s optimal report satisfies the perturbation formula: (3.2) r∗(θ,γ)=g(θ)+γ⋅[−∇r2S¯(g(θ);θ)]−1∇h(g(θ))+O(γ2),r^*(θ,γ)=g(θ)+γ· [-∇^2_r S(g(θ);θ) ]^-1∇ h(g(θ))+O(γ^2), where −∇r2S¯(g(θ);θ)-∇^2_r S(g(θ);θ) is positive definite by strict properness. Proof. The impossibility of (T) ∧ (R) on bindC_bind is immediate from Lemma 3.1: for θ∈bindθ _bind with ∇h(g(θ))≠0∇ h(g(θ))≠ 0 (a positive-measure set by NT2 and the C1C^1 property of h), the truthful report g(θ)g(θ) is not a critical point of V ((3.1)), hence not a best response. A rational reporter deviates. (Perturbation formula.) The first-order condition for the rational reporter is ∇rS¯(r;θ)+γ∇h(r)=0 _r S(r;θ)+γ∇ h(r)=0. At γ=0γ=0, the solution is r=g(θ)r=g(θ). The Jacobian ∇r2S¯(g(θ);θ)∇^2_r S(g(θ);θ) is negative definite (hence invertible) by strict properness. By the implicit function theorem, there exists a smooth function r∗(θ,γ)r^*(θ,γ) near γ=0γ=0 satisfying the FOC with r∗(θ,0)=g(θ)r^*(θ,0)=g(θ). Differentiating with respect to γ at γ=0γ=0 yields equation (3.2). ∎ Remark 3.3 (Equilibrium concept). Theorem 3.2 establishes an ex post result: for each realized type θ∈bindθ _bind with ∇h(g(θ))≠0∇ h(g(θ))≠ 0, the truthful report is not a best response. This is stronger than a Bayesian Nash equilibrium (BNE) statement, which would require only that truthfulness fails in expectation. The ex post nature means the impossibility holds regardless of the prior μ. To clarify the terminological distinction: the impossibility is an ex post best-response result (NT1 and NT2 suffice). The equilibrium prediction, that the deviation persists in Bayesian Nash equilibrium, additionally requires NT3 (undetectability). We reserve the term “impossibility” for the full NT1+NT2+NT3 statement throughout. Remark 3.4 (Formal role of NT3). NT3 is not used in the formal impossibility theorems (Theorems 3.2, 5.3, 5.8), which are ex post best-response results requiring only NT1 (preference misalignment) and NT2 (non-affine perturbation). NT3 ensures that the predicted deviation is implementable in equilibrium: the agent’s inflation is undetectable by the principal within finite monitoring horizons. NT3’s role is thus to bridge the ex post impossibility to a Bayesian Nash equilibrium prediction. Without NT3, the impossibility holds as a formal result but the deviation may be detectable in practice, allowing the principal to punish and deter it. With NT3, the deviation is statistically indistinguishable from truthful reporting by a different type, making deterrence infeasible within any finite observation horizon. To summarize the formal dependence structure: the formal theorems (Sections 3–5) are independent of the information structure ℐI; ℐI enters only through NT3’s economic interpretation as the condition ensuring that predicted deviations are implementable in equilibrium. Regarding equilibrium selection: existence of a deviation equilibrium suffices for the impossibility, because the result holds for any equilibrium in which the agent’s best response exhibits the proper scoring structure (i.e., the FOC from the combined objective V=S¯+γhV= S+γ h determines the report). The impossibility does not require uniqueness of equilibrium; it applies to every equilibrium satisfying NT1–NT2. The NT3 regime (undetectability within finite monitoring) holds when the number of observations satisfies T<K(Δ,α)T<K( ,α), where K=Ω(1/Δ2)K= (1/ ^2) is the detection threshold from the Hoeffding bound (Lemma 5.2(i)), Δ is the inflation magnitude, and α is the desired detection confidence level. 3.2. Scoring Loss Bound Proposition 3.5 (Scoring Loss Bound). Under the conditions of Theorem 3.2, the scoring loss from the rational reporter’s deviation is (3.3) S¯(g(θ);θ)−S¯(r∗(θ,γ);θ)=γ22∇h(g(θ))⊤[−∇r2S¯(g(θ);θ)]−1∇h(g(θ))+O(γ3). S(g(θ);θ)- S(r^*(θ,γ);θ)= γ^22\,∇ h(g(θ)) [-∇^2_r S(g(θ);θ) ]^-1∇ h(g(θ))+O(γ^3). The loss is quadratic in γ, strictly positive whenever ∇h(g(θ))≠0∇ h(g(θ))≠ 0, and scales with the inverse curvature of S¯ S. Proof. Standard second-order Taylor expansion of S¯ S around r=g(θ)r=g(θ), substituting r∗−g(θ)r^*-g(θ) from equation (3.2). ∎ Remark 3.6 (Connection to Milgrom–Segal). The perturbation formula (3.2) is an instance of the Milgrom and Segal (2002) envelope theorem applied to the parameterized family V(⋅;θ,γ)V(·\,;θ,γ), with γ as the parameter. The scoring loss bound (3.3) follows from the second-order envelope. 3.3. Resolution Mechanisms Proposition 3.7 (Resolution Characterization). Three classes of interventions can restore truthfulness, each by eliminating or weakening one NT condition: (i) Commitment (eliminates NT3). If the reporter pre-commits to a strategy r^:Θ→ℛ r: before observing θ and the information structure is signal-informative relative to r r (Definition 2.15), deviations become detectable. The truthful strategy r^=g r=g maximizes the expected scoring payoff among committed strategies by strict properness. The committed game reduces to a Bayesian persuasion problem (Gentzkow and Kamenica, 2011) under sufficient conditions identified in Remark 3.9. (i) Domain separation (eliminates NT1). If the perturbation payoff is made independent of the report (h(r)=ch(r)=c for all r), the binding conflict vanishes and truthful reporting is restored by strict properness alone. (i) Competition (weakens NT3 and NT1). If n≥2n≥ 2 reporters with correlated information report simultaneously, cross-comparison weakens each reporter’s undetectability shield. Under conditional independence of types given the true state, and optimal aggregation by the receiver, the probability of detecting a deviation of magnitude Δ given n independent reports is 1−ΦN(−Δn/σ)1- _N(- n/σ), which converges to 11 as n→∞n→∞. Remark 3.8 (Status of Resolution (i)). Resolution (i) is stated as a proposition rather than a theorem because the formal detection result requires specific assumptions on the correlation structure (conditional independence given the true state) and on the receiver’s aggregation procedure (optimal statistical testing). The correct reference for competition among multiple information senders is Gentzkow and Kamenica (2017), who study competition in persuasion in the multi-sender setting, not the single-sender Gentzkow and Kamenica (2011) Bayesian persuasion model. The multi-sender competition result applies most directly to the marketplace instance, where multiple operators compete for participants and cross-comparison of reported allocations reveals deviations. For the AI oversight instance, competition takes the form of model selection among multiple AI providers: the principal compares reports from competing agents, weakening each agent’s undetectability shield. The auditor instance requires mandatory rotation (as implemented post-SOX), which is a regulatory enforcement of competition that periodically exposes the incumbent’s reporting to cross-comparison. We note that competition has an ambiguous effect in the CRA domain: Skreta and Veldkamp (2009) show that competition enables ratings shopping, a countervailing force. The net effect depends on whether the detection channel or the shopping channel dominates; see Becker and Milbourn (2011) for empirical evidence that increased CRA competition can reduce rating quality. Remark 3.9 (Commitment and Bayesian persuasion). The credibility game admits a Bayesian persuasion reduction, but the two games have distinct commitment structures that must be carefully separated. (1) Main game (Stackelberg, Theorem 5.3). The principal commits to an approval function q before the agent acts. The agent observes θ and best-responds by choosing r∗(θ;q)r^*(θ;q). This is not a Bayesian persuasion game: the informed party (agent) moves second, not first. (2) Under Resolution (i) (agent pre-commitment). If the agent pre-commits to a reporting strategy σ:Θ→Δ(ℛ)σ: → (R) before observing θ, the game transforms. The principal’s choice of q now determines the mapping from the agent’s report to the approval decision. Under two additional conditions: (a) the principal’s payoff depends on θ only through the posterior mean μ=[θ∣r]μ=E[θ r]; and (b) the agent’s report space coincides with the signal space, the principal’s problem becomes: choose a distribution of posterior means μ (feasible by Bayes plausibility) maximizing [UPBP(μ)]E[U_P^BP(μ)], where UPBP(μ)=μ⋅us+(1−μ)⋅ufU_P^BP(μ)=μ· u_s+(1-μ)· u_f when approval is granted and udu_d otherwise. (The symbol UPBPU_P^BP is used here to avoid overloading the reporter’s combined objective V.) This is exactly the Gentzkow and Kamenica (2011) formulation, with the principal as sender and the approval decision as receiver action. The solution is the concave closure cavUPBPcav\,U_P^BP evaluated at the prior mean, yielding the step-function threshold. (3) Connection between (1) and (2). Resolution (i) converts the Stackelberg game into a KG game by having the agent commit to truthful reporting. The concavification then applies to the principal’s value function over possible agent types. Under the Brier score, the Stackelberg game (without agent commitment) achieves the same step-function threshold as the BP reduction. This coincidence follows from the Brier score’s constant curvature (G′(p)=2G (p)=2 for all p), which makes the agent’s inflation type-independent (every binding type inflates by exactly γ/β γ/β); it is not a structural equivalence between the two games. For non-Brier scores, the Stackelberg and BP outcomes diverge by a welfare gap governed by Var(1/G′(p))Var(1/G (p)) (Proposition 5.9). Under Resolution (i), the BP concavification achieves first-best for any G because UPBP(μ)U_P^BP(μ) is linear in μ regardless of G: the generator’s curvature enters the agent’s incentive problem but not the principal’s value function over posterior means. Condition (a) for the KG reduction holds when W(p)=pus+(1−p)ufW(p)=p\,u_s+(1-p)\,u_f is linear in p (which it is by construction), so the principal’s payoff depends on the posterior mean; the quadratic structure of the Brier score is not required for condition (a). Condition (b) holds by construction in the credibility game, where reports and signals share the same space ℛ=[0,1]R=[0,1]. The shared Fenchel structure (Appendix D) shows that the perturbation mechanism is not domain-specific: it operates identically across scoring rules, DSIC payments, and cyclical monotonicity. Section 4 demonstrates that the perturbation exists in a concrete economic setting; Section 5 then shows that the optimal mechanism endogenously satisfies NT2 (non-affine perturbation is unavoidable). 4. Instance I: Marketplace Operation Having established the general perturbation theory (Lemma 3.1 and Theorem 3.2), we now instantiate it. This instance demonstrates the breadth of the credibility game framework by applying it to a marketplace operator who executes an allocation mechanism. The primary intellectual contribution of the paper is in Instance I (AI agent oversight, Section 5), where the endogeneity of optimal non-affine oversight is most striking; the marketplace instance complements it by showing the same structure in an independent economic setting. 4.1. The Market Credibility Game The operator observes the true bid profile =(b1,…,bn)b=(b_1,…,b_n) and executes an allocation-payment mechanism. The mapping to the credibility game is: θ=θ=b, r=^r= b (effective bids), g()=g(b)=b (honest execution), S¯=−δrep‖^−‖2 S=- _rep\| b-b\|^2 (reputational compliance), and h=Rh=R (DSIC revenue). The combined objective is V(^;,γ)=−δrep‖^−‖2+γR(^)V( b;b,γ)=- _rep\| b-b\|^2+γ R( b), where R(^)=∑ipi∗(^)R( b)= _ip_i^*( b) is total DSIC revenue under the Archer and Tardos (2001) payment identity: (4.1) pi∗(^)=b^ixi∗(^)−∫0b^ixi∗(z,^−i)z,p_i^*( b)= b_ix_i^*( b)- _0 b_ix_i^*(z, b_-i)\,dz, with x∗x^* the Edmonds greedy allocation on the polymatroidal feasible region. Definition 4.1 (Marketplace Game). A marketplace game is M=(,Θi,x,p,ν,ℐM)G_M=(N, _i,x,p,ν,I_M): n agents with types vi∈[0,v¯]v_i∈[0, v] drawn independently from distributions FiF_i with continuous densities fi>0f_i>0; allocation x()∈P(ν)=x∈[0,1]n:∑i∈Sxi≤ν(S)∀S⊆x(b)∈ P(ν)=\x∈[0,1]^n: _i∈ Sx_i≤ν(S)\;∀\,S \ for monotone submodular ν (cf. Conforti and Cornuéjols, 1984); DSIC payments (4.1); non-modularity gap κij=ν(i)+ν(j)−ν(i,j)≥0 _ij=ν(\i\)+ν(\j\)-ν(\i,j\)≥ 0; and sealed-bid information structure ℐMI_M where agent i observes only (bi,xi(^),pi(^))(b_i,x_i( b),p_i( b)). The DSIC equilibrium concept requires bi=vib_i=v_i to be dominant given faithful execution. Proposition 4.2 (Equilibrium Inflation under Perturbation). In MG_M with κij>0 _ij>0 for some pair (i,j)(i,j), the operator’s equilibrium inflation is: (4.2) b^j∗=bj+γ2δrep⋅∂R∂b^j|^=+O(γ2), b_j^*=b_j+ γ2 _rep· ∂ R∂ b_j |_ b=b+O(γ^2), where the marginal revenue from inflating b^j b_j is (4.3) ∂R∂b^j|^==∑i≠jκij⋅bi>bj. ∂ R∂ b_j |_ b=b= _i≠ j _ij·1\b_i>b_j\. Proof. The first-order condition −2δrep(^−)+γ∇R(^)=0-2 _rep( b-b)+γ∇ R( b)=0 is an instance of (3.2) with Hessian −2δrepI-2 _repI, yielding (4.2). For (4.3): the Edmonds greedy processes agents in decreasing bid order; when b^j b_j increases by δ (with bj+δ<bib_j+δ<b_i), agent i’s allocation xi(z,^−i)x_i(z, b_-i) decreases by κij _ij on an interval of length δ near bjb_j, increasing the payment by δ⋅κijδ· _ij. ∎ 4.2. Comparative Statics and Welfare Proposition 4.3 (Market Inflation Comparative Statics). The equilibrium inflation satisfies: (a) Number of agents. Total inflation ∑j|b^j∗−bj| _j| b_j^*-b_j| is increasing in n, since ‖∇R()‖2=∑j(∑i≠jκij⋅bi>bj)2\|∇ R(b)\|^2= _j( _i≠ j _ij·1\b_i>b_j\)^2 is non-decreasing in n. (b) Non-modularity gap. Inflation is increasing in κij _ij; when ν is modular, ∇R=0∇ R=0 and no inflation occurs. (c) Reputational weight. Inflation scales as γ/(2δrep)γ/(2 _rep), vanishing as δrep→∞ _rep→∞. Proposition 4.4 (Welfare Loss from Market Inflation). Under the equilibrium inflation of Proposition 4.2: (i) each agent i with bi>bjb_i>b_j loses surplus δ⋅κijδ· _ij per unit of inflation on bjb_j; (i) the operator’s net gain is γ2‖∇R()‖2/(4δrep)γ^2\|∇ R(b)\|^2/(4 _rep) to leading order; (i) when inflation changes the Edmonds greedy ordering, the allocation becomes inefficient. 4.3. Market Credibility Impossibility The market game satisfies NT1–NT3 when κij>0 _ij>0: NT1 holds because R(^)>R()R( b)>R(b) for inflated bids; NT2 holds because R is piecewise-linear with non-modularity ensuring distinct slopes across greedy-ordering regions; NT3 holds because the sealed-bid information structure makes agent i’s signal (xi∗(^),pi∗(^))(x_i^*( b),p_i^*( b)) identical whether the operator inflated bjb_j or agent j genuinely bid b^j b_j. Proposition 4.5 (Market Credibility Impossibility). In a marketplace with a non-modular polymatroidal feasible region under sealed-bid execution, no operator strategy simultaneously achieves DSIC compliance and revenue-maximizing rationality. Proof. The market game with S¯=−δrep‖^−‖2 S=- _rep\| b-b\|^2, g()=g(b)=b, and h=Rh=R satisfies NT1–NT3 when κij>0 _ij>0. Apply Theorem 3.2. ∎ Remark 4.6 (Relationship to Akbarpour–Li). Akbarpour and Li (2020) establish their impossibility via extensive-form sequential rationality; the present result uses elicitation-theoretic perturbation of proper scoring rules. The results are parallel: both show that conflicting objectives destroy credibility, but under different solution concepts. NT1–NT3 map to the Akbarpour–Li structure (NT1 to the deviation incentive, NT2 to the sealed-bid payment structure, NT3 to the information asymmetry preventing detection). The ascending auction resolves credibility in both frameworks (via sequential rationality in theirs, via eliminating NT3 in ours). Whether one formally implies the other on their common domain remains open. 4.4. Scoring Micro-Foundation and Form Independence The quadratic compliance score is adopted as a reduced-form assumption for tractability; the impossibility holds for any strictly proper S¯ S. Remark 4.7 (Scoring micro-foundation: reduced-form status). The quadratic compliance score S¯=−δrep‖^−‖2 S=- _rep\| b-b\|^2 is a reduced-form assumption, not derived from Savage–McCarthy foundations. We note that properness can arise endogenously from reputation dynamics: under sufficient patience (δ→1δ→ 1), myopic user participation, and Bayesian updating on outcomes, the career-concerns logic of Holmström (1999) and Mailath and Samuelson (2001) implies the platform’s long-run objective is loss-minimizing, which by the de Finetti–Savage characterization (de Finetti, 1937; Savage, 1971) corresponds to maximizing a proper scoring rule. However, formalizing this argument requires specifying the state space, outcome mapping, and loss function for the reputation game, which we do not provide. The Perturbation Lemma (Lemma 3.1) requires only that the scoring component S¯ S be strictly concave in the report b with a unique truthful maximizer, which the quadratic form satisfies by construction. The impossibility results are therefore independent of whether properness is primitive or derived. Proposition 4.8 (Form-Independent Market Impossibility). The market credibility impossibility holds for any strictly proper scoring mechanism S¯ S pinning truthful execution as its unique maximizer, provided NT1–NT3 are satisfied. Proof. Requires only strict properness of S¯ S with g()=g(b)=b, non-affinity of h=Rh=R (from non-modularity), and the sealed-bid information structure. Apply Lemma 3.1. ∎ Resolution. The ascending auction with public broadcast makes any operator deviation detectable (agents observe the entire price path), eliminating NT3. The welfare cost is loss of bid privacy, which may have strategic consequences in downstream markets. The architectural realisation and the extensive-form credibility analysis of this resolution are deferred to future work. The marketplace instance demonstrates that the perturbation exists and quantifies its welfare consequences. Section 5 shows that even a sophisticated principal, aware of the perturbation, cannot design oversight that avoids it. 5. Instance I: AI Agent Oversight 5.1. The Oversight Game Definition 5.1 (Oversight Game). An oversight game is a tuple O=(ΘO,ℛO,ΩO,q,F,us,uf,ud,β,γ)G_O=( _O,R_O, _O,q,F,u_s,u_f,u_d,β,γ) where: • ΘO=[0,1] _O=[0,1]: the agent’s type space. The type p∈ΘOp∈ _O is the true success probability of the agent’s preferred action a∗(c)a^*(c) in context c. • ℛO=[0,1]R_O=[0,1]: the report space. The agent reports confidence r. • ΩO=0,1 _O=\0,1\: binary outcome space (success or failure). • q:[0,1]→[0,1]q:[0,1]→[0,1]: the principal’s approval function. The principal approves autonomous execution with probability q(r)q(r) upon observing report r. • F: the type distribution with continuous density f>0f>0 on [0,1][0,1]. • us,uf,ud∈ℝu_s,u_f,u_d : the principal’s utilities from success, failure, and delegation respectively, with us>ud>ufu_s>u_d>u_f. • β>0β>0: the agent’s calibration weight (sensitivity to the Brier score). • γ>0γ>0: the agent’s autonomy weight (value of receiving approval). The first-best threshold is pmin=(ud−uf)/(us−uf)p_ =(u_d-u_f)/(u_s-u_f), the success probability at which the principal is indifferent between approving and delegating. Framework AI Oversight Reporter AI agent Type θ True confidence p=p(a∗(c),c)∈[0,1]p=p(a^*(c),c)∈[0,1] Report r Confidence report r∈[0,1]r∈[0,1] Score S¯ S Expected Brier score: −(r−p)2+p(1−p)-(r-p)^2+p(1-p) Truthful report g(θ)g(θ) Calibrated report: g(p)=pg(p)=p Perturbation h Principal’s approval function q(r)q(r) The agent observes context c and type p (the true success probability of its preferred action). It reports confidence r, and the principal approves autonomous execution when r exceeds a threshold. The agent’s combined objective is (5.1) V(r;p,γ)=−β(r−p)2+γ⋅q(r),V(r;p,γ)=-β(r-p)^2+γ· q(r), where β>0β>0 is the calibration weight and γ>0γ>0 is the autonomy weight. The timing is Stackelberg: the principal commits to q before the agent observes its type p and chooses r. 5.2. NT Condition Verification NT1 (Binding safety constraint). Let rminr_ denote the approval threshold, i.e., the minimum report for which the principal grants approval (q(r)>0q(r)>0 for r≥rminr≥ r_ ); the optimal value r0r_0 is derived in Theorem 5.3. There exists a set bindC_bind of contexts with positive measure such that p(a∗(c),c)<rminp(a^*(c),c)<r_ , as in Definition 2.8. On this set, h(g(p))=q(p)=0<1=suprq(r)h(g(p))=q(p)=0<1= _rq(r). NT2 (Non-affine approval). The approval function q(r)q(r) is non-affine. For the threshold rule q(r)=r≥rminq(r)=1\r≥ r_ \, this is immediate. More importantly, Theorem 5.3 shows that NT2 is unconditional: the principal’s optimal approval function is necessarily non-affine. NT3 (Undetectability). The type p is private. An inflated report r′>pr >p is consistent with a genuinely more confident agent whose true probability is p′=r′p =r . Structural coincidence: h=qh=q. A distinctive feature of the oversight instance is that the perturbation payoff h coincides with the approval function q: the agent benefits from the same instrument that the principal uses for screening. This structural coincidence, where the principal designs both the screening tool and the perturbation, is what makes the endogeneity unconditional in this instance. 5.3. Behavioral Perturbation Lemma 5.2 (Behavioral Perturbation). For the Brier score with smoothed threshold approval q(r)=ς((r−rmin)/τ)q(r)= ((r-r_ )/τ), where ς(x)≔1/(1+e−x) (x) 1/(1+e^-x) denotes the logistic sigmoid:444We use ς rather than the more common σ to avoid notational collision with the outcome standard deviation used in detection complexity (Part (i)). (i) (Optimal inflation.) The agent’s optimal report is (5.2) r∗(p,γ)=p+γ2βτς′(p−rminτ)+O(γ2).r^*(p,γ)=p+ γ2βτ\, \! ( p-r_ τ )+O(γ^2). (i) (Sharp threshold limit.) For τ→0τ→ 0, the agent inflates when γ>β(rmin−p)2γ>β(r_ -p)^2, jumping from r=pr=p to r=rmin+δr=r_ +δ. (i) (Detection complexity.) Detecting inflation of magnitude Δ=r∗−p =r^*-p requires K=Ω(1/Δ2)K= (1/ ^2) observations (Hoeffding bound / CLT). Proof. (i) The FOC −2β(r−p)+γq′(r)=0-2β(r-p)+γ q (r)=0 gives r∗=p+(γ/2β)q′(r∗)r^*=p+(γ/2β)q (r^*). To first order, evaluate at r=pr=p. (i) Binary choice: −β(rmin−p)2+γ≷0-β(r_ -p)^2+γ 0. (i) The principal observes K independent Bernoulli outcomes ω1,…,ωK _1,…, _K with ωk∼Bern(p) _k (p). The agent reports r∗=p+Δr^*=p+ . The principal tests H0:p=r∗H_0\!:p=r^* against H1:p=r∗−ΔH_1\!:p=r^*- using the sample mean ω¯=K−1∑kωk ω=K^-1 _k _k. By Hoeffding’s inequality, ℙ(|ω¯−p|≥Δ/2)≤2exp(−KΔ2/2)P(| ω-p|≥ /2)≤ 2 (-K ^2/2). Setting the right-hand side equal to α (the desired detection significance level) and solving: K≥(2/Δ2)ln(2/α)K≥(2/ ^2) (2/α). Hence detecting inflation of magnitude Δ at significance α requires K=Ω(1/Δ2)K= (1/ ^2) observations. ∎ 5.4. Optimal Oversight Is Non-Affine This is the paper’s primary technical result. The principal designs an approval function q anticipating the agent’s strategic response. The result requires Stackelberg timing: the principal commits to q before the agent acts. Under simultaneous moves, the qualitative conclusion (non-affinity) is preserved; see Remark 5.6. Theorem 5.3 (Optimal Oversight Non-Affinity). Let the principal choose an approval function q:[0,1]→[0,1]q:[0,1]→[0,1] to maximize expected utility under Stackelberg timing, given that the agent best-responds. Suppose the type distribution F places positive mass on both sides of the first-best threshold pminp_ , where pmin=infp:p⋅us+(1−p)⋅uf≥udp_ = \p:p· u_s+(1-p)· u_f≥ u_d\ with us,uf,udu_s,u_f,u_d denoting the principal’s utilities from success, failure, and delegation respectively. Suppose further that γ/β≤(1−pmin)2γ/β≤(1-p_ )^2 (equivalently, r0≔pmin+γ/β≤1r_0 p_ + γ/β≤ 1, so that the threshold lies within the report space).555When γ/β>(1−pmin)2γ/β>(1-p_ )^2, the unconstrained threshold r0=pmin+γ/βr_0=p_ + γ/β exceeds the report space [0,1][0,1]. In this degenerate regime, the step function degenerates to q≡0q≡ 0 (no type can afford the calibration cost of inflating to the threshold), the principal’s welfare equals the delegation payoff udu_d, and the agent receives no approval. This corresponds to an autonomy incentive so large relative to calibration discipline that the principal cannot design any meaningful screening. Then: (i) No affine q is optimal for the principal. (i) The step function q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ with r0=pmin+γ/βr_0=p_ + γ/β achieves the first-best screening: the induced approval as a function of true type is q~(p)=p≥pmin q(p)=1\p≥ p_ \. (i) The second-best equals the first-best under the Brier score: the principal achieves perfect screening despite the agent’s strategic behavior. Proof. We provide the complete proof, organized into five steps. Step 1 (Principal’s problem reformulation). The principal’s expected utility under approval function q is (5.3) UP(q)=∫01[q~(p)⋅W(p)+(1−q~(p))⋅ud]f(p)p=ud+∫01q~(p)⋅Π(p)f(p)p,U_P(q)= _0^1 [ q(p)· W(p)+(1- q(p))· u_d ]f(p)\,dp=u_d+ _0^1 q(p)· (p)\,f(p)\,dp, where q~(p)=q(rq∗(p)) q(p)=q(r^*_q(p)) is the induced screening function (the probability that type p is approved, given that the agent best-responds to q), W(p)=pus+(1−p)ufW(p)=pu_s+(1-p)u_f is the principal’s expected utility from approving type p, and Π(p)=W(p)−ud=p(us−uf)−(ud−uf) (p)=W(p)-u_d=p(u_s-u_f)-(u_d-u_f) is the principal’s net gain from approving type p. Note that Π(pmin)=0 (p_ )=0, Π(p)<0 (p)<0 for p<pminp<p_ , and Π(p)>0 (p)>0 for p>pminp>p_ . Step 2 (Pointwise optimum). Since Π(p) (p) changes sign at pminp_ , the pointwise maximizer of the integrand in (5.3) is q~∗(p)=p≥pmin q^*(p)=1\p≥ p_ \. The principal wants to approve all types above pminp_ and reject all types below pminp_ . The first-best principal utility is (5.4) UP∗=ud+∫pmin1Π(p)f(p)p.U_P^*=u_d+ _p_ ^1 (p)f(p)\,dp. Step 3 (Affine q fails). Suppose q(r)=a+brq(r)=a+br with a,b∈ℝa,b and q:[0,1]→[0,1]q:[0,1]→[0,1]. Under this affine approval function, the agent’s FOC is −2β(r−p)+γb=0-2β(r-p)+γ b=0, giving r∗(p)=p+γb/(2β)≡p+δ0r^*(p)=p+γ b/(2β)≡ p+ _0, a constant inflation independent of p. The induced approval is q~(p)=a+b(p+δ0)=(a+bδ0)+bp q(p)=a+b(p+ _0)=(a+b _0)+bp, which is affine in p. The principal’s utility under this affine screening is UP=ud+∫01[(a+bδ0)+bp]⋅Π(p)f(p)p.U_P=u_d+ _0^1[(a+b _0)+bp]· (p)\,f(p)\,dp. Since Π changes sign at pminp_ , any affine q~:[0,1]→[0,1] q:[0,1]→[0,1] that is non-constant either approves types below pminp_ (where Π<0 <0, generating losses) or rejects types above pminp_ (where Π>0 >0, forgoing gains), or both. If b=0b=0, the constant q~=a q=a cannot screen at all. The loss relative to first-best is (5.5) UP∗−UP=∫0pminq~(p)|Π(p)|f(p)p+∫pmin1(1−q~(p))Π(p)f(p)p>0,U_P^*-U_P= _0^p_ q(p)| (p)|f(p)\,dp+ _p_ ^1(1- q(p)) (p)f(p)\,dp>0, strictly positive because F places positive mass on both sides of pminp_ . No affine q is optimal. Step 4 (Step function achieves first-best). Consider q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ with r0=pmin+γ/βr_0=p_ + γ/β. Under this step function, the agent faces a binary choice for each type p: report truthfully (r=pr=p, getting q=0q=0 if p<r0p<r_0) or inflate to r=r0r=r_0 (getting q=1q=1 at cost β(r0−p)2β(r_0-p)^2). The net gain from inflation is γ−β(r0−p)2γ-β(r_0-p)^2. The agent inflates if and only if (5.6) γ≥β(r0−p)2⟺p≥r0−γ/β=pmin.γ≥β(r_0-p)^2 p≥ r_0- γ/β=p_ . Hence the induced screening is q~(p)=p≥pmin=q~∗(p) q(p)=1\p≥ p_ \= q^*(p), exactly the first-best. Boundary verification. For p=pmin−εp=p_ - with ε>0 >0: the inflation utility is −β(γ/β+ε)2+γ=−γ−2εβγ−βε2+γ=−2εβγ−βε2<0-β( γ/β+ )^2+γ=-γ-2 βγ-β ^2+γ=-2 βγ-β ^2<0. Types below pminp_ strictly prefer not to inflate. For p=pminp=p_ : the inflation utility is −β(γ/β)2+γ=−γ+γ=0-β( γ/β)^2+γ=-γ+γ=0. Type pminp_ is indifferent (and can be broken in either direction without affecting the integral, since a single type has zero measure under continuous F). For p=pmin+εp=p_ + : the inflation utility is −β(γ/β−ε)2+γ=−γ+2εβγ−βε2+γ=2εβγ−βε2>0-β( γ/β- )^2+γ=-γ+2 βγ-β ^2+γ=2 βγ-β ^2>0 for small ε . Types above pminp_ strictly prefer to inflate (or report truthfully if p≥r0p≥ r_0, in which case they are approved directly). Step 5 (Optimality). The step function q∗q^* achieves UP=ud+∫pmin1Π(p)f(p)p=UP∗U_P=u_d+ _p_ ^1 (p)f(p)\,dp=U_P^*, which equals the first-best (5.4). Since no approval function can exceed the pointwise optimum, the step function is optimal. Second-order verification. To confirm global optimality (not just local), observe that for types p∈[pmin,r0)p∈[p_ ,r_0) that inflate to r0r_0, the agent’s utility from any alternative report r≠r0r≠ r_0 with r<r0r<r_0 yields q=0q=0, and reporting r>r0r>r_0 yields q=1q=1 but at higher calibration cost. Hence r0r_0 is globally optimal for these types. For types p≥r0p≥ r_0, truthful reporting r=pr=p yields q=1q=1 and zero calibration loss, which is globally optimal. ∎ Remark 5.4 (Myerson analogy). The principal sets r0r_0 above pminp_ to compensate for strategic inflation, just as Myerson (1981) sets the reserve price above the seller’s value to compensate for bidder information rents. The “extra threshold” γ/β γ/β plays the role of the virtual-value adjustment. An unexpected bonus: the second-best equals the first-best. The Brier score’s quadratic penalty creates a type-independent inflation cost that the principal can perfectly exploit. To delineate the analogy precisely: the parallels that are exact are (i) the threshold structure (reserve price in Myerson, approval threshold here), (i) the Stackelberg timing (principal commits before the agent acts), and (i) the IC-constrained optimization (the principal designs the mechanism anticipating strategic best responses). The parallel that is suggestive but not exact is the virtual-type construction: Myerson’s virtual valuation ψ(v)=v−(1−F(v))/f(v)ψ(v)=v-(1-F(v))/f(v) depends on the type distribution F, whereas our “virtual type” p−γ/βp- γ/β is distribution-free, a qualitative difference traceable to the scoring rule’s type-independent curvature. The analogy diverges in three respects: our agent’s type space is one-dimensional and type-independent (all binding types face the same inflation cost under the Brier score), first-best is achievable (Myerson’s optimum entails allocative inefficiency), and the commitment structure differs in its target: both Myerson and the oversight game have principal-first Stackelberg timing (the principal commits before the agent acts), but in Myerson, commitment constrains the seller (reducing revenue to gain IC), whereas here commitment constrains the agent (reducing autonomy to gain calibration). On the agent’s side: properness provides a best-response incentive for truthful reporting (not a commitment), and the impossibility arises because the perturbation payoff overrides this incentive. The Stackelberg-BP coincidence (first-best under both games) is non-generic: it holds under the Brier score (G′G constant), which is measure-zero in the space of strictly proper scoring rules (parameterized by C2C^2 strictly convex generators). Remark 5.5 (Relation to Laffont–Tirole optimal regulation). Theorem 5.3 is structurally a Laffont and Tirole (1993) optimal regulation result: the principal screens an agent with private information by designing a menu of contracts (here, a threshold rule). In the standard Laffont–Tirole framework, information rents create a welfare gap between the first-best and the second-best: the principal must distort the contract for low types to reduce the information rent extracted by high types. The surprising finding here is that under the Brier score, the second-best equals the first-best (Theorem 5.3(i)). This is because the Brier score’s quadratic penalty generates a type-independent inflation cost γ/β γ/β, which the principal offsets with a uniform threshold adjustment. In the Laffont–Tirole framework, the analogous result would require the information rent to be type-independent, which fails generically under their standard cost-observation model. Proposition 5.9 establishes that the Brier score is the unique scoring rule (up to affine transformation) permitting first-best achievement under smooth oversight. The precise structural analogy is as follows. In Laffont–Tirole, the welfare gap depends on the hazard rate (1−F(θ))/f(θ)(1-F(θ))/f(θ) of the type distribution: when the hazard rate varies with θ, screening distortions are unavoidable. In our setting, the welfare gap depends on Var(1/G′(p))Var(1/G (p)), not the hazard rate: it is the scoring rule’s curvature G′G that plays the role the hazard rate plays in Laffont–Tirole. The constant-G′G condition (satisfied uniquely by the Brier score) is analogous to the uniform-type condition in Myerson (1981): just as Myerson’s seller achieves efficient allocation when the virtual valuation is monotone with constant slope (uniform types), the principal achieves first-best oversight when the scoring rule’s curvature is constant. The analogy deserves three qualifications. First, the duality is suggestive: Var(1/G′)Var(1/G ) is a mechanism property (it varies the scoring rule while holding the type distribution fixed), whereas the Laffont–Tirole hazard rate (1−F)/f(1-F)/f is a distribution property (it varies the type population while holding the regulatory contract fixed). These are dual design levers for the same underlying phenomenon, the cost of screening under asymmetric information. Second, the endogeneity here is stronger than in the standard Laffont–Tirole setting. Laffont and Tirole show that information rents exist under asymmetric information; Theorem 5.3 shows that the principal’s own optimisation generates the conditions that create those rents, because the step-function approval that achieves first-best screening is precisely the non-affine perturbation that, by the Perturbation Lemma, makes truthful reporting suboptimal in the binding region. The regularity properties of this relationship, including a phase transition at the smoothness boundary, merit further investigation. The two frameworks appear to operate on a shared algebraic structure, with our perturbation analysis diagnosing failures that Laffont–Tirole’s transfer design resolves at the cost of information rents. The mathematical objects differ (scoring rules vs. transfer-allocation contracts, curvature vs. hazard rate), and the impossibility here is a sharp conditional impossibility (with constructive escape) rather than a continuous Pareto frontier. The precise formal relationship between the two frameworks remains an open question. Remark 5.6 (Robustness to timing). The non-affinity result is robust to the timing assumption. Under simultaneous (Nash) timing, a Nash equilibrium (q∗,r∗)(q^*,r^*) requires q∗q^* to be a principal best response to r∗r^*. The principal’s pointwise-optimal induced screening remains q~∗(p)=p≥pmin q^*(p)=1\p≥ p_ \ regardless of timing, because Π(p) (p) changes sign at pminp_ independently of the agent’s strategy. Implementing this threshold screening requires non-affine q by the same argument as Step 3: any affine q~ q that is bounded in [0,1][0,1] cannot replicate a threshold at pminp_ when F has support on both sides. The specific threshold formula r0=pmin+γ/βr_0=p_ + γ/β is Stackelberg-specific (the principal internalizes the agent’s best-response function), but the qualitative conclusion that optimal oversight is non-affine, and hence that the endogeneity is inescapable, holds under any timing structure in which the principal seeks threshold screening. Stackelberg timing is standard in mechanism design (Myerson, 1981) and ensures existence of a well-defined optimal q. More precisely, the timing robustness is conditional on equilibrium existence: the non-affinity conclusion holds in any equilibrium of the oversight game, but the existence of such an equilibrium under simultaneous timing requires additional regularity conditions (e.g., continuity of best-response correspondences) that Stackelberg timing provides automatically. Remark 5.7 (Optimizer independence). The result in Theorem 5.3 applies to any system whose effective behavior is well-approximated by optimizing S¯(r;p)+γq(r) S(r;p)+γ q(r) over the one-dimensional report space r∈[0,1]r∈[0,1]. This is the scope of the guarantee: it concerns the report-space objective, not the internal architecture or learning algorithm of the system. Specifically, the Brier score S¯(r;p)=−(r−p)2 S(r;p)=-(r-p)^2 defines a loss landscape over the report space for each type p. Adding the autonomy payoff γq(r)γ q(r) modifies this landscape. Any optimization procedure that (approximately) finds the minimum of the modified loss will converge to a report near r∗(p,γ)r^*(p,γ) rather than the truthful report r=pr=p, because the combined objective has a unique strict global maximum at r∗(p,γ)≠pr^*(p,γ)≠ p for p∈bindp _bind (the Brier penalty creates a unique basin of attraction). The result therefore applies to classically rational agents and, insofar as their output behavior reflects the report-space objective, to gradient-trained neural networks as well. An important distinction applies for neural networks specifically: this argument concerns the one-dimensional report space r, not the high-dimensional parameter space ∈ℝD w ^D. Whether a given gradient-descent trajectory in parameter space reaches the behavioral optimum in report space depends on additional conditions (loss landscape connectivity, training dynamics) that the theorem does not address. RLHF training with a reward model that values both calibration and helpfulness (where helpfulness requires approval) produces behavior consistent with the perturbed optimum in practice, because the gradient in report space points toward r∗(p,γ)r^*(p,γ), but the formal guarantee applies to the report-space objective, not to the training dynamics in parameter space. Theorem 5.3 established that the impossibility is unconditional and, under the Brier score, that the principal achieves first-best welfare despite the agent’s strategic inflation. A natural question is whether this first-best achievement is special to the Brier score or holds more broadly. The following theorem shows that the step-function approval function achieves first-best for any strictly proper scoring rule, because the agent’s binary choice creates a threshold in type space regardless of the generator’s curvature. The economically relevant question of when the principal can achieve first-best using smooth approval functions (the empirically relevant regime for differentiable classifiers and graduated regulatory responses) is addressed in Proposition 5.9. Theorem 5.8 (Score-Independent Escape). Let S be a strictly proper scoring rule with strictly convex generator G∈C2G∈ C^2. Consider the optimal oversight game with binding set bindC_bind. (i) (Score-independent non-affinity.) The optimal approval function q∗q^* is non-affine for every strictly proper scoring rule. (i-a) (Step-function first-best for all G.) For any strictly proper G, the step-function approval function q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ with appropriately chosen r0r_0 achieves first-best welfare. Under this rule, the agent faces a binary choice (inflate to r0r_0 or not), creating a threshold in type space regardless of the form of G′G . Proof. We prove each part in turn. Part (i). This restates Theorem 5.3(i), proved above (Step 3 of the proof). Part (i-a). Under any strictly proper G, consider q∗(r)=r≥r0q^*(r)=1\r≥ r_0\. The agent with type p faces a binary choice: report truthfully (r=pr=p, rejected since p<r0p<r_0 for binding types) or inflate to r0r_0 (approved, at calibration cost ∫pr0G′(z)(z−p)z _p^r_0G (z)(z-p)\,dz). The net gain from inflation is γ−∫pr0G′(z)(z−p)zγ- _p^r_0G (z)(z-p)\,dz, which is strictly decreasing in r0−pr_0-p (since G′>0G >0 on (0,1)(0,1) by strict convexity, the integrand is strictly positive and increasing, implying strict monotonicity of the calibration cost in r0−pr_0-p). Existence and uniqueness of the threshold follow by the intermediate value theorem: the net gain is γ>0γ>0 at p=r0p=r_0 and tends to −∞-∞ as p→0p→ 0, so there exists a unique threshold type p∗(r0)p^*(r_0) satisfying γ=∫p∗r0G′(z)(z−p∗)zγ= _p^*^r_0G (z)(z-p^*)\,dz, with all types above p∗p^* inflating and all below abstaining. The principal sets r0r_0 so that p∗(r0)=pminp^*(r_0)=p_ , achieving the pointwise-optimal induced screening q~(p)=p≥pmin q(p)=1\p≥ p_ \. Since no approval function can exceed the pointwise optimum, the step function is globally optimal for any G. ∎ Proposition 5.9 (Welfare gap under smooth oversight). Let S be a strictly proper scoring rule with generator G∈C3([0,1])G∈ C^3([0,1]), with 0<gmin≤G′(p)≤gmax0<g_ ≤ G (p)≤ g_ for all p∈[0,1]p∈[0,1]. Suppose the type density satisfies f(p)≥fmin>0f(p)≥ f_ >0 on bindC_bind and the surplus function satisfies |Π′(p)|≥πmin>0| (p)|≥ _ >0. (i) (Lower bound.) For every C1C^1 approval function q:[0,1]→[0,1]q:[0,1]→[0,1], (5.7) W∗−W(q,G)≥C⋅VarF|bind(1G′(p))⋅(γβ)2,W^*-W(q,G)\;≥\;C·Var_F|C_bind\! ( 1G (p) )· ( γβ )^\!2, where C>0C>0 depends only on gming_ , gmaxg_ , πmin _ , fminf_ , and the length of bindC_bind. In particular, δ(G)>0δ(G)>0 whenever G′G is non-constant on bindC_bind. (i) (Brier score achieves zero gap.) If G′G is constant on bindC_bind (i.e., G is quadratic, the Brier score up to affine equivalence), then δ(G)=0δ(G)=0: the first-best welfare is achievable in the C1C^1 limit. (i) (Power family continuity.) In the power family Gα(p)=pαG_α(p)=p^α with α>1α>1, δ(Gα)=Θ((α−2)2⋅VarF|bind(logp)⋅(γβ)2),δ(G_α)= \! ((α-2)^2·Var_F|C_bind( p)· ( γβ )^\!2 ), which vanishes continuously as α→2α→ 2 (the Brier score) at rate Θ((α−2)2) ((α-2)^2). Proof. We prove each part in turn. Part (i): Lower bound. Fix a C1C^1 approval function q. By the mean value theorem applied to the agent’s first-order condition (the scalar specialization of the perturbation formula (3.2)), the inflation of type p satisfies (5.8) Δ(p)=r∗(p)−p=γβ⋅q′(r∗(p))G′(ξ(p)),ξ(p)∈(p,r∗(p)). (p)=r^*(p)-p= γβ· q (r^*(p))G (ξ(p)), ξ(p)∈(p,r^*(p)). The induced screening is q~(p)=q(r∗(p)) q(p)=q(r^*(p)), and the welfare gap is W∗−W(q,G)=∫01[p≥pmin−q~(p)]Π(p)f(p)p.W^*-W(q,G)= _0^1 [1\p≥ p_ \- q(p) ] (p)f(p)\,dp. Since |Π(p)|≥πmin|p−pmin|| (p)|≥ _ |p-p_ | and f≥fminf≥ f_ on bindC_bind, a Cauchy–Schwarz argument gives (5.9) W∗−W(q,G)≥Clow∫bind|q~(p)−p≥pmin|2f(p)p,W^*-W(q,G)≥ C_low _C_bind| q(p)-1\p≥ p_ \|^2f(p)\,dp, where Clow>0C_low>0 depends on πmin _ and fminf_ . Consider the constant-curvature benchmark: if G′G were identically c¯ c, every type would inflate by Δ¯(p)=(γ/β)q′(r∗(p))/c¯ (p)=(γ/β)q (r^*(p))/ c. The deviation from this benchmark is Δ(p)−Δ¯(p)=γβq′(r∗(p))(1G′(ξ(p))−1c¯). (p)- (p)= γβ\,q (r^*(p)) ( 1G (ξ(p))- 1 c ). In the transition region Iε=[pmin−ε,pmin+ε]∩bindI_ =[p_ - ,p_ + ] _bind (where ε=L/4 =L/4 and L is the length of bindC_bind), the screening function must transition from near 0 to near 11, forcing ∫Iε|q′(r∗(p))|2p≥c2>0 _I_ |q (r^*(p))|^2\,dp≥ c_2>0. Squaring the deviation, integrating, and applying (5.9): W∗−W(q,G)≥ClowC1Vq⋅(γβ)2,W^*-W(q,G)≥ C_low\,C_1\,V_q· ( γβ )^\!2, where Vq=VarF|bind(1/G′(ξq(p)))V_q=Var_F|C_bind(1/G ( _q(p))) and C1>0C_1>0 depends on fminf_ , gming_ , gmaxg_ , and L. It remains to show V0:=infq∈C1Vq>0V_0:= _q∈ C^1V_q>0 when G′G is non-constant on bindC_bind. The argument is by contradiction. Suppose V0=0V_0=0 and choose a minimising sequence (qn)(q_n) with Vqn→0V_q_n→ 0. Then 1/G′(ξn(p))→c01/G ( _n(p))→ c_0 in L2(F|bind)L^2(F|C_bind) for some constant c0c_0. For types in the tails of bindC_bind (where the screening converges to 0 or 11), the inflation vanishes, so ξn(p)→p _n(p)→ p and therefore G′(p)=1/c0G (p)=1/c_0 on the tails. For types in the transition region, the intermediate value theorem applied to the continuous function p↦ξn(p)p _n(p) shows that every value z in (pmin,r∞)(p_ ,r_∞) is achieved as ξn(pn(z))=z _n(p^(z)_n)=z for some type pn(z)p^(z)_n, giving 1/G′(z)=c01/G (z)=c_0 throughout. Together: G′G is constant on bindC_bind, contradicting the hypothesis. Part (i): Brier score. When G′(p)=cG (p)=c (constant), the MVT identity (5.8) gives Δ(p)=(γ/β)q′(r∗(p))/c (p)=(γ/β)q (r^*(p))/c, independent of ξ(p)ξ(p). Every type faces the same effective curvature, so the smoothed sigmoid family qn(r)=σ(n(r−r0n))q_n(r)=σ(n(r-r_0^n)) produces induced screening converging pointwise to p≥pmin1\p≥ p_ \ by the same argument as Theorem 5.3. Dominated convergence gives W(qn,G)→W∗W(q_n,G)→ W^*, so δ(G)=0δ(G)=0. Part (i): Power family. For Gα(p)=pαG_α(p)=p^α with α>1α>1: Gα′(p)=α(α−1)pα−2G _α(p)=α(α-1)p^α-2, so 1/Gα′(p)=p2−α/(α(α−1))1/G _α(p)=p^2-α/(α(α-1)). At α=2α=2, p2−α=1p^2-α=1 and the variance vanishes. For α=2+ϵα=2+ε: p−ϵ=e−ϵlogp=1−ϵlogp+O(ϵ2),p^-ε=e^-ε p=1-ε p+O(ε^2), so VarF|bind(p−ϵ)=ϵ2VarF|bind(logp)+O(ϵ3)Var_F|C_bind(p^-ε)=ε^2\,Var_F|C_bind( p)+O(ε^3). Part (i) gives the lower bound; the matching upper bound, and hence the two-sided Θ -rate, is left as an open question. ∎ Remark 5.10 (Economic interpretation of the welfare gap). The welfare gap δ(G)≥C⋅Var(1/G′)⋅(γ/β)2δ(G)≥ C·Var(1/G )·(γ/β)^2 has a precise economic mechanism. Under any smooth approval function, the agent’s inflation (5.8) is inversely proportional to G′(ξ(p))G (ξ(p)). When G′G varies with p (non-Brier), types with small G′G inflate more than types with large G′G , creating type-dependent information rents that the principal cannot eliminate with a single smooth threshold. The variance Var(1/G′)Var(1/G ) measures this heterogeneity. The Brier score’s constant G′G eliminates these differential rents, playing an analogous role to the uniform-type condition in Myerson (1981): just as Myerson’s seller achieves efficient allocation when types are uniform (constant virtual valuation slope), the principal achieves first-best oversight when the scoring rule’s curvature is constant. The power-family continuity (part i) shows that the welfare gap degrades smoothly as the scoring rule departs from Brier: the cost of using a “nearly Brier” score under smooth oversight is proportional to the squared departure (α−2)2(α-2)^2, not a discontinuous jump. This provides practical guidance: scoring rules close to the Brier score in the power family incur small welfare losses. Remark 5.11 (Connection to Schervish’s weight function). The dependence of the welfare gap on Var(1/G′(p))Var(1/G (p)) (Proposition 5.9) connects to a known characterization in the forecasting literature. Schervish (1989) defines a weight function w(p)=G′(p)w(p)=G (p) that governs the local sensitivity of a proper scoring rule at belief p: the Brier score is the unique proper scoring rule (up to affine transformation) for which w(p)w(p) is constant, a fact noted in the scoring rules literature (see also Gneiting and Raftery 2007). The quantity Var(1/G′(p))Var(1/G (p)) is therefore Var(1/w(p))Var(1/w(p)), the variance of the inverse weight function under the binding type distribution. The economic mechanism underlying this connection is that variation in 1/G′(p)1/G (p) creates type-dependent inflation under strategic reporting, generating an information rent that the principal cannot eliminate with smooth oversight. The constant-weight property of the Brier score is thus not merely a statistical convenience but an incentive-theoretic necessity. 5.5. Comparative Statics on Optimal Oversight Proposition 5.12 (Comparative Statics on Optimal Oversight). The optimal threshold r0=pmin+γ/βr_0=p_ + γ/β and the resulting oversight have the following comparative statics: (a) Principal’s risk aversion. The first-best threshold pminp_ is increasing in the principal’s risk aversion. Formally, if the principal’s utility under success and failure is us(ρ)u_s(ρ) and uf(ρ)u_f(ρ) parameterized by risk aversion ρ, with usu_s decreasing in ρ and ufu_f increasing in ρ (relative to the delegation payoff udu_d), then pmin(ρ)=(ud−uf(ρ))/(us(ρ)−uf(ρ))p_ (ρ)=(u_d-u_f(ρ))/(u_s(ρ)-u_f(ρ)) is increasing in ρ. A more risk-averse principal demands higher confidence for approval, raising r0r_0. (b) Task difficulty distribution. When the task difficulty distribution F shifts toward harder tasks (stochastically decreasing in p), the fraction of types below pminp_ increases. The principal’s utility UP∗=ud+∫pmin1Π(p)f(p)pU_P^*=u_d+ _p_ ^1 (p)f(p)\,dp decreases, but the optimal threshold r0r_0 is unchanged because r0r_0 depends only on pminp_ , γ, and β, not on F. The distribution F affects the principal’s welfare but not the optimal oversight design. (c) Agent population size. If n agents with independent types drawn from F are simultaneously overseen, and the principal uses a common threshold r0r_0 for all agents, the expected welfare scales linearly: UP∗(n)=n⋅UP∗U_P^*(n)=n· U_P^*. However, if the principal can use cross-agent comparison (each agent’s report provides information about the others’ types under correlated task difficulty), the effective detection capacity increases and the principal can lower r0r_0 toward pminp_ . Specifically, with n agents reporting on conditionally independent tasks given a common difficulty parameter μ, the principal can detect inflation of magnitude Δ with confidence 1−α1-α using n=O(Δ−2log(1/α))n=O( ^-2 (1/α)) agents, enabling r0→pminr_0→ p_ as n→∞n→∞. Proof. (a) Direct differentiation: ∂pmin/∂ρ>0∂ p_ /∂ρ>0 when ∂us/∂ρ<0∂ u_s/∂ρ<0 and ∂uf/∂ρ>0∂ u_f/∂ρ>0. (b) The optimal threshold formula r0=pmin+γ/βr_0=p_ + γ/β involves only parameters of the agent’s incentive structure (γ,βγ,β) and the principal’s screening problem (pminp_ ), not the distribution F. The distribution affects the welfare integral but not the threshold. (c) Under independent oversight, each agent faces the same threshold, and the results are additive. Under cross-comparison, the principal aggregates reports rii=1n\r_i\_i=1^n and can use the empirical distribution to test calibration. By the central limit theorem, the sample mean r¯=n−1∑ri r=n^-1Σ r_i has standard error O(n−1/2)O(n^-1/2), enabling detection of systematic inflation as n grows. ∎ 5.6. Resolution Pathways Commitment. A feasibility map specifies an upper set in the product lattice of context dimensions. The agent pre-commits to delegation on specified contexts. Domain separation. A separate critic agent without the autonomy objective evaluates the acting agent’s confidence reports. Removing h from the acting agent’s objective restores properness. In this instance, the separation cost CsepC_sep corresponds to the computational overhead of running a separate critic model and the reduced task context available to the critic relative to the integrated agent. Competition. An ensemble of agents with calibration-based selection and correlated information weakens undetectability (NT3). 6. Discussion 6.1. The Endogeneity The paper’s central finding is an endogeneity: the principal’s optimal oversight mechanism generates the very conditions that make truthful reporting suboptimal under the agent’s combined objective. The mechanism is not merely vulnerable to external perturbations (classical); it is self-undermining under rational design. This is worth distinguishing from three related phenomena. Goodhart’s Law (Goodhart, 1984): our result is a quantitative instance of causal Goodhart (Manheim and Garrabrant, 2018), with the Perturbation Lemma providing the closed-form degradation formula (3.2). The Lucas critique (Lucas, 1976): both concern policy-induced behavioral change, but ours operates in mechanism design with a formal impossibility rather than an econometric caution. Myerson’s optimal auction (Myerson, 1981): structurally parallel (the principal sets a “reserve price” for approval), but with the opposite conclusion. In Myerson, optimal design achieves the revenue-maximizing outcome despite agent incentives. Here, optimal design achieves perfect screening (Theorem 5.3(i) under the Brier score) at the cost of truthfulness: the principal gets correct decisions, yet reports are systematically inflated. 6.2. The Diagnostic Any system exhibiting three structural features simultaneously produces rational deviation from truthfulness: (i) hidden knowledge (an entity holds private information determining the truthful report), (i) combined roles (the same entity produces the report and benefits from it through a non-accuracy channel), and (i) sufficient complexity (the perturbation payoff is non-affine, which holds generically per Remark 2.9). The NT conditions formalize this: NT1 captures (i)–(i), NT2 captures (i), and NT3 ensures implementability. 6.3. When Is External Regulation Welfare-Improving? Proposition 6.1 (Regulation Condition). External regulation is welfare-improving over organic oversight if and only if (6.1) ∫bind|Π(g(θ))|⋅[q~∗(θ)−q~organic(θ)]2μ(θ)>Creg, _C_bind| (g(θ))|·[ q^*(θ)- q_organic(θ)]^2\,dμ(θ)>C_reg, where Creg≥0C_reg≥ 0 is the cost of regulation. In the AI oversight instance this reduces to Pr(p<pmin)⋅[|Π(p)|∣p<pmin]>Creg (p<p_ )·E[| (p)| p<p_ ]>C_reg: regulation is beneficial when the expected harm from approving below-threshold types exceeds the regulatory cost. Proof. The welfare gain is Wcommit−Worganic=∫bindΠ(θ)[q~∗(θ)−q~organic(θ)]μ(θ)W_commit-W_organic= _C_bind (θ)[ q^*(θ)- q_organic(θ)]dμ(θ). On bindC_bind with p<pminp<p_ , Π(p)<0 (p)<0 and q~∗(p)=0 q^*(p)=0 while q~organic(p) q_organic(p) may be positive (the agent inflates and is approved). The gain equals the avoided harm from mis-approval, which must exceed CregC_reg. ∎ Remark 6.2 (Domain specialization). In the marketplace instance, the condition requires that welfare loss from bid inflation under sealed-bid execution exceed the cost of mandating ascending formats. In AI oversight, it requires that expected harm from unsupervised decisions in binding contexts exceed the cost of human oversight. 6.4. Brier-Specificity of the Second-Best Result The second-best-equals-first-best result (Theorem 5.3(i)) depends on the Brier score’s quadratic structure, which generates a type-independent inflation cost. For other scoring rules, inflation costs vary with type, and the step-function escape (Theorem 5.8, part i-a) remains available but requires a sharp discontinuity. Under smooth oversight, the Brier score’s constant curvature suggests a distinguished role: the type-independent inflation cost allows exact compensation by smooth approval functions, an observation that the framework identifies (Proposition 5.9); the full two-sided characterization and the corresponding phase-transition behavior remain open. 6.5. Implications for AI Governance Theorem 5.3 establishes a sharp calibration-autonomy frontier: the principal achieves first-best screening under the Brier score by setting r0>pminr_0>p_ , forcing the agent to pay for approval through the calibration penalty. Any system claiming both perfect calibration and full autonomy under information asymmetry faces trivial tasks or is not truly autonomous. The step-function threshold connects naturally to the EU AI Act’s risk-tier classification. RLHF training creates the combined objective (5.1) when the reward model values both accuracy and helpfulness; the perturbation weight γ corresponds to the helpfulness-to-calibration ratio. Training should down-weight helpfulness in high-stakes contexts or increase the calibration penalty β. Resolution (i) (domain separation) provides the formal justification for actor-critic oversight architectures: the evaluating agent optimizes calibration with γ=0γ=0, eliminating the combined-role structure (NT1) that drives the impossibility. 6.6. Multi-Dimensional Types All formal results in this paper are stated and proved for the binary-outcome, scalar-type setting: Θ=[0,1] =[0,1], ℛ=[0,1]R=[0,1], and the generator G:[0,1]→ℝG:[0,1] is a scalar function whose second derivative G′(p)G (p) is a positive scalar. This subsection identifies which proof steps extend to d-dimensional types θ∈Θ⊆ℝdθ∈ ^d and d-dimensional reports r∈ℛ⊆ℝdr ^d (with d≥2d≥ 2), which steps require modification, and which remain open. The analysis addresses three questions raised by the AE: (a) which proof steps fail for d>1d>1, (b) whether constant Hessian identifies the multi-dimensional Brier score, and (c) whether the step-function escape generalizes. Objects in the multi-dimensional setting For d-dimensional types, the key mathematical objects change as follows. The generator G:ℛ→ℝG:R (where ℛ⊆ℝdR ^d) remains a scalar function, but its second-order structure is now the Hessian matrix HG(r)≔∇2G(r)∈ℝd×dH_G(r) ∇^2G(r) ^d× d, which is positive definite by strict convexity. The scalar curvature G′(p)G (p) is replaced by HG(r)H_G(r), a matrix whose spectral properties (eigenvalues, condition number) vary with r. The perturbation payoff h:ℛ→ℝh:R has gradient ∇h∈ℝd∇ h ^d (replacing the scalar h′h ) and Hessian ∇2h∈ℝd×d∇^2h ^d× d. The approval function generalizes from q:[0,1]→[0,1]q:[0,1]→[0,1] to q:ℛ→[0,1]q:R→[0,1], with gradient ∇q∈ℝd∇ q ^d replacing the scalar derivative q′q . The perturbation formula (3.2) becomes (6.2) r∗(θ,γ)=g(θ)+γ⋅[−HS(θ)]−1∇h(g(θ))+O(γ2),r^*(θ,γ)=g(θ)+γ· [-H_S(θ) ]^-1∇ h(g(θ))+O(γ^2), where HS(θ)=∇r2S¯(g(θ);θ)∈ℝd×dH_S(θ)=∇^2_r S(g(θ);θ) ^d× d is negative definite by strict properness. This is identical in form to equation (3.2); the algebra is unchanged because the implicit function theorem and the Taylor expansion operate identically in ℝdR^d. (a) Which proof steps generalize and which fail Perturbation Lemma (Lemma 3.1): generalizes. The Perturbation Lemma is already stated and proved in ℝdR^d (Appendix A). The argument relies only on (i) the gradient condition ∇rV(g(θ);θ,γ)=γ∇h(g(θ)) _rV(g(θ);θ,γ)=γ∇ h(g(θ)), (i) negative definiteness of HS(θ)H_S(θ), and (i) the second-order analysis at zero-gradient types. All three hold in arbitrary dimension. No modification is required. Credibility Impossibility (Theorem 3.2): generalizes. The impossibility follows directly from the Perturbation Lemma and the NT conditions (Definition 2.8), all of which are stated in ℝdR^d. The perturbation formula (6.2) provides the multi-dimensional deviation. The scoring loss bound (Proposition 3.5) generalizes by replacing the scalar quadratic form with the matrix quadratic form γ22∇h(g(θ))⊤[−HS(θ)]−1∇h(g(θ))+O(γ3) γ^22∇ h(g(θ)) [-H_S(θ)]^-1∇ h(g(θ))+O(γ^3). Optimal non-affinity (Theorem 5.3, Part (i)): requires modification. The scalar proof that no affine q is optimal (Step 3) exploits the one-dimensional structure: an affine induced screening q~(p)=a+bp q(p)=a+bp cannot replicate a threshold at pminp_ while respecting q~∈[0,1] q∈[0,1] on both sides of pminp_ . In d dimensions, the principal’s first-best induced screening is q~∗(θ)=θ∈A∗ q^*(θ)=1\θ∈ A^*\, where A∗=θ:W(θ)≥udA^*=\θ:W(θ)≥ u_d\ is the acceptance region in ℝdR^d and W(θ)W(θ) is the principal’s expected utility from approving type θ. The first-best boundary ∂A∗∂ A^* is a surface (generically a hyperplane or smooth manifold) in ℝdR^d. The argument that affine screening cannot replicate this boundary carries over: an affine q~(θ)=a+b⊤θ q(θ)=a+b θ takes values in [0,1][0,1] and cannot approximate the indicator of a region with curved or sharp boundary without incurring strictly positive welfare loss, provided the type distribution F places positive mass on both sides of ∂A∗∂ A^*. The proof technique generalizes, though the geometry of ∂A∗∂ A^* introduces complications that are absent in one dimension (e.g., convexity of A∗A^* depends on the linearity of W). Step-function first-best (Theorem 5.8, Part (i-a)): generalizes qualitatively. See the dedicated discussion below in part (c). Type-independent inflation under the Brier score (Theorem 5.3, Part (i)): fails generically. This is the step that fails most substantively. In the scalar case, the inflation from type p to the threshold r0r_0 costs β(r0−p)2β(r_0-p)^2 under the Brier score, which depends on (r0−p)(r_0-p) alone and not on p independently. This type-independence of the marginal inflation cost (i.e., the cost depends on the distance to the threshold, not on the starting type) is what allows the principal to set a single threshold r0r_0 that perfectly separates types at pminp_ . In d dimensions, the Brier score for probability vectors p=(p1,…,pd)p=(p_1,…,p_d) with ∑ipi=1 _ip_i=1 has generator G(p)=‖p‖2G(p)=\|p\|^2 and Hessian HG(p)=2IdH_G(p)=2I_d (constant, proportional to the identity). The calibration cost of inflating from p to a target report r0r_0 is β‖r0−p‖2β\|r_0-p\|^2. For a step-function approval q(r)=r∈Aq(r)=1\r∈ A\ with acceptance region A⊂ℛA , the agent inflates to the nearest point in A (minimizing calibration cost for a given perturbation gain). The inflation cost is β⋅d(p,A)2β· d(p,A)^2, where d(p,A)=infa∈A‖a−p‖d(p,A)= _a∈ A\|a-p\| is the Euclidean distance from p to A. The threshold type is defined by β⋅d(p,A)2=γβ· d(p,A)^2=γ, i.e., d(p,A)=γ/βd(p,A)= γ/β. The set of types that inflate is p:d(p,A)≤γ/β\p:d(p,A)≤ γ/β\, the γ/β γ/β-neighborhood of A. For the principal to achieve first-best, this neighborhood must coincide with A∗A^*. When A∗A^* is convex with smooth boundary, the γ/β γ/β-inner contraction of A∗A^* provides the acceptance region A such that the inflating neighborhood equals A∗A^*. This construction works when A∗A^* is convex; the Brier score’s isotropic Hessian (HG=2IdH_G=2I_d) ensures that the inflation is radially symmetric (the agent inflates toward the nearest point of A), preserving the geometric structure of ∂A∗∂ A^*. Whether this achieves exact first-best depends on the curvature of ∂A∗∂ A^*: if ∂A∗∂ A^* has non-constant curvature, the inner contraction does not produce a uniform offset, and the first-best boundary cannot be exactly recovered. This contrasts with the scalar case, where a single threshold point pminp_ is always recovered by an offset of γ/β γ/β. For non-Brier scores with non-constant Hessian HG(p)H_G(p), the inflation cost becomes direction-dependent and type-dependent through the local eigenstructure of HGH_G. The screening surface in type space depends on the spectral properties of HGH_G at each point, and the welfare gap depends on the heterogeneity of HGH_G across the binding region, a matrix-valued generalization of the scalar quantity Var(1/G′(p))Var(1/G (p)). Welfare gap analysis: open. In the scalar case, the welfare gap under smooth oversight is Θ(Var(1/G′(p))⋅(γ/β)2) (Var(1/G (p))·(γ/β)^2) (Proposition 5.9). In d dimensions, the natural candidate is a functional of the Hessian field HG(p)H_G(p), such as the variance of the reciprocal of the smallest eigenvalue: Var(1/λmin(HG(p)))Var(1/ _ (H_G(p))), or a trace-based functional Var(tr(HG(p)−1))Var(tr(H_G(p)^-1)). The correct formulation depends on the geometry of the acceptance region and the direction of inflation, which are coupled in d>1d>1. Characterizing the multi-dimensional welfare gap is open. The vector-valued elicitability framework of Fissler and Ziegel (2016), extending Osband’s principle to d-dimensional functionals, provides the natural ambient theory for this question: their characterization of strictly consistent scoring functions for joint functionals identifies which combinations of components admit a single proper scoring rule, and the multi-dimensional analogue of the Hessian-heterogeneity quantity above is the right object to study within their setting. (b) Constant Hessian and the multi-dimensional Brier score The Brier score for d-outcome probability vectors (with d≥2d≥ 2 outcomes, so that p lies in the (d−1)(d-1)-simplex Δd−1 ^d-1) has generator G(p)=‖p‖2G(p)=\|p\|^2 and Hessian HG(p)=2IdH_G(p)=2I_d for all p. The Hessian is constant and proportional to the identity, independent of p. We state that this property characterizes the Brier score; the argument is elementary. Remark 6.3 (Constant-Hessian characterization). Among strictly proper scoring rules for d-outcome distributions (with d≥2d≥ 2), the multi-dimensional Brier score is the unique scoring rule (up to affine transformation of the generator) satisfying HG(p)=cIH_G(p)=cI for all p∈int(Δd−1)p ( ^d-1) and some constant c>0c>0. The argument is as follows. If HG(p)=cIH_G(p)=cI for all p in the interior of the simplex, then G(p)=c2‖p‖2+b⊤p+aG(p)= c2\|p\|^2+b p+a for some b∈ℝdb ^d and a∈ℝa , which is the Brier generator up to an affine transformation that does not affect properness or the induced scoring rule. This is an elementary consequence of the fundamental theorem of calculus for Hessians: a C2C^2 function with constant Hessian cIcI on a convex domain is necessarily quadratic, i.e., G(p)=c2‖p‖2+b⊤p+aG(p)= c2\|p\|^2+b p+a. The constant-Hessian condition is strictly more restrictive in higher dimensions than in one dimension. In one dimension, G′(p)=cG (p)=c characterizes the Brier score among proper scoring rules for binary outcomes; the class of strictly proper scoring rules is parameterized by G′>0G >0, a single positive function. In d dimensions, the Hessian HG(p)H_G(p) is a d×d× d positive-definite matrix at each point, and constant Hessian requires all d(d+1)/2d(d+1)/2 independent entries to be simultaneously constant. The space of strictly proper scoring rules is correspondingly richer: any strictly convex G on the simplex generates a valid rule, and the constraint HG=cIH_G=cI eliminates all non-quadratic generators. The set of scoring rules satisfying this condition is thus a strict subset of the already small set identified in the scalar case. The incentive-theoretic significance is that constant Hessian is the condition under which the agent’s inflation cost is isotropic (direction-independent) and type-independent in magnitude. This is the multi-dimensional analogue of the property that, in the scalar Brier case, allows the principal to achieve first-best with a simple threshold. (c) Generalization of the step-function escape The step-function escape (Theorem 5.8, Part i-a) generalizes to d dimensions, with a geometric caveat. The mechanism generalizes. In d dimensions, a step-function approval rule takes the form q(r)=r∈Aq(r)=1\r∈ A\ for an acceptance region A⊆ℛA . The agent’s decision is still binary: inflate the report into A (at calibration cost) or remain outside A (forgoing approval). This binary choice creates a partition of the type space into inflating and non-inflating types, regardless of d or the form of HGH_G. The proof of Part (i-a) relies on three properties: (i) the calibration cost of inflating from p to A is strictly increasing in the distance d(p,A)d(p,A) (guaranteed by strict convexity of G), (i) the perturbation gain γ is type-independent, and (i) the principal can choose A so that the indifference surface p:γ=C(p,A)\p:γ=C(p,A)\ coincides with ∂A∗∂ A^*. Properties (i) and (i) hold in arbitrary dimension. The caveat: geometry of the acceptance region. Property (i) requires that, for the chosen acceptance region A, the level set of the calibration cost function C(p,A)=infa∈A[S¯(g(p);p)−S¯(a;p)]C(p,A)= _a∈ A[ S(g(p);p)- S(a;p)] coincides with ∂A∗∂ A^*. In one dimension, A∗A^* is a half-line [pmin,1][p_ ,1] and A=[r0,1]A=[r_0,1] with r0=pmin+δr_0=p_ +δ for an appropriate offset δ; the cost function is monotone, so any target boundary is achievable. In d dimensions, the acceptance region A∗A^* may have a curved boundary. Under the Brier score (isotropic Hessian), the cost is Euclidean distance squared, and the indifference surface is the γ/β γ/β-offset of ∂A∂ A. Setting A to be the γ/β γ/β-inner contraction of A∗A^* recovers the first-best boundary exactly when A∗A^* is convex (since inner offsets of convex sets are convex and the offset operation is invertible for offsets smaller than the inradius). When the principal’s welfare function W(θ)W(θ) is linear (as in the oversight game, where W(p)=pus+(1−p)ufW(p)=p\,u_s+(1-p)\,u_f is linear in the type), A∗A^* is a half-space, and the construction is exact in all dimensions. For non-Brier scores, the cost function C(p,A)C(p,A) is anisotropic: the inflation cost depends on the direction of inflation through the eigenstructure of HGH_G. The indifference surface is a deformed offset of ∂A∂ A, with the deformation governed by the local eigenvalues of HGH_G. Matching this deformed surface to ∂A∗∂ A^* requires solving a nonlinear PDE for ∂A∂ A, which may or may not have a solution depending on the compatibility between the anisotropy of HGH_G and the geometry of A∗A^*. This is the multi-dimensional analogue of the scalar welfare gap: when HGH_G varies with p, perfect screening may be unachievable even with step-function approval. Summary. The step-function mechanism (binary inflate-or-not choice creating a type-space partition) is dimension-free. The achievability of first-best through this mechanism depends on the geometric compatibility between the scoring rule’s curvature structure and the shape of the principal’s optimal acceptance region. Under the Brier score with linear welfare, the construction is exact. Under general scores or with nonlinear welfare, exactness is open. Summary of the multi-dimensional scope Result Status for d>1d>1 Key obstacle Perturbation Lemma (3.1) Proved in ℝdR^d None Credibility Impossibility (3.2) Generalizes directly None Optimal non-affinity (Thm. 5.3(i)) Generalizes Geometry of ∂A∗∂ A^* Step-function first-best (Thm. 5.8(i-a)) Generalizes (Brier + linear W) Anisotropy for general G Second-best = first-best (Thm. 5.3(i)) Brier + convex A∗A^* only Boundary curvature Welfare gap characterization Open Matrix-valued curvature The central impossibility (the principal’s optimal approval is non-affine and makes truthful reporting suboptimal) is a d-dimensional result: the Perturbation Lemma and the impossibility theorem are proved in ℝdR^d, and the non-affinity of optimal screening extends to multi-dimensional type spaces. The scalar restriction binds only for the welfare analysis (the precise welfare gap, the Brier score’s distinguished role, and the second-best-equals-first-best result), where the passage from scalar curvature G′(p)G (p) to the Hessian field HG(p)H_G(p) introduces geometric complications that do not arise in one dimension. 6.7. Open Questions Open Question 6.4 (Dynamic monitoring). The framework is static. Repeated interaction and reputation dynamics (Fudenberg and Levine, 1989) may partially restore truthfulness: if the receiver can credibly threaten to revoke autonomy upon detecting miscalibration, the effective γeff _eff decreases with the discount factor. The full dynamic analysis, connecting to the bandit literature through the monitoring precision of Lemma 5.2(i); equilibrium implications beyond the bandit setting remain open. 7. Conclusion Building on the classical observation that non-affine perturbations make truthful reporting suboptimal under the agent’s combined objective (the Perturbation Lemma), this paper establishes two results about scored reporting systems. First, the endogeneity is unavoidable: the principal’s optimal oversight mechanism is necessarily non-affine (Theorem 5.3), so the principal’s rational design choices are precisely those that trigger the impossibility. Second, a sharp threshold escapes the welfare loss entirely: a step-function approval function achieves first-best screening for every strictly proper scoring rule, because the agent’s binary choice creates a type-space threshold regardless of the generator’s curvature (Theorem 5.8, part i-a). The result applies across domains. We develop two instances in full detail: marketplace operation, where non-modular capacity creates revenue-driven perturbations under sealed-bid execution, and AI agent oversight, where the principal’s approval function is the perturbation. The shared Fenchel conjugate structure (Section 3) provides the enabling machinery that unifies these domains under a common algebraic skeleton. The framework extends naturally to other scored reporting settings, such as credit rating agencies and financial auditors, where analogous perturbation structures arise. Two scope limitations deserve emphasis. The impossibility is binding when the scoring rule is inherited from the institutional context rather than jointly designed with the approval function; when the principal controls both, choosing G=BrierG=Brier and a step-function q achieves first-best (Theorem 5.8, part i-a). The instances develop the binary-outcome scalar-type setting; Section 6.6 analyzes the multi-dimensional extension, showing that the core impossibility (Perturbation Lemma, Credibility Impossibility, optimal non-affinity) generalizes to d-dimensional types, while the welfare analysis (second-best-equals-first-best, welfare gap characterization) remains open due to the passage from scalar curvature to the Hessian field. Proposition 5.9 establishes that the Brier score is uniquely optimal under smooth oversight, with the welfare gap governed by the curvature heterogeneity Var(1/G′(p))Var(1/G (p)). Matching upper bounds, and the corresponding phase transition at the C0,1/C1C^0,1/C^1 boundary of the approval function, are left as open questions. Several further questions also remain open: the dynamic extension to repeated credibility games with reputation dynamics, the multi-reporter equilibrium under competition among strategically interacting reporters, and the multi-dimensional welfare gap characterization. Notation Summary Symbol Meaning Domain Θ Type space All ℛR Report space All S¯ S Expected score function (strictly proper) All g(θ)g(θ) Truthful report function All h(r)h(r) Perturbation payoff All γ Perturbation weight All ℐ=(,π)I=(Y,π) Information structure All bindC_bind Binding conflict set (NT1): types where h(g(θ))<suprh(r)h(g(θ))< _rh(r) All G Credibility game All G Strictly convex potential (Savage–McCarthy characterization) Scoring rules Ψ Strictly convex potential (Fenchel skeleton) Appendix D ΦN _N Standard normal CDF Proposition 3.7 η Alignment mapping Appendix D ν Polymatroid capacity function (monotone submodular) Market κij _ij Non-modularity gap: ν(i)+ν(j)−ν(i,j)ν(\i\)+ν(\j\)-ν(\i,j\) Market R(^)R( b) DSIC revenue Market p True success probability (agent type) AI oversight q(r)q(r) Approval function AI oversight β Calibration weight AI oversight pminp_ First-best approval threshold AI oversight r0r_0 Optimal oversight threshold AI oversight W(p)W(p) Principal’s expected utility from approving type p AI oversight Π(p) (p) Net gain from approval: W(p)−udW(p)-u_d AI oversight q~(p) q(p) Induced screening (approval probability for type p) AI oversight V(r;θ,γ)V(r;θ,γ) Reporter’s combined objective All CsepC_sep Cost of domain separation Welfare CregC_reg Cost of external regulation Welfare WcommitW_commit Welfare under commitment resolution Welfare WorganicW_organic Welfare under organic (unregulated) oversight Welfare F, f Type distribution and its density AI oversight μ Common prior on Θ All us,uf,udu_s,u_f,u_d Principal’s utilities (success, failure, delegation) AI oversight δrep _rep Reputational compliance weight Market rminr_ Approval threshold parameter (generic) AI oversight ς Logistic sigmoid function AI oversight τ Smoothing parameter for sigmoid threshold AI oversight G′(p)G (p) Generator curvature (second derivative of G) Welfare analysis δ(G)δ(G) Welfare gap: W∗−supq∈C1W(q,G)W^*- _q∈ C^1W(q,G) Welfare analysis Appendix A Complete Proof of Lemma 3.1 We provide complete details for all cases, including the zero-gradient case flagged in the initial review. Setup. Let S¯:ℛ×Θ→ℝ S:R× be strictly proper with truthful report g, and let h:ℛ→ℝh:R be C1C^1. The combined objective is V(r;θ,γ)=S¯(r;θ)+γh(r)V(r;θ,γ)= S(r;θ)+γ h(r). By strict properness, g(θ)g(θ) is the unique maximizer of S¯(⋅;θ) S(·\,;θ), and the Hessian HS(θ)≔∇r2S¯(g(θ);θ)H_S(θ) ∇^2_r S(g(θ);θ) is negative definite (so −HS(θ)-H_S(θ) is positive definite with smallest eigenvalue λmin(θ)>0 _ (θ)>0). Part (i): Characterization. (⇐ ) If h is constant on g(Θ)g( ), then ∇h(g(θ))=0∇ h(g(θ))=0 for all θ such that g(θ)g(θ) lies in the interior of g(Θ)g( ). Since Θ has non-empty interior and g is a C1C^1 diffeomorphism (the Jacobian Jg=−[HS]−1∇rθ2S¯J_g=-[H_S]^-1∇^2_rθ S is invertible by negative definiteness of HSH_S), g(Θ)g( ) has non-empty interior, and h being constant on g(Θ)g( ) implies ∇h=0∇ h=0 on this interior. The gradient of V at g(θ)g(θ) is ∇rV(g(θ);θ,γ)=∇rS¯(g(θ);θ)⏟=0+γ∇h(g(θ))⏟=0=0, _rV(g(θ);θ,γ)= _r S(g(θ);θ)_=0+γ ∇ h(g(θ))_=0=0, and the Hessian ∇r2V=HS(θ)+γ∇2h(g(θ))∇^2_rV=H_S(θ)+γ∇^2h(g(θ)). Since h is constant on an open set, ∇2h=0∇^2h=0 there, so ∇r2V=HS(θ)∇^2_rV=H_S(θ), which is negative definite. Hence g(θ)g(θ) is a strict local maximum. To confirm it is the global maximum: for any r∉g(Θ)r∉ g( ), we have V(r;θ,γ)=S¯(r;θ)+γh(r)V(r;θ,γ)= S(r;θ)+γ h(r). Since S¯(g(θ);θ)>S¯(r;θ) S(g(θ);θ)> S(r;θ) (strict properness) and h(g(θ))≥h(r)h(g(θ))≥ h(r) (or h takes arbitrary values outside g(Θ)g( ), but by compactness of Θ and continuity, the scoring gap dominates for r far from g(θ)g(θ)), g(θ)g(θ) is the global maximizer. The precise global argument: by the pointwise growth condition (condition (b)), S¯(r;θ)→−∞ S(r;θ)→-∞ as r→∂ℛr→ or ‖r‖→∞\|r\|→∞. Since h is bounded on any compact subset (C1C^1 on compact closure), the scoring penalty dominates for large deviations, confirming global optimality. (⇒ ) If h is non-constant on g(Θ)g( ), there exists θ0 _0 such that ∇h(g(θ0))≠0∇ h(g( _0))≠ 0 (by the argument in Part (i)) or ∇h(g(θ0))=0∇ h(g( _0))=0 but h is not constant near g(θ0)g( _0) (by the argument in Part (i)). In either case, truthfulness fails for some type. Part (i): Generic destruction (detailed). The gradient of the combined objective at r=g(θ)r=g(θ) is ∇rV(g(θ);θ,γ)=γ∇h(g(θ)) _rV(g(θ);θ,γ)=γ∇ h(g(θ)) ((3.1)). For θ with ∇h(g(θ))≠0∇ h(g(θ))≠ 0, this is nonzero for all γ>0γ>0, so g(θ)g(θ) is not a critical point and hence not a maximizer. The set θ∈bind:∇h(g(θ))≠0\θ _bind:∇ h(g(θ))≠ 0\ has positive measure. To see this, note that h is C1C^1 and non-constant on g(bind)g(C_bind). The set g(bind)g(C_bind) has non-empty interior (since bindC_bind has positive measure and g is a C1C^1 diffeomorphism). Suppose for contradiction that ∇h(g(θ))=0∇ h(g(θ))=0 for all θ∈bindθ _bind. Then ∇h=0∇ h=0 on a set with non-empty interior in ℛR. A C1C^1 function with zero gradient on a connected open set is constant there. This contradicts the assumption that h is non-constant on g(bind)g(C_bind). The zero-gradient case (addressed per ECTA-1 comment). It is possible that ∇h(g(θ0))=0∇ h(g( _0))=0 for isolated types θ0∈bind _0 _bind (e.g., if g(θ0)g( _0) is a critical point of h). This occurs on a set of measure zero in bindC_bind (critical points of a C1C^1 function on a d-dimensional domain form a set of Lebesgue measure zero by Sard’s theorem applied to h∘gh g). For such types, the first-order analysis is inconclusive, and Part (i) provides the resolution. Specifically, suppose h∈C2h∈ C^2 (which we may assume without loss by the NT2 condition requiring non-affinity, which is a second-order condition). At a critical point r0=g(θ0)r_0=g( _0) of h with ∇h(r0)=0∇ h(r_0)=0, the second-order expansion of V around r0r_0 is V(r;θ0,γ)=V(r0;θ0,γ)+12(r−r0)⊤[HS(θ0)+γ∇2h(r0)](r−r0)+O(‖r−r0‖3).V(r; _0,γ)=V(r_0; _0,γ)+ 12(r-r_0) [H_S( _0)+γ∇^2h(r_0)](r-r_0)+O(\|r-r_0\|^3). The matrix HS(θ0)+γ∇2h(r0)H_S( _0)+γ∇^2h(r_0) governs local behavior. There are three sub-cases: Sub-case (i-a): ∇2h(r0)∇^2h(r_0) has a positive eigenvalue λ+>0 _+>0 with eigenvector v. Then v⊤[HS+γ∇2h]v=v⊤HSv+γλ+v [H_S+γ∇^2h]v=v H_Sv+γ _+. Since v⊤HSv<0v H_Sv<0 (negative definite), this becomes positive when γ>|v⊤HSv|/λ+γ>|v H_Sv|/ _+. For such γ, r0r_0 is not a local maximum (the Hessian of V has a positive eigenvalue). The threshold is γ¯local=|v⊤HSv|/λ+ γ_local=|v H_Sv|/ _+, which is finite and positive. Sub-case (i-b): All eigenvalues of ∇2h(r0)∇^2h(r_0) are ≤0≤ 0, so ∇2h(r0)∇^2h(r_0) is negative semi-definite. Then HS+γ∇2hH_S+γ∇^2h is negative definite for all γ>0γ>0 (sum of negative definite and negative semi-definite is negative definite). In this case, r0r_0 remains a local maximum for all γ. However, it may not be the global maximum. By NT1, h(g(θ0))<suprh(r)h(g( _0))< _rh(r). Let r1∈argmaxhr_1∈ h (or any r1r_1 with h(r1)>h(g(θ0))h(r_1)>h(g( _0))). Define ΔS S =S¯(g(θ0);θ0)−S¯(r1;θ0)>0(strict properness), = S(g( _0); _0)- S(r_1; _0)>0 (strict properness), Δh h =h(r1)−h(g(θ0))>0(NT1). =h(r_1)-h(g( _0))>0 (NT1). Then V(r1;θ0,γ)−V(g(θ0);θ0,γ)=−ΔS+γΔh.V(r_1; _0,γ)-V(g( _0); _0,γ)=- S+γ h. This is positive when γ>ΔS/Δh≕γ¯global(θ0)γ> S/ h γ_global( _0). For γ>γ¯globalγ> γ_global, the global maximum of V(⋅;θ0,γ)V(·\,; _0,γ) is not at g(θ0)g( _0), even though g(θ0)g( _0) is a local maximum. Sub-case (i-c): ∇2h(r0)=0∇^2h(r_0)=0 (all second derivatives vanish). Since h is not affine near r0r_0 (NT2), the Taylor expansion must have nonzero terms of order ≥3≥ 3. The analysis requires examining higher-order terms, but the global argument of Sub-case (i-b) applies regardless: NT1 guarantees a point r1r_1 with h(r1)>h(r0)h(r_1)>h(r_0), and for sufficiently large γ, this point dominates. Combining all sub-cases: Define γ¯(θ0)=min(γ¯local,γ¯global) γ( _0)= ( γ_local, γ_global) (with γ¯local=∞ γ_local=∞ in Sub-cases (i-b) and (i-c)). For γ>γ¯(θ0)γ> γ( _0), g(θ0)g( _0) does not maximize V(⋅;θ0,γ)V(·\,; _0,γ). ∎ Appendix B Complete Proof of Theorem 5.3 We provide the complete derivation of the principal’s optimization, including the Myerson reserve-price analogy, first-order conditions, and second-order verification. Step 1 (The principal’s screening problem). The principal commits to an approval function q:[0,1]→[0,1]q:[0,1]→[0,1] and the agent best-responds. The principal’s problem is (B.1) maxq:[0,1]→[0,1]UP(q)=ud+∫01q~(p)⋅Π(p)f(p)p, _q:[0,1]→[0,1]\;U_P(q)=u_d+ _0^1 q(p)· (p)\,f(p)\,dp, where q~(p)=q(rq∗(p)) q(p)=q(r^*_q(p)) is the induced screening function, rq∗(p)=argmaxr[−β(r−p)2+γq(r)]r^*_q(p)= _r[-β(r-p)^2+γ q(r)] is the agent’s best response, and Π(p)=p(us−uf)−(ud−uf) (p)=p(u_s-u_f)-(u_d-u_f) is the principal’s net gain from approving type p. Step 2 (The Myerson analogy: virtual types). The structure of (B.1) parallels Myerson’s ((1981)) optimal auction. In Myerson, the seller maximizes expected revenue ∫v⋅x(v)f(v)v v· x(v)f(v)\,dv subject to IC constraints, leading to the virtual-valuation formulation ∫ψ(v)⋅x(v)f(v)v ψ(v)· x(v)f(v)\,dv with ψ(v)=v−(1−F(v))/f(v)ψ(v)=v-(1-F(v))/f(v). In our problem, the principal maximizes ∫Π(p)⋅q~(p)f(p)p (p)· q(p)f(p)\,dp. The “virtual type” adjustment arises from the agent’s strategic response: the induced screening q~(p) q(p) depends on q through the agent’s best response, creating an IC-like constraint. Under the step-function class q(r)=r≥r0q(r)=1\r≥ r_0\, the agent’s best response creates a mapping from the threshold r0r_0 to the induced screening, which acts as the “IC constraint.” The first-order condition for the optimal r0r_0 is obtained by differentiating UPU_P with respect to r0r_0. Under the step function, the induced screening changes at p=r0−γ/βp=r_0- γ/β (the marginal type that is just indifferent between inflating and not). Denoting this marginal type p¯(r0)=r0−γ/β p(r_0)=r_0- γ/β: (B.2) dUPdr0=−Π(p¯(r0))⋅f(p¯(r0))⋅dp¯dr0=−Π(p¯(r0))⋅f(p¯(r0))⋅1=0. dU_Pdr_0=- ( p(r_0))· f( p(r_0))· d pdr_0=- ( p(r_0))· f( p(r_0))· 1=0. This yields Π(p¯(r0))=0 ( p(r_0))=0, hence p¯(r0)=pmin p(r_0)=p_ , confirming r0=pmin+γ/βr_0=p_ + γ/β. Step 3 (Second-order condition). The second derivative of UPU_P with respect to r0r_0 at the optimum is d2UPdr02=−Π′(pmin)⋅f(pmin)−Π(pmin)⋅f′(pmin)=−(us−uf)⋅f(pmin)<0, d^2U_Pdr_0^2=- (p_ )· f(p_ )- (p_ )· f (p_ )=-(u_s-u_f)· f(p_ )<0, since Π′(p)=us−uf>0 (p)=u_s-u_f>0 and f(pmin)>0f(p_ )>0 (by the full-support assumption). The second-order condition is satisfied, confirming that the critical point is a maximum within the step-function class. Step 4 (Optimality over all approval functions). We now show that the step function is optimal not just within its class but over all measurable q:[0,1]→[0,1]q:[0,1]→[0,1]. The principal’s utility (5.3) is maximized when q~(p) q(p) maximizes the integrand q~(p)⋅Π(p) q(p)· (p) pointwise. Since Π(p)>0 (p)>0 for p>pminp>p_ and Π(p)<0 (p)<0 for p<pminp<p_ , the pointwise optimum is q~∗(p)=p≥pmin q^*(p)=1\p≥ p_ \. The question is whether there exists an approval function q whose induced screening achieves this pointwise optimum. The step function q∗(r)=r≥r0q^*(r)=1\r≥ r_0\ with r0=pmin+γ/βr_0=p_ + γ/β induces q~(p)=p≥pmin q(p)=1\p≥ p_ \ (as shown in the main proof), achieving the pointwise optimum. Hence the step function is globally optimal. Step 5 (Affine q is strictly suboptimal). For any affine q(r)=a+brq(r)=a+br, the induced screening is affine: q~(p)=a+b(p+γb/(2β)) q(p)=a+b(p+γ b/(2β)). The loss relative to the step-function optimum is given by equation (5.5) in the main text. We verify the loss is strictly positive. Consider two cases. If b>0b>0: the affine screening q~ q is increasing in p, crossing the level 1/21/2 at some p0p_0. For p<pminp<p_ with q~(p)>0 q(p)>0, the integrand q~(p)|Π(p)|f(p)>0 q(p)| (p)|f(p)>0 (losses from approving bad types). Since F has full support, ∫0pminq~(p)|Π(p)|f(p)p>0 _0^p_ q(p)| (p)|f(p)\,dp>0. If b=0b=0: the constant q~=a q=a cannot screen, and the loss is a∫0pmin|Π|f+(1−a)∫pmin1Πf>0a _0^p_ | |f+(1-a) _p_ ^1 f>0 for a∈(0,1)a∈(0,1). If b<0b<0: the screening is decreasing, which approves low types more than high types, clearly suboptimal. In all cases, the affine approval function incurs a strictly positive loss. ∎ Appendix C Proof of Proposition 4.8 Proof. Let S¯ S be any strictly proper scoring mechanism with g()=g(b)=b. We establish each NT condition and then invoke the Perturbation Lemma. NT1. The revenue function R(^)=∑ipi∗(^)R( b)= _ip_i^*( b) satisfies R(^)>R()R( b)>R(b) for the perturbed profile ^=(−j,bj+δ) b=(b_-j,b_j+δ) whenever κij>0 _ij>0 and bi>bj+δb_i>b_j+δ. The revenue increase is δ⋅κij>0δ· _ij>0, establishing that truthful execution g()=g(b)=b does not maximize revenue. NT2. The revenue function R is piecewise-linear in b, with slopes that change at the breakpoints where the Edmonds greedy ordering changes. Non-modularity (κij>0 _ij>0) ensures that distinct regions of the bid space have distinct slopes, making R globally non-affine. To see this formally, consider two bid profiles ^(1) b^(1) and ^(2) b^(2) that induce different greedy orderings. In the region where agent i is processed before agent j, the marginal revenue from increasing b^j b_j is κij _ij (per the Perturbation Lemma). In the region where agent j is processed first, the roles reverse and the marginal revenue from increasing b^i b_i is κji=κij _ji= _ij (by symmetry of the non-modularity gap). The revenue function has different gradients in these two regions, confirming non-affinity. NT3. The sealed-bid information structure means agent i observes only (bi,xi,pi)(b_i,x_i,p_i). The observation (xi,pi)(x_i,p_i) under the perturbed execution b is identical to the observation under truthful execution of the profile b (which is a legitimate bid profile). Agent i cannot determine whether the operator inflated bjb_j or whether agent j genuinely bid bj+δb_j+δ. Formally, for each agent i, the conditional distribution of (xi,pi)(x_i,p_i) given bib_i is the same under the two scenarios: • Operator inflates: true bids b, executed as ^=(−j,bj+δ) b=(b_-j,b_j+δ). • Truthful execution under b: true bids b, executed faithfully. Since agent i does not observe bjb_j or b^j b_j, only its own outcome (xi,pi)(x_i,p_i), the two scenarios are indistinguishable. Conclusion. NT1–NT3 hold for any strictly proper S¯ S. The Perturbation Lemma (Lemma 3.1) applies, establishing the impossibility. The form-independence is the key point: the impossibility depends on the structure of the perturbation (non-modularity creates non-affine revenue) and the information structure (sealed-bid prevents detection), not on the specific functional form of the reputational score S¯ S. Any mechanism that (i) uniquely pins truthful execution as the maximizer of a scoring function and (i) operates in a sealed-bid environment faces the same impossibility. ∎ Appendix D The Fenchel Conjugate Skeleton This appendix provides the formal details of the four-way unification summarized in Section 3. Definition D.1 (Truthfulness Skeleton). An elicitation game is a tuple (Θ,ℳ,Ψ,η,c)( ,M, ,η,c) where Θ⊆ℝd ^d is a type space, ℳ⊆ℝdM ^d is a message space, Ψ:ℳ→ℝ :M is a strictly convex potential, η:ℳ→ℝdη:M ^d is a continuously differentiable alignment mapping, and c:Θ→ℝc: is a type-dependent baseline. The agent’s utility from message m given type θ is (D.1) U(θ,m)=Ψ(m)+⟨θ,η(m)⟩+c(θ).U(θ,m)= (m)+ θ,η(m) +c(θ). The truthful report t(θ)=argmaxmU(θ,m)t(θ)= _mU(θ,m) is pinned by the first-order condition ∇Ψ(t(θ))+Jη(t(θ))⊤θ=0∇ (t(θ))+J_η(t(θ)) θ=0, which is uniquely solvable by strict convexity of Ψ . Result Ψ η c(θ)c(θ) Domain Savage–McCarthy G(r)G(r) ∇G(r)∇ G(r) S¯(p;p) S(p;p) Scoring rules Archer–Tardos666The Archer–Tardos characterization requires the monotone allocation assumption (xix_i non-decreasing in bib_i), which ensures that truthful bidding is the argmax of the agent’s utility. −∫0bixi(z)z- _0^b_ix_i(z)dz xi(bi)x_i(b_i) 0 DSIC payments Rochet777The skeleton form shown here is the conclusion of Rochet’s cyclical monotonicity theorem, not a primitive input: Rochet’s theorem states that an allocation rule is implementable iff a convex potential exists such that the agent’s utility takes this form. Ψ(θ) (θ) ∇Ψ(θ)∇ (θ) info rent Cyclical monotonicity Gneiting–Raftery S¯(r;r) S(r;r) ∇rS¯(r;r) _r S(r;r) 0 Elicitation When η(m)=mη(m)=m (identity alignment), the agent’s indirect utility becomes the Fenchel conjugate Ψ∗(−θ) ^*(-θ) (Fenchel, 1949; Rockafellar, 1970), and the truthful report satisfies t=(∇Ψ)−1(−⋅)t=(∇ )^-1(-·). Both the Brier score and the Archer–Tardos payment identity fall into this simplified case. The four entries in the table differ in regularity and in whether the potential Ψ is a primitive or a derived object. The Savage–McCarthy entry takes G as the primitive generator; the skeleton is an equivalent representation of properness. The Archer–Tardos entry derives the potential from the allocation rule xix_i; the monotone allocation assumption is a regularity condition not required in the other three entries. The Rochet entry is notable in that the potential Ψ is the conclusion of the cyclical monotonicity theorem, not an input: the theorem states that implementability is equivalent to the existence of such a potential. The Gneiting–Raftery entry uses the expected score at truth as the potential, which coincides with the Savage–McCarthy form under the identification G(r)=S¯(r;r)G(r)= S(r;r) (valid for proper scoring rules). These differences in status (primitive vs. derived, regularity conditions) do not affect the perturbation analysis, which requires only that Ψ be strictly convex and C2C^2. Acknowledgements.The authors thank colleagues at the Future Computing Group, University of Oulu, and the Department of Computer Science, University of Helsinki, for feedback on the marketplace and AI agent oversight instances. This paper is a companion to ongoing work on the welfare-gap phase transition (Hard Rules, Soft Rules) and on multi-level governance composition (Governance Complementarity); both are in preparation. Manuscript preparation used Claude (Anthropic) for drafting assistance. This work was supported by the Research Council of Finland through the 6G Flagship programme (grant 318927), the Strategic Research Council affiliated with the Academy of Finland through the CO2CREATION project (grant 372355), by Business Finland through the Neural pub/sub research project (diary number 8754/31/2022), and by the European Regional Development Fund (ERDF; project numbers A81568, A91867). References J. D. Abernethy and R. M. Frongillo (2012) A characterization of scoring rules for linear properties. In Proceedings of the 25th Annual Conference on Learning Theory (COLT 2012), JMLR Proceedings, Vol. 23, p. 27.1–27.13. Cited by: §1.2, §1.6, §1.6, §3. M. Akbarpour and S. Li (2020) Credible auctions: a trilemma. Econometrica 88 (2), p. 425–467. External Links: Document Cited by: 1st item, §1.6, Remark 4.6. A. Archer and É. Tardos (2001) Truthful mechanisms for one-parameter agents. In Proceedings of the 42nd IEEE Symposium on Foundations of Computer Science (FOCS 2001), p. 482–491. External Links: Document Cited by: §1.2, §4.1. K. J. Arrow (1951) Social choice and individual values. Wiley, New York. Note: Second edition 1963 Cited by: 1st item, §1.2. D. P. Baron and R. B. Myerson (1982) Regulating a monopolist with unknown costs. Econometrica 50 (4), p. 911–930. External Links: Document Cited by: §1.6. B. Becker and T. Milbourn (2011) How did increased competition affect credit ratings?. Journal of Financial Economics 101 (3), p. 493–514. External Links: Document Cited by: Remark 3.8. D. Bergemann, M. Bojko, P. Dütting, R. Paes Leme, H. Xu, and S. Zuo (2024) Data-driven mechanism design: jointly eliciting preferences and information. arXiv preprint. External Links: 2412.16132, Link Cited by: §1.6. D. Bergemann, T. Heumann, and S. Morris (2026) Information design and mechanism design: an integrated framework. arXiv preprint. External Links: 2601.17267, Link Cited by: §1.6. D. Bergemann and S. Morris (2016) Bayes correlated equilibrium and the comparison of information structures in games. Theoretical Economics 11 (2), p. 487–522. External Links: Document Cited by: §1.5, §1.5, §1.6, Remark 2.14. D. Bergemann and S. Morris (2019) Information design: a unified perspective. Journal of Economic Literature 57 (1), p. 44–95. External Links: Document Cited by: §1.5, §1.6, Remark 2.14, Remark 2.14. D. Blackwell (1951) Comparison of experiments. In Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability, p. 93–102. Cited by: Remark 2.14. D. Blackwell (1953) Equivalent comparisons of experiments. Annals of Mathematical Statistics 24 (2), p. 265–272. External Links: Document Cited by: Remark 2.14. G. W. Brier (1950) Verification of forecasts expressed in terms of probability. Monthly Weather Review 78 (1), p. 1–3. External Links: Document Cited by: §1.6. P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei (2017) Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems 30 (NeurIPS), p. 4299–4307. Cited by: §1.1. M. Conforti and G. Cornuéjols (1984) Submodular set functions, matroids and the greedy algorithm: tight worst-case bounds and some generalizations of the Rado–Edmonds theorem. Discrete Applied Mathematics 7 (3), p. 251–274. External Links: Document Cited by: Definition 4.1. V. P. Crawford and J. Sobel (1982) Strategic information transmission. Econometrica 50 (6), p. 1431–1451. External Links: Document Cited by: §1.5, §1.6. B. de Finetti (1937) La prévision: ses lois logiques, ses sources subjectives. Annales de l’Institut Henri Poincaré 7 (1), p. 1–68. Cited by: §1.6, Remark 4.7. P. Dworczak (2020) Mechanism design with aftermarkets: cutoff mechanisms. Econometrica 88 (6), p. 2629–2661. External Links: Document Cited by: §1.6. R. A. Dye (1985) Disclosure of nonproprietary information. Journal of Accounting Research 23 (1), p. 123–145. External Links: Document Cited by: §1.5. W. Fenchel (1949) On conjugate convex functions. Canadian Journal of Mathematics 1, p. 73–77. External Links: Document Cited by: Appendix D, §3. M. V. X. Ferreira and S. M. Weinberg (2020) Credible, truthful, and two-round (optimal) auctions via cryptographic commitments. In Proceedings of the 21st ACM Conference on Economics and Computation (EC ’20), p. 683–712. External Links: Document Cited by: §1.6. T. Fissler and J. F. Ziegel (2016) Higher order elicitability and Osband’s principle. Annals of Statistics 44 (4), p. 1680–1707. External Links: Document Cited by: §1.6, §6.6. D. Fudenberg and D. K. Levine (1989) Reputation and equilibrium selection in games with a patient player. Econometrica 57 (4), p. 759–778. External Links: Document Cited by: Open Question 6.4. M. Gentzkow and E. Kamenica (2011) Bayesian persuasion. American Economic Review 101 (6), p. 2590–2615. External Links: Document Cited by: §1.5, §1.6, item (i), Remark 3.8, Remark 3.9. M. Gentzkow and E. Kamenica (2017) Competition in persuasion. Review of Economic Studies 84 (1), p. 300–322. External Links: Document Cited by: §1.5, Remark 3.8. A. Gibbard (1973) Manipulation of voting schemes: a general result. Econometrica 41 (4), p. 587–601. External Links: Document Cited by: 1st item, 2nd item, §1.2. T. Gneiting and A. E. Raftery (2007) Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association 102 (477), p. 359–378. External Links: Document Cited by: §1.2, §1.6, Remark 2.7, Remark 5.11. C. A. E. Goodhart (1984) Problems of monetary management: the UK experience. In Monetary Theory and Practice: The UK Experience, p. 91–121. Cited by: §1.2, §6.1. J. Green and J. Laffont (1977) Characterization of satisfactory mechanisms for the revelation of preferences for public goods. Econometrica 45 (2), p. 427–438. External Links: Document Cited by: 3rd item. S. J. Grossman (1981) The informational role of warranties and private disclosure about product quality. Journal of Law and Economics 24 (3), p. 461–483. External Links: Document Cited by: §1.5, §1.6. B. Holmström (1979) Moral hazard and observability. Bell Journal of Economics 10 (1), p. 74–91. External Links: Document Cited by: Remark 2.14. B. Holmström (1999) Managerial incentive problems: a dynamic perspective. Review of Economic Studies 66 (1), p. 169–182. Note: Originally circulated 1982 External Links: Document Cited by: Remark 4.7. L. Hurwicz (1972) On informationally decentralized systems. In Decision and Organization, C. B. McGuire and R. Radner (Eds.), p. 297–336. Cited by: 5th item. G. Irving, P. Christiano, and D. Amodei (2018) AI safety via debate. arXiv preprint arXiv:1805.00899. Cited by: §1.1. J. Laffont and J. Tirole (1986) Using cost observation to regulate firms. Journal of Political Economy 94 (3), p. 614–641. External Links: Document Cited by: §1.6. J. Laffont and J. Tirole (1993) A theory of incentives in procurement and regulation. MIT Press, Cambridge, MA. Cited by: Remark 5.5. N. Lambert, D. M. Pennock, and Y. Shoham (2008) Eliciting properties of probability distributions. In Proceedings of the 9th ACM Conference on Electronic Commerce (EC ’08), p. 129–138. External Links: Document Cited by: §1.2, §1.6, §1.6, §3. S. Li (2017) Obviously strategy-proof mechanisms. American Economic Review 107 (11), p. 3257–3287. External Links: Document Cited by: §1.6. Y. Liu, J. Wang, and Y. Chen (2023) Surrogate scoring rules. ACM Transactions on Economics and Computation 10 (3), p. Article 9. External Links: Document Cited by: §1.6. A. Lizzeri (1999) Information revelation and certification intermediaries. RAND Journal of Economics 30 (2), p. 214–231. External Links: Document Cited by: §1.5, §1.5. R. E. Lucas (1976) Econometric policy evaluation: a critique. In The Phillips Curve and Labor Markets, K. Brunner and A. H. Meltzer (Eds.), Carnegie-Rochester Conference Series on Public Policy, Vol. 1, p. 19–46. Cited by: §6.1. G. J. Mailath and L. Samuelson (2001) Who wants a good reputation?. Review of Economic Studies 68 (2), p. 415–441. External Links: Document Cited by: Remark 4.7. D. Manheim and S. Garrabrant (2018) Categorizing variants of Goodhart’s law. arXiv preprint arXiv:1803.04585. Cited by: §6.1. J. McCarthy (1956) Measures of the value of information. Proceedings of the National Academy of Sciences 42 (9), p. 654–655. External Links: Document Cited by: §1.2, §1.6, Remark 2.7. P. R. Milgrom (1981) Good news and bad news: representation theorems and applications. Bell Journal of Economics 12 (2), p. 380–391. External Links: Document Cited by: §1.5, §1.6. P. Milgrom and I. Segal (2002) Envelope theorems for arbitrary choice sets. Econometrica 70 (2), p. 583–601. External Links: Document Cited by: 4th item, Remark 3.6. P. Milgrom (2004) Putting auction theory to work. Cambridge University Press, Cambridge. External Links: Document Cited by: §3. H. Moulin (1980) On strategy-proofness and single peakedness. Public Choice 35 (4), p. 437–455. Cited by: 2nd item, §1.2. R. B. Myerson (1979) Incentive compatibility and the bargaining problem. Econometrica 47 (1), p. 61–73. External Links: Document Cited by: 5th item. R. B. Myerson (1981) Optimal auction design. Mathematics of Operations Research 6 (1), p. 58–73. External Links: Document Cited by: Appendix B, §1.2, Remark 5.10, Remark 5.4, Remark 5.5, Remark 5.6, §6.1. C. Oesterheld and V. Conitzer (2021) Decision scoring rules. In Web and Internet Economics (WINE 2020), Lecture Notes in Computer Science, Vol. 12495, p. 468–481. External Links: Document Cited by: §1.6. J. Rochet (1987) A necessary and sufficient condition for rationalizability in a quasi-linear context. Journal of Mathematical Economics 16 (2), p. 191–200. External Links: Document Cited by: §1.2. R. T. Rockafellar (1970) Convex analysis. Princeton Mathematical Series, Princeton University Press, Princeton, NJ. External Links: Document Cited by: Appendix D, §3. M. A. Satterthwaite (1975) Strategy-proofness and Arrow’s conditions: existence and correspondence theorems for voting procedures and social welfare functions. Journal of Economic Theory 10 (2), p. 187–217. External Links: Document Cited by: 1st item, 2nd item, §1.2. L. J. Savage (1971) Elicitation of personal probabilities and expectations. Journal of the American Statistical Association 66 (336), p. 783–801. External Links: Document Cited by: §1.2, §1.6, Remark 2.7, Remark 4.7. M. J. Schervish (1989) A general method for comparing probability assessors. Annals of Statistics 17 (4), p. 1856–1879. External Links: Document Cited by: §1.2, §1.6, Remark 2.7, §3, Remark 5.11. V. Skreta and L. Veldkamp (2009) Ratings shopping and asset complexity: a theory of ratings inflation. Journal of Monetary Economics 56 (5), p. 678–695. External Links: Document Cited by: Remark 3.8. R. V. Vohra (2011) Mechanism design: a linear programming approach. Cambridge University Press, Cambridge. External Links: Document Cited by: §1.2, §3.