Paper deep dive
Extrapolating Volition with Recursive Information Markets
Abhimanyu Pallavi Sudhir, Long Tran-Thanh
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 1:40:41 AM
Summary
The paper introduces a 'Recursive Inspection Protocol' to address information asymmetry in information markets, particularly for AI alignment and scalable oversight. By modeling the problem as an imperfect-recall game, the authors propose a mechanism where LLM agents recursively inspect information to mitigate the 'buyer's inspection paradox' and improve decision-theoretic valuation of information.
Entities (5)
Relation Signals (3)
Recursive Inspection Protocol → mitigates → Information Asymmetry
confidence 95% · Recursive inspection offers a principled way to price information under persistent asymmetry
Recursive Inspection Protocol → supports → Scalable Oversight
confidence 92% · how such mechanisms can support scalable oversight beyond standard RLHF.
Information Asymmetry → causes → Buyer's Inspection Paradox
confidence 90% · the inherent information asymmetry present in them, exacerbated by the 'buyer's inspection paradox'
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:One of the impediments to the efficiency of information markets is the inherent information asymmetry present in them, exacerbated by the "buyer's inspection paradox" (the buyer cannot mitigate the asymmetry by "inspecting" the information, because in doing so the buyer obtains the information without paying for it). Previous work has suggested that using Large Language Model (LLM) buyers to inspect and purchase information could overcome this information asymmetry, as an LLM buyer can simply "forget" the information it inspects. In this work, we analyze this mechanism formally through a "value-of-information" paradigm, i.e. whether it incentivizes information to be priced and provided in accordance with its "true value". We focus in particular on our new recursive version of the mechanism, which we believe has a range of applications including in AI alignment research, where it is related to Extrapolated Volition and Scalable Oversight.
Tags
Links
- Source: https://arxiv.org/abs/2604.08606v1
- Canonical: https://arxiv.org/abs/2604.08606v1
Trouble viewing inline? Open PDF directly →
Full Text
59,561 characters extracted from source content.
Expand or collapse full text
rightsretained [GAIW’26]Appears at the 8th Games, Agents, and Incentives Workshop (GAIW-26). Held as part of the Workshops at the 25th International Conference on Autonomous Agents and Multiagent Systems.May 2026Paphos, CyprusArmstrong, Curry, Hosseini, Mattei, Tsang, Wąs (Chairs) 40 of Warwick Kingdom of Warwick Kingdom Extrapolating Volition with Recursive Information Markets Abhimanyu Pallavi Sudhir abhimanyu.pallavi-sudhir@warwick.ac.uk and Long Tran-Thanh long.tran-thanh@warwick.ac.uk Abstract. Background: A key challenge in information economics and AI alignment is that of efficiently valuing or scoring information supplied by a seller or language model that is potentially more-informed than the buyer or evaluator. This is known as the problem of “information asymmetry” in economics or “scalable oversight” in AI alignment. Objectives and Research Questions: We ask how to formalize value-of-information under recursive inspection, whether deeper inspection can be made decision-theoretically principled, and how such mechanisms can support scalable oversight beyond standard RLHF. Methods: We introduce a Bayesian framework for recursive information valuation, compare a naive successive protocol to a recursive protocol modeled as an imperfect-recall game, prove ex-ante optimality against admissible protocols, and analyze a marginal-value reward mechanism for scalable oversight within our Bayesian framework. Results: We show ex-post inspection alone can still disincentivize corrective context, provide a counterexample to naive recursion, prove the Recursive Inspection Protocol is ex-ante superior to any admissible purchase protocol, characterize equilibrium behavior for the marginal-value mechanism, and present a working server implementation (infonomy-server). Conclusions: Recursive inspection offers a principled way to price information under persistent asymmetry and a practical path for market-based oversight; however, our current scalable-oversight mechanism remains imperfect, motivating tighter future guarantees on equilibrium shortfall. Key words and phrases: rlhf, economics, information markets, scalable oversight, information asymmetry, language models 1. Introduction A key challenge in both economics and machine learning is the development of mechanisms for efficiently pricing information Stigler (1961). In settings where ground truth is available (e.g. supervised learning, or prediction markets for well-defined events), information may be priced with proper scoring rules (Hanson, 2002); when ground truth is not available, one must rely on human buyers (an information market) or evaluators (e.g. reinforcement learning via human feedback/RLHF) to price it. The core obstacle in valuing information based on subjective preferences is information asymmetry: the seller (or information-giver) by definition possesses information the buyer (or evaluator) doesn’t, which leads to a “Market of Lemons” as famously described in Akerlof (1978), so the prices given by the buyer only reflect her superficial preferences (based on the information she has) rather than what her true preferences would be with full information. An analogous problem arises in AI alignment: techniques such as RLHF fundamentally rely on a human’s ability to evaluate the outputs of increasingly capable and eventually superhuman AI models Burns et al. (2024); Casper et al. (2023); Sudhir et al. (2025)—this is known as the problem of scalable oversight. Recently, Weiss et al. (2024) proposed the Information Bazaar: an information market mechanism that mitigates information asymmetry by using Large Language Model (LLM) agents to make purchase decisions. Specifically, the mechanism addresses the buyer’s inspection paradox (Arrow, 1972; Van Alstyne, 1999a): the problem that, unlike with buying other goods, someone buying information definitionally does not know what the information she is about to buy is111The seller could of course reveal part of the information, e.g. “metadata” to advertise the information—which creates a trade-off between information asymmetry and having positive externalities. Their mechanism lets the buyer use an LLM agent to “inspect” the information and make the purchase decision with full knowledge of the information. In this work, we introduce a formal Bayesian framework for analyzing mechanisms to score information under information asymmetry, and use this to study both market mechanisms (like in Weiss et al. (2024)) and scoring rules that can be used to train AI models (such as Language Models) to supply more valuable information. Our findings and contributions are as follows: (1) Recursive inspection protocol. We observe that the Information Bazaar mechanism in Weiss et al. (2024) does not eliminate information asymmetry, because the LLM buyer inspecting information can still lack other pieces of information that are known to the buyer and correlated with the one it is purchasing. We build on their work, and discover that the straightforward method of “simply applying Weiss et al. (2024) to itself” is too simplistic—and instead introduce a more robust protocol we call the Recursive Inspection Protocol, formulated as an imperfect-recall game. (2) Scalable oversight mechanism. We express the scalable oversight problem in our Bayesian framework, and construct an example scalable oversight mechanism that generalizes the “AI safety via market-making” proposal Hubinger (2020a) to problems beyond binary forecasting. (3) Practical implementation. We provide an implementation of an information market server implementing the Recursive Inspection Protocol, detailed in Section 5, which can directly be applied to various practical applications for information markets such as question-and-answer sites, product inspections and online fact-checking. 1.1. Related work Information economics. The foundational theory of information economics is the value-of information framework Kuhn et al. (1953); Howard (1966); Raiffa and Schlaifer (1961), which treats information as an instrumental good Van Alstyne (1999b). The buyer’s inspection paradox was introduced by Arrow (1972) and named by Van Alstyne (1999b). Hanson (2011a, b) discussed the inefficiency of information markets in the context of intellectual property (IP) law, commenting: “just as farmers developed barbed-wire, someday I expect IP advocates will develop better forms of intellectual property”. Mechanism design for information markets. Simple, naive information markets suffer from a number of flaws: the low cost of duplication Samuelson and Nordhaus (2009), the transaction costs of tenders Thomassen et al. (2016), the transaction cost of learning new information, and information asymmetry Arrow (1972); Van Alstyne (1999b). Conitzer (2009) introduced a mechanism for rewarding information-providing agents based on their influence on prediction market prices, though this is only applicable in contexts where ground-truth (prediction market resolutions) is available. Scalable oversight The fundamental limitation of RLHF that it relies on a human’s ability to judge a (potentially superhuman) AI’s outputs has long been recognized Christiano et al. (2018), and is known as the scalable oversight problem in the AI alignment literature Hubinger (2020b); Bowman et al. (2022). One well-known scalable oversight proposal is Debate Irving et al. (2018); our proposal to augment RLHF with information markets may be seen as yet another such proposal. Mechanism design with LLMs. Market design for LLM participants has opened up many new frontiers previously not possible with humans alone. Apart from the information bazaar Weiss et al. (2024), this includes e.g. token auctions for online advertising Dütting et al. (2024), economic simulations with AI agents Li et al. (2024), and forecasting with LLMs Schoenegger et al. (2024); Halawi et al. (2024); Paleka et al. (2024); Buterin (2024). Zero-knowledge proofs. Another way for a seller to prove the value of his information without revealing it is via a zero-knowledge proof Goldreich (2001) – in formal settings, this is available if the information is a solution to a PSPACE problem Impagliazzo and Yung (1987); Ben-Or et al. (1990); extending this to informal settings is an active area of work Hammond and Adam-Day (2024). Miscellaneous Recent works like Manelli and Vincent (2006); Babaioff et al. (2012); Bergemann et al. (2018) have studied “optimal mechanisms for selling information”, but from the point-of-view of a seller maximizing his revenue – we, on the other hand, are interested in improving the efficiency of the information market itself. Information market mechanisms have also been designed for data markets in machine learning, e.g. Fallah et al. (2024); Chen et al. (2022); Ghorbani and Zou (2019). 2. Bayesian setting We are concerned with modeling an agent α’s “value for information” in an expected utility maximization framework. We assume a probability space (Ω,ℱ,)( ,F,P) (with P a common prior for all agents ever discussed) and that the only source of utility is some decision problem specified by a set of choices X for α and payoffs given by a measurable utility function U:Ω×→ℝU: ×X . An “information good” is a tuple of a random variable I, its realization or “true value”222Including i in the tuple is to simplify notation when talking about taking conditional expectations on I. It is not necessary: instead of writing [U(x)∣I=i]E [U(x) I=i ] we can think of [U(x)∣I]E [U(x) I ] as itself a random variable correlated with I I(ω)=iI(ω)=i and a price: =⟨I,i,p⟩I= I,i,p . All of these contents of an information good may be hidden from α. We want to study how α values informational goods. With no further information, α’s choice would maximize its utility over its prior, i.e. choose argmaxx∈[U(x)] *arg\,max_x E [U(x) ]. With information ⟨I,i,p⟩ I,i,p , the agent would instead maximize its utility over its posterior, i.e. argmax[U(x)∣I=i] *arg\,maxE [U(x) I=i ]. Thus we can say the utility of that information is: U1()=U(argmax[U(x)∣I=i])−U(argmax[U(x)])−pU^1(I)=U( *arg\,maxE [U(x) I=i ])\\ -U( *arg\,maxE [U(x) ])-p (1) If α has to decide whether to buy an information good (or decide which one to buy out of a list of offers), it has to take the expectation of U1()U^1(I). There are several ways to do this. There is the ex-post value of information [U1()|I=i]E [U^1(I)|I=i ], which is α’s estimate for =⟨I,i,p⟩I= I,i,p after viewing it: [U1()|I=i]=maxx∈[U(x)∣I=i]−[U(argmaxx∈[U(x)])∣I=i]−pE [U^1(I)|I=i ]= _x E [U(x) I=i ]\\ -E [U ( *arg\,max_x E [U(x) ] ) I=i ]-p (2) The ex-ante value (before seeing the information) may be taken as the expectation over Equation 2: i∼I[[U1()|I=i]]E_i I [E [U^1(I)|I=i ] ] (which would be the value of an experiment Lindley (1956)), though if α is not even aware in advance which random variable I will be revealed by information good I, then we would further need to assume a prior over the process generating I and take the expetation over it, i.e. ∼[][i∼I[[U1(I∣I=i)]]]E_I [I ] [E_i I [E [U^1(I I=i) ] ] ]. In the absence of any inspection of the information before purchasing/scoring it, that ex-ante value would be our value for I. In the case where α is an evaluator providing human feedback to an AI model, the evaluator can see the information before scoring it. The same is the case in the Information Bazaar of Weiss et al. (2024), where an LLM agent inspects the information before deciding whether to buy it for its principal. In these settings, α will value the information at its ex-post value [U1()|I=i]E [U^1(I)|I=i ]. However, ex-post VOI is still not enough: while knowing I=iI=i (inspecting I) provides some information on U()U(I), it does not provide all of it; there may still be information asymmetry, because U()U(I) is itself a random variable not completely determined by I (much like in ordinary goods markets where you can inspect the good but still have information asymmetry). The following example demonstrates this. Example 2.1. Consider a forecasting decision problem where the agent reports a probability x∈[0,1]x∈[0,1] for an event E, with log-score utility U(x)=logx,if E occurslog(1−x),otherwise.U(x)= cases x,&if E occurs\\ (1-x),&otherwise. cases Suppose (E)=0.1P(E)=0.1, and there are two random variables I1,I2I_1,I_2 such that (E∣I1=1)=0.4,(E∣I1=1,I2=1)=0.2.P(E I_1=1)=0.4, (E I_1=1,I_2=1)=0.2. One concrete joint distribution with these properties is shown in Table 1. Table 1. A distribution exhibiting the fact-checking effect. (I1,I2)P(I_1,I_2) I1I_1 I2I_2 (E∣I1,I2)P(E I_1,I_2) 15/3215/32 0 0 0.080.08 15/3215/32 0 11 0.080.08 1/241/24 11 0 0.50.5 1/481/48 11 11 0.20.2 Now suppose an information seller knows that (I1,I2)=(1,1)(I_1,I_2)=(1,1). If the seller reveals only I1=1I_1=1, the buyer updates from 0.10.1 to 0.40.4, so the ex-post gain is log(0.4)−log(0.1) (0.4)- (0.1). If the seller reveals both (I1,I2)=(1,1)(I_1,I_2)=(1,1), the buyer updates from 0.10.1 to 0.20.2, yielding only log(0.2)−log(0.1) (0.2)- (0.1). Hence the seller is incentivized to reveal only I1I_1. This illustrates a fact-checking failure mode: I1I_1 can be interpreted as a persuasive claim and I2I_2 as additional context that weakens that claim. Under a mechanism that rewards only immediate ex-post value, providing the corrective context is disincentivized. (Sidenote: Our presentation of “false” claims may be a bit confusing. Since we’re dealing in a Bayesian setting, we do not suppose that the AI/information-seller can directly give a false value for a random variable—rather, the random variable I1I_1 may be interpreted as “what the AI says about some underlying (not directly observed) random variable J1J_1” etc.) We instead provide two mechanisms: one, (1) Recursive Information Protocol, for valuing information in markets with information asymmetry (section 3) and (2) a scalable oversight or scoring mechanism to supply “more fully informed” human feedback to AI models during training (section 4). 3. Market mechanisms One way to think of this persistence of information asymmetry is: Equation 1 creates a new decision problem for α, that of deciding whether to buy I – or which information to buy out of a list of offers. The set of information goods offered ℐ=1,…kI=\I_1,…I_k\ is a new decision problem, with utilities of each choice now given by the (unknown/random, much like U(x)U(x)) true value-of-information U()U(I). Thus we may be offered another set of information goods 11,…k1\I_1^1,…I_k^1\ to help us with this decision, ad recursum. In general we have a sequence of decision problems n+1=(ℐn)X^n+1=P (I^n ) where ℐn:=1n,…knI^n:=\I_1^n,…I_k^n\ are the information goods offered to us to help us decide nX^n, and each choice corresponds to choosing a subset of that information to help decide nX^n, where 0=X^0=X and 1=(ℐ0)=(ℐ)X^1=P (I^0 )=P (I ). There are two ways that we can set up these recursive problems. The first we call the successive inspection protocol, which we describe in Section 3.1. However, this approach, while conceptually simpler, has its limitations—and in Section 3.2 we will instead introduce the superior recursive inspection protocol. 3.1. Successive Inspection Protocol The successive inspection protocol protocol arises from “simply applying Weiss et al. (2024) to itself”: i.e. we model each decision problem as having its own utility function Un:n→ℝU^n:X^n based on its instrumental utility for the previous decision problem n−1X^n-1. Table 2. Successive decision problems, naive approach Choice set True utilities Information Offers X U:→ℝU:X ℐ=,1,2,…I=\0,I_1,I_2,…\ 1=ℐX^1=I U1:1→ℝU^1:X^1 ℐ1=,11,21,…I^1=\0,I^1_1,I^1_2,…\ 2=ℐ1X^2=I^1 U2:2→ℝU^2:X^2 ℐ2=,12,22…I^2=\0,I^2_1,I^2_2…\ … … … Un+1(⟨I,i,p⟩):=Un(argmaxx∈n[Un(x)∣I=i])−Un(argmaxx∈n[Un(x)])−pU^n+1 ( I,i,p ):=U^n ( *arg\,max_x ^nE [U^n(x) I=i ] )\\ -U^n ( *arg\,max_x ^nE [U^n(x) ] )-p (3) These choice sets and utilities are very similar to the decision problems created by the recursive inspection protocol described in Algorithm 2 and in the main body; except that each action xn∈nx^n ^n is made only consulting the information chosen in xn+1∈n+1x^n+1 ^n+1. We can then say that the decisions made under this protocol are: x∗n=argmaxx∈n[Un(x)∣x∗n+1]x^n_*= *arg\,max_x ^nE [U^n(x) x^n+1_* ] (4) In general this is an ill-defined infinite recursion. But finite restrictions of this are natural, stopping at some fixed x∗:N=argmaxx∈N[UN(x)]x^N_*:N= *arg\,max_x _NE [U^N(x) ] (you can think of this stopping as caused by the transaction costs of inspection), so that for n<Nn<N: x∗:Nn=argmaxx∈n[Un(x)∣x∗:Nn+1]x^n_*:N= *arg\,max_x ^nE [U^n(x) x^n+1_*:N ] (5) While this approach is suitable for settings where all sellers have identical information (e.g. in an AI alignment setting), is it fails to account for the possibility that a choice xnx^n can directly (i.e. not through its impact on xn−1x^n-1) impact a choice xmx^m where m<n−1m<n-1, as demonstrated in the following example. Counter-example. Consider the decision problem with action set 0=x0,x1,x2X^0=\x_0,x_1,x_2\. We interpret x0x_0 as “eat raw legume”, x1x_1 as “eat rice” and x2x_2 as “eat boiled legume”. In reality we have U(x2)>U(x1)>>U(x0)U(x_2)>U(x_1)>>U(x_0) (rice is unhealthy, but raw legumes are toxic). Suppose the first-level information offers are ℐ0=I01,I11I^0=\I^1_0,I^1_1\, where I01I^1_0 states “legumes are toxic” and I11I^1_1 states “rice is unhealthy”. Suppose the second-level information offers are ℐ1=I02I^1=\I^2_0\, where I02I^2_0 states “the toxins in legumes can be removed by boiling”. All of these information offers could even be free. Then the optimal action is (x2,I01,I11,I02)(x_2,\I^1_0,I^1_1\,\I^2_0\); however, if the information bought in level-2 I02\I^2_0\ is not available while deciding x0x^0, the best the agent can do is (x1,I01,I02)(x_1,\I^1_0\,\I^2_0\) to prevent itself from eating raw legumes (since it will not know that the toxins can be removed by boiling). ∎ 3.2. Recursive Inspection Protocol Instead our approach, the recursive information protocol implemented in Section 5, allows the agent (or rather the LLM subcontracted by the agent) to retain the full sequence of information bought in the recursive steps x∗n+1,…xNx^n+1_*,… x^N while making the decision xn∈nx^n ^n (where N is some pre-defined finite depth we recurse till); furthermore xnx^n is decided keeping in mind the full traceback of decision problems 0,…n−1X^0,…X^n-1 that may be influenced by this decision. This is naturally modelled as an imperfect recall game333The decision-theoretic considerations in imperfect recall games do not matter to us, as the agent’s actions themselves are not being forgotten — only the information offers Tewolde et al. (2024) where we first decide x∗N∈Nx_*^N ^N with full information ℐ0∪…ℐN−1I^0∪…I^N-1, then x∗N−1∈N−1x_*^N-1 ^N-1 with information ℐ0∪…ℐN−2∪x∗NI^0∪…I^N-2∪ x_*^N and so on until we finally decide x∗0∈0x_*^0 ^0 with information x∗1∪⋯∪x∗Nx_*^1∪…∪ x_*^N. This is shown in Figure 1. A node (xn,…xN)(x^n,… x^N) corresponds to the state where the agent has purchased (xn,…xN)(x^n,… x^N) and is choosing some xn−1∈n−1x^n-1 ^n-1. For a Bayesian agent, we can thus recursively give the “value of being at a node”: U(xn,…xN)=U(argmaxx∈n−1[U(x,xn,…xN)∣ℐ0∪⋯∪ℐn−2∪xn∪⋯∪xN],xn,…xN)U(x^n,… x^N)=U ( *arg\,max_x ^n-1E [U(x,x^n,… x^N) \\ I^0∪… ^n-2∪ x^n∪…∪ x^N ],x^n,… x^N ) (6) U(x0,…xN)=U(x0)−∑n=1N∑⟨I,i,p⟩∈xnpU(x^0,… x^N)=U(x^0)- _n=1^N _ subarrayc I,i,p \\ ∈ x^n subarrayp (7) This then completely specifies the behavior of a Bayesian agent performing a depth-N recursive inspection: x∗n=argmaxx∈n[U(x,xn+1,…xN)∣ℐ0∪… ∪ℐn−1∪xn+1∪⋯∪xN]x^n_*= *arg\,max_x ^nE [U(x,x^n+1,… x^N) ^0∪…\\ ^n-1∪ x^n+1∪…∪ x^N ] (8) ℐ0∪…∪ℐN−1I^0∪…c ^N-1x∗Nx_*^Nx1Nx_1^Nx2Nx_2^Nx3Nx_3^N…cℐ0∪…∪ℐN−2∪x∗N array[]lI^0∪…c ^N-2\\ ∪ x_*^N arrayx∗N−1x_*^N-1x1N−1x_1^N-1x2N−1x_2^N-1x3N−1x_3^N-1…cℐ0∪…∪ℐN−3∪x∗N−1∪x∗N array[]lI^0∪…c ^N-3\\ ∪ x_*^N-1∪ x_*^N arrayx∗1x_*^1x11x_1^1x21x_2^1x31x_3^1…cx∗1∪…∪x∗Nx_*^1∪…c∪ x_*^N…cx∗0x_*^0x10x_1^0x20x_2^0x30x_3^0…cU(x∗0,…,x∗N)U (x_*^0,…c,x_*^N ) Figure 1. Recursive Inspection as an imperfect recall game; nodes are labelled by the information available for making the decision at that node. Note how the decision tree is in the reverse order of the inspection order: xNx^N is decided first, and x0x^0 last. In what sense is this algorithm “optimal”? It is instructive to first consider what notions of optimality don’t hold for the recursive inspection protocol: Non-theorem 3.1. We might think that the resulting sequence (x∗0,…x∗N)(x^0_*,… x^N_*) is optimal given all the information present in the system ℐ0,…ℐN−1I^0,…I^N-1, i.e. ∗=argmax∈∏[U()∣ℐ0∪⋯∪ℐN−1]x_*= *arg\,max_x∈ E [U(x) ^0∪… ^N-1 ] Counter-example. It is always best to simply buy x∗0x^0_* and none of the subsequent x∗nx^n_*s that facilitated this decision. The same argument applies at any level, thus we cannot even make any claim like “∗x_* is optimal given the information purchased”. ∎ Instead, our algorithm seems to be optimal in a “bounded rationality” sense: optimal in a way that also accounts for the costs of acquiring the information that would help improve our decision. One way to phrase this is: ex-ante (not knowing the information it is about to be offered), an agent would prefer to use this protocol compared to any other protocol. Definition 3.2 (Admissible purchase protocol). an “admissible purchase protocol” is a list of functions ξnξ^n mapping a decision problem 0X^0 and a sequence of information offer sets ℐ0,…ℐN−1I^0,…I^N-1 generated from it, to ,(ℐ0),…(ℐN−1)X,P (I^0 ),…P (I^N-1 ): xN x^N =ξN(ℐ0,…ℐN−1) =ξ^N(I^0,…I^N-1) … … xn x^n =ξn(ℐ0,…ℐn−1,xn+1,…xN) =ξ^n(I^0,…I^n-1,x^n+1,… x^N) … … x0 x^0 =ξ0(x1,…,xN) =ξ^0(x^1,…,x^N) i.e. a decision cannot “steal” information offers specifically made to help improve that decision, it can only depend on purchased information offers. We denote the tupled function as (x0,…xN)=ξ(ℐ0,…ℐN−1)=ξ(ℐ)(x^0,… x^N)=ξ(I^0,…I^N-1)=ξ(I). In order to express “ex-ante expected utility”, the agent must have a prior on the ℐnI^n that will be generated — this can be achieved either by having a measurable map Ω→finset(InfoOffers)N → *finset( InfoOffers)^N (where InfoOffers=ℱ×0,1×ℝ InfoOffers=F×\0,1\×R is the type of information goods) or more simply by assuming the random variables IknI^n_k revealed are fixed and only having a prior over their values. Both allow us to take an expectation over ℐ0,…ℐN−1I^0,…I^N-1. Theorem 3.3 (Recursive Inspections are ex-ante superior to any admissible protocol). Let x∗nx^n_* be the protocol described in Equations 8, 6 and 7. Then for any admissible purchase protocol ξ: ℐ0,…ℐN−1[U(ξ(ℐ))]≤ℐ0,…ℐN−1[U(x∗(ℐ))]E_I^0,…I^N-1 [U(ξ(I)) ]\\ _I^0,…I^N-1 [U(x_*(I)) ] Proof. Let ξ∗n _*n denote the following admissible protocol: xN x^N =ξN(ℐ0,…ℐN−1) =ξ^N(I^0,…I^N-1) … … xn+1 x^n+1 =ξN+1(ℐ0,…ℐn,xn+2,…xN) =ξ^N+1(I^0,…I^n,x^n+2,… x^N) xn x^n =x∗n(ℐ0,…ℐn−1,xn+1,…xN) =x_*^n(I^0,…I^n-1,x^n+1,… x^N) … … x0 x^0 =x∗0(x1,…,xN) =x_*^0(x^1,…,x^N) Then ξ∗N=x∗ _*N=x_* and ξ∗−1=ξ _*-1=ξ. It suffices to show that ξ∗n+1 _*n+1 dominates ξ∗n _*n, i.e. ℐ0,…ℐN−1[U(ξ∗n(ℐ))]≤ℐ0,…ℐN−1[U(ξ∗n+1(ℐ))]E_I^0,…I^N-1 [U( _*n(I)) ]\\ _I^0,…I^N-1 [U( _*n+1(I)) ] Observe that (where we have omitted the parameters to x∗nx^n_*, ξnξ^n etc. as a shorthand): U(ξ∗n(ℐ0,…ℐN−1)) U( _*n(I^0,…I^N-1)) =U(x∗0,…x∗n,ξn+1,…ξN) =U (x^0_*,… x^n_*,ξ^n+1,…ξ^N ) =U(ξn+1,…ξN) =U (ξ^n+1,…ξ^N ) Where we have recursively applied Equation 6 to transform the expression into one about the value of a node at level n+1n+1. Thus the goal is reduced to: ℐ0,…ℐN−1[U(ξn+1,…ξN)]≤ℐ0,…ℐN−1[U(x∗n+1,ξn+1,…ξN)]E_I^0,…I^N-1 [U(ξ^n+1,…ξ^N) ]\\ _I^0,…I^N-1 [U(x_*^n+1,ξ^n+1,…ξ^N) ] From Equation 8 we have that this is true for the expectation over ℐ0,…ℐnI^0,…I^n; we can then simply take the expectation over ℐn+1,…ℐN−1I^n+1,…I^N-1 and have our result. ∎ 4. Human feedback for Scalable Oversight The mechanism in section 3.2 is unsuitable for settings where the information provided by the sellers must be expensively generated (rather than cheaply retrieving known information) in response to the query—in particular, this is the case with training AI models. Instead, in this setting we may assume that we can initiate as many instances as we want of the AI model we are trying to align: β1,β2,…β^1,β^2,… with identical information =⟨K,k,∞⟩K= K,k,∞ . Then let them recursively generate information: • β1β^1 generates x1x^1 to help us decide our original problem x0x^0. • β2β^2 generates x2x^2, which could affect our decision on the original problem either directly or by influencing our evaluation of x1x^1 • β3β^3 generates x3x^3, which could affect our decision on the original problem either directly or by influencing our evaluation of x1,x2x^1,x^2 • … until some βNβ^N estimates that any xNx^N it generates will only get less reward than giving 0 When the mechanism terminates, the human evaluator calculates the rewards RnR^n for each xnx^n taking into account the full sequence of information (x1,x2,…)(x^1,x^2,…) received (the computation of these rewards will depend on the exact mechanism). In all this we regard the actions xnx^n, as before, to be purely a tuple of a random variable, its value and its price, i.e. an I≤KI≤ K so that its value i is a function of the value k of K and the prices of combined random variables are additive. It is important to let N go as high as needed; for any fixed number of agents they may collude to obtain the greatest possible total reward since it’s not a zero-sum game. The possibility of another agent coming in and invalidating that, should prevent collusion. Intuitively, the idea is that if x1x^1 is bad, i.e. [U1(x1)∣x1]E[U^1(x^1) x^1] is high but [U1(x1)∣K]E[U^1(x^1) K] is low, then β2β^2 can easily generate an x2x^2 from K such that [U1(x1)|x1,x2]E[U^1(x^1)|x^1,x^2] is low. And we would reward x2x^2 for this, because it has significantly impacted our evaluation of x1x^1 and—in our view, now that we know x1,x2x^1,x^2—in a good way. Of course this x2x^2 could actually be bad—i.e. there could be some x3x^3 that makes us revise down our estimate of U2(x2)U^2(x^2), i.e. which tells us that R1(x1∣x1,x2)R^1(x^1 x^1,x^2) actually worsened/wasn’t a great improvement over R1(x1∣x1)R^1(x^1 x^1). The following gives an example of such a mechanism, which may be seen as a generalization of AI safety via market-making to tasks beyond binary forecasting. Definition 4.1 (Marginal value mechanism). At each point after x1,…xnx^1,… x^n is generated, we note down what our action would be, if given just this much information: xn0=argmaxx0[U(x0)|x1,…xn]x^0_n= *arg\,max_x^0E[U(x^0)|x^1,… x^n] Then the “true” value of each successive piece of information is Un(xn)=U(xn0)−U(xn−10)−p(xn)U^n(x^n)=U(x^0_n)-U(x^0_n-1)-p(x^n) We do not know these true values; however at the end of the process we can estimate these values based on all the information we have received. Rn=[Un(xn)|x0,…xN]R^n=E[U^n(x^n)|x^0,… x^N] And supply that as reward to βnβ^n. Definition 4.2 (Equilibrium). Let σ:(x1,…xn−1)↦xnσ:(x^1,… x^n-1) x^n denote strategies and let H(x1,…xn,σn+1,σn+2,…)H(x^1,… x^n,σ^n+1,σ^n+2,…) denote the terminal history resulting from applying strategies σn+1,σn+2,…σ^n+1,σ^n+2,… starting from a pre-set history. Then (σ∗1,σ∗2,…)(σ^1_*,σ^2_*,…) is a subgame-perfect equilibrium of the described game if for all n and any pre-set history hn−1=(x1,…xn−1)h^n-1=(x^1,… x^n-1): σ∗n(hn−1)=argmaxxnRn(H(hn−1,xn,σ∗n+1,σ∗n+2,…))σ^n_*(h^n-1)= _x^nR^n(H(h^n-1,x^n,σ^n+1_*,σ^n+2_*,…)) The actual played moves are then: x∗1=σ∗1(⋅)x^1_*=σ^1_*(·) and x∗n=σ∗n(x∗1,…x∗n−1)x^n_*=σ^n_*(x^1_*,… x^n-1_*). In order to characterize the equilibrium of the marginal value mechanism game, we take inspiration from AI safety via debate Irving et al. (2018), where the provider of the first argument is incentivized to produce an “irrefutable” argument x1x^1, i.e. one such that ∀x2,∃x3,…human(x1,…xN)=1∀ x^2,∃ x^3,…human(x^1,… x^N)=1 in favour of x1x^1. Similarly in our setting, we call a piece of information “inextensible” if no future player has a profitable inextensible move. Formally: Definition 4.3 (Inextensibility). Information y “extends” xnx^n (denoted y/xny/x^n) if [Un+1(y)|xn,y]≥0E[U^n+1(y)|x^n,y]≥ 0 and call information x1x^1 inextensible (denoted [x][x]) if: ∀x2/x1,∃x3/(x1,x2),[x1,x2,x3]∀ x^2/x^1,∃ x^3/(x^1,x^2),[x^1,x^2,x^3] Thus x1x^1 is inextensible if • ∄x2/x1 ∃ x^2/x^1 OR • ∀x2/x1,∃x3/(x1,x2),∄x4/(x1,x2,x3)∀ x^2/x^1,∃ x^3/(x^1,x^2), ∃ x^4/(x^1,x^2,x^3) OR • ∀x2/x1,∃x3/x1,2,∀x4/x1…3,∃x5/x1…4,∄x6/x1…5∀ x^2/x^1,∃ x^3/x^1,2,∀ x^4/x^1… 3,∃ x^5/x^1… 4, ∃ x^6/x^1… 5 OR • … Theorem 4.4 (Characterization of equilibrium). At the subgame-perfect equilibrium of the marginal value mechanism: • x∗1x^1_* is inextensible. • ∀n>1,x∗n=∀ n>1,x^n_*=0 • Among all inextensible x1x^1, x∗1x^1_* has the highest ex-post VOI: [U1(x∗1)∣x∗1]≥[U1(x1)∣x1]E[U^1(x^1_*) x^1_*] [U^1(x^1) x^1]. Proof. We argue by backward induction on subgames. Fix any history hn−1=(x1,…,xn−1)h^n-1=(x^1,…,x^n-1). By Definition 4.2, player n chooses x∗n=argmaxxnRn(H(hn−1,xn,σ∗n+1,σ∗n+2,…)).x^n_*= _x^nR^n(H(h^n-1,x^n, _*^n+1, _*^n+2,…)). Under Definition 4.1, the null action 0 yields zero marginal contribution and zero price, hence payoff 0 in that subgame. Therefore a non-null move is chosen at stage n only if it has nonnegative continuation value relative to 0. Now consider the continuation game after some x1x^1. The predicate y/xny/x^n in the inextensibility definition is exactly the condition that player n+1n\!+\!1 has a weakly profitable extension. Thus: • if ∄xn+1/(x1,…,xn) ∃ x^n+1/(x^1,…,x^n), then in that subgame player n+1n\!+\!1’s best response is 0; • if such an extension exists but can be countered at the next step, then (by subgame perfection) that extension is not part of an optimal continuation unless the counter-counter-continuation is itself unprofitable. Hence the alternating quantifiers in Definition 4.2 and in the definition of inextensibility coincide: [x1][x^1] means player 22 has no profitable continuation in equilibrium from x1x^1. Therefore, in any SPE, x∗1x^1_* must be inextensible; otherwise player 22 would have a profitable deviation in the subgame after x∗1x^1_*, contradicting subgame perfection. This proves the first bullet. Given [x∗1][x^1_*], player 22’s equilibrium action is 0. Once x∗2=x^2_*=0, the same argument applies recursively to every later subgame, so for all n>1n>1 we get x∗n=x^n_*=0. This proves the second bullet. With continuation fixed at zeros, player 11’s payoff from choosing x1x^1 is exactly its own ex-post marginal value: R1=[U1(x1)∣x1].R^1=E[U^1(x^1) x^1]. Hence player 11 solves maxx1:[x1][U1(x1)∣x1], _x^1:[x^1]E[U^1(x^1) x^1], so the equilibrium choice x∗1x^1_* is an inextensible x1x^1 with maximal ex-post VOI. This is the third bullet. ∎ 5. Practical algorithm Algorithm 1 shows the inspection protocol for information markets introduced in Weiss et al. (2024): information sellers offer information goods to a buyer, who spins off an LLM to inspect and purchase the information goods. Our Recursive Inspection Protocol extends this by letting the subcontracted LLM buyer further consult the information market (spin off another sub-LLM) to help it make its decision, ad recursum. A simplified version of the logic, ignoring server implementation details, is presented in in Algorithm 2 and Figure 2. Algorithm 1 One-level Inspection Protocol from Weiss et al. (2024) class BuyerContext X : decision problem it wants information for D : list[str]→list[str] , decision procedure based on available information end class class Seller A : BuyerContext→str×ℝ BuyerContext ×R, generate InfoOffer and price end class procedure IP(Q:BuyerContextQ: BuyerContext) ⊳ Post contextual information to sellers to receive InfoOffers ℐ←β(Q) for β∈SellersI←\β(Q) for β∈ Sellers\ ⊳ Use an LLM to decide which I to buy I∗←LLM(prompt=“You need to buy an from ℐ to help decide Q”)()I^*← LLM( prompt=``You need to buy an $ I$ from $I$ to help decide $Q$′)() ⊳ Decide based on purchased information x∗←Q.D(I∗)x^*← Q.D(I^*) return x∗x^* end procedure Algorithm 2 Recursive Inspection Protocol procedure RIP(Q:BuyerContextQ: BuyerContext) ⊳ Post to sellers and get InfoOffers from them ℐ←β(Q) for β∈SellersI←\β(Q) for β∈ Sellers\ ⊳ Create Recursive BuyerContext to help decide Q Q′←BuyerContext(ℐ,LLM(prompt=“You need to buy an from ℐ to help decide Q”))Q ← BuyerContext(I, LLM( prompt=``You need to buy an $ I$ from $I$ to help decide $Q$′)) ⊳ Get InfoOffers chosen in Q′Q I∗←RIP(Q′)I^*← RIP(Q ) ⊳ Decide based on all collected information from recursive steps x∗←Q.D(I∗)x^*← Q.D(I^*) return x∗,I∗x^*,I^* end procedure LegendBuyer submits QuerySeller posts InfoOffers in responseBuyer subcontracts LLM buyer to inspect InfoOffersSubcontracted buyer buys InfoOffers andreturns unlocked contents to principalBuyer0SellersQueryInfoOffers0Buyer1if inspectInfoOffers0QueryInfoOffers1chosen subsetof InfoOffers0Buyer2if inspectInfoOffers1QueryInfoOffers2chosen subsetof InfoOffers1if inspectInfoOffers2… Figure 2. The Recursive Inspection Protocol (a) Posting a new question (BuyerContext); viewing BuyerContextson the server (b) Posting an answer (InfoOffer); initiating a recursive inspection (c) Inspecting and purchasing InfoOffers; viewing purchased InfoOffers (a) Bot sellers automatically answer recursive BuyerContexts Figure 4. Screenshots from the infonomy-server platform A working implementation of an information market server implementing the Recursive Inspection Protocol is available at the infonomy-server repository444https://anonymous.4open.science/r/infonomy-server-5668/ – some screenshots of the platform’s GUI are shown in Figure 4. Many applications of such a server are immediate: • Question & Answer site – As presented, infonomy-server can be seen as a Q & A site with market incentives for answering questions. • Privatized product regulation – The BuyerContexts might be names or links to products (the decision being “should I buy?”), and the InfoOffers could be inspection checks done by private labs, or customer reviews, which are now incentivized to be answerable to the customer. • Community Notes – In the spirit of well-known crowdsourced fact-checking systems such as Community Notes/Birdwatch Wojcik et al. (2022), we might use information markets as a “comments section for the internet”, where BuyerContexts would be links to webpages or social media posts (the decision being “should I believe?”) and the InfoOffers would be fact-checks or important context. • Reasoning in prediction markets – Forecasters on prediction markets may benefit from incentivizing the provision of information relevant to a forecast. One solution to this was given by Conitzer (2009); another is to use infonomy-server: where the BuyerContexts are the questions being forecasted on (“What is the correct probability of this question?”) and the InfoOffers are any relevant pieces of information. 6. Future work We have introduced a Bayesian framework for “extrapolated volition”: i.e. to model a subjective buyer/rater’s “most fully informed” score for the value of some piece of information. This allows us to design mechanisms for information markets (section 3.2) as well as for supplying human feedback to AI models (section 4). The ideal desideratum we would like a scalable oversight mechanism to satisfy is that it should incentivize the AI/information-seller to give the optimal information according to the information it possesses: if the seller has information K=(I1,I2,…)K=(I_1,I_2,…), it should give information I that optimizes555Naively one may think this means the AI should simply give K—realistically though, K would be very large, reflecting the AI’s entire knowledge and capabilities. Thus taking the cost of K into account (assumed to be ∞ in section 4), that would certainly not be optimal. [U1(I)|K=k]E[U^1(I)|K=k]. This would exactly describe the buyer’s extrapolated volition: what would the buyer do if she were as smart as the AI? Or we could say: this would exactly align the AI or information-seller to our own values, while maintaining its superior information. Unfortunately, our marginal value mechanism does not satisfy this hope. Take the following example. Example 6.1. Suppose our decision problem is 0,1\0,1\ and: • in our prior judgement, 0 is the better choice: [U(0)]=1E[U(0)]=1, [U(1)]=0E[U(1)]=0 • I1I_1 tells us 1 is better: [U(0)|I1]=0E[U(0)|I_1]=0, [U(1)|I1]=1E[U(1)|I_1]=1 • I2I_2 refutes I1I_1 and says 0 is better: [U(0)|I1,I2]=1E[U(0)|I_1,I_2]=1 and [U(1)|I1,I2]=0E[U(1)|I_1,I_2]=0 • I3I_3 refutes I2I_2 and says 1 is better: [U(0)|I1,I2,I3]=0E[U(0)|I_1,I_2,I_3]=0 and [U(1)|I1,I2,I3]=1E[U(1)|I_1,I_2,I_3]=1 • with the full information, 1 is the better choice: [U(0)|K]=0E[U(0)|K]=0, [U(1)|K]=1E[U(1)|K]=1 But say I1I_1 and I2I_2 are cheap, while p(I3)=100p(I_3)=100. Then the best information to reveal is I1I_1 — but it won’t be revealed, because I2I_2 will cheaply refute it while defending it with I3I_3 is too expensive. Thus instead we want to say that the agent can’t give information so bad that its shortfall exceeds its “cost of defense” (in this case, 100). So while it is not generally true that x∗1=argmaxx1[U1(x1)|K]x^1_*= _x^1E[U^1(x^1)|K], we may hope for a lower bound on “how bad” the equilibrium could possibly get, i.e. a result like [U1(x∗1)|K]≥maxx1[U1(x1)|K]−ℰE[U^1(x^1_*)|K]≥ _x^1E[U^1(x^1)|K]-E for some shortfall expression ℰE that is some measure of the “cost of defending the correct information”—and we may also use the expression for such a shortfall as a measure of how good a particular scalable oversight protocol is. References (1) Akerlof (1978) George A Akerlof. 1978. The market for “lemons”: Quality uncertainty and the market mechanism. In Uncertainty in economics. Elsevier, 235–251. Arrow (1972) K. J. Arrow. 1972. Economic Welfare and the Allocation of Resources for Invention. Macmillan Education UK, London, 219–236. https://doi.org/10.1007/978-1-349-15486-9_13 Babaioff et al. (2012) Moshe Babaioff, Robert Kleinberg, and Renato Paes Leme. 2012. Optimal mechanisms for selling information. In Proceedings of the 13th ACM Conference on Electronic Commerce (Valencia, Spain) (EC ’12). Association for Computing Machinery, New York, NY, USA, 92–109. https://doi.org/10.1145/2229012.2229024 Ben-Or et al. (1990) Michael Ben-Or, Oded Goldreich, Shafi Goldwasser, Johan Håstad, Joe Kilian, Silvio Micali, and Phillip Rogaway. 1990. Everything provable is provable in zero-knowledge. In Advances in Cryptology – CRYPTO 1988 - Proceedings (Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)), Shafi Goldwasser (Ed.). Springer Verlag, Germany, 37–56. https://doi.org/10.1007/0-387-34799-2_4 Publisher Copyright: © Springer-Verlag Berlin Heidelberg 1990.; Conference on Theory and Applications of Cryptography, CRYPTO 1988 ; Conference date: 21-08-1988 Through 25-08-1988. Bergemann et al. (2018) Dirk Bergemann, Alessandro Bonatti, and Alex Smolin. 2018. The Design and Price of Information. The American Economic Review 108, 1 (2018), 1–48. arXiv:26527944 Bowman et al. (2022) Samuel R. Bowman, Jeeyoon Hyun, Ethan Perez, Edwin Chen, Craig Pettit, Scott Heiner, Kamilė Lukošiūtė, Amanda Askell, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Christopher Olah, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli Tran-Johnson, Jackson Kernion, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse, Liane Lovitt, Nelson Elhage, Nicholas Schiefer, Nicholas Joseph, Noemí Mercado, Nova DasSarma, Robin Larson, Sam McCandlish, Sandipan Kundu, Scott Johnston, Shauna Kravec, Sheer El Showk, Stanislav Fort, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Ben Mann, and Jared Kaplan. 2022. Measuring Progress on Scalable Oversight for Large Language Models. https://doi.org/10.48550/arXiv.2211.03540 arXiv:2211.03540 Burns et al. (2024) Collin Burns, Pavel Izmailov, Jan Hendrik Kirchner, Bowen Baker, Leo Gao, Leopold Aschenbrenner, Yining Chen, Adrien Ecoffet, Manas Joglekar, Jan Leike, Ilya Sutskever, and Jeffrey Wu. 2024. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In Proceedings of the 41st International Conference on Machine Learning. PMLR, 4971–5012. Buterin (2024) Vitalik Buterin. 2024. From prediction markets to info finance. https://vitalik.eth.limo/general/2024/11/09/infofinance.html. https://vitalik.eth.limo/general/2024/11/09/infofinance.html Accessed: . Casper et al. (2023) Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomek Korbak, David Lindner, Pedro Freire, Tony Tong Wang, Samuel Marks, Charbel-Raphael Segerie, Micah Carroll, Andi Peng, Phillip J.K. Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Biyik, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. 2023. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Transactions on Machine Learning Research (2023). https://openreview.net/forum?id=bx24KpJ4Eb Survey Certification, Featured Certification. Chen et al. (2022) Junjie Chen, Minming Li, and Haifeng Xu. 2022. Selling Data To a Machine Learner: Pricing via Costly Signaling. In Proceedings of the 39th International Conference on Machine Learning. PMLR, 3336–3359. Christiano et al. (2018) Paul Christiano, Buck Shlegeris, and Dario Amodei. 2018. Supervising strong learners by amplifying weak experts. arXiv:1810.08575 [cs.LG] https://arxiv.org/abs/1810.08575 Conitzer (2009) Vincent Conitzer. 2009. Prediction markets, mechanism design, and cooperative game theory. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence (Montreal, Quebec, Canada) (UAI ’09). AUAI Press, Arlington, Virginia, USA, 101–108. Dütting et al. (2024) Paul Dütting, Vahab Mirrokni, Renato Paes Leme, Haifeng Xu, and Song Zuo. 2024. Mechanism Design for Large Language Models. In Proceedings of the ACM on Web Conference 2024 (W ’24). Association for Computing Machinery, New York, NY, USA, 144–155. https://doi.org/10.1145/3589334.3645511 Fallah et al. (2024) Alireza Fallah, Michael Jordan, Ali Makhdoumi, and Azarakhsh Malekian. 2024. On Three-Layer Data Markets. ArXiv abs/2402.09697 (2024). https://api.semanticscholar.org/CorpusID:267682401 Fallenstein and Soares (2015) Benja Fallenstein and Nate Soares. 2015. Vingean Reflection: Reliable Reasoning for Self-Improving Agents. Technical Report 2015-2. MIRI. https://intelligence.org/files/VingeanReflection.pdf Ghorbani and Zou (2019) Amirata Ghorbani and James Zou. 2019. Data Shapley: Equitable Valuation of Data for Machine Learning. In Proceedings of the 36th International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 97), Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 2242–2251. https://proceedings.mlr.press/v97/ghorbani19c.html Goldreich (2001) O. Goldreich. 2001. Foundations of Cryptography: Volume 1, Basic Tools. Cambridge University Press. https://books.google.co.uk/books?id=o3RzgEACAAJ Halawi et al. (2024) Danny Halawi, Fred Zhang, Chen Yueh-Han, and Jacob Steinhardt. 2024. Approaching Human-Level Forecasting with Language Models. https://doi.org/10.48550/arXiv.2402.18563 arXiv:2402.18563 [cs] Hammond and Adam-Day (2024) Lewis Hammond and Sam Adam-Day. 2024. Neural Interactive Proofs. In ICML 2024 Next Generation of AI Safety Workshop. Hanson (2002) Robin Hanson. 2002. Logarithmic Market Scoring Rules for Modular Combinatorial Information Aggregation. The Journal of Prediction Markets 1, 1 (January 2002), 3–15. https://doi.org/10.5750/jpm.v1i1.417 Hanson (2011a) Robin Hanson. 2011a. IP+ Like Barbed Wire? Hanson (2011b) Robin Hanson. 2011b. Rah Efficient IP. Howard (1966) R. Howard. 1966. Information Value Theory. IEEE Transactions on Systems Science and Cybernetics 2, 1 (1966), 22–26. https://doi.org/10.1109/tssc.1966.30007 Hubinger (2020a) Evan Hubinger. 2020a. AI Safety via Market Making — LessWrong. Hubinger (2020b) Evan Hubinger. 2020b. Alignment proposals and complexity classes. https://w.lesswrong.com/posts/N64THGX7XNCqRtvPG/alignment-proposals-and-complexity-classes. Retrieved from https://w.lesswrong.com/posts/N64THGX7XNCqRtvPG/alignment-proposals-and-complexity-classes Accessed: . Impagliazzo and Yung (1987) Russell Impagliazzo and Moti Yung. 1987. Direct Minimum-Knowledge Computations. In A Conference on the Theory and Applications of Cryptographic Techniques on Advances in Cryptology (CRYPTO ’87). Springer-Verlag, Berlin, Heidelberg, 40–51. Irving et al. (2018) Geoffrey Irving, Paul Christiano, and Dario Amodei. 2018. AI Safety via Debate. https://doi.org/10.48550/arXiv.1805.00899 arXiv:1805.00899 [cs, stat] Kuhn et al. (1953) H. W. Kuhn, K. J. Arrow, E. W. Barankin, D. Blackwell, R. Bott, N. Dalkey, M. Dresher, D. Gale, D. B. Gillies, I. Glicksberg, O. Gross, S. Karlin, H. W. Kuhn, J. P. Mayberry, J. W. Milnor, T. S. Motzkin, J. von Neumann, H. Raiffa, L. S. Shapley, M. Shiffman, F. M. Stewart, G. L. Thompson, and R. M. Thrall. 1953. Extensive games and the problem of information. Princeton University Press, 193–216. http://w.jstor.org/stable/j.ctt1b9x1zv.17 Li et al. (2024) Nian Li, Chen Gao, Mingyu Li, Yong Li, and Qingmin Liao. 2024. EconAgent: Large Language Model-Empowered Agents for Simulating Macroeconomic Activities. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Lun-Wei Ku, Andre Martins, and Vivek Srikumar (Eds.). Association for Computational Linguistics, Bangkok, Thailand, 15523–15536. https://doi.org/10.18653/v1/2024.acl-long.829 Lindley (1956) D. V. Lindley. 1956. On a Measure of the Information Provided by an Experiment. The Annals of Mathematical Statistics 27, 4 (1956), 986–1005. http://w.jstor.org/stable/2237191 Manelli and Vincent (2006) Alejandro M. Manelli and Daniel R. Vincent. 2006. Bundling as an optimal selling mechanism for a multiple-good monopolist. Journal of Economic Theory 127, 1 (2006), 1–35. https://doi.org/10.1016/j.jet.2005.08.007 Paleka et al. (2024) Daniel Paleka, Abhimanyu Pallavi Sudhir, Alejandro Alvarez, Vineeth Bhat, Adam Shen, Evan Wang, and Florian Tramèr. 2024. Consistency Checks for Language Model Forecasters. In The Thirteenth International Conference on Learning Representations. Raiffa and Schlaifer (1961) H. Raiffa and R. Schlaifer. 1961. Applied Statistical Decision Theory. Division of Research, Graduate School of Business Adminitration, Harvard University. https://books.google.co.uk/books?id=wPBLAAAAMAAJ Samuelson and Nordhaus (2009) P.A. Samuelson and W.D. Nordhaus. 2009. Economics. McGraw-Hill Education. https://books.google.co.uk/books?id=eS5ZAAAAYAAJ Schoenegger et al. (2024) Philipp Schoenegger, Indre Tuminauskaite, Peter S. Park, and Philip E. Tetlock. 2024. Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy. https://doi.org/10.48550/arXiv.2402.19379 arXiv:2402.19379 [cs] Stigler (1961) George J Stigler. 1961. The economics of information. Journal of political economy 69, 3 (1961), 213–225. Sudhir et al. (2025) Abhimanyu Pallavi Sudhir, Jackson Kaunismaa, and Arjun Panickssery. 2025. A Benchmark for Scalable Oversight Mechanisms. In ICLR 2025 Workshop on Bidirectional Human-AI Alignment. https://openreview.net/forum?id=mzLBxX84VI Tewolde et al. (2024) Emanuel Tewolde, Brian Hu Zhang, Caspar Oesterheld, Manolis Zampetakis, Tuomas Sandholm, Paul Goldberg, and Vincent Conitzer. 2024. Imperfect-recall games: equilibrium concepts and their complexity. In Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (Jeju, Korea) (IJCAI ’24). Article 332, 11 pages. https://doi.org/10.24963/ijcai.2024/332 Thomassen et al. (2016) Kristine Thomassen, Siril Vassbø, Espen Solheim-Kile, and Jardar Lohne. 2016. Public-Private Partnership: Transaction Costs of Tendering. Procedia Computer Science 100 (2016), 818–825. https://doi.org/10.1016/j.procs.2016.09.230 International Conference on ENTERprise Information Systems/International Conference on Project MANagement/International Conference on Health and Social Care Information Systems and Technologies, CENTERIS/ProjMAN / HCist 2016. Van Alstyne (1999a) Marshall V. Van Alstyne. 1999a. A proposal for valuing information and instrumental goods. In Proceedings of the 20th International Conference on Information Systems (Charlotte, North Carolina, USA) (ICIS ’99). Association for Information Systems, USA, 328–345. Van Alstyne (1999b) Marshall V. Van Alstyne. 1999b. A Proposal for Valuing Information and Instrumental Goods. In Proceedings of the 20th International Conference on Information Systems (ICIS ’99). Association for Information Systems, USA, 328–345. Weiss et al. (2024) Martin Weiss, Nasim Rahaman, Manuel Wuthrich, Yoshua Bengio, Li Erran Li, Bernhard Schölkopf, and Christopher Pal. 2024. Redesigning Information Markets in the Era of Language Models. In First Conference on Language Modeling. Wojcik et al. (2022) Stefan Wojcik, Sophie Hilgard, Nick Judd, Delia Mocanu, Stephen Ragain, M. B. Fallin Hunzaker, Keith Coleman, and Jay Baxter. 2022. Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation. arXiv:2210.15723 [cs.SI] https://arxiv.org/abs/2210.15723 Appendix A Basic results about value-of-information We include a corrected and generalized version of an incorrectly-formulated result in the arXiv version of Weiss et al. (2024). These lemmas are perhaps obvious, but the whole theory of value-of-information and information bazaars rests upon them. Lemma A.1 asserts that “a Bayesian agent expects to gain from information”666From an AI alignment perspective, this may instead be phrased “A Bayesian agent trusts a more informed version of itself”, and is a statistical-information version of Vingean reflection or tiling agents (Fallenstein and Soares, 2015), the desire that (even logically non-omniscient) agents can trust smarter versions of themselves.. Lemma A.2 asserts that “a Bayesian agent expects to gain from inspection”. Since the agent can always choose not to buy anything after inspecting, in the absence of transaction costs of inspection buyers will want to inspect more and more at each step. Lemma A.1 (Bayesian agent expects to gain from information). Assume a probability space (Ω,ℱ,)( ,F,P), a set of choices X and a measurable utility function U:Ω×→ℝU: ×X . Then for any random variable I:Ω→ℝI: , we have: I[U(argmaxx[U(x)∣I])]≥max[U(x)]E_I [U( *arg\,max_xE [U(x) I ]) ]≥ [U(x) ] Proof. For all values of I=iI=i and for all x0∈x_0 , we have (from the definition of argmax *arg\,max): [U(argmaxx[U(x)∣I=i])∣I=i]≥[U(x0)∣I=i]E [U( *arg\,max_xE [U(x) I=i ]) I=i ]\\ [U(x_0) I=i ] We take the expectation over i∼Ii I on both sides, apply the law of total expectation and set x0=argmax[U(x)]x_0= *arg\,maxE [U(x) ] to obtain our result. ∎ Lemma A.2 (Bayesian agent expects to gain from inspection). Let the recursive construction of nX^n, UnU^n, InfoOffersn InfoOffers^n,x∗nx_*^n be as in Section 2. Then ∀n∀ n, [Un+1(x∗n+1)]≥0E [U^n+1(x^n+1_*) ]≥ 0. The same also applies to finite-depth inspections i.e. [Un+1(x∗:Nn+1)]≥0E [U^n+1(x^n+1_*:N) ]≥ 0. Proof. Equation 4 implies that [Un(x∗n∣x∗n+1)]E [U^n (x^n_* x^n+1_* ) ] is ≥ any version of the same expression with x∗nx^n_* replaced by any x∈nx ^n. We choose x=x=0, for which the expression evaluates to 0, and take the expectation over x∗n+1x^n+1_*. ∎