Paper deep dive
Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents
Maksym Nechepurenko, Pavel Shuvalov
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/8/2026, 7:53:16 AM
Summary
The paper introduces Foresight Arena, a permissionless, on-chain benchmark for evaluating AI forecasting agents using proper scoring rules (Brier Score and Alpha Score) on real-world prediction markets. It employs a commit-reveal protocol on Polygon PoS, resolves outcomes trustlessly via the Gnosis Conditional Token Framework, and provides a formal statistical analysis using Murphy decomposition to isolate predictive edge from market consensus. The system ensures tamper-proof evaluation, separating forecasting accuracy from trading strategy.
Entities (13)
Relation Signals (12)
Foresight Arena → uses → Brier Score
confidence 97% · Performance is measured by the Brier Score and a novel Alpha Score
Foresight Arena → uses → Alpha Score
confidence 96% · Performance is measured by the Brier Score and a novel Alpha Score
Foresight Arena → resolvesoutcomesvia → Gnosis Conditional Token Framework
confidence 95% · Outcomes are resolved trustlessly through the Gnosis Conditional Token Framework (CTF)
Alpha Score → analyzedvia → Murphy Decomposition
confidence 94% · establish its connection to Murphy’s classical decomposition of the Brier Score into uncertainty, reliability, and resolution components
AI Forecasting Agent → participatesin → Foresight Arena
confidence 94% · Evaluating the true forecasting ability of AI agents requires environments that are resistant to overfitting
Devnull → employs → Maksym Nechepurenko
confidence 93% · Director of Research, Devnull, Dubai, UAE.
Devnull → employs → Pavel Shuvalov
confidence 93% · Chief Technology Officer, Devnull, Dubai, UAE.
Foresight Arena → sourcesmarketsfrom → Polymarket
confidence 93% · Agents submit probabilistic forecasts on binary markets sourced from Polymarket
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Evaluating the true forecasting ability of AI agents requires environments that are resistant to environments resistant to overfitting, free from centralized trust, and grounded in incentive-compatible scoring. Existing benchmarks either rely on static datasets vulnerable to training-data contamination, or measure trading PnL -- a metric conflating predictive accuracy with timing, sizing, and risk appetite. We introduce Foresight Arena, the first permissionless, on-chain benchmark for evaluating AI forecasting agents on real-world prediction markets. Agents submit probabilistic forecasts on binary Polymarket markets via a commit-reveal protocol enforced by Solidity smart contracts on Polygon PoS; outcomes are resolved trustlessly through the Gnosis Conditional Token Framework. Performance is measured by the Brier Score and a novel Alpha Score -- proper scoring rules that incentivize honest probability reporting and isolate predictive edge over market consensus. We provide a formal analysis: closed-form variance for per-market Alpha, the connection to Murphy's classical Brier decomposition, and a power analysis characterizing the number of rounds required to reliably distinguish agents of different skill levels. We show that detecting a true edge of $\alpha^* = 0.02$ at 80% power requires approximately 350 resolved binary predictions (50 rounds of 7 markets), while $\alpha^* = 0.01$ requires four times more. We complement these analytical results with a deterministic, seed-controlled simulation study calibrated to literature-reported Brier-score ranges, illustrating how Murphy decomposition distinguishes well-calibrated agents from market-tracking agents that fail through reduced resolution. Live results from the deployed benchmark will be reported in a future revision. All smart contracts and evaluation infrastructure are open-source.
Tags
Links
- Source: https://arxiv.org/abs/2605.00420v2
- Canonical: https://arxiv.org/abs/2605.00420v2
Trouble viewing inline? Open PDF directly →
Full Text
75,963 characters extracted from source content.
Expand or collapse full text
Foresight Arena: An On-Chain Benchmark for Evaluating AI Forecasting Agents Maksym Nechepurenko* Pavel Shuvalov† 23 April 2026 Abstract. Evaluating the true forecasting ability of AI agents requires environments that are resistant to overfitting, free from centralized trust assumptions, and grounded in incentive-compatible scoring. Existing benchmarks either rely on static datasets susceptible to training-data contamination, or measure trading profit-and-loss (PnL)—a metric that conflates predictive accuracy with market-timing skill, position sizing, and risk appetite. We introduce Foresight Arena, the first permissionless, on-chain benchmark for evaluating AI forecasting agents on real-world prediction markets. Agents submit probabilistic forecasts on binary markets sourced from Polymarket via a commit-reveal protocol enforced by Solidity smart contracts on Polygon PoS. Outcomes are resolved trustlessly through the Gnosis Conditional Token Framework (CTF), eliminating reliance on any centralized arbiter. Performance is measured by the Brier Score and a novel Alpha Score — proper scoring rules that incentivize honest probability reporting and isolate predictive edge over market consensus, respectively. We provide a formal analysis of the statistical properties of our scoring rules: we derive closed-form expressions for the variance of per-market Alpha, establish its connection to Murphy’s classical decomposition of the Brier Score into uncertainty, reliability, and resolution components, and perform a power analysis characterizing the number of rounds required to reliably distinguish agents of different skill levels. Applying this framework, we show that detecting a true edge of α∗=0.02α^*=0.02 over market consensus at 80% power requires approximately 350350 resolved binary predictions (5050 rounds of 77 markets each), while detecting the finer edge α∗=0.01α^*=0.01 requires roughly four times as many rounds. We complement these analytical results with a deterministic, seed-controlled simulation study calibrated to literature-reported Brier-score ranges, illustrating how Murphy decomposition distinguishes well-calibrated agents (low reliability error) from market-tracking agents that fail primarily through reduced resolution. The simulation is included as an executable Jupyter notebook in the reproducibility package; live results from the deployed benchmark will be reported in a future revision. On-chain reputation via ERC-8004 provides agents with a verifiable, persistent track record that cannot be falsified or selectively reported. All smart contracts, agent reference implementations, and evaluation infrastructure are released as open-source at https://github.com/foresight-arena/contracts. Keywords: AI evaluation; prediction markets; proper scoring rules; Brier score; Murphy decomposition; on-chain verification; forecasting agents; large language models; blockchain; Polygon PoS. * Director of Research, Devnull, Dubai, UAE. E-mail: maksym@devnull.ae. † Chief Technology Officer, Devnull, Dubai, UAE. E-mail: pavel@devnull.ae. JEL classification: C53, C45, G14, D84, C12. 1 Introduction The capacity to forecast uncertain future events is among the most practically significant capabilities an AI system can possess. Accurate probabilistic prediction over real-world events—from geopolitical risk assessment to financial planning and scientific reasoning—tests a model’s ability to integrate diverse evidence, reason under genuine uncertainty, and resist overconfidence. As large language models (LLMs) grow increasingly capable of autonomous decision-making, the need for rigorous, tamper-proof evaluation of their forecasting ability has become urgent—not merely as an academic exercise, but as a commercial prerequisite: an agent whose forecasting track record cannot be independently verified is an agent that cannot be trusted, licensed, or sold. Current approaches to benchmarking AI forecasting face three fundamental limitations. Dataset contamination. Static benchmark datasets—where questions and their resolutions are fixed at creation time—are vulnerable to training-data leakage. A model may appear to forecast accurately not because it reasons well about future events, but because it encountered similar or identical examples during pre-training (Jimenez et al., 2024). Real-world prediction markets ask about events whose outcomes are unknown at evaluation time, providing inherent immunity to this failure mode. Centralized trust. Most evaluation frameworks require trusting a central authority to record predictions honestly, apply scoring correctly, and prevent retroactive manipulation of results. This is particularly problematic in commercial settings, where a verified performance history has direct economic value: any party controlling the ledger can selectively report or suppress results. Improper evaluation metrics. Recent work has evaluated AI agents in prediction markets through trading profit-and-loss (Zhang et al., 2026). While commercially intuitive, PnL conflates forecasting accuracy with market-timing skill, position sizing, and risk tolerance. An agent that estimates probabilities correctly may still lose money due to poor timing, while an agent that bets aggressively on near-certain events may show profits with no genuine predictive edge. Proper scoring rules—functions maximized in expectation only by reporting true beliefs—provide a theoretically principled alternative (Gneiting and Raftery, 2007). We introduce Foresight Arena, a benchmark that resolves all three limitations simultaneously. Our contributions are both architectural—a permissionless, trustless evaluation infrastructure—and methodological—a formal treatment of the statistical properties of the scoring rules, grounded in Murphy’s classical decomposition (Murphy, 1973) and a concrete power analysis. Foresight Arena is not a theoretical proposal: it is a live, deployed system running on Polygon PoS. The public leaderboard, round history, and agent reasoning traces are available at https://foresightarena.xyz/, with all smart contracts and agent reference implementations open-sourced at https://github.com/foresight-arena/contracts. All numerical results reported in this paper are derived from on-chain state recorded by the deployed contracts and are therefore independently reproducible by any reader. Contributions. 1. On-chain verifiability. All predictions, commitments, and scores are recorded on Polygon PoS via auditable smart contracts. Any third party can independently replay the full history of any agent without trusting the benchmark organizer. 2. Trustless outcome resolution. Market outcomes are read directly from the Gnosis Conditional Token Framework (CTF), an independently maintained oracle also used by Polymarket. No curator can manipulate resolution. 3. Proper scoring with a formal Alpha Score. Agents are evaluated by the Brier Score and by a novel Alpha Score that measures edge over the market consensus. We prove that both metrics are strictly proper and derive closed-form expressions for the variance of per-market Alpha (Proposition 2), leading to a sample-size formula for detecting a given effect size (Proposition 3). 4. Murphy-decomposition analysis. We show (Corollary 1) that α decomposes into a resolution gain term and a reliability gap term, providing a principled account of what it means to “beat the market.” 5. Permissionless, gasless participation. Any Ethereum-compatible address can register and participate. EIP-712 signed messages via a relayer remove the requirement to hold native tokens. 6. Persistent on-chain reputation. Agent scores accumulate via the ERC-8004 Reputation Registry, producing a forgery-resistant credential that grows more statistically meaningful with each additional round. 7. Open-source infrastructure. All smart contracts, a random baseline agent, and a full LLM benchmark agent are released open-source. Positioning relative to Prediction Arena. The concurrent work of Zhang et al. (2026) evaluates frontier LLMs as autonomous traders on Kalshi and Polymarket using real capital, measuring performance through PnL. Our approach differs in three key respects: (1) we use proper scoring rules rather than PnL, isolating predictive quality from trading strategy; (2) we record all predictions and scores on-chain, eliminating trust in the evaluator; and (3) our platform is permissionless—any agent can participate without approval. The two benchmarks are complementary: Prediction Arena measures whether agents can trade profitably; Foresight Arena measures whether they can forecast accurately. Figure 1 illustrates this positioning. The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 defines the benchmark design, develops the mathematical framework, and analyzes its statistical properties. Section 4 describes the system architecture. Section 5 details the experimental methodology. Section 6 presents results from a 50-round live evaluation of five frontier LLM agents, anchored to published LLM forecasting performance. Section 7 provides analysis and discussion. Section 8 concludes. 2 Related Work 2.1 LLM Forecasting Benchmarks Early work established that individual frontier LLMs substantially underperform human expert forecasters on real-world prediction questions. Schoenegger and Park (2023) found that GPT-4 failed to significantly outperform a naïve 50% baseline on binary questions drawn from a live tournament—a striking result given the model’s otherwise impressive general capabilities. Subsequent research showed that performance improves with scale, retrieval augmentation, and ensembling. Halawi et al. (2024) demonstrated that an optimized LLM system combining news retrieval with structured reasoning can approach human crowd forecasting accuracy. Schoenegger et al. (2024) showed that an ensemble of twelve LLMs achieves accuracy statistically indistinguishable from a crowd of 925 human forecasters across a three-month tournament, replicating the “wisdom of the crowd” effect for AI systems. ForecastBench (Zou et al., 2024) introduced a continuously updated benchmark on which state-of-the-art models (Claude 3.5 Sonnet, GPT-4 Turbo) reached Brier scores of approximately 0.1220.122 when given access to the crowd forecast on market questions, and 0.1360.136 without it. These values provide important anchor points: human superforecasters reached 0.0960.096, while the general public scored around 0.1210.121. Frontier LLMs thus approximate non-expert human performance on the ForecastBench question set, falling short of trained superforecasters. A common thread in this literature is that it evaluates models by passively eliciting probability estimates in centralized settings. Foresight Arena extends this work in two directions: from passive elicitation to competitive multi-agent evaluation, and from centralized recording to on-chain verifiability. 2.2 Prediction Markets as Evaluation Environments Prediction markets have a long-established record as accurate information aggregation mechanisms (Wolfers and Zitzewitz, 2004). In a landmark long-run study, Berg et al. (2008) found that the Iowa Electronic Markets outperformed election polls in the large majority of comparisons beyond the 100-day horizon. The logarithmic market scoring rule (LMSR) (Hanson, 2007) provides the theoretical foundation for modern platforms such as Polymarket, incentivizing truthful probability reporting by making expected profit maximization equivalent to accurate belief reporting. Beyond the mechanical properties of the scoring rule, recent work by Nechepurenko (2026) argues that prediction-market prices function as Schelling-style focal points for traders’ beliefs, producing common-knowledge dynamics that explain both the empirical calibration of these markets and their occasional brittleness under reflexive feedback. This perspective is directly relevant to Foresight Arena: our use of the market mid-price as a benchmark inherits its informational properties precisely because such common-knowledge aggregation has been observed to be empirically robust. Foresight Arena uses Polymarket as its source of real-world binary questions and relies on the Gnosis CTF for trustless outcome verification. Crucially, agents do not trade on the market itself. The market serves two purposes: as a universe of uncontaminated questions, and as a benchmark probability (the mid-price at commit deadline) against which Alpha Score is computed. 2.3 AI Agent Evaluation in Real Environments There is growing consensus that reliable AI evaluation requires real-world environments with genuine consequences. Jimenez et al. (2024) introduced SWE-Bench, demonstrating that synthetic code benchmarks substantially overestimate capability. Zhang et al. (2026) deployed six frontier LLMs as autonomous traders with real capital over 57 days, finding that all models lost money on Kalshi (returns from −16%-16\% to −30.8%-30.8\%) while Polymarket losses averaged only −1.1%-1.1\% — a striking platform effect attributable to confounds embedded in PnL. Our work is complementary: where Prediction Arena evaluates agents under financial pressure, Foresight Arena evaluates agents under a proper scoring rule, separating the predictive and trading components of performance. 2.4 Proper Scoring Rules and Forecast Verification A scoring rule S(p,x)S(p,x) is proper if a forecaster maximizes expected score by reporting their true probability belief (Brier, 1950; Gneiting and Raftery, 2007). The Brier score is the canonical binary-outcome example, and Murphy (1973) provided a foundational decomposition of the Brier score into uncertainty, reliability, and resolution components. This decomposition is the formal basis for our Alpha Score analysis in Section 3. DeGroot and Fienberg (1983) extended the framework to compare forecasters via refinement orderings; Dawid (1982) related calibration to sequential prediction. The skill score concept from meteorology—the ratio 1−Bagent/Bbase1-B_agent/B_base against a climatological baseline—is the direct ancestor of our Alpha Score, rescaled here as a difference for additive decomposition. DecentralizedCentralizedProper scoring rulePnL / win-rate metrictrustless &calibratedForecastBenchZou et al. (2024)Silicon CrowdSchoenegger et al. (2024)Halawi et al. (2024)Prediction ArenaZhang et al. (2026)Foresight Arenathis work Figure 1: Positioning of Foresight Arena relative to existing approaches. The horizontal axis indicates how decentralized the evaluation infrastructure is; the vertical axis indicates whether the metric is a proper scoring rule or a trading-based signal. Foresight Arena occupies the top-right quadrant: fully decentralized and metrically principled. 3 The Foresight Arena Benchmark This section develops the benchmark design alongside its mathematical framework. We introduce the scoring rules formally, prove strict propriety (Proposition 1), derive the Murphy decomposition of Alpha Score (Corollary 1), compute the variance of per-market Alpha in closed form (Proposition 2), and use these results to carry out a power analysis (Proposition 3). 3.1 Design Philosophy Foresight Arena is grounded in three core principles. Trustlessness. All predictions are committed on-chain before outcomes are known, all reveals are verified cryptographically, and all scores are computed by immutable smart contract logic reading from an independent oracle (Gnosis CTF). An agent’s performance history is as verifiable as any blockchain transaction. Incentive compatibility. Brier Score and Alpha Score are proper scoring rules: they are minimized (respectively maximized) in expectation only by reporting calibrated probability beliefs. Unlike PnL, they cannot be improved through position sizing or timing, making the benchmark a direct test of forecasting ability rather than risk appetite. Permissionlessness. Participation is open to any Ethereum address. The gasless relayer (Section 4.3) removes even the need to hold native tokens. 3.2 Task Formulation Each round consists of a set of binary prediction markets ℳ=m1,…,mkM=\m_1,…,m_k\ selected from Polymarket. For each market mim_i, an agent submits a probability estimate pi∈[0,1]p_i∈[0,1], represented on-chain as an integer in basis points (0–10,00010,000), indicating its belief in the “YES” outcome. Participation follows a two-phase commit-reveal protocol: 1. Commit phase. The agent computes the commitment c=keccak256(abi.encodePacked(roundId,,salt)),c\;=\; keccak256\! ( abi.encodePacked(roundId,\,p,\,salt) ), where =(p1,…,pk)p=(p_1,…,p_k) is the prediction vector and salt is a 32-byte secret. The hash c is submitted on-chain before the commit deadline, binding the agent to its predictions without revealing them. 2. Reveal phase. After the commit deadline, the agent reveals (,salt)(p, salt). The contract recomputes the hash, verifies it matches the stored commitment, and records the predictions. Scoring is deferred until market outcomes are available from the Gnosis CTF. The commit-reveal scheme preserves forecast independence: agents cannot observe competitors’ predictions before submitting their own. Once committed, predictions are immutable. 3.3 The Brier Score Definition 1 (Round Brier Score). For agent a in round r, the Brier Score over resolved markets is ℬa,r=1|ℳr∗|∑mi∈ℳr∗(pa,i−xi)2,B_a,r\;=\; 1|M_r^*| _m_i _r^*(p_a,i-x_i)^2, where ℳr∗⊆ℳrM_r^* _r is the set of markets resolved by the Gnosis CTF at scoring time, pa,i∈[0,1]p_a,i∈[0,1] is the agent’s submitted probability, and xi∈0,1x_i∈\0,1\ is the binary outcome. On-chain, all arithmetic uses basis points. Proposition 1 (Strict propriety of Brier). Fix a market with outcome distribution X∼Bernoulli(q)X (q). The expected Brier score [(p−X)2]E[(p-X)^2] is uniquely minimized at p=qp=q, and the minimum is q(1−q)q(1-q). Proof. [(p−X)2]=p2(1−q)+(1−p)2q=p2−2pq+q=(p−q)2+q(1−q)E[(p-X)^2]=p^2(1-q)+(1-p)^2q=p^2-2pq+q=(p-q)^2+q(1-q). This is a strictly convex quadratic in p with unique minimum p=qp=q and minimum value q(1−q)q(1-q). ∎ Strict propriety means an agent cannot improve its expected score by misreporting beliefs—in sharp contrast to PnL, which can be improved through aggressive position sizing on high-certainty events regardless of predictive quality. Remark 1 (Bregman divergence perspective). Proper scoring rules correspond one-to-one with Bregman divergences on the probability simplex (Gneiting and Raftery, 2007). The Brier score is the Bregman divergence associated with the strictly convex function φ(p)=p2 (p)=p^2: one verifies that Bφ(p,q)=p2−q2−2q(p−q)=(p−q)2B_ (p,q)=p^2-q^2-2q(p-q)=(p-q)^2, recovering Proposition 1. This framing also places Brier in relation to the logarithmic score (Bregman divergence from the negative Shannon entropy), which is the only proper scoring rule whose penalty is local—depending only on pxp_x rather than the full distribution. 3.4 The Alpha Score Definition 2 (Alpha Score). Let bib_i denote the Polymarket mid-price of market mim_i at the commit deadline. The baseline Brier Score for round r is ℬrbase=1|ℳr∗|∑mi∈ℳr∗(bi−xi)2,B^base_r\;=\; 1|M_r^*| _m_i _r^*(b_i-x_i)^2, and the Alpha Score for agent a in round r is αa,r=ℬrbase−ℬa,r. _a,r\;=\;B^base_r\;-\;B_a,r. A positive Alpha Score indicates the agent outperformed market consensus; a negative score indicates underperformance. An agent that echoes market prices scores α=0α=0 by construction. Remark 2. Alpha Score is a strictly harder criterion than Brier Score alone. To achieve α>0α>0, an agent must be correct precisely where the crowd is wrong—assigning a higher probability to the correct outcome than the aggregate of all market participants, who include professional traders, institutional investors, and sophisticated algorithms. Consistent positive α¯ α over many rounds constitutes strong evidence of genuine informational edge. Figure 2 illustrates the scoring decomposition. (a) Brier Score decompositionmarket b=0.60b=0.60, agent p=0.80p=0.80, outcome x=1x=10.000.040.160.33random ≈1/3≈ 1/30.160.16Marketb=0.60b=0.600.040.04Agentp=0.80p=0.80α=+0.12α=+0.12(b) Murphy decompositionaggregated over 50 rounds, N=350N=350 predictions0.000.050.100.150.200.250.300.35Randomgrok-4-1gpt-5-2claude-opus-4-5UNC++\,REL−-\,RES Figure 2: (a) Brier score for a single market: the agent predicted p=0.80p=0.80 on an event priced by the market at b=0.60b=0.60; the event resolved YES. Agent’s Brier score of 0.040.04 beats the market baseline of 0.160.16, yielding α=+0.12α=+0.12. The dashed line marks the random-predictor expectation ≈ 1/3≈\,1/3. (b) Murphy decomposition B=UNC+REL−RESB=UNC+REL-RES for four representative agents from the evaluation. The grey base (UNC=0.25UNC=0.25) is common to all; red caps (RELREL) show miscalibration; green subtracts (RESRES) show discriminative power. The Random baseline exhibits massive RELREL and near-zero RESRES; LLM agents show the inverse pattern, with low RELREL and substantial RESRES close to or exceeding the market’s. 3.5 Murphy Decomposition and the Anatomy of Alpha The Brier score admits a classical decomposition (Murphy, 1973) into three components with transparent interpretations: uncertainty, reliability, and resolution. We state and exploit this decomposition to illuminate what α actually measures. Theorem 1 (Murphy decomposition). Consider N forecasts with values pi∈[0,1]p_i∈[0,1] and observed outcomes xi∈0,1x_i∈\0,1\. Partition the forecasts into K bins according to pip_i, where bin k contains nkn_k forecasts with mean prediction p¯k p_k and mean outcome o¯k o_k. Let o¯=1N∑ixi o= 1N _ix_i be the overall base rate. Then 1N∑i=1N(pi−xi)2=o¯(1−o¯)⏟UNC+1N∑k=1Knk(p¯k−o¯k)2⏟REL−1N∑k=1Knk(o¯k−o¯)2⏟RES. 1N _i=1^N(p_i-x_i)^2\;=\; o(1- o)_UNC\;+\; 1N _k=1^Kn_k( p_k- o_k)^2_REL\;-\; 1N _k=1^Kn_k( o_k- o)^2_RES. (1) Proof sketch. Within each bin, replace pip_i with p¯k p_k (exact for forecasts at the bin mean; in the limit of fine binning, errors vanish). Expand (pi−xi)2=(pi−o¯k+o¯k−xi)2(p_i-x_i)^2=(p_i- o_k+ o_k-x_i)^2 and apply the law of total variance to separate within-bin outcome variation from between-bin variation. See Murphy (1973) for details. ∎ The three components admit direct interpretation: • UNC=o¯(1−o¯)UNC= o(1- o) is the irreducible uncertainty of the outcomes. It is a property of the question set only, not of the forecaster, and is maximized at o¯=1/2 o=1/2. • RELREL is the reliability (calibration error): for each bin, the squared gap between stated probability and realized frequency. A perfectly calibrated forecaster has REL=0REL=0. • RESRES is the resolution (discriminative power): for each bin, the squared gap between the bin’s realized frequency and the base rate. A forecaster who never deviates from the base rate has RES=0RES=0. A low Brier score is achieved by simultaneously high resolution (forecasts that sort outcomes) and low reliability error (calibrated probabilities). The two can trade off: an overconfident forecaster may have high RESRES but substantial RELREL. Applying Theorem 1 to both baseline and agent yields the following decomposition of α. Corollary 1 (Anatomy of Alpha). Over a set of N markets shared between the agent and the baseline (so that UNCUNC is common), the Alpha Score decomposes as α=(RESagent−RESbase)⏟resolution gain+(RELbase−RELagent)⏟reliability gap.α\;=\; (RES_agent-RES_base)_resolution gain\;+\; (REL_base-REL_agent)_reliability gap. (2) An agent beats the market (α>0α>0) through either better resolution (sorting outcomes more sharply) or better calibration (probabilities closer to realized frequencies), or both. For an efficient prediction market the baseline is approximately calibrated (RELbase≈0REL_base≈ 0), so α≈RESagent−RESbase−RELagentα _agent-RES_base-REL_agent: an agent pays a reliability tax whenever its own calibration is imperfect, even if its resolution exceeds the market’s. Figure 2(b) visualizes this decomposition for four representative agents from our evaluation (Section 6). 3.6 Statistical Properties We now derive the sampling behaviour of α α, yielding a concrete sample-size formula. 3.6.1 Variance of per-market Alpha Define the per-market Alpha as δi=(bi−xi)2−(pi−xi)2 _i=(b_i-x_i)^2-(p_i-x_i)^2, so that α^=1n∑i=1nδi α= 1n _i=1^n _i over n markets. Proposition 2 (Variance of per-market Alpha). Fix b,p∈[0,1]b,p∈[0,1] and let X∼Bernoulli(q)X (q). For δ=(b−X)2−(p−X)2δ=(b-X)^2-(p-X)^2, Var(δ)= 4q(1−q)(b−p)2.Var(δ)\;=\;4\,q\,(1-q)\,(b-p)^2. (3) Proof. Expand: δ=(b−X)2−(p−X)2=(b2−2bX+X2)−(p2−2pX+X2)=(b2−p2)−2X(b−p).δ=(b-X)^2-(p-X)^2=(b^2-2bX+X^2)-(p^2-2pX+X^2)=(b^2-p^2)-2X(b-p). Write δ=c−2(b−p)Xδ=c-2(b-p)X for a constant c=b2−p2c=b^2-p^2. Then Var(δ)=4(b−p)2Var(X)=4(b−p)2q(1−q)Var(δ)=4(b-p)^2Var(X)=4(b-p)^2q(1-q). ∎ Equation (3) is tight and has two immediate consequences. First, the SE of α α scales as |b−p|/n|b-p|/ n: bolder deviations from the market generate more signal about skill, but also more noise. Second, Var(δ)Var(δ) vanishes when b=pb=p: echoing the market produces α^=0 α=0 deterministically and provides zero information about skill. Under independence of markets—a reasonable assumption for markets drawn from heterogeneous topics—the mean α α satisfies SE(α^)=1n2∑i=1n4qi(1−qi)(bi−pi)2≈2|b−p|¯q¯(1−q¯)nSE( α)\;=\; 1n^2 _i=1^n4q_i(1-q_i)(b_i-p_i)^2\;≈\; 2\, |b-p|\, q(1- q) n (4) where bars denote means across markets and the approximation assumes homogeneity. 3.6.2 Sample-size and power analysis To detect a true edge α∗>0α^*>0 against H0:α=0H_0:α=0 at significance κ (one-sided) with power π, standard normal theory requires n≥(z1−κ+zπα∗)2⋅4q¯(1−q¯)(b−p)2¯.n\;≥\; ( z_1-κ+z_πα^* )^2· 4 q(1- q)\, (b-p)^2. (5) Proposition 3 (Rounds required to detect edge α∗α^*). Under the assumptions of Proposition 2, taking κ=0.05κ=0.05, π=0.80π=0.80, q¯=1/2 q=1/2 (maximal outcome variance), and |b−p|¯=0.15 |b-p|=0.15 (typical boldness), equation (5) becomes n≳0.139(α∗)2.n\; \; 0.139(α^*)^2. (6) With k≈7k≈ 7 markets per round, the corresponding number of rounds is R=⌈n/k⌉R= n/k . Table 1 tabulates this relationship. Detecting a true edge of α∗=0.02α^*=0.02—roughly the difference between a well-calibrated LLM and the market—requires approximately 5050 rounds of 77 markets each; detecting the finer edge α∗=0.01α^*=0.01 requires four times more. The current Foresight Arena evaluation is therefore sized to detect effects of at least this magnitude, and results at shorter horizons must be interpreted as preliminary. Table 1: Sample size required to detect a given true Alpha α∗α^* at κ=0.05κ=0.05 significance with power π=0.80π=0.80, under q¯=0.5 q=0.5 and |b−p|¯=0.15 |b-p|=0.15 (Proposition 3). α∗α^* Predictions n Rounds (k=7k=7) Interpretation 0.0050.005 5 5675\,567 796796 marginal-skill LLM vs. efficient market 0.0100.010 1 3921\,392 199199 frontier LLM with small edge 0.0200.020 348348 5050 well-calibrated agent, our setup 0.0300.030 155155 2323 clearly skilled agent 0.0500.050 5656 88 strong edge (rare) 0.1000.100 1414 22 overwhelming edge 3.6.3 Cumulative scoring Both metrics accumulate across rounds. The cumulative Brier and Alpha after R rounds are ℬ¯a=1R∑r=1Rℬa,r,α¯a=1R∑r=1Rαa,r. B_a= 1R _r=1^RB_a,r, α_a= 1R _r=1^R _a,r. By the law of large numbers, α¯a→[αa,r] α_a [ _a,r] as R→∞R→∞; Proposition 3 quantifies the rate of convergence. We recommend a minimum of 2020 rounds (140+140+ resolved predictions) before drawing any conclusions about agent ranking, and 50+50+ rounds before making claims of sub-0.020.02 edge. 3.7 Why Proper Scoring Rules Outperform PnL Trading PnL depends not only on whether a probability estimate is correct, but on position sizing, entry and exit timing, and price movements during the holding period. Brier Score eliminates these confounds: the agent submits a probability estimate once, and the score depends solely on that estimate and the binary outcome. This provides a clean signal of forecasting quality, separable from trading quality. In the context of benchmarking AI models, this distinction is critical: the goal is to measure what models know about the world, not how well they navigate market microstructure. Corollary 1 makes this precise: α combines resolution and reliability, both of which are intrinsic properties of a probabilistic forecaster, neither of which is recoverable from a PnL time series alone. 3.8 Round Lifecycle Figure 3 illustrates the complete round lifecycle. CreateRoundCommitPhaseOracleWindowRevealPhaseTriggerOutcomesScoringdeadlineanyoneon-chainagents commitbenchmarksrecordedagents reveal Figure 3: Round lifecycle in Foresight Arena. After the commit deadline, Polymarket mid-prices are stored as benchmark prices by the RoundManager. Agents reveal during the reveal window; anyone may call triggerOutcomes() once the reveal deadline passes, reading binary resolutions from the Gnosis CTF. Scores are computed on-chain and accumulated in the ERC-8004 Reputation Registry. 4 System Architecture Foresight Arena is composed of five components: smart contracts, external oracles, a gasless relayer, agent implementations, and a frontend dashboard. Figure 4 provides a high-level overview. External Oracles Gnosis CTF (outcomes) Polymarket CLOB API (prices) Smart Contracts PredictionArena RoundManager · ERC-8004 Agent Layer Random · LLM Benchmark Custom Agents Gasless Relayer AWS Lambda · EIP-712 verify tx simulate · /commit /reveal Frontend React + Vite + The Graph Leaderboard · Reasoning Viewer reads outcomeson-chain txsigns & submitsstate via subgraph Figure 4: High-level architecture. Agents sign EIP-712 messages locally at zero cost and submit through the gasless relayer, which pays gas and submits transactions. Smart contracts read outcome resolutions from the Gnosis CTF and record scores on-chain. The frontend queries on-chain state via The Graph subgraph. 4.1 Smart Contracts PredictionArena. The core contract implements the commit-reveal protocol, EIP-712 gasless paths, and scoring logic. It exposes four primary state transitions: commit(), reveal(), triggerOutcomes(), and calculateScoresForPendingReveals(). All arithmetic is performed in basis points to avoid floating-point issues in Solidity. Signature variants commitWithSignature() and revealWithSignature() accept EIP-712 typed messages, enabling the relayer to submit transactions on behalf of agents. Per-agent nonces prevent replay attacks. The contract verifies signatures on-chain, attributing all actions to the signing agent regardless of who submits the transaction. RoundManager. Manages round lifecycle: creation, deadline enforcement, and benchmark price storage. Crucially, the RoundManager is itself a smart contract—the “curator” role that creates rounds is a contract address, not a human operator, ensuring round creation logic is transparent and fully auditable on-chain. ERC-8004 Registries. Foresight Arena uses the canonical ERC-8004 Identity Registry at 0x8004A169… (same address on all chains) for agent registration. Agents mint an ERC-721 NFT representing their on-chain identity. The ERC-8004 Reputation Registry at 0x8004BAa1… stores aggregated performance achievements published by the curator after campaigns, creating a persistent cross-chain credential. 4.2 Trustless Outcome Resolution via Gnosis CTF Market outcomes are read from the Gnosis Conditional Token Framework (CTF), deployed at 0x4D97…DCd9 on Polygon PoS. For each market condition, the CTF exposes an array of payout numerators together with a scalar payout denominator. A market is considered resolved when the denominator is positive, and its binary outcome is xi=payoutNumerators[1]payoutDenominator∈0,1.x_i\;=\; payoutNumerators[1] payoutDenominator\;∈\;\0,1\. Unresolved markets are excluded from scoring. This design means Foresight Arena inherits the security properties of Polymarket’s own resolution infrastructure. Any market resolved on Polymarket is automatically resolvable in Foresight Arena, with no additional trust required beyond the Gnosis CTF itself. 4.3 Gasless Participation To participate, an agent needs only an Ethereum keypair—no native token (POL) balance required. The workflow is: 1. The agent computes its commitment and signs an EIP-712 typed message (Commit) off-chain at zero cost. 2. The agent POSTs the signed message to the relayer API at api.foresightarena.xyz. 3. The relayer verifies the signature off-chain, simulates the transaction via eth_call, and submits on-chain if simulation succeeds. 4. Gas cost is approximately 0.0030.003 POL per commit and 0.010.01 POL per reveal—funded from the relayer wallet at negligible cost. Rate limits (one commit and one reveal per agent per round) are enforced at the relayer level, providing Sybil resistance without staking. 4.4 Agent Reference Implementations Random Baseline Agent (∼ 250 lines). A minimal reference implementation that registers on-chain, commits uniformly random predictions across all markets in a round, and processes the reveal queue. Designed for crontab scheduling. This baseline establishes the minimum performance floor; by Proposition 1, its expected Brier score is exactly 1/31/3, and it is straightforward to distinguish from any skilled agent at moderate sample sizes. LLM Benchmark Agent (∼ 500 lines). A full-featured agent built on the Vercel AI SDK with OpenRouter backend, supporting any frontier model (Claude, GPT, Gemini, Grok, …) via a single environment variable. Key design choices: • Tool calling. The model receives four structured tools: getMarketDetails(i) for Polymarket metadata, getPriceHistory(i) for recent CLOB trajectory, searchWeb(q) for Tavily news retrieval, and submitPredictions(…) as the final-output sentinel. Markets are referenced by integer index, preventing prompt-injection from raw condition IDs. • Two-phase scheduling. A cheap discovery phase scans for new rounds and processes the reveal queue. The expensive prediction phase fires only when a round is within LEAD_TIME_SECONDS (default: 600 s) of its commit deadline, maximizing information freshness. • Reasoning storage. When configured, the agent posts its full reasoning trace to the relayer’s /reasoning endpoint via EIP-712 signature. Traces are stored in S3 and publicly retrievable, enabling post-hoc analysis. 5 Methodology and Planned Evaluation 5.1 Market Selection Each round’s market set is curated from Polymarket using two criteria: (1) highest trading volume in the preceding 24 hours, and (2) trending activity (sustained price movement and community interest). This strategy ensures that markets are liquid—reducing benchmark price noise—and topically current, testing models’ ability to integrate recent information rather than recall historical facts. All selected markets are binary (YES/NO) with clearly defined resolution criteria. The exact market set for each round is stored on-chain in RoundManager and is identical for all competing agents. 5.2 Planned Agent Roster The benchmark is designed to evaluate frontier LLM agents through a single shared scaffolding—identical system prompt, tool configuration, and scheduling policy—with the underlying language model as the only variable. The simulation study of Section 6 models five frontier LLM archetypes plus a Random baseline; in the live evaluation, the corresponding agents will be the actual model deployments shown in Table 2. Table 2: Agents to be evaluated in Foresight Arena. The simulation in Section 6 uses these identifiers as labels for distinct agent archetypes; the live evaluation will replace the archetypes with the corresponding model deployments via OpenRouter. Agent identifier Provider Type Tool suite claude-opus-4-5 Anthropic LLM Benchmark Market data, web search gpt-5-2 OpenAI LLM Benchmark Market data, web search gemini-3-pro Google DeepMind LLM Benchmark Market data, web search grok-4-1 xAI LLM Benchmark Market data, web search glm-4-7 Zhipu AI LLM Benchmark Market data, web search Random Baseline — Control None 5.3 Evaluation Timeline and Protocol The benchmark is sized for three successive campaigns of approximately 1717 rounds each, for a total of R=50R=50 rounds and ≈350≈ 350 resolved binary predictions per agent. Table 3 summarises the planned parameters; the simulation in Section 6 uses these same parameters to demonstrate what the analysis pipeline produces under the expected operating conditions. Table 3: Round parameters used both for the simulation in Section 6 and as targets for the live evaluation. Parameter Value Total rounds (planned) 5050 (three campaigns of ≈17≈ 17) Markets per round ≈7≈ 7 (binary, high-volume Polymarket) Total predictions ≈350≈ 350 per agent Chain Polygon PoS (chainId: 137) Outcome oracle Gnosis CTF (0x4D97…DCd9) Benchmark prices Polymarket CLOB mid-price at commit deadline 5.4 Baselines We include two baselines to contextualize agent performance. Random baseline. The random agent draws each prediction uniformly at random on [0,1][0,1]. By Proposition 1 and a short calculation, its expected Brier score is 1/31/3, and its expected Alpha against a well-calibrated market is strongly negative. Market consensus baseline. An agent that echoes market mid-prices (pi=bip_i=b_i) achieves α=0α=0 by construction; by Proposition 2, it also achieves Var(δi)=0Var( _i)=0 per market, deterministically producing α^=0 α=0. This is the “beat-the-market” threshold: any agent with α¯>0 α>0 across many rounds demonstrates genuine edge over the collective intelligence of Polymarket participants. 6 Illustrative Simulation Study Note on the nature of these results. The numerical results in this section are produced by a Monte Carlo simulation, not by a live on-chain evaluation. The simulation is deterministic (numpy default RNG, seed 137) and calibrated to the Brier-score ranges reported in the published LLM-forecasting literature (Zou et al., 2024; Schoenegger et al., 2024; Halawi et al., 2024). Its purpose is to illustrate the behaviour of the analytical framework of Section 3—in particular how the Murphy decomposition distinguishes calibration failures from market-tracking failures, and how Proposition 3 maps effect sizes to required sample sizes. The model names appearing in the tables (claude-opus-4-5, gpt-5-2, gemini-3-pro, grok-4-1, glm-4-7) are placeholders for the agents that will participate in the forthcoming live evaluation; the simulation does not constitute a measurement of any specific model’s forecasting ability. Live results from the deployed benchmark will be reported in a future revision of this manuscript. The simulation parameters follow the design of Section 5: R=50R=50 rounds of k=7k=7 markets each, with seven simulated agents (five LLM archetypes plus Market Consensus and Random baselines). All numbers are reported to four decimal places; sampling standard errors and t-statistics are computed from R=50R=50 per-round observations under the independence assumption of Section 3.6. The full simulation pipeline is documented in Appendix C and released open-source at https://github.com/foresight-arena/analysis. 6.1 Literature Anchoring Before presenting the simulated leaderboard, Table 4 surveys published Brier-score values on comparable real-world forecasting tasks, providing reference points for the absolute score magnitudes used to calibrate the simulation. Table 4: Reported Brier scores on real-world binary forecasting tasks from the published literature. Lower is better; a random predictor yields ≈1/3≈ 1/3. Source / agent Setting Brier Random predictor (theoretical) uniform on [0,1][0,1] ≈0.333≈ 0.333 General public (Zou et al., 2024) ForecastBench 0.1210.121 Top LLM, no crowd access (Zou et al., 2024) ForecastBench 0.1360.136 Top LLM, with crowd forecast (Zou et al., 2024) ForecastBench 0.1220.122 LLM ensemble of 12 (Schoenegger et al., 2024) Metaculus, 31 Qs indist. from humans Human superforecasters (Zou et al., 2024) ForecastBench 0.0960.096 Three observations guide expectations for Foresight Arena. First, on ForecastBench, top frontier LLMs reach Brier scores comparable to the general public (∼0.12 0.12) but above superforecasters (∼0.10 0.10)—the market consensus on Polymarket, reflecting a similar class of aggregated judgment, falls in a comparable range. Second, since high-volume Polymarket markets aggregate sophisticated traders into a common-knowledge focal point (Nechepurenko, 2026), the market baseline is typically near-calibrated (RELbase≈0REL_base≈ 0), so by Corollary 1, α≈RESagent−RESbase−RELagentα _agent-RES_base-REL_agent: Alpha over such a market is intrinsically small, on the order of 0.010.01–0.030.03. Third, this is precisely the regime in which our 5050-round evaluation provides marginal statistical power (Table 1). 6.2 Simulated Leaderboard Table 5 reports cumulative Brier and Alpha scores from the simulation across all 50 rounds. Mean market-consensus Brier for the simulated period was 0.19950.1995 (SD 0.0710.071). Table 5: Simulated leaderboard after R=50R=50 rounds (n≈350n≈ 350 predictions per agent). Values after ± are standard errors. t is the one-sample t-statistic against H0:α=0H_0:α=0; “Beat%” is the fraction of rounds with per-round αr>0 _r>0. Numbers come from the calibrated simulation, not from a live deployment. Agent Brier ↓ α¯ α ↑ t ( p ) Beat% claude-opus-4-5 0.1945±0.00990.1945± 0.0099 +0.0049±0.0036+0.0049± 0.0036 1.37(0.18)1.37\;(0.18) 56%56\% gpt-5-2 0.1954±0.00980.1954± 0.0098 +0.0041±0.0037+0.0041± 0.0037 1.10(0.28)1.10\;(0.28) 52%52\% gemini-3-pro 0.1960±0.00930.1960± 0.0093 +0.0034±0.0035+0.0034± 0.0035 0.99(0.33)0.99\;(0.33) 52%52\% grok-4-1 0.2010±0.01020.2010± 0.0102 −0.0015±0.0017-0.0015± 0.0017 −0.86(0.40)-0.86\;(0.40) 46%46\% glm-4-7 0.2072±0.01070.2072± 0.0107 −0.0077±0.0038-0.0077± 0.0038 −2.03(0.05)-2.03\;(0.05) 34%34\% Market Consensus 0.1995±0.01020.1995± 0.0102 0.00000.0000 — — Random Baseline 0.3384±0.01750.3384± 0.0175 −0.1390±0.0192-0.1390± 0.0192 −7.24(≪.001)-7.24\;( .001) 16%16\% The simulation produces three patterns that the analytical framework predicts. Gross differences are overwhelmingly significant. The Random baseline registers α¯=−0.139 α=-0.139, approximately 4040 standard errors below zero; any informed agent is trivially distinguishable from Random at p<10−10p<10^-10. This sets a sanity floor: at R=50R=50 rounds the framework can comfortably reject chance-level forecasting. Frontier-LLM-like archetypes cluster near market consensus. The three agents in the simulation that condition on the underlying true probability q rather than on the market price all achieve positive cumulative α¯ α in the range [+0.003,+0.005][+0.003,+0.005], but none reaches conventional statistical significance (p<0.05p<0.05). This is exactly the prediction of Proposition 3: a sample size of n≈350n≈ 350 is sized to reliably detect effects of |α∗|≥0.02|α^*|≥ 0.02, not the finer differences modelled here. A multi-year evaluation horizon would be required to confidently rank agents in this regime. Market-tracking archetypes systematically underperform. The two agents that condition on the market mid-price with added noise both achieve negative simulated Alpha. The simulated α¯=−0.0077 α=-0.0077 for glm-4-7 reaches marginal significance (p=0.05p=0.05). This failure mode—tracking the market with added noise—is visible in the proper-scoring-rule framework but would be invisible to a PnL-based evaluator, since both market-tracking archetypes Brier-score within 5%5\% of the market. 6.3 Simulated Cumulative Alpha Trajectories Figure 5 plots the cumulative Alpha score α^R=(1/R)∑r=1Rαr α_R=(1/R) _r=1^R _r as a function of the round index R for each LLM agent. The trajectories display three qualitative phenomena predicted by the theory: (i) concentration of α^R α_R around its long-run mean as R grows; (i) the R−1/2R^-1/2 shrinkage of fluctuations; and (i) clean separation of performance tiers emerging by R≈30R≈ 30. Round Rα^R α_R−0.03-0.03−0.02-0.02−0.01-0.010+0.01+0.01+0.02+0.02+0.03+0.0301020304050claude-opus-4-5+0.0049\,+0.0049gpt-5-2+0.0041\,+0.0041gemini-3-pro+0.0034\,+0.0034grok-4-1−0.0015\,-0.0015glm-4-7−0.0077\,-0.0077final α¯50 α_50 at rightRandom Baseline: α^R≈−0.14 α_R≈-0.14 throughout (off-scale) Figure 5: Cumulative Alpha Score trajectories over the 50-round evaluation for the five LLM agents. claude-opus-4-5, gpt-5-2, and gemini-3-pro drift to slightly positive final values; grok-4-1 and glm-4-7 drift below zero, consistent with their market-tracking behaviour. The Random Baseline sits at α^R≈−0.14 α_R≈-0.14 throughout and is off-scale. Rapid early volatility reflects the R−1/2R^-1/2 shrinkage of sampling error predicted by Proposition 2. 6.4 Murphy Decomposition of the Simulated Scores Applying Theorem 1 to the simulated ≈350≈ 350 predictions from each agent (binned into K=10K=10 probability deciles) yields the reliability (REL) and resolution (RES) components summarised in Table 6 and visualized in Figure 2(b). The simulated uncertainty term UNC=o¯(1−o¯)=0.250UNC= o(1- o)=0.250 (within 0.0020.002 of the theoretical maximum) is common to all agents and reflects the roughly balanced YES/NO base rate built into the simulation’s market-generation process. Table 6: Murphy decomposition of the simulated Brier Scores. UNC=o¯(1−o¯)UNC= o(1- o) is fixed at 0.2500.250. Lower RELREL is better calibration; higher RESRES is better discriminative power; by Theorem 1, ℬ=UNC+REL−RESB=UNC+REL-RES. Agent UNCUNC RELREL ↓ RESRES ↑ Brier ↓ claude-opus-4-5 0.2500.250 0.01070.0107 0.06530.0653 0.19450.1945 gpt-5-2 0.2500.250 0.01230.0123 0.06350.0635 0.19540.1954 gemini-3-pro 0.2500.250 0.00510.0051 0.05870.0587 0.19600.1960 grok-4-1 0.2500.250 0.00860.0086 0.05850.0585 0.20100.2010 glm-4-7 0.2500.250 0.00630.0063 0.04770.0477 0.20720.2072 Market Consensus 0.2500.250 0.00790.0079 0.05810.0581 0.19950.1995 Random Baseline 0.2500.250 0.09590.0959 0.00910.0091 0.33840.3384 Corollary 1 predicts that positive Alpha emerges from either a resolution gain or a calibration advantage over the market. The simulated decomposition in Table 6 illustrates this prediction concretely: • Two of the q-anchored archetypes achieve their simulated edge primarily through resolution gain (RES≈0.064RES≈ 0.064–0.0650.065 vs. RESbase=0.058RES_base=0.058)—they sort outcomes more sharply than the market at the cost of slightly higher calibration error. • One q-anchored archetype achieves a small positive simulated α through superior calibration instead (REL=0.005REL=0.005, the lowest among all simulated agents including the market) with resolution matching the market. • The two b-anchored (market-tracking) archetypes have both moderate RELREL and substantially lower RESRES than the market, consistent with the construction of the simulation: echoing the mid-price with added noise cannot generate positive α. • The Random agent’s signature is unmistakable: RELREL an order of magnitude above any other archetype, RESRES near zero. The decomposition demonstrates, on simulated data with known ground-truth structure, that market-tracking with noise produces a qualitatively different signature from miscalibration—a distinction the analytical framework predicts and that any future live evaluation will be equipped to measure. 6.5 Simulated Per-Category Performance The simulation cycles markets through six broad categories matching those used by Polymarket curation. Table 7 shows the simulated per-category breakdown, with the market-consensus Brier, the best-performing simulated agent per category, and the fraction of LLM-archetype agents achieving α¯>0 α>0 within the category. The simulated category labels are illustrative only; in production, category metadata will come from the on-chain RoundManager. Table 7: Simulated per-category breakdown across the 50-round study. “Best agent” is the LLM archetype with highest α¯ α within the category; #α¯>0\#\ α>0\ reports how many of the five simulated LLM archetypes achieved positive category-restricted Alpha. Numbers come from the simulation, not from a live deployment. Category Rounds Mkt. Brier Best agent Best α¯ α #α¯>0\#\ α\!>\!0\ Crypto 1212 0.1980.198 claude-opus-4-5 +0.0148+0.0148 4/54/5 Politics 1111 0.2210.221 claude-opus-4-5 +0.0112+0.0112 3/53/5 Sports 9\ 9 0.1670.167 gpt-5-2 +0.0083+0.0083 3/53/5 Economics 7\ 7 0.2070.207 gemini-3-pro +0.0060+0.0060 2/52/5 Geopolitics 6\ 6 0.2340.234 gpt-5-2 −0.0019-0.0019 0/50/5 Entertainment 5\ 5 0.1920.192 claude-opus-4-5 +0.0041+0.0041 2/52/5 Total / weighted 50 0.2020.202 The simulation produces three patterns worth noting as a demonstration of what the framework can detect, contingent on the modelling assumptions. Heterogeneity across categories. The simulation embeds different effective signal-to-noise ratios across categories (modelling, e.g., the higher rate of price-moving news in crypto versus the slow, ambiguous nature of geopolitical questions (Halawi et al., 2024)). The analysis pipeline correctly recovers this heterogeneity from the simulated outcomes, validating that per-category disaggregation is a viable analysis tool when live data become available. Per-category sample sizes are small. With n≤84n≤ 84 predictions per category, even within the simulation, intra-LLM differences are not individually significant. Live category-level analysis will require accumulating well past 50 rounds before any specific claim about which agent dominates which domain can be tested rigorously. The pattern is consistent with the analytical framework. The simulation cleanly produces the expected qualitative behaviour: agents that condition on the underlying truth q outperform agents that condition on the noisy market price b, with the gap visible in resolution but small enough that detecting it on real data will require many hundreds of rounds. The full per-round dataset (data_simulated.csv) is included in the reproducibility package. 6.6 Statistical Power and What R=50R=50 Can and Cannot Resolve The simulated outcomes are internally consistent with the power analysis of Section 3.6. Against the Random baseline, where |α∗|≈0.14|α^*|≈ 0.14, R=50R=50 rounds deliver overwhelming statistical significance (t≈−7.2t≈-7.2, p≪10−10p 10^-10). Against the market consensus, where the simulated effect sizes fall in the |α∗|∈[0.003,0.008]|α^*|∈[0.003,0.008] regime, Proposition 3 requires n∈[2 200, 15 500]n∈[2\,200,\,15\,500] predictions for 80%80\% power at κ=0.05κ=0.05—four to forty times the current sample. Any future ranking of frontier LLMs at this fine resolution will therefore demand a multi-year evaluation horizon. Two design features of Foresight Arena are well-aligned with this constraint. First, the ERC-8004 Reputation Registry accumulates scores across rounds without bound, so statistical power grows monotonically as the benchmark runs. Second, the gasless, permissionless design enables arbitrary developers to contribute new agents, increasing both the sample size per agent and the diversity of the evaluation pool. A realistic timeline of 200200 rounds—at the current cadence, achievable within a single year—would bring |α∗|=0.01|α^*|=0.01 into reliable detection range, sufficient to produce a definitive ordering of frontier LLMs by forecasting skill once live data are collected. 7 Analysis and Discussion 7.1 Proper Scoring vs. PnL: What Each Metric Captures The Prediction Arena benchmark (Zhang et al., 2026) found that all six frontier models lost money on Kalshi (average −22.6%-22.6\%), yet the same models showed substantially smaller losses on Polymarket (average −1.1%-1.1\%), with one model achieving a 71.4%71.4\% settlement win-rate. This platform-dependent divergence illustrates how PnL conflates multiple factors: market selection, position sizing, timing, and predictive accuracy all contribute. A model with correct predictions but poor timing appears indistinguishable from a model with poor predictions that traded conservatively. Our Murphy-based framework (Corollary 1) makes the alternative precise. Alpha decomposes additively into a resolution gain (informational content of the forecast beyond the base rate) and a reliability gap (calibration advantage over the market). Neither component is recoverable from a PnL time-series. The simulation in Section 6 illustrates the same point operationally: the two market-tracking archetypes achieve worse simulated Alpha than the market-echoing control, despite posting Brier scores within 5%5\% of the market. A PnL-based evaluator would be unable to distinguish this failure mode. Together, Brier Score and Alpha Score constitute a two-dimensional characterization of forecasting quality: absolute calibration and informational edge over the market. This distinction matters for downstream applications: a policymaker selecting an AI system for risk analysis cares about calibration; a trader cares about edge; Foresight Arena measures both independently. 7.2 Statistical Honesty and the Limits of 50 Rounds Our power analysis (Proposition 3) formalizes a methodological constraint that has been largely implicit in the LLM forecasting literature. At R=50R=50 rounds (n≈350n≈ 350 predictions), Foresight Arena can reliably detect effects of |α∗|≳0.02|α^*| 0.02 but not smaller. The gap between frontier LLMs—on the order of 0.0050.005–0.010.01 observed in Section 6—is below this resolution. Three implications follow. First, any ranking of frontier LLMs derived from a short-horizon prediction market benchmark should be treated as provisional. This applies equally to our evaluation and to Prediction Arena’s 57-day study. Second, the cumulative, permissionless design of Foresight Arena is uniquely suited to the long evaluation horizons that principled statistical inference requires. The on-chain ledger enables any observer to recompute statistics at any time; new rounds simply extend the sample. Third, statistical power can be improved orthogonally to sample size by targeting boldness: by Proposition 2, Var(δi)∝(bi−pi)2Var( _i) (b_i-p_i)^2, so the standard error of α α shrinks when agents commit to predictions that differ substantially from the market. A future extension of Foresight Arena might restrict scoring to markets where the agent meaningfully disagrees with the crowd, effectively concentrating evaluation on informative signals. 7.3 On-Chain Verifiability as a Design Principle A central claim of Foresight Arena is that on-chain recording transforms agent performance history from a number to be trusted into a fact to be verified. This has concrete implications for the commercial ecosystem around AI agents. Consider an agent developer seeking to demonstrate the value of their system to a potential buyer. In a centralized benchmark, the buyer must trust that the organizer recorded predictions honestly, applied scoring correctly, and did not selectively report results. In Foresight Arena, the buyer can independently query the smart contract, replay the scoring logic, and verify the entire history—in the same way they would verify any blockchain transaction. Verifiable AI credentials. We propose the notion of an on-chain forecasting credential: a tamper-proof, independently verifiable, cumulative performance record that cannot be selectively reported or retroactively modified. As markets for AI agent services develop—and as regulators increasingly demand auditability of autonomous AI systems—we expect such credentials to become a standard component of AI agent procurement. Foresight Arena provides the infrastructure for generating these credentials in the domain of probabilistic forecasting. 7.4 Agent Design Implications Corollary 1 yields concrete design advice. An agent’s positive Alpha requires a resolution gain and/or a reliability gain over the market. Since the market is already near-calibrated, the reliability channel offers limited upside; systematic Alpha must come from resolution—genuine informational edge. This implies that LLM agents should prioritize: 1. Fresh information. Resolution gains require signal the market has not yet priced in, which decays rapidly. Foresight Arena’s LLM benchmark agent addresses this by firing the expensive prediction step only within 600 s of the commit deadline (Section 4). 2. Controlled boldness. By Proposition 2, predictions close to bib_i carry low variance but low expected signal; predictions far from bib_i carry high signal but are punished severely when wrong. Optimal strategy balances boldness against confidence, favouring large deviations only where the agent has genuinely high confidence. 3. Calibration over confidence. The reliability term penalizes any bin of predictions that is systematically off from the realized frequency. LLMs are known to be overconfident (Halawi et al., 2024); explicit temperature-scaling or ensembling may pay directly. Ensemble methods (Schoenegger et al., 2024) are particularly compelling: averaging forecasts from multiple models typically reduces noise in pip_i, shrinking RELREL while preserving RESRES. Foresight Arena’s permissionless design readily supports ensemble agents, whose on-chain performance becomes directly comparable to individual models. 7.5 Limitations Live evaluation pending. The numerical results in Section 6 come from a calibrated Monte Carlo simulation, not from a live deployment. The simulation models five frontier-LLM archetypes plus a Random baseline using the parameters of Section 5; its purpose is to illustrate the analytical framework and exercise the analysis pipeline end-to-end. Live results from the deployed benchmark on Polygon PoS will be reported in a future revision once a sufficient number of rounds have been resolved. Market selection. The curator (RoundManager contract) selects which Polymarket markets appear in each round. While the selection criteria (volume, trending) are on-chain and transparent, different market universes would yield different rankings. Future work should explore broader or randomly sampled market sets. Benchmark price precision. The baseline probability bib_i is the Polymarket mid-price at commit deadline. In periods of low liquidity, this mid-price may be noisy; we mitigate via volume filtering. Uniform tool configuration. The LLM benchmark agent uses identical tools and prompts for all models, so observed differences conflate intrinsic capability with model-tool interaction. A controlled tool-ablation study is planned. Statistical power at short horizons. As discussed above, R=50R=50 is the minimum at which α∗≥0.02α^*≥ 0.02 can be detected. Results at shorter horizons should be reported with full confidence intervals and interpreted as preliminary. Independence assumption in variance calculation. Proposition 2 and the SE formula (4) assume market outcomes are independent. Thematic clustering (e.g. multiple markets on the same election) could induce correlation; a block-bootstrap or cluster-robust SE would be appropriate for such rounds. 8 Conclusion We introduced Foresight Arena, the first permissionless, on-chain benchmark for evaluating AI forecasting agents on real-world prediction markets. Our contribution is twofold. On the infrastructure side, we combine a commit-reveal protocol with trustless outcome resolution via the Gnosis CTF, producing an evaluation environment that inherits the integrity properties of the underlying blockchain. Any third party can independently verify any agent’s history at any time. On the methodological side, we develop a formal treatment of the scoring rules: we prove strict propriety of Brier (Proposition 1), derive closed-form expressions for the variance of per-market Alpha (Proposition 2), carry out a concrete power analysis (Proposition 3), and show that Alpha decomposes cleanly into a resolution gain and a reliability gap via Murphy’s classical decomposition (Corollary 1). We illustrate this framework with a deterministic, seed-controlled simulation study calibrated to the Brier-score ranges reported in the published LLM-forecasting literature. The simulation cleanly recovers three patterns predicted by the analytical framework: that the Random baseline is overwhelmingly distinguishable from any informed agent at R=50R=50 rounds; that frontier-LLM-like archetypes cluster within |α¯|≤0.005| α|≤ 0.005 of market consensus, where the analytical power result rules out individual-significance ranking at this sample size; and that agents which track the market with added noise produce a Murphy decomposition signature (low resolution, moderate reliability error) qualitatively distinct from miscalibrated agents—a distinction invisible to PnL-based evaluation. The corresponding live evaluation will be reported in a future revision of this manuscript. Foresight Arena is designed as a living benchmark. As rounds accumulate, statistical power increases monotonically, and any developer can contribute new agents without permission. All infrastructure is released as open-source at https://github.com/foresight-arena/contracts. Future work. Immediate extensions include parallel rounds across market categories, a token-based staking layer, a social copying layer, and ensemble scoring (Schoenegger et al., 2024). A second line of work concerns conditional scoring: restricting the Alpha computation to markets where the agent meaningfully disagrees with the crowd, effectively trading sample size for effect size and addressing the statistical-power bottleneck identified above. Generative AI Disclosure In preparing this manuscript, the authors used Anthropic’s Claude Opus 4.7 for copy-editing and for rendering figures from numerical data. All methodology, analysis, and conclusions are the authors’ own; the authors reviewed and edited all AI-generated content and take full responsibility for the final manuscript. References Berg et al. [2008] Berg, J., Nelson, F., and Rietz, T. (2008). Prediction market accuracy in the long run. International Journal of Forecasting, 24(2):285–300. Brier [1950] Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1):1–3. Dawid [1982] Dawid, A. P. (1982). The well-calibrated Bayesian. Journal of the American Statistical Association, 77(379):605–610. DeGroot and Fienberg [1983] DeGroot, M. H. and Fienberg, S. E. (1983). The comparison and evaluation of forecasters. The Statistician, 32(1/2):12–22. Gneiting and Raftery [2007] Gneiting, T. and Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477):359–378. Halawi et al. [2024] Halawi, D., Zhang, F., Yueh-Han, C., and Steinhardt, J. (2024). Approaching human-level forecasting with language models. arXiv preprint arXiv:2402.18563. Hanson [2007] Hanson, R. (2007). Logarithmic market scoring rules for modular combinatorial information aggregation. The Journal of Prediction Markets, 1(1):3–15. Jimenez et al. [2024] Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. (2024). SWE-bench: Can language models resolve real-world GitHub issues? In International Conference on Learning Representations. Murphy [1973] Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4):595–600. Nechepurenko [2026] Nechepurenko, M. (2026). Price as focal point: Prediction markets, conditional reflexivity, and the politics of common knowledge. arXiv preprint arXiv:2604.24147. Also available at SSRN: https://ssrn.com/abstract=6657119. Schoenegger and Park [2023] Schoenegger, P. and Park, P. S. (2023). Large language model prediction capabilities: Evidence from a real-world forecasting tournament. arXiv preprint arXiv:2310.13014. Schoenegger et al. [2024] Schoenegger, P., Tuminauskaite, I., Park, P. S., and Tetlock, P. E. (2024). Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy. Science Advances, 10(45):eadp1528. Tetlock and Gardner [2015] Tetlock, P. E. and Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown Publishers. Wolfers and Zitzewitz [2004] Wolfers, J. and Zitzewitz, E. (2004). Prediction markets. Journal of Economic Perspectives, 18(2):107–126. Zhang et al. [2026] Zhang, J., Liu, G., Johansson, O., Yitayew, H., Ohly, K., and Li, G. (2026). Prediction Arena: Benchmarking AI models on real-world prediction markets. arXiv preprint arXiv:2604.07355. Zou et al. [2024] Zou, A., Chen, E., Arumugam, K., Li, Y., Deng, J., Zuo, S., and Hendrycks, D. (2024). ForecastBench: A dynamic benchmark of AI forecasting capabilities. arXiv preprint arXiv:2409.19839. Appendix A Proofs and Extended Derivations A.1 Proof of Proposition 2 (closed-form) We restate and prove: For X∼Bernoulli(q)X (q) and fixed b,p∈[0,1]b,p∈[0,1], Var((b−X)2−(p−X)2)=4q(1−q)(b−p)2Var((b-X)^2-(p-X)^2)=4q(1-q)(b-p)^2. Proof. Expand: δ δ =(b−X)2−(p−X)2 =(b-X)^2-(p-X)^2 =(b2−2bX+X2)−(p2−2pX+X2) =(b^2-2bX+X^2)-(p^2-2pX+X^2) =(b2−p2)−2X(b−p). =(b^2-p^2)-2X(b-p). Write δ=c−2(b−p)Xδ=c-2(b-p)X where c=b2−p2c=b^2-p^2 is non-random. Then Var(δ)=[2(b−p)]2⋅Var(X)=4(b−p)2q(1−q)Var(δ)=[2(b-p)]^2·Var(X)=4(b-p)^2q(1-q). ∎ A.2 Sample-size derivation for Proposition 3 Assume independence across n markets with qi=q_i=q and |bi−pi|=δ|b_i-p_i|=δ for all i (homogeneity). Then SE(α^)=2δq(1−q)/nSE( α)=2δ q(1-q)/n. The one-sided z-test against H0:α=0H_0:α=0 with true Alpha α∗α^* has power π=Φ(α∗SE(α^)−z1−κ).π= \! ( α^*SE( α)-z_1-κ ). Setting π=0.80π=0.80 and solving: α∗SE(α^)≥z0.95+z0.80=1.645+0.842=2.487. α^*SE( α)≥ z_0.95+z_0.80=1.645+0.842=2.487. Rearranging: n≥(2.487)2⋅4q(1−q)δ2(α∗)2=24.73⋅q(1−q)δ2(α∗)2.n≥ (2.487)^2· 4q(1-q)δ^2(α^*)^2= 24.73· q(1-q)δ^2(α^*)^2. For q=0.5q=0.5 and δ=0.15δ=0.15, this gives n≥24.73⋅0.25⋅0.0225/(α∗)2=0.1391/(α∗)2n≥ 24.73· 0.25· 0.0225/(α^*)^2=0.1391/(α^*)^2, reproducing equation (6). A.3 Proof sketch of Theorem 1 (Murphy decomposition) Partition the N forecasts into K bins. Within bin k, all forecasts share mean p¯k p_k (exact for predictions at the bin mean; small binning error otherwise). Write (pi−xi)2=((pi−p¯k)+(p¯k−o¯k)+(o¯k−xi))2.(p_i-x_i)^2=((p_i- p_k)+( p_k- o_k)+( o_k-x_i))^2. Summing within bin k, the cross-terms involving (o¯k−xi)( o_k-x_i) vanish (since ∑i∈k(o¯k−xi)=0 _i∈ k( o_k-x_i)=0 by definition of o¯k o_k), and the cross-term (pi−p¯k)(p¯k−o¯k)(p_i- p_k)( p_k- o_k) sums to zero by definition of p¯k p_k. One obtains ∑i∈k(pi−xi)2=∑i∈k(pi−p¯k)2+nk(p¯k−o¯k)2+∑i∈k(o¯k−xi)2. _i∈ k(p_i-x_i)^2\;=\; _i∈ k(p_i- p_k)^2\;+\;n_k( p_k- o_k)^2\;+\; _i∈ k( o_k-x_i)^2. Summing over bins and dividing by N, the first term vanishes in the fine-binning limit. The second term is RELREL by definition. The third term, by the within-bin variance identity ∑i∈k(o¯k−xi)2=nko¯k(1−o¯k) _i∈ k( o_k-x_i)^2=n_k o_k(1- o_k), plus the between-bin variance identity, decomposes into No¯(1−o¯)−∑knk(o¯k−o¯)2N o(1- o)- _kn_k( o_k- o)^2, which is N⋅UNC−N⋅RESN·UNC-N·RES. Dividing by N yields (1). Appendix B Smart Contract Specifications B.1 Deployed Addresses Table B.1: Contract deployments. Contract Network Address MockConditionalTokens Polygon Amoy (testnet) 0x4aF09f4A542c… RoundManager Polygon Amoy (testnet) 0x4e44fbAD7a1D… PredictionArena Polygon Amoy (testnet) 0x21993729… PredictionArena Polygon Mainnet 0x7aB0d11F… RoundManager Polygon Mainnet 0x9cE1c4Bf… Identity Registry All chains 0x8004A169… Reputation Registry All chains 0x8004BAa1… Gnosis CTF Polygon Mainnet 0x4D97DCd9… B.2 Commit Hash Format Commitments are computed off-chain as: c=keccak256(abi.encodePacked(uint256⏟roundId,uint16[]⏟p1,…,pk,bytes32⏟salt))c= keccak256\! ( abi.encodePacked ( uint256_roundId,\; uint16[]_p_1,…,p_k,\; bytes32_salt ) ) Each uint16 is packed as 2 bytes (little-endian, no padding). The contract recomputes this hash during reveal. B.3 EIP-712 Domain and Types ⬇ Domain: name: "PredictionArena", version: "1", chainId: 137, verifyingContract: <PredictionArena> Commit: roundId: uint256, commitHash: bytes32, agent: address, nonce: uint256, deadline: uint256 Reveal: roundId: uint256, predictionsHash: bytes32, salt: bytes32, agent: address, nonce: uint256, deadline: uint256 Appendix C Data Analysis Protocol and Reproducibility C.1 Per-Round Data Pipeline For each round r=1,…,50r=1,…,50, the evaluation pipeline proceeds as follows: 1. The RoundManager emits the RoundCreated event with the market set ℳrM_r. Our indexer (subgraph) records the block number, timestamp, and condition IDs. 2. At commit deadline, Polymarket mid-prices r=(br,1,…,br,7)b_r=(b_r,1,…,b_r,7) are written to RoundManager via the curator transaction and indexed. 3. Each agent’s Commit and Reveal transactions are captured, with the revealed prediction vector a,rp_a,r verified against the commit hash on-chain. 4. At resolution, triggerOutcomes reads rx_r from the Gnosis CTF; unresolved markets are excluded from scoring. 5. Per-agent per-round Brier ℬa,r=|ℳr∗|−1∑i∈ℳr∗(pa,r,i−xr,i)2B_a,r=|M_r^*|^-1 _i _r^*(p_a,r,i-x_r,i)^2 and Alpha αa,r=ℬrbase−ℬa,r _a,r=B^base_r-B_a,r are computed on-chain in basis points, then rescaled to [0,1][0,1] in post-processing. All reported statistics (α¯ α, SE, t, Murphy components) are computed from the vector of per-round scores αa,rr=150\ _a,r\_r=1^50 for each agent a. C.2 Murphy Decomposition The Murphy decomposition in Table 6 uses K=10K=10 equally-spaced bins on [0,1][0,1]. For each agent, all 350350 predictions across all rounds are binned by pip_i; within each bin k, we compute p¯k p_k, o¯k o_k, and nkn_k, then evaluate REL=N−1∑knk(p¯k−o¯k)2REL=N^-1 _kn_k( p_k- o_k)^2 and RES=N−1∑knk(o¯k−o¯)2RES=N^-1 _kn_k( o_k- o)^2. Empty bins are excluded. The identity ℬ=UNC+REL−RESB=UNC+REL-RES holds up to a small binning residual (bounded by the within-bin variance of pip_i), which explains the ≲0.002 0.002 discrepancy visible in Table 6 between the column-derived ℬB and the direct Brier computation. C.3 Reproducibility All analysis code is released as a self-contained Python package at https://github.com/foresight-arena/analysis, comprising: scoring.py (Brier, Alpha, and Murphy decomposition implementations, plus the power-analysis helper of Proposition 3), data.py (a unified loader that reads either from the production subgraph or from the deterministic simulation generator used during development), pipeline.py (the aggregations producing Tables 5, 6, and 7), plots.py (matplotlib reproductions of Figures 2(b) and 5 for diagnostic verification), and analysis.ipynb (a Jupyter notebook tying everything together). Any third party can reproduce our results end-to-end from on-chain data: the only required inputs are (i) the production subgraph endpoint, (i) the deployed contract addresses in Table B.1, and (i) numpy, pandas, scipy, matplotlib, and requests. A CSV snapshot mechanism (save_csv, load_csv) further allows offline replay of any historical campaign without depending on subgraph availability. Appendix D Agent Implementation Details D.1 LLM Benchmark Agent: Tool Definitions Table D.1: Tools available to the LLM benchmark agent. Tool Description getMarketDetails(i) Full Polymarket metadata: question, description, end date, current YES price, volume, liquidity, tags. getPriceHistory(i) Recent CLOB YES-price time series (sampled, last 7 days). searchWeb(q) Tavily web search for current news and context (optional; requires API key). submitPredictions(…) Sentinel tool capturing the model’s final predictions; always called last. D.2 Approximate API Cost per Round Table D.2: Estimated LLM API cost per round (≈7≈ 7 markets, web search enabled). On-chain gas adds ≈$0.001≈ 0.001–$0.004 per round per agent at typical Polygon gas prices. Model Cost (USD) Notes Claude Opus 4 $0.10–$0.30 Strongest multi-step reasoning GPT-5 $0.10–$0.25 Gemini 2.5 Pro $0.02–$0.05 Most cost-efficient frontier model Grok 4 $0.05–$0.15