Paper deep dive
Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting
Vivienne Ming
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/5/2026, 1:15:56 PM
Summary
This pilot study investigates how human-AI collaboration affects forecasting accuracy using a real-money prediction market (Polymarket) as a benchmark. The research identifies a trimodal distribution of hybrid performance: 'Automators' (match the model), 'Validators' (perform worse than the model/human alone), and 'Cyborgs' (exceed both the model and the market). Crucially, the study finds that while raw cognitive ability (g) predicts solo human performance, it does not predict hybrid success. Instead, 'collaborative human capital'âspecifically perspective-taking, intellectual humility, and curiosityâis the primary predictor of reaching the high-performing 'Cyborg' mode.
Entities (10)
Relation Signals (5)
Cyborg â characterizedby â Perspective-taking
confidence 90% · Cyborgs exceeded other hybrid forecasters on intellectual humility, curiosity, and perspective-taking.
Cyborg â outperforms â Polymarket
confidence 90% · Cyborg forecasters matched or exceeded (i.e., achieved lower error than) the market benchmark of 3.5
Perspective-taking â predicts â Accuracy
confidence 90% · perspective-taking predicted lower error (r = -0.32, p = .04)
Intellectual Humility â predicts â Accuracy
confidence 90% · curiosity (r = -0.26) and intellectual humility (r = -0.25) trending [as predictors of lower error]
Automator â performsworsethan â Llama 3.1
confidence 85% · Automators reproduced roughly their toolâs behavior but not its accuracy... worse than the AI-only baseline
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an objective, externally resolved benchmark, this pilot shows that the value of human-AI collaboration depends on a specific, measurable form of human capital. Analyzed at the level of the individual forecaster, hybrid performance is trimodal: most people either deferred to the model (matching it) or used it to rubber-stamp a prior guess (performing worse than the model alone), while a minority engaged in genuine complementary reasoning and reached accuracy matching or even exceeding (i.e., lower error than) the market itself. Collaborative traits (perspective-taking, intellectual humility, and curiosity) rather than raw cognitive ability or model benchmarks, distinguished who reached that mode. The results are preliminary but statistically robust, and motivate a pre-registered replication now in preparation.
Tags
Links
- Source: https://arxiv.org/abs/2607.02467v1
- Canonical: https://arxiv.org/abs/2607.02467v1
Trouble viewing inline? Open PDF directly â
Full Text
14,076 characters extracted from source content.
Expand or collapse full text
Human Capital, Not Model Benchmarks, Predicts Hybrid Intelligence in Forecasting Vivienne Ming The Human Trust; Possibility Science; UCL Global Business School for Health Preprint â pilot study. Findings are preliminary and intended to motivate a pre-registered replication. Significance. Whether pairing people with AI helps or hurts is usually reported as a single average effect. Using a real-money prediction market (Polymarket) as an objective, externally resolved benchmark, this pilot shows that the value of humanâAI collaboration depends on a specific, measurable form of human capital. Analyzed at the level of the individual forecaster, hybrid performance is trimodal: most people either deferred to the model (matching it) or used it to rubber-stamp a prior guess (performing worse than the model alone), while a minority engaged in genuine complementary reasoning and reached accuracy matching or even exceeding (i.e., lower error than) the market itself. Collaborative traitsâperspective-taking, intellectual humility, and curiosityârather than raw cognitive ability or model benchmarks, distinguished who reached that mode. The results are preliminary but statistically robust, and motivate a pre-registered replication now in preparation. Introduction Studies of humanâAI collaboration report conflicting effects: some find augmentation of lower-skilled workers (1, 2), others find the benefits accrue mainly to experienced users (3), still others find that access to capable models degrades human performance relative to the model working alone (4, 5), and others show substantial returns on complex, elite tasks (6). A recurring confound is evaluation. In some cases AI systems may simply have encoded the problem and its solution, so that measured âabilityâ reflects contamination rather than reasoning (7, 8). And even on less constrained tasks, âqualityâ is often scored by expert raters who are themselves swayed by the fluency of AI-assisted prose rather than the substance of the work (9, 10). Forecasting against a real-money prediction market sidesteps both problems: every prediction is scored against an externally resolved ground truth on a common, style-free metric. I report a pilot designed to estimate (i) whether human capital moderates the value of AI assistance in forecasting, and (i) whether distinct interaction styles mediate that moderation. Because each participant entered their own probabilistic forecast, all analyses are conducted at the level of the individual forecaster. I frame the work as hypothesis-generating: several cells are small, and the aim is to obtain effect-size estimatesânow statistically supportedâto power a confirmatory trial in preparation. Methods Participants and design. One hundred eight adults were recruited by flyer in Berkeley, CA and compensated $20 per session for three sessions. The main study comprised 78 participants (42 UC Berkeley students; 36 community adults; mean age 28.7, range 18â60) organized into 26 three-person teams and assigned either to a Human-only condition (12 teams) or a Hybrid condition (14 teams) with access to one of four large language models. A separate pool of 30 participants (10 teams) worked with a âSocraticâ model that withholds answers, reported separately (Supplementary S1). Participants collaborated in teams but each entered an individual forecast; analyses are at the participant level. One participant was excluded for an out-of-range Brier score (â0.005), leaving 77 analyzed in the main pool. Forecasting task and ground truth. The 30 questions were live Polymarket contracts resolving between November 2025 and January 2026, spanning economics, international relations, and business. Each questionâs externally certified resolution (o â 0,1) provided ground truth, and each team forecast 10 questions drawn at random from the pool. The four models also forecast all 30 questions independently (AI-only baseline): Llama 3.1 (8B), Qwen3 (8B), GPT-4o, and Gemini 3 Pro. As an external reference I recorded the marketâs implied probability (scaled Brier 3.5). Outcome metric. Forecast accuracy is the Brier score, the mean squared error of the probabilistic forecast, B = (1/N) ÎŁ (pᔹ â oᔹ)ÂČ, reported scaled Ă100 (range 0â100; lower is better; an uninformative 0.5 forecast scores 25). Human-capital measures. Session 1 administered Ravenâs Advanced Progressive Matrices Set I (fluid reasoning), the ICAR public-domain g proxy, the IRI Perspective-Taking subscale, the Comprehensive Intellectual Humility Scale, an epistemic-curiosity scale, the Brief Resilience Scale, and the TIPI Big-Five inventory. Interaction styles. Hybrid participants were classified post hoc from their process into three emergent styles: Automators (adopt the modelâs answer with little added reasoning), Validators (use the model to check a prior guess), and Cyborgs (iterative complementary reasoning). Because style is emergent rather than assigned, it is collinear with human capital; comparisons across styles are descriptive. Analysis. I report participant-level condition means with 95% confidence intervals and individual points (Fig. 1). Given heterogeneous within-condition variances I use a non-parametric omnibus test (KruskalâWallis) with Welch's t-tests for post-hoc pairwise comparisons, and Pearson correlations of human capital with accuracy. Effect sizes are Cohenâs d. Results Figure 1. Participant-level performance and the human capital of Cyborgs. (A) Scaled Brier score (Ă100; lower is better) by condition; bars are participant means, dots individual forecasters, whiskers 95% CIs. Dashed line, Polymarket benchmark (3.5); dotted line, mean AI-only baseline (5.5). (B) Cyborgs versus other hybrid forecasters on four human-capital traits (group means, minâmax normalized; raw means printed). The Socratic pool is reported separately (Supplementary S1); one participant was excluded for an out-of-range score (see Limitations). A trimodal âK-shape.â Forecasting error differed sharply across conditions (KruskalâWallis ÏÂČ across the four groups, p < 10â»ÂčÂł; Fig. 1A). Human-only forecasters scored 14.7 (95% CI 14.4â15.0, n = 35), well above chance but far behind the models. Crucially, hybridization did not uniformly help. Automators reproduced roughly their toolâs behavior but not its accuracy (10.4, CI 9.7â11.0, n = 24)âbetter than humans alone (p < .001) yet worse than the AI-only baseline (5.5). Validatorsâthe classic âhuman-in-the-loopâ who uses AI to check a prior guessâwere significantly worse than humans working alone (31.7, CI 20.6â42.7, n = 9; vs. human-only p = .017). Only Cyborgs improved on the machine: 3.8 (CI 3.5â4.1, n = 9), lower error than every individual AI model and statistically indistinguishable fromâindeed numerically below the mean ofâthe strongest models, significantly better than Automators (p < .001). Relation to the market. Cyborg forecasters matched or exceeded (i.e., achieved lower error than) the market benchmark of 3.5, whereas Automators and Validators did not approach it. The comparison of Cyborgs to the pooled AI-only baseline was in the same direction but only marginal (p = .06), reflecting the four-model baseline rather than a weak effect; Cyborg error was below each of the four models individually. Human capital predicts hybrid accuracyâbenchmarks do not. Among the four models, AI-only accuracy tracked scale and benchmark standing (GPT-4o and Gemini 3 Pro, 4.6; Qwen3, 5.5; Llama 3.1, 7.1). Yet among hybrid forecasters, cognitive abilityâthe human analog of a benchmark scoreâdid not predict accuracy (g, r = â0.10; fluid reasoning, r = â0.08; both ns). Collaborative human capital did: perspective-taking predicted lower error (r = â0.32, p = .04), with curiosity (r = â0.26) and intellectual humility (r = â0.25) trending. The contrast with solo human forecasting is striking: among Human-only participants, every trait strongly predicted accuracyâintellectual humility (r = â0.66, p < .001), and g, fluid reasoning, curiosity, and perspective-taking all near r = â0.45 (p < .01). Access to AI thus decouples raw ability from outcomeâthe model supplies a floorâwhile collaborative capacity instead governs who transcends that floor. Who becomes a Cyborg. Human capital sharply gated emergence of the Cyborg mode. Cyborgs exceeded other hybrid forecasters on intellectual humility (6.5 vs. 4.4, d = 4.1), curiosity (6.7 vs. 5.1, d = 4.0), and perspective-taking (6.5 vs. 4.6, d = 3.2), with smaller but reliable differences in fluid reasoning (d = 1.5) and g (d = 1.7) (all p < .01; Fig. 1B). The separation was strong enough that a logistic model of Cyborg emergence on these traits was not identifiable at this sample sizeâevidence of a very strong association rather than a null. Qualitative complementarity. Process notes suggest a division of labour in Cyborg pairs: humans drove exploration of the ill-posed, contingent parts of a question (what could intervene, what the market might be missing) while the model anchored the well-posed, data-bound parts. Automators outsourced both; Validators used the model to crystallize, rather than challenge, their own reasoning. Discussion These results reframe the question of AI and productivity from whether AI helps to when and how it helps. Pooling all hybrid interaction styles (Automators, Validators, and Cyborgs) into a single average âhybridâ cell hides a strongly trimodal outcomeâincluding a condition (Validator) that underperforms the AI alone and even the unaided human. The moderator that organizes the pattern is collaborative human capital, especially perspective-taking, curiosity, and intellectual humility, consistent with broader findings that heterogeneity in collaboration ability drives the returns to hybrid intelligence (11â14). Notably, the human capital that matters is not the human analog of a model benchmark: raw cognitive ability predicted solo human accuracy and tracked the modelsâ own benchmark ordering, but did not predict who succeeded with AI. What makes a model score well, or a person reason well alone, is not what makes a humanâAI team effective. Limitations. This is a pilot and should be read as such. The Cyborg and Validator cells are small (n = 9 each); interaction style is emergent and collinear with human capital, so style comparisons are descriptive, not causal; and within-condition outcome variance is low, which inflates some standardized effect sizesâinterpretation here rests on the rank-based omnibus test and confidence intervals rather than on d. The market comparison is limited by a four-model baseline and per-question calibration was not independently re-audited. A pre-registered, adequately powered replication (in preparation) will (i) pre-register the style taxonomy and a human-capital composite, (i) manipulate rather than merely measure interaction style where feasible, and (i) fix the forecast-elicitation timestamp relative to market state. Supplementary notes (incidental, not part of the pre-specified design) S1. A âSocraticâ model that never answers. A separate pool of 30 participants (10 teams) worked with the lowest-tier open model re-tuned to withhold direct answers and instead prompt reasoning. It failed every accuracy benchmark, yet shifted behavior in the predicted direction: the Cyborg rate rose to 30% (vs. 21% in the main hybrid pool), and would-be Automators, unable to copy an answer, fell to human-level accuracy (14.9 vs. 10.4 in the main pool). Cyborgs remained accurate (4.2). User satisfaction, however, was uniformly poor (post-task CSAT â 1â2 on a 10-point scale). S2. EEG cognitive engagement. Midway through data collection, a mobile EEG (UC Berkeley Neurotech Collider Lab) recorded relative gamma power as a proxy for cognitive effort in volunteers. Automators showed â43% lower relative gamma than human-only participants (0.57 vs. 1.0). This is an opportunistic, unblinded observation in 7 participants, with gamma-as-engagement a contested operationalization; it is reported only to motivate instrumented follow-up. References 1. E. Brynjolfsson, D. Li, L. Raymond, Generative AI at work. Q. J. Econ. 140, 889â942 (2025). 2. S. Noy, W. Zhang, Experimental evidence on the productivity effects of generative artificial intelligence. Science 381, 187â192 (2023). 3. S. Daniotti, J. Wachs, X. Feng, F. Neffke, Who is using AI to code? Global diffusion and impact of generative AI. Science 391, 831â835 (2026). 4. M. Vaccaro, A. Almaatouq, T. Malone, When combinations of humans and AI are useful: A systematic review and meta-analysis. Nat. Hum. Behav. 8, 2293â2303 (2024). 5. F. A. Csaszar, A. Peterson, D. Wilde, The strategic foresight of LLMs: Evidence from a fully prospective venture tournament. arXiv [econ.GN] (2026). 6. N. Zöller, et al., Human-AI collectives most accurately diagnose clinical vignettes. Proc. Natl. Acad. Sci. U. S. A. 122, e2426153122 (2025). 7. S. Kapoor, P. Henderson, A. Narayanan, Promises and pitfalls of artificial intelligence for legal applications. arXiv [cs.CY] (2024). 8. A. M. Bean, et al., Measuring what matters: Construct validity in large language model benchmarks. arXiv [cs.CL] (2025). 9. F. DellâAcqua, et al., The cybernetic teammate: A field experiment on generative AI and teamwork. Organ. Sci. (2026). https://doi.org/10.1287/orsc.2025.20702. 10. F. DellâAcqua, et al., Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality. SSRN Electron. J. (2023). https://doi.org/10.2139/ssrn.4573321. 11. How Claude Code is used in practice. Available at: https://w.anthropic.com/research/claude-code-expertise [Accessed 29 June 2026]. 12. J. H. Shen, A. Tamkin, How AI impacts skill formation. arXiv [cs.CY] (2026). 13. Z. Zhou, et al., Group-AI collaboration enhances creativity performance: The roles of perspective-taking and AI utilisation strategies. J. Comput. Assist. Learn. 42 (2026). 14. V. Ming, Robot-proof: When machines have all the answers, build better people (John Wiley & Sons, 2026).