Paper deep dive
Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals
ZhaoBin Li, Mark Steyvers
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/26/2026, 1:37:19 AM
Summary
This paper investigates how humans mentally recalibrate their trust in AI confidence signals through repeated experience. Using a behavioral experiment (N=200) across four calibration conditions (standard, overconfidence, underconfidence, and reverse confidence), the authors demonstrate that humans can learn to adapt their reliance strategies. A computational model based on a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule explains these dynamics, showing that humans update baseline trust and confidence sensitivity, though they face significant challenges in overriding inductive biases in counterintuitive 'reverse confidence' scenarios.
Entities (5)
Relation Signals (3)
Computational Model â utilizes â Rescorla-Wagner learning rule
confidence 99% · We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule
Humans â mentallyrecalibrate â AI confidence signals
confidence 95% · We investigate whether humans can learn to mentally recalibrate AI confidence signals through repeated experience.
Reverse Confidence Condition â challenges â Human Inductive Bias
confidence 92% · the reverse confidence scenario, where a substantial proportion of participants struggled to override initial inductive biases.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to mentally recalibrate AI confidence signals through repeated experience. In a behavioral experiment (N = 200), participants predicted the AI's correctness across four AI calibration conditions: standard, overconfidence, underconfidence, and a counterintuitive "reverse confidence" mapping. Results demonstrate robust learning across all conditions, with participants significantly improving their accuracy, discrimination, and calibration alignment over 50 trials. We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule to explain these dynamics. The model reveals that humans adapt by updating their baseline trust and confidence sensitivity, using asymmetric learning rates to prioritize the most informative errors. While humans can compensate for monotonic miscalibration, we identify a significant boundary in the reverse confidence scenario, where a substantial proportion of participants struggled to override initial inductive biases. These findings provide a mechanistic account of how humans adapt their trust in AI confidence signals through experience.
Tags
Links
- Source: https://arxiv.org/abs/2603.22634v1
- Canonical: https://arxiv.org/abs/2603.22634v1
Trouble viewing inline? Open PDF directly â
Full Text
43,999 characters extracted from source content.
Expand or collapse full text
Learning to Trust: How Humans Mentally Recalibrate AI Confidence Signals ZhaoBin Li (zhaobin.li@uci.edu) Department of Cognitive Sciences, University of California, Irvine Mark Steyvers (mark.steyvers@uci.edu) Department of Cognitive Sciences, University of California, Irvine Abstract Productive human-AI collaboration requires appropriate reliance, yet contemporary AI systems are often miscalibrated, exhibiting systematic overconfidence or underconfidence. We investigate whether humans can learn to mentally recalibrate AI confidence signals through repeated experience. In a behavioral experiment (N=200N=200), participants predicted the AIâs correctness across four AI calibration conditions: standard, overconfidence, underconfidence, and a counterintuitive âreverse confidenceâ mapping. Results demonstrate robust learning across all conditions, with participants significantly improving their accuracy, discrimination, and calibration alignment over 50 trials. We present a computational model utilizing a linear-in-log-odds (LLO) transformation and a Rescorla-Wagner learning rule to explain these dynamics. The model reveals that humans adapt by updating their baseline trust and confidence sensitivity, using asymmetric learning rates to prioritize the most informative errors. While humans can compensate for monotonic miscalibration, we identify a significant boundary in the reverse confidence scenario, where a substantial proportion of participants struggled to override initial inductive biases. These findings provide a mechanistic account of how humans adapt their trust in AI confidence signals through experience. Keywords: Human-AI Collaboration; Trust Calibration; Social Metacognition; Reinforcement Learning; Introduction From everyday scenarios like vacation planning to critical settings like medical triage, humans are increasingly relying on artificial intelligence to aid decision-making [undefap, undefah]. The key to productive human-AI collaboration is knowing when to accept or reject the AIâs recommendationsâa phenomenon known as appropriate reliance [undefay, undefaak]. To support this, AI systems often provide confidence scores to help users gauge the reliability of a given suggestion on a case-by-case basis [undefaao]. Consequently, understanding how humans interpret and calibrate their trust in these uncertainty signals is essential to optimizing human-AI collaboration. Prior research suggests humans often use AI confidence scores as direct indicators of reliability, trusting predictions when confidence is high and rejecting them when low [undefaao, undefaag]. Critically, AI systems frequently exhibit systematic miscalibrationâexpressing overconfidence or underconfidence relative to their actual accuracy. Contemporary AI systems, including neural networks and large language models, are particularly prone to overconfidence [undefaq, undefas, undefaam]. This mismatch between reported and actual reliability causes humans to over- or under-rely on AI advice [undefaab], degrading both task performance and possibly trust in AI systems [undefaz]. However, in real-world settings, humans are not static observers but engage with AI repeatedly over time, creating opportunities to improve their collaboration [undefai, undefaak]. This raises an important question about experiential learning: can people learn to mentally calibrate AI confidence scores, adapting their reliance strategies to compensate for systematic biases in uncertainty signals? This question bears directly on both the cognitive science of learning and the practical design of AI systems. From a cognitive perspective, mentally calibrating AI confidence scores poses a nontrivial learning challenge: humans need to aggregate prediction errors across repeated interactions, gradually updating an internal mapping between the AIâs reported confidence and its actual probability of being correct, and then subsequently adjusting their reliance strategies in accordance based on their evolving mental model. This adaptive process naturally suggests the involvement of reinforcement-learning mechanisms, yet existing research has not directly investigated this topic. The practical implications are equally significant: if humans cannot easily compensate for AI miscalibration through experience, then system designers need to prioritize calibrationâpotentially at the expense of other important metrics like accuracy [undefaac] and fairness [undefav], or requiring increased computational and data complexity [undefaam, undefar]. Conversely, if humans can learn to recalibrate AI confidence, then designers may deploy AI systems with greater miscalibration without sacrificing downstream performance, shifting design priorities toward supporting humanâAI collaboration rather than purely AI calibration. We investigate this question using a controlled behavioral experiment and computational modeling. In the experiment, participants learned over repeated interactions to predict whether an AIâs response was correct based solely on its reported confidence. By manipulating the AIâs calibrationâincluding overâconfidence, underâconfidence, and a challenging reverseâconfidence condition in which higher AI confidence signaled lower accuracyâwe show that participants exhibit robust learning across all conditions, steadily improving their accuracy, discrimination, and calibration alignment over trials. To explain the mechanisms underlying this adaptation, we develop a computational model demonstrating that participants mentally recalibrate AI confidence by updating their baseline trust and confidence sensitivity. Together, these results provide a principled account of human adaptation to unreliable information sources. Figure 1: Experimental trial structure: participants viewed a 1-second colored dots animation, and then received the AIâs prediction on which color has the most dots and confidence score rounded to the nearest 10%. Participants then judge whether the AI was correct or wrong, and receive immediate feedback on their accuracy. Experiment Methods To investigate whether humans can learn to mentally calibrate AI confidence signals through experience, we conducted an online behavioral experiment. Participants were tasked with learning to predict the correctness of an AIâs response based solely on its reported confidence level. To simulate a realistic AI-assisted decision-making environment, we employed a cover story where participants performed a visual task and verified the AIâs assessment of it, adapting established paradigms in human-AI collaboration [undefaak, undefaaa]. Participants We recruited 200 participants via Prolific (Mage=43M_age=43, SâDage=13SD_age=13; 55% women, 45% men). The study protocol was approved by our universityâs Institutional Review Board (IRB). The median time to complete the experiment was approximately 9 minutes. Participants received a base compensation of $1.25 and were incentivized with an additional $0.01 bonus for every correct answer, up to a maximum of $0.50. Materials and Procedure The online experiment was implemented using JavaScript and the steps are illustrated in Figure 1. After providing informed consent and completing an onboarding tutorial and practice trial, participants entered the main experimental phase consisting of 50 trials. On each trial, participants viewed a 1-second animation of moving dots and received the AIâs prediction (which color had the most dots) along with its confidence score (0â100%, rounded to nearest 10%). Participants then judged whether the AI was âCorrectâ or âWrong.â In reality, the dot counts were identical, so the only way to succeed was by learning the AIâs calibration pattern. Participants were not told the AIâs underlying accuracy or confidence distribution. Immediate feedback was provided after each judgment, showing whether the participant was correct, the change in their score, and their accumulated bonus. Figure 2: Probability densities of AI confidence distributions across the four conditions for correct and wrong decisions. The probabilities are rounded to the nearest 10% as shown to the experimental interface and normalized across both distributions. Experimental Conditions We utilized a between-subjects design in which participants were randomly assigned to one of four AI calibration conditions: standard confidence, overconfidence, underconfidence, and reverse confidence. In all conditions, the AIâs overall accuracy was set at 50% to ensure that any performance above chance resulted from genuine learning rather than simply always judging the AI as correct or wrong. In all four conditions, the confidence scores were generated using a signal detection model in which confidence values for correct and wrong decisions were drawn from logit-normal distributions (Ï=0.5Ï=0.5) with means separated by 2 units on the logit scale to promote learning (an ideal observer could judge the AIâs correctness with 97% accuracy). The confidence distributions are plotted in FigureË2 and the parameters are listed below: 1. Standard Confidence: Means for correct and wrong decisions were 1 and â1-1, respectively; the optimal decision criterion was 0.5. 2. Overconfidence: Means were shifted to 2 and 0, resulting in a higher optimal decision criterion of 0.75. 3. Underconfidence: Means were shifted to 0 and â2-2, resulting in a lower optimal decision criterion of 0.25. 4. Reverse Confidence: Means were inverted relative to the standard confidence condition (â1-1 for correct, 1 for wrong); the criterion remained at 0.5 but the AI was more likely correct when reporting low confidence. In the experiment, we pre-generated 10,000 confidence scores per condition and randomly sampled from them per trial. Results Accuracy Improved Across Trials As shown in FigureË3, participants demonstrated clear learning across all four calibration conditions. Accuracy increased substantially over the 50 experimental trials, rising from early (trials 1â10) to late (trials 41â50) phases: 62% to 86% in standard confidence, 55% to 87% in overconfidence, 62% to 88% in underconfidence, and 42% to 69% in the reverse confidence condition. Even in the reverse confidence conditionâdesigned to be maximally counterintuitiveâparticipants improved markedly. Multilevel logistic regression predicting accuracy from trial number confirmed significant positive learning slopes across all conditions (p<0.001p<0.001, ÎČ from 0.049 to 0.064). Figure 3: Accuracy improvement across trial blocks for human participants and cognitive model in all four AI calibration conditions. Error bars represent 95% confidence intervals and the dashed line indicates 50% chance accuracy. Figure 4: Hit rate increases and false alarm rate decreases across trial blocks for human participants and cognitive model in all four AI calibration conditions. Error bars represent 95% confidence intervals and the dashed line indicates 50% chance level. Hit and False Alarm Rates Improved Across Trials To better understand the mechanisms underlying these accuracy gains, we examined changes in hit rate (HR)âthe proportion of trials where participants correctly identified the AI as âcorrectââand false alarm rate (FAR)âthe proportion where participants mistakenly identified a wrong AI response as âcorrect.â As shown in FigureË4, HR and FAR improved significantly in all four conditions. In the standard confidence condition, HR increased from 72% to 90% and FAR decreased from 48% to 18%, demonstrating substantially improved discrimination. Participants in the overconfidence condition were initially misled by AI overconfidence, with early FAR at 62%, yet successfully learned to correct for its bias, achieving 92% HR and 19% FAR in late trials. A similar pattern appeared in the underconfidence condition. The most substantial learning occurred in the reverse confidence condition. Early on, participantsâ decisions aligned with the misleading confidence signals, resulting in 51% HR and 68% FAR. Nevertheless, by late trials, participants had partially inverted their decision rule, achieving 73% HR and 35% FAR. Multilevel probit regression confirmed significant increases in HR (ÎČ from 0.018 to 0.027, p<0.001p<0.001) and decreases in FAR (ÎČ from â0.027-0.027 to â0.036-0.036, p<0.001p<0.001) across all conditions. Critically, sensitivity (dâČd ) increased significantly in all conditions (ÎČ from 0.045 to 0.062, p<0.001p<0.001), indicating that participants became more successful at distinguishing correct from wrong AI responses. Human Calibration to AI Improved Across Trials Participants also improved their calibration to the AIâs confidence signals. We use the cognitive model in the next section to visualize the human calibration curves, because the pooled estimate is less noisy. To quantify the improvement in the raw data, we calculated the Expected Calibration Error (ECE) between the AIâs accuracy and perceived AI accuracy per condition and phase [undefaae]: ECE=âm=1M|Bm|nâ|accAIâ(Bm)âacchumanâ(Bm)|ECE= _m=1^M |B_m|n|acc_AI(B_m)-acc_human(B_m)|, where M=11M=11 bins (confidence rounded to the nearest 0.1), |Bm||B_m| is the number of trials in bin m, n is the total number of trials, accAIâ(Bm)acc_AI(B_m) is the AIâs actual accuracy in that bin, and acchumanâ(Bm)acc_human(B_m) is the perceived AI accuracy (the proportion of trials where participants believed the AI was correct). Paired t-tests confirmed significant ECE reductions from early to late trials (p<0.001p<0.001 for all conditions). Taken together, these results demonstrate that participants exhibited robust learning across conditions, not only improving in overall accuracy but also becoming more discriminating and substantially better calibrated to the AIâeven when the AIâs calibration was systematically distorted or inverted. Cognitive Modeling The behavioral results demonstrate clear learning but do not reveal the underlying cognitive mechanisms. Participants need to form beliefs about how AI confidence relates to its correctness, update these beliefs trial-by-trial based on feedback, and do so in ways that vary across individuals. We model participantsâ judgments using a linear-in-log-odds (LLO) transformation with two evolving parameters: an intercept capturing baseline trust in the AI, and a slope capturing confidence sensitivity. Both parameters are updated via a RescorlaâWagner learning rule and estimated using Bayesian multilevel inference, enabling us to explain trial-by-trial dynamics, prior beliefs, and individual differences. Modeling Framework Linear-in-Log-Odds Calibration Consistent with the literature on human probability weighting [undefao, undefaan, undefaal], we assume that participants mentally recalibrate the AIâs reported confidence using a linear-in-log-odds (LLO) transformationâtheir perceived AI accuracy vtv_t on trial t is linear with respect to the AIâs reported confidence ctc_t on a log odds scale: logâĄ(vt1âvt)=bt+wtâ logâĄ(ct1âct) ( v_t1-v_t )=b_t+w_t· ( c_t1-c_t ) This transformation maps confidence to perceived accuracy with an S-shaped curve (examples seen in FigureË5) and includes two interpretable parameters to capture shifts in participant beliefs. First, the intercept btb_t shifts the curve vertically and represents the participantâs baseline propensity to trust the AI. We expect it to increase in the underconfidence condition and decrease in the overconfidence condition if participants learn to compensate for the AIâs systematic bias. Second, the slope wtw_t captures human sensitivity to changes in the AIâs confidence scores, with the sign of wtw_t representing the perceived direction of correlation between the AIâs confidence and accuracy. We expect wtw_t to be positive a priori, reflecting the intuitive belief that higher AI confidence signals higher accuracy, and to become negative in the reverse confidence condition if participants learn that higher AI confidence signals lower accuracy. RescorlaâWagner Learning Rule We assume that both parametersâbaseline trust btb_t and confidence sensitivity wtw_tâare updated trial by trial using a RescorlaâWagner learning rule [undefaah], equivalent to the TD(0) update used in reinforcement learning [undefaaj]. After every trial, participants observe whether the AIâs prediction was correct (gt=1g_t=1) or wrong (gt=0g_t=0) and compute a prediction error ÎŽt=gtâvt _t=g_t-v_t. Learning rates for the two parameters b and w were kept separate and asymmetric for positive and negative errors, reflecting established patterns in human reward learning [undefat, undefan]. As the results will show, the learning rates vary substantially across b and w and by the sign of the errors. Because the sign of the error corresponds to the AIâs accuracy, the learning rates are indexed by gtg_t. To minimize the binary cross-entropy loss, the parameters are updated according to: bt+1 b_t+1 =bt+αb,gtâ ÎŽt =b_t+ _b,g_t· _t wt+1 w_t+1 =wt+αw,gtâ ÎŽtâ logâĄ(ct1âct) =w_t+ _w,g_t· _t· ( c_t1-c_t ) In total, six parameters are estimated per participant: the initial bias b0b_0, the initial weight w0w_0, and four learning rates (αb,gt=1 _b,g_t=1, αb,gt=0 _b,g_t=0, αw,gt=1 _w,g_t=1, and αw,gt=0 _w,g_t=0). Bayesian Multilevel Inference Parameters were estimated using multilevel Bayesian modeling to improve precision and capture individual variability while obtaining group-level trends. On every trial t, participant iâs binary response yt,iy_t,i (coded as 1 if the participant judged the AI correct, 0 otherwise) is modeled as a Bernoulli draw with parameter vt,iv_t,i: yt,iâŒBernoulliâ(vt,i).y_t,i (v_t,i). Because participants were randomly assigned to conditions, we applied an experiment-wide prior to their initial biases and weights: b0,iâŒNâ(ÎŒb0,Ïb0)b_0,i N( _b_0, _b_0), ÎŒb0âŒNâ(0,1) _b_0 N(0,1), and Ïb0âŒExponentialâ(1) _b_0 (1). This prior centers the initial bias at 50% perceived accuracy, providing a 95% prior interval of [12%, 88%] to allow for wide variability in baseline trust. An identical prior structure was applied to w0w_0. Furthermore, given that learning rates are likely influenced by the specific AI condition c, we utilized separate condition-wide priors: logâĄ(αb,gt,i,c1âαb,gt,i,c)âŒNâ(Όαb,gt,c,Ïαb,gt,c) ( _b,g_t,i,c1- _b,g_t,i,c ) N( _ _b,g_t,c, _ _b,g_t,c), Όαb,gt,câŒNâ(â1.5,1.5) _ _b,g_t,c N(-1.5,1.5), and Ïαb,gt,câŒExponentialâ(1) _ _b,g_t,c (1). The hyper-mean was set to center the learning rate around â0.18â 0.18. The log-odds transformation ensures that all α values remain within the [0,1][0,1] interval. The model was implemented in Stan [undefaj] using Markov Chain Monte Carlo (MCMC) with 4 chains and 2,000 samples per chain. Convergence was verified via trace plots and by ensuring R^<1.01 R<1.01 and ESS >400>400. Modeling Results Cognitive model predicts behavioral trends The behavioral results demonstrate clear learning, but to understand the underlying mechanisms, we need a computational account. Our cognitive model captures the key behavioral patterns across all four conditions, including improvements in overall accuracy (FigureË3) and changes in hit and false alarm rates (FigureË4). On average, model predictions match human decisions on 75% of trials (mean log-likelihood per trial =â0.38=-0.38, McFaddenâs pseudo-R2=0.45R^2=0.45), substantially exceeding 50% chance performance. This level of agreement indicates that a simple belief-mapping framework with trial-by-trial updating is sufficient to reproduce participantsâ aggregate reliance behavior. Figure 5: Calibration curves showing AI accuracy (orange) and human perceived AI accuracy (blue dashed) across conditions. Left panel shows prior calibration at trial 1 from experiment-wide priors. Subsequent panels show posterior calibration at trial 50 from participant-level posteriors per experimental condition. Black diagonal dashed line shows perfect calibration. Shaded regions show 95% credible intervals. Changes in belief mapping of AI confidence and perceived accuracy over time The behavioral analysis established that calibration improves, but the noise in raw data prevents us from recovering participantsâ initial beliefs or their learned mappings. To address this, we use the cognitive model to infer how participants mentally translate AI confidence into perceived accuracy. To estimate a priori calibration, we sampled 1,000 values of the initial parameters (b0b_0 and w0w_0) from the experiment-wide priors and computed perceived AI accuracy across all confidence values. FigureË5 (left panel) shows these inferred prior calibration curves. As expected, participants begin with positive confidence sensitivity (M=0.69M=0.69, SâD=1.8SD=1.8), believing that higher AI confidence signals higher accuracy. They also exhibit positive baseline trust (M=0.60M=0.60, SâD=0.56SD=0.56), initially believing that the AI is approximately 65% accurate when it reports 50% confidence. To obtain posterior calibration at the end of learning, we used participant-level posterior means of the parameters at trial 50 (b50b_50 and w50w_50). The resulting curves in FigureË5 (subsequent panels) reveal successful adaptation: learned calibration closely tracks the AI calibration in the standard, overconfidence, and underconfidence conditions. However, the reverse confidence condition shows only partial alignment, reflecting substantial individual differences we explore next. Figure 6: Individual differences in the reverse-confidence condition. Left: Calibration curves at trial 50 showing AI accuracy (dashed green) and human perceived AI accuracy for learners (blue) and non-learners (orange). Right: Learning rates for confidence sensitivity (w) when the AI was correct (αw,gt=1 _w,g_t=1) and wrong (αw,gt=0 _w,g_t=0). Shaded regions and error bars represent 95% credible intervals. Explaining Individual Differences in Reverse Confidence The reverse confidence condition posed a uniquely counterintuitive challengeâparticipants had to learn that higher AI confidence actually signaled lower accuracy. While participants generally improved over time, a detailed analysis of individual accuracy revealed that a substantial proportion struggled to adapt. Categorizing participants by whether their late-stage accuracy (trials 31â50) exceeded 60%, we found that 44% in the reverse confidence condition were non-learnersâa stark contrast to the other three conditions, where non-learners comprised less than 15% of the sample. FigureË6 (left panel) shows the estimated calibration curves at trial 50, separated by learner status. Learners eventually matched the negative slope of the AIâs actual calibration (mean w50=â2.15w_50=-2.15), successfully inverting their initial beliefs. In contrast, non-learnersâ calibration remained largely flat (mean w50=0.01w_50=0.01), retaining a near-zero sensitivity to confidence signals. Critically, both groups began with similar positive priors (w0>0w_0>0), yet only learners managed to override their initial inductive bias. The cognitive model reveals the mechanism behind this divergence. FigureË6 (right panel) shows that learners exhibited dramatically higher learning rates for confidence sensitivity (w) than non-learners. When the AI was correct, learners updated their sensitivity nearly 3Ă faster (αw,gt=1=0.14 _w,g_t=1=0.14 vs. 0.040.04); when the AI was wrong, this multiple expanded to 30Ă (0.500.50 vs. 0.0150.015). These results suggest that learners rapidly incorporated prediction errors to update their confidence sensitivity, while non-learners remained resistant to changing their initial mental calibration. Importantly, learning rates for baseline trust (b) were similar between groups (p>0.05p>0.05), confirming that the model successfully isolates the specific cognitive mechanism responsible for individual differences. Figure 7: Changes in baseline trust btb_t (solid blue) and confidence sensitivity wtw_t (dashed orange) across trials in all four conditions. Shaded regions represent 95% credible intervals. Participants adapt baseline trust and confidence sensitivity To understand how learning evolves, we examined changes in the modelâs two key parameters across trials (FigureË7). Confidence sensitivity (wtw_t) captures how participants weight AI confidence signals. In the standard, overconfidence, and underconfidence conditions, wtw_t increased substantiallyâfrom 0.45â0.65 at trial 1 to 2.7â2.9 at trial 50âdemonstrating that participants became more attuned to AI confidence as a diagnostic signal. In the reverse confidence condition, wtw_t decreased from 0.59 to â1.2-1.2, indicating that participants successfully learned the inverse mapping where higher confidence signaled lower accuracy. Baseline trust (btb_t) represents participantsâ tendency to trust the AI, independent of its reported confidence. This parameter adapted to compensate for systematic AI biases. In the underconfidence condition, baseline trust increased substantially (from 0.570.57 to 2.52.5) as participants learned that the AI was more accurate than it claimed. Conversely, in the overconfidence condition, trust decreased markedly (from 0.620.62 to â1.9-1.9), reflecting learned skepticism toward AI overconfidence. In the standard and reverse conditions, trust decreased slightly (from 0.610.61 to 0.270.27 and 0.60.6 to 0.450.45, respectively) to correct for initial overtrust. All changes were statistically significant (p<0.001p<0.001). Asymmetric learning rates prioritize the most informative errors A key insight from the model is that participants do not update their beliefs uniformlyâinstead, they exhibit asymmetric learning that prioritizes the most diagnostic feedback. These asymmetries vary by condition, reflecting which outcomes are most informative for learning each AIâs calibration pattern. In the overconfidence condition, the most informative signal is when a highly confident AI makes an error. Accordingly, updates for baseline trust are substantially higher when the AI is wrong (αb,gt=0=0.46 _b,g_t=0=0.46) than when it is correct (αb,gt=1=0.29 _b,g_t=1=0.29), indicating participants rapidly lose trust when observing confident mistakes. Conversely, updates for confidence sensitivity are higher when the AI is correct (αw,gt=1=0.55 _w,g_t=1=0.55 vs. αw,gt=0=0.05 _w,g_t=0=0.05)âwhen an extremely confident AI proves correct, participants learn to give significantly more weight to its confidence signals. As expected, these patterns reverse in the underconfidence condition, where the opposite pattern emerges. Baseline trust updates more when the AI is correct (αb,gt=1=0.51 _b,g_t=1=0.51 vs. αb,gt=0=0.14 _b,g_t=0=0.14), while confidence sensitivity updates more strongly when the AI is wrong (αw,gt=0=0.49 _w,g_t=0=0.49 vs. αw,gt=1=0.04 _w,g_t=1=0.04). The standard and reverse confidence conditions show more symmetric learning rates, likely because both correct and incorrect outcomes provide comparable information about the AIâs calibration. Together, these results demonstrate that human learning strategically prioritizes feedback that best reveals the AIâs underlying calibration structure. Discussion Our results demonstrate that humans can learn to mentally recalibrate AI confidence signals through repeated experience. Across four calibration conditions, participants substantially improved their accuracy, discrimination, and calibration alignment over only 50 trials. Our computational model reveals that this adaptation occurs through trial-by-trial updates to baseline trust and confidence sensitivity, with asymmetric learning rates that prioritize the most informative errors. While prior research has examined how people calibrate their own confidence in human-AI contexts [undefak, undefaab] or recalibrate human advisers [undefaaf, undefaai], we demonstrate that humans adaptively recalibrate confidence signals generated by AI systems. This extends established frameworks of metacognition [undefam, undefaw] to human-AI collaboration and provides a computational account of how appropriate reliance develops through experience [undefay, undefaad]. The success of our reinforcement-learning model suggests that this calibration process engages error-based learning mechanisms [undefal, undefaaj], consistent with recent findings that humans learn their own confidence through prediction errors [undefax]. Our model reveals that baseline trust and confidence sensitivity are key components of human mental models of AI calibration [undefai, undefau]. By decomposing trust into the two interpretable parameters, we provide a clear way to characterize how users miscalibrate AI confidence signals. This decomposition helps identify whether errors arise from overall over- or under-trust versus misinterpretation of confidence levels. There are several limitations worth noting about our empirical approach. First, participants completed only 50 trials in a single session, and longer-term dynamics including memory consolidation, decay, or transfer remain unexplored. Second, immediate feedback was provided, but learning would likely be more challenging with delayed or absent ground truth. Third, participants in the task did not have the opportunity to make independent predictions. They could only learn from the AIâs confidence and feedback, but real-world collaboration often allows some degree of independent verification. Future research should address these limitations. Conclusion We demonstrate that humans possess substantial capacity to adapt to AI confidence signals through reinforcement learning mechanisms. By showing that humans can compensate for monotonic miscalibration but struggle with inverse mappings, we identify both the potential and boundaries of human adaptability in collaborative AI systems. These findings contribute to theories of metacognition and social learning in human-AI collaboration while providing insights for designing systems that support appropriate reliance. References [undef] Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar and Tobias Salz âCombining human expertise with artificial intelligence: Experimental evidence from radiologyâ Cambridge, MA: National Bureau of Economic Research, 2023 [undefa] Gagan Bansal et al. âBeyond accuracy: The role of mental models in human-AI team performanceâ In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 7 Association for the Advancement of Artificial Intelligence (AAAI), 2019, p. 2â11 [undefb] Bob Carpenter et al. âStan: A probabilistic programming languageâ In J. Stat. Softw. 76.1 Foundation for Open Access Statistic, 2017, p. 1â32 [undefc] Leah Chong et al. âHuman confidence in artificial intelligence and in themselves: The evolution and impact of confidence on adoption of AI adviceâ In Comput. Human Behav. 127.107018 Elsevier BV, 2022, p. 107018 [undefd] Nathaniel D Daw and Kenji Doya âThe computational neurobiology of learning and rewardâ In Curr. Opin. Neurobiol. 16.2 Elsevier BV, 2006, p. 199â204 [undefe] Stephen M Fleming, Raymond J Dolan and Christopher D Frith âMetacognition: computation, biology and functionâ In Philos. Trans. R. Soc. Lond. B Biol. Sci. 367.1594 The Royal Society, 2012, p. 1280â1286 [undeff] Samuel J Gershman âDo learning rates adapt to the distribution of rewards?â In Psychon. Bull. Rev. 22.5 Springer ScienceBusiness Media LLC, 2015, p. 1320â1327 [undefg] R Gonzalez and G Wu âOn the shape of the probability weighting functionâ In Cogn. Psychol. 38.1 Elsevier BV, 1999, p. 129â166 [undefh] Ben Green and Yiling Chen âThe principles and limits of algorithm-in-the-loop decision makingâ In Proc. ACM Hum. Comput. Interact. 3.CSCW Association for Computing Machinery (ACM), 2019, p. 1â24 [undefi] Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Q Weinberger âOn calibration of modern neural networksâ In ICML abs/1706.04599, 2017 [undefj] Yingxiang Huang et al. âA tutorial on calibration measurements and calibration models for clinical prediction modelsâ In J. Am. Med. Inform. Assoc. 27.4 Oxford University Press (OUP), 2020, p. 621â633 [undefk] Zhengbao Jiang, J Araki, Haibo Ding and Graham Neubig âHow can we know when language models know? On the calibration of language models for question answeringâ In Trans. Assoc. Comput. Linguist. 9, 2020, p. 962â977 [undefl] Kentaro Katahira âThe statistical structures of reinforcement learning with asymmetric value updatesâ In J. Math. Psychol. 87 Elsevier BV, 2018, p. 31â45 [undefm] Markelle Kelly, Aakriti Kumar, Padhraic Smyth and Mark Steyvers âCapturing humansâ mental models of AI: An item response theory approachâ In 2023 ACM Conference on Fairness Accountability and Transparency New York, NY, USA: ACM, 2023, p. 1723â1734 [undefn] Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan âInherent trade-offs in the fair determination of risk scoresâ In arXiv [cs.LG], 2016 [undefo] Asher Koriat âThe self-consistency model of subjective confidenceâ In Psychol. Rev. 119.1 American Psychological Association (APA), 2012, p. 80â113 [undefp] Pierre Le Denmat, Kobe Desender and Tom Verguts âLearning to be confident: How agents learn confidence based on prediction errorsâ In Cognition 266.106332 Elsevier BV, 2026, p. 106332 [undefq] John D Lee and Katrina A See âTrust in automation: designing for appropriate relianceâ In Hum. Factors 46.1 Oxford University Press (OUP), 2004, p. 50â80 [undefr] Jingshu Li et al. âUnderstanding the effects of miscalibrated AI confidence on user trust, reliance, and decision efficacyâ In arXiv [cs.AI], 2025 [undefs] Garston Liang, Jennifer F Sloane, Christopher Donkin and Ben R Newell âAdapting to the algorithm: how accuracy comparisons promote the use of a decision aidâ In Cogn. Res. Princ. Implic. 7.1 Springer ScienceBusiness Media LLC, 2022, p. 14 [undeft] Shuai Ma et al. âare you really sure?â understanding the effects of human self-confidence calibration in AI-assisted decision makingâ In Proceedings of the CHI Conference on Human Factors in Computing Systems 63 New York, NY, USA: ACM, 2024, p. 1â20 [undefu] Matthias Minderer et al. âRevisiting the calibration of modern neural networksâ In arXiv [cs.LG], 2021, p. 15682â15694 [undefv] Bonnie M Muir âTrust in automation: Part I. Theoretical issues in the study of trust and human intervention in automated systemsâ In Ergonomics 37.11 Informa UK Limited, 1994, p. 1905â1922 [undefw] Mahdi Pakdaman Naeini, Gregory Cooper and Milos Hauskrecht âObtaining well calibrated probabilities using Bayesian Binningâ In Proc. Conf. AAAI Artif. Intell. 29.1 Association for the Advancement of Artificial Intelligence (AAAI), 2015 [undefx] Niccolo Pescetelli and Nick Yeung âThe role of decision confidence in advice-taking and trust formationâ In arXiv [cs.SI], 2019 [undefy] Amy Rechkemmer and Ming Yin âWhen confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learning modelsâ In CHI Conference on Human Factors in Computing Systems New York, NY, USA: ACM, 2022, p. 1â14 [undefz] R Rescorla âA theory of Pavlovian conditioning : Variations in the effectiveness of reinforcement and nonreinforcementâ In Classical conditioning, Current research and theory cir.nii.ac.jp, 1972, p. 64â99 [undefaa] O Stanciu and J Fiser âDo humans recalibrate the confidence of advisers or take their confidence at face value?â In CogSci 44.44 escholarship.org, 2022 [undefab] Richard S Sutton and Andrew G Barto âReinforcement Learning: An Introductionâ, Adaptive Computation and Machine Learning series Cambridge, MA: Bradford Books, 2018 [undefac] Heliodoro Tejeda, Aakriti Kumar, Padhraic Smyth and M Steyvers âAI-assisted decision-making: A cognitive modeling approach to infer latent reliance strategiesâ In Comput. Brain Behav. 5.4 Springer ScienceBusiness Media LLC, 2022, p. 491â508 [undefad] Brandon M Turner et al. âForecast aggregation via recalibrationâ In Mach. Learn. 95.3 Springer ScienceBusiness Media LLC, 2014, p. 261â289 [undefae] Cheng Wang âCalibration in deep learning: A survey of the state-of-the-artâ In arXiv [cs.LG], 2023 [undefaf] Hang Zhang and Laurence T Maloney âUbiquitous log odds: a common representation of probability and frequency distortion in perception, action, and cognitionâ In Front. Neurosci. 6 Frontiers Media SA, 2012, p. 1 [undefag] Yunfeng Zhang, Q Vera Liao and Rachel K E Bellamy âEffect of confidence and explanation on accuracy and trust calibration in AI-assisted decision makingâ In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency New York, NY, USA: ACM, 2020 References [undefah] Nikhil Agarwal, Alex Moehring, Pranav Rajpurkar and Tobias Salz âCombining human expertise with artificial intelligence: Experimental evidence from radiologyâ Cambridge, MA: National Bureau of Economic Research, 2023 [undefai] Gagan Bansal et al. âBeyond accuracy: The role of mental models in human-AI team performanceâ In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing 7 Association for the Advancement of Artificial Intelligence (AAAI), 2019, p. 2â11 [undefaj] Bob Carpenter et al. âStan: A probabilistic programming languageâ In J. Stat. Softw. 76.1 Foundation for Open Access Statistic, 2017, p. 1â32 [undefak] Leah Chong et al. âHuman confidence in artificial intelligence and in themselves: The evolution and impact of confidence on adoption of AI adviceâ In Comput. Human Behav. 127.107018 Elsevier BV, 2022, p. 107018 [undefal] Nathaniel D Daw and Kenji Doya âThe computational neurobiology of learning and rewardâ In Curr. Opin. Neurobiol. 16.2 Elsevier BV, 2006, p. 199â204 [undefam] Stephen M Fleming, Raymond J Dolan and Christopher D Frith âMetacognition: computation, biology and functionâ In Philos. Trans. R. Soc. Lond. B Biol. Sci. 367.1594 The Royal Society, 2012, p. 1280â1286 [undefan] Samuel J Gershman âDo learning rates adapt to the distribution of rewards?â In Psychon. Bull. Rev. 22.5 Springer ScienceBusiness Media LLC, 2015, p. 1320â1327 [undefao] R Gonzalez and G Wu âOn the shape of the probability weighting functionâ In Cogn. Psychol. 38.1 Elsevier BV, 1999, p. 129â166 [undefap] Ben Green and Yiling Chen âThe principles and limits of algorithm-in-the-loop decision makingâ In Proc. ACM Hum. Comput. Interact. 3.CSCW Association for Computing Machinery (ACM), 2019, p. 1â24 [undefaq] Chuan Guo, Geoff Pleiss, Yu Sun and Kilian Q Weinberger âOn calibration of modern neural networksâ In ICML abs/1706.04599, 2017 [undefar] Yingxiang Huang et al. âA tutorial on calibration measurements and calibration models for clinical prediction modelsâ In J. Am. Med. Inform. Assoc. 27.4 Oxford University Press (OUP), 2020, p. 621â633 [undefas] Zhengbao Jiang, J Araki, Haibo Ding and Graham Neubig âHow can we know when language models know? On the calibration of language models for question answeringâ In Trans. Assoc. Comput. Linguist. 9, 2020, p. 962â977 [undefat] Kentaro Katahira âThe statistical structures of reinforcement learning with asymmetric value updatesâ In J. Math. Psychol. 87 Elsevier BV, 2018, p. 31â45 [undefau] Markelle Kelly, Aakriti Kumar, Padhraic Smyth and Mark Steyvers âCapturing humansâ mental models of AI: An item response theory approachâ In 2023 ACM Conference on Fairness Accountability and Transparency New York, NY, USA: ACM, 2023, p. 1723â1734 [undefav] Jon Kleinberg, Sendhil Mullainathan and Manish Raghavan âInherent trade-offs in the fair determination of risk scoresâ In arXiv [cs.LG], 2016 [undefaw] Asher Koriat âThe self-consistency model of subjective confidenceâ In Psychol. Rev. 119.1 American Psychological Association (APA), 2012, p. 80â113 [undefax] Pierre Le Denmat, Kobe Desender and Tom Verguts âLearning to be confident: How agents learn confidence based on prediction errorsâ In Cognition 266.106332 Elsevier BV, 2026, p. 106332 [undefay] John D Lee and Katrina A See âTrust in automation: designing for appropriate relianceâ In Hum. Factors 46.1 Oxford University Press (OUP), 2004, p. 50â80 [undefaz] Jingshu Li et al. âUnderstanding the effects of miscalibrated AI confidence on user trust, reliance, and decision efficacyâ In arXiv [cs.AI], 2025 [undefaaa] Garston Liang, Jennifer F Sloane, Christopher Donkin and Ben R Newell âAdapting to the algorithm: how accuracy comparisons promote the use of a decision aidâ In Cogn. Res. Princ. Implic. 7.1 Springer ScienceBusiness Media LLC, 2022, p. 14 [undefaab] Shuai Ma et al. âare you really sure?â understanding the effects of human self-confidence calibration in AI-assisted decision makingâ In Proceedings of the CHI Conference on Human Factors in Computing Systems 63 New York, NY, USA: ACM, 2024, p. 1â20 [undefaac] Matthias Minderer et al. âRevisiting the calibration of modern neural networksâ In arXiv [cs.LG], 2021, p. 15682â15694 [undefaad] Bonnie M Muir âTrust in automation: Part I. Theoretical issues in the study of trust and human intervention in automated systemsâ In Ergonomics 37.11 Informa UK Limited, 1994, p. 1905â1922 [undefaae] Mahdi Pakdaman Naeini, Gregory Cooper and Milos Hauskrecht âObtaining well calibrated probabilities using Bayesian Binningâ In Proc. Conf. AAAI Artif. Intell. 29.1 Association for the Advancement of Artificial Intelligence (AAAI), 2015 [undefaaf] Niccolo Pescetelli and Nick Yeung âThe role of decision confidence in advice-taking and trust formationâ In arXiv [cs.SI], 2019 [undefaag] Amy Rechkemmer and Ming Yin âWhen confidence meets accuracy: Exploring the effects of multiple performance indicators on trust in machine learning modelsâ In CHI Conference on Human Factors in Computing Systems New York, NY, USA: ACM, 2022, p. 1â14 [undefaah] R Rescorla âA theory of Pavlovian conditioning : Variations in the effectiveness of reinforcement and nonreinforcementâ In Classical conditioning, Current research and theory cir.nii.ac.jp, 1972, p. 64â99 [undefaai] O Stanciu and J Fiser âDo humans recalibrate the confidence of advisers or take their confidence at face value?â In CogSci 44.44 escholarship.org, 2022 [undefaaj] Richard S Sutton and Andrew G Barto âReinforcement Learning: An Introductionâ, Adaptive Computation and Machine Learning series Cambridge, MA: Bradford Books, 2018 [undefaak] Heliodoro Tejeda, Aakriti Kumar, Padhraic Smyth and M Steyvers âAI-assisted decision-making: A cognitive modeling approach to infer latent reliance strategiesâ In Comput. Brain Behav. 5.4 Springer ScienceBusiness Media LLC, 2022, p. 491â508 [undefaal] Brandon M Turner et al. âForecast aggregation via recalibrationâ In Mach. Learn. 95.3 Springer ScienceBusiness Media LLC, 2014, p. 261â289 [undefaam] Cheng Wang âCalibration in deep learning: A survey of the state-of-the-artâ In arXiv [cs.LG], 2023 [undefaan] Hang Zhang and Laurence T Maloney âUbiquitous log odds: a common representation of probability and frequency distortion in perception, action, and cognitionâ In Front. Neurosci. 6 Frontiers Media SA, 2012, p. 1 [undefaao] Yunfeng Zhang, Q Vera Liao and Rachel K E Bellamy âEffect of confidence and explanation on accuracy and trust calibration in AI-assisted decision makingâ In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency New York, NY, USA: ACM, 2020