Paper deep dive
Learning under noisy supervision is governed by a feedback-truth gap
Elan Schonfeld, Elias Wisnia
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/21/2026, 1:07:02 AM
Summary
The paper introduces the 'feedback-truth gap,' a universal phenomenon where learners over-commit to noisy feedback when integrated faster than truth can be evaluated. Using a two-timescale model and empirical tests across neural networks, human behavioral tasks, and EEG, the study demonstrates that while the gap is inevitable under timescale mismatch, its consequences depend on system regulation: dense networks accumulate it as memorization, sparse architectures suppress it, and humans actively recover from transient over-commitment.
Entities (8)
Relation Signals (6)
Feedback-Truth Gap â causedby â Two-Timescale Model
confidence 95% · The gap is inevitable when feedback is integrated on a faster timescale than the learnerâs evaluation of the task structure.
Human Probabilistic Reversal Learning â exhibits â Over-Commitment
confidence 92% · humans generated transient over-commitment that was actively recovered.
Dense Neural Networks â exhibits â Memorization
confidence 90% · In dense neural networks, the gap grew progressively and became a permanent feature of the networkâs behavior â memorization.
Sparse-Residual Architectures â suppresses â Feedback-Truth Gap
confidence 88% · sparse-residual architectures and noise-robust training, the gap was reduced or suppressed
EEG Decoding â measures â Feedback-Truth Gap
confidence 85% · the difference between post-feedback neural decoding of the feedback value and the participantâs pre-feedback expectation
Memorization â negativelypredicts â Generalization Accuracy
confidence 80% · Higher memorization resulted in lower generalization; the gap was detrimental.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:When feedback is absorbed faster than task structure can be evaluated, the learner will favor feedback over truth. A two-timescale model shows this feedback-truth gap is inevitable whenever the two rates differ and vanishes only when they match. We test this prediction across neural networks trained with noisy labels (30 datasets, 2,700 runs), human probabilistic reversal learning (N = 292), and human reward/punishment learning with concurrent EEG (N = 25). In each system, truth is defined operationally: held-out labels, the objectively correct option, or the participant's pre-feedback expectation - the only non-circular reference decodable from post-feedback EEG. The gap appeared universally but was regulated differently: dense networks accumulated it as memorization; sparse-residual scaffolding suppressed it; humans generated transient over-commitment that was actively recovered. Neural over-commitment (~0.04-0.10) was amplified tenfold into behavioral commitment (d = 3.3-3.9). The gap is a fundamental constraint on learning under noisy supervision; its consequences depend on the regulation each system employs.
Tags
Links
- Source: https://arxiv.org/abs/2602.16829v1
- Canonical: https://arxiv.org/abs/2602.16829v1
Trouble viewing inline? Open PDF directly â
Full Text
85,717 characters extracted from source content.
Expand or collapse full text
Learning under noisy supervision is governed by a feedbackâtruth gap Elan Schonfeld 1,* Elias Wisnia 2 1 Department of Biology, Columbia University, New York, NY, USA 2 Department of Chemistry, Columbia University, New York, NY, USA * Corresponding author: elan.schonfeld@columbia.edu Abstract When feedback is absorbed faster than task structure can be evaluated, the learner will favor feedback over truth. A two-timescale model shows this feedbackâtruth gap is inevitable whenever the two rates dier and vanishes only when they match. We test this prediction across neural networks trained with noisy labels (30 datasets, 2,700 runs), human probabilistic reversal learning (N= 292), and human reward/punishment learning with concurrent EEG (N= 25). In each system, truth is dened operationally: held-out labels, the objectively correct option, or the participantâs pre-feedback expectation â the only non-circular reference decodable from post-feedback EEG. The gap appeared universally but was regulated dierently: dense networks accumulated it as memorization; sparse-residual scaolding suppressed it; humans generated transient over-commitment that was actively recovered. Neural over-commitment (âŒ0.04â0.10) was amplied tenfold into behavioral commitment (d= 3.3â3.9). The gap is a fundamental constraint on learning under noisy supervision; its consequences depend on the regulation each system employs. Introduction In many domains, learners do not receive truth as a direct instructional cue. Rather, learners update their internal representations of the world based upon feedback that only imperfectly reects the truth. In machine learning, noisy labels can lead to memorization and poor generalizability 1â4 . In human decision-making, probabilistic outcomes can blur the distinction between contingency and chance, requiring the learner to act on noisy feedback 5;6 . While these settings seem quite dierent, both represent examples of noisy supervision. Several groups have argued for deeper integration between AI and neuroscience 7â10 ; however, we have lacked a common quantitative framework for understanding how noisy supervision aects learning processes in articial and biological systems. Here, we describe such a framework â the feedbackâtruth gap â the extent to which a learnerâs internal representation of the environment tracks feedback more than it tracks truth. We operationally dene this term in three systems: for a neural network, the training-minus-validation accuracy dierence as a function of epoch; for a human reversal task, the distance between feedback-aligned and truth-aligned responses to contingency reversals; and for an EEG experiment, the dierence between post-feedback neural decoding of the feedback value and the participantâs pre-feedback expectation (see Methods). Critically, in the EEG system the truth reference is the participantâs prior expectation of correctness â an operational proxy, not the latent task state â because no non-circular reconstruction of objective environmental truth was decodable from the post-feedback signal (Extended Data Fig. 10). While the denitions vary across systems, they share the same 1 arXiv:2602.16829v1 [cs.LG] 18 Feb 2026 formal properties: the learner revises its internal state based upon feedback before it becomes possible to resolve the discrepancy between the learnerâs representation and the true state of the environment. Memorization dynamics in neural networks have been studied extensively 11;12 , and so have perseverative errors after contingency reversals in animals and humans 13â16 . Computational models of human learning have focused on asymmetric learning rates for reward and punishment 17â19 . However, none of these previous studies have provided a necessary condition for when a learner will inevitably become over-committed to feedback, nor have they linked post-reversal over-commitment in humans to the generalization gap in machine learning. Using a minimal two-timescale learner, we establish that a feedbackâtruth gap is inevitable when feedback is integrated on a faster timescale than the learnerâs evaluation of the task structure, and that the gap disappears only when the two timescales are equal (Methods; Fig. 2). We then empirically test this prediction across three independent systems â neural networks trained with label noise (30 datasets, 2,700 runs), human probabilistic reversal learning (N= 292), and human reward/punishment learning with concurrent EEG (N= 25) â using the same gap metrics (AUG pos , T â ) across all three systems. We found that the gap appeared in all three systems but was regulated in very dierent ways (Fig. 1). In dense neural networks, the gap grew progressively and became a permanent feature of the networkâs behavior â memorization. In sparse-residual architectures and noise-robust training, the gap was reduced or suppressed (Fig. 4). In human reversal learning, the gap increased immediately following contingency reversals and was subsequently decreased through active behavioral regulation; the relationship between the size of the gap and the recovery time varied among individuals (Fig. 3). In EEG, post-feedback activity represented feedback valence and prior expectations, and individual dierences in the neural gap predicted the amount of behavioral commitment each individual exhibited (Fig. 5). The same measurable quantity that indexed damage to learning in unregulated neural networks had the opposite implications for performance in regulated human learning. Results The feedbackâtruth gap appears across learning systems We quantied the feedbackâtruth gap in three independent systems: articial neural networks, human probabilistic reversal learners, and human reward/punishment learners with concurrent EEG (Fig. 1). The gap represents the same formal property of the learnerâs behavior â feedback-aligned minus truth-aligned performance â but was measured using each systemâs own notions of feedback and truth. The gap is therefore a common measurement construct, not a claim of common mechanism. In neural networks trained on tabular data with 40% symmetric label noise, training accuracy rose faster than validation accuracy, resulting in a persistent positive gap that grew throughout training (Fig. 1a). This gap resulted from progressive memorization of the noisy labels; the network tracked feedback â including noise â at the cost of generalizing to new, unseen data. No mechanism in either the training process or the network architecture prevented the growth of the gap. Across ve benchmark datasets (Ionosphere, Glass, Sonar, WDBC, Vehicle) and six architectures, including label smoothing, L 2 regularization, and residual variants as controls, the gap was observed in greater than 90% of model-noise combinations at 40% label noise. The gap increased with the label noise rate and disappeared completely when the label noise rate was 0% (Extended Data Figs. 1 and 2). The unregulated dense neural network thus provides the reference prediction of the two-timescale model, while the control conditions verify that the gap persists under conventional remediation and is not specic to any particular architecture. In human probabilistic reversal learning (N= 292, ages 8â30; ref. 11), participants chose between 2 two options with 75%/25% reward probabilities that reversed periodically (empirical noise rate: 17.1%). We calculated the same gap metric â feedback-aligned minus truth-aligned choice accuracy â on a rolling trial-by-trial basis, synchronized with the contingency reversals. Following each reversal, feedback-aligned accuracy recovered more quickly than truth-aligned accuracy, resulting in a transient positive gap that peaked within approximately ve trials and resolved over 15â20 trials (Fig. 1b). This signal was extremely reliable (AUG pos = 0.049±0.032,t(291) = 26.4,d= 1.55,p <10 â75 ) and nearly universal: all 292 subjects demonstrated positive over-commitment following a reversal, and 99.3% eventually regained stable truth-aligned responding (meanT â = 9.9trials). Unlike in networks, the human gap resolves: active behavioral regulation closes it. In concurrent EEG recordings during reward and punishment learning (N= 25; ref. 12), we decoded post-feedback feedback valence and each subjectâs pre-feedback expectation from the post- feedback EEG (Methods). Only one of four candidate truth references â the participantâs prior expectation of correctness â was reliably decodable in this window; objective reconstructions of the task state based on Bayesian and reinforcement-learning models were not decodable (Extended Data Fig. 10). The expectation-truth decoder performed above chance in both tasks (REW: AUROC = 0 .541,t(23) = 2.43,p= 0.012; PUN: AUROC= 0.545,t(24) = 3.84,p <10 â3 ). The primary endpoint is not individual decoder accuracy but the trial-by-trial gap between the two decoders â how much more condently the EEG signal represents feedback valence than expected accuracy. The neural gap was positive in all subjects included in the analysis (reward:N= 24; punishment: N= 25; Extended Data Fig. 3), and feedback-dominant signals in the EEG diminished over trials while expectation-related signals followed feedback (Fig. 1c), consistent with progressive adaptation of the neural response. Fig. 1 thus illustrates a hierarchy of regulation. Over-commitment to feedback occurred in all three systems, but the progression of the gap through time diered. Dense networks drifted further from truth; human choices returned to truth; and EEG signals adapted at an intermediate timescale. Two questions follow: why does the gap occur at all, and what determines whether the system regulates it? The gap is inevitable once timescales dier We analyzed the minimal two-timescale learner to determine why the gap occurs so uniformly in our experiments. The model uses a fast feedback-integrating channel (rate α fast ) and a slower truth-integrating channel (rateα slow ) (Fig. 2; Methods). Noise in the feedback causes any change in the underlying state to expose the dierence in timescales: the fast channel moves before the slow channel, producing a temporary separation of feedback tracking and truth tracking. Using a closed-form solution (Supplementary Derivation), the gap after a state change of sizeâ att= 0is: gap(t) = â· ( e âα slow ·t âe âα fast ·t ) which is positive for allt >0whenever the two rates dier (Fig. 2d). The peak gap magnitude depends on the timescale ratior=α fast /α slow : it approachesâasrgrows large and shrinks to zero asrapproaches 1. If the two timescales match (r= 1), the gap is identically zero regardless of the noise level. In simulation, atr= 1with 20% noise, the mean peak gap stayed below10 â5 (Fig. 2d); across 40 timescale ratios the correspondence between theory and simulation gaveR 2 >0.99. Based on these ndings, we created a phase diagram of the two-timescale learner with respect to timescale ratio and noise level (Extended Data Fig. 2c). The only no-gap boundary in timescale space is when the two channels operate at the same rate (r= 1); simulations conrm the predicted gap sizes throughout the entire parameter range. All three empirical systems can be located on this diagram: dense networks fall in the high-ratio, high-noise region (râ10â20), explaining their 3 large persistent gaps; humans occupy a moderate-ratio regime (râ3â5) where regulation can close the gap; sparse-residual networks have reduced eective ratios and correspondingly smaller gaps. All systems examined empirically demonstrated that removing label noise collapsed the gap and eliminated the advantage of sparse bottlenecks (Extended Data Fig. 5). This result describes what must occur based on the two-timescale structure; it does not describe how any individual system implements learning. Once the feedback is noisy and is integrated faster than truth, the only way to prevent transient over-commitment is to delay, suppress, or compensate the gap through downstream regulation. Because the gap is deterministic given the timescale ratio, it is predictable from early data: the rst 15 epochs of network training suce to classify the eventual memorization regime (epochs 35+) with AUC= 0.96, and early-to-late prediction also works in both human paradigms (Extended Data Fig. 4). The signicance of the gap depends upon regulation Although the same gap construct appeared in every system examined, the relationship of the gap to performance diered qualitatively among the systems (Fig. 3). The signicance of the gap is determined entirely by whether and how the system regulates it. In neural networks under synthetic label noise, the cumulative gap magnitude (AUG norm ) was a strong negative predictor of validation accuracy across 150 datasetâmodel pairs (ÎČ=â0.641, p <10 â6 , xed-eects regression; Fig. 3a). Higher memorization resulted in lower generalization; the gap was detrimental. Under natural noise (crowd-labelled CIFAR-10N), this relationship weakened as model rankings compressed, and onset timingT â became the informative metric (r= 0.672, p= 0.048; Extended Data Fig. 2), demonstrating regime dependence in which metrics carry diagnostic information. The relationship was reversed in human reversal learning. Subjects who exhibited larger transient gaps after contingency reversals also exhibited faster subsequent recovery of truth-aligned responding (Fig. 3b). The gap did not indicate damage; instead, it indicated the extent to which the subject engaged with the changed contingency, and a larger initial discrepancy between feedback tracking and truth tracking was associated with more rapid realignment. This does not imply that the gap was benecial; rather, it implies that the downstream eects of the gap depended on whether regulatory mechanisms existed to close it. This reversal (negative gapâperformance relationship in networks, positive gapârecovery relationship in humans) is the central dissociation: the same measurable quantity had opposite implications for performance depending on regulatory capacity. The EEG data identied a neural basis for this dissociation. The neural gap was spatially organized, with feedback-dominant components at frontal electrodes and expectation-dominant components at posterior sites (Extended Data Fig. 3). Instead of a uniform scalar measure, the post-feedback signal separated evaluative and expectation-related components (Fig. 3c), consistent with the established role of medial frontal cortex in feedback monitoring 20;21 and with the theta, P300, and ERN literatures on outcome evaluation 22â24 . The results thus dened three regulatory regimes. Dense networks lacked regulation, and the gap accumulated without restriction, directly indexing damage. Additional noise-robust methods (co-teaching, forward loss correction) suppressed but did not abolish the gap (Extended Data Table 1). Architectural constraints limited the ability to overt the feedback, thereby suppressing the gap. Humans exhibited transient gaps and then recovered through behavioral adjustment. Although the phenomenon occurred in every system, the nature of the regulatory response determined the outcome. 4 Regulation is tunable and dissociates pressure from recovery Two complementary analyses assessed whether the regulation of the gap could be systematically controlled: architectural intervention in neural networks and cross-level integration in human learners (Figs. 4â5). The strength of a sparse-residual architectural bottleneck was varied (αâ 0.1,0.25,0.5,1.0) across ve datasets with 40% label noise (10 seeds per conguration). Both AUG norm andT â increased monotonically withα(SpearmanÏ= 1.0for both metrics; Fig. 4aâb), demonstrating graded causal control over the memorization dynamics. At the dataset level, this monotonic suppression held in all datasets with measurable gaps (AUG norm >0.005;Ï= 1.0in 3/3), whereas datasets at the noise oor demonstrated no systematic trends because there was no gap to suppress (Extended Data Fig. 5). However, test accuracy did not follow this monotonic pattern: it peaked at α= 0.1â0.25and decreased at higher values (Fig. 4c). This non-monotonic dissociation demonstrated a capacityâmemorization trade-o in which congurations that maximized suppression of the gap did not maximize generalization. (For comparisons with label smoothing, weight decay, and loss- correction methods 25;26 , see Methods and Supplementary Information.) Additional validation across 30 OpenML datasets conrmed that sparse bottlenecks improved accuracy on 93% of datasets at 40% noise but provided no reliable advantage at 0% noise (Extended Data Fig. 5; Extended Data Table 2). Suppression is noise-dependent. Sparse-residual architectures served as a causal probe for graded, monotonic control over the gap, not as a proposed method for practical noise robustness. Behavioral over-commitment to feedback was large and consistent in the EEG dataset (N= 25; ref. 12). All 25 subjects in both reward and punishment conditions showed positive over-commitment: rewardt(24) = 16.6,d= 3.32,p <10 â14 ; punishmentt(24) = 19.6,d= 3.92,p <10 â15 (Fig. 5a). Decomposition into win-stay ( âŒ0.80â0.83) and lose-shift (âŒ0.20â0.23) components demonstrated that over-commitment was largely due to high persistence in rewarding actions, not failure to switch after negative outcomes. The concurrent neural gap, as assessed by EEG decoders, was considerably smaller (neuralAUG pos â0.04â0.10; Fig. 5b), demonstrating approximately tenfold amplication from neural representation to behavioral strategy. The neural gap was highly reproducible (split-half correlation between even and odd trials: r= 0.35,p <10 â5 for reward;r= 0.45,p <10 â8 for punishment) and decreased from early to late trials in the punishment condition ( d= 0.82,p <0.001; Fig. 3d), suggesting that the neural gap tracks a learning-related process rather than noise or drift (Extended Data Fig. 3). At the individual level, subjects with larger neural gaps also exhibited stronger behavioral commitment (reward: Ï= 0.43,p= 0.036; punishment:Ï= 0.42,p= 0.042; Fig. 5c). Neither EEG channel signicantly predicts single-trial behavioral choice; rather, both channels appear to provide noisy measures of the same slow-changing, task-dependent learning state. The amplication is concordant across levels and across individuals (see Extended Data Fig. 6 for further individual-dierences analyses, Extended Data Fig. 9 for robustness to the class-balance exclusion threshold, and Extended Data Table 4 for reinforcement-learning parameter associations). The gap is suppressed dierently in each system. Architectural bottlenecks suppress the gap by reducing the capacity of the network. Human brains employ a dynamic process that amplies modest neural over-representation into large behavioral commitments and then corrects them. While the gap exists in all three systems, the systems dier in how they regulate it. Discussion The feedbackâtruth gap is inherently mathematical: when a learning system integrates feedback before it can verify the true state of aairs, over-commitment to the feedback is inevitable. This gap 5 was consistently evident in all of the systems evaluated. More importantly, the critical nding is that systems dier categorically in how they respond. Networks without regulation generate the gap as damage; architectural constraints produce static suppression; and humans generate active recovery, with concordant amplication at the neural and behavioral levels. This taxonomy â unregulated, suppressed, recovered â provides a conceptual framework for evaluating how dierent systems manage a common computational constraint. Scope.We do not contend that humans learn in the same way as articial neural networks. The shared aspect of the research is a framework for measurement â the feedbackâtruth gap â applied to the same formal manipulation (perturbing the agreement between feedback and ground truth at each learning step) across substrates. The gap is an inherent consequence of timescale mismatch, not something we advocate as adaptive or optimal. The sparse-residual architectures serve as causal probes; the human neural evidence is correlational. In the machine-learning studies, the architectural manipulations serve as controlled probes rather than proposed general-purpose noise-resistant methods; topology ablations indicate that the sparse-residual scaolding, not the graph structure itself, suppresses the gap (Extended Data Table 3). Shared framework, not shared mechanism.Convergence across substrates represents shared measurement, not shared mechanism. The same quantitative metrics ( AUG pos ,T â , com- mitment indices) capture gap dynamics in network training curves, human reversal learning, and reward/punishment strategies. The formal structure is identical across domains: noisy supervision, time-varying commitment, and eventual alignment with truth. However, the substrate-specic process that generates and regulates the gap is dierent in each case, and convergent measurement does not imply convergent implementation. Regulation as the dierentiator.What converges is the principle that raw gap pressure is modulated before the gap is expressed behaviorally. In networks, architectural regularization delays and suppresses the gap. In humans, neural over-representation (Ï= 0.43in reward,0.42in punishment; bothp <0.05) is amplied roughly tenfold into behavioral over-commitment. This concordance holds at subject and block timescales; at the single-trial level, both neural and behavioral signals provide noisy estimates of a shared, slow-changing, task-dependent learning state. In networks the regulatory mechanisms are weight decay and capacity constraints. In humans, prefrontal control circuits are the natural candidate, and dopaminergic signaling â which modulates learning from prediction errors 27;28 and dierentially shapes reward versus punishment learning 29â31 â is a plausible substrate for the amplication we observe. EEG dissociation.EEG provided a distinct dissociation: post-feedback features decode feedback valence and the participantâs pre-feedback expectation, but do not reconstruct the latent task state (Extended Data Fig. 10). Post-feedback ERPs, especially the P300, are responsive to surprise (outcome minus expectation), resulting in a residue of prior belief rather than a direct representation of the environment. Thus, the neural feedbackâtruth gap represents the balance between outcome evaluation and expectation. By the time this balance is reached, the prior beliefs are already partially regulated; feedback and expectation representations are nearly equalized, so the much larger behavioral commitment is likely mediated by downstream control rather than simply scaled versions of the same neural signal. Regime dependence.Neither the gap nor its diagnostic value is uniform across conditions. In networks,AUG norm predicts accuracy under synthetic noise but loses discrimination under natural noise, whereT â becomes informative instead. In humans, the neural gap is modest (dâ1.0) but is amplied manyfold at the behavioral level (d= 3.3â3.9), with signicant cross-modal correlations in both tasks (Ï= 0.43,0.42;p <0.05). The phase diagram (Extended Data Fig. 2c) captures this continuous variation: gap magnitude scales with noise rate and timescale ratio and vanishes at specic boundaries. No single metric is universally diagnostic; the informative metric depends on 6 the noise regime. The capacityâmemorization tradeo connects to the double-descent literature 32 , mathematical theories of learning dynamics 33 , and work on architecture and generalization 34 . Episodic memory 35 may also play a role, particularly where long delays separate feedback from truth verication. This study has several limitations. Primarily, the human neural evidence is correlational. An exploratory pharmacological reanalysis is consistent with dopaminergic modulation of gap regulation: in a double-blind crossover (N= 27; 36 ), the D2/D3 agonist cabergoline attened within-session gap regulation relative to placebo (t(26) = 2.31,d= 0.44, one-sidedp= 0.015; Wilcoxonp= 0.016; Extended Data Fig. 7), and cumulative gap level showed a directionally consistent reduction ( d=â0.32, one-sidedp= 0.053). However, this result is preliminary; proper causal tests will require stimulation or lesion studies. A developmental analysis of the reversal-learning dataset (N= 292; Extended Data Fig. 8) indicates a shift from more exibility in children to more persistent use of feedback in adults, but future studies should assess these relationships in more diverse populations. The human samples are Western adults performing laboratory tasks from public repositories; generalizability to naturalistic or developmental settings is untested. The two-timescale model determines when the gap should appear, not how specic systems create or manage it; attention, working memory, and normalization dynamics all modulate the gap in ways that timescale separation alone does not capture. Expanding the framework to reinforcement-learning agents, developmental populations where prefrontal regulation is still maturing 37â39 , and clinical groups with abnormal feedback processing 40â44 will provide evidence for the frameworkâs broader utility. The gap between feedback and truth should arise wherever there is noisy feedback and latent truth: scientic inference, social learning 45 , organizational adaptation, cultural transmission 46 . The structure echoes foundational ideas in associative learning 47;48 . Understanding what regulates the gap â across architectures, development, and pathology â may provide insight into developing both robust machine learning and human adaptive behaviors. Methods Two-timescale learner model The simplest model exhibiting a feedbackâtruth gap tracks a binary environment via two exponential- moving-average channels with update ratesα fast andα slow (α fast > α slow >0). The observed feedback is noisy: on each trial, the observed signal matches the true state with probability1âΔ. After a state reversal of magnitudeâatt= 0, the gap between the fast and slow estimates is gap(t) = â·(e âα slow ·t âe âα fast ·t ), strictly positive for allt >0(Supplementary Derivation), with peak magnitudegap max = â·r â1/(râ1) ·(1âr â1 )wherer=α fast /α slow . Simulations used 20 seeds per parameter setting withα slow = 0.02, noiseâ 0.1,0.2,0.4, andrranging from 1 to 63. Phase diagrams were constructed from20Ă20grids of noise rateĂtimescale ratio. Machine learning experiments Five tabular datasets from UCI/OpenML (Ionosphere, Glass, Sonar, WDBC, Vehicle) were selected because memorization is clearly observable within 100 epochs at 40% symmetric label noise. An extended benchmark of 30 OpenML datasets across three noise rates, three noise-type congurations, and ten seeds (2,700 runs total) conrmed generality. CIFAR-10 with synthetic noise and CIFAR-10N (human annotator noise) served as boundary cases. Fully connected networks provided the dense baseline. The sparse-residual variant combines an identity shortcut with a xed sparse bottleneck wired as an expander graph, reducing the dense layerâs 7 multiply-accumulate operations to roughly 22%. Bottleneck strengthαâ 0.1,0.25,0.5,1.0controls the capacity of the sparse branch. Three control architectures isolate specic ingredients: dense with label smoothing (Dense+LS), dense with a residual shortcut but no sparsity (Dense+Residual), and dense with strongL 2 (Dense+StrongReg). Two metrics quantify the gap.AUG norm is the normalized area under the positive portion of the training-minus-validation accuracy curve;T â is the rst epoch at which the gap exceedsÏ= 0.05 for three consecutive epochs. Both were computed per random seed and averaged. Statistical inference relied on xed-eects regression across 150 datasetâmodel pairs, Spearman correlations for monotonicity, and 95% condence intervals throughout. Each conguration had 10 seeds in the Full stage. Human probabilistic reversal learning Trial-level behavioral data from 292 participants (ages 8â30) who performed a probabilistic reversal learning task 49 were obtained from a publicly accessible repository (OSF;https://osf.io/7wuh4/). Participants played a two-armed bandit with 75%/25% reward probabilities that reversed periodically (4â9 reversals per session, 125â131 trials total). The empirical noise rate was 17.1%. Rolling accuracy was computed over an 8-trial window independently for truth-aligned and feedback-aligned responses; the gap is the dierence (feedback minus truth), time-locked to each reversal.AUG pos is the baseline-corrected area under the positive portion of this curve from trials 0 to 25 post-reversal.T â is the rst post-reversal trial at which truth accuracy exceeded 65% for 15 consecutive trials.AUG pos was tested against zero with a one-samplet-test, reporting Cohenâsd and bootstrap 95% CIs. Human reward/punishment learning with EEG Behavioral and 32-channel EEG data from 26 healthy adults who performed separate reward and punishment probabilistic learning tasks 50 were obtained from OpenNeuro (ds004295). One subject was dropped owing to an EEG recording malfunction, leavingN= 25for behavioral analyses; reinforcement-learning model ts usedN= 23. Artifact-based epoch rejection reduced the neural decoder samples toN REW = 24andN PUN = 25; a stricter class-balance exclusion for the prevalence analysis in Fig. 2c yieldedN= 22(see Extended Data Fig. 9 for sensitivity to this threshold). Both tasks were two-armed bandits withâŒ65â70% contingency validity. Correct responses in the reward condition earned+10cents (otherwise 0); incorrect responses in the punishment condition triggered an aversive noise burst. Each task ran for approximately 280 trials. Behavioral commitment was dened as win-stay minus lose-shift probability and tested against zero with a one-sample t-test. EEG preprocessing consisted of bandpass ltering at 1â40 Hz, average referencing, epoching â200 to 600 ms surrounding feedback onset, baseline correction, and rejection of epochs exceeding ±150ÎŒV (implemented in MNE-Python). Three feedback-locked features entered the decoder: theta power (4â8 Hz at FCz, 200â400 ms), frontal beta power (13â30 Hz, frontal channels, 200â400 ms), and P300 amplitude (Pz, 250â450 ms). Logistic regression decoders (3-fold cross-validation) predicted feedback valence and pre-feedback expectation separately; the neural gap on held-out folds was P(feedback+|EEG)âP(expectation+|EEG). Four candidate truth denitions were evaluated â slow exponential moving average of feedback, pre-feedback expectation ratings, Bayesian HMM-inferred correct side, and tted RL model Q-values â and the pre-feedback expectation was selected as the only non-circular truth label decodable above chance in both conditions (Extended Data Fig. 10). Cross-modal associations were assessed with Spearman correlation at the subject level (behavioral commitment vs. neuralAUG pos ) and Pearson correlation on smoothed trial-level time series (15-trial 8 moving average). Dopaminergic modulation (Extended Data Fig. 7) Study 2 of the Cavanagh & Frank dataset 36 includedN= 27healthy adults in a double-blind crossover design. Each participant completed the Probabilistic Selection Task under cabergoline (1.25 mg, a D2/D3 agonist) and placebo in separate sessions, with order counterbalanced. Three stimulus pairs oered graded noise regimes (20%, 30%, 40%). The primary measure was the gap regulation slope â the linear trend of the feedback-minus-truth gap time series (20-trial rolling window) over trials â quantifying how fast the gap resolves within a session.AUG pos (cumulative over-commitment) served as a secondary measure. Drug eects were examined with a one-sided paired t-test (directional hypothesis: a dopamine agonist should atten the regulation slope and reduce over-commitment). Statistical reporting All tests are two-tailed unless otherwise noted; the exception is the dopaminergic modulation analysis (Extended Data Fig. 7), where a directional hypothesis was specied before examining the data. Eect sizes are Cohenâsdthroughout: group mean divided by pooled SD for between-group comparisons, mean dierence divided by SD of dierences for paired and one-sample tests. Correlations are Spearman Ï, each accompanied by 95% CIs. Normality was assessed with ShapiroâWilk tests; when distributions deviated, non-parametric alternatives were run alongside the parametric tests (Wilcoxon signed-rank for the drug analysis, MannâWhitneyUfor developmental comparisons). No multiple-comparison correction was applied to the primary conrmatory tests (AUG pos >0 in each system), because each is a single pre-specied hypothesis tested independently per system. Exploratory correlations (individual dierences, cross-modal, developmental) carry exactp-values without correction and should be interpreted accordingly. The only covariate in any analysis is the empirical noise rate in the individual-dierences partial correlation (Extended Data Fig. 6b). Sample sizes were determined by the publicly available datasets; no power analyses were conducted in advance. âTruthâ is dened dierently across systems: in the ML and PRL analyses, truth is objective (held-out labels and the objectively better option, respectively). In the EEG analysis, the truth reference is the participantâs pre-feedback expectation of correctness â the only non-circular, non-feedback label decodable above chance in both conditions (Extended Data Fig. 10). This is an operational decision and does not imply that the participantâs condence is equivalent to objective truth. Data Availability All datasets are publicly available: Ecksteinet al.(2022) PRL data (https://osf.io/7wuh4/), Stolz et al.(2022) EEG data (https://openneuro.org/datasets/ds004295), Cavanagh & Frank (2014) PST data (https://openneuro.org/datasets/ds004532), tabular datasets via UCI/OpenML, CIFAR-10N ( https://github.com/UCSC-REAL/cifar-10-100n). Processed analysis artifacts are provided in the companion repository. Code Availability Analysis scripts for all results are available at [GitHub repository URL upon acceptance]. 9 References [1] Arpit, D.et al.A Closer Look at Memorization in Deep Networks (2017). URLhttp: //arxiv.org/abs/1706.05394. ArXiv:1706.05394 [stat]. [2] Zhang, C., Bengio, S., Hardt, M., Recht, B. & Vinyals, O. Understanding deep learning requires rethinking generalization (2017). URL http://arxiv.org/abs/1611.03530. ArXiv:1611.03530 [cs]. [3] Frenay, B. & Verleysen, M. Classication in the Presence of Label Noise: A Survey.IEEE Transactions on Neural Networks and Learning Systems25, 845â869 (2014). URL http: //ieeexplore.ieee.org/document/6685834/. [4] Song, H., Kim, M., Park, D., Shin, Y. & Lee, J.-G. Learning From Noisy Labels With Deep Neural Networks: A Survey.IEEE Transactions on Neural Networks and Learning Systems34, 8135â8153 (2023). URLhttps://ieeexplore.ieee.org/document/9729424/. [5] Daw, N. D., OâDoherty, J. P., Dayan, P., Seymour, B. & Dolan, R. J. Cortical substrates for exploratory decisions in humans.Nature441, 876â879 (2006). URLhttps://w.nature. com/articles/nature04766. [6] Sutton, R. S. & Barto, A.Reinforcement learning: an introduction. Adaptive computation and machine learning (The MIT Press, Cambridge, Massachusetts London, England, 2020), second edition edn. [7] Hassabis, D., Kumaran, D., Summereld, C. & Botvinick, M. Neuroscience-Inspired Articial Intelligence.Neuron95, 245â258 (2017). URLhttps://linkinghub.elsevier.com/retrieve/ pii/S0896627317305093. [8] Lake, B. M., Ullman, T. D., Tenenbaum, J. B. & Gershman, S. J. Building machines that learn and think like people.Behavioral and Brain Sciences40, e253 (2017). URLhttps://w. cambridge.org/core/product/identifier/S0140525X16001837/type/journal_article. [9] Richards, B. A.et al.A deep learning framework for neuroscience.Nature Neuroscience22, 1761â1770 (2019). URLhttps://w.nature.com/articles/s41593-019-0520-2. [10] Marblestone, A. H., Wayne, G. & Kording, K. P. Toward an Integration of Deep Learning and Neuroscience.Frontiers in Computational Neuroscience10(2016). URLhttp://journal. frontiersin.org/Article/10.3389/fncom.2016.00094/abstract. [11] Feldman, V. & Zhang, C. What Neural Networks Memorize and Why: Discovering the Long Tail via Inuence Estimation (2020). URLhttp://arxiv.org/abs/2008.03703. ArXiv:2008.03703 [cs]. [12]Toneva, M.et al.An Empirical Study of Example Forgetting during Deep Neural Network Learning (2019). URLhttp://arxiv.org/abs/1812.05159. ArXiv:1812.05159 [cs]. [13] Cools, R., Clark, L., Owen, A. M. & Robbins, T. W. Dening the Neural Mechanisms of Probabilistic Reversal Learning Using Event-Related Functional Magnetic Resonance Imaging. The Journal of Neuroscience22, 4563â4567 (2002). URLhttps://w.jneurosci.org/lookup/ doi/10.1523/JNEUROSCI.22-11-04563.2002. 10 [14]Izquierdo, A., Brigman, J., Radke, A., Rudebeck, P. & Holmes, A. The neural basis of reversal learning: An updated perspective.Neuroscience345, 12â26 (2017). URL https: //linkinghub.elsevier.com/retrieve/pii/S030645221600244X. [15] Schoenbaum, G., Roesch, M. R., Stalnaker, T. A. & Takahashi, Y. K. A new perspective on the role of the orbitofrontal cortex in adaptive behaviour.Nature Reviews Neuroscience10, 885â892 (2009). URLhttps://w.nature.com/articles/nrn2753. [16]Clark, L., Cools, R. & Robbins, T. The neuropsychology of ventral prefrontal cortex: Decision- making and reversal learning.Brain and Cognition55, 41â53 (2004). URLhttps://linkinghub. elsevier.com/retrieve/pii/S0278262603002847. [17] Collins, A. G. E. & Frank, M. J. How much of reinforcement learning is working memory, not reinforcement learning? A behavioral, computational, and neurogenetic analysis.European Journal of Neuroscience35, 1024â1035 (2012). URL https://onlinelibrary.wiley.com/ doi/10.1111/j.1460-9568.2011.07980.x. [18] Niv, Y., Edlund, J. A., Dayan, P. & OâDoherty, J. P. Neural Prediction Errors Reveal a Risk- Sensitive Reinforcement-Learning Process in the Human Brain.The Journal of Neuroscience 32, 551â562 (2012). URLhttps://w.jneurosci.org/lookup/doi/10.1523/JNEUROSCI. 5498-10.2012. [19] Dayan, P. & Balleine, B. W. Reward, Motivation, and Reinforcement Learning.Neuron36, 285â 298 (2002). URLhttps://linkinghub.elsevier.com/retrieve/pii/S0896627302009637. [20]Cavanagh, J. F. & Frank, M. J. Frontal theta as a mechanism for cognitive control.Trends in Cognitive Sciences18, 414â421 (2014). URLhttps://linkinghub.elsevier.com/retrieve/ pii/S1364661314001077. [21] Ridderinkhof, K. R., Ullsperger, M., Crone, E. A. & Nieuwenhuis, S. The Role of the Medial Frontal Cortex in Cognitive Control.Science306, 443â447 (2004). URLhttps://w.science. org/doi/10.1126/science.1100301. [22] Polich, J. Updating P300: An integrative theory of P3a and P3b.Clinical Neurophys- iology118, 2128â2148 (2007). URL https://linkinghub.elsevier.com/retrieve/pii/ S1388245707001897. [23] Holroyd, C. B. & Coles, M. G. H. The neural basis of human error processing: Reinforcement learning, dopamine, and the error-related negativity.Psychological Review109, 679â709 (2002). URLhttps://doi.apa.org/doi/10.1037/0033-295X.109.4.679. [24]Walsh, M. M. & Anderson, J. R. Learning from experience: Event-related potential correlates of reward processing, neural adaptation, and behavioral choice.Neuroscience & Biobehavioral Reviews36, 1870â1884 (2012). URLhttps://linkinghub.elsevier.com/retrieve/pii/ S0149763412000875. [25] Patrini, G., Rozza, A., Menon, A. K., Nock, R. & Qu, L. Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2233â2241 (IEEE, Honolulu, HI, 2017). URLhttp: //ieeexplore.ieee.org/document/8099723/. 11 [26]Natarajan, N., Dhillon, I. S., Ravikumar, P. & Tewari, A. Learning with noisy labels. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1, NIPSâ13, 1196â1204 (Curran Associates Inc., Red Hook, NY, USA, 2013). Event-place: Lake Tahoe, Nevada. [27]Schultz, W., Dayan, P. & Montague, P. R. A Neural Substrate of Prediction and Reward. Science275, 1593â1599 (1997). URLhttps://w.science.org/doi/10.1126/science.275. 5306.1593. [28] Schultz, W. Dopamine reward prediction-error signalling: a two-component response.Nature Reviews Neuroscience17, 183â195 (2016). URLhttps://w.nature.com/articles/nrn. 2015.26. [29] Frank, M. J., Seeberger, L. C. & OâReilly, R. C. By Carrot or by Stick: Cognitive Reinforcement Learning in Parkinsonism.Science306, 1940â1943 (2004). URLhttps://w.science.org/ doi/10.1126/science.1102941. [30] Frank, M. J. & OâReilly, R. C. A mechanistic account of striatal dopamine function in human cognition: Psychopharmacological studies with cabergoline and haloperidol.Behavioral Neuroscience120, 497â517 (2006). URLhttps://doi.apa.org/doi/10.1037/0735-7044.120. 3.497. [31] Cools, R. Dopaminergic modulation of cognitive function-implications for l-DOPA treatment in Parkinsonâs disease.Neuroscience & Biobehavioral Reviews30, 1â23 (2006). URLhttps: //linkinghub.elsevier.com/retrieve/pii/S0149763405000540. [32] Belkin, M., Hsu, D., Ma, S. & Mandal, S. Reconciling modern machine-learning practice and the classical biasâvariance trade-o.Proceedings of the National Academy of Sciences116, 15849â15854 (2019). URLhttps://pnas.org/doi/full/10.1073/pnas.1903070116. [33]Saxe, A. M., McClelland, J. L. & Ganguli, S. A mathematical theory of semantic development in deep neural networks.Proceedings of the National Academy of Sciences116, 11537â11546 (2019). URLhttps://pnas.org/doi/full/10.1073/pnas.1820226116. [34]Neyshabur, B., Bhojanapalli, S., McAllester, D. & Srebro, N. Exploring Generalization in Deep Learning (2017). URLhttp://arxiv.org/abs/1706.08947. ArXiv:1706.08947 [cs]. [35] Gershman, S. J. & Daw, N. D. Reinforcement Learning and Episodic Memory in Humans and Animals: An Integrative Framework.Annual Review of Psychology68, 101â128 (2017). URL https://w.annualreviews.org/doi/10.1146/annurev-psych-122414-033625. [36]Cavanagh, J. F., Masters, S. E., Bath, K. & Frank, M. J. Conict acts as an implicit cost in reinforcement learning.Nature Communications5, 5394 (2014). URLhttps://w.nature. com/articles/ncomms6394. [37] Casey, B. J. Beyond Simple Models of Self-Control to Circuit-Based Accounts of Adolescent Behavior.Annual Review of Psychology66, 295â319 (2015). URLhttps://w.annualreviews. org/doi/10.1146/annurev-psych-010814-015156. [38] Somerville, L. H. & Casey, B. Developmental neurobiology of cognitive control and motivational systems.Current Opinion in Neurobiology20, 236â241 (2010). URLhttps://linkinghub. elsevier.com/retrieve/pii/S0959438810000073. 12 [39]Hartley, C. A. & Somerville, L. H. The neuroscience of adolescent decision-making.Current Opinion in Behavioral Sciences5, 108â115 (2015). URLhttps://linkinghub.elsevier.com/ retrieve/pii/S2352154615001205. [40] Waltz, J. A. & Gold, J. M. Probabilistic reversal learning impairments in schizophrenia: Further evidence of orbitofrontal dysfunction.Schizophrenia Research93, 296â303 (2007). URL https://linkinghub.elsevier.com/retrieve/pii/S092099640700120X. [41]Peterson, D.et al.Probabilistic reversal learning is impaired in Parkinsonâs disease.Neu- roscience163, 1092â1101 (2009). URL https://linkinghub.elsevier.com/retrieve/pii/ S0306452209012068. [42] Maia, T. V. & Frank, M. J. From reinforcement learning models to psychiatric and neurological disorders.Nature Neuroscience14, 154â162 (2011). URL https://w.nature.com/articles/ n.2723. [43] Murray, G. K.et al.Substantia nigra/ventral tegmental reward prediction error disruption in psychosis.Molecular Psychiatry13, 267â276 (2008). URL https://w.nature.com/ articles/4002058. [44] Waltz, J. A.et al.Patients with Schizophrenia have a Reduced Neural Response to Both Unpredictable and Predictable Primary Reinforcers.Neuropsychopharmacology34, 1567â1577 (2009). URLhttps://w.nature.com/articles/npp2008214. [45] Rendell, L.et al.Why Copy Others? Insights from the Social Learning Strategies Tourna- ment.Science328, 208â213 (2010). URLhttps://w.science.org/doi/10.1126/science. 1184719. [46] Richerson, P. J. & Boyd, R.Culture and the evolutionary process(University of Chicago press, Chicago London, 1985). [47] Rescorla, R. & Wagner, A. A theory of Pavlovian conditioning: Variations in the eectiveness of reinforcement and nonreinforcement. InClassical Conditioning I: Current Research and Theory, vol. Vol. 2 (1972). Journal Abbreviation: Classical Conditioning I: Current Research and Theory. [48]Siegel, S. & Allan, L. G. The widespread inuence of the Rescorla-Wagner model.Psychonomic Bulletin & Review3, 314â321 (1996). URL http://link.springer.com/10.3758/BF03210755. [49]Eckstein, M. K., Master, S. L., Dahl, R. E., Wilbrecht, L. & Collins, A. G. Reinforcement learning and Bayesian inference provide complementary models for the unique advantage of adolescents in stochastic reversal.Developmental Cognitive Neuroscience55, 101106 (2022). URLhttps://linkinghub.elsevier.com/retrieve/pii/S1878929322000494. [50]Stolz, C., Pickering, A. & Mueller, E. M. Reward gain and punishment avoidance reversal learning (2022). URLhttps://openneuro.org/datasets/ds004295/versions/1.0.0. 13 Figure 1|The feedbackâtruth gap appears across machines and humans. Each panel plots feedback-aligned minus truth-aligned performance in a system that receives noisy supervision. All three systems produce the gap; only the temporal trajectory diers.a,An unregularized dense neural network trained on tabular data with 40% label noise (representative run at the 46th percentile of gap magnitude acrossn= 90model-noise congurations; Ionosphere dataset; see Extended Data Fig. 1). Training accuracy (grey) exceeds validation accuracy (teal), and the resulting gap (shaded) persists and grows throughout training with no sign of self-correction. b,Human reversal learning ( N= 292subjects, 1,764 reversal epochs from the Ecksteinet al. dataset; ref. 11). The baseline-corrected feedbackâtruth gapâgap(t)is aligned to reversal (t= 0). A transient positive deection (over-commitment) appears immediately post-reversal and returns toward baseline withinâŒ5â10 trials (8-trial rolling window; baseline=mean pre-reversal gap).c, Decoder-based neural gap during reward learning (n= 24subjects from the Stolzet al.EEG dataset; ref. 12). Logistic regression decoders trained on theta power, beta power, and P300 amplitude yield P(feedback+|EEG) (grey) and P(expectation+|EEG) (teal), where expectation = pre-feedback condence in the chosen option. The gap between decoder probabilities diminishes across trials (15-trial rolling mean), consistent with progressive neural adaptation; the shading transition marks the onset of gap narrowing. Color convention for all panels: teal=truth, grey=feedback, shaded =learning gap. 14 Figure 2|The feedbackâtruth gap is universal across systems and mathematically inevitable. Gap prevalence across machines and humans, together with the two-timescale model that predicts it.a,ML prevalence across 90 conditions (5 datasetsĂ6 architecturesĂ3 noise rates). Stacked bars classify each condition as strong gap (dark purple, AUGâ„0.05), weak (light purple,0.005†AUG<0.05), or absent (grey, AUG<0.005). Across all 90 conditions, 77 (86%) show a measurable gap. Gap prevalence increases with noise rate and exceeds 90% at 40% noise.b,Cumulative distribution of over-commitment gap (AUG pos ) acrossN= 292PRL subjects. Every subject shows gap>0(292/292, 100%). Vertical dashed lines mark quartiles; median AUG pos = 0.043.c, Per-subject neural gap (AUG+fraction, feedback dominance) in the reward task (N= 22after class-balance exclusion; see Methods and Extended Data Fig. 9), sorted by magnitude. 21/22 subjects (95%) show feedback-dominant neural representation (mean±SEM= 0.82±0.03). Orange bars are individual subjects; white dashed line marks 50% (parity).d,Peak gap magnitude from the closed-form solution plotted against timescale ratior=α fast /α slow (log scale). Purple curve: gap max =Ύ·r â1/(râ1) ·(1âr â1 ). The gap is positive for allr >1and equals zero exactly atr= 1 (matched timescales, black circle). Empirical systems overlaid: Dense network (star), Human PRL (circle), Regularized network (square). See Extended Data Fig. 2c for the full phase diagram. 15 Figure 3|The gap indexes damage in unregulated networks but structured engagement in regulated humans. â0.2â0.10.00.10.20.30.40.5 Final trainâval gap 0.3 0.4 0.5 0.6 0.7 0.8 Final validation accuracy Ï = -0.74 Dense (n = 150) ResEx (n = 100) a ML: gap indexes damage 0.0000.0250.0500.0750.1000.1250.1500.175 Over-commitment gap (AUG+) 0 5 10 15 20 25 30 Recovery time T* (trials) Ï = +0.20, p < 10â»Âł b PRL: gap indexes engagement (N = 288) â0.8â0.6â0.4â0.20.00.20.40.6 Spearman Ï (gap vs outcome) Dense ML ResEx ML PRL behav. NeuralĂbehav. Ï = -0.74 Ï = -0.37 Ï = +0.20 Ï = +0.43 c Cross-system sign flip 12345 Trial block 0.00 0.02 0.04 0.06 0.08 0.10 0.12 Neural gap (fb â truth) REW d = 0.36 PUN d = 0.96 d Neural gap declines over time REW PUN Whether the gap predicts harm or engagement depends on whether the system regulates it.a,In unregulated dense networks, a larger nal train-validation gap predicts lower validation accuracy (150 dense runs, black; 100 sparse-residual runs, teal). Dense networks: Spearman Ï=â0.74. Regression line with 95% CI shown for dense networks.b,In human PRL the relationship reverses (N= 288 with sucient post-reversal data). Larger over-commitment (AUG pos ) predicts faster recovery (T â in trials; SpearmanÏ= +0.20,p <10 â3 ), suggesting that the gap reects contingency engagement rather than damage. Regression line with 95% CI; point sizes proportional to density.c,Forest plot of gapâoutcome Spearman correlations across four systems. Dense ML:Ï=â0.74(harm); ResEx ML:Ï=â0.37(attenuated); PRL behavioral:Ï= +0.20(sign reversal); NeuralĂbehavioral: Ï= +0.43(concordant amplication). 95% CIs shown.d,Mean neural gap (feedback minus truth decoder probability) across ve trial blocks. Reward (orange):d= 0.36; punishment (purple): d= 0.82. Error bars:±1SEM. The declining trajectory indicates progressive neural recalibration toward truth, distinct from the persistent accumulation seen in networks and the behavioral recovery observed in reversal learning. 16 Figure 4|The gap is causally tunable via architectural intervention. The sparse-residual bottleneck parameterαprovides graded, monotonic control over memorization dynamics and reveals a dissociation between gap suppression and generalization.a,Normalized cumulative gap (AUG norm ) increases monotonically withα(SpearmanÏ= 1.0across all datasets with measurable gaps). The dense baseline (grey diamond) shows the largest gap. Light points are individual runs; diamonds are group means with 95% CI. Five datasets, 40% label noise, 10 seeds per conguration.b,Gap onset time ( T â in epochs) also increases monotonically withα(Ï= 1.0): more constrained architectures delay memorization. Dense baseline (grey diamond) shows earliest onset. Error bars: 95% CI.c,Test accuracy plotted against gap magnitude (AUG norm ). The relationship is non-monotonic â accuracy peaks atα= 0.25(teal), not at maximal gap suppression. Atα= 1.0 (dark teal), accuracy falls below the moderate-suppression level, exposing a capacityâmemorization trade-o. Dense baseline (grey diamond) achieves lower accuracy with larger gap. Point sizes proportional to sample count; error bars: 95% CI. 17 Figure 5|Neuralâbehavioral concordance in the feedbackâtruth gap. EEG dataset (N= 25; ref. 12). The roughly 10-fold amplication from neural representation to overt behavior is the central observation.a,Behavioral commitment (win-stay minus lose-shift) for reward (orange,d= 3.32) and punishment (purple,d= 3.92). Every subject shows positive commitment in both conditions. Connected lines are paired within-subject values; box plots show median and IQR. Dierence between tasks:d= 0.40,p= 0.053(two-sided pairedt-test).b,Neural gap magnitude (AUG pos from EEG decoders) is roughly an order of magnitude smaller than the behavioral eect. Reward:d= 1.03(N= 24); punishment:d= 1.82(N= 25). All subjects positive in both conditions (24/24 reward, 25/25 punishment). Box plots with individual data; error bars: ±1 SEM. Dierence:d= 0.39,p= 0.061.c,Subjects with larger neural over-representation also show stronger behavioral commitment. Reward: SpearmanÏ= +0.43,p= 0.036(n= 24); punishment: Ï= +0.42,p= 0.042(n= 24). Regression lines with 95% CI shown per task. Orange circles: reward; purple squares: punishment. 18 Extended Data Figure 1|Gap emergence rates across systems. Under noisy conditions the gap is near-universal. Gap=feedback-aligned minus truth-aligned performance throughout.a,Fraction of ML conditions showing gap (AUGâ„0.005) at each noise rate (5 datasetsĂ6 architecturesĂ3 noise rates= 90conditions). Across all 90 conditions, 77 (86%) show a measurable gap; the proportion exceeds 90% at 40% noise. Bars are proportions; no error bars (each condition is a single count).b,Cumulative distribution ofAUG pos acrossN= 292 PRL subjects (Eckstein dataset; ref. 11). Virtually all subjects show positive over-commitment (median AUG pos = 0.043). Feedback=reward received; truth=objectively better option.c, Per-subject AUG+fraction (feedback dominance;N= 22non-excluded subjects from Stolz EEG dataset; ref. 12). 95% of subjects exceed the 50% parity line (mean±SEM= 0.82±0.03). Feedback =reinforcer valence; truth=pre-feedback expectation. 19 Extended Data Figure 2|Boundary conditions: when the gap vanishes. The inevitability claim is falsiable: the gap disappears under specic counterfactual conditions. Gap=feedback-aligned minus truth-aligned performance throughout.a,AUG by noise rate and architecture across 90 ML conditions (5 datasetsĂ6 architecturesĂ3 noise rates). The gap scales with noise and vanishes under strong regularization or 0% noise; 15/90 conditions show near-zero gap (AUG<0.01). Points are individual conditions.b,Distribution ofAUG pos acrossN= 292PRL subjects. 23 subjects (7.9%) show near-zero gap, consistent with low-noise or matched-timescale subgroups. Feedback=reward received; truth=objectively better option.c,Phase diagram showing AUG across noise rateĂarchitecture. Color scale is mean normalized AUG; cells with white squares have AUG<0.01(gap absent). The gap vanishes under clean labels, strong regularization, or capacity suppression. 20 Extended Data Figure 3|Neural validation details. Subject-level neural gap data, decoder validation, and cross-modal dynamics. Feedback=reinforcer valence (reward or punishment); truth=prior expectation (pre-feedback condence; see Methods). aâb,Individual neuralAUG pos values for reward (N= 24) and punishment (N= 25), sorted by magnitude. All values are positive. Bars are individual subjects.câd,An example subject (s6, reward task) illustrates decoder probability trajectories for feedback and expectation (15-trial rolling mean). Shading marks the gap between decoder curves.e,Distributions of the three EEG features by feedback type: theta power, beta power, and P300 amplitude. Box plots show median and IQR; whiskers extend to1.5ĂIQR.f,The memorization index is approximately 10-fold larger at the behavioral than neural level (in standardized eect size), illustrating the amplication from cortical representation to overt strategy. Error bars:±1SEM. 21 Extended Data Figure 4|Early training dynamics predict later memorization regime. Features measured early in training (epochs 1â15) predict the later memorization regime (epochs 35+) with a strict 20-epoch temporal buer, indicating that gap dynamics are deterministic rather than stochastic.a,Cross-validated AUC acrossN= 100ML runs (5 datasets, 40% noise). Logistic regression on 14 early features achieves AUC= 0.96(Brier score= 0.064); the single best feature (gap_slope) alone reaches AUC= 0.88. Bars are AUC; error metric is cross-validation standard error. Inset: generalization across datasets.b,In PRL (N= 291; one subject excluded for insucient post-reversal data), early post-reversal gap predicts later perseveration ( Ï= 0.85,p <0.001) but not recovery time T â (Ï=â0.04, n.s.). Early pressure and downstream regulation are therefore separable.c,In the EEG sample ( N= 22), early neural gap predicts later signal persistence ( Ï= 0.45,p= 0.035). The sample is small, but the eect is directionally consistent with the behavioral and ML results. 22 Extended Data Figure 5|Breadth validation across 30 OpenML datasets. How often the sparse bottleneck outperforms the dense baseline, plotted against noise rate across 30 OpenML datasets (n= 2,700total runs; 30 datasetsĂ3 noise ratesĂ10 seedsĂ3 congurations). Truth=held-out test labels (uncorrupted); feedback=noisy training labels. At 0% noise: 57% (no reliable advantage). At 40% noise the win fraction rises to 93% (medianâ = +3.28p). Error bars: 95% binomial CI. The advantage is noise-dependent, not architecture-dominant. 23 Extended Data Figure 6|Individual dierences in gap dynamics. Individual variation in gap magnitude and recovery speed reveals heterogeneity in regulation. Gap= feedback-aligned minus truth-aligned performance; feedback=reward received; truth=objectively better option (PRL) or pre-feedback expectation (EEG).a,In PRL (N= 292), noise rate positively predicts gap magnitude (Ï= 0.20,p= 0.001) but explains only 0.6% of variance â most of the variation lies between individuals, not between task conditions. Each point is one subject; line is a linear t.b,Gap magnitude predicts recovery speed (partialÏ= 0.14,p= 0.021, controlling for noise rate) and recovery timeT â (partialÏ= 0.18,p= 0.002). PRL,N= 292. Each point is one subject. c,Subjects classied as high-regulation versus low-regulation phenotypes show separable gap proles ( d= 0.25). Shading:±1SEM.d,In the EEG sample (N= 22), neuralAUG pos correlates negatively with lose-shift rate (Ï=â0.48,p= 0.026): subjects with stronger neural over-representation are less likely to shift after losses. Each point is one subject; line is a linear t. 24 Extended Data Figure 7|Dopaminergic modulation of the feedbackâtruth gap. Cabergoline, a dopamine D2/D3 agonist, attens within-session gap regulation in a double-blind crossover ( N= 27; Cavanagh & Frank dataset; 36 ). Gap=feedback-aligned minus truth-aligned accuracy (20-trial rolling window).a,The gap regulation slope (linear trend of gap over trials) is signicantly atter under cabergoline than placebo (one-sided pairedt-test,d= 0.44,p= 0.015; Wilcoxon p= 0.016). Connected dots are individual subjects; red squares are group means.b, Group-averaged gap time courses under cabergoline (red) and placebo (blue), with SEM shading. Under placebo the gap declines more steeply (stronger regulation); cabergoline blunts this trajectory. c,Cumulative over-commitment (AUG pos ) is directionally lower under cabergoline (d=â0.32, one-sidedp= 0.053), suggesting that the drug eect extends to both the dynamics and the overall level of the gap. 25 Extended Data Figure 8|Developmental gradient in feedback persistence and gap regulation. â10â50510152025 Trials from reversal â0.050 â0.025 0.000 0.025 0.050 0.075 0.100 0.125 Baseline-corrected gap a Minors (8-17) Young Adults (18-24) Adults (25-30) MinorsYoung Adults Adults 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 AUG pos Minors vs Adults d = 0.04, p = 0.330 b MinorsYoung Adults Adults 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Win-stay rate Minors vs Adults d = -0.42, p = 0.0002 c 1020304050 Age (years) 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Win-stay rate Ï = 0.32 p = 2.8e-08 N = 285 d With development, the balance between exibility and persistence shifts in a direction that parallels the pharmacological eect of dopamine agonism. Gap=feedback-aligned minus truth-aligned accuracy; feedback=reward received; truth=objectively better option.a,Reversal-locked gap dynamics for two developmental cohorts (sID<300: minors ages 8â17,N= 178; sIDâ„300: adults ages 18â30,N= 114; Ecksteinet al.dataset; ref. 11). Mean baseline-correctedâgap(t), aligned to reversal ( t= 0). Both groups show a positive post-reversal gap; adults show a slightly atter recovery trajectory. The faster gap closure in minors reects reduced feedback persistence rather than enhanced compensatory control. Shading is SEM.b,Total over-commitment (AUG pos ) does not dier signicantly between minors and adults (d= 0.04,p= 0.66, two-sided MannâWhitneyU), so the dierent dynamics inado not translate into a dierence in cumulative magnitude. Violin plots with individual data points; error bars are 95% bootstrap CI.c,Win-stay rate (probability of repeating a rewarded choice) increases from minors to adults ( d=â0.42,p <0.001). The direction parallels the cabergoline eect (Extended Data Fig. 7): both developmental maturation and acute dopamine agonism increase feedback persistence.d,Treating age as a continuous variable conrms the pattern (SpearmanÏ= 0.32,p <0.0001; age range 7â50 years,N= 285with matched demographics). The grouped result incis not an artifact of binning. 26 Extended Data Figure 9|Cross-modal correlation robustness to exclusion threshold. The cross-modal Spearman correlation (behavioral commitment versus neuralAUG pos ) is stable across class-balance exclusion thresholds. Subjects whose expectation label imbalance exceeds the threshold are excluded at each step; thresholds from 0.60 to 1.00 were tested. The pooled correlation (reward+punishment) ranges fromÏ= 0.51toÏ= 0.59and remains signicant (p <0.01) throughout. The 0.80 threshold used in the main analysis falls in the middle of this stable range. Per-condition correlations (reward, punishment) are shown separately.Nper threshold varies from 14 to 38 subjects. 27 Extended Data Figure 10|Non-circular truth reconstruction and decoder comparison. Post-feedback EEG features decode feedback valence and prior expectation (pre-feedback condence) but not objective environmental correctness inferred from Bayesian or reinforcement-learning models. We tested four truth denitions: slow exponential moving average of feedback (circular baseline), pre-feedback expectation ratings, Bayesian HMM-inferred correct side (forwardâbackward with 70/30 emission structure), and tted RL modelQ-values (M5 parameters from ref. 12).a,HMM posterior trajectory for an example subject (s1, REW task): P(left correct|all data), with overlaid expectation ratings (scaled). Purple shading marks epochs where the HMM infers right-correct state.b,Agreement between truth denitions. Expectation truth is nearly independent of the other three (âŒ50%agreement), whereas EMA, HMM, and RL cluster together (âŒ65â72%). This 28 is consistent with expectation ratings capturing a genuinely distinct signal.c,Decoder AUROC across truth denitions and tasks. Only the expectation-truth decoder signicantly exceeds chance (REW: AUROC= 0.541,t(23) = 2.43,p= 0.012; PUN: AUROC= 0.545,t(24) = 3.84,p <10 â3 ). The feedback decoder is shown for reference.d,Neural gap (feedback minus truth AUROC) by truth denition. With non-circular expectation truth the gap is near-zero in REW (gap= +0.006, p= 0.39) and small in PUN (gap= +0.038,p= 0.054), indicating that feedback and expectation representations are nearly balanced in post-feedback EEG.e,ROI decomposition for Bayesian truth (all features, frontal-only, posterior-only). No feature subset rescues objective-truth decoding.f, Bayesian truth decoder AUROC acrossp switch values (0.01â0.10). The null result for objective truth is stable.g,Individual-subject scatter of feedback AUROC versus Bayesian truth AUROC. The null is not driven by outlier subjects.h,Real versus null AUROC distributions for the Bayesian truth decoder (REW task). HMM validation: reward rate when HMM-inferred correct= 0.86, when incorrect= 0.28, conrming reconstruction validity against the known 70/30 contingency structure. 29 Extended Data Table 1|Gap dynamics across noise-robust training methods. MethodMedian AUG norm MedianT â Frac. AUG>0.005Median val. acc. Dense (baseline)0.12124.073%0.646 Sparse-residual0.00050.0 (censored)30%0.770 Co-teaching0.00050.0 (censored)23%0.806 Forward loss correction0.00948.857%0.705 Four training methods applied to 30 OpenML datasets with 40% symmetric label noise (3 seeds each). Median AUG norm andT â computed per dataset (mean across seeds), then median across datasets. All methods produce a measurable gap in a substantial fraction of datasets; no method eliminates it entirely. The gap is suppressed but not abolished by algorithmic intervention, consistent with inevitability under the two-timescale framework. 30 Extended Data Table 2|Sparse-residual gap suppression across 30 datasets. Noise rateNdatasets MedianâAUG Frac.âAUG<0Medianâ acc Frac.â acc >0T â delay frac. 0.230â0.05580% [63â93%]+0.05287% [73â97%]80% 0.430â0.11997% [90â100%]+0.08193% [83â100%]91% Dense baseline versus sparse-residual architecture (hidden dimensions= 256, degree= 5,α= 0.25) across 30 OpenML datasets at two noise rates (3 noise typesĂ10 seeds= 18runs per dataset per architecture).âAUG=sparse-residual minus dense (negative=suppression).â acc =sparse- residual minus dense (positive=improvement).T â delay fraction: proportion of datasets where sparse-residual delays the onset of memorization. Sparse-residual models use approximately 5% of dense-baseline compute (median MACs ratio= 0.052). 31 Extended Data Table 3|Topology null study. MetricValue Ndatasets10 MedianâAUG norm (randomâexpander)0.000 Frac.|âAUG norm |<0.01100% Medianâ test acc â0.0004 Frac. ordering consistent (randomâ„expander)0% Expander-graph versus random-regular graph wiring within the same sparse-residual scaold (10 datasets, symmetric noise at 40%, 3 seeds).âAUG norm =random-regular minus expander. Graph topology produces no detectable dierence in gap dynamics: medianâAUG norm = 0.000and 100% of datasets fall within|â|<0.01. The sparse-residual scaolding, not specic graph structure, is the operative ingredient for gap suppression. 32 Extended Data Table 4|Reinforcement-learning parameters selectively predict gap dynamics. Outcome PredictorÏpResidualizedÏResidualizedp AUG pos α(standard RW)0.291<10 â6 0.0570.328 AUG pos Ï(perseveration)â0.264<10 â5 â0.0830.159 AUG pos α â (asymmetric)0.291<10 â6 â AUG pos α + âα â (asymmetry)â0.222<10 â4 â Gap slopeα(standard RW)0.333<10 â8 0.367<10 â10 Gap slopeÏ(perseveration)â0.1490.011â0.207<10 â3 Gap slopeα â (asymmetric)0.314<10 â7 â Gap slopeα + âα â (asymmetry)â0.1290.027â0.1130.054 Spearman correlations between tted RL parameters and gap metrics (N= 292subjects, Eckstein 2022 dataset). Three models tted: standard RescorlaâWagner (RW), RW with perseveration (Ï), and asymmetric RW (α + ,α â ). The standard learning rateαpredicts both gap magnitude (AUG pos ) and recovery speed (gap slope), consistent with the two-timescale framework. Perseveration Ïshows inverse relationships. These associations survive residualization for task-level confounds (noise rate, trial count, reversal count). 33 Supplementary Information Learning under noisy supervision is governed by a feedbackâtruth gap Elan Schonfeld 1 Elias Wisnia 2 Contents 1 Supplementary Derivation: Two-Timescale Gap2 1.1 Model denition . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.2 Gap dynamics after a state change . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.3 Proof of strict positivity . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.4 Peak timing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 2 1.5 Peak magnitude . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.6 Limiting behavior . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 1.7 Cumulative gap . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 3 2 Supplementary Methods4 2.1 Machine learning: extended experimental details . . . . . . . . . . . . . . . . . . . . 4 2.1.1 Dataset specications . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.1.2 Architecture details . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.1.3 Memorization metric denitions . . . . . . . . . . . . . . . . . . . . . . . . . . 4 2.1.4 Alpha-grid causal intervention . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2 Human probabilistic reversal learning: extended methods . . . . . . . . . . . . . . . 5 2.2.1 Dataset and participants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2.2 Task structure . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2.3 Reversal-locked analysis . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.2.4 AUG pos computation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 5 2.3 Human reward/punishment learning with EEG: extended methods . . . . . . . . . . 6 2.3.1 Dataset and participants . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.3.2 EEG preprocessing . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.3.3 Neural features . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.3.4 Neural gap decoder . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6 2.3.5 Source localization . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 2.3.6 Cross-modal correlation . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 7 3 Supplementary Results8 3.1 Complete statistical summary . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 8 3.2 Reinforcement learning model parameters . . . . . . . . . . . . . . . . . . . . . . . . 8 3.3 Positioning relative to noise-robust machine learning . . . . . . . . . . . . . . . . . . 8 4 Supplementary Figures10 1 1 Supplementary Derivation: Two-Timescale Gap Below we work out the gap in closed form for the two-timescale learner, show that it is strictly positive whenever the timescales dier, and derive the peak magnitude. 1.1 Model denition A learner tracks an environment statesâ0,1via two channels, each updated by an exponential moving average: x fast (t+ 1) =x fast (t) +α fast · ( o(t)âx fast (t) ) (1) x slow (t+ 1) =x slow (t) +α slow · ( s(t)âx slow (t) ) (2) whereo(t)is the observed feedback signal (equal tos(t)with probability1âΔand1âs(t)with probability Δ) and0< α slow < α fast â€1. In the continuous-time limit, these become exponential decay processes with ratesα fast andα slow respectively. 1.2 Gap dynamics after a state change Consider a state reversal of magnitudeâat timet= 0. Both channels begin with the same initial displacementâfrom the new true state. In expectation, the residual error of each channel decays exponentially: E[error fast (t)] = â·e âα fast ·t (3) E[error slow (t)] = â·e âα slow ·t (4) The feedbackâtruth gap is the dierence in alignment between the two channels (equivalently, the dierence in residual errors): gap(t) = â· ( e âα slow ·t âe âα fast ·t ) (5) 1.3 Proof of strict positivity Proposition 1.Ifα fast > α slow >0, then gap(t)>0for allt >0. Proof.Sinceα fast > α slow , we haveα fast ·t > α slow ·tfor allt >0. Becausef(x) =e âx is strictly decreasing,e âα fast ·t < e âα slow ·t . Thereforee âα slow ·t âe âα fast ·t >0, and sinceâ>0, gap(t)>0.⥠Corollary (Counterfactual).Whenα fast =α slow ,gap(t) = 0for allt, regardless of noise level Δ. 1.4 Peak timing Dierentiating Eq. (5) with respect totand setting the derivative to zero: d dt gap(t) = â· ( âα slow ·e âα slow ·t +α fast ·e âα fast ·t ) = 0(6) This yields: α fast ·e âα fast ·t â =α slow ·e âα slow ·t â (7) Taking logarithms: ln ( α fast α slow ) = (α fast âα slow )·t â (8) 2 Dening the timescale ratior=α fast /α slow : t â = lnr α slow ·(râ1) (9) 1.5 Peak magnitude Substitutingt â back into Eq. (5): e âα slow ·t â =e âlnr/(râ1) =r â1/(râ1) (10) e âα fast ·t â =e âr·lnr/(râ1) =r âr/(râ1) (11) Therefore: gap max = â· ( r â1/(râ1) âr âr/(râ1) ) (12) This can be simplied by noting thatr âr/(râ1) =r â1/(râ1) ·r â1 : gap max = â·r â1/(râ1) · ( 1â 1 r ) (13) 1.6 Limiting behavior Large timescale mismatch(rââ): Asrââ,r â1/(râ1) â1and(1â1/r)â1, sogap max ââ. The peak gap approaches the full magnitude of the state change. Matched timescales(râ1 + ): Writingr= 1 +Δand expanding: r â1/(râ1) = (1 +Δ) â1/Δ âe â1 asΔâ0(14) while(1â1/r) =Δ/(1 +Δ)â0. Thusgap max â0, conrming that the gap vanishes continuously as timescales converge. 1.7 Cumulative gap The cumulative positive gap, which corresponds to theAUG pos metric used in the main text, integrates Eq. (5) over the post-reversal period: AUG pos = â« T 0 gap(t)dt= â· ( 1 α slow â 1 α fast ) · ( 1âresidual(T) ) (15) where the residual term vanishes asTââ. In the innite-horizon limit: AUG pos = â· râ1 α fast (16) This conrms that cumulative over-commitment scales linearly with timescale ratio and inversely with the fast learning rate. 3 2 Supplementary Methods 2.1 Machine learning: extended experimental details 2.1.1 Dataset specications The core experiments used ve tabular benchmarks from UCI/OpenML: Table 1: Primary benchmark datasets. DatasetSamples Features Classes OpenML ID Ionosphere35134259 Glass2149641 Sonar20860240 WDBC5693021510 Vehicle84618454 To inject noise, we ipped each training label to a uniformly random class with probability 0.4 (symmetric noise). Validation and test labels were left untouched. The breadth validation (Extended Data Fig. 5) drew on 30 additional OpenML datasets, spanning 150â10,000 samples, 4â784 features, and 2â26 classes. 2.1.2 Architecture details Dense baseline.Two hidden layers (128 and 64 units), ReLU, softmax output. Trained with Adam at lr= 0.001, batch size 32, for 50 epochs. Nothing fancy. Sparse-residual (ResEx).The rst hidden layer is replaced by a residual block with a xed sparse branch: h 1 =x+α·Ï(W sparse ·x)(17) W sparse is an expander-graph adjacency matrix â about 22% dense, withd= 3random neighbours per node. Whenα= 0the block is a pass-through; turningαup lets the sparse branch reshape the representation. The point is to have a single knob that smoothly dials memorization capacity. Controls.Three architectures isolate individual ingredients: âąDense+LS â the dense baseline plus label smoothing (Δ= 0.1) âąDense+Residual â dense with an identity shortcut butnosparse bottleneck âąDense+StrongReg â dense with aggressiveL 2 (λ= 0.01) 2.1.3 Memorization metric denitions The generalization gap at epochtis simply training accuracy minus validation accuracy: gap(t) =acc train (t)âacc val (t)(18) We then dene two summary statistics. The rst,AUG norm , captures cumulative memorization pressure: AUG norm = 1 T T â t=1 max(0,gap(t))(19) 4 whereTis the total number of epochs â soAUG norm is an average over the whole run, not just the tail. The second,T â , marks onset: the rst epoch where the gap exceedsÏ= 0.05for three epochs running. If that never happens, we setT â =T(right-censored). Both metrics are computed per seed and averaged. 2.1.4 Alpha-grid causal intervention The logic is straightforward. If the sparse branch really controls memorization, then increasingα should push bothAUG norm andT â up monotonically â but test accuracy should peak somewhere in the middle, because at highαthe bottleneck starts hurting useful capacity too. We swept αâ0.1,0.25,0.5,1.0across all ve benchmarks (10 seeds each, 200 runs). SpearmanÏbetween αand the group-median metric was 1.0 for bothAUG norm andT â (p <10 â3 ): perfect monotonic control. Test accuracy peaked atα= 0.25, exactly the kind of inverted-U you would expect from a capacityâmemorization tradeo. 2.2 Human probabilistic reversal learning: extended methods 2.2.1 Dataset and participants We used trial-level data from Ecksteinet al.(2022) 1 : 292 participants, ages 8â30, performing a probabilistic reversal learning task. The data are on OSF athttps://osf.io/7wuh4/. 2.2.2 Task structure On each trial, participants picked one of two options and got binary feedback (reward or no reward). One option paid o 75% of the time, the other 25%. These probabilities ipped 4â9 times per session (125â131 trials total). Because the âgoodâ option still fails a quarter of the time, the eective noise rate works out to 17.1% â lower than the nominal 25%, since the reward asymmetry partly disambiguates the contingency. 2.2.3 Reversal-locked analysis Around each reversal we extracted a window from 8 trials before to 25 after. Within each window we computed two rolling accuracies (Gaussian-smoothed,Ï= 2trials): âąTruth accuracyâ fraction of choices matching the objectively better (post-reversal) option âąFeedback accuracyâ fraction of choices that happened to receive positive feedback The gap is the dierence: feedback accuracy minus truth accuracy, trial by trial. We averaged across reversals within each subject rst, then across subjects. 2.2.4 AUG pos computation We wanted a single number capturing how much each subject over-committed to feedback after a reversal.AUG pos is the baseline-corrected area under the positive portion of the gap curve from trial 0 to trial 25: AUG pos = 1 25 25 â t=0 max(0,gap(t)âgap baseline )(20) 5 The baseline (gap baseline ) is the average gap in the 8 pre-reversal trials, which removes any tonic oset that was already present before the contingency switched. 2.3 Human reward/punishment learning with EEG: extended methods 2.3.1 Dataset and participants The dataset comes from Stolz, Endres & Mueller (2022) 2 : 26 healthy adults recorded with 32-channel EEG while doing separate reward and punishment probabilistic learning tasks (OpenNeuro ds004295). We dropped one subject whose EEG recording cut out mid-session, leavingN= 25for behavior andN= 23for RL model ts. Two more subjects lost too many epochs to artifact rejection in the reward condition (one in punishment), so the nal neural samples areN REW = 24andN PUN = 25. 2.3.2 EEG preprocessing Raw.setles were read into MNE-Python (v1.6+). We bandpass-ltered at 1â40 Hz (zero-phase FIR), re-referenced to the average, cut feedback-locked epochs fromâ200to+600ms, subtracted the pre-stimulus baseline (â200to 0 ms), and threw out any epoch where the voltage swing exceeded 150ÎŒV. Nothing unusual here â the pipeline is deliberately standard so the interesting part is what comes next. 2.3.3 Neural features We picked three feedback-locked features, all of which have solid literatures linking them to outcome processing: 1. Theta power (4â8 Hz) at FCz.Morlet wavelets (n cycles = 4), averaged 200â400 ms post- feedback. Midline frontal theta is the canonical EEG correlate of prediction-error signaling and feedback monitoring. 2.Frontal beta (13â30 Hz).Same time window, averaged across F3, F4, Fz, FC1, FC2 (Morlet, n cycles = 7). Beta in this region has been tied to decision condence and motor planning. 3. P300 at Pz.Mean ERP voltage from 250â450 ms â a textbook marker of context updating after informative outcomes. All three werez-scored within subject and task before entering the decoder. 2.3.4 Neural gap decoder We trained two logistic regressions per subject per task, using 3-fold cross-validation: âą Afeedback decoderthat predicts whether feedback was positive or negative from the three EEG features âą Atruth decoderthat predicts whether the subject expected their choice to be correct before feedback (pre-feedback expectation), from the same features On held-out folds, the two decoders produce probability estimates. The neural gap on trialiis just the dierence: neural_gap(i) =P(feedback=positive|EEG i )âP(expected correct|EEG i )(21) NeuralAUG pos is the mean ofmax(0,neural_gap(i))across trials â how much, on average, the brain over-represents feedback relative to expectation. 6 2.3.5 Source localization To get a rough sense of where the signal lives on the scalp, we computed the positive-minus-negative feedback ERP dierence in the 200â400 ms window at each electrode, averaging across theN= 10 subjects with enough trials in both valence conditions. We split electrodes into a frontal group (F3, F4, F7, F8, Fz, FC1, FC2, FC5, FC6, FCz, AF3, AF4, Fp1, Fp2) and a posterior group (P3, P4, P7, P8, Pz, PO3, PO4, O1, O2, Oz). Both tasks showed a frontal-greater-than-posterior pattern, consistent with the well-known role of medial frontal cortex in feedback evaluation. 2.3.6 Cross-modal correlation At the subject level, we correlated each personâs behavioral commitment (win-stay minus lose-shift) with their neuralAUG pos using SpearmanâsÏ. The correlations were signicant in both tasks: rewardÏ= 0.43,p= 0.036; punishmentÏ= 0.42,p= 0.042(Fig. 5c). People whose brains over- represented feedback more strongly also showed more behavioral over-commitment â the amplication is concordant across levels. At the trial level, we smoothed both the neural gap and the behavioral gap with a 15-trial moving average, then correlated them (Pearsonr) within each subject and averaged. Reward:r= 0.67, p <0.001; punishment:r= 0.53,p <10 â178 . 7 3 Supplementary Results 3.1 Complete statistical summary Table 2: Summary of all primary statistical results across experiments. AnalysisNMeasureStatisticp-valueEect size Human probabilistic reversal learning (Eckstein 2022) PRL AUG pos 2920.049±0.032t(291) = 26.4 1.1Ă10 â79 d= 1.55 PRLT â recovery2929.9±13.9trials 99.3%â Human reward/punishment learning (Stolz 2022) REW commitment250.567±0.171t(24) = 16.6 5.9Ă10 â15 d= 3.32 PUN commitment250.628±0.160t(24) = 19.6 1.4Ă10 â16 d= 3.92 REW win-stay250.799â PUN win-stay250.829â REW lose-shift250.232â PUN lose-shift250.201â Neural gap analysis (EEG) REW neural AUG pos 240.041±0.040100% positive âd= 1.03 PUN neural AUG pos 250.059±0.033100% positive âd= 1.82 Cross-modal correlations REW behavâneural 24Ï= 0.43Spearman0.036Ï= 0.43 PUN behavâneural 24Ï= 0.42Spearman0.042Ï= 0.42 ML alpha-grid causal intervention AUG norm vsα200 runsmonotonicÏ= 1.0<10 â3 â T â vsα200 runsmonotonicÏ= 1.0<10 â3 â 3.2 Reinforcement learning model parameters We tted standard RescorlaâWagner models to each participant in the Stolz (2022) dataset, with a single learning rate and softmax choice rule: Q t+1 (a) =Q t (a) +α·(r t âQ t (a))(22) P(a t =a) = exp(ÎČ·Q t (a)) â a âČ exp(ÎČ·Q t (a âČ )) (23) AcrossN= 23subjects, the tted parameters were: âąREW:α= 0.439±0.314,ÎČ= 3.851±1.822, decay= 0.607±0.365 âąPUN:α= 0.431±0.278,ÎČ= 4.137±1.222, decay= 0.582±0.374 The reward and punishment parameters are strikingly similar, which ts with the parallel behavioral and neural signatures we observe across conditions. 3.3 Positioning relative to noise-robust machine learning We are not trying to beat existing noise-robust training methods. There is already a large toolbox for that: 8 âąLoss correction 3;4 â forward/backward correction, condent learning â these adjust the loss to account for estimated noise transitions. âąSample selectionâ Co-teaching, DivideMix, and relatives that identify and downweight suspicious examples during training. âą Implicit regularizationâ MixUp, AugMax, label smoothing â approaches that reduce memorization as a side eect of data augmentation or soft targets. Our question is dierent: nothowto x label noise butwhenandwhyarchitectural changes alter memorization dynamics. The sparse-residual network is a probe, not a proposed solution. What matters for the paper is the monotonic, causal link between capacity and gap dynamics â and the fact that the same measurement framework extends to human learning. 9 4 Supplementary Figures â1001020 Trials from reversal â0.025 0.000 0.025 0.050 0.075 0.100 0.125 Î gap (baseline-corrected) a w = 4 w = 6 w = 8 w = 10 w = 12 4681012 Smoothing window (trials) 0.00 0.01 0.02 0.03 0.04 0.05 AUG_pos (mean) b canonical (w=8) 4681012 Smoothing window (trials) â0.006 â0.005 â0.004 â0.003 â0.002 â0.001 0.000 Post-reversal slope c 0.0 0.2 0.4 0.6 0.8 1.0 Fraction positive Figure 1:Supplementary Figure 1|PRL window sensitivity analysis.Robustness of the feedbackâtruth gap to rolling window size in the probabilistic reversal learning dataset (N= 292). Gap magnitude (AUG pos ) and temporal structure remain qualitatively consistent across window sizes from 4 to 16 trials, conrming that the observed dynamics are not artifacts of the chosen smoothing parameter. Main text uses 8-trial window (highlighted). References [1] Eckstein, M. K., Master, S. L., Dahl, R. E., Wilbrecht, L. & Collins, A. G. Reinforcement learning and Bayesian inference provide complementary models for the unique advantage of adolescents in stochastic reversal.Developmental Cognitive Neuroscience55, 101106 (2022). URLhttps://linkinghub.elsevier.com/retrieve/pii/S1878929322000494. [2] Stolz, C., Pickering, A. & Mueller, E. M. Reward gain and punishment avoidance reversal learning (2022). URLhttps://openneuro.org/datasets/ds004295/versions/1.0.0. [3] Patrini, G., Rozza, A., Menon, A. K., Nock, R. & Qu, L. Making Deep Neural Networks Robust to Label Noise: A Loss Correction Approach. In2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2233â2241 (IEEE, Honolulu, HI, 2017). URLhttp: //ieeexplore.ieee.org/document/8099723/. [4] Natarajan, N., Dhillon, I. S., Ravikumar, P. & Tewari, A. Learning with noisy labels. In Proceedings of the 27th International Conference on Neural Information Processing Systems - Volume 1, NIPSâ13, 1196â1204 (Curran Associates Inc., Red Hook, NY, USA, 2013). Event-place: Lake Tahoe, Nevada. 10