Paper deep dive
On Emergences of Non-Classical Statistical Characteristics in Classical Neural Networks
Hanyu Zhao, Yang Wu, Yuexian Hou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 7:41:13 AM
Summary
The paper introduces Non-Classical Network (NCnet), a classical neural architecture that exhibits non-classical statistical behaviors analogous to quantum Bell inequality violations. The study demonstrates that the CHSH statistic S, which measures these correlations, arises from gradient competition among shared hidden-layer neurons in multi-task learning. Key findings include that S increases with model resources, temporarily exceeding the classical bound of 2 in a critical regime of insufficient capacity, and correlates positively with generalization performance. The authors validate this using both a simplified XOR-based network and complex models like Multilingual BERT with LoRA adapters.
Entities (10)
Relation Signals (8)
S Statistic → measures → non-classicality
confidence 96% · non-classicality, measured by the S statistic of CHSH inequality
NCnet → exhibits → non-classical statistical behaviors
confidence 95% · NCnet... stably exhibits non-classical statistical behaviors under typical and interpretable experimental setups.
Gradient Competition → causes → non-classicality
confidence 94% · non-classicality... arises from gradient competitions of hidden-layer neurons shared by multi-tasks.
NCnet → uses → CHSH Inequality
confidence 93% · We present the first approach that maps the CHSH statistic S onto multi-task models
S Statistic → exceeds → classical upper-bound
confidence 92% · S may temporarily exceed 2... As resources continue to grow, S then asymptotically decays down to and fluctuates around 2.
S Statistic → correlateswith → generalization performance
confidence 91% · S is positively correlated with generalization performance
Multilingual BERT → isusedin → Real-World Experiments
confidence 90% · We adopt Multilingual BERT (mBERT)... as the base models... This experiment aims to further verify that non-classical features may also emerge within complex neural networks
LoRA → modulates →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Inspired by measurement incompatibility and Bell-family inequalities in quantum mechanics, we propose the Non-Classical Network (NCnet), a simple classical neural architecture that stably exhibits non-classical statistical behaviors under typical and interpretable experimental setups. We find non-classicality, measured by the $S$ statistic of CHSH inequality, arises from gradient competitions of hidden-layer neurons shared by multi-tasks. Remarkably, even without physical links supporting explicit communication, one task head can implicitly sense the training task of other task heads via local loss oscillations, leading to non-local correlations in their training outcomes. Specifically, in the low-resource regime, the value of $S$ increases gradually with increasing resources and approaches toward its classical upper-bound 2, which implies that underfitting is alleviated with resources increase. As the model nears the critical scale required for adequate performance, $S$ may temporarily exceed 2. As resources continue to grow, $S$ then asymptotically decays down to and fluctuates around 2. Empirically, when model capacity is insufficient, $S$ is positively correlated with generalization performance, and the regime where $S$ first approaches $2$ often corresponding to good generalization. Overall, our results suggest that non-classical statistics can provide a novel perspective for understanding internal interactions and training dynamics of deep networks.
Tags
Links
- Source: https://arxiv.org/abs/2603.04451v1
- Canonical: https://arxiv.org/abs/2603.04451v1
Trouble viewing inline? Open PDF directly →
Full Text
42,452 characters extracted from source content.
Expand or collapse full text
On Emergences of Non-Classical Statistical Characteristics in Classical Neural Networks Hanyu Zhao ∗ , Yang Wu ∗ and Yuexian Hou † Tianjin University, China hanyzhao, yangwu, yxhou@tju.edu.cn Abstract Inspired by measurement incompatibility and Bell- family inequalities in quantum mechanics, we pro- pose the Non-Classical Network (NCnet), a sim- ple classical neural architecture that stably exhibits non-classical statistical behaviors under typical and interpretable experimental setups. We find non- classicality, measured by the S statistic of CHSH inequality, arises from gradient competitions of hidden-layer neurons shared by multi-tasks. Re- markably, even without physical links supporting explicit communication, one task head can implic- itly sense the training task of other task heads via local loss oscillations, leading to non-local corre- lations in their training outcomes. Specifically, in the low-resource regime, the value of S increases gradually with increasing resources and approaches toward its classical upper-bound 2, which implies that underfitting is alleviated with resources in- crease. As the model nears the critical scale re- quired for adequate performance, S may temporar- ily exceed 2. As resources continue to grow, S then asymptotically decays down to and fluctuates around 2. Empirically, when model capacity is in- sufficient, S is positively correlated with general- ization performance, and the regime where S first approaches 2 often corresponding to good gener- alization. Overall, our results suggest that non- classical statistics can provide a novel perspective for understanding internal interactions and training dynamics of deep networks. 1 Introduction Driven by the concurrent growth of model scale, data availability, and computational resources, Transformer-based architectures( [ Devlin et al., 2019; Liu et al., 2019; Achiam et al., 2023; Dubey et al., 2024 ] ), particularly large lan- guage models(LLMs), have shown strong capabilities in natu- ral language understanding( [ Wei et al., 2025; Liu et al., 2024; Zhu et al., 2025 ] ), generation( [ Mondshine et al., 2025; ∗ Equal contribution. † Corresponding author. Savoldi et al., 2025; Kulkarni et al., 2024 ] ), reasoning( [ Dong et al., 2024; Wang et al., 2023 ] ), and cross-task transfer. However, the rapid expansion of model capabilities has also made reliable model evaluation an increasingly prominent challenge. On the one hand, traditional evaluation paradigms rely primarily on single-task benchmarks and performance metrics( [ Papineni et al., 2002; Van Rijsbergen, 1979; Baner- jee and Lavie, 2004 ] ), in which score improvements do not necessarily correspond to genuine gains in overall model ca- pability. On the other hand, as application scenarios continue to evolve, evaluation increasingly depends on performance after fine-tuning or instruction alignment. More critically, such metrics typically characterize only the statistical prop- erties of individual tasks and do not provide effective means for assessing internal feature representations or inter-task re- lationships. Dynamic multi-task performance evaluation is closely aligned with the core requirements of artificial general intelligence (AGI), which posits that a general-purpose agent should be able to transfer flexibly across diverse tasks and their compositions. In practice, such transitions often neces- sitate partial retraining, typically via parameter fine-tuning. From the perspective of fine-tuning( [ Devlin et al., 2019; Houlsby et al., 2019; Li and Liang, 2021 ] ), when a model is adapted to different tasks, its parameter space often ex- hibits systematic variations, indicating that each fine-tuned model effectively corresponds to a distinct measurement con- text within the same semantic space. However, these con- texts do not necessarily admit a common global parameter configuration that can satisfy them simultaneously. Differ- ent tasks impose optimization objectives( [ Sener and Koltun, 2018 ] ) on representations and decision boundaries that may conflict within shared parameter dimensions, such that pa- rameter updates optimized for one task inevitably degrade performance on another. This phenomenon admits a natu- ral analogy to measurement incompatibility( [ Von Neumann, 2013; Heisenberg, 1927 ] ) in quantum physics, where non- commuting observables cannot be simultaneously determined and consequently lead to violations of constraints that are oth- erwise satisfiable under classical assumptions, such as Bell inequalities( [ Bell, 1964 ] ). Machine learning techniques have also introduced new an- alytical paradigms for the study of non-classical correlations. Tam ́ as Kriv ́ achy et al.( [ Kriv ́ achy et al., 2020 ] ) employed neu- ral networks to reconstruct classical strategies in order to arXiv:2603.04451v1 [cs.LG] 27 Feb 2026 verify nonlocal correlations in quantum networks. Zhang et al.( [ Zhang et al., 2024 ] )developed a detection framework for non-classical correlations in trilocal networks using stacked deep neural networks. However, the effectiveness of exist- ing approaches relies on an implicit assumption that, in the absence of explicit connections, classical feedforward neural networks cannot generate non-classical correlations between outputs, as their operations are describable by local hidden- variable models. We recognize that these assumptions have limitations and propose NCnet, which is designed to stably exhibit non-classical statistical behaviors under specific and interpretable experimental settings. In addition, we intro- duce the CHSH statistic as an externally observable diagnos- tic tool, providing a new perspective for understanding in- ternal interaction structures in deep neural networks and for evaluating the performance of large scale models. Our con- tributions are as follows : • Methodological innovation: We present the first ap- proach that maps the CHSH statistic S onto multi-task models, enabling a quantitative characterization of task cooperation and competition via the perspective of non- classical statistical analysis. • Architectural Contribution: We introduce NCnet, a classical neural architecture comprising a shared hid- den layer and dual task-specific heads, which consis- tently exhibits non-classical statistical behavior under well-defined and reproducible experimental conditions. • Mechanistic Insight: We demonstrate that the violation of the CHSH inequality does not arise from any explicit information channel, but is instead driven by gradient competition induced by shared parameters in multi-task learning. Moreover, we show that this phenomenon is most pronounced in a critical regime where model ca- pacity is nearly sufficient yet not redundant. 2 Preliminaries 2.1 Local Realism and Local Hidden Variable Models Locality implies that two spacelike-separated regions are in- formationally isolated during the considered time interval, as no signal can exceed the speed of light, thereby excluding any form of instantaneous “action at a distance”. Realism holds that the properties of a physical system, such as position, momentum, or spin, exist objectively prior to measurement and independently of it. Taken together, locality and realism constitute local realism( [ Einstein et al., 1971; Lalo ̈ e, 2019; Laudisa, 2023 ] ), which stipulates that events in one region cannot immediately affect the physical reality of another, and that any causal influence must respect the finite speed of prop- agation. Within the historical debates surrounding the completeness of quantum mechanics, the framework of local hidden vari- able (LHV) models( [ Einstein et al., 1935; Bell, 1964 ] ) was proposed under the premise of local realism, representing an attempt to account for quantum phenomena within a classical worldview. Concretely, for a particle p and an observable O, the measurement outcome is fully determined by the observ- able O together with an intrinsic property λ of the particle. In other words, when measuring O on particle p, the outcome is represented as O(λ), where λ denotes the local hidden vari- able associated with p. Within the feedforward neural networks framework con- sidered in this paper, two top-level neurons between which no (explicit) information link exists are regarded as mutually local (i.e., informationally isolated). In this setting, event in- formation occurring at one neuron cannot be transmitted to the other within the time interval under consideration. This definition is logically consistent with the notion of locality as defined in the philosophy of physics. 2.2 The Classical Upper-Bound of the CHSH Statistic To examine whether there exists a fundamental conflict be- tween quantum mechanics and local realism, John Bell pro- posed the correlation-based Bell inequalities( [ Bell, 1964 ] ). Clauser, Horne, Shimony, and Holt introduced the CHSH inequality( [ Clauser et al., 1969 ] ), a specific formulation amenable to experimental verification. Consider four binary observable variables A i ,B j ∈ −1, +1, for i,j ∈ 0, 1, where A 0 and A 1 denote the measurement outcomes of one party, Alice, and B 0 and B 1 denote the outcomes of the other party, Bob. The experimen- tal setup satisfies the basic conditions of the CHSH inequal- ity: each party independently selects an observable variable in each trial, producing binary outcomes that yield four sets of correlation data. Theorem 1 (Classical Upper-Bound of the CHSH Statistic). For any local hidden-variable theory, the CHSH value S sat- isfies |S| =|C A 0 B 0 + C A 0 B 1 + C A 1 B 0 − C A 1 B 1 |≤ 2(1) where C A i B j =E[A i B j ] is an expectation statistic indi- cates the association between the measurement outcomes of Alice and Bob. If supposing that A i and B j are centered and normalized, it turns out that C A i B j simply becomes Person’s correlation coefficient. The detailed proof and the corresponding derivation within LHV framework can be found in the literature( [ Brunner et al., 2014 ] ). The CHSH inequality provides an operational crite- rion to distinguish between LHV models and non-classical probabilistic correlations. Theoretically, this inequality must hold for LHV models, as well as for all classical probabilistic models grounded in the assumptions of locality and realism. 3 Related Work Neural Network–Based Representation and Quantifica- tion of Non-Classical Correlations. Machine learning has emerged as a powerful analytical paradigm for the study of non-classical correlations. Askery Canabarro et al.( [ Can- abarro et al., 2019 ] ) developed hybrid models combining multilayer perceptrons with genetic algorithms to achieve high-precision quantification of non-classical correlations in Bell scenarios and to discover novel correlation structures. +1 -1 -1 -1 +1 +1 -1 -1 +1 ReLU ReLUReLU X 1 X2 (a) XORnet X 1 X2X3X4 XORnetXORnet a i b j (b) NCnet Figure 1: The framework of the proposed NCnet. (a) XORnet illustrates the basic network structure for modeling the XOR function with ReLU activations. (b) NCnet is constructed by integrating two XORnets. The red node denotes a shared neuron where gradient competition is prone to occur. Luo et al. ( [ Luo, 2018 ] ) constructed nonlinear Bell inequal- ities based on bipartite graph matching to characterize mul- tipartite correlations in quantum networks. There also exist some works( [ Kriv ́ achy et al., 2020; Polino et al., 2023 ] ) are typically grounded in the assumption that if the correlations present in the data can be perfectly modeled by classical neu- ral networks, then the data contain only classical correlations; conversely, if these correlations cannot be captured by those models, the data are taken to exhibit non-classical correla- tions. Table 1: Parameter Mapping Between NCnet and Bell Experiments. NCnet ParameterBell Experiment Parameter Alice’s task α i Alice’s measurement basis i Bob’s task β j Bob’s measurement basis j Consistencyindicator A i (B j ) Measurement result A i (B j ) Correlation of measurement outcomes C(A i ,B j ) Correlation C A i B j 4 Methodology To more clearly demonstrate the emergence of non-classical statistical phenomena in neural networks, we propose a Non- Classical Network (NCnet) derived from the XORnet. In this paper, the “non-classical network” is defined as a classical neural network that can robustly exhibit non-classical statis- tical features. As shown in Figure 1, this architecture pro- vides an insightful and controlled framework for analyzing non-classical phenomena. 4.1 Framework of NCnet and Task Definition NCnet takes four binary inputs, X 1 through X 4 , after which the network produces two output decisions corresponding to the measurement outcomes at Alice’s and Bob’s sides. We define the tasks on Alice’s side and Bob’s side as follows: Alice : α 1 = X 1 α 2 = X 1 ⊕ X 2 , (2) Bob : β 1 = X 3 β 2 = X 3 ⊕ X 4 (3) In this setup, α 1 and β 1 perform identity mappings, while α 2 and β 2 realize XOR logical operations. A i (B j ) denotes the binary measurement outcome as the consistency indica- tor for task α i (β j ), taking value +1 if NCnet’s prediction is correct and −1 otherwise (similarly for B j ). For each task pair (α i ,β j ), the correlation of their measurement outcomes is defined as C(A i ,B j ) = ⟨A i · B j ⟩,i,j ∈ 1, 2, where⟨·⟩ denotes the mathematical expectation, approximated by the arithmetic average in the experiment. The correspondence between NCnet’s prediction tasks and the Bell experimen- tal parameters is summarized in Table 1. Finally, the CHSH statistic S defined as S = C(A 1 ,B 1 )+C(A 1 ,B 2 )+C(A 2 ,B 1 )−C(A 2 ,B 2 ) (4) Indeed, if the random variables A i and B j both have a mean of 0 and a standard deviation of 1, then the formula for their correlation C(A i ,B j ) reduces to ρ A i B j i.e., Pear- son correlation coefficient. If the experiment yields S > 2, it indicates the presence of non-classical correlations in the system that cannot be explained by any LHV model. 4.2 CHSH Inequality Violation Based on NCnet To verify the emergence of non-classical features in NC- net under specific conditions, we trained four models corre- sponding to the four task combinations formulated in Equa- tions (2) and (3), and calculated the statistic S. Furthermore, to analyze the impact of model capacity on S, we consid- ered three scenarios with the number of hidden neurons set to n = 2, 3, 4. All experiments were repeated 50 times to ensure stability and reliability. As shown in Figure 2, the statistic S exhibits a pronounced nonlinear dependence on the number of hidden units n. When n = 2, the values of S are generally low: in most runs S < 1.5, which is well below the LHV bound of 2.0. When n = 3, S attains its maximum. Almost all repeated runs yield values that significantly exceed the LHV upper-bound. In some cases, the observed values even surpass the Tsirelson bound( [ Cirel’son, 1980 ] ) 2 √ 2 ≈ 2.828, reaching approxi- mately S ≈ 3.5. When n = 4, S decreases and stabi- lizes around 2, and the apparent violation of CHSH inequal- ity disappears. In summary, the nonlinear relationship be- tween S and the number of hidden layer neurons n may rep- resent a universally applicable phenomenon. Specifically, the (a) Scatter Distribution of S(b) Average Correlation Terms and S Figure 2: Overall results of NCnet with different hidden-layer sizes. (a) Scatter distribution of S obtained from 50 independent runs of NCnet for hidden-layer sizes n = 2, 3, 4. The red dashed line indicates the classical upper-bound of the CHSH statistic, and the blue dot-dashed line marks the Tsirelson bound. (b) Mean correlation values C(A i ,B j ) and the corresponding average statistic S for each value of n. most pronounced non-classical features emerge near a criti- cal regime in which the network’s representational capacity is nearing an adequate level, while remaining insufficient. 4.3 Causes of Non-Classical Features In the experiments described above, when the number of neu- rons was set to three, the statistic S exceeded the LHV upper- bound in all 50 repeated trials. Formally, this behavior is reminiscent of “nonlocal” correlations. When Alice and Bob perform any task combination other than α 2 and β 2 , the hid- den layer of NCnet is capable of representing the features re- quired by all necessary tasks; that is, there exists at least one set of weights that allows NCnet to achieve perfect training performance on these combinations. In contrast, when Alice and Bob perform the task combination α 2 and β 2 , the hidden layer of NCnet fails to capture at least one of the necessary features, meaning that no set of weights can be found that enables NCnet to train perfectly on this combination. Under this condition, the first three terms of the CHSH statistic S are close to 1, whereas the last term is far smaller than 1, which leads to the emergence of non-classical statistical features. Here the fundamental reason for non-classicality is that the top-layer neurons realize implicit communication via gradi- ent competition in the absence of an explicit information- transmission pathway. Gradient competition occurs during backpropagation when a hidden-layer neuron receives con- flicting gradient updates from two top-layer neurons. Note that gradient competition is inevitable if the number of neu- rons in the feature representation layer is insufficient. Persis- tent gradient competition induces oscillations in the loss func- tions associated with the top-layer neurons, impeding conver- gence. In principle, an observer confined to a single top-layer neuron could infer that the other top-layer neuron is attempt- ing to optimize a more difficult training task by monitoring oscillations in its local training cost function. This constitutes the mechanism by which implicit communication emerges. This finding carries significant research implications. Given that feedforward neural networks constitute fundamen- tal building blocks of modern deep learning systems and are widely employed across diverse tasks and architectures, it is reasonable to hypothesize that non-classical statistical fea- tures may be pervasive in contemporary deep learning models. Moreover, Bell inequalities and their general- ized forms (i.e., the family of Bell inequalities) provide a theoretical foundation for establishing a novel analyti- cal framework for neural networks. This framework not only reveals implicit coupling relationships among different tasks or modules but also serves as a valuable complement to conventional methods of performance evaluation, offering a new analytical perspective for understanding the internal representational mechanisms and training dynamics of neural networks. 5 Real-World Experiments This experiment aims to further verify that non-classical fea- tures may also emerge within complex neural networks, and to explore the introduction of the CHSH statistic as a novel model evaluation metric. This metric serves to quantita- tively assess a neural network’s representational capacity and generalization performance. Unlike the earlier model based on NCnet, this section focuses on multi-layer architectures that more closely resemble real-world task scenarios, such as contemporary large language models. The investigation examines how non-classical correlations manifest in higher- dimensional and more complex learning tasks. 5.1 Experiments Setup Model Architecture and Datasets. We adopt Multilingual BERT (mBERT) ( [ Pires et al., 2019 ] ) and BERT as the base models, both consisting of 12 Transformer encoder layers, 12 self-attention heads, and 768-dimensional hidden representa- tions. Following the previously described NCnet architecture, we adopt a standard multi-task learning framework with hard parameter sharing, in which independent task-specific heads are attached to a shared semantic representation layer, corre- sponding to the observational outputs of Alice and Bob. To enable fine-grained control over model capacity and support efficient multi-task adaptation, we inject trainable low-rank matrices( [ Hu et al., 2022 ] ) into the query and value mod- ules of the self-attention mechanism, adjusting the rank r ∈ 1, 2, 4, 8, 16, 32, 64, 128, 256, 512 to modulate the effec- tive capacity. This allows us to systematically examine how model size influences coupling and competition among tasks during training. Throughout this paper, model capacity refers Table 2: Accuracy Acc A i (Acc B j ), Correlation C(A i ,B j ), and CHSH Statistics S for Multilingual Training and Mixed Reasoning Tasks across Varying LoRA Ranks r. rTask Multilingual TrainingMixed Reasoning Tasks Acc A i Acc B j C(A i ,B j )SAcc A i Acc B j C(A i ,B j )S 1 A 1 B 1 95.44%91.31%0.7501 1.502 99.68%99.82%0.9899 1.902 A 1 B 2 95.65%89.34%0.719298.92%76.16%0.5144 A 2 B 1 93.68%91.28%0.720281.79%98.88%0.6197 A 2 B 2 93.80%89.18%0.687179.89%68.91%0.2219 2 A 1 B 1 97.32%94.67%0.8453 1.713 99.92%99.96%0.9976 2.191 A 1 B 2 97.26%94.41%0.844599.46%93.89%0.8670 A 2 B 1 96.31%95.35%0.839796.71%99.66%0.9299 A 2 B 2 96.29%93.91%0.816193.33%84.58%0.6038 4 A 1 B 1 98.79%97.83%0.9335 1.870 99.98%99.99%0.9996 2.103 A 1 B 2 98.75%97.21%0.920599.94%97.38%0.9463 A 2 B 1 97.97%97.64%0.914199.58%99.96%0.9908 A 2 B 2 97.96%96.82%0.898397.48%93.77%0.8337 8 A 1 B 1 99.50%98.88%0.9678 1.931 99.98%99.99%0.9995 2.001 A 1 B 2 99.32%98.40%0.954799.90%99.54%0.9887 A 2 B 1 98.98%98.79%0.955799.88%99.98%0.9971 A 2 B 2 98.87%98.50%0.948399.86%99.34%0.9839 16 A 1 B 1 99.51%98.90%0.9692 1.938 99.72%99.62%0.9867 2.020 A 1 B 2 99.44%98.73%0.963399.96%99.86%0.9963 A 2 B 1 98.97%98.84%0.957199.98%99.98%0.9995 A 2 B 2 98.94%98.60%0.951599.44%98.68%0.9623 32 A 1 B 1 99.94%99.68%0.9931 1.985 99.66%99.56%0.9844 2.021 A 1 B 2 99.96%99.66%0.990399.96%99.76%0.9944 A 2 B 1 99.72%99.60%0.986398.72%99.64%0.9671 A 2 B 2 99.64%99.70%0.986798.78%97.38%0.9247 to the total number of trainable parameters introduced by the LoRA modules together with those in the fully connected lay- ers of the task-specific heads. The specific number of train- able parameters for r ranging from 1 to 32 is provided in Ta- ble 4. In the following experiments, we construct a multi- task setting using the PAWS-X( [ Yang et al., 2019 ] ), SST- 2( [ Wang et al., 2018 ] , MRPC, CommonsenseQA( [ Talmor et al., 2019 ] ), and MathQA( [ Amini et al., 2019 ] ) datasets. The specific task assignments for Alice and Bob are summarized in the following Table 3. Table 3: Different Task Combinations of Experiments TaskMultilingual TrainingMixed Reasoning Tasks A 1 PAWS-X-enSST-2 A 2 PAWS-X-frCommonsenseQA B 1 PAWS-X-jaMRPC B 2 PAWS-X-koMathQA Statistical Definitions. Within this framework, we adopt the previously defined correlation measure C(A i ,B j ) and the CHSH statistic S, applying them to the subsequent ex- perimental analysis. The measure C(A i ,B j ) quantifies task compatibility: values approaching +1 indicate strong positive correlation between tasks. Conversely, when the value sig- nificantly falls below 1, it implies competition or interference between tasks. The absolute value of the CHSH statistic S can serve as a sufficient criterion for the presence of non- Figure 3: The trend of the CHSH statistic S as a function of the rank r. The plot includes two sets of experimental results: the blue curve (Multilingual Training) and the orange curve (Mixed Reason- ing Tasks). The red dashed line represents the classical upper-bound S = 2. classicality. When|S| significantly exceeds 2, it provides ev- idence for the robust existence of non-classical correlations. 5.2 Experiments Results We trained models based on the settings in Table 3 to com- pute the measurement correlations C(A i ,B j ) and the CHSH statistic S. Detailed numerical results are provided in Ta- ble 2. Figure 3 illustrates the variation of the CHSH statistic S with the LoRA rank r. In the Mixed Reasoning Tasks, we observe a significant violation of the classical upper-bound (a) Evolution of S over Training Epochs(b) Convergence Rate Figure 4: The CHSH statistic S convergence across different ranks under Mixed Reasoning Tasks. (a) Higher ranks r lead to faster conver- gence of S. (b) The bar chart reports μ ∇S , the arithmetic mean of the instantaneous slopes of S over epochs 0–80; larger bars indicate a faster convergence rate. (S = 2) at low ranks (r = 2 and r = 4), where S reaches its peak. As r increases beyond 4, S declines and asymp- totically converges to 2. In contrast, within the Multilingual Training scenario, the value of S increases monotonically as r increases and gradually converges to 2, with no instances significantly exceeding 2 observed. In this setting, the four task combinations exhibit relatively balanced difficulty lev- els, which suggests that non-classical correlations induced by gradient competition during multi-task training may be mod- erated. This hypothesis is clearly supported by the results in the low-rank region of Table 2, where the C(A i ,B j ) values across different task combinations remain numerically close. This explains why S does not significantly exceed 2. The experimental results on Mixed Reasoning Tasks demonstrate that, because the difficulty of different task com- binations varies, classical neural networks for real-world tasks can exhibit pronounced non-classical statistical charac- teristics, significantly violating the CHSH inequality. This observation suggests that the non-classicality captured by the CHCH statistic S may be pervasive since real-world tasks typically exhibit varying levels of difficulty. Note also that, because S > 2 represents only a sufficient, rather than a necessary, condition for the existence of non-classicality, there may in principle exist alternative criteria for identify- ing non-classical behavior that correspond to different non- classicality generation scenarios. Consequently, this observa- tion offers a new perspective that motivates a re-examination of the implicit assumptions underlying certain models for distinguishing non-classical correlations, particularly the as- sumption that classical neural networks are incapable of gen- erating non-classical correlations. 5.3 Experimental Summary and Discussion Dynamic Analysis of the Training Process In this section, we investigate the relationship between the LoRA rank parameter and the average gradient of the CHSH statistic S during multi-task training. The average gradient is computed as the arithmetic mean of the instantaneous slopes of S across training epochs. As shown in Figure 4, when r = 1, the average gradient of S attains its minimum value, with μ ∇S = 0.0208. As r increases from 2 onward, the average gradient exhibits an overall upward trend. In the r=1r=2r=4r=8 r=16r=32r=64 r=128r=256r=512 0.700 0.725 0.750 0.775 0.800 0.825 0.850 0.875 Accuracy 1.0 1.2 1.4 1.6 1.8 2.0 S Acc (A 1 , B 1 ) Acc (A 1 , B 2 ) Acc (A 2 , B 1 ) Acc (A 2 , B 2 ) Acc combo_avg S Classical upper-bound S = 2 Figure 5: Generalization performance and the CHSH statistic S across ranks r in Multilingual Training. Bars show task-pair mean accuracy Acc(A i ,B j ) = Acc(A i )+Acc(B j ) 2 .The purple curve shows the combination average Acc combavg (the mean of the four bars) per rank. The red curve denotes S value, reflecting the non- classical coupling strength of their learned representations. range r = 8 to r = 32, although the rate of increase di- minishes, the average gradient remains at a relatively high level. Overall, the average gradient of the CHSH statistic S is positively correlated with the trainable parameter capacity in- troduced by LoRA. As r increases, S converges more rapidly during training. A plausible explanation is that higher ranks introduce additional trainable parameters, thereby enhancing the model’s representational capacity and alleviating gradient competition, which in turn improves the efficiency of cross- task gradient propagation. Generalization Ability and Non-classical Correlation Analysis under Different LoRA Ranks In this section, we investigate the relationship between the model’s generalization capability and the CHSH statistic S. The generalization capability is characterized by the model’s testing accuracy. From the Figure 5 , we observe a strong pos- itive correlation between the model’s generalization capabil- ity and the statistical indicator S within the range 1≤ r ≤ 32. The model capacity at which r first approaches 2 corresponds to a scale that is sufficient but not redundant. According to classical statistical learning theory (e.g., the bias-variance trade-off), this specific model capacity is expected to yield near-optimal generalization performance. Table 4: Parameter Statistics for Different LoRA Rank Parameters r Across Experiments r Multilingual TrainingMixed Reasoning Tasks Total Params Trainable Params Trainable Ratio (%) Total Params Trainable Params Trainable Ratio (%) 1177,893,38039,9400.022109,522,18039,9400.036 2 177,930,24476,8040.043109,559,04476,8040.070 4178,003,972150,5320.085109,632,772150,5320.137 8 178,151,428297,9880.167109,780,228297,9880.271 16178,446,340592,9000.332110,075,140592,9000.539 32 179,036,9321183,4920.661110,664,9641,182,7241.070 However, in the Mixed Reasoning task, the model exhibits consistently low average testing accuracy across sub-tasks, with no significant variation as r increases. This phenomenon is attributed to the high complexity and open-ended nature of the tasks. Consequently, testing performance is statistically independent of the model capacity, implying that there is also no statistical correlation between S and generalization capa- bility at the scale of the model under consideration. Resource Competition Hypotheses for Non-classicality Witnessed by the CHSH Statistic S To further elucidate the mechanism by which variations in the LoRA rank parameter influence the CHSH statistic S, we propose the following hypotheses based on our experimental observations. Assumption 1: Different task pairs (A i ,B j ) exhibit het- erogeneous demands for model resources (e.g., trainable pa- rameters). Specifically, certain task pairs achieve effective coordination with limited resources, whereas others require substantially higher parameter capacity to converge. Assumption 2: Under resource-limited conditions, sub- tasks within a task pair compete for shared resources. This competition increases the possibility that the product A i · B j takes negative values, thereby reducing the magnitude of C(A i ,B j ). Assumption 3: When resources become abundant, inter- task competition is substantially mitigated. Consequently, A i and B j tend to take positive values simultaneously, driving C(A i ,B j ) close to 1. Based on the above hypothesis and experimental results, we draw the following conclusion: • When S ≪ 2 1 , the model is underfitting and per- forms poorly across nearly all task combinations, in- dicating insufficient model capacity or limited ability to form shared representations. • When S ≈ 2 and C(A i ,B j ) for all task pairs are close to 1, the model has converged well across all combi- nations. This suggests that the trainable parameter ca- pacity is well-matched to the representational demands, potentially even exhibiting redundancy. • When S ≫ 2, the model has well converged across most task pairs but fails to fully converge on one spe- 1 In this paper, the symbols “≪” and “≫” are used in a qualitative sense to indicate that S is significantly smaller or larger than the bound 2. cific combination due to persistent gradient competi- tion. In this regime, salient non-classical statistical fea- tures emerge, suggesting that the model operates near a critical point where representational capacity is ap- proaching sufficient, but remains insufficient. 6 Conclusion In conclusion, this paper introduces NCnet, a classical neu- ral network architecture demonstrated to robustly exhibit non-classical statistical characteristics within the theoretical framework of measurement (probabilistic) incompatibility. The emergence of these non-classical correlations arises from implicit communication between top-layer neurons in the ab- sence of explicit information exchange pathways. An ob- server confined to a single top-layer neuron could infer, solely from oscillations in the neuron’s local training loss, that the other top-layer neuron is engaged in a more challenging opti- mization task. Through extended experiments on complex architectures and real-world datasets, we further validate the generality and stability of this phenomenon. In parallel, we propose employ- ing the CHSH inequality as a useful analytical metric for neu- ral networks. Note that the CHSH statistic S is only one of the numerous statistical measures for deciding the existence of non-classical statistical features. In principle, we can se- lect or construct appropriate deciding statistics based on the specific problem settings of the task at hand, which indicates a general applicability of non-classical statistical methods in the performance analysis of deep learning models. Unlike conventional metrics based on performance or loss, this ap- proach reveals the internal information-coupling structure of neural systems by quantitatively analyzing the statistical be- havior of observable outputs. References [ Achiam et al., 2023 ] Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Alt- man, Shyamal Anadkat, et al.Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023. [ Amini et al., 2019 ] Aida Amini, Saadia Gabriel, Shanchuan Lin, Rik Koncel-Kedziorski, Yejin Choi, and Hannaneh Hajishirzi.Mathqa: Towards interpretable math word problem solving with operation-based formalisms.In North American Chapter of the Association for Compu- tational Linguistics, 2019. [ Banerjee and Lavie, 2004 ] Satanjeev Banerjee and Alon Lavie. Meteor: an automatic metric for mt evaluation with high levels of correlation with human judgments. Proceed- ings of ACL-WMT, pages 65–72, 2004. [ Bell, 1964 ] John Stewart Bell. On the einstein-podolsky- rosen paradox. Physics, 1:195–200, 1964. [ Brunner et al., 2014 ] Nicolas Brunner, Daniel Cavalcanti, Stefano Pironio, Valerio Scarani, and Stephanie Wehner. Bell nonlocality. Reviews of modern physics, 86(2):419– 478, 2014. [ Canabarro et al., 2019 ] Askery Canabarro, Samura ́ ı Brito, and Rafael Chaves. Machine learning nonlocal correla- tions. Physical review letters, 122(20):200401, 2019. [ Cirel’son, 1980 ] Boris S Cirel’son. Quantum generaliza- tions of bell’s inequality. Letters in Mathematical Physics, 4(2):93–100, 1980. [ Clauser et al., 1969 ] John F Clauser, Michael A Horne, Ab- ner Shimony, and Richard A Holt. Proposed experiment to test local hidden-variable theories. Physical review letters, 23(15):880, 1969. [ Devlin et al., 2019 ] Jacob Devlin, Ming-Wei Chang, Ken- ton Lee, and Kristina Toutanova. Bert: Pre-training of deep bidirectional transformers for language understand- ing. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL- HLT 2019), pages 4171–4186, 2019. [ Dong et al., 2024 ] Qingxiu Dong, Lei Li, Damai Dai, Ce Zheng, Jingyuan Ma, Rui Li, Heming Xia, Jingjing Xu, Zhiyong Wu, Baobao Chang, Xu Sun, Lei Li, and Zhi- fang Sui. A survey on in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1107–1128, 2024. [ Dubey et al., 2024 ] Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783, 2024. [ Einstein et al., 1935 ] Albert Einstein, Boris Podolsky, and Nathan Rosen. Can quantum-mechanical description of physical reality be considered complete? Physical review, 47(10):777, 1935. [ Einstein et al., 1971 ] Albert Einstein, Max Born, and Hed- wig Born. The born-einstein letters: Correspondence be- tween albert einstein and max and hedwig born from 1916- 1955, with commentaries by max born. (No Title), 1971. [ Heisenberg, 1927 ] Werner Heisenberg. ̈ Uber den an- schaulichen inhalt der quantentheoretischen kinematik und mechanik. Zeitschrift f ̈ ur Physik, 43(3):172–198, 1927. [ Houlsby et al., 2019 ] Neil Houlsby, Andrei Giurgiu, Stanis- law Jastrzebski, Bruna Morrone, Quentin De Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. Parameter-efficient transfer learning for nlp. In Interna- tional conference on machine learning, pages 2790–2799. PMLR, 2019. [ Hu et al., 2022 ] Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al. Lora: Low-rank adaptation of large language models. ICLR, 1(2):3, 2022. [ Kriv ́ achy et al., 2020 ] Tam ́ as Kriv ́ achy, Yu Cai, Daniel Cav- alcanti, Arash Tavakoli, Nicolas Gisin, and Nicolas Brun- ner. A neural network oracle for quantum nonlocality problems in networks. npj Quantum Information, 6(1):70, 2020. [ Kulkarni et al., 2024 ] Atharva Kulkarni, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Dhivya Piraviperumal, Hong Yu, and Shruti Bhargava. SynthDST: Synthetic data is all you need for few-shot dialog state tracking. In Proceed- ings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1988–2001, 2024. [ Lalo ̈ e, 2019 ] Franck Lalo ̈ e. Do we really understand quan- tum mechanics? Cambridge University Press, 2019. [ Laudisa, 2023 ] Federico Laudisa. How and when did local- ity become ‘local realism’? a historical and critical anal- ysis (1963–1978). Studies in History and Philosophy of Science, 97:44–57, 2023. [ Li and Liang, 2021 ] Xiang Lisa Li and Percy Liang. Prefix- tuning: Optimizing continuous prompts for generation. arXiv preprint arXiv:2101.00190, 2021. [ Liu et al., 2019 ] Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. Roberta: A robustly optimized bert pretraining approach. ArXiv, abs/1907.11692, 2019. [ Liu et al., 2024 ] Zhiwei Liu, Kailai Yang, Qianqian Xie, Tianlin Zhang, and Sophia Ananiadou. Emollms: A se- ries of emotional large language models and annotation tools for comprehensive affective analysis. In Proceed- ings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, pages 5487–5496, 2024. [ Luo, 2018 ] Ming-Xing Luo. Computationally efficient non- linear bell inequalities for quantum networks. Physical Review Letters, 120(14):140402, 2018. [ Mondshine et al., 2025 ] ItaiMondshine,TzufPaz- Argaman, and Reut Tsarfaty.Beyond n-grams: Re- thinking evaluation metrics and strategies for multilingual abstractive summarization. In Proceedings of the 63rd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 19019–19035, 2025. [ Papineni et al., 2002 ] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for au- tomatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Compu- tational Linguistics, pages 311–318, 2002. [ Pires et al., 2019 ] Telmo Pires, Eva Schlinger, and Dan Gar- rette.How multilingual is multilingual bert?arXiv preprint arXiv:1906.01502, 2019. [ Polino et al., 2023 ] Emanuele Polino, Davide Poderini, Giovanni Rodari, Iris Agresti, Alessia Suprano, Gonzalo Carvacho, Elie Wolfe, Askery Canabarro, George Moreno, Giorgio Milani, et al. Experimental nonclassicality in a causal network without assuming freedom of choice. Na- ture Communications, 14(1):909, 2023. [ Savoldi et al., 2025 ] Beatrice Savoldi, Giuseppe Attanasio, Eleonora Cupin, Eleni Gkovedarou, Janic ̧a Hackenbuch- ner, Anne Lauscher, Matteo Negri, Andrea Piergentili, Manjinder Thind, and Luisa Bentivogli. Mind the inclu- sivity gap: Multilingual gender-neutral translation evalua- tion with mgente. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 13709–13731, 2025. [ Sener and Koltun, 2018 ] Ozan Sener and Vladlen Koltun. Multi-task learning as multi-objective optimization. Ad- vances in neural information processing systems, 31, 2018. [ Talmor et al., 2019 ] AlonTalmor,JonathanHerzig, Nicholas Lourie, and Jonathan Berant. Commonsenseqa: A question answering challenge targeting commonsense knowledge. ArXiv, abs/1811.00937, 2019. [ Van Rijsbergen, 1979 ] Cornelis Joost Van Rijsbergen. In- formation retrieval. No Title, 1979. [ Von Neumann, 2013 ] John Von Neumann. Mathematische grundlagen der quantenmechanik, volume 38. Springer- Verlag, 2013. [ Wang et al., 2018 ] Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, 2018. [ Wang et al., 2023 ] Xinyi Wang, Wanrong Zhu, Michael Saxon, Mark Steyvers, and William Yang Wang. Large language models are latent variable models: Explain- ing and finding good demonstrations for in-context learn- ing. Advances in Neural Information Processing Systems, 36:15614–15638, 2023. [ Wei et al., 2025 ] Xiao Wei, Xiaobao Wang, Ning Zhuang, Chenyang Wang, Longbiao Wang, et al. Integration of old and new knowledge for generalized intent discov- ery: A consistency-driven prototype-prompting frame- work. arXiv preprint arXiv:2506.08490, 2025. [ Yang et al., 2019 ] Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge.Paws-x: A cross-lingual adversar- ial dataset for paraphrase identification. arXiv preprint arXiv:1908.11828, 2019. [ Zhang et al., 2024 ] Ying Zhang, Jin-chuan Hou, and Kan He. Detecting nonlocal correlations in chain-shaped quan- tum networks via a deep-learning method. Physical Re- view A, 110(6):062609, 2024. [ Zhu et al., 2025 ] Linlin Zhu, Heli Sun, Xiaoyong Huang, Qi Zhang, Ruichen Cao, and Liang He.Sentiment- enhanced multi-hop connected graph attention network for multimodal aspect-based sentiment analysis. In Proceed- ings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, pages 8393–8401, 2025.