Paper deep dive
Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting
Huiyu Li, Weibo Liu, Xinru Xu, Dongchen Gao, Meng Zhang, Junhua Hu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/14/2026, 6:11:44 AM
Summary
The paper proposes a unified evidence reasoning framework for multi-source decision making under Dempster-Shafer theory (DST). It introduces a Chaos-Conflict Measurement (CCM) to jointly quantify cross-evidence conflict and intra-evidence non-specificity, and a historical-experience driven weighting scheme using spectral clustering and regret theory to compute context-specific reliability profiles. These components feed into a hybrid combination rule that balances uncertainty preservation and weighted consensus, followed by a belief-interval decision strategy. The framework is evaluated on 16 real-world benchmark datasets, outperforming DST-based baselines and gradient boosting methods.
Entities (9)
Relation Signals (8)
Chaos-Conflict Measurement → partof → Unified Evidence Reasoning Framework
confidence 95% · This paper proposes a unified evidence reasoning framework... Specifically, a chaos-conflict measurement is introduced
Historical-Experience Weighting → partof → Unified Evidence Reasoning Framework
confidence 95% · A historical experience driven weighting scheme... These mechanisms feed into a hybrid combination rule
Chaos-Conflict Measurement → quantifies → Cross-evidence conflict
confidence 94% · chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity
Chaos-Conflict Measurement → quantifies → Intra-evidence non-specificity
confidence 94% · chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity
Regret Theory → usedin → Historical-Experience Weighting
confidence 93% · applies regret theory to compute context-specific reliability profiles from past fusion outcomes
Spectral clustering → usedin → Historical-Experience Weighting
confidence 92% · historical experience driven weighting scheme partitions the decision space via spectral clustering
Unified Evidence Reasoning Framework → outperforms → Gradient boosting methods
confidence 90% · outperforming eight DST-based baselines and three gradient boosting methods
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisns without exploiting their long-term reliability across diverse decision contexts. This paper proposes a unified evidence reasoning framework that addresses both limitations. Specifically, a chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity, with five formally proven properties ensuring consistent assessment. A historical experience driven weighting scheme partitions the decision space via spectral clustering and applies regret theory to compute context-specific reliability profiles from past fusion outcomes. These mechanisms feed into a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, controlled by the global conflict level, followed by a belief-interval decision strategy that enables robust classification without discarding epistemic uncertainty. Experiments on 16 real-world benchmark datasets demonstrate that the proposed framework achieves an average F1 score of 85.78 and a mean AUC of 93.30, outperforming eight DST-based baselines and three gradient boosting methods. Ablation analysis confirms the contribution of each component we proposed. The framework offers an effective approach for adaptive evidence fusion in multi-source decision making.
Tags
Links
- Source: https://arxiv.org/abs/2608.13108v1
- Canonical: https://arxiv.org/abs/2608.13108v1
Trouble viewing inline? Open PDF directly →
Full Text
124,807 characters extracted from source content.
Expand or collapse full text
Robust Dempster-Shafer Evidence Fusion with Chaos-Conflict Measurement and Historical-Experience Weighting Huiyu Li1,2 Weibo Liu2 Xinru Xu2 Dongchen Gao2 Meng Zhang2 Junhua Hu1 1School of Business, Central South University, Changsha, Hunan, People’s Republic of China, 410083 2School of Management, Shandong University, Jinan, Shandong, People’s Republic of China, 250100 Corresponding author: Junhua Hu (hujunhua@csu.edu.cn) Abstract Multi-source evidence fusion under Dempster-Shafer theory faces two persistent challenges: existing conflict measures assess inter-evidence inconsistency and intra-evidence uncertainty independently, yielding incomplete evaluations, and current fusion methods evaluate evidence sources exclusively through instantaneous comparisons without exploiting their long-term reliability across diverse decision contexts. This paper proposes a unified evidence reasoning framework that addresses both limitations. Specifically, a chaos-conflict measurement is introduced to jointly quantify cross-evidence conflict and intra-evidence non-specificity, with five formally proven properties ensuring consistent assessment. A historical experience driven weighting scheme partitions the decision space via spectral clustering and applies regret theory to compute context-specific reliability profiles from past fusion outcomes. These mechanisms feed into a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, controlled by the global conflict level, followed by a belief-interval decision strategy that enables robust classification without discarding epistemic uncertainty. Experiments on 16 real-world benchmark datasets demonstrate that the proposed framework achieves an average F1 score of 85.78 and a mean AUC of 93.30, outperforming eight DST-based baselines and three gradient boosting methods. Ablation analysis confirms the contribution of each component we proposed. The framework offers an effective approach for adaptive evidence fusion in multi-source decision making. 1 Introduction Multi-source information fusion plays a critical role in modern decision-making systems, where heterogeneous evidence must be integrated to support reliable judgments under uncertainty (38; 13). Applications ranging from industrial fault diagnosis to pattern recognition and environmental monitoring share a common challenge: multiple evidence sources often provide assessments that are incomplete, imprecise, and mutually conflicting (41; 28; 20). The ability to fuse such evidence while preserving uncertainty information remains a fundamental problem in the fields of information systems and decision science. Dempster-Shafer theory (DST) provides a generalization of Bayesian probability that accommodates partial belief and ignorance through the assignment of basic probability assignments (BPAs) to subsets of a hypothesis space (31). Unlike classical probability, DST assigns mass to multi-element subsets, thereby explicitly representing epistemic uncertainty. Despite its theoretical elegance, DST faces well-documented challenges when evidence sources exhibit high conflict (4; 46). The Zadeh paradox (42) demonstrated that Dempster’s rule can produce counterintuitive results when two sources strongly disagree, and this limitation has motivated decades of research into conflict-aware fusion strategies. Two broad categories of solutions have emerged. The first modifies the combination rule itself: 40 reassigns conflict mass to the frame of discernment, 12 propose a disjunctive rule that preserves uncertainty rather than normalizing it away, and various authors have introduced weighted, cautious, or transferable belief models to handle specific failure modes. The second category addresses the evidence sources directly, applying distance-based or similarity-based weighting schemes before fusion. 24 averages BPAs uniformly; 9 weights sources using Jousselme distance; subsequent methods incorporate entropy, trust, or reliability coefficients to discount suspect evidence. Recent studies have attempted to combine preprocessing with adaptive fusion strategies. While these approaches have achieved incremental improvements, two persistent limitations remain. First, existing conflict measures (10) treat uncertainty and conflict as independent dimensions, yielding assessments that are either incomplete or redundant. Second, all of the above methods evaluate evidence solely on the basis of instantaneous BPA comparisons, ignoring the long-term track record of evidence sources across diverse decision contexts (39). These limitations are consequential in practice. Existing models predict that evidence sources operating under similar conditions should exhibit consistent reliability (18), yet empirical observation reveals the opposite: a source that produces misleading assessments in one scenario may be highly dependable in another, depending on the characteristics of the specific decision context (14). Methods that assign static weights or that evaluate evidence exclusively through current pairwise comparisons cannot capture such context-dependent reliability (26), leading to two failure modes. Informative sources that are occasionally conflicting but generally dependable are indiscriminately suppressed, while systematically unreliable sources receive insufficient penalization (43; 15). Furthermore, conventional conflict measures such as Dempster’s conflict coefficient K consider only the total mass assigned to empty intersections (35), disregarding the internal structure of multi-element focal elements. This simplification leads to coarse assessments of evidence quality, particularly when sources allocate substantial mass to composite hypotheses that express genuine ignorance rather than true disagreement. This paper proposes a unified evidence reasoning framework that addresses both limitations through two complementary mechanisms. The first is a chaos-conflict measurement (CCM) that jointly evaluates cross-evidence association and intra-evidence non-specificity within a single scalar quantity. The CCM is grounded in a novel evidence similarity measure whose properties of boundedness, symmetry, monotonicity, extreme consistency, and refinement insensitivity are formally proven. The second mechanism is a historical-experience-driven evidence weighting scheme that exploits past decision outcomes to learn context-dependent reliability profiles. Spectral clustering partitions the decision space into heterogeneous contexts, and regret theory is employed to quantify, for each evidence source, the counterfactual contribution when a fusion decision deviates from ground truth. Aggregated regret and rejoice scores are normalized to produce context-specific weights that reflect long-term source reliability. These two mechanisms feed into a hybrid combination rule that adaptively balances a conservative uncertainty preservation term against a weighted consensus component, with the tradeoff controlled by the global conflict level. A belief-interval decision rule then scores each hypothesis by combining its belief lower bound with a stability-weighted plausibility upper bound, enabling robust classification without forcing artificial redistribution of mass from multi-element focal elements. The proposed framework is evaluated on 16 real-world benchmark datasets spanning biology, medicine, geography, and demography, using decision trees as evidence sources. Compared against eight DST-based fusion methods and three gradient boosting baselines, the framework achieves the best average rank across all evaluation metrics. Robustness analyses demonstrate stable performance under feature noise, label noise, hybrid noise scenarios, varying numbers of evidence sources, and substitution of the base evidence generator. Ablation experiments confirm that historical experience weighting provides the largest individual contribution to performance, reducing AUC by 5.03% when removed. These results establish the proposed framework as an effective approach for evidence fusion that balances uncertainty preservation with adaptive reliability assessment, with direct applicability to any multi-source decision setting where evidence sources operate repeatedly across heterogeneous contexts. The remainder of this paper is organized as follows. Section 2 reviews the preliminary concepts of Dempster-Shafer theory, regret theory, and spectral clustering. Section 3 presents the proposed framework, including the chaos-conflict measurement, the historical experience driven weighting mechanism, the hybrid combination rule, and the belief-interval decision rule. Section 4 details the algorithmic design of the offline training and online inference phases. Section 5 reports the experimental setup, comparative results, robustness analyses, parameter sensitivity, and ablation study. Section 6 concludes the paper with a summary of contributions, a discussion of limitations, and directions for future research. 2 Preliminaries 2.1 Dempster-Shafer theory DST, also known as belief function theory, is a mathematical framework for representing and reasoning with uncertain, imprecise, and incomplete information. Originating from the work of 8 and formalized by 32, DST extends classical probability theory by allowing probability masses to be assigned not only to singletons but also to subsets of the hypothesis space. This provides a richer and more flexible representation of uncertainty, especially when prior knowledge is limited or heterogeneous evidence sources yield conflicting assessments. Definition 1. Frame of Discernment (FoD). Let Θ=θ1,θ2,⋯,θi,⋯,θn = \θ_1,θ_2,·s,θ_i,·s,θ_n \ denote the frame of discernment, i.e., a set of mutually exclusive and collectively exhaustive hypotheses. DST operates on the power set: 2Θ=∅,θ1,θ2,θ3,…,θ1,θ2,θ1,θ3,…,θ1,θ2,θ3,…,Θ2 = \ ,\;\ _1\,\;\ _2\,\;\ _3\,\;…,\\ \ _1, _2\,\;\ _1, _3\,\;…,\;\ _1, _2, _3\,\;…,\; \ (1) where each subset represents a proposition whose truth is uncertain. Here, ∅ is an empty set. If S∈2Θ,S∈2 ,S is called a hypothesis. Definition 2. Basic Probability Assignment. A BPA or mass function m:2Θ→[0,1]m:2 → [0,1 ] satisfies: ∑S⊆2Θm(S)=1m(∅)=0 \ array[]l Σ _S 2 m (S )=1\\ m ( )=0 array . (2) A subset S with m(S)>0m (S )>0 is called a focal element (FE). Mass assigned to multi-element sets (N-Focal Elements, NFEs) expresses epistemic uncertainty or ignorance, since it indicates that the exact hypothesis cannot be precisely identified. Conversely, single-element focal elements (SFEs) represent precise hypotheses. Definition 3. Belief and plausibility functions. Based on BPA, DST defines two dual measures that jointly describe the degree of support for a proposition. The belief function, representing the minimal support for S : Bel(S)=∑A⊆Sm(A)Bel (S )= Σ _A Sm (A ) (3) The plausibility function, representing the maximal potential support for S : Pl(S)=1−Bel(S¯)=∑A∩S≠∅m(A),A⊆ΘPl (S )=1-Bel ( S )= Σ _A∩ S≠ m (A ),A (4) Together, [Bel(S),Pl(S)] [Bel (S ),Pl (S ) ] forms the belief interval, which provides a natural expression of uncertainty. A wider interval indicates greater ambiguity, while a narrow interval signifies strong confidence. Definition 4. Dempster’s combination rule. One of DST’s core strengths is its ability to integrate independent pieces of evidence. Given n mass functions m1,m2,…,mnm_1,m_2,…,m_n defined on the same frame, Dempster’s rule combines them as m1⊕2⊕…⊕n(S)=∑∩Ai=S,Ai⊆2Θ∏jmj(Ai)1−K=∑⋂Ai=∅,Ai⊆2Θm1(A1)⋅m2(A2)⋯mn(An) \ array[]lm_1 2 … n (S )= Σ _∩A_i=S,A_i 2 Π _jm_j (A_i )1-K\\ K= Σ _ A_i= ,A_i 2 m_1 (A_1 )·m_2 (A_2 )·sm_n (A_n ) array . (5) where ⊕ denotes Dempster’s orthogonal sum, which aggregates independent pieces of evidence through combination. K∈[0,1]K∈ [0,1 ] is the conflict coefficient. A higher K implies stronger disagreement between evidence sources. While Dempster’s rule is elegant and widely used, its behavior in high-conflict situations may yield counterintuitive results, motivating a large body of research on conflict management and evidence fusion strategies. 2.2 Regret theory Regret theory (RT) explains choice under uncertainty by embedding the anticipatory emotions of regret and rejoicing. Whereas expected utility theory treats agents as evaluators of absolute payoffs, RT posits that people also compare the realized outcome with what they would have received had they chosen differently. Regret arises when the forgone alternative would have delivered a superior outcome; rejoicing emerges when it would have delivered an inferior one. Definition 5. Basic formulation of RT. Let AC1AC_1 and AC2AC_2 be two possible actions. If AC1AC_1 yields outcome x1x_1 and AC2AC_2 yields outcome x2x_2 , the decision maker (DM) does not merely consider the utility values u(x1)u (x_1 ) and u(x2)u (x_2 ) but also the psychological difference between them. The utility obtained by choosing action AC1AC_1 and thereby forgoing AC2AC_2 is (22): v=u(x1)+r(u(x1)−u(x2))v=u (x_1 )+r (u (x_1 )-u (x_2 ) ) (6) where u(∙)u ( ) is a monotonic utility function, r(Δu)r ( u ) is an increasing and typically S-shaped regret-rejoice function and is defined as follow: Definition 6. Regret-rejoice function. The utility function commonly used in regret theory is u(xi)=1−e−θxiθu (x_i )= 1-e^-θx_iθ (7) where θ∈(0,1)θ∈ (0,1 ) represents the DM’s degree of risk aversion. A widely used parametric representation for regret-rejoice function (2) is: r(Δu)=1−e−δΔur ( u )=1-e^-δ u (8) where δ∈[0,+∞)δ∈[0,+∞) denotes the regret-rejoice sensitivity coefficient. Within this study, RT is repurposed to audit the past behavior of evidence sources. Whenever the fused decision is wrong, each contributor is retrospectively credited with either a regret or a rejoice score: regret if its testimony steered the ensemble away from the truth, rejoice if it pulled the ensemble toward the truth. Aggregating these scores across historical instances quantifies the long-run reliability of every evidence source. 2.3 Spectral clustering Spectral clustering (SC) treats grouping as a graph-partitioning problem: it recovers the intrinsic geometry of data from the eigenvectors of an affinity matrix rather than from distances (25). The algorithm readily separates strongly connected vertex subsets that are only weakly linked to the rest of the graph, giving it a decisive advantage over conventional methods such as k-means when the underlying distribution is high-dimensional (3). Definition 7. Graph construction. Given a dataset X=x1,x2,…,xnX= \x_1,x_2,…,x_n \ , each data instance is treated as a vertex in an undirected weighted graph G=⟨V,E,W⟩G= V,E,W . Edges encode pairwise similarity, and the similarity matrix W=[wij]W= [w_ij ] is constructed through a Gaussian kernel: wij=exp(−‖xi−xj‖222σ2)w_ij= (- Vmatrixx_i-x_j Vmatrix_2^22σ^2 ) (9) where σ is a scale parameter controlling locality. Definition 8. Graph Laplacian Formulation. Given the affinity matrix W∈ℝn×nW∈R^n× n constructed in Definition 7, define the degree matrix D as a diagonal matrix with entries Dii=∑jwijD_i= Σ _jw_ij . The pair (D,W) (D,W ) induces related Laplacian operators that encode the structural properties of the graph. The symmetric normalized Laplacian is given by L=I−D−1/2WD−1/2L=I-D^-1/2WD^-1/2 (10) Definition 9. Graph Cut Objective. Given a weighted graph G=(V,E,W)G= (V,E,W ) , clustering can be viewed as partitioning the vertex set into k disjoint subsets g1,g2,…,gkg_1,g_2,…,g_k . A desirable partition should minimize the similarity between different subsets while maintaining high within-cluster affinity. To formalize this principle, 33 introduced the Normalized Cut (Ncut) criterion, defined for a k -way partition as NCutq=ming1,g2,…,gk⊆V12∑i=1kWa(gi,gi¯)vol(gi)NCut_q= _g_1,g_2,…,g_k V 12 Σ _i=1^k W_a (g_i, g_i )vol (g_i ) (11) where W(gi,gi¯)=∑u∈gi∑v∈giwuvW (g_i, g_i )= Σ _u∈g_i Σ _v∈g_iw_uv denotes the total weight of edges crossing the boundary of gig_i , and vol(gi)=∑u∈giDuuvol (g_i )= Σ _u∈g_iD_u is the degree mass of subset gig_i . The Ncut criterion balances the cut value by the size of each subset, thereby preventing trivial solutions in which a small set is separated from the rest of the graph. 3 Proposed approach 3.1 Overview We propose a DST evidence reasoning framework that combines chaos-conflict measurement with historical-experience weighting for robust multi-source decision making. The framework operates in two phases: offline training and online inference, as illustrated in Figure 1. In the offline phase, historical BPAs from multiple evidence sources are processed along two parallel paths. SC partitions the decision space into context scenarios based on Sample feature representations. Meanwhile, Dempster’s rule fuses the historical BPAs to identify decision failures. For each failure, the regret-rejoice mechanism evaluates individual evidence sources: the optimal evidence receives a rejoice score proportional to its conflict with the erroneous fusion, while each candidate erroneous source receives a regret score proportional to the distortion it introduces. Aggregated regret and rejoice scores are normalized via softmax to produce context-specific historical experience weights. In the online phase, a test sample generates BPAs from the same evidence sources. The chaos-conflict measurement computes pairwise association and similarity across all evidence pairs, yielding a global chaos-conflict degree that captures both inter-evidence inconsistency and intra-evidence uncertainty. The test BPAs are then weighted using the context-indexed historical weights retrieved from offline training. A hybrid combination rule adaptively balances a conservative Dubois term with the weighted consensus evidence, controlled by the global conflict level. When conflict is high, the rule emphasizes uncertainty preservation; when conflict is mild, it relies on the consensus term. Finally, a belief-interval decision rule scores each hypothesis by combining its belief lower bound with a stability-weighted plausibility upper bound, outputting the predicted class label. Figure 1: Overall architecture of the proposed framework. 3.2 Chaos-conflict measurement In multi-source evidence fusion, conflict reflects discrepancies between evidence sources, whereas uncertainty stems from NFEs. Existing measures treat them independently, yielding incomplete assessments: conflict metrics overlook internal ambiguity, while uncertainty metrics fail to capture cross-evidence inconsistency. To better characterize overall evidential stability, we propose a unified measurement named chaos-conflict measurement (CCM). By jointly evaluating FE set’s compatibility and intrinsic uncertainty, CCM offers a more coherent basis for fusion. Definition 10. Evidence association measure. Let mim_i and mjm_j be two mass functions defined on the same FoD Θ=θ1,θ2,…,θn = \θ_1,θ_2,…,θ_n \ , denote by S1,…,S2n \S_1,…,S_2^n \ the set of all focal subsets of Θ . The pairwise association measure between mim_i and mjm_j is defined as k(mi,mj)=∑i=12n∑j=12n2m(Si)m(Sj)|Si∩Sj||Si||Sj|(|Si|+|Sj|)k (m_i,m_j )= Σ _i=1^2^n Σ _j=1^2^n2 m (S_i )m (S_j ) |S_i∩S_j | |S_i | |S_j | ( |S_i |+ |S_j | ) (12) =∑Si∩Sj≠∅m(Si)|Si|⋅m(Sj)|Sj|⋅2|Si∩Sj||Si|+|Sj|= Σ _S_i∩S_j≠ m (S_i ) |S_i |· m (S_j ) |S_j |· 2 |S_i∩S_j | |S_i |+ |S_j | (13) Corollary 1. For a collection of mass functions m1,…,mJ \m_1,…,m_J \ defined on the same FoD Θ , their global association measure is given by k(⊕j=1mj)=∑∩Si≠∅(∏jmj(Si)|Si|)|⋂iSi|∑i|Si|k ( j=1 J m_j )= Σ _∩S_i≠ ( Π _j m_j (S_i ) |S_i | ) | _iS_i | Σ _i |S_i | (13a) Corollary 1 follows immediately from Definition 10 by extending the pairwise association to all mass functions on the frame. It aggregates compatible FEs across all evidence sources. Remark 1. The global association measure incorporates both the non-specificity of NFEs and their compatibility structure. In particular, mass assigned to NFEs reflects epistemic uncertainty or ignorance, which weakens the reliability of the corresponding BPA. The formulation therefore down-weights such contributions and adjusts the association strength according to the degree of focal-set overlap. As a result, the measure captures both the uncertainty in NFE allocations and the compatibility across evidence sources, yielding a more faithful characterization of their mutual association. Definition 11. Evidence similarity. For two mass functions mim_i and mjm_j defined on the same Θ , their similarity is quantified by: S(mi,mj)=k(mi,mj)1−k(mi,mj)+k(mi,mi)k(mj,mj)S (m_i,m_j )= k (m_i,m_j )1-k (m_i,m_j )+k (m_i,m_i )k (m_j,m_j ) (14) where S(⋅)S (· ) denotes the similarity between the two mass functions. Before establishing the main properties of the similarity measure in Definition 11, we first introduce auxiliary lemma that facilitates the subsequent proofs. Lemma 1. Cauchy-Schwarz inequality. For any real vectors x=(x1,…,xn)Tx= (x_1,…,x_n )^T and y=(y1,…,yn)Ty= (y_1,…,y_n )^T , the following inequality holds: (∑i=1nxiyi)2≤(∑i=1nxi2)(∑i=1nyi2) ( Σ _i=1^nx_iy_i )^2≤ ( Σ _i=1^nx_i^2 ) ( Σ _i=1^ny_i^2 ) (15) with equality if and only if x=λyx=λ y for some scalar λ∈Rλ∈ R . Proof. The result is classical and follows directly from the fact that the quadratic form ∑i=1n(xi−λyi)2≥0 Σ _i=1^n (x_i-λy_i )^2≥ 0 holds for all λ . ∎ Lemma 2. Positive semi-definiteness of the compatibility matrix. Let M¯ij=2|Si∩Sj||Si||Sj|(|Si|+|Sj|),i,j=1,…,2|Θ|, M_ij= 2 |S_i∩S_j | |S_i | |S_j | ( |S_i |+ |S_j | ),\;i,j=1,…,2 | |, then M¯ M is a real symmetric positive semi-definite matrix. Proof. We begin by fixing an ordering of the power set 2Θ,2Θ=S1,S2,…,S2|Θ|2 ,2 = \S_1,S_2,…,S_2 | | \ . It is immediate that M¯ij=M¯ji M_ij= M_ji , hence M¯ M is real and symmetric. Consider any nonzero vector z=(z1,…,z2|Θ|)Tz= (z_1,…,z_2 | | )^T , the associated quadratic form is zTM¯z=∑i=12|Θ|∑j=12|Θ|zizj2|Si∩Sj||Si||Sj|(|Si|+|Sj|).z^T Mz= Σ _i=1^2 | | Σ _j=1^2 | |z_iz_j 2 |S_i∩S_j | |S_i | |S_j | ( |S_i |+ |S_j | ). Because every focal element is compatible with itself, we have |Si∩Si|=|Si|>0⇒M¯ii=2|Si||Si||Si|(|Si|+|Si|)=1|Si|2>0. |S_i∩S_i |= |S_i |>0 M_i= 2 |S_i | |S_i | |S_i | ( |S_i |+ |S_i | )= 1 |S_i |^2>0. Hence each diagonal entry is strictly positive. To show that M¯ M is positive semi-definite, we substitute the identity 2|Si|+|Sj|=2∫0∞e−(|Si|+|Sj|)tt 2 S_i + S_j =2 _0^∞e^- ( S_i + S_j )t\,dt into the quadratic form, which yields zTM¯z=2∫0∞∑θ∈Θ(∑i∋θzie−|Si|t|Si|)2t≥0z^T Mz=2 _0^∞ _θ∈ ( _i θz_i e^- S_i t S_i )^2\,dt≥ 0 for every vector z, because the integrand is a sum of squares. Hence zTM¯z≥0z^T Mz≥ 0 for all z, which proves that M¯ M is positive semi-definite. ∎ We adopted the set of attributes that a conflict measure ought to satisfy as stipulated by 10 and supply the corresponding proof below. Property 1. Boundedness, 0≤S(mi,mj)≤10≤ S (m_i,m_j )≤ 1 . Proof. Since M¯ M is positive semi-definite, there exists a lower-triangular matrix L such that M¯=LL⊤ M=LL . Define =L,=Lx=vL,y=bL . Then k(mi,mj)=2LL⊤⊤=2⊤,k(mi,mi)=2⊤k(mj,mj)=2⊤k (m_i,m_j )=2vLL b =2xy ,k (m_i,m_i )=2x \;k (m_j,m_j )=2y . Set a=⊤,b=⊤,c=⊤a=xy ,b=x ,c=y . Then S(mi,mj)=2a1−2a+4bc.S (m_i,m_j )= 2a1-2a+4bc. By construction, all entries of M¯ M and of the mass vectors are non-negative, hence a,b,c≥0a,b,c≥ 0 . Moreover 0≤k(mi,mj)≤10≤ k (m_i,m_j )≤ 1 and k(mi,mj)=2ak (m_i,m_j )=2a imply a∈[0,1/2]a∈ [0,1/2 ] . Thus the denominator satisfies 1−2a+4bc≥1−2a≥01-2a+4bc≥ 1-2a≥ 0 , and the numerator 2a≥02a≥ 0 ; therefore S(mi,mj)≥0S (m_i,m_j )≥ 0 . By Lemma 1, a2=(⊤)2≤(⊤)(⊤)=bca^2= (xy )^2≤ (x ) (y )=bc . Let u=bc≥0u=bc≥ 0 and 0≤a≤u0≤ a≤ u . To show S(mi,mj)≤1S (m_i,m_j )≤ 1 , it suffices to prove 2a≤1−2a+4bc⇔0≤u−u+1/42a≤ 1-2a+4bc 0≤ u- u+1/4 . Consider the function f(u)=u−u+1/4f (u )=u- u+1/4 , Its derivative is f′(u)=1−12u,u∈(0,1].f (u )=1- 12 u,u∈(0,1]. The equation f′(u)=0f (u )=0 has the unique solution u∗=1/4u^*=1/4 . Since f(1/4)=0f (1/4 )=0 , by continuity we obtain f(u)≥0,∀u∈[0,1]f (u )≥ 0,\;∀ u∈ [0,1 ] . Therefore S(mi,mj)≤1S (m_i,m_j )≤ 1 . ∎ Property 2. Symmetry, S(mi,mj)=S(mj,mi)S (m_i,m_j )=S (m_j,m_i ) . Proof. S(mi,mj)=k(mi,mj)1−k(mi,mj)+k(mi,mi)⋅k(mj,mj)S (m_i,m_j )= k (m_i,m_j )1-k (m_i,m_j )+k (m_i,m_i )· k (m_j,m_j ) =∑i=12n∑j=12n2m(Si)m(Sj)|Si∩Sj||Si||Sj|(|Si|+|Sj|)=∑i=12n∑j=12n2m(Sj)m(Si)|Sj∩Si||Sj||Si|(|Sj|+|Si|)=S(mj,mi)= Σ _i=1^2^n Σ _j=1^2^n2 m (S_i )m (S_j ) |S_i∩S_j | |S_i | |S_j | ( |S_i |+ |S_j | )\\ = Σ _i=1^2^n Σ _j=1^2^n2 m (S_j )m (S_i ) |S_j∩S_i | |S_j | |S_i | ( |S_j |+ |S_i | )=S (m_j,m_i ) Thus, the similarity measure is symmetric. ∎ Property 3. Monotonicity. Let m1m_1 and m2m_2 be two mass functions defined on the same FoD Θ . Without loss of generality, assume k(m2,m2)≥k(m1,m1)k (m_2,m_2 )≥ k (m_1,m_1 ) . For t∈[0,1]t∈ [0,1 ] , consider the convex combination m3(t)=(1−t)m2+tm1m_3 (t )= (1-t )m_2+tm_1 , define ζ(t) := S(m1,m3(t))ζ (t ) := S (m_1,m_3 (t ) ) . Then ζ(t)ζ (t ) is non-decreasing on [0,1] [0,1 ] . Moreover, if m1≠m2m_1≠m_2 , then ζ(t)ζ (t ) is strictly increasing on (0,1) (0,1 ) . Proof. By Property 1, we have ∂S∂a=2+8bc(1−2a+4bc)2>0. ∂ S∂ a= 2+8bc (1-2a+4bc )^2>0. Thus, S(mi,mj)S (m_i,m_j ) is a strictly increasing transform of k(mi,mj)k (m_i,m_j ) . Meanwhile, k(m1,m3(t))=2a(t)∝xy(t)⊤=x((1−t)y+tx)⊤k (m_1,m_3 (t ) )=2a (t ) xy (t ) =x ( (1-t )y+tx ) =(1−t)xy⊤+txx⊤=(1−t)a+tb=⟨m1,(1−t)m2+tm1⟩D.= (1-t )xy +txx = (1-t )a+tb= m_1, (1-t )m_2+tm_1 _D. Differentiating gives a′(t)=b−aa (t )=b-a . Because a≤bca≤ bc and b≥c,a′(t)>0,mi≠mjb≥ c,a (t )>0,\;m_i≠m_j . The composition S∘a\;S a\; is strictly increasing. Consequently, t1<t2⇒S(m1,m3(t1))<S(m1,m3(t2)).t_1<t_2 S (m_1,m_3 (t_1 ) )<S (m_1,m_3 (t_2 ) ). ∎ Property 4. Extreme Consistency. (1) If and only if mi=mjm_i=m_j and the BPAs of mim_i and mjm_j are all concentrated on the same SFE, S(mi,mj)=1S (m_i,m_j )=1 ;(2)S(mi,mj)=0⇔(⋃Si)∩(⋃Sj)=∅.S (m_i,m_j )=0 ( S_i )∩ ( S_j )= . Proof. For (1), by Property 1, we have a=1/4+bca=1/4+bc and a2≤bca^2≤bc . When u∗=1/4u^*=1/4 , we have a=b=c=1/2a=b=c=1/2 , then x=yx=y and k(mi,mj)=k(mi,mi)=k(mj,mj)=1k (m_i,m_j )=k (m_i,m_i )=k (m_j,m_j )=1 , condition mi=mjm_i=m_j is satisfied. Due to Eq. 2, let the unique FE be A , then their masses are mi(A)=mj(A)=1m_i (A )=m_j (A )=1 . If A is a NFE, (1/|A|)2<1 (1/ |A | )^2<1 , which contradicts k(mi,mj)=1k (m_i,m_j )=1 . Thus A must be a SFE. For (2), by the Property 1,S(mi,mj)=0⇔a=01,S (m_i,m_j )=0 a=0 . Due to the construction, we have a∝∑Si∩Sj≠∅mi(Si)mj(Sj)|Si∩Sj|/(|Si|+|Sj|)a Σ _S_i∩S_j≠ m_i (S_i )m_j (S_j ) |S_i∩S_j |/ ( |S_i |+ |S_j | ) and all terms in the summation are nonnegative. Therefore a can be zero only if |Si∩Sj|=0 |S_i∩S_j |=0 , which is equivalent to (∪Si)∩(∪Sj)=∅ . (∪S_i )∩ (∪S_j )= . ∎ Property 5. Refinement Insensitivity. When the FoD is refined from Θ to Θ′,S(mi,mj) ,S (m_i,m_j ) remains unchanged. Proof. A refinement of the frame preserves the structure of both mass functions. Let S⊆Θ′S denote such a refined FE. Under refinement, the masses satisfy mi(S)=mj(S)=0m_i (S )=m_j (S )=0 , and these terms appear symmetrically in the computation of k(mi,mj)k (m_i,m_j ) , contributing additively as zeros. Removing these null contributions does not affect any terms involved in the definition of the similarity measure. Consequently, the value of S(mi,mj)S (m_i,m_j ) is invariant under refinement of the frame. ∎ To further illustrate the proposed similarity measure, we consider two numerical examples. Example 1. Three BPA functions m1,m2m_1,m_2 and m3m_3 are defined on the same FoD Θ and constructed as follows: m1(θ1)=x,m1(θ2)=1−xm_1 (θ_1 )=x,m_1 (θ_2 )=1-x m2(θ1)=1−x,m2(θ2)=xm_2 (θ_1 )=1-x,m_2 (θ_2 )=x m3(θ1)=1−x,m3(θ1,θ2)=xm_3 (θ_1 )=1-x,m_3 (θ_1,θ_2 )=x where x∈[0,1]x∈ [0,1 ] . Using Eqs. 12 and 14, the similarity values are computed and the results are shown in Figure 2. Several observations can be made: (1) When x=0x=0 or x=1x=1 , the BPAs become completely conflicting, the similarity correctly returns S(⋅)=0S (· )=0 ; (2) When x=0.5x=0.5, m1m_1 and m2m_2 become identical. However, the resulting similarity does not reach 1. This is because S(⋅)S (· ) not only considers the closeness between the two BPAs, but also incorporates their intrinsic uncertainty (nonspecificity). At this moment, both BPAs carry the maximum amount of ambiguity and contribute no effective decision-making information. This observation is consistent with Property 4, which establishes that similarity reaches 1 only under defined conditions. And moreover: (3) We observe that S(m1,m3)≤S(m1,m2)S (m_1,m_3 )≤ S (m_1,m_2 ) . Compared with m2,m3m_2,m_3 assigns part of its mass to NFE. This allocation increases the level of nonspecificity, resulting in greater intrinsic uncertainty. Consequently, even when the BPA values appear similar, the uncertainty prevents the similarity from reaching a higher value. Figure 2: Similarity for Example 1 Example 2. m1(θ1)=x,m1(θ2)=1−xm_1 (θ_1 )=x,m_1 (θ_2 )=1-x m2(θ1)=x,m2(θ2)=1−xm_2 (θ_1 )=x,m_2 (θ_2 )=1-x m3(θ1)=x,m3(θ1,θ2)=1−xm_3 (θ_1 )=x,m_3 (θ_1,θ_2 )=1-x where x∈[0,1]x∈ [0,1 ] and the similarity results are shown in Figure 3. From Figure 3, it can be observed that even when two evidences follow identical BPA assignment, their similarity values may still differ significantly. Specifically, when the BPA becomes more dispersed, the associated uncertainty increases accordingly. In such cases, it becomes difficult to reliably assess the similarity between two evidences, as higher uncertainty weakens the confidence in their consistency. Furthermore, when BPA values are transferred from SFEs to their supersets, as in the case of m3m_3 , additional uncertainty is introduced. Since a NFE contains multiple propositions, the belief mass assigned to it cannot be fully committed to any single hypothesis. As a result, the evidence provides less information for decision-making, leading to a lower similarity. Figure 3: Similarity for Example 2 In summary, the proposed similarity measure quantifies the agreement between evidences by jointly comparing their BPA while fully accounting for their intrinsic uncertainty. On this basis, we introduce the CCM of evidence as follows: Definition 12. CCM. Let mim_i and mjm_j be two arbitrary mass functions defined on the same FoD Θ and the CCM is defined as K^(mi,mj)=1−S(mi,mj)=1−2k(mi,mj)+k(mi,mi)⋅k(mj,mj)1−k(mi,mj)+k(mi,mi)⋅k(mj,mj) K (m_i,m_j )=1-S (m_i,m_j )\\ = 1-2k (m_i,m_j )+k (m_i,m_i )· k (m_j,m_j )1-k (m_i,m_j )+k (m_i,m_i )· k (m_j,m_j ) (2) It is straightforward to verify that the CCM inherits the same five desirable properties as the similarity measure. Corollary 2. Global chaos-conflict degree (GCCD). Let m1,m2,…,mn \m_1,m_2,…,m_n \ be a set of mass functions defined on the same FoD. The GCCD of the evidence set is defined as K^=1n(n−1)/2∑i<jK^(mi,mj). K= 1n (n-1 )/2 Σ _i<j K (m_i,m_j ). (16) Corollary 3. Relative stability reliability. Given a set of mass functions m1,m2,…,mn \m_1,m_2,…,m_n \ defined on the same FoD, the relative stability reliability of an individual evidence source mim_i is defined as SR(mi)=∑j≠iS(mi,mj)1−K^.SR (m_i )= Σ _j≠ iS (m_i,m_j )1- K. (17) While the above definitions and corollaries enable instantaneous evidence assessment of conflict and stability, they remain limited to the current decision context. In practical applications where evidence sources are repeatedly reused under heterogeneous conditions, exploiting historical experience becomes essential for achieving robust and adaptive evidence fusion, which is the focus of the next section. 3.3 Evidence weighting driven by historical experience Existing strategies predominantly assess evidential inconsistency based on the current hypothesis space and the configuration of BPAs. Such approaches, however, implicitly assume that all evidence sources are equally reliable across different decision contexts (30), thereby overlooking the informative value embedded in their long-term historical performance. In real-world applications, evidence sources are rarely one-shot contributors. The generation of BPAs is often constrained by complex environmental factors, and the reliability of an evidence source typically exhibits within specific scene. Relying exclusively on instantaneous conflict evaluation may therefore lead to biased reliability assessment and suboptimal evidence weighting (43; 16). A representative example can be found in complex diagnostic scenarios, where different evidence sources exhibit markedly different detection capabilities across specific conditions (23). In such cases, certain sources may systematically outperform others in particular contexts, despite occasionally producing conflicting assessments. Treating these discrepancies solely as instantaneous conflict may overlook the historically validated reliability of informative evidence sources and lead to their unjustified suppression during fusion (37). Motivated by these observations, we introduce a historical experience driven evidence weighting mechanism that complements the CCM proposed above. Rather than relying exclusively on instantaneous conflict evaluation, the proposed approach leverages historical decision outcomes to learn context-dependent evidence weights, which are subsequently integrated into a hybrid combination rule that adaptively balances conservative uncertainty preservation with weighted consensus. Specifically, heterogeneous decision contexts are first identified via clustering, after which the historical behavior of each evidence source is evaluated through regret and rejoice measures derived from regret theory. These measures quantify the counterfactual contribution of individual evidence sources when fusion outcomes deviate from ground truth, enabling continuous credit assignment that rewards evidence sources aligned with successful outcomes and penalizes those contributing to fusion errors. The accumulated historical experience is then translated into adaptive evidence weights, which are subsequently incorporated into the fusion process. By explicitly integrating long-term evidence behavior with instantaneous conflict assessment, our approach avoids indiscriminate suppression of conflicting evidences and yields a more reliable fusion strategy. RT is employed to analyze the potential regret effects that may arise when decision outcomes deviate from expectations. This theory not only focuses on the absolute gains of decisions but also emphasizes the relative gains compared to other potential choices. Introducing regret theory into the calculation of historical experience for evidence synthesis provides a quantitative and systematic method for evaluating the applicability of different pieces of evidence in historical contexts, thereby aiding in the optimization of future evidence selection and integration strategies. Among these, the calculation method for regret-rejoicing values, which assesses the applicability of evidence, will serve as the core research focus of this section. Next, we will introduce the regret-rejoicing measurement of evidence bodies during the evidence synthesis process. Let Θ=θ1,…,θn = \θ_1,…,θ_n \ denote the FoD and M=m1,m2,…,mhM= \m_1,m_2,…,m_h \ be the set of available evidence sources defined on 2Θ2 . Definition 13. Optimal evidence. For a given sample whose ground-truth hypothesis is known, the optimal evidence is defined as the evidence source whose BPA is most consistent with the truth according to the similarity measure defined in Eq. 14: m best =argminmi∈MK^(mi,p target )m_ best = _m_i∈ M K (m_i,p_ target ) (18) Here, p target p_ target is a one-hot vector of dimension 2n2^n . Definition 14. Error-removal combination. Given the evidence set M , the initial fused result m⊕m_ is first obtained using Dempster’s rule of combination as defined in Eq. 5. If the decision derived from m⊕m_ is inconsistent with the truth, a decision error is considered to have occurred. In such a case, the evidence set is assumed to contain at least one erroneous evidence source, whose mass function fails to assign its maximum support to the true hypothesis. Let me∈Mm_e∈ M denote a candidate erroneous evidence source. By removing mem_e from the M and recombining the remaining evidence, the error-removal combination result m⊕−em_ ^-e is defined as m⊕−e(S)=∑∩Sj=S,Sj⊆2Θ∏i≠emi(Sj)1−K=∑⋂Sj=∅,Sj⊆2Θ∏i≠emi(Sj). \ array[]lm_ ^-e (S )= Σ _∩S_j=S,S_j 2 Π _i≠ em_i (S_j )1-K\\ K= Σ _ S_j= ,S_j 2 Π _i≠ em_i (S_j ) array. . (19) Definition 15. Rejoice value of the optimal evidence. Based on the CCM and the formulation of RT, the rejoice value associated with the m best m_ best is defined as j(m best )=1−exp(−ηK^(m best ,m⊕)),j (m_ best )=1- (-η K (m_ best ,m_ ) ), (20) where η∈(0,1]η∈(0,1] is a sensitivity parameter and controls the responsiveness of the rejoice value to variations in conflict intensity. Remark 2. The underlying intuition of Definition 15 is as follows. When the fused evidence leads to an incorrect decision, it indicates that the collective judgment derived from evidence aggregation is unreliable. In such cases, the decision supported by the optimal evidence may provide a more faithful reflection of ground truth. Consequently, a larger conflict between m best m_ best and the erroneous fusion result m⊕m_ implies that m best m_ best is closer to the true outcome, and should therefore be assigned a higher rejoice value. Definition 16. Regret value of erroneous evidence. Analogous to Definition 15, the regret value associated with an erroneous evidence source mem_e is defined as r(me)=1−exp(−γK^(m⊕−e,m⊕)),r (m_e )=1- (-γ K (m_ ^-e,m_ ) ), (21) where γ∈(0,1]γ∈(0,1] is a sensitivity parameter analogous to η , which governs the responsiveness of the regret value to variations in conflict intensity. Remark 3. The intuitive interpretation of Definition 16 is that the discrepancy between the m⊕m_ and the m⊕−em_ ^-e reflects the extent to which the erroneous evidence distorts the collective decision. A larger conflict between these two fusion outcomes indicates a stronger negative impact of mem_e on the fusion process, and consequently corresponds to a higher regret value. In practical decision systems, the reliability of an evidence source is rarely invariant across all operating conditions. Instead, evidence behavior typically exhibits strong context dependency: a source that is highly informative in one scenario may become less reliable in another due to changes in acquisition conditions, latent nuisance factors, or feature distribution shifts (18; 44). To reflect such heterogeneity, we partition the historical samples into a set of decision contexts (clusters), and assign each sample to a context. Within each context, we then aggregate regret and rejoice values induced by decision failures, thereby obtaining context-conditioned historical experience for each evidence. Definition 17. Historical experience-based evidence weight. Let ℋ=(xt,yt)t=1NH= \ (x_t,y_t ) \_t=1^N be the historical set and =c1,…,cLC= \c_1,…,c_L \ the context partition induced by clustering algorithm. c(xt)∈c (x_t ) denotes the cluster membership of xtx_t . For an evidence source mim_i and a target context c∈c , define the failure indicator δt=[argmaxθk∈ΘBelm⊕(⋅∣xt)(θk)≠yt]δ_t=I [ _θ_k∈ Bel_m_ (· x_t ) (θ_k )≠y_t ] (22) where (⋅)I (· ) denotes the indicator function and δt=1δ_t=1 indicates that the final fusion decision at sample xtx_t is incorrect. Let jt(mi)j_t (m_i ) and rt(mi)r_t (m_i ) denote the rejoice and regret values of evidence source mim_i computed according to Eqs. 20 and 21, the weight score of mim_i under context c is defined as: JRWic=∑t=1Nδt[c(xt)=c]jt(mi)−∑t=1Nδt[c(xt)=c]rt(mi).JRW_i^c= Σ _t=1^Nδ_tI [c (x_t )=c ]j_t (m_i )- Σ _t=1^Nδ_tI [c (x_t )=c ]r_t (m_i ). (23) Then, softmax normalization JRWic~=exp(JRWic)∑j=1Mexp(JRWjc) JRW_i^c= (JRW_i^c ) Σ _j=1^M (JRW_j^c ) (24) is implemented to guarantee the final weight JRWic~≥0 JRW_i^c≥ 0 and ∑iJRWic~=1 Σ _i JRW_i^c=1 . 3.4 A robust fusion-and-decision rule Building upon the historical experience-driven evidence weighting mechanism developed in Section 3.3, we now construct a robust fusion rule that integrates long-term source reliability with instantaneous evidential conflict. Specifically, Section 3.3 yields context-dependent reliability profiles learned from historical decision outcomes, which enable the construction of a weighted evidence representation reflecting the heterogeneous performance of different evidence sources. While such historical calibration mitigates the risk of indiscriminately suppressing informative evidence, robust fusion further requires an explicit treatment of the conflict structure exhibited by the current evidence set. To this end, Section 3.2 introduced the CCM, which characterizes not only the degree of evidential inconsistency but also the uncertainty embedded in evidence formation. Motivated by these observations, this section proposes a new robust evidence fusion rule following a principled two-level design. First, instead of aggregating evidence via the conventional simple averaging strategy in 24’s method, we employ reliability discounting guided by historical experience to obtain a calibrated weighted evidence representation. This calibrated evidence is subsequently used as the consensus component in the fusion process. Second, the CCM is computed to quantify the instantaneous inconsistency more faithfully. Finally, a conflict-adaptive hybrid combination is performed, which smoothly balances informative consensus and uncertainty according to the estimated conflict level. In this way, the proposed rule achieves improved robustness under heterogeneous contexts. The following is the historical experience weighted evidence we obtained: mw(S∣x)=∑iJRWic(x)~⋅mi(S∣x)m_w (S x )= Σ _i JRW_i^c(x)·m_i (S x ) (25) where x denotes the target sample for which evidence fusion is performed, and the notation mi(S∣x)m_i (S x ) indicates that the BPA provided by the i¯ i -th evidence source. The corresponding historical experience weight JRWic(x)~ JRW_i^c(x) is therefore selected, ensuring that the contribution of each evidence is modulated according to its long-term performance under similar conditions. Consequently, the weighted evidence mw(S∣x)m_w (S x ) represents a context-aware aggregation of the original BPAs associated with the sample x . After obtaining weighted evidence based on historical experience, we give a new evidence fusion rule defined as follows. Definition 18. Hybrid combination rule. To construct belief intervals and to allocate the aggregated mass more informatively, we define a hybrid combination rule as follows: m⊕(S)=(1−e−K^)(∑A1⋂⋯⋂AM=S≠∅∏i=1Mmi(Ai)+∑A1⋂⋯⋂AM=∅A1⋃⋯⋃AM=S∏i=1Mmi(Ai))+e−K^mw(S)m (S )= (1-e^- K ) ( Σ _A_1 ·s A_M=S≠ Π _i=1^Mm_i (A_i )\\ + Σ _ subarraycA_1 ·s A_M= \\ A_1 ·s A_M=S subarray Π _i=1^Mm_i (A_i ) )+e^- Km_w (S ) (3) Eq. 3 consists of two complementary terms. The first term is a refined Dubois combination rule 21, which aggregates both conjunctive contributions and union-based allocations so as to explicitly preserve uncertainty under severe disagreement. The second term mw(S)m_w (S ) represents the consensus evidence obtained by weighted averaging, which is beneficial when the evidence sources are largely consistent. The mixing factor e−K^∈(0,1]e^- K∈(0,1] adaptively controls the trade-off: when the overall conflict is high, e−K^e^- K becomes small and the rule places greater emphasis on the conservative term, thereby pushing more mass toward NFEs and retaining uncertainty for subsequent interval construction; when the conflict is mild, the rule increasingly relies on the consensus term, yielding a sharper and more decisive fused belief assignment. Having obtained the final fused mass function through the proposed Definition 18, a principled decision rule is still required to output a unique hypothesis. This step is nontrivial in the presence of composite FEs, since the fused evidence may preserve a nonnegligible amount of epistemic uncertainty. To this end, we adopt a belief interval-based decision strategy, which directly exploits the confidence bounds induced. The resulting score provides a measure for each singleton hypothesis and enables a deterministic decision. Definition 19. Decision rule induced by belief intervals. Let m⊕(S∣x)m (S x ) denote the final fused mass function obtained by the proposed hybrid combination rule, and let Bel(⋅∣x)Bel (· x ) and Pl(⋅∣x)Pl (· x ) be the standard belief and plausibility functions induced by m⊕(S∣x)m (S x ) . The decision score for each singleton hypothesis SiS_i is defined as Pti(Si∣x)=(1−Pl(Si∣x)−Bel(Si∣x)Plmax−Belmin)×(Pl(Si∣x)−Bel(Si∣x))+(Bel(Si∣x)−Belmin) splitPti (S_i x )&= (1- Pl (S_i x )-Bel (S_i x )Pl_ -Bel_ )\\ & × (Pl (S_i x )-Bel (S_i x ) )+ (Bel (S_i x )-Bel_ ) split (27) and the final decision is made by y^(x)=argmaxSi=θiPti(Si∣x). y (x )= _S_i= \θ_i \Pti (S_i x ). (28) This decision rule provides an interval-aware transformation. Specifically, BelminBel_ and PlmaxPl_ denote the minimum credibility and the maximum plausibility among all singleton hypotheses respectively. Global spread Plmax−BelminPl_ -Bel_ serve as a common reference scale for the current fused evidence. The coefficient 1−(Pl(Si∣x)−Bel(Si∣x))/(Plmax−Belmin)1- (Pl (S_i x )-Bel (S_i x ) )/ (Pl_ -Bel_ ) is then introduced as a stability indicator of the credibility interval: it decreases monotonically with the interval width Pl(Si∣x)−Bel(Si∣x)Pl (S_i x )-Bel (S_i x ) , thereby penalizing hypotheses whose support is highly ambiguous or weakly concentrated. Accordingly, Pti(Si)Pti (S_i ) combines two complementary contributions. The term Bel(Si∣x)−BelminBel (S_i x )-Bel_ quantifies the relative advantage of the lower bound, reflecting the amount of commitment that is guaranteed for SiS_i beyond the weakest supported hypothesis. The second term, (1−Pl(Si∣x)−Bel(Si∣x)Plmax−Belmin)⋅(Pl(Si∣x)−Bel(Si∣x)) (1- Pl(S_i x)-Bel(S_i x)Pl_ -Bel_ )·(Pl(S_i x)-Bel(S_i x)), incorporates the potential support upper bound while adaptively discounting it by the interval stability. This design prevents overly optimistic decisions driven by large plausibility values that are accompanied by wide uncertainty intervals, and at the same time preserves the informative role of plausibility when the evidence is consistent and well-concentrated. As a result, the proposed score yields a robust Bayesian-style decision criterion that remains well-defined in the presence of non-specific focal elements, without forcing an artificial redistribution of NFE mass onto singleton hypotheses. 4 Algorithmic design This section presents the complete algorithmic workflow of the proposed framework. The procedure consists of two distinct phases, namely offline training on historical experience (Algorithm 1) and online inference (Algorithm 2), which correspond to the learning and utilization of historical experience, respectively. Input: Historical dataset ℋ=(xt,yt)t=1NH=\(x_t,y_t)\_t=1^N; evidence sources M=m1,…,mhM=\m_1,…,m_h\ defined on FoD Θ ; sensitivity parameters η,γη,γ Output: Context-specific historical experience weights JRW~ici=1h,c∈\ JRW_i^c\_i=1^h,\,c /* Phase 1: Context partition via spectral clustering */ Construct BPA feature matrix ∈ℝN×(h⋅|2Θ|)B ^N×(h·|2 |) from all evidence sources; 1 Perform spectral clustering on B to obtain context partition =c1,…,cLC=\c_1,…,c_L\; 2 /* Phase 2: Regret--rejoice computation on historical samples */ foreach sample (xt,yt)∈(x_t,y_t) do 3 Fuse all evidence via Dempster’s rule: m⊕(⋅∣xt)←m1⊕⋯⊕mhm (· x_t)← m_1 ·s m_h; 4 y^t←argmaxθk∈ΘBelm⊕(θk∣xt) y_t← _ _k∈ Bel_m ( _k x_t); 5 if y^t=yt y_t=y_t then δt←0 _t← 0 // correct decision, skip; 6 else δt←1 _t← 1; 7 if δt=1 _t=1 then 8 Identify optimal evidence: mbest←argminmi∈MK^(mi,ptarget)m_best← _m_i∈ M K(m_i,p_target) ; // Eq. 18 Compute rejoice value: j(mbest)←1−exp(−ηK^(mbest,m⊕))j(m_best)← 1- \! (-η\, K(m_best,m ) ) ; // Eq. 20 foreach candidate erroneous evidence me∈Mm_e∈ M do 9 Compute error-removal combination m⊕,−em ,-e via Eq. 19; 10 Compute regret value: r(me)←1−exp(−γK^(m⊕,−e,m⊕))r(m_e)← 1- \! (-γ\, K(m ,-e,m ) ) ; // Eq. 21 /* Phase 3: Aggregate and normalize historical weights */ foreach context c∈c do 11 foreach evidence source mi∈Mm_i∈ M do 12 JRWic←∑t=1δt=1,c(xt)=cNjt(mi)−∑t=1δt=1,c(xt)=cNrt(mi)JRW_i^c← _ subarrayct=1\\ _t=1,\,c(x_t)=c subarray^Nj_t(m_i)- _ subarrayct=1\\ _t=1,\,c(x_t)=c subarray^Nr_t(m_i) ; // Eq. 23 Normalize: JRW~ic←exp(JRWic)∑j=1hexp(JRWjc) JRW_i^c← (JRW_i^c) _j=1^h (JRW_j^c) for all i ; // Eq. 24 return JRW~cii=1h,c∈\ JRW_i^c\_i=1^h,\,c Algorithm 1 Offline Training: Historical Experience Modeling In the training phase (historical experience modeling): given a historical dataset ℋ=(xt,yt)t=1NH= \ (x_t,y_t ) \_t=1^N , SC is first performed on the BPA feature representations of individual samples to partition the heterogeneous decision space into L context scenarios C=c1,…,cLC= \c_1,…,c_L \ . For each historical sample xtx_t , the BPA results provided by multiple evidence sources M=m1,…,mhM= \m_1,…,m_h \ are fused via Dempster’s combination to obtain the aggregated result m⊕(⋅∣xt)m (· x_t ) . When the fused decision conflicts with the ground truth label yty_t , the regret-rejoice computation is triggered: the optimal evidence m best m_ best is identified according to Eq. 18, and its rejoice value j(m best )j (m_ best ) is calculated using Eq. 20; subsequently, for each candidate erroneous evidence me∈Mm_e∈ M , the regret value r(me)r (m_e ) is computed as per Eq. 21. The cumulative historical score JRWicJRW_i^c of each evidence within scenario c is defined as the algebraic sum of rejoice and regret, which is then normalized via softmax to yield the scenario-specific historical experience weight. In the inference phase (evidence fusion and decision-making): for a test sample xx , its belonging cluster c(x)c (x ) is first determined, and the corresponding historical experience weights JRWic(x)~ JRW_i^c(x) are retrieved. The CCM matrix of the current evidence set is computed to obtain the GCCD K K . The historically weighted evidence mwm_w is constructed according to Eq. 25, followed by the application of the hybrid combination rule given in Eq. 3, which adaptively adjusts the weight allocation between the conflict and consensus terms via e−K^e^- K . The mwm_w is then processed through the belief intervals interval decision rule defined in Eqs. 27-28 to output the final class label. Input: Test sample x; evidence sources M=m1,…,mhM=\m_1,…,m_h\ on FoD Θ ; historical weights JRW~ic\ JRW_i^c\; sensitivity parameters η,γη,γ Output: Predicted class label y y Determine cluster membership: c(x)←c(x)← assign x to nearest context in C; 1 Retrieve historical weights JRW~ic(x)i=1h\ JRW_i^c(x)\_i=1^h; 2 /* Step 1: Compute CCM and GCCD */ foreach pair (mi,mj)(m_i,m_j), 1≤i<j≤h1≤ i<j≤ h do 3 Compute association k(mi,mj)k(m_i,m_j) via Eq. 12; 4 Compute similarity S(mi,mj)S(m_i,m_j) via Eq. 14; 5 Compute chaos-conflict K^(mi,mj)←1−S(mi,mj) K(m_i,m_j)← 1-S(m_i,m_j) ; // Eq. 2 K^←2h(h−1)∑i<jK^(mi,mj) K← 2h(h-1) _i<j K(m_i,m_j) ; // GCCD, Eq. 16 /* Step 2: Construct historical-experience weighted evidence */ mw(S∣x)←∑i=1hJRW~ic(x)⋅mi(S∣x),∀S⊆2Θm_w(S x)← _i=1^h JRW_i^c(x)· m_i(S x), ∀ S 2 ; // Eq. 25 /* Step 3: Hybrid combination rule */ α←1−e−K^α← 1-e^- K ; // conflict weight β←e−K^β← e^- K ; // consensus weight mconj(S)←∑A1∩⋯∩Ah=S≠∅∏i=1hmi(Ai)m_conj(S)← _A_1∩·s∩ A_h=S≠ _i=1^hm_i(A_i) ; // conjunctive term mdisj(S)←∑A1∩⋯∩Ah=∅A1∪⋯∪Ah=S∏i=1hmi(Ai)m_disj(S)← _ subarraycA_1∩·s∩ A_h= \\ A_1∪·s∪ A_h=S subarray _i=1^hm_i(A_i) ; // disjunctive term m⊕(S)←α⋅(mconj(S)+mdisj(S))+β⋅mw(S)m (S)←α·(m_conj(S)+m_disj(S))+β· m_w(S) ; // Eq. 3 /* Step 4: Belief-interval decision */ foreach singleton hypothesis θk∈Θ _k∈ do 6 Bel(θk∣x)←∑A⊆θkm⊕(A∣x)Bel( _k x)← _A \ _k\m (A x); 7 Pl(θk∣x)←∑A∩θk≠∅m⊕(A∣x)Pl( _k x)← _A∩\ _k\≠ m (A x); 8 Belmin←minθkBel(θk∣x)Bel_ ← _ _kBel( _k x); Plmax←maxθkPl(θk∣x)Pl_ ← _ _kPl( _k x); 9 foreach θk∈Θ _k∈ do 10 Pti(θk∣x)←(1−Pl(θk∣x)−Bel(θk∣x)Plmax−Belmin)(Pl(θk∣x)−Bel(θk∣x))+(Bel(θk∣x)−Belmin)Pti( _k x)← (1- Pl( _k x)-Bel( _k x)Pl_ -Bel_ )\! (Pl( _k x)-Bel( _k x) )+ (Bel( _k x)-Bel_ ) ; // Eq. 27 y^←argmaxθkPti(θk∣x) y← _ _kPti( _k x) ; // Eq. 28 return y y Algorithm 2 Online Inference: Robust Evidence Fusion and Decision 5 Experiment 5.1 Experimental settings 5.1.1 Datasets Table 1 summarises the 16 real world datasets used in the experiments. All were obtained from the UCI Machine Learning Repository and the NIH-CIP Repository, which collectively cover a broad spectrum of characteristics: high or low dimensional feature spaces, binary or multi class targets, and balanced or imbalanced distributions. The data covers fields such as biology, medicine, geography, and demography analysis. To prevent scale-dependent bias in the classifiers, every dataset was standardized before use. Table 1: Experimental datasets Datasets Sample Class (Ratio) Feature (Continuous/Discrete) Accent 329 6(0.5:0.14:0.09:0.09:0.09:0.09) 12(12/0) BMNP 63 3(0.62:0.25:0.13) 36(16/20) Diabetes 520 2(0.62:0.38) 16(1/15) Forest 523 4(0.37:0.3:0.16:0.16) 27(27/0) German 1000 2(0.7:0.3) 24(3/21) Hfcr 299 2(0.68:0.32) 12(8/4) Ionosphere 351 2(0.64:0.36) 34(32/2) ISPY 215 2(0.73:0.27) 18(4/14) Sonar 208 2(0.53:0.47) 60(60/0) Sports 1000 2(0.64:0.36) 59(49/10) Turkish 400 4(0.25:0.25:0.25:0.25) 50(50/0) Vehicle 846 4(0.26:0.26:0.25:0.24) 18(18/0) Wine 178 3(0.33:0.4:0.27) 13(13/0) WDBC 569 2(0.63:0.37) 30(30/0) WPBC 198 2(0.76:0.24) 33(31/2) Zoo 101 7(0.41:0.2:0.13:0.1:0.08:0.05:0.04) 16(0/16) 5.1.2 Baseline and hyperparameter settings To demonstrate the superior inference capability of the proposed framework, we benchmark it against two representative classes of methods: evidential reasoning and ensemble learning. Within the evidential reasoning category, the comparative methods include Dempster’s rule (8), Yager’s rule (40), Dubois’ rule (12), Murphy’s averaging rule (24), Deng’s weighted rule (9), Belief Hellinger Fusion method (BHF) (47), EC-FMDS (5), and the Hierarchical Evidence Fusion method (HEF) (43). For ensemble learning, comprehensive comparative experiments are conducted against three gradient boosting methods, namely eXtreme Gradient Boosting (XGBoost) (7), Light Gradient Boosting Machine (LightGBM) (19), and Categorical Boosting (CatBoost) (27). To ensure a fair comparison, given that the ensemble learning methods utilize decision trees (DTs) as base learners, we similarly employ DTs as evidence bodies within the DST framework. Random split selection is adopted to introduce necessary stochasticity into each evidence body (i.e., each individual DT), while hyperparameters such as tree depth are configured according to recommendations to promote generalization (11). During the comparative experiments, the number of base DTs in the ensemble methods is set equal to the number of evidence bodies in the DST methods. Detailed parameter settings are provided in Table 2; parameters not listed in the table follow the default values of the official Python libraries. Table 2: Hyperparameter settings Type Methods Hyperparameter settings Evidence DT splitter: random, max_depth: 8, min_samples_leaf: 2, min_samples_split: 4, max_features: ”sqrt”, class_weight: ”balanced” Cluster SC n_clusters: number of categories, affinity: rbf, gamma: 1, assign_labels: kmeans, eigen_solver:arpack, n_init: 20 Ensemble learning CatBoost learning_rate: 0.1, depth: 8, l2_leaf_reg: 1 LightGBM learning_rate: 0.1, max_depth: 8, subsample: 1.0, colsample_bytree: 1.0 XGBoost learning_rate: 0.1, max_depth: 8, subsample: 1.0, colsample_bytree: 1.0, reg_lambda:1e-5 DST Dempster - Deng Distance: Jousselme Dubois - EC-FMDS q=2 HEF - Murphy - Yager - BHF - Our η: 0.5, γ: 0.5, |C||C|: number of class 5.1.3 Procedure and metrics All experiments were conducted on a system with an Intel Core Ultra 5 225H CPU, 32 GB RAM, and Windows 11 25H2. The code was implemented in Python 3.9.25, and all stochastic procedures were fixed with a random seed of 42 to ensure reproducibility. To ensure fair and comprehensive evaluation, all methods were assessed under a unified experimental protocol across the 16 datasets listed in Table 1. Each dataset was randomly partitioned using 5-fold cross-validation. Reported results are presented as mean values, standard deviations, and rankings (where applicable) across all folds to ensure statistical reliability and robustness. To further validate the proposed method, we conducted additional analyses encompassing robustness evaluation under noisy conditions, adjustment of the number of evidence bodies, and substitution of the evidence body base model, together with sensitivity analysis and ablation studies of hyperparameters. These complementary experiments provide deeper insights into the stability, generalization capability, and component contributions of the proposed framework. Detailed experimental procedures are presented in the corresponding sections below. For comprehensive assessment of classification performance, we employed four widely adopted evaluation metrics: accuracy (ACC), precision (PRE), recall (REC), and F1-score (F1). ACC measures the overall proportion of correctly classified samples, whereas PRE and REC evaluate the correctness and completeness of positive predictions, respectively. The F1, defined as the harmonic mean of precision and recall, offers a balanced assessment particularly suited to imbalanced datasets. We adopted macro averaging for all metrics. Additionally, receiver operating characteristic (ROC) curves and the area under the curve (AUC) were employed to evaluate the discriminative capability of different methods. The predictions from all folds were pooled to generate a single ROC curve and calculate the AUC. 5.2 Results 5.2.1 Comparative analysis This section uniformly employed three decision trees as evidence sources to construct the DST framework, with the number of base learners matching that of the compared ensemble methods. Table 3 summarizes the mean performance, standard deviation, and average rank of each method in terms of ACC, PRE, REC, and F1. Our method achieved optimal mean values across all evaluation metrics (ACC: 87.01, PRE: 85.44, REC: 87.51, F1: 85.78), with average ranks of 3.31, 3.52, 2.89, and 3.22, respectively. Detailed performance results for each dataset are provided in Appendix A. Compared with the DST model, the proposed framework demonstrates superior capability in quantifying and resolving high-conflict evidence. The performance degradation of Dempster’s rule (ACC: 80.30, F1: 76.03) stems from its rigid normalization mechanism, which tends to produce counterintuitive fusion results. Yager’s approach reallocates conflict mass to the frame of discernment; however, this conservative strategy indiscriminately discards valuable informative discrimination signals, manifesting as the second-lowest F1 score among all methods (74.87). Murphy’s average rule enhances stability through uniform evidence weighting, yet its assumption of source consistency fails to account for context-dependent reliability variations; despite its smoothing effect, it shows no substantial advantage over Dempster’s rule (F1: 77.74 vs. 76.03), indicating inadequate handling of such uncertainty. Deng’s method introduces source differentiation through Jousselme distance-based weighting, but measures only static BPA discrepancies without modeling the internal uncertainty structure of FEs, limiting its adaptability under epistemic uncertainty. The BHF method employs belief Hellinger distance to quantify evidence dissimilarity and incorporates entropy to characterize ambiguity, thereby establishing a weighted average fusion strategy. However, its weight generation relies exclusively on inter-evidence distance metrics and intrinsic uncertainty, without leveraging historical decision feedback; consequently, it fails to capture reliability variations of evidence sources across diverse decision contexts. EC-FMDS attempts to integrate feature selection preprocessing with fusion mechanisms, yet their independent operational modes cannot resolve conflicts arising from evidence-level inconsistency, as corroborated by its substantial performance ranking variance. HEF’s hierarchical structure mitigates conflict and reduces computational complexity through averaged evidence sources; however, its strict fusion sequence and evidence averaging mechanism may lead to premature determination of unreliable averaged evidence system combinations when evidence bodies are scarce and homogeneous (e.g., three decision trees in this experiment), evidenced by its lowest F1 (75.36). Compared with gradient boosting frameworks, the proposed method exhibits marked superiority in F1 score (85.78 versus 75.05, 77.87, and 78.35). Although all three ensemble methods construct additive models through gradient optimization to improve fitting accuracy by minimizing empirical risk, they lack explicit modeling of inter-evidence uncertainty. In contrast, the proposed framework characterizes decision boundary uncertainty through belief and plausibility functions, thereby preserving more discriminative information in regions with imbalanced class distributions or overlapping feature spaces and achieving superior predictive performance. The proposed framework overcomes these limitations through a unified CCM mechanism. CCM simultaneously quantifies cross-evidence conflict and intra-evidence non-specificity without presupposing fusion order or structural assumptions; furthermore, it introduces historical experience-driven weighting based on regret theory, enabling evidence reliability assessment to adapt dynamically to current sample conditions and maintaining fine-grained reliability discrimination even under evidence-scarce scenarios. This end-to-end design avoids the decoupling of preprocessing and fusion, achieving collaborative optimization of evidence quality assessment and fusion decision. According to the Friedman test results and Nemenyi post-hoc test (CD = 2.08) presented in Figure 4, the proposed framework achieved the highest average rank among all compared methods, demonstrating statistically significant performance advantages over every alternative approach. Among the competing methods, LightGBM, Murphy, XGBoost, EC-FMDS, and Deng formed a relatively competitive cluster with no statistically significant differences detected within this group, whereas Dempster, Dubois, CatBoost, BHF, Yager, and HEF occupied lower overall rankings. Table 3: Performance comparison (mean ± std) of different methods with average rank Type Model ACC ACC_rank PRE PRE_rank REC REC_rank F1 F1_rank Ensemble learning CatBoost 80.21±13.30 7.69±4.78 77.36±17.02 8.44±4.76 74.93±18.11 10.52±4.58 75.05±18.25 9.94±4.68 LightGBM 83.10±9.90 4.39±3.32 79.08±15.41 4.95±3.53 78.00±15.00 5.53±3.60 77.87±15.54 5.28±3.58 XGBoost 82.26±11.20 4.97±2.98 79.15±14.50 5.48±3.03 78.48±14.74 5.50±3.20 78.35±14.83 5.36±3.15 DST Dempster 80.30±12.90 8.64±4.16 77.44±17.30 9.72±4.45 76.61±16.80 9.84±4.15 76.03±17.56 10.09±4.13 Deng 80.44±13.55 8.11±4.25 77.50±16.02 9.31±4.65 77.66±16.06 8.70±4.36 77.22±16.19 8.92±4.38 Dubois 80.18±12.57 8.52±4.07 76.60±15.91 10.11±4.05 76.85±15.76 9.44±4.07 76.20±16.06 9.58±3.99 EC-FMDS 80.31±14.06 5.48±3.01 77.55±16.13 6.08±3.24 77.79±16.09 5.41±3.19 77.14±16.36 5.62±3.25 HEF 78.51±13.33 9.98±4.35 75.36±16.36 10.64±4.10 74.36±16.54 11.33±3.83 73.63±17.15 11.53±3.77 Murphy 80.89±13.97 5.48±3.26 78.35±16.06 5.89±3.26 78.11±16.56 5.47±2.97 77.74±16.61 5.55±3.01 Yager 78.80±13.54 7.56±3.56 75.65±16.59 7.70±3.67 75.45±16.31 7.58±3.58 74.87±16.86 7.83±3.51 BHF 79.34±13.84 9.08±4.04 75.89±17.12 10.34±3.88 76.80±16.55 9.38±3.83 75.87±16.95 9.88±3.92 Our 87.01±13.93 3.31±3.78 85.44±15.29 3.52±3.90 87.51±14.75 2.89±3.51 85.78±15.42 3.22±3.77 Figure 4: CD diagram of all models. Figure 5 presents the ROC curves for all methods across the 16 datasets. The proposed method (red curve) consistently occupies the uppermost position in the majority of cases, exhibiting substantially higher true positive rates at low false positive rates. This indicates superior discriminative capability and well-calibrated uncertainty estimation, particularly in scenarios characterized by ambiguous decision boundaries or class imbalance. Closer examination reveals two distinct patterns. On datasets with inherently high separability, such as Diabetes, ISPY, and Wine, performance differences among methods diminish considerably, suggesting that simple evidence aggregation strategies suffice when classification tasks are relatively straightforward. Conversely, on challenging datasets with substantial class overlap, including BMNP, German, and WPBC, the proposed framework demonstrates marked performance advantages over competing approaches. This pronounced separation in difficult scenarios underscores the efficacy of our uncertainty-aware evidence fusion mechanism. Consequently, the proposed framework achieves the highest overall mean AUC of 93.3, with its robust generalization across diverse datasets attributable primarily to exceptional performance in discriminatively demanding conditions rather than incremental gains on already well-resolved tasks. Figure 5: Comparison of ROC curves for different models on 16 datasets 5.2.2 Robustness analysis To verify the robustness of the proposed framework under diverse conditions, we conducted a series of experiments, which included introducing data noise and varying both the types and the number of evidence sources. 5.2.2.1 Noise To comprehensively evaluate the robustness of the proposed framework under noisy conditions, we introduced multiple types of data noise during the training phase to emulate interference scenarios commonly encountered in real-world applications. Specifically, feature-level noise, label-level noise, and their combinations as hybrid noise were considered to assess the stability and generalization capability of the framework across varying noise intensities. For feature-level perturbations, two noise injection strategies were adopted. The first randomly replaced a proportion of feature values with the corresponding feature means, thereby modeling data missingness during the acquisition process. The second randomly added noise sampled from a zero-mean Gaussian distribution to selected features, simulating stochastic interference with a normal distribution. For label-level noise, a specified ratio of sample labels was randomly reassigned to alternative classes, reflecting typical annotation errors in practical datasets. To further examine the compound effects of concurrent feature and label corruption, two hybrid noise configurations were designed: uniform feature noise combined with random label noise, and Gaussian-distributed feature noise combined with random label noise. Noise intensities were set to 5%, 10%, 15%, and 20%, enabling a systematic analysis of performance degradation as noise levels increased. Figure 6 and Figure 7 present the variations in F1 performance of the compared methods under two mixed-noise scenarios, while Appendix B presents the corresponding F1 results under the three basic noise settings. Overall, as the noise level gradually increases from 5% to 20%, the performance of all methods declines to varying degrees. Nevertheless, the proposed method maintains a relatively high F1 score under most noise settings. This result indicates that the proposed approach is not effective only in relatively ideal data environments; rather, it can still sustain stable decision performance when the training samples are affected by feature contamination, label shifts, or their combined interference. Figure 6: F1-score under hybrid noise (gaussian feature noise and random label) Figure 7: F1-score under hybrid noise (mean feature noise and random label) In contrast to most existing DST-based methods, which typically rely on static processing according to the differences among the current BPAs, they struggle to distinguish anomalous evidence caused by incidental noise from valid evidence with long-term stability. In specific contexts, the proposed method exhibits greater robustness under noise. On the one hand, it employs an adaptive dual-weighting mechanism based on historical experience and current conflict. On the other hand, its hybrid combination rule preserves uncertain mass instead of forcibly compressing it into the SFE, thereby reducing the dominance of anomalous evidence and the associated noise in the final fusion result while maintaining relatively stable performance. It should nevertheless be acknowledged that, compared with ensemble tree-based methods, most DST-based approaches, including the proposed method in some cases, are generally less stable under noisy conditions. This difference is directly related to the distinct modeling priorities of the two methodological paradigms. Ensemble tree methods are fundamentally developed within a supervised learning framework, where classification errors are continuously fitted through iterative empirical risk optimization or the parallel integration of multiple heterogeneous tree structures, and their objective functions are therefore closely aligned with the final decision outcome (36; 1). Ensemble strategies and recursive splitting will mitigate the instability of individual trees and promote complementary effects across learners (17). By contrast, the main strength of DST methods lies in the explicit representation of uncertainty and the management of conflict, rather than in the direct minimization of classification error. The fusion process in evidence theory is typically centered on BPA construction and combination rule design, with the primary emphasis placed on integrating support relationships among multiple evidence sources. However, this process does not inherently guarantee sufficiently strong discriminative learning of class boundaries, and fusion rules mainly redistribute information under existing quality constraints rather than iteratively correcting prediction bias to improve generalization, as ensemble tree models do (29). For this reason, the proposed method further incorporates historical experience modeling and conflict-adaptive fusion within the DST framework, thereby compensating for the limited discriminative learning capability of conventional evidence fusion methods and enabling it to narrow the gap with, or even outperform, ensemble models under noisy conditions. Interestingly, under mixed-noise scenarios, the performance of most methods declines relatively gradually as noise intensity increases, whereas under single-noise conditions, performance exhibits more pronounced fluctuations and irregularity. This observation suggests that perturbations do not necessarily exacerbate model degradation in a purely monotonic manner; rather, they may encourage the formation of more robust decision boundaries, thereby improving adaptability to distribution shifts and indirectly enhancing generalization. This finding is consistent with the conclusions of 21, who reported that moderate noise exposure helps mitigate overfitting and improves model stability in complex and uncertain environments. Figure 8 further illustrates the distribution of AUC values for different methods across varying noise types and intensities, providing a complementary perspective on robustness from the standpoint of ranking-based discriminative ability. Overall, as the noise level increases, the median AUC of most methods shifts downward and the interquartile range tends to widen, indicating that noise affects not only average discriminative performance but also cross-dataset stability. The proposed method maintains a relatively competitive AUC distribution under most noise settings, as evidenced by a higher median and mean, a narrower interquartile range, and fewer outliers (corresponding to markedly poor performance). This suggests that its advantages are not confined to threshold-dependent classification metrics, but also extend to sample-level discriminability. Figure 8 also shows that the boxplots of AUC obtained under mixed-noise conditions appear more scattered than those under single-noise conditions, which is broadly consistent with the pattern observed for F1. Figure 8: Boxplots of average AUC for different models under various noise types and levels 5.2.2.2 Different number of evidences Figure 9 shows that as the number of basic evidences (trees) increases, the performance of most methods generally follows a pattern of improvement followed by stabilization. When the number of evidences is small, individual tree sources exhibit strong randomness, and the fusion result is more susceptible to incidental bias. As the number of evidences increases, both tree diversity and collective stability improve, leading to better classification performance. However, with further increases, the complementary information provided by additional trees gradually diminishes, and the performance gains tend to saturate. Certain approaches, such as Dempster and HEF, fail to handle the increased conflict effectively as the number of evidence sources grows, resulting in performance degradation. By contrast, the proposed method maintains consistently strong and more stable performance across different numbers of evidences, indicating its ability to balance evidence complementarity with conflict suppression effectively. Figure 9: Comparative performance of models across subtree sizes 5.2.2.3 Different base model To further examine the proposed framework’s dependence on the evidence generators, this section keeps all other experimental settings unchanged and replaces the original combination of multiple random DTs used to construct the evidence bodies with a set of seven different base evidence generators: DT, Support Vector Classification (SVC), Logistic Regression (LR), Gaussian Process Classifier (GPC), K-Nearest Neighbors (KNN), Multilayer Perceptron (MLP), and Gaussian Naive Bayes (GNB). All models use the default parameters provided by scikit-learn, and the fusion results of different DST methods are compared accordingly. Table 4 presents the corresponding results. Compared with the relatively stable DT-based setting, replacing the base models introduces greater overall performance fluctuations, indicating that evidence quality, output diversity, and fusion compatibility all affect the final decision. Under this setting, the proposed framework achieves an ACC of 83.42, a PRE of 81.21, a REC of 78.34, and an F1 score of 78.43, with corresponding average ranks of 4.44, 4.91, 4.64, and 4.81. Overall, it remains competitive and outperforms most of the compared DST methods. Although the set of base evidence generators includes several relatively weak models, almost all DST methods achieve favorable performance after evidence fusion. Notably, the proposed framework maintains comparatively stable overall performance after the replacement of base models, as reflected by its lower standard deviations in both performance and ranking, demonstrating the framework’s strong adaptive evidence fusion capability when confronted with heterogeneous base models of varying quality. Table 4: Performance comparison under different base models (mean ± std) with average rank Type Model ACC ACC_rank PRE PRE_rank REC REC_rank F1 F1_rank Base evidence SVC 76.81±17.45 9.25±5.71 76.48±16.14 8.98±5.09 75.21±16.85 8.00±5.43 74.34±17.69 8.22±5.46 LR 75.50±16.66 7.97±4.48 68.75±22.90 8.73±4.18 70.60±17.21 8.31±4.31 67.74±21.09 8.62±4.15 GPC 63.95±21.67 10.64±4.98 60.58±23.39 10.95±4.69 56.61±21.23 11.19±4.81 50.00±25.85 11.61±4.41 DT 78.20±13.42 9.34±5.55 74.27±16.59 9.75±5.39 74.39±15.96 9.06±5.45 73.93±16.41 8.95±5.39 KNN 75.00±14.10 9.67±4.16 64.22±22.77 10.69±4.58 66.10±18.17 10.80±4.35 63.36±20.76 10.97±4.26 MLP 63.95±21.67 10.64±4.98 60.58±23.39 10.95±4.69 56.61±21.23 11.19±4.81 50.00±25.85 11.61±4.41 GNB 80.32±13.44 7.58±5.02 78.30±15.16 7.69±4.71 75.88±16.08 7.69±4.98 76.20±15.97 7.55±4.87 DST Deng 80.51±11.33 7.36±4.73 77.33±15.21 7.67±4.49 75.30±15.02 7.41±4.51 75.09±15.46 7.36±4.46 HEF 80.20±13.50 5.97±3.55 75.84±19.71 6.86±3.55 75.32±16.52 6.41±3.38 73.97±18.88 6.66±3.53 Dubois 78.22±15.74 5.94±3.77 72.91±22.17 6.95±3.50 73.25±16.85 6.52±3.39 71.23±20.31 6.88±3.51 Dempster 80.14±13.49 6.16±3.56 75.79±19.69 7.00±3.58 75.29±16.48 6.56±3.39 73.93±18.85 6.81±3.51 EC-FMDS 82.89±11.03 4.33±2.52 81.38±14.63 5.00±3.07 77.76±14.51 5.02±2.50 77.92±15.13 4.97±2.70 Murphy 78.08±15.29 6.45±4.43 74.02±20.00 6.05±3.86 72.59±17.23 6.88±4.14 71.29±19.61 6.59±4.24 BHF 81.72±11.10 5.52±3.02 81.46±14.36 5.75±3.26 75.77±15.57 6.52±3.02 75.62±16.40 6.48±3.06 Yager 77.70±16.21 6.31±4.10 72.07±22.49 7.36±3.77 72.77±17.23 6.86±3.67 70.66±20.84 7.20±3.76 Our 83.42±10.13 4.44±3.06 81.21±13.71 4.91±2.98 78.34±13.90 4.64±2.85 78.43±14.28 4.81±3.04 5.2.3 Parameter Sensitivity To evaluate the influence of key parameters on the performance of the proposed framework, this section conducts parameter sensitivity analysis from two perspectives: historical context partitioning and experience-weight generation. Specifically, all other hyperparameters are fixed while varying the number of clusters obtained by spectral clustering, and the resulting changes in ACC, PRE, REC, and F1 under different clustering granularities are examined to assess the sensitivity of historical experience modeling to the context partition scale. Subsequently, a joint grid search is performed over the two sensitivity coefficients, η and γ , in the regret-rejoice mechanism. By plotting F1 heatmaps and ACC three-dimensional surfaces across multiple datasets, we further investigate the stability of the experience-weight mechanism and the distribution of high-performing regions under different parameter combinations. Figure 10 presents the performance variation of the proposed model under different numbers of clusters. Overall, model performance does not increase monotonically with the number of clusters, but instead exhibits dataset-dependent local optima associated with sample size, class structure, and feature dimensionality. For small-sample datasets such as BMNP, Wine, and Zoo, the performance curves fluctuate more noticeably, indicating that when limited samples are available for historical experience statistics, changes in the number of clusters directly affect the stability of the regret-rejoice values within each context. An overly fine-grained partition may lead to insufficient samples within individual clusters, making the estimation of experience weights more susceptible to local sample distributions. In contrast, larger datasets such as Diabetes, German, Sports, and Vehicle show relatively stable performance across a wider range of cluster numbers, suggesting that sufficient historical samples can mitigate the statistical uncertainty introduced by changes in context partitioning. In addition, multi-class and imbalanced datasets such as Accent and Zoo are also more sensitive to the number of clusters, possibly because minority-class samples are more easily diluted across different clusters, thereby affecting local reliability modeling. These observations suggest that the role of the cluster number is to balance contextual expressiveness and statistical stability: too few clusters may fail to capture reliability differences across sample regions, whereas too many clusters may amplify estimation fluctuations in small-sample or minority-class scenarios. A moderate clustering granularity is therefore generally more conducive to stable fusion performance. Figure 10: Sensitivity analysis of clustering performance with respect to the number of clusters Figure 11 further reports the joint sensitivity results for η and γ . The F1 heatmaps on the left show that the optimal parameter combinations vary across datasets, indicating that the reward and penalty intensities should be adapted to the sample size, class distribution, and feature structure of each dataset. For datasets with relatively sufficient samples, such as WDBC, Diabetes, Sports, and Vehicle, high performance can be maintained over a relatively broad range of parameter combinations, forming continuous high-performance regions. This suggests that the historical experience weights in these scenarios are relatively tolerant to variations in η and γ . In contrast, small-sample datasets such as BMNP, Wine, WPBC, and Zoo exhibit more pronounced changes in the heatmaps, indicating that when historical samples are limited or class proportions are imbalanced, variations in the reward and penalty coefficients can more easily amplify local statistical errors in experience estimation and thus lead to performance fluctuations. For multi-class datasets such as Accent and Turkish, parameter sensitivity may also be jointly affected by inter-class separability and the reliability estimation of minority classes. The ACC surfaces on the right show a similar pattern: most datasets exhibit local fluctuations rather than a consistent increase along a single parameter direction, suggesting that excessively strong rewards or penalties may disrupt the balance of evidence weights. Overall, the effective regions of η and γ are usually not isolated points but are distributed across several neighboring parameter combinations, demonstrating a certain degree of parameter robustness in the proposed historical experience-driven mechanism. Meanwhile, the differences in optimal regions across datasets indicate that parameter selection should be moderately adjusted according to data scale, class imbalance, and feature complexity, so as to achieve a more appropriate balance between preserving historically reliable evidence and suppressing misleading evidence. Figure 11: Joint sensitivity analysis of hyperparameters η and γ across multiple datasets 5.2.4 Ablation study To examine the independent contribution of the key modules in the proposed framework, we conducted ablation experiments under the same experimental setting as that used in Section 5.2.1. The ablation variants include: replacing CCM with Dempster’s conflict coefficient K; removing the clusters obtained by spectral clustering; separately removing the weights associated with rejoice and regret; completely removing historical experience weighting; retaining only the Dubois combination term or only the historical experience weighting term; and replacing the proposed decision rule with pignistic decision. All results are summarized in terms of the average F1 and AUC over the 16 datasets. The absolute and relative changes with respect to Full are also reported to characterize the influence of each component on classification performance and discriminative ability. As shown in Table 5, the Full model achieves the best performance in both F1 and AUC. From the perspective of module contribution, historical experience weighting has the most pronounced impact. w/o History Weighting leads to a 3.21% decrease in F1 and a 5.03% decrease in AUC, producing the largest AUC loss among all variants. This suggests that historical feedback improves not only the final classification results but also the framework’s overall ability to discriminate among evidence sources. In comparison, w/o Rejoice and w/o Regret result in F1 losses of 1.54% and 1.98%, respectively, indicating that both positive reinforcement and negative penalty contribute to reliability estimation. The larger effect of the regret term suggests that suppressing misleading evidence may be more critical in conflict-aware evidence fusion. w/o Cluster decreases F1 and AUC by 2.05% and 1.57%, respectively, further demonstrating the contextual dependency of evidence-source reliability. The degradation observed for w/ Dempster K indicates that the traditional conflict coefficient is insufficient to fully capture the internal uncertainty introduced by non-singleton focal elements, whereas the proposed CCM can more finely characterize both inter-evidence conflict and evidence non-specificity. Dubois Only suffers the largest F1 loss, suggesting that relying solely on conservative uncertainty preservation weakens the decisiveness of final classification. Although History Weighted Only can exploit historical reliability information, it still fails to achieve the stability of the full model without Dubois-type uncertainty preservation and conflict-adaptive regulation; consequently, its AUC degradation is larger than that of Dubois Only. Overall, the ablation results support the core design logic of the proposed framework: CCM evaluates the conflict and uncertainty of the evidence set, historical experience weighting characterizes the long-term reliability of different evidence sources under similar contexts, and the hybrid fusion rule adaptively balances reliable consensus with uncertainty preservation. These components jointly improve the classification and discriminative performance of the model across heterogeneous datasets. Table 5: Ablation study results averaged over 16 datasets Variant F1 Δ 1 (abs.) Δ 1 (rel.) AUC Δ (abs.) Δ (rel.) Full Model 85.78±15.42 - - 93.30±5.39 - - w/ Pignistic 84.62±15.63 -1.16 -1.35% 92.91±5.83 -0.39 -0.42% w/ Dempster K 84.31±16.03 -1.47 -1.71% 92.14±5.75 -1.16 -1.24% w/o History Weighting 83.03±15.56 -2.75 -3.21% 88.61±5.87 -4.69 -5.03% w/o Cluster 83.73±16.17 -2.05 -2.39% 91.73±5.83 -1.57 -1.68% w/o Regret 84.08±16.02 -1.70 -1.98% 92.08±5.78 -1.22 -1.31% w/o Rejoice 84.46±15.25 -1.32 -1.54% 92.51±5.79 -0.79 -0.85% Dubois Only 82.97±16.02 -2.81 -3.28% 92.10±5.82 -1.20 -1.29% History Weighted Only 83.69±16.20 -2.09 -2.44% 91.36±5.72 -1.94 -2.08% 6 Conclusion This paper has presented a unified evidence reasoning framework that combines CCM with historical-experience-driven weighting for robust multi-source decision making under uncertainty. The central premise of this work is that two persistent gaps in existing DST methods, namely the treatment of conflict and uncertainty as independent quantities and the neglect of long-term evidence source behavior, can be jointly addressed through complementary statistical and behavioral mechanisms. The CCM provides a unified scalar assessment of both inter-evidence inconsistency and intra-evidence non-specificity, grounded in a similarity measure whose mathematical properties are formally established. The historical experience weighting scheme exploits the observation that evidence sources operating repeatedly across heterogeneous decision contexts carry learnable reliability profiles, which RT makes it possible to quantify in terms of counterfactual regret and rejoice when fusion decisions deviate from ground truth. These two mechanisms are integrated through a hybrid combination rule that adaptively balances uncertainty preservation against weighted consensus, followed by a belief-interval decision strategy that produces deterministic classifications without discarding the epistemic uncertainty preserved by the fusion process. The theoretical contributions of this work are threefold. First, the CCM offers a more faithful characterization of evidential stability than existing measures that examine conflict and uncertainty in isolation, and its five proven properties ensure consistent behavior across frame refinements and extreme cases. Second, the integration of SC with RT establishes a principled paradigm for translating historical decision feedback into context-dependent evidence weights, extending the scope of DST beyond instantaneous assessment toward adaptive reliability modeling. Third, the hybrid combination rule provides a conflict-aware mechanism for balancing the conservatism of fusion with the decisiveness of weighted consensus, parameterized by a single GCCD conflict indicator that requires no manual tuning of mixture coefficients. On the practical side, experiments conducted across 16 heterogeneous datasets demonstrate that the proposed framework achieves competitive performance among both DST-based and ensemble learning methods, with the highest average F1 score (85.78) and mean AUC (93.30). The framework maintains its advantage under noisy training conditions, varying numbers of evidence sources, and different base evidence generators, indicating that the gains stem from the fusion architecture rather than from favorable data conditions. Ablation analysis confirms that historical experience weighting, CCM-based conflict assessment, and the hybrid combination rule each contribute meaningfully to performance, with the historical experience component exerting the strongest individual effect. Several limitations suggest directions for further research. The current framework constructs BPAs through standard classifiers rather than dedicated evidence generation methods, and the quality of these assignments directly affects downstream fusion performance (45); developing domain-specific BPA generation strategies could improve both accuracy and computational efficiency. The computational cost of the CCM scales with the number of focal element pairs and may become prohibitive when the frame of discernment is large or the number of evidence sources is high (6); approximate or distributed computation strategies would extend the framework’s applicability to real-time decision settings. The sensitivity parameters governing the regret-rejoice mechanism, while demonstrating reasonable robustness across the datasets examined, may benefit from data-adaptive calibration rather than uniform specification. Future work should also explore the extension of the historical experience paradigm to heterogeneous multi-modal data environments, where evidence sources of fundamentally different types must be fused, and to nonstationary settings where the decision context itself evolves over time (34; 43). Appendix A Table A.1: Comparison of ACC on 16 datasets Model BHF CatBoost Dempster Deng Dubois EC-FMDS HEF LightGBM Murphy Our XGBoost Yager Accent 69.01±3.95 60.79±4.36 67.47±4.33 69.00±2.00 65.96±1.64 65.95±2.31 62.00±3.52 72.94±3.47 66.57±6.11 69.05±12.42 71.41±5.50 60.82±6.98 BMNP 49.06±6.72 55.62±4.15 58.85±10.26 52.40±7.82 57.19±5.44 47.71±4.77 55.73±10.67 66.67±5.89 54.27±17.33 71.35±34.12 61.77±8.15 52.40±7.82 Diabetes 95.77±0.77 93.08±2.26 93.85±1.88 95.58±3.40 96.73±1.46 97.31±0.77 96.73±1.71 91.15±1.47 95.00±0.99 97.69±1.26 94.23±2.55 95.96±0.74 Forest 84.89±2.93 86.22±3.31 86.23±4.71 86.03±2.94 87.56±2.50 85.46±3.19 82.98±4.40 87.75±4.89 85.46±4.34 88.72±5.11 88.14±3.59 84.71±1.36 German 66.60±1.80 73.20±2.38 69.50±4.12 68.70±1.74 70.10±2.76 69.20±4.88 68.60±2.58 74.40±2.19 70.00±2.79 74.80±3.68 74.00±4.36 68.20±2.39 Hfcr 79.95±3.85 79.59±3.89 82.96±7.38 81.29±4.80 80.94±7.79 83.63±5.85 79.60±7.25 84.96±7.54 81.95±6.16 84.29±4.22 80.95±8.50 82.62±4.70 Ionosphere 89.18±4.66 89.46±3.28 90.33±5.03 91.75±2.97 88.90±3.73 91.75±4.84 90.03±3.76 91.46±3.51 93.17±3.33 94.03±5.44 89.76±4.62 89.47±6.02 ISPY 96.73±1.82 99.07±1.08 95.80±2.83 95.35±1.84 94.40±3.45 92.49±10.24 92.57±5.44 99.07±1.08 97.67±1.79 94.38±5.14 99.07±1.08 90.23±3.81 Sonar 71.63±1.84 73.08±4.15 68.75±5.06 77.40±4.54 74.04±4.00 75.96±7.11 70.67±5.95 77.88±5.98 76.92±6.84 90.87±12.10 75.48±3.28 75.48±6.35 Sports 77.70±3.24 80.30±1.71 78.30±2.68 77.20±1.18 77.20±3.12 79.30±2.78 77.10±2.43 80.50±3.17 80.30±3.88 87.70±7.14 79.20±1.57 74.80±3.41 Turkish 67.25±2.22 65.50±4.04 67.75±5.85 67.25±3.59 66.25±4.11 67.75±3.40 63.50±5.80 78.25±2.50 69.25±0.96 86.00±14.21 74.50±4.80 66.50±7.59 Vehicle 73.41±2.80 68.56±1.17 69.74±1.52 72.82±1.35 70.93±4.20 74.00±1.32 68.20±2.39 72.93±2.02 73.88±2.64 77.91±7.89 73.41±1.88 69.38±2.51 WDBC 95.25±1.56 94.55±1.84 93.67±1.73 95.43±1.35 93.15±1.04 94.73±1.67 93.67±2.64 95.08±2.99 94.90±2.33 98.42±1.05 94.03±3.06 95.08±0.99 Wine 93.23±4.93 94.95±3.86 93.80±5.00 95.52±3.18 93.24±3.19 94.37±6.01 89.87±4.37 92.68±3.88 96.06±3.88 98.89±2.22 93.27±3.17 88.79±4.45 WPBC 68.71±6.12 79.29±4.55 73.74±4.33 68.17±7.87 73.23±4.14 73.24±4.36 73.73±3.72 75.78±5.10 66.70±6.45 80.89±14.21 74.78±3.96 73.23±1.88 Zoo 91.12±6.77 90.04±9.55 94.04±6.94 93.15±6.56 93.08±1.95 92.12±7.22 91.12±3.71 88.15±4.45 92.12±5.55 97.12±5.77 92.12±4.49 93.15±8.03 Table A.2: Comparison of PRE on 16 datasets Model BHF CatBoost Dempster Deng Dubois EC-FMDS HEF LightGBM Murphy Our XGBoost Yager Accent 62.90±3.39 60.51±13.33 68.26±4.36 63.51±3.49 60.13±2.21 61.48±3.65 63.55±4.92 68.89±8.55 60.92±7.25 66.04±11.57 68.75±4.72 54.61±8.64 BMNP 34.28±7.75 34.31±4.61 36.71±23.19 45.16±10.35 43.11±12.96 42.34±11.65 37.67±15.76 39.06±6.48 51.11±24.70 68.65±36.95 45.33±12.11 41.44±10.14 Diabetes 95.30±0.70 92.53±2.56 93.21±2.03 95.31±3.84 96.31±1.72 97.04±0.91 96.32±1.97 91.00±1.90 94.50±1.19 97.38±1.56 93.80±2.81 95.48±0.94 Forest 84.95±1.90 86.59±2.94 86.71±4.58 86.30±1.56 87.63±0.91 85.38±2.55 85.36±4.41 88.39±3.66 84.88±3.64 89.07±5.92 88.67±2.77 85.00±2.77 German 62.54±3.11 67.56±3.83 64.73±4.11 63.09±2.58 65.24±3.25 64.26±4.79 63.51±3.25 69.15±3.25 65.38±2.90 71.69±3.81 68.74±5.56 63.63±2.55 Hfcr 77.46±4.67 77.76±4.20 80.63±8.30 78.99±5.00 78.04±8.99 81.13±6.62 76.92±8.25 82.96±8.59 79.12±6.88 82.32±4.50 78.34±9.80 80.61±4.83 Ionosphere 89.07±4.99 89.76±3.52 89.52±5.46 90.85±2.64 88.11±4.05 91.21±5.14 88.77±3.78 91.88±4.10 93.05±3.37 93.52±6.04 90.16±5.40 89.07±6.16 ISPY 95.35±3.03 99.38±0.72 96.01±2.76 94.55±1.37 93.56±4.54 92.03±12.77 93.43±6.98 99.38±0.72 97.00±2.82 92.06±5.94 99.38±0.72 88.67±3.58 Sonar 71.58±1.91 74.12±2.66 70.23±5.61 77.40±4.55 74.58±3.34 76.37±7.39 71.77±7.24 78.24±6.35 77.05±6.96 91.00±12.12 75.46±3.31 75.47±6.28 Sports 75.98±3.43 79.38±2.17 77.42±2.88 75.58±1.26 75.80±3.66 77.77±2.98 75.49±2.79 79.27±3.30 78.76±4.19 86.68±7.64 77.81±1.93 72.90±3.79 Turkish 68.04±3.38 66.44±4.37 68.36±6.03 67.25±3.99 67.33±3.03 69.78±2.62 66.48±4.82 79.41±2.24 69.63±1.00 86.04±14.46 74.80±4.44 69.87±7.70 Vehicle 73.11±3.37 66.73±0.72 68.92±2.02 72.32±1.46 70.50±4.96 73.56±1.70 67.43±2.83 71.88±2.65 73.23±2.77 77.42±8.69 73.03±1.88 68.61±2.48 WDBC 94.92±1.80 94.72±1.95 93.00±1.99 95.08±1.59 92.46±1.22 94.25±1.90 93.04±3.02 94.92±3.66 94.83±2.78 98.12±1.11 93.45±3.35 94.48±1.02 Wine 93.35±4.99 95.32±3.82 94.71±3.90 95.54±3.16 93.82±2.78 94.69±5.67 91.23±3.67 93.30±4.01 96.18±3.74 98.90±2.21 93.72±3.34 89.59±4.42 WPBC 55.97±8.11 70.35±10.70 60.78±7.40 55.68±9.47 56.80±10.59 60.98±8.46 58.91±8.63 66.32±8.17 56.10±6.07 75.24±16.54 64.07±5.59 57.90±7.36 Zoo 79.46±14.33 82.29±12.16 89.88±14.07 83.33±16.64 82.14±9.86 78.51±16.32 75.95±3.05 71.21±10.93 81.85±12.35 92.86±14.29 80.89±9.89 83.04±22.09 Table A.3: Comparison of REC on 16 datasets Model BHF CatBoost Dempster Deng Dubois EC-FMDS HEF LightGBM Murphy Our XGBoost Yager Accent 65.16±5.77 41.79±7.27 61.05±8.66 63.24±6.96 62.17±5.32 62.86±4.20 52.56±7.72 62.17±4.57 59.75±6.52 76.71±16.43 63.38±7.60 56.99±7.94 BMNP 37.96±7.22 38.15±3.79 39.49±13.90 44.31±2.37 44.72±11.59 40.56±6.48 40.32±11.50 43.24±5.86 44.49±19.53 72.87±40.50 45.88±9.22 42.22±9.65 Diabetes 96.00±1.13 93.16±2.21 94.44±1.47 95.66±3.00 97.06±1.05 97.34±0.77 96.97±1.47 90.47±0.99 95.28±0.57 97.84±1.04 94.38±2.07 96.25±0.62 Forest 83.73±5.00 85.46±5.14 84.92±4.85 84.91±4.69 86.66±4.24 85.28±5.26 81.59±5.29 86.66±6.19 85.48±5.93 89.48±4.37 86.44±4.89 84.98±0.91 German 64.05±4.07 62.86±3.00 65.64±3.91 63.55±3.26 66.26±3.61 64.76±5.10 64.43±3.80 65.05±3.93 66.57±3.08 74.67±4.54 66.57±5.47 64.81±3.20 Hfcr 80.01±6.28 73.17±5.71 79.20±10.48 81.01±4.07 77.45±9.89 82.46±7.00 75.38±9.34 81.79±9.62 80.12±7.78 83.77±2.84 76.91±10.66 78.69±7.58 Ionosphere 87.24±5.37 87.27±3.90 90.60±5.10 91.50±4.13 88.60±3.98 90.99±5.66 90.33±4.66 89.71±3.87 92.09±3.90 93.60±5.68 87.67±5.00 88.84±5.73 ISPY 96.69±2.06 98.27±2.00 93.36±5.16 93.61±3.26 92.97±5.18 90.98±9.41 87.17±8.53 98.27±2.00 97.31±2.03 96.16±3.48 98.27±2.00 86.32±7.08 Sonar 71.52±1.85 72.35±5.00 69.45±5.06 77.14±4.76 74.27±3.81 75.86±7.06 71.13±6.35 77.59±5.89 77.10±6.96 90.84±12.01 75.34±3.31 75.42±6.28 Sports 76.57±4.02 77.33±1.63 74.59±3.48 75.30±2.49 74.37±3.53 77.54±3.28 74.52±2.59 78.14±4.04 78.72±4.32 87.05±7.43 76.87±1.44 71.72±4.03 Turkish 67.25±2.22 65.50±4.04 67.75±5.85 67.25±3.59 66.25±4.11 67.75±3.40 63.50±5.80 78.25±2.50 69.25±0.96 86.00±14.21 74.50±4.80 66.50±7.59 Vehicle 73.70±2.82 68.81±1.17 70.03±1.66 73.14±1.48 71.21±4.16 74.31±1.41 68.63±2.37 73.26±2.05 74.13±2.70 78.06±8.00 73.68±1.94 69.77±2.50 WDBC 94.97±1.52 93.65±2.13 93.71±1.57 95.21±1.25 93.10±0.96 94.55±1.64 93.71±2.28 94.74±2.48 94.31±2.13 98.55±1.16 93.99±2.92 95.22±1.17 Wine 93.73±4.53 95.24±3.45 93.49±5.80 96.11±2.71 93.73±3.00 94.78±5.31 90.83±3.37 93.06±3.50 96.20±3.32 98.98±2.04 93.53±2.47 88.87±4.46 WPBC 55.36±6.70 63.12±10.62 57.95±5.95 55.00±8.74 55.50±7.72 60.76±9.73 57.24±7.37 60.09±5.80 56.26±5.99 80.93±20.04 62.48±7.71 55.23±5.76 Zoo 84.88±10.30 82.74±9.76 90.12±10.78 85.60±14.13 85.36±9.43 83.93±12.20 81.43±3.50 75.60±4.91 82.74±12.20 94.64±10.71 85.71±5.83 85.36±17.70 Table A.4: Comparison of F1 on 16 datasets Model BHF CatBoost Dempster Deng Dubois EC-FMDS HEF LightGBM Murphy Our XGBoost Yager Accent 62.90±4.48 44.45±7.66 59.40±7.33 61.94±4.29 59.94±3.02 60.81±1.79 49.54±7.05 63.15±6.40 59.06±6.93 68.33±14.48 64.44±6.79 54.58±8.38 BMNP 35.48±7.32 36.06±4.25 35.30±14.95 43.54±4.79 42.77±11.46 40.05±7.24 37.91±13.18 40.53±6.06 44.73±19.78 68.34±38.34 44.39±10.32 39.91±8.39 Diabetes 95.57±0.82 92.76±2.36 93.62±1.90 95.38±3.52 96.59±1.48 97.17±0.81 96.58±1.76 90.62±1.42 94.78±0.98 97.58±1.31 93.98±2.59 95.78±0.74 Forest 83.90±3.72 85.80±4.27 85.59±4.75 85.22±3.45 86.66±3.17 85.00±4.00 82.50±4.55 87.20±5.32 84.93±4.93 88.77±5.44 87.25±4.04 84.57±1.26 German 62.62±3.01 63.75±3.39 65.02±4.14 63.25±2.84 65.59±3.36 64.25±4.90 63.82±3.41 65.99±4.22 65.75±3.00 72.23±4.09 67.24±5.58 63.85±2.69 Hfcr 78.03±4.83 74.49±5.52 79.55±9.67 79.52±4.83 77.68±9.49 81.62±6.68 75.79±8.70 82.25±9.16 79.55±7.25 82.63±4.09 77.38±10.31 79.13±6.82 Ionosphere 87.97±5.20 88.19±3.77 89.73±5.32 91.06±3.39 88.08±3.87 91.01±5.27 89.37±4.07 90.52±3.85 92.50±3.66 93.53±5.87 88.58±5.10 88.67±6.19 ISPY 95.91±2.30 98.79±1.40 94.42±3.89 94.02±2.44 92.85±4.33 91.03±11.58 89.56±8.02 98.79±1.40 97.09±2.23 93.46±5.68 98.79±1.40 87.07±5.50 Sonar 71.49±1.82 72.05±5.51 68.61±4.97 77.18±4.74 73.90±4.10 75.77±7.04 70.55±5.92 77.63±6.03 76.89±6.89 90.85±12.09 75.33±3.32 75.37±6.34 Sports 76.19±3.65 78.07±1.77 75.42±3.39 75.27±1.80 74.79±3.50 77.58±3.12 74.87±2.61 78.53±3.78 78.73±4.23 86.84±7.56 77.23±1.56 72.11±3.99 Turkish 66.65±1.90 65.42±4.12 66.94±6.13 66.75±3.81 66.14±4.06 67.20±3.34 63.61±5.73 78.19±2.30 69.10±1.14 85.92±14.38 74.28±4.40 66.90±7.31 Vehicle 73.23±3.20 67.23±0.86 68.99±1.77 72.51±1.58 70.68±4.62 73.77±1.53 66.68±2.45 72.34±2.42 73.54±2.80 77.22±8.35 73.18±1.90 68.57±2.84 WDBC 94.93±1.64 94.10±2.01 93.30±1.79 95.13±1.42 92.73±1.08 94.39±1.76 93.31±2.72 94.77±3.10 94.53±2.44 98.32±1.12 93.67±3.19 94.78±1.05 Wine 93.28±4.98 95.08±3.76 93.62±5.32 95.63±3.10 93.48±3.00 94.38±5.90 90.46±3.99 92.71±3.83 95.87±3.85 98.92±2.17 93.35±2.97 88.67±4.58 WPBC 55.41±7.19 63.27±12.49 58.14±6.87 55.27±9.03 55.01±9.58 60.38±9.13 57.02±8.81 60.67±6.64 56.06±6.16 76.62±17.67 62.51±6.72 54.69±6.52 Zoo 80.38±13.36 81.28±11.22 88.86±12.64 83.82±15.59 82.28±10.43 79.78±15.17 76.57±3.18 71.99±8.23 80.71±13.46 92.86±14.29 82.01±7.35 83.23±20.70 Appendix B Figure B.1: F1-score under Gaussian feature noise Figure B.2: F1-score under mean feature noise Figure B.3: F1-score under random label noise References Alzubaidi et al. (2021) L. Alzubaidi, J. Zhang, A. J. Humaidi, A. Al-Dujaili, Y. Duan, O. Al-Shamma, J. Santamaría, M. A. Fadhel, M. Al-Amidie, and L. Farhan Review of deep learning: concepts, cnn architectures, challenges, applications, future directions. Journal of Big Data 8 (1). External Links: ISSN 2196-1115, Link, Document Cited by: §5.2.2. Bleichrodt et al. (2010) H. Bleichrodt, A. Cillo, and E. Diecidue A quantitative measurement of regret theory. Management Science 56 (1), p. 161–175 (en). External Links: ISSN 0025-1909, Link, Document Cited by: §2.2. Cai and Ma (2022) T. T. Cai and R. Ma Theoretical foundations of t-SNE for visualizing high-dimensional clustered data. Journal of Machine Learning Research 23 (301), p. 1–54 (en). External Links: ISSN 1533-7928, Link Cited by: §2.3. Calderwood et al. (2016) S. Calderwood, K. McAreavey, W. Liu, and J. Hong Context-dependent combination of sensor information in dempster–shafer theory for bdi. Knowledge and Information Systems 51 (1), p. 259–285. External Links: ISSN 0219-3116, Link, Document Cited by: §1. Chamlal et al. (2025) H. Chamlal, F. E. Rebbah, and T. Ouaderhman An ensemble classifier combining dempster–shafer theory and feature selection methods aggregation strategy. Applied Soft Computing 180, p. 113306 (en). External Links: ISSN 15684946, Link, Document Cited by: §5.1.2. Chen et al. (2023) L. Chen, Z. Zhang, G. Yang, Q. Zhou, Y. Xia, and C. Jiang Evidence-theory-based reliability analysis from the perspective of focal element classification using deep learning approach. Journal of Mechanical Design 145 (7). External Links: ISSN 1528-9001, Link, Document Cited by: §6. Chen and Guestrin (2016) T. Chen and C. Guestrin XGBoost: a scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco California USA, p. 785–794 (en). Note: arXiv:1603.02754 [cs] External Links: ISBN 978-1-4503-4232-2, Link, Document Cited by: §5.1.2. Dempster (1967) A. P. Dempster Upper and lower probabilities induced by a multivalued mapping. Annals of Mathematical Statistics 38 (2), p. 57–72 (en). External Links: Link, Document Cited by: §2.1, §5.1.2. Deng et al. (2004) Y. Deng, W. Shi, Z. Zhu, and Q. Liu Combining belief functions based on distance of evidence. Decision Support Systems 38 (3), p. 489–493 (en). External Links: ISSN 0167-9236, Link, Document Cited by: §1, §5.1.2. Destercke and Burger (2013) S. Destercke and T. Burger Toward an axiomatic definition of conflict between belief functions. IEEE Transactions on Cybernetics 43 (2), p. 585–596 (en). External Links: ISSN 2168-2267, 2168-2275, Link, Document Cited by: §1, §3.2. Dhanka et al. (2026) S. Dhanka, A. Sharma, A. Kumar, S. Maini, and H. Vundavilli Advancements in hybrid machine learning models for biomedical disease classification using integration of hyperparameter-tuning and feature selection methodologies: a comprehensive review. Archives of Computational Methods in Engineering 33 (1), p. 289–324 (en). External Links: ISSN 1134-3060, 1886-1784, Link, Document Cited by: §5.1.2. Dubois and Prade (1988) D. Dubois and H. Prade Representation and combination of uncertainty with belief functions and possibility measures. Computational Intelligence 4 (3), p. 244–264 (en). External Links: ISSN 0824-7935, 1467-8640, Link, Document Cited by: §1, §5.1.2. El-Din et al. (2024) D. M. El-Din, A. E. Hassanein, and E. E. Hassanien An adaptive and late multifusion framework in contextual representation based on evidential deep learning and dempster–shafer theory. Knowledge and Information Systems 66 (11), p. 6881–6932. External Links: ISSN 0219-3116, Link, Document Cited by: §1. Gao and Pan (2025) X. Gao and L. Pan An information fusion model of mutual influence between focal elements: a perspective on interference effects in dempster–shafer evidence theory. Information Fusion 124, p. 103286. External Links: ISSN 1566-2535, Link, Document Cited by: §1. Ghorbanzadeh et al. (2021) O. Ghorbanzadeh, S. R. Meena, H. Shahabi Sorman Abadi, S. Tavakkoli Piralilou, Z. Lv, and T. Blaschke Landslide mapping using two main deep-learning convolution neural network streams combined by the dempster–shafer model. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing 14, p. 452–463. External Links: ISSN 2151-1535, Link, Document Cited by: §1. Guo et al. (2026) B. Guo, J. Shen, G. Zhou, D. Gao, and Z. Wu A multi-damage fusion diagnostic method based on dempster-shafer evidence theory for guided vision-adaptive detection using guided waves. Measurement 287, p. 122435. External Links: ISSN 0263-2241, Link, Document Cited by: §3.3. Hasan et al. (2024) M. Hasan, M. Z. Abedin, P. Hajek, K. Coussement, Md. N. Sultan, and B. Lucey A blending ensemble learning model for crude oil price forecasting. Annals of Operations Research 353 (2), p. 485–515. External Links: ISSN 1572-9338, Link, Document Cited by: §5.2.2. Huang et al. (2025) L. Huang, S. Ruan, P. Decazes, and T. Denœux Deep evidential fusion with uncertainty quantification and reliability learning for multimodal medical image segmentation. Information Fusion 113, p. 102648. External Links: ISSN 1566-2535, Link, Document Cited by: §1, §3.3. Ke et al. (2017) G. Ke, Q. Meng, T. Finley, T. Wang, W. Chen, W. Ma, Q. Ye, and T. Liu LightGBM: a highly efficient gradient boosting decision tree. In Advances in Neural Information Processing Systems, Vol. 30, p. 3149–3157 (en). External Links: Link Cited by: §5.1.2. Li et al. (2026) Q. Li, X. Wan, H. Guo, W. Wang, M. Liu, and Y. Cui Uncertainty-aware multi-modal time series anomaly detection via reduced-order evidence fusion. Pattern Recognition 180, p. 114212. External Links: ISSN 0031-3203, Link, Document Cited by: §1. Liu et al. (2023) Q. Liu, W. Lee, M. Huang, and Q. Wu Synergy between stock prices and investor sentiment in social media. Borsa Istanbul Review 23 (1), p. 76–92 (en). Note: JCR分区: Q1 中科院分区升级版: 经济学2区 影响因子: 7.1 5年影响因子: 6.3 南农高质量: A External Links: ISSN 2214-8450, Link, Document Cited by: §3.4, §5.2.2. Loomes and Sugden (1982) G. Loomes and R. Sugden Regret theory: an alternative theory of rational choice under uncertainty. The Economic Journal 92 (368), p. 805–824 (en). External Links: ISSN 0013-0133, Link, Document Cited by: Definition 5. Lu and Zhu (2026) Y. Lu and W. Zhu Soft likelihood-based multi-source fusion with reliability modeling and optimization under dempster–shafer theory. Engineering Applications of Artificial Intelligence 163, p. 112993. External Links: ISSN 0952-1976, Link, Document Cited by: §3.3. Murphy (2000) C. K. Murphy Combining belief functions when evidence conflicts. Decision Support Systems 29 (1), p. 1–9 (en). External Links: ISSN 0167-9236, Link, Document Cited by: §1, §3.4, §5.1.2. Ng et al. (2001) A. Ng, M. Jordan, and Y. Weiss On spectral clustering: analysis and an algorithm. In Advances in Neural Information Processing Systems, Cambridge, MA, USA, p. 849–856 (en). External Links: Link Cited by: §2.3. Park (2025) J. Park Estimation of vessel collision risk under uncertainty using interval type-2 fuzzy inference systems and dempster–shafer evidence theory. Journal of Marine Science and Engineering 14 (1), p. 34. External Links: ISSN 2077-1312, Link, Document Cited by: §1. Prokhorenkova et al. (2018) L. Prokhorenkova, G. Gusev, A. Vorobev, A. V. Dorogush, and A. Gulin CatBoost: unbiased boosting with categorical features. In Advances in Neural Information Processing Systems, Vol. 31, Red Hook, NY, USA, p. 6639–6649 (en). External Links: Link Cited by: §5.1.2. Qiao et al. (2023a) S. Qiao, Y. Fan, G. Wang, and H. Zhang Multi-sensor data fusion method based on improved evidence theory. Journal of Marine Science and Engineering 11 (6), p. 1142. External Links: ISSN 2077-1312, Link, Document Cited by: §1. Qiao et al. (2023b) S. Qiao, B. Song, Y. Fan, and G. Wang A fuzzy dempster–shafer evidence theory method with belief divergence for unmanned surface vehicle multi-sensor data fusion. Journal of Marine Science and Engineering 11 (8), p. 1596. External Links: ISSN 2077-1312, Link, Document Cited by: §5.2.2. Qiu et al. (2025) Z. Qiu, Y. Qin, Z. Chen, L. Zeng, and R. Cai Overcoming negative weighting in uncertainty-based methods: a multi-uncertainty clustering method for evidence fusion. Complex & Intelligent Systems 11 (9). External Links: ISSN 2198-6053, Link, Document Cited by: §3.3. Radzvilas et al. (2024) M. Radzvilas, W. Peden, D. Tortoli, and F. De Pretis A comparison of imprecise bayesianism and dempster–shafer theory for automated decisions under ambiguity. Journal of Logic and Computation 35 (8). External Links: ISSN 1465-363X, Link, Document Cited by: §1. Shafer (1976) G. Shafer A mathematical theory of evidence. Princeton University Press, Princeton, USA (en). Note: Google-Books-ID: wug9DwAAQBAJ External Links: ISBN 978-0-691-10042-5, Document Cited by: §2.1. Shi and Malik (2000) J. Shi and J. Malik Normalized cuts and image segmentation. IEEE Transactions on Pattern Analysis and Machine Intelligence 22 (8), p. 888–905 (en). Note: Conference Name: IEEE Transactions on Pattern Analysis and Machine Intelligence External Links: ISSN 1939-3539, Link, Document Cited by: Definition 9. Strelet et al. (2025) E. Strelet, I. Castillo, Y. Peng, and M. S. Reis Data fusion: integrating heterogeneous information sources in the chemical processing industry. Journal of Chemometrics 39 (11). External Links: ISSN 1099-128X, Link, Document Cited by: §6. Urbani et al. (2023) M. Urbani, G. Gasparini, and M. Brunelli A numerical comparative study of uncertainty measures in the dempster–shafer evidence theory. Information Sciences 632, p. 119027. External Links: ISSN 0020-0255, Link, Document Cited by: §1. van Engelen and Hoos (2019) J. E. van Engelen and H. H. Hoos A survey on semi-supervised learning. Machine Learning 109 (2), p. 373–440. External Links: ISSN 1573-0565, Link, Document Cited by: §5.2.2. Wang et al. (2021) H. Wang, J. Wang, and G. Wang Clustering validity function fusion method of fcm clustering algorithm based on dempster–shafer evidence theory. International Journal of Fuzzy Systems 24 (1), p. 650–675. External Links: ISSN 2199-3211, Link, Document Cited by: §3.3. Xiao et al. (2026) Y. Xiao, X. Ma, and J. Zhan Group decision-making in heterogeneous multi-scale information fusion: integrating overconfident and non-cooperative behaviors. Information Fusion 125, p. 103401. External Links: ISSN 1566-2535, Link, Document Cited by: §1. Xu et al. (2016) X. Xu, Z. Zhang, D. Xu, and Y. Chen Interval-valued evidence updating with reliability and sensitivity analysis for fault diagnosis. International Journal of Computational Intelligence Systems 9 (3), p. 396. External Links: ISSN 1875-6883, Link, Document Cited by: §1. Yager (1987) R. R. Yager On the dempster-shafer framework and new combination rules. Information Sciences 41 (2), p. 93–137 (en). External Links: ISSN 0020-0255, Link, Document Cited by: §1, §5.1.2. Yu et al. (2026) Y. Yu, H. R. Karimi, L. Gelman, J. Tian, and P. Mei A novel multi-source sensor correlation adaptive fusion framework with uncertainty quantification for intelligent fault diagnosis. Reliability Engineering & System Safety 267, p. 111812. External Links: ISSN 0951-8320, Link, Document Cited by: §1. Zadeh (1986) L. A. Zadeh A simple view of the dempster-shafer theory of evidence and its implication for the rule of combination. AI Magazine 7 (2), p. 85–91 (en). Note: Number: 2 External Links: ISSN 2371-9621, Link, Document Cited by: §1. Zhang et al. (2025) Q. Zhang, P. Zhang, and T. Li Information fusion for large-scale multi-source data based on the dempster-shafer evidence theory. Information Fusion 115, p. 102754. External Links: ISSN 1566-2535, Link, Document Cited by: §1, §3.3, §5.1.2, §6. Zhao et al. (2025) Z. Zhao, R. Wang, W. Pang, and Z. Li Feature selection for label distribution learning using dempster-shafer evidence theory. Applied Intelligence 55 (4). External Links: ISSN 1573-7497, Link, Document Cited by: §3.3. Zhou et al. (2019) J. Zhou, X. Hong, and P. Jin Information fusion for multi-source material data: progress and challenges. Applied Sciences 9 (17), p. 3473. External Links: ISSN 2076-3417, Link, Document Cited by: §6. Zhou and Xiao (2026) Z. Zhou and F. Xiao Conflict management in sequential evidence combination. Information Sciences 734, p. 122958. External Links: ISSN 0020-0255, Link, Document Cited by: §1. Zhu and Xiao (2021) C. Zhu and F. Xiao A belief hellinger distance for D–S evidence theory and its application in pattern recognition. Engineering Applications of Artificial Intelligence 106, p. 104452 (en). External Links: ISSN 09521976, Link, Document Cited by: §5.1.2.