Paper deep dive
LAF-Based Evaluation and UTTL-Based Learning Strategies with MIATTs
Yongquan Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/26/2026, 6:20:58 PM
Summary
The paper introduces the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework, designed for machine learning scenarios where a single, objective 'ground truth' is undefined or ambiguous. The framework consists of two core mechanisms: LAF (Logical Assessment Formula)-based evaluation, which uses logical aggregation (conjunction, disjunction, fuzzy operations) to assess models via multiple partially correct targets, and UTTL (Undefinable True Target Learning)-based learning, which optimizes models using multi-target strategies (e.g., Dice or Cross-Entropy loss) to handle epistemic uncertainty. The research analyzes how the structural properties of task-specific MIATTsβspecifically coverage and diversityβinfluence the soundness of logical evaluation and the stability of statistical optimization.
Entities (8)
Relation Signals (5)
Yongquan Yang β affiliatedwith β Institute of Sciences for AI
confidence 100% Β· Yongquan Yang 1* 1 Institute of Sciences for AI
EL-MIATTs β comprises β LAF
confidence 100% Β· the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework has been proposed... we develop two complementary mechanisms: LAF... and UTTL...
EL-MIATTs β comprises β UTTL
confidence 100% Β· the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework has been proposed... we develop two complementary mechanisms: LAF... and UTTL...
LAF β evaluates β MIATTs
confidence 100% Β· LAF (Logical Assessment Formula)-based evaluation algorithms... with MIATTs
UTTL β learnswith β MIATTs
confidence 100% Β· UTTL (Undefinable True Target Learning)-based learning strategies with MIATTs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In many real-world machine learning (ML) applications, the true target cannot be precisely defined due to ambiguity or subjectivity information. To address this challenge, under the assumption that the true target for a given ML task is not assumed to exist objectively in the real world, the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework has been proposed. Bridging theory and practice in implementing EL-MIATTs, in this paper, we develop two complementary mechanisms: LAF (Logical Assessment Formula)-based evaluation algorithms and UTTL (Undefinable True Target Learning)-based learning strategies with MIATTs, which together enable logically coherent and practically feasible modeling under uncertain supervision. We first analyze task-specific MIATTs, examining how their coverage and diversity determine their structural property and influence downstream evaluation and learning. Based on this understanding, we formulate LAF-grounded evaluation algorithms that operate either on original MIATTs or on ternary targets synthesized from them, balancing interpretability, soundness, and completeness. For model training, we introduce UTTL-grounded learning strategies using Dice and cross-entropy loss functions, comparing per-target and aggregated optimization schemes. We also discuss how the integration of LAF and UTTL bridges the gap between logical semantics and statistical optimization. Together, these components provide a coherent pathway for implementing EL-MIATTs, offering a principled foundation for developing ML systems in scenarios where the notion of "ground truth" is inherently uncertain. An application of this work's results is presented as part of the study available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.20944v1
- Canonical: https://arxiv.org/abs/2604.20944v1
Trouble viewing inline? Open PDF directly β
Full Text
67,925 characters extracted from source content.
Expand or collapse full text
LAF-Based Evaluation and UTTL-Based Learning Strategies with MIATTs Yongquan Yang 1* 1 Institute of Sciences for AI, Chengdu, Sichuan, China * Corresponding author (Email: remy_yang@foxmail.com or yongquan.yang@sciences4ai.com, ORCiD: 0000-0002-3965-4816) In many real-world machine learning (ML) applications, the true target cannot be precisely defined due to ambiguity or subjectivity information. To address this challenge, under the assumption that the true target for a given ML task is not assumed to exist objectively in the real world, the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework has been proposed. Bridging theory and practice in implementing EL-MIATTs, in this paper, we develop two complementary mechanisms: LAF (Logical Assessment Formula)-based evaluation algorithms and UTTL (Undefinable True Target Learning)-based learning strategies with MIATTs, which together enable logically coherent and practically feasible modeling under uncertain supervision. We first analyze task-specific MIATTs, examining how their coverage and diversity determine its structural property and influence downstream evaluation and learning. Based on this understanding, we formulate LAF-grounded evaluation algorithms that operate either on original MIATTs or on ternary targets synthesized from them, balancing interpretability, soundness, and completeness. For model training, we introduce UTTL-grounded learning strategies using Dice and Cross- Entropy loss functions, comparing per-target and aggregated optimization schemes. We also discuss how the integration of LAF and UTTL bridges the gap between logical semantics and statistical optimization. Together, these components provide a coherent pathway for implementing EL-MIATTs, offering a principled foundation for developing ML systems in scenarios where the notion of βground truthβ is inherently uncertain. An application of this workβs results is presented as part of the study available at https://w.qeios.com/read/EZWLSN. 1. Introduction Modern machine learning (ML) systems are increasingly deployed in real-world environments where the notion of a single, precisely defined true target is often ill-posed or even nonexistent. In many practical domains, such as medical diagnosis, social behavior analysis, and open-world perception, ground truth labels arise from subjective, incomplete, or conflicting human or model-generated sources. These conditions challenge the classical assumption of a deterministic and accurate true target (ATT), which underpins most conventional evaluation [1β5] and learning [6β10] paradigms. To bridge this theoreticalβ practical gap, under the explicitly posited assumption that the true target for a given ML task is not assumed to exist as a well-defined object in the real world, the EL-MIATTs (Evaluation and Learning with Multiple Inaccurate True Targets) framework has been proposed as a means to model, analyze, and utilize imperfect yet informative approximations of the underlying truth [11]. While the theoretical formulation of EL-MIATTs establishes a principled basis for addressing epistemic uncertainty, its practical implementation, which extends beyond the generation and assessment of MIATTs [12], critically depends on two complementary mechanisms that operationalize the framework in real-world ML tasks: LAF (Logical Assessment Formula [13])-based evaluation and UTTL (Undefinable True Target Learning [14])-based learning with MIATTs. Accordingly, this paper presents applicable LAF-grounded evaluation algorithms and UTTL-grounded learning strategies as the operational realization of these two mechanisms within the EL-MIATTs framework. The former enables logically consistent assessment of models when multiple partially correct targets coexist; the latter allows model training to proceed effectively even when the true target is fundamentally undefinable. Together, these mechanisms transform EL-MIATTs from a conceptual framework into a deployable methodology for complex, real-world machine learning tasks. Their integration is thus crucial for achieving both theoretical rigor and practical utility in uncertain supervision settings. Preliminaries related to EL-MIATTs framework are provided in Section2. A central component of this study is the analysis of task-specific MIATTs, which are generated from diverse task-related AI models (AIM) retrievable from real-world resources [12]. Each MIATT represents a partial, probabilistic, or context-specific approximation of the latent true target. The quality of a task-specific MIATTs set can be characterized by its coverage (mean of PartialRepresentation) and diversity (1βRedundancy), together determining its representational fidelity. High-quality MIATTs achieve a balance between completeness and consistency, covering the semantic scope of the true target while minimizing redundant or contradictory information. As this structural property is essential for designing reliable evaluation and learning processes under EL-MIATTs, we also analyze the downstream influence of task-specific MIATTs on evaluation and learning. A comprehensive analysis of the structural property of task-specific MIATTs and its downstream influence on evaluation and learning is presented in Section 3. LAF [13] provides a theoretical and algorithmic foundation for evaluating predictive models with MIATTs. Building upon its principle, we extend classical logic-based and fuzzy operations [15β19] to accommodate evaluation with MIATTs, thereby supporting both parallel multi-perspective and ternary synthesized evaluation schemes. In the parallel multi- perspective evaluation scheme, each MIATT preserves its individual partial truth and contributes to the overall assessment through logical aggregation operations such as conjunction, disjunction, t-norm, and t-conorm. This approach maintains detailed information and enables fine-grained interpretability. In contrast, the ternary synthesized evaluation scheme compresses the MIATTs set into a single three-valued representation 0,0.5,1, facilitating unified scoring and computational simplicity at the expense of some informational granularity. Together, these LAF-based evaluation algorithms achieve a balance between logical completeness and practical interpretability, approximating conventional ATT- based evaluation in complex, ill-defined tasks while reflecting intrinsic ambiguity in simpler or underdefined scenarios. Further methodological details regarding these two LAF-based evaluation algorithms are presented in Section 4. Complementing LAF, UTTL [14] addresses the challenge of model training when the true target is uncertain or inherently undefined. UTTL can be operationalized by treating MIATTs as multiple weakly reliable surrogates of the ground truth, enabling the learning process to proceed through multi-target optimization [14]. Building on this principle, two main strategies can be derived: (1) Per-target then Aggregate, which computes loss functions (e.g., Dice [20] or Cross Entropy [21]) for each MIATT individually before aggregating the results; and (2) Aggregate then Single Loss, which first synthesizes a composite target and then computes a single loss against it. Depending on the loss function used for optimization, these two strategies exhibit different theoretical behaviors and learning biasesβone emphasizing robustness to diversity, the other promoting consistency and stability. Collectively, they establish a flexible paradigm for learning under epistemic uncertainty, fostering the development of adaptive and explainable ML systems. Further theoretical analyses and practical guidelines for applying these two UTTL-based learning strategies are presented in Section 5. Synthesizing the structural property of task-specific MIATTs and its downstream influence on evaluation and learning, and the theoretical implications and practical insights derived from LAF-based evaluation and UTTL-based learning, we further discuss how the integration of LAF and UTTL bridges the gap between logical semantics and statistical optimization. This discussion examines the inherent trade-offs among coverage, consistency, and interpretability within multi-valued logic systems, extending the EL-MIATTs framework toward paraconsistent reasoning and adaptive weighting mechanisms. Comprehensive discussion and analytical results are presented in Section 6. In summary, this work advances the theoretical and practical development of the EL- MIATTs framework by introducing concrete algorithmic and strategic realizations for both evaluation and learning with MIATTs, establishing a foundation for reliable, interpretable, and uncertainty-aware machine learning. The key contributions of this work are summarized as follows: β« We analyze the structural properties of task-specific MIATTs using coverage- and diversity-based indicators and discuss its downstream influence on evaluation and learning. β« We propose two LAF-grounded algorithms for evaluation with MIATTs, supporting both parallel multi-perspective and ternary synthesized schemes. β« We develop and summarize two UTTL-grounded learning strategies for model training with MIATTs, based on Dice and Cross-Entropy losses. β« We further discuss these three contributions to bring out the integration of LAF- based evaluation with and UTTL-based learning with MIATTs, extending the EL- MIATTs framework to bridge logical semantics with statistical optimization while enhancing robustness, interpretability, and adaptability in uncertainty-aware learning. The remainder of this paper is structured as follows: Section 2 introduces the preliminaries related to EL-MIATTs. Section 3 analyzes the structural properties of task-specific MIATTs with respect to their assessment indicators. Section 4 presents the two proposed LAF-grounded algorithms for evaluation with MIATTs. Section 5 describes two UTTL-grounded learning strategies for model training with MIATTs. Section 6 discusses the integration of LAF-based evaluation and UTTL-based learning to extend the EL-MIATTs framework. Finally, Section 7 concludes the paper, highlighting limitations and directions for future research. 2. Preliminary In this section, building on previous studies [11, 12], we briefly introduce the definition and core concept of MIATTs, their task-specific generation and evaluation processes, as well as LAF [13] and UTTL [14] for evaluation and learning of predictive models with MIATTS. 2.1 Definition and essence of MIATTs Building on the core premise that the true target for a given machine learning task is not assumed to exist as a well-defined object in the real world, the notion of MIATTs can be introduced as: Let ν‘ β be the underlying (possibly undefinable) true target and ννΉ ( ν‘ β ) its set of semantic facts. A MIATTs set is ννΌν΄νν = ν‘ ν β |νβ 1,β―,ν ,νβ₯2 , where each ν‘ ν β satisfies ννΉ ( ν‘ ν β ) βννΉ ( ν‘ β ) and β ννΉ ( ν‘ ν β ) ν ν=1 βννΉ(ν‘ β ). (1) In other words, each ν‘ ν β reflects only part of the semantics of ν‘ β , while the collection as a whole provides a broader approximation of it. Fundamentally, MIATTs capture the insight that supervision in real-world machine learning is inherently partial and noisy [11]. While any single inaccurate true target conveys just a fragment of the underlying semantic structure, the set collectively reconstructs it with wider coverage. This reframing shifts supervision away from a strict single-target view toward a distributional, multi-perspective paradigm, thereby laying a principled basis for the EL- MIATTs framework to support robust evaluation and learning under uncertainty. 2.2 Generation and assessment of task-specific MIATTs Based on this definition and essence of MIATTs, logic-driven algorithms have been proposed for generation and assessment of MIATTS [12]. In this section, we summarize two simplified logic-driven solutions using retrievable real-world resources for generation and assessment of task-specific MIATTs. 2.2.1 Generation of task-specific MIATTs For specific tasks, existing resourcesβsuch as accumulated task-specific datasets, pretrained models, or related large AI modelsβcan be readily utilized. By leveraging these resources, we construct a set of task-specific AI models (AIM), integrating both prior task- specific models and relevant large AI models. This AIM set maps each instance to multiple predicted true targets corresponding to its underlying ground truth, collectively forming the generated MIATTs. Let ν΄νΌν= ν 1 ,ν 2 ,...,ν ν denote the set of predictive models and tools derived from these resources. The stepwise procedure of this approach [12] is outlined as follows: Input: β’ Raw data instances νΌ. β’ A set of AI models ν΄νΌν= ν 1 ,ν 2 ,...,ν ν‘ for mapping an instance into predicted multiple true targets for the underlying true target ν‘ β . Step 1: Prediction of multiple potential true targets β’ Predict multiple potential true targets for νΌ with ν΄νΌν: ννΌν΄νν =ν΄νΌν ( νΌ ) = ν 1 ( νΌ ) ,ν 2 ( νΌ ) ,...,ν ν ( νΌ ) =ν‘ ν β |νβ 1,β―,ν , νβ₯2. (2) Output: β’ Each instance in νΌ is assigned an MIATTs set ν‘ ν β ν=1 νβ₯2 , each being a partial but informative approximation of the corresponding underlying true target ν‘ β . 2.2.2 Assessment of task-specific MIATTs For a specific task, once an MIATTs set is generated from task-specific AIM retrieved from real-world resources, a probable true target can be derived to approximate the underlying ground truth. Since each IATT in the MIATTs set covers a partial yet reliable aspect of the true target, integrating these partial coverages yields a synthesized target that captures the essential semantics of the ground truth. This summarized probable target then serves as a reference for evaluating the MIATTs set itself, ensuring a self-consistent and task-adaptive assessment process. The stepwise procedure [12] is outlined as follows: Input: β’ An MIATTs set ν‘ ν β ν=1 νβ₯2 generated by task-specific AIM. Step 1: Approximate probable true target ν‘ β Μ =νννν ( ν‘ ν β ν=1 νβ₯2 ) . (3) Step 2: Represent IATT as Boolean vectors β’ Represent each IATT ν‘ ν β as a Boolean vector ν£ ν β0,1 ν , where: ν£ ν [ ν ] = 1, νν ννν ( ν‘ ν β ( ν ) βν‘ β Μ ( ν ) ) <νΏ 0, νν‘βννν€νν ν . (4) Step 3: Assess partial representation (Per-IATT quality) β’ For each IATT ν‘ ν β : νννν‘νννν ννννν ννν‘νν‘ννν ( ν‘ ν β ) = β ν£ ν [ ν ] ν ν . (5) Step 4: Assess redundancy / diversity β’ Compute pairwise intersections (logical AND) between MIATTs: νΌνν‘ννν ννν‘ννν(ν‘ ν β ,ν‘ ν β )=ν£ ν β§ν£ ν . β’ Measure redundancy ratio: ν ννν’νννννν¦= β β£ν£ ν β§ν£ ν β£ ν<ν β β£ν£ ν β£ ν . (6) Step 5: Overall quality score β’ Combine the metrics into an aggregate score: ν ννΌν΄νν =νΌβ νννν(νννν‘νννν ννννν ννν‘νν‘ννν)βνΎβ ν ννν’νννννν¦. (7) Output: β’ The computed ν ννΌν΄νν for assessing the overall quality score of the MIATTs set. 2.3 LAF and UTTL Operating under a relaxed but shared assumption that the true target for a given ML task is not assumed to exist as a well-defined object in the real world [11], LAF [13] and UTTL [14] respectively provide the theoretical foundations for the evaluation and learning of predictive models with MIATTs. 2.3.1 LAF LAF provides a logical framework for evaluating predictive models using MIATTs by aggregating multiple partially correct targets [13]. The principle of LAF is: Based on MIATTs, LAF can approximate conventional ATT-based evaluation reasonably well in complex tasks, while potentially exhibiting greater deviations in simpler ones [11, 13]. Thus, LAF can be implemented through logical operations (e.g., conjunction, disjunction, or fuzzy aggregation). It assesses predictive model performance from a multi-perspective logical viewpoint, emphasizing coverage and consistency rather than relying on a single ground truth. This enables fine-grained diagnosis of correctness across different incomplete or uncertain targets. 2.3.2 UTTL UTTL provides a learning paradigm that optimizes predictive models without relying on a single accurate true target [14]. The principle of UTTL is: Based on MIATTs, UTTL can be effectively implemented within a multi-target learning framework [11, 14]. Thus, UTTL can be implemented by learning from multiple incomplete or uncertain targets within the MIATTs set by aligning their shared and reliable components, thereby approximating the underlying true target in a self-consistent and task-adaptive manner. 3. Analysis of Task-Specific MIATTs Regarding the generation and assessment of MIATTs for a specific task, this section analyzes the possible qualities and structural patterns of task-specific MIATTs, along with their downstream influence on evaluation and learning. 3.1 Possible qualities of task-specific MIATTs with respect to assessment indicators Based on the task-specific AIM, multiple diverse MIATTs can be generated according to Formula (2). With respect to the assessment indicators defined in Formulas (5) and (6), the potential qualities of these MIATTs are summarized in Table 1. Specifically, in Table 1, the mean of PartialRepresentation reflects the coverage of MIATTs with respect to the underlying true target, while (1 β Redundancy) represents their diversity in capturing the true target. Table 1. Possible qualities of task-specific MIATTs. 1-Redundancy=0 1-Redundancy=0.5 1-Redundancy=1 mean(PartialRepresentation)=0 Worst Quality mean(PartialRepresentation)=0.5 Median Quality mean(PartialRepresentation)=1 Best Quality As shown in Table 1, the quality of task-specific MIATTs is determined jointly by coverageβmeasured by the mean of PartialRepresentationβand diversityβmeasured by (1βRedundancy). High coverage ensures that each IATT reliably captures essential aspects of the underlying true target, while high diversity provides complementary perspectives that reduce shared bias. Low values of both indicators lead to the worst quality, where MIATTs fail to represent the underlying true target effectively. Moderate levels yield median quality, offering partial but useful diagnostic information. The best quality is achieved when both coverage and diversity are high, producing MIATTs that are accurate, comprehensive, and mutually reinforcing. This configuration enables robust, fine-grained evaluation and learning, making the MIATTs set a faithful and informative surrogate for the unattainable accurate true target. 3.2 Structural patterns of task-specific MIATTs The structural patterns of task-specific MIATTs corresponding to the quality levels summarized in Table 1 are illustrated in Figure 1. A task-specific MIATTs set can be formed by including all feasible patterns, excluding the unattainable pattern depicted in Figure 1. Each included pattern contributes to capturing different aspects of the underlying true target, with variations in coverage and diversity that collectively determine the overall informativeness and reliability of the MIATTs set. By combining these patterns, the MIATTs set provides a comprehensive, multi-perspective approximation of the underlying true target, enabling robust evaluation and learning. Figure 1. Possible patterns of task-specific MIATTs. Blue circles represent the underlying true target, while orange circles represent the generated MIATTs for approximating it. Blue bounding boxes indicate MIATTs with no coverage. Blue bounding boxes indicate MIATTs with adequate coverage but limited diversity. Red bounding boxes indicate MIATTs with both adequate coverage and diversity. The black bounding boxe indicates the pattern that is unattainable. 3.3 Summary of task-specific MIATTs For a specific task, if the number of MIAs used to generate the task-specific MIATTs is sufficiently large, it is likely that the union of the generated MIATTs set, β ν‘ ν β ν ν=1 , will approximate full coverage of the underlying true target ν‘ β . At the same time, this union inevitably introduces some additional noise. Consequently, the resulting MIATTs set provides a near-complete, multi-perspective representation of ν‘ β , capturing its essential semantic structure while satisfying the balance between coverage and redundancy. Thus, the task-specific MIATTs set is subject to ννΉ ( ν‘ β ) β β ννΉ ( ν‘ ν β ) ν ν=1 βννΉ(ν‘ β )βͺν. (8) And, it can be geometrically visualized as Fig. 2. Figure 2. Geometric visualization of task-specific MIATTs. The relationship depicted by Formula (8) and Fig. 2 reflects a coverageβnoise trade-off that propagates into subsequent evaluation and learning processes. 3.4 Downstream influence on evaluation and learning 3.4.1 Influence on evaluation (LAF-based) For evaluation with MIATTs, the coverage completeness of the MIATTs set determines the soundness and stability of LAF aggregation results. High coverage with moderate noise allows the LAF frameworkβthrough conjunction, disjunction, or fuzzy t-norm/t-conorm operationsβto approximate traditional accurate-target (ATT) evaluation while preserving interpretability. Insufficient coverage leads to incomplete logical representation, increasing the risk of false negatives (correct predictions misjudged as errors). Excessive redundancy or noise, on the other hand, weakens discriminability, potentially smoothing out subtle correctness patterns. Hence, well-balanced task-specific MIATTs yield evaluation outcomes that are both logically comprehensive and diagnostically informative, aligning with LAFβs goal of interpretable uncertainty reasoning. 3.4.2 Influence on learning (UTTL-based) In UTTL-grounded learning, MIATTs act as multiple weak surrogates of the ground truth. When MIATTs exhibit broad but non-redundant coverage, the learning process benefits from multi-target regularization, enhancing robustness and generalization under epistemic uncertainty. Conversely, overlapping or noisy MIATTs may cause inconsistent gradient directions among targets, potentially slowing convergence or leading to bias toward majority patterns. In the extreme case of sparse MIATT coverage, models may underfit due to insufficient supervision signal, emphasizing the need for adaptive weighting or noise correction mechanisms. Therefore, the statistical behavior of learning with MIATTs depends on the geometric and semantic distribution of the MIATTs setβwhether it provides a well-distributed approximation of ν‘ β or an imbalanced, noisy surrogate. 3.4.3 Joint implication The task-specific MIATTs set serves as the epistemic substrate upon which both evaluation and learning are built. Its structural balance between coverage, diversity, and noise governs: the logical fidelity of LAF-based evaluation and the optimization stability of UTTL- based learning. Achieving this balance ensures that EL-MIATTs not only approximate ν‘ β effectively but also sustain coherent interactions between logical semantics and statistical optimization across the full pipeline. 4. Evaluation with MIATTs: LAF-Grounded Strategies Assuming that the true target for a given ML task is not assumed to exist as a well-defined object in the real world [11], LAF provides a principled framework that, based on MIATTs, can approximate ATT-based evaluation effectively in complex settings but may diverge in simpler ones [11, 13]. Accordingly, this section presents LAF-grounded evaluation algorithms employing logical operators such as conjunction, disjunction, and fuzzy aggregation. Fundamentally, there are two approaches to conducting evaluation with MIATTs: (1) evaluation using the original MIATTs, and (2) evaluation using a ternary target synthesized from MIATTs. The first approach preserves the multi-perspective and partially true nature of each MIATT, enabling detailed diagnostic analysis, whereas the second consolidates these perspectives into a unified, scoreable three-valued target for simplified assessment. In the following section, we systematically discuss and compare these two approaches in terms of their logical foundations, evaluation granularity, interpretability, and computational characteristics. 4.1 Evaluation with original MIATTs Treating semantic facts in MIATTs as computable logical formulas, we use (multi-valued) logical semantics to evaluate the model's satisfaction with partial true goals and perform reasonable aggregation across multiple incomplete goals. This approach relies on Kleene's three-valued logic (K3) [17β19] and fuzzy logic's t-norm (β§) and t-conorm (β¨) [15, 16], along with paraconsistency testing [22, 23] and coverage correction [24, 25]. 4.1.1 Evaluation metric design (based on logic) Assume that each inaccurate true target ν‘ ν β is represented by a set of semantic facts (formulas). Each fact ν for a single sample (ν,ν‘ Μ ) gives a three-valued truth value ν£β0, 1/2 ,1: β’ 1: Satisfied (True); β’ 0: Violated (False); β’ 1/2: Unknown/Not applicable (Undefined). 1) Formula level (fact-level): β’ K3 semantics: Β¬ν=1βν; Conjunction (β§) uses GΓΆdel t-norm: νβ§ν=ννν(ν,ν); Disjunction (β¨) using t-conorm: νβ¨ν=ννν₯(ν,ν); Implication: νβν=ννν₯(1βν,ν). β’ Applicability ν΄(ν): Whether the fact is βdecidableβ for the sample. In practice: ν΄(ν)=1[ν£β 1/2]. 2) Satisfaction of a single IATT ν ν β : β’ Intra-fact Aggregation (emphasizing that "partial representations" must be satisfied simultaneously): ν ν ( ν ) =min νβν· ν ν£ ν ( ν,ν‘ Μ ) . (9) Or a weighted version ν ν ( ν ) =min ν (νΌ ν βν£ ν ) (default is equal weighting, β is the weighting rule; simpler options include weighted minimum or weighted geometric mean). β’ Applicability coverage of this IATT: νΆ ν ( ν ) = β ν΄ ( ν ) νβν· ν | ν· ν | = β 1 [ ν£β 1/2 ] νβν· ν | ν· ν | . (10) 3) Collective coverage of aggregation across multiple MIATTs: β’ We want to reflect "collectively covering more true target facts." Aggregation using t-conorm: ν ννΌν΄νν ( ν ) =max ν=1... ν ν ν ( ν ) . (11) Explanation: As long as all the core facts of a model are satisfied, the model is correct in that aspect. β’ Alternatively, a probabilistic "at least one aspect is correct" can be used: ν ννΌν΄νν νννν ν¦νν =1β β (1βν ν ( ν ) ) ν ν=1 . (12) β’ Overall applicability: νΆ ννΌν΄νν ( ν ) = 1 ν β νΆ ν ( ν ) ν ν=1 . (13) 4) Paraconsistency penalty: When two facts (possibly across MIATTs) with mutually exclusive requirements on the same sample are both "satisfied" to 1 (or close to 1), this is counted as a contradictory hit. β’ Define a set of mutually exclusive pairs ν= ( ν ν ,ν ν ) . β’ Sample-level contradiction rate: νΎ ννΌν΄νν ( ν ) = 1 | ν | β 1[ν£ ν ν =1β§ν£ ν ν =1] ( ν ν ,ν ν ) βν . (14) 5) Final sample score (with coverage correction and consistency penalty): The final sample score can be expressed as ννννν ννΌν΄νν ( ν ) = (νν ννΌν΄νν ( ν ) + ( 1βν ) ν ννΌν΄νν νννν ν¦νν )βνΆ ννΌν΄νν ( ν ) β(1βνΎνΎ ννΌν΄νν ( ν ) ), (15) where νβ[0,1] controls "strict correctness in one aspect" vs. "correctness in at least one aspect"; νΎβ[0,1] controls the intensity of the contradiction penalty. 6) Dataset-level metrics: β’ Average: ννννν Μ ννΌν΄νν = 1 | ν· | β ννννν ννΌν΄νν ( ν ) νβν· . (16) β’ Simultaneously output νΆ Μ ννΌν΄νν = 1 | ν· | β νΆ ννΌν΄νν ( ν ) νβν· ("decidability" of the evaluation), νΎ Μ ννΌν΄νν = 1 | ν· | β νΎ ννΌν΄νν ( ν ) νβν· (overall contradiction rate). 4.1.2 Methodological key points and scalability This logic-based metric design method for evaluation with original MIATTs has following key points and scalability. 1) Aligned with the definition: β’ "Partial representation" β Use conjunction (min) within a single IATT to force the key facts in that aspect to hold simultaneously. β’ "Collective coverage" β Use disjunction (max / NoisyOR) between MIATTs to indicate that "together they cover more aspects of the truth." 2) Unknown/Not applicable: Use ννννννννν=0.5 to maintain the good algebraic properties of K3; the coverage νΆ allows you to determine "how many facts were used" in the evaluation. 3) Conflict resolution: Mutually Exclusive Group + Penalty Term ( 1βνΎνΎ ννΌν΄νν ) is a simple "quasi-parallel consistency" approach; complex scenarios can be replaced with Belnap four-valued logic ( β₯,β€,ννν‘β,νννν‘βνν ) or a constraint solver [26, 27]. 4) Weight learning: Weights can be given by an expert or meta-learned using a validation set (e.g., using the ννννν Μ ννΌν΄νν as the target and using Bayesian Optimization/Differential Evolution to find weights). 5) Interpretability: Each sample has a corresponding set of multiple true targets (MIATTs) for evaluation, which naturally provides an explanation of the hotspots of "where it satisfies/does not satisfy." 4.2 Evaluation with ternary target synthesized from MIATTs We combine MIATTs into a three-valued logical "synthetic true target" ν‘ β via logical merging, then we use this ν‘ β to evaluate the quality of the model's predictions. 4.2.1 Logical merging idea MIATTs are characterized by the following characteristics: each MIATT ν‘ ν β captures only part of the facts of the true target. When merging, we hope to form a single three-valued logical target ν‘ β , which can express: β’ ννν’ν (1): All MIATTs agree that this is true for this sample; β’ νΉννν ν (0): All MIATTs agree that this is false for this sample; β’ νννννν€ν (1/2): When MIATTs have incomplete or conflicting information. Then the formalization of this idea can be: Suppose that each MIATT ν‘ ν β is given a true value ν£ ν β0,0.5,1 on the sample (ν,ν‘ Μ ). The true value of the synthetic target ν‘ β is: ν‘ β ( ν,ν‘ Μ ) = 1, νν βν,ν£ ν =1 0, νν βν,ν£ ν =0 0.5, νν‘βννν€νν ν (ννν₯νν/ν’ννννν€ν) . (17) Thus ν‘ β becomes a three-valued logic true target and can be used directly to evaluate the model. 4.2.2 Evaluation algorithm based on three-valued logic true targets 1) Fact computation β’ For each sample (ν,ν‘ Μ ), calculate the true value of all MIATTs ν£ ν ; β’ Use the above rules to generate the composite truth value ν‘ β ( ν,ν‘ Μ ) . 2) Model score definition β’ For a sample, the score is: ννννν ν‘ β ( ν ) =ν‘ β ( ν,ν‘ Μ ) . (18) That is: ο² If target = 1, the prediction is completely correct β score 1; ο² If target = 0, the prediction is completely wrong β score 0; ο² If target = Β½, the prediction is "uncertain/insufficient" given the partial true target β score 0.5. β’ Overall score for the dataset: ννννν Μ ν‘ β = 1 | ν· | β ννννν ν‘ β ( ν ) νβν· . (19) 4.3 Comparison between the two methods for evaluation with MIATTs Evaluation with the original MIATTs relies on a parallel multi-perspective logic, emphasizing coverage and multifaceted correctness, and enabling fine-grained diagnostic analysis. In contrast, evaluation with a ternary target synthesized from MIATTs is grounded in a synthetic three-valued logic goal, yielding a compressed single truth target that trades detailed information for a unified, directly scoreable measure. The former resembles logical set semantics, where multiple partial truths coexist in parallel, while the latter approximates a ground truth within a multi-valued logic framework. A detailed comparison of the two approaches to MIATTs-based evaluation is provided in Table 2. Table 2. Comparison of two approaches to MIATTs-based evaluation Feature Evaluation with original MIATTs Evaluation with ternary target synthesized from MIATTs Granularity Preserves the βpartial truthβ of each MIATT, showing where the model performs well/poorly Merged into a single unified target, losing detailed aspects Truth space Multiple values (0, 0.5, 1) + aggregation logic (min/max/Noisy-OR) Single value (0, 0.5, 1) Information loss Low; allows diagnosing βon which MIATT the model failsβ High; conflicts/partial coverage compressed into β0.5β Interpretability Strong; directly explains βon which inaccurate target the model does well/poorlyβ Weak; only states βcorrect/incorrect/uncertainβ Computational simplicity Complex; requires handling coverage, contradiction rate, etc. Simple; only needs merging rules Use case Suitable for analysis/research scenarios (high interpretability) Suitable for scenarios requiring a single metric (e.g., quick model comparison) 4.4 Relation between MIATTs-based and ATT-based evaluation We examine the logical and mathematical relations between the two MIATTs-based evaluation approaches and the standard accurate true target (ATT)-based evaluation. 4.4.1 Baseline: ATT ν β Assuming the underlying true target ν‘ β is fully definable, the evaluation is the same as normal evaluation with ATT: ννννν ν‘ β ( ν ) = 1 ν‘ Μ =ν‘ β ( ν ) 0 ν‘ Μ β ν‘ β ( ν ) . (20) Or in the continuous value scenario, use | ν‘ Μ βν‘ β ( ν )| , log-likelihood, etc. Logically, this is equivalent to having a complete Boolean truth function, without ambiguity or unknowns. 4.4.2 Relation between original MIATTs and baseline Logical Relationship: The whole set of original MIATTs are partial projections of ν‘ β and with additional noise (ννΉ ( ν‘ β ) β β ννΉ ( ν‘ ν β ) ν ν=1 βννΉ(ν‘ β )βͺν). Evaluation is equivalent to checking whether the model satisfies several logical clauses. The aggregation method (min/max/NoisyOR) determines whether we approximate ν‘ β more conservatively or more loosely. Mathematically: The whole set of original MIATTs is an upper approximation of ν‘ β with additional noise (ννΉ ( ν‘ β ) β β ννΉ ( ν‘ ν β ) ν ν=1 βννΉ(ν‘ β )βͺν). If the model is correct or error on all MIATTs, it must be correct or error on ν‘ β (a sufficient condition). But if the model is correct or error on a particular IATT, it does not necessarily mean it is correct or error on ν‘ β (it may simply have missed another aspect). Thus, evaluation with original MIATTs is complete but unsound: while it preserves the full information of ν‘ β , it can sometimes misclassify correct results as errors and errors as correct. 4.4.3 Relation between ternary target synthesized from MIATTs (ν β ) and baseline Logical Relationship: The ternary target synthesized from MIATTs (ν‘ β ) combines the results of multiple partial targets ν‘ ν β into a single three-valued domain: β’ 1: All aspects are consistently true β consistent with ν‘ β ; β’ 0: All aspects are consistently false β consistent with ν‘ β ; β’ 0.5: Inconsistent or insufficient information β indicates that we cannot determine the relationship with ν‘ β . Mathematically: The ν‘ β is equivalent to introducing an "uncertain" intermediate value on the classical truth value set 0,1. Thus, it is equivalent to constructing a three-valued upper approximation of ν‘ β with uncertainty: β’ If ν‘ β =1, then ν‘ β =1 (reliable); β’ If ν‘ β =0, then ν‘ β =0 (reliable); β’ If ν‘ β =0.5, then ν‘ β may be 0 or 1 (uncertain). So ν‘ β is partial but consistent: it will not mistake an error for a correct answer, nor a correct answer for an error, but it may abandon the judgment. 4.4.4 Relation between the three Intuitive analogy: β’ ν‘ β : Given a complete map, determine whether the path is correct. β’ MIATTS: Given only a few partial maps with noise, you can say "right/wrong" on these partial paths, but you can't guarantee the overall picture. β’ ν‘ β : Overlay all partial maps to create a simplified map; any conflicts or omissions are marked as "unknown." More detailed relation between the three are provided in Table 3. Table 3. Summarized relation between the three methods for evaluation. Method Mathematical relation to ν β Property Misjudgment risk Information loss ν β Exact equivalence Precise ground truth None None MIATTs Upper approximation (with extra noisy zone) Complete but unsound May sometimes yield misclassify correct results as errors and errors as correct Very high; probably can preserve the entire information of ν‘ β Synthesized ternary ν β Upper approximation (with uncertainty zone) Partial but consistent No misjudgment, but produces βuncertainβ cases High; some details compressed into uncertain (0.5) The geometry relation between the three can be shown as Fig. 3. Figure 3. Geometry relation between the three. Left: Upper approximation (with extra noisy zone) of MIATTs to ν‘ β . Right: Upper approximation (with uncertain zone) of synthesized ternary ν‘ β to ν‘ β . 4.4.5 Summary Direct evaluation with ATT (ν‘ β ) is the gold standard (but often unavailable). Evaluation with original MIATTs sacrifices unsoundness for completeness and interpretability under "multiple partial knowledge" with additional noise. Evaluation with triple-valued synthetic target (ν‘ β ) sacrifices detail for a single, unified, three-valued judgment. Therefore, evaluation with original MIATTs is more suitable for diagnostic analysis (which aspects are correct/incorrect). Triple-valued synthetic method is better for quickly comparing models overall, but they discard judgments on some samples. Both can be viewed as approximate logical alternatives to true evaluation when the ν‘ β is missing. 5. Learning with MIATTs: UTTL-Grounded Strategies Assuming that the true target for a given ML task is not assumed to exist as a well-defined object in the real world [11], UTTL provides a principled framework for learning from MIATTs, which indicates that UTTL can be effectively implemented through a multi-target learning paradigm [11, 14]. Accordingly, this section introduces UTTL-grounded learning strategies constructed upon two commonly used loss functionsβDice [20] and Cross Entropy (CE) [21]. When handling multiple targets (e.g., multiple annotators or probabilistic labels), two general aggregation strategies can be adopted: Method A (Per-target then Aggregate): compute the loss for each target separately and then aggregate the individual losses (e.g., through averaging or robust weighting); and Method B (Aggregate then Single Loss): first aggregate all targets (e.g., by averaging) and then compute a single loss with respect to the aggregated target. Although these two approaches may appear similar in form, their theoretical behavior diverges depending on the specific loss function. In the following discussion, we analyze and summarize the characteristics and implications of both methods under Dice and CE loss functions. 5.1 Comparison of Dice-based loss formulations for Methods A and B Let ν‘ Μ β(0,1) denote the model prediction for a given pixel (or soft mask value), and let ννΌν΄νν =ν‘ ν β |νβ 1,β―,ν β 0,1 denote for multiple inaccurate true targets (e.g. multiple soft targets or multiple measurement repeats). Define the per-pair Dice similarity ν·ννν ( ν‘ Μ ,ν‘ β ) = 2ν‘ Μ ν‘ β ν‘ Μ 2 +ν‘ β 2 . (21) We compare the following two training objectives (losses rewritten in βscoreβ form ν·ννν; if you use 1βν·ννν as the loss the inequalities below reverse accordingly): β’ Averaged per-target Dice (Avg-Dice, Method A): ν ν΄ (ν‘ Μ ,ννΌν΄νν =ν‘ ν β |νβ 1,β―,ν )= 1 ν β ν·ννν ( ν‘ Μ ,ν‘ ν β ) ν ν=1 . (22) β’ Dice of averaged targets (Dice-of-Mean, Method B): ν ν΅ (ν‘ Μ ,ννΌν΄νν =ν‘ ν β |νβ 1,β―,ν )=ν·ννν ( ν‘ Μ ,ν‘ β Μ ) , ν‘ β Μ = 1 ν β ν‘ ν β ν ν=1 . Below we analyze their mathematical relationships, optimization implications, robustness and practical recommendations. 5.1.1 Mathematical relation (Jensen bias) between Methods A and B Treating ν‘ Μ as fixed, consider ν·ννν as a function of ν‘ β . Its second derivative (with respect to ν‘ β ) is ν 2 ν·ννν ν 2 ν‘ β 2 = 4ν‘ Μ ν‘ β ( ν‘ β 2 β3ν‘ Μ 2 ) ( ν‘ Μ 2 +ν‘ β 2 ) 3 . (23) Consequently ν·ννν is concave in ν‘ β for ν‘ β 2 <3ν‘ Μ 2 (i.e. ν‘ β < β 3ν‘ Μ ) and convex for ν‘ β 2 >3ν‘ Μ 2 . By Jensenβs inequality: β’ If all ν‘ ν β lie in a concave region for the give ν‘ Μ then ν ν΅ =ν·ννν ( ν‘ Μ ,ν‘ β Μ ) β₯ 1 ν β β ( ν‘ Μ ,ν‘ ν β ) ν ν=1 =ν ν΄ . (24) Equivalently, using losses 1βν·ννν, Dice-of-Mean underestimates the average loss (is overly optimistic). β’ If all ν‘ ν β lie in a convex region, the inequality reverses: ν ν΅ β€ν ν΄ , (25) so Dice-of-Mean is more conservative. β’ If the ν‘ ν β span both regions (some below and some above β 3ν‘ Μ ), no universal ordering holds; numerical comparison is required. Equality ν ν΅ =ν ν΄ holds if ν‘ 1 β =ν‘ 2 β =β―=ν‘ ν β (or ν·ννν is affine over the convex hull of ν‘ ν β , which here only occurs at a degenerate set). Dice-of-Mean (Method B) introduces a systematic bias relative to the mean per-target Dice (Method A) because ν·ννν is nonlinear in ν‘ β ; the direction of bias depends on where the ν‘ ν β lie relative to the threshold β 3ν‘ Μ . 5.1.2 Gradient and optimization dynamics Compute derivative of ν·ννν with respect to ν‘ Μ : νν·ννν ( ν‘ Μ ,ν‘ β ) νν‘ Μ = 2ν‘ β ( ν‘ β 2 βν‘ Μ 2 ) ( ν‘ Μ 2 +ν‘ β 2 ) 2 . (26) Thus, gradients under the two schemes are β’ Avg-Dice (Method A): ν» ν‘ Μ ν ν΄ = 1 ν β 2ν‘ ν β ( ν‘ ν β 2 βν‘ Μ 2 ) ( ν‘ Μ 2 +ν‘ ν β 2 ) 2 ν ν=1 . (27) β’ Dice-of-Mean (Method B): ν» ν‘ Μ ν ν΅ = 2ν‘ β Μ ( ν‘ β Μ 2 βν‘ Μ 2 ) ( ν‘ Μ 2 +ν‘ β Μ 2 ) 2 . (28) Consequences:ν ν΄ aggregates per-target gradients and therefore preserves heterogeneity: extremes in individual ν‘ ν β produce strong corrective signals for ν‘ Μ . This supports faster correction of individual mismatches and is beneficial when multiple targets are informative but heterogeneous. ν ν΅ uses a single gradient computed at ν‘ β Μ . Strong individual signals from outlying ν‘ ν β are diluted by averaging; in effect, Dice-of-Mean yields lower gradient variance but may slow correction of biased/erroneous targets or rare but important modes. 5.1.3 Robustness to target noise and outliers Avg-Dice (Method A) is sensitive to outliers: a single erroneous ν‘ ν β contributes directly to the average and can dominate the update if its gradient magnitude is large. However, this sensitivity is controllable, one can apply per-target weights, trimmed means, or robust aggregators to mitigate outliers while retaining per-target supervision. Dice-of-Mean (Method B) naturally smooths isolated target noise via the averaging step, so single outliers have reduced immediate impact. Nevertheless, because of the nonlinearity of ν·ννν, a sufficiently large outlier may still shift ν‘ β Μ across the convexity threshold and induce a nontrivial bias. Recommendation: If annotator reliability is unknown and outliers are expected, use ν ν΄ with robust aggregation or weighted averaging; use ν ν΅ only when target noise is approximately zero-mean and targets are exchangeable. 5.1.4 Interpretability and diagnostic capability ν ν΄ provides per-target loss terms that are directly interpretable and diagnostic: one can detect which targets systematically disagree with the model and apply target-specific strategies (reweighting, calibration, or retraining). ν ν΅ obviates per-target diagnostics because only the aggregate label is used; this simplifies implementation but loses visibility into target heterogeneity 5.1.5 Computational considerations Computational cost difference is minor in typical deep learning settings (vectorized implementations make multiple evaluations cheap). ν ν΅ does one Dice computation; ν ν΄ does ν. If the number of targets is very large, cost matters and one may prefer an aggregated target. Otherwise, cost should not drive the choice. 5.1.6 Practical recommendations (1) When to prefer ν ν΄ (Avg-Dice) o Multiple targets with unknown or varying reliability. o Need for per-target diagnostics, reweighting or curriculum learning. o Desire to preserve strong corrective signals from individual labels (e.g. rare features or small structures). (2) When to prefer ν ν΅ (Dice-of-Mean) o Labels are many noisy observations drawn from a common unbiased process and computational simplicity is desired. o The training objective must match an inference scheme that uses averaged labels/predictions (i.e. you care about optimizing the mean output directly). (3) Hybrid and robust schemes o Use ν ν΄ together with a robust aggregator on the loss (e.g. trimmed mean, median of means). o Aggregate labels with a reliability-weighted mean ν‘ β Μ = β ν€ ν ν‘ ν β ν ν=1 (learn or estimate ν€ ν ), then apply ν·ννν ( ν‘ Μ ,ν‘ β Μ ) . This can combine the smoothing of ν ν΅ with weighted robustness. o Combine both terms in a composite loss: νΏ=Ξ»(1βν ν΄ )+(1βΞ»)(1βν ν΅ ), (29) where νβ[0,1] trades per-target supervision against ensemble-level optimization. (4) Validation monitoring o Monitor per-target metrics and ensemble metrics on a held-out set to detect Jensen bias and divergence between ν ν΄ and ν ν΅ behaviour. 5.1.7 Concise takeaway Because Dice similarity is nonlinear in targets, averaging before applying Dice (Dice-of- Mean, Method B) is not algebraically equivalent to averaging per-target Dice (Avg-Dice, Method A). The two choices encode a biasβvariance trade-off: β ν ν΄ preserves target-specific corrective signals (higher variance, greater fidelity), whereas ν ν΅ reduces variance via averaging but can introduce a systematic bias whose sign depends on the local concavity/convexity of ν·ννν relative to ν‘ Μ . β In most practical annotation settingsβwhere targets differ and diagnostics are desirableβν ν΄ (possibly with robust aggregation or weighting) is the safer default; ν ν΅ may be chosen when targets are exchangeable, noise is symmetric, and simplicity or inference-matching is prioritized. 5.2 Comparison of CE-based loss formulations for Methods A and B Letβs focus our discussion on binary/multiclass cross-entropy (CE) and rigorously compare two approaches: β’ Averaged per-target CE (Avg-CE, Method A): ν ν΄ (ν‘ Μ ,ννΌν΄νν =ν‘ ν β |νβ 1,β―,ν )= 1 ν β νΆνΈ ( ν‘ Μ ,ν‘ ν β ) ν ν=1 . (30) β’ Dice of averaged targets (CE-of-Mean, Method B): ν ν΅ (ν‘ Μ ,ννΌν΄νν =ν‘ ν β |νβ 1,β―,ν )=νΆνΈ ( ν‘ Μ ,ν‘ β Μ ) , ν‘ β Μ = 1 ν β ν‘ ν β ν ν=1 . (31) Below are the conclusions, algebraic proofs, and practical considerations. 5.2.1 Conclusions For standard cross entropy (no additional nonlinear transformations, no weighting terms, and targets with probability distributions or soft labels), Method A and Method B are completely equivalent: their numerical values, gradients, and optimization impact are identical. Therefore, in this case, there is no Jensen bias issue, as with Dice. However, in terms of implementation and engineering, Method A still retains more diagnostic/weighting flexibility (for example, the ability to weight or prune individual targets), while Method B is more concise but tends to βobscureβ individual information. 5.2.2 Algebraic proof (Taking binary CE as an example) Define binary CE: νΆνΈ ( ν‘ Μ ,ν‘ β ) =β [ ν‘ β logν‘ Μ + ( 1βν‘ β ) log ( 1βν‘ Μ )] , (32) where ν‘ Μ β(0,1) is the model output probability, and ν‘ β β[0,1] is the target probability (which can be a soft target or target average). Expand νΆνΈ ( ν‘ Μ ,ν‘ β ) by ν‘ β : νΆνΈ(ν‘ Μ ,ν‘ β )=βν‘ β logν‘ Μ β ( 1βν‘ β ) log ( 1βν‘ Μ ) =ν‘ β (βlogν‘ Μ +log ( 1βν‘ Μ ) )βlog ( 1βν‘ Μ ) . It can be seen that νΆνΈ ( ν‘ Μ ,ν‘ β ) is an affine (linear) function with respect to ν‘ β of the form ν ( ν‘ Μ ) ν‘ β +ν(ν‘ Μ ). Thus, for a set of ν‘ ν β : νΆνΈ ( ν‘ Μ ,ν‘ β Μ ) =ν ( ν‘ Μ ) ν‘ β Μ +ν ( ν‘ Μ ) =ν ( ν‘ Μ ) 1 ν β ν‘ ν β ν ν=1 +ν ( ν‘ Μ ) = 1 ν β (ν ( ν‘ Μ ) ν‘ ν β +ν ( ν‘ Μ ) ) ν ν=1 = 1 ν β νΆνΈ ( ν‘ Μ ,ν‘ ν β ) ν ν=1 . (33) That is ν ν΄ =ν ν΅ . (34) 5.2.3 Gradient level (Proving consistency) Taking the derivative with respect to ν‘ Μ , the derivative of the binary CE is linear with respect to ν‘ β : ννΆνΈ ( ν‘ Μ ,ν‘ β ) νν‘ Μ = ν‘ Μ βν‘ β ν‘ Μ ( 1βν‘ Μ ) . (35) Therefore, ν» ν‘ Μ ν ν΄ = 1 ν β ν‘ Μ βν‘ ν β ν‘ Μ ( 1βν‘ Μ ) ν ν=1 = ν‘ Μ βν‘ β Μ ν‘ Μ ( 1βν‘ Μ ) =ν» ν‘ Μ ν ν΅ . (36) So the gradients are exactly the same, and the training behavior is equivalent. 5.2.4 Multi-class case If we use multi-class cross entropy (categorical CE) and consider as class probability vectors, the same holds true: CE is linear in the target distribution, so νΆνΈ ( ν‘ Μ , 1 ν β ν‘ ν β ν ν=1 ) = 1 ν β νΆνΈ ( ν‘ Μ ,ν‘ ν β ) ν ν=1 . (37) 5.2.5 Practical notes and exceptions Although mathematically equivalent, engineering/practical considerations require the following: (1) Additive weights or nonlinear postprocessing can violate the equivalence: If different weights ν€ ν are introduced to each target's loss, then νΆνΈ ( ν‘ Μ ,ν‘ β Μ ) is no longer equal to β ν€ ν νΆνΈ ( ν‘ Μ ,ν‘ ν β ) ν ν=1 (unless the weights are equal and ν‘ β Μ are weighted equally). If nonlinear transformations (such as truncation, thresholding, temperature scaling, or logit-space operations) are performed on ν‘ ν β before computing the loss, the equivalence disappears. (2) Robust aggregation and outlier handling: If robust aggregations on outliers, such as de-extinction, mean truncating, or median truncation are performed, aggregating first and then performing CE will produce different results than performing CE first and then performing robust operations (because robust operations are often nonlinear). While mathematically equivalent, from an engineering perspective, you may want to retain the loss of a single target for diagnostic purposes, which is more convenient with Method A (although computationally, vectorization can recover the same results). (3) Target calibration/temperature/target smoothing: If target smoothing or temperature scaling is performed on the averaged soft targets, the behavior will differ from performing CE first and then averaging on each ν‘ ν β (this can differ depending on the order). (4) Precision and numerical stability: Due to floating-point arithmetic, the numerical errors between the two implementations may differ very slightly at extreme scales, but these are usually negligible. 5.2.6 Concise takeaway Advices for usage of the CE-based Methods A and B can be summarized as follows: β If the standard (unweighted, no additional nonlinearity) cross-entropy is used, the two methods are theoretically and optimizationally equivalent, so Methods A or B can be chosen based on implementation preference (averaging the labels first and then computing the loss is convenient for saving memory or simplifying code). β If diagnosing target behavior, weighting individual targets, or providing robustness to individual losses are needed, method A (single loss first, then aggregation) is more flexibleβeven though the mathematical equivalent is true, method A is easier to scale and debug. β If post-processing of the target (thresholding, truncation, temperature, weighting, etc.) is included in pipeline, be sure to re-analyze which order makes sense, as this will break the equivalence. 5.3 Summarization of Methods A and B under Dice and CE For Dice loss, Method A should be regarded as the default due to its preservation of target-specific corrective signals and its compatibility with weighting or robust aggregation schemes. While Method B can reduce variance and may simplify implementation, it inevitably introduces systematic bias, making it less reliable in heterogeneous annotation settings For CE loss, the situation differs fundamentally: under the standard form (no weighting, no nonlinear label transformations), Method A and Method B are mathematically and optimizationally equivalent. Thus, practical considerations such as memory footprint, code simplicity, or diagnostic flexibility should guide the choice. Importantly, once additional operations (e.g., temperature scaling, thresholding, or per-target weighting) are applied, the equivalence is broken, and Method A again becomes the safer and more generalizable choice. See Table 4. Table 4. Summarization of Methods A and B under Dice and CE loss Loss Method Equivalence & Bias Recommended Usage A B Dice Computes Dice per target and aggregates thereafter; preserves target-specific corrective signals, maintaining fidelity at the cost of higher variance. Aggregates targets first, then computes a single Dice score; variance is reduced but bias is introduced. Not equivalent. Due to nonlinearity of Dice, Method B incurs Jensen-type bias. The bias direction depends on local concavity/convexity of Dice with respect to the averaged target. Method A is safer in most practical annotation settings, especially when target-specific diagnostics or robustness are needed. Method B may be acceptable when targets are exchangeable, annotation noise is symmetric, and implementation simplicity is prioritized. CE (standard) Computes CE per target (with probability distributions or soft labels), then aggregates. Provides hooks for weighting, pruning, or diagnostics. Aggregates targets/labels first, then computes a single CE. Compact and memory- efficient. Completely equivalent under standard CE: identical values, gradients, and optimization behavior. No Jensen bias issue. Either method is acceptable for standard CE; choice can follow implementation preference. Method A remains advantageous when post-hoc weighting, target-specific robustness, or diagnostic analysis are required. Equivalence breaks if nonlinear post-processing (e.g., thresholding, temperature scaling, weighting) is applied to the targets. 6. Discussion The concept of MIATTs provides a pragmatic response to the epistemological limitation that the true target of a ML task is not assumed to exist as a well-defined object in the real world. Instead, MIATTs represent diverse, partially correct, and complementary views derived from task-specific AI models (AIM). The quality of a task-specific MIATTs set can be characterized by two key indicators: mean(PartialRepresentation), representing its coverage of the underlying true target, and (1-Redundancy), reflecting its diversity. As summarized in Table 1, higher coverage and lower redundancy together yield higher-quality MIATTs, while limited diversity or insufficient coverage leads to degraded representational fidelity. When the number of MIAs involved in generating MIATTs increases, the union of all MIATTs ( β ν‘ ν β ν ν=1 ) tends to approximate full coverage of the latent true target ν‘ β , though this inevitably introduces noise. Hence, an optimal MIATTs configuration should achieve a balance between completeness and consistencyβmaximizing semantic coverage while minimizing overlap and contradiction. This balance ensures that MIATTs collectively provide a robust, multi-perspective representation suitable for both evaluation and learning. The downstream influence of task-specific MIATTs is twofold. For evaluation, the balance between coverage and redundancy directly affects the soundnessβcompleteness trade-off: higher coverage enhances completeness but may reduce soundness due to noise, while moderate redundancy supports more stable, interpretable scoring under LAF-based evaluation. For learning, MIATTs diversity provides richer supervision signals that improve generalization and robustness in UTTL-grounded strategies, though excessive overlap or noise can introduce gradient conflicts or bias. Hence, maintaining a balanced MIATTs structureβadequate coverage with controlled redundancyβis essential for stable and reliable model evaluation and training within the EL-MIATTs framework. From the evaluation perspective, LAF offers a principled foundation for assessing models when ground truth cannot be assumed as a single precise label. Under LAF, evaluation can proceed in two main forms: (1) Evaluation using original MIATTs, which retains each MIATTβs partial truth and employs logical or fuzzy operators (e.g., conjunction, disjunction, t-norms, and t-conorms) for aggregation; and (2) Evaluation using a ternary target synthesized from MIATTs, which compresses multi-perspective truth into a unified three-valued label 0,0.5,1. The former provides fine-grained interpretability and preserves detailed coverage information, enabling diagnosis of where and why a model performs well or poorly. However, it is computationally complex and may yield inconsistent judgments across MIATTs. The latter approach simplifies computation and facilitates single-metric comparison but sacrifices information granularity. When using original MIATTs, LAF-grounded evaluation can be viewed as complete but not perfectly soundβit captures the full informational spectrum of ν‘ β yet may misinterpret partially correct predictions as errors or vice versa. Nevertheless, when combined with appropriate fuzzy aggregation (e.g., using min/max or Noisy-OR operators), LAF enables a balanced trade-off between interpretability, consistency, and practicality. In contrast, when using a ternary target synthesized from MIATTs, evaluation becomes sound but incompleteβ the unified three-valued representation ensures consistency and ease of comparison but inevitably compresses detailed multi-perspective information into a simplified truth structure. This trade-off enhances computational efficiency and interpretability at the global level but reduces diagnostic depth regarding which MIATT contributed to a specific outcome. Overall, LAF-grounded evaluationβwhether using original MIATTs or a synthesized ternary targetβ offers complementary perspectives. In complex, ill-defined tasks, the former better approximates traditional accurate-target evaluation by capturing nuanced partial truths, whereas in simpler or more deterministic tasks, the latter provides a more stable and pragmatic alternative. Deviations observed in either case largely reflect the intrinsic ambiguity of the task itself rather than methodological inadequacy. On the learning side, UTTL establishes a theoretical foundation for training models when the true target cannot be precisely defined. Within UTTL, MIATTs are treated as multiple weakly reliable surrogates of the underlying truth, allowing learning to proceed through multi-target optimization rather than single-target supervision. Two general strategies can be adopted: Method A (Per-target then Aggregate) computes the loss with respect to each MIATT separately and aggregates the individual losses (e.g., via averaging or weighted combination). Method B (Aggregate then Single Loss) first aggregates all MIATTs into a composite target and then computes a single loss against it. When implemented with Dice and Cross Entropy (CE) loss functions, these strategies exhibit distinct theoretical behaviors. Method A emphasizes robustness and diversity by allowing individual MIATT contributions to remain explicit, making it suitable when MIATTs vary substantially in reliability. Method B, by contrast, enforces global consistency and is computationally simpler but risks losing fine-grained information. UTTL therefore provides a flexible unifying framework in which both strategies can be adapted to task complexity, data uncertainty, and desired learning bias. Together, LAF-based evaluation and UTTL-based learning bridge the theoretical gap between logical semantics and statistical optimization in machine learning under uncertain supervision. The two are mutually reinforcing components within the EL-MIATTs paradigm. LAF provides a logic-grounded framework for evaluating model outputs against multiple partially true targets, ensuring that model assessment remains consistent with multi-valued truth semantics rather than collapsing ambiguity into binary correctness. UTTL, in turn, translates these same logical principles into a statistical learning mechanism, optimizing model parameters in a way that reflects the graded or composite nature of truth encoded by MIATTs. Through this integration, logical reasoningβexpressed via conjunction, disjunction, and fuzzy aggregation in LAFβis aligned with the probabilistic gradient-based optimization used in UTTL. This alignment allows evaluation and learning to operate within a shared semantic space, where logical consistency and statistical efficiency jointly contribute to model improvement. Consequently, EL-MIATTs not only evaluate but also train models in a manner that respects uncertainty, disagreement, and partial correctness, offering a unified pathway toward explainable, epistemically grounded, and practically effective machine learning. 7. Conclusion, Limitation, and Future Work This paper presents a systematic effort to bridge the theoretical formulation and practical implementation of the EL-MIATTs framework [11], emphasizing how LAF-based evaluation algorithms and UTTL-based learning strategies enable reliable model assessment and training under epistemic uncertainty. By grounding the evaluation process in the LAF and the learning process in the UTTL paradigm [13, 14], we operationalize EL-MIATTs into two coherent, complementary mechanisms. Specifically, we analyzed the structural and semantic properties of task-specific MIATTs [12] and its downstream influence on evaluation and learning, proposed two LAF-grounded evaluation schemes (parallel multi-perspective and ternary synthesized), and developed two UTTL-grounded learning strategies (Per-target then Aggregate and Aggregate then Single Loss) applicable to common loss functions such as Dice and Cross Entropy. Together, these developments enhance the robustness, interpretability, and adaptability of predictive models in settings where the βtrueβ target is uncertain, incomplete, or inherently multi-faceted. Furthermore, the discussion about the integration of LAF and UTTL illustrates how logical semantics and statistical optimization can be unified within the EL-MIATTs framework, offering a principled pathway toward more explainable and uncertainty-aware machine learning. Despite its theoretical soundness and practical flexibility, the current implementation of EL-MIATTs remains subject to several limitations. First, the construction of task-specific MIATTs depends on the availability and diversity of task-relevant AI models (AIM), which may constrain generalizability in data-sparse domains. Second, the proposed LAF-based aggregation and UTTL-based optimization methods may involve additional computational overhead due to multi-target representation and aggregation operations. Third, while the framework provides strong interpretability in uncertainty reasoning, its performance may vary depending on the alignment between the MIATTs structure and the underlying target distribution, requiring careful calibration of aggregation and weighting parameters. Finally, the current work focuses primarily on non-specific task settings; extension to classification, segmentation, sequential, generative, or reinforcement learning scenarios remains an open challenge. Future research will extend EL-MIATTs along several promising directions: (1) Adaptive MIATT Generation and Weighting: develop self-adaptive mechanisms for generating and weighting MIATTs based on dynamic task characteristics or model uncertainty; (2) Paraconsistent and Multi-valued Reasoning: integrate paraconsistent logic and higher-order multi-valued semantics into LAF to better handle conflicting or contradictory MIATTs; (3) Scalable Implementation: explore efficient computational approximations or distributed architectures to reduce the cost of MIATT-based aggregation and optimization; and (4) Cross-domain and Real-world Proof-of-Concept: validate EL-MIATTs in broader real-world contexts, such as medical imaging, autonomous systems, and human-in-the-loop AI, where uncertainty and partial truth are intrinsic. In summary, this work lays a theoretical and algorithmic foundation for bridging logical reasoning and statistical learning in uncertainty-aware machine learning. Future advancements in EL-MIATTs are expected to further strengthen its capacity to support robust, interpretable, and epistemically grounded AI systems. An application of this workβs results is presented as part of the study available at https://w.qeios.com/read/EZWLSN. Reference 1. Deng W, Zheng L. Are Labels Always Necessary for Classifier Accuracy Evaluation? Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021. p. 15069β78. 2. Jung HJ, Lease M. Evaluating Classifiers Without Expert Labels. 2012. https://doi.org/10.48550/arxiv.1212.0960. 3. Chang H, Zhuang AH, Valentino DJ, Chu WC. Performance measure characterization for evaluating neuroimage segmentation algorithms. NeuroImage. 2009. https://doi.org/10.1016/j.neuroimage.2009.03.068. 4. Taha A, Hanbury A. Metrics for evaluating 3D medical image segmentation: analysis, selection, and tool. BMC Medical Imaging. 2015;15:29. https://doi.org/10.1186/s12880- 015-0068-x. 5. M H, M.N S. A Review on Evaluation Metrics for Data Classification Evaluations. International Journal of Data Mining & Knowledge Management Process. 2015;5:01β11. https://doi.org/10.5121/ijdkp.2015.5201. 6. Song H, Kim M, Park D, Shin Y, Lee J-G. Learning From Noisy Labels With Deep Neural Networks: A Survey. IEEE Trans Neural Netw Learning Syst. 2023;34:8135β53. https://doi.org/10.1109/TNNLS.2022.3152527. 7. Fatima S, Ali S, Kim H-C. A Comprehensive Review on Multiple Instance Learning. Electronics. 2023;12:4323. https://doi.org/10.3390/electronics12204323. 8. Yang X, Song Z, King I, Xu Z. A Survey on Deep Semi-Supervised Learning. IEEE Trans Knowl Data Eng. 2023;35:8934β54. https://doi.org/10.1109/tkde.2022.3220219. 9. Supervised Learning. Web Data Mining. Berlin, Heidelberg: Springer Berlin Heidelberg; 2011. p. 63β132. https://doi.org/10.1007/978-3-642-19460-3_3. 10. Yang Y. Moderately supervised learning: definition, framework and generality. Artif Intell Rev. 2024;57:37. https://doi.org/10.1007/s10462-023-10654-6. 11. Yang Y. EL-MIATTs: Evaluation and Learning with Multiple Inaccurate True Targets. 2026. https://doi.org/10.32388/UMHEFG.4. 12. Yang Y. Bridging Theory and Practice in Implementing EL-MIATTs: Logic-Driven Algorithms for MIATTs Generation and Assessment. 2025. https://doi.org/10.32388/0UD1AN. 13. Yang Y. Logical assessment formula and its principles for evaluations with inaccurate ground-truth labels. Knowl Inf Syst. 2024. https://doi.org/10.1007/s10115-023-02047-6. 14. Yang Y. Undefinable True Target Learning: Towards Learning with Democratic Supervision. 2025. https://doi.org/10.32388/KBK3P8.5. 15. Fahmi A, Amin F, Aslam M, Yaqoob N, Shaukat S. T-norms and T-conorms hesitant fuzzy Einstein aggregation operator and its application to decision making. Soft Comput. 2021;25:47β71. https://doi.org/10.1007/s00500-020-05426-1. 16. Boixader D, Recasens J. Vague and fuzzy t-norms and t-conorms. Fuzzy Sets and Systems. 2022;433:156β75. https://doi.org/10.1016/j.fss.2021.07.008. 17. Negri M. An algebraic completeness proof for Kleeneβs 3-valued logic. Bollettino DellβUnione Matematica Italiana. 2002;5-B:447β67. 18. Robles G, MΓ©ndez JM. Natural implicative expansions of variants of Kleeneβs strong 3- valued logic with GΓΆdel-type and dual GΓΆdel-type negation. Journal of Applied Non- Classical Logics. 2021;31:130β53. https://doi.org/10.1080/11663081.2021.1948285. 19. Rasiowa H. Stephen Cole Kleene. Introduction to metamathematics. North-Holland Publishing Co., Amsterdam, and P. Noordhoff, Groningen, 1952; D. van Nostrand Company, New York and Toronto 1952; X + 550 p. J Symb Log. 1954;19:215β6. https://doi.org/10.2307/2268620. 20. Milletari F, Navab N, Ahmadi S-A. V-net: Fully convolutional neural networks for volumetric medical image segmentation. 2016 fourth international conference on 3D vision (3DV). Ieee; 2016. p. 565β71. 21. Mao A, Mohri M, Zhong Y. Cross-Entropy Loss Functions: Theoretical Analysis and Applications. In: Krause A, Brunskill E, Cho K, Engelhardt B, Sabato S, Scarlett J, editors. Proceedings of the 40th International Conference on Machine Learning, vol. 202. PMLR; 2023. p. 23803β28. 22. Barrio E, Da Re B. Paraconsistency and its Philosophical Interpretations. AJL. 2018;15:151β 70. https://doi.org/10.26686/ajl.v15i2.4860. 23. Ciuciura J. Literal and Controllable Paraconsistency. LLP. 2024;1β19. https://doi.org/10.12775/LLP.2024.027. 24. Miranda B, Bertolino A. An assessment of operational coverage as both an adequacy and a selection criterion for operational profile based testing. Software Qual J. 2018;26:1571β 94. https://doi.org/10.1007/s11219-017-9388-0. 25. Hossain SB, Dwyer MB. A Brief Survey on Oracle-based Test Adequacy Metrics. 2022. https://doi.org/10.48550/ARXIV.2212.06118. 26. Bienvenu M, Inoue K, Kozhemiachenko D. Abductive Reasoning in a Paraconsistent Framework. Proceedings of the TwentyFirst International Conference on Principles of Knowledge Representation and Reasoning. Hanoi, Vietnam: International Joint Conferences on Artificial Intelligence Organization; 2024. p. 134β44. https://doi.org/10.24963/kr.2024/13. 27. Belnap ND. How a Computer Should Think. In: Omori H, Wansing H, editors. New Essays on Belnap-ΒDunn Logic, vol. 418. Cham: Springer International Publishing; 2019. p. 35β 53. https://doi.org/10.1007/978-3-030-31136-0_4.