Paper deep dive
Demonstration of the common dual-channel feature decoupling characteristic of front-door mediation causal inference methods in whole-slice image classification
Zhirui Zhang, Tianhang Nan, Yong Ding, Zhuolun Song, Dayu Hu, Xiaoyu Cui
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/18/2026, 11:02:40 AM
Summary
The paper proposes and validates two hypotheses regarding front-door intervention-based Multi-Instance Learning (MIL) for Whole Slide Image (WSI) classification. It argues that these methods inherently create a dual-channel architecture: a baseline channel capturing biased observational features and an independent causal channel executing front-door intervention. The effectiveness of eliminating false correlations is determined by the feature divergence between these channels, which increases deep feature diversity and suppresses spurious associations caused by unobserved confounders like staining protocols.
Entities (10)
Relation Signals (7)
CAMELYON16 β contains β Breast Cancer
confidence 97% Β· Camelyon16 (dataset of breast cancer sentinel lymph node metastases
Front-door intervention β uses β Multi-instance learning
confidence 95% Β· Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images
Front-door intervention β targets β unobserved confounders
confidence 94% Β· eliminating unobserved confounders in MIL
Feature Divergence β eliminates β False correlations
confidence 93% Β· Greater difference between features extracted by the new and baseline channels increases effectiveness in eliminating false correlations
Front-door intervention β creates β Dual-channel architecture
confidence 92% Β· mainstream approaches ... essentially construct a dual-pathway classification architecture
Dual-channel architecture β consistsof β Causal channel
confidence 90% Β· an independent causal channel executing the interventional expectation
Dual-channel architecture β consistsof β Baseline channel
confidence 90% Β· a baseline channel capturing the biased observational distribution
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in digital pathology. These methods adjust feature distributions of subtle evidence sub-images to correctly associate them with WSI-level diagnoses. We propose and prove 2 hypotheses for evaluating such methods: 1) Causal inference MIL introduces an independent classification channel that effectively completes WSI classification; 2) Greater difference between features extracted by the new and baseline channels increases effectiveness in eliminating false correlations. This hypothesis describes the core of causal inference MILs: overlaying parallel, independent channels to eliminate false associations between WSI-level diagnostic and non-diagnostic evidence sub-images by increasing deep feature diversity. Based on these hypotheses, we evaluated several causal inference MILs on breast cancer and non-small cell lung cancer datasets. This hypothesis provides a new theoretical perspective for applying causal inference to WSI analysis.
Tags
Links
- Source: https://arxiv.org/abs/2607.12376v1
- Canonical: https://arxiv.org/abs/2607.12376v1
Trouble viewing inline? Open PDF directly β
Full Text
37,814 characters extracted from source content.
Expand or collapse full text
DEMONSTRATING THE DUAL-CHANNEL FEATURE DECOUPLING CHARACTERISTIC COMMON TO FRONT-DOOR INTERVENTION CAUSAL INFERENCE METHODS IN WHOLE SLIDE IMAGE CLASSIFICATION Zhirui Zhang 1# Tianhang Nan 1# , Yong Ding 1 , Zhuolun Song 1 , Dayu Hu 1* , Xiaoyu Cui 1* Authors masked with # have the same contribution. 1. College of Medicine and Biological Information Engineering, Northeastern University, China ABSTRACT Causal inference using front door intervention and multi-instance learning (MIL) has advanced the analysis of Whole Slide Images (WSI) in digital pathology. These methods adjust feature distributions of subtle evidence sub-images to correctly associate them with WSI-level diagnoses. We propose and prove 2 hypotheses for evaluating such methods: 1) Causal inference MIL introduces an independent classification channel that effectively completes WSI classification; 2) Greater difference between features extracted by the new and baseline channels increases effectiveness in eliminating false correlations. This hypothesis describes the core of causal inference MILs: overlaying parallel, independent channels to eliminate false associations between WSI-level diagnostic and non- diagnostic evidence sub-images by increasing deep feature diversity. Based on these hypotheses, we evaluated several causal inference MILs on breast cancer and non-small cell lung cancer datasets. This hypothesis provides a new theoretical perspective for applying causal inference to WSI analysis. Index TermsβCausal Inference, Interpretability Evaluation, Multi Instance Learning, Whole Slide Image Diagnosis 1. INTRODUCTION As the βgold standardβ for disease diagnosis, the precision of pathology profoundly influences clinical decision-making [1-3]. The maturation of whole-slide imaging (WSI) technology has propelled histopathological analysis into the digital era, enabling computational analysis of gigapixel-level images. Multiple instance learning (MIL), as a canonical paradigm of weakly supervised learning [4, 5], addresses these challenges by decomposing WSIs into thousands of image patches (instances) and training classification models using only WSI-level labels [6-8]. This approach assumes that at least one critical instance (e.g., a cancerous region) is relevant to the global label and aggregates instance features via attention mechanisms to achieve global predictions[9]. Although MIL significantly reduces annotation requirements, it remains a correlation-based statistical learning method with the following limitations: 1. Traditional aggregators treat all instance features equally, failing to distinguish causal features (e.g., cellular atypia) from confounding features (e.g., staining intensity). 2. When training data exhibit selection bias (e.g., a hospital using different staining protocols for slides of different diseases), models may erroneously solidify non-causal associations (e.g., staining style- diagnosis outcome) as decision-making criteria[10]. Recent advances in front-door intervention-based causal inference have provided a novel pathway for eliminating unobserved confounders in MIL. According to Pearlβs causal hierarchy theory [11], the front-door criterion establishes a causal chain νβνβν via a mediator variable, making it particularly suitable for scenarios involving unmeasurable confounders (e.g., institutional staining protocols) [12]. In MIL tasks, this framework treats image patches as input variables ν. Local histopathological features (ν) influence diagnostic outcomes (ν) through deep learning representations (ν), while confounding factors such as scanner variability (ν), constrained by the feature extractor and aggregator) can only indirectly affect predictions via the νβν pathway (Figure 1.a). By applying the front-door adjustment [11- 13], the model can block the confounding path νβνβνβν without directly observing ν (Figure 1.b & c). Although front-door intervention-based causal inference methods provide mathematical guarantees for eliminating unobserved confounders at the theoretical level[14], their clinical application faces significant challenges: pathologists remain skeptical about the interpretability of black-box causal modeling. Existing studies predominantly validate method effectiveness through classification performance metrics (e.g., accuracy, AUC) [15, 16]. However, due to their failure to elucidate the correspondence between causal intervention mechanisms and histopathological evidence, as well as the lack of explainable causal relationships in their results, these studies suffer from insufficient clinical trustworthiness. By deconstructing the front-door adjustment formula and the forward propagation paths of network parameters, we found that mainstream approaches (e.g., feature re-weighting based on the do- operator[17], stratified sampling of mediator variables[18]) essentially construct a dual-pathway classification architecture: the baseline pathway learns correlation patterns in the original feature space, while the causal pathway extracts deconfounded causal patterns through front-door intervention (Figure 1.d). We establish a universal theoretical framework for front-door intervened MILs grounded in Structural Causal Models (SCM) and Information Theory. We rigorously derive and validate two core principles: 1. The causal pathway can operate independently of the baseline pathway while maintaining diagnostic efficacy. Computing the interventional distribution ν(ν|νν(ν)) necessitates a marginalization over a global prior P(X'). We prove that this mathematical operation strictly requires the model to bifurcate into a dual-pathway architecture: a baseline channel capturing the biased observational distribution, and an independent causal channel executing the interventional expectation. 2. The degree of feature divergence ν₯ between the dual pathways determines the deconfounding effect. Defining the inner product of feature vectors as the inter-channel feature divergence, the causal pathway preferentially attends to instances exhibiting greater WSI-level feature discrepancies compared to the baseline pathway (Figure 1.e). This represents an information-complementary decoupling that mitigates spurious correlations arising from the baseline pathway's overreliance on input dependencies. We demonstrate that the geometric feature divergence ν₯ between the dual pathways is not a mere empirical similarity metric, but a strict variational upper bound for confounder isolation. Let ν and νΎ denote the representations extracted by the baseline and causal channels, respectively. According to the Data Processing Inequality (DPI), suppressing the mutual information νΌ(νΎ; ν) stringently limits the entanglement between the causal feature νΎ and the unobserved confounder ν. Consequently, we establish the following optimization equivalence: maxΞ ( ν,νΎ ) βΊminνΌ ( νΎ;ν ) βminUpperBound(νΌ ( νΎ;ν ) ) This formulation quantifies the confounding-blocking effect: maximizing ν₯ mathematically forces the causal channel to project features orthogonally to the ν-dependent subspace, thereby fulfilling the unconfoundedness prerequisite of the front-door criterion. A larger ν₯ indicates that ν has a more significant effect from ν in the baseline pathway, while the causal pathway successfully eliminates the confounding interference. On Camelyon16 (dataset of breast cancer sentinel lymph node metastases with manual pixelwise annotations), these instances exhibited substantial overlap with ground truth annotations, demonstrating that the causal pathway establishes genuine associations between model predictions and Figure 1.a-c. Causal diagrams for no intervention, back-door intervention and front-door intervention. d. Workflow of front-door intervention. e. Patches relevant to diagnosis exhibit substantial dissimilarity from the WSI-level features in the baseline channel. diagnostically evidential patches. These two principles elucidate the underlying mechanism of front-door-intervened causal MIL: by constructing parallel yet orthogonal feature representation spaces, feature diversity suppresses spurious correlations. Unlike conventional methods that solely optimize feature aggregation mechanisms[6, 7], this structural innovation shifts the modelβs focus from βhow to fuseβ to βhow to decoupleβ, better aligning with the clinical reasoning logic of βcritical evidence identificationβ in pathological diagnosis. Through systematic validation of the aforementioned dual- channel hypotheses, this study elucidates the interpretability formation mechanism in front-door-intervened causal MIL. Thus, we develop a rigorous and broadly applicable evaluation framework for establishing explanatory rationality in practical pathological image analysisβan aspect that previously lacked standardized assessment. Within this framework, we conducted a unified evaluation of existing front-door-intervened causal MIL methods, demonstrating their effectiveness and interpretability from this novel perspective. The findings reveal both the common limitations and potential improvement directions for front-door intervention methods at the interpretability level, thereby providing new empirical foundations for the emerging field of causality-aware interpretable WSI-assisted diagnosis. 2. METHOD 2.1 Backdoor Path in Whole Slice Image Diagnosis Tasks. In a standard Whole Slide Image (WSI) classification task, let the macroscopic variable ν denote the complete pathological image and ν denote the definitive diagnostic outcome (e.g., benign or malignant). In an ideal, unbiased environment devoid of any data acquisition artifacts, the diagnostic prediction relies exclusively on genuine morphological evidence present in the image. This establishes a direct, unconfounded causal pathway from the WSI to the diagnosis: ν βν. Under this assumption, the objective of the predictive model is to fit a deterministic function ν ( β ) that maps the WSI directly to the predicted diagnosis ν: ν=ν ( ν )( 1 ) In this ideal state, the model's predictive distribution ν ( ν | ν ) perfectly aligns with the true causal mechanism, capturing only the pathological variations without any spurious interference. However, such ideal assumptions may not always hold in real- world observational data. To mathematically formalize the realistic data generating mechanism without introducing ambiguous exogenous noise variables, this task can be described as a causal graph based on causal theory and the Structural Causal Model (SCM) framework as shown in Figure 1.a. Let ν denote a shared, unobserved variable representing global confounders (e.g., varying institutional staining protocols, scanner color profiles, and tissue processing artifacts). The causal relationships among the pathological image ν, the diagnosis ν, and the confounder ν are structurally defined by the following topological links: ν βν: The true causal effect where genuine morphological features dictate the diagnosis. ν βν: The unobserved confounder affects the observed appearance distribution of the WSI, including staining style, scanner-dependent color profiles, tissue-processing artifacts, and other non-diagnostic visual variations. Although these factors do not constitute genuine pathological morphology, they are embedded in the input image ν. Due to the intrinsic representational bias and limited pathological understanding of neural image encoders, the model may fail to completely separate these ν-induced visual variations from true diagnostic evidence. ν βν: Selection bias in data collection (e.g., a specific hospital with a unique staining protocol ν predominantly collects severe malignant cases ν), establishing a correlation between the confounder and the diagnostic label. Based on this causal DAG, the presence of the shared confounder ν inherently derives a spurious link, known in causal graph theory as the backdoor path: ν βν βν. The presence of this path implies that, in practice, when a model attempts to learn the prediction of ν Μ from input ν, it inadvertently captures correlations confounded by variable ν. Consequently, in realistic settings, the resulting observed distribution can be expanded as: P(ν Μ β£ β£ ν)=νΌ β€βΌβ(β€ β£ β£ ν) [ν(ν Μ β£ β£ ν,ν)] ( 2 ) This equation reveals that the learned ν Μ is effectively an approximation biased by ν. Driven by the term ν(νβ£ν), the model tends to exploit the spurious correlation between ν and ν Μ , thereby learning a predictive "shortcut" rather than capturing the true causal mechanism defined in Equation (1). 2.2 Micro Confounding and Bias in the MIL Framework To reveal the formation mechanism of predictive bias, the macro- level process outlined in Section 2.1 is formally decomposed into instance-level variables under the MIL framework. MIL models are tasked with detecting the presence of positive instances (e.g., localized tumor regions) within a bag containing numerous instances. If at least one positive instance is present, the bag is classified as positive; otherwise, it is classified as negative. Assuming a WSI is represented as a bag ν= ν₯ 1 ,ν₯ 2 ,...,ν₯ ν comprising ν instance-level patches, with corresponding unobserved instance-level labels ν¦ 1 ,ν¦ 2 ,...,ν¦ ν where ν¦ ν β0,1. According to the standard MIL assumption, if at least one positive instance is present, the bag is classified as positive; otherwise, it is classified as negative. The WSI-level bag label ν β 0,1 can then be strictly formulated as follows: Y= 0, νν βν¦ ν =0 ν ν=1 1, νν‘βννν€νν ν ( 3 ) In practice, the network cannot directly access the biological ground truth ν¦ ν . Instead, it approximates the true diagnostic label ν through a parameterized predictive model. Under the standard MIL neural instantiation, the prediction ν Μ is derived by utilizing an aggregator ΞΈ ν ( β ) to fuse instance-level features, followed by a classification module ΞΈ ν ( β ) to map the aggregated bag-level representation to the final diagnosis: Y Μ = ΞΈ c (ΞΈ ν ( ν₯ 1 , ν₯ 2 ,...,ν₯ ν ) ) ( 4 ) In a perfectly unconfounded scenario, the modules ΞΈ ν and ΞΈ ν would exclusively process genuine pathological morphologies, seamlessly aligning ν Μ with the true ν. To mathematically formalize this ideal alignment and the unconfounded causal mechanism, we can expand the true predictive distribution ν ( ν | ν ) by marginalizing over the entire latent instance-level label space ν¦ 1 ,ν¦ 2 ,...,ν¦ ν . According to the law of total probability and the structural characteristics of biological tissues, this marginalization is governed by two critical properties: 1. As established in Equation (3), the WSI-level label ν is strictly and solely determined by the latent instance labels ν¦ ν . This renders the raw instance features ν₯ ν conditionally independent of ν given ν¦ ν , meaning ν ( ν β£ β£ ν¦ 1 ,...,ν¦ ν ,ν₯ 1 ,...,ν₯ ν ) = ν ( ν β£ β£ ν¦ 1 ,...,ν¦ ν ) . This term functions mathematically as a deterministic indicator. 2. In real biological tissues, pathological regions exhibit spatial continuity and microenvironmental correlations. The true biological states of the patches are not mutually independent. Therefore, the joint probability ν ( ν¦ 1 ,...,ν¦ ν β£ β£ ν₯ 1 ,...,ν₯ ν ) holistically captures the true underlying pathological state driven by the spatially correlated morphological features across the entire WSI. By incorporating the conditional independence from the deterministic aggregation property directly into the total probability expansion, the ideal, unconfounded MIL predictive distribution can be elegantly formulated as: ν ( ν | ν ) =βν ( ν β£ β£ ν¦ 1 ,...,ν¦ ν ,ν₯ 1 ,...,ν₯ ν ) ν ( ν¦ 1 ,...,ν¦ ν β£ β£ ν₯ 1 ,...,ν₯ ν ) ν¦ ν β0,1 ν =βν ( ν β£ β£ ν¦ 1 ,...,ν¦ ν ) ν ( ν¦ 1 ,...,ν¦ ν β£ β£ ν₯ 1 ,...,ν₯ ν ) ν¦ ν β0,1 ν ( 5 ) This establishes the unconfounded, ideal micro-causal chain: the true instance morphology ν₯ ν determines ν¦ ν , which collectively determines ν. However, the empirically fitted ν Μ is inevitably influenced by the macroscopic confounder ν. To formalize this microscopic confounding mechanism, the global confounder ν is structurally decomposed into a hierarchical set: ν= ν β² ,ν§ 1 ,ν§ 2 ,...,ν§ ν .Crucially, these decomposed components strictly correspond to the spurious paths defined in our macroscopic causal DAG: ν ν ,ν ν ,...,ν ν : These correspond to the localized manifestations of the νβνpath at the patch level. Each ν§ ν denotes the non- diagnostic visual variation embedded in the corresponding patch ν₯ ν , such as local staining fluctuation, scanner-induced color shift, tissue-processing artifact, or background texture bias. Due to the intrinsic representational limitations of neural image encoders, such local confounding factors may be entangled with genuine pathological morphology during feature extraction, leading to biased instance-level representations. ν β² :This corresponds to the ν βν path. νβ² represents systemic selection biases injected during the MIL aggregation and classification process (e.g., the model erroneously associating a hospital's specific background stroma color with severe malignancy). Consequently, from an instantiated perspective, the network's actual prediction process is no longer a pure function of ν. The feature extraction is coupled with ν§ ν , and the classification is modulated by ν β² . The formulation of the network prediction structurally degenerates into a confounded mapping: ν Μ =ΞΈ ν ( ΞΈ ν ( ( ν₯ ν ,ν§ ν ) ν=1 ν ) ,ν β² )( 7 ) Building upon the instantiated network architecture, we can now expand the observational predictive distribution ν(ν Μ |ν) for the given input bag ν=ν₯ 1 ,ν₯ 2 ,...,ν₯ ν . By marginalizing over the entire latent confounder spaceβcomprising both the local instance- level artifacts ν§ ν and the global bag-level bias ν β² βthe law of total probability yields: ν(ν Μ β£ β£ ν₯ 1 ,...,ν₯ ν ) =βν ν§ 1 ,...,ν§ ν ν β² (ν Μ β£ β£ ν₯ 1 ,...,ν₯ ν ,ν§ 1 ,...,ν§ ν ,ν β² )ν ( ν§ 1 ,...,ν§ ν ,ν β² β£ β£ ν₯ 1 ,...,ν₯ ν ) (8) Specifically, the term ν(ν Μ β£ν₯ 1 ,...,ν₯ ν ,ν§ 1 ,...,ν§ ν ,ν β² )corresponds to the network computation process detailed above, demonstrating that the genuine morphological features ν₯ ν become inevitably entangled with confounding variables during both feature extraction and aggregation. Meanwhile, the term ν(ν§ 1 ,...,ν§ ν ,ν β² β£ ν₯ 1 ,...,ν₯ ν )reveals that, due to its inherent lack of causal awareness, the model resorts to shortcut learning by directly deriving spurious correlations from the confounders. For instance, the network may infer the hospital-level global staining bias ν β² by capturing local color artifacts ν§ ν within a given patch ν₯ ν , and subsequently exploit this non-causal association to generate the diagnostic prediction ν Μ . By comparing the ideal causal predictive distribution in Equation (5) with the instantiated confounded distribution in Equation (8), we can formally define the core optimization objective of a causal MIL framework. The fundamental goal is to eliminate the predictive bias by minimizing the discrepancy between the network's confounded empirical prediction ν(ν Μ |ν) and the true, unconfounded causal distribution ν(ν|ν): min|ν(ν Μ |ν)βν ( ν | ν ) | =ννν|βν(ν Μ | ν₯ ν ,ν§ ν ,νβ²)ν(ν§ ν ,νβ² | ν₯ ν ) νβ²,ν§ ν ββν(ν | ν¦ ν )ν(ν¦ ν | ν₯ ν ) ν¦ ν | ( 9 ) This mathematical expansion transparently exposes the structural root of the predictive bias. The right-hand term represents the ideal causal mechanism defined in Equation (5), which is driven exclusively by the latent instance-level diagnostic states ν²= ν¦ ν ν=1 ν . In contrast, the left-hand term represents the networkβs empirically fitted observational prediction, where the prediction is additionally modulated by the posterior distribution of the latent confounders ν=(ν§ ν ν=1 ν ,ν β² ). In the SCM, the predictive bias arises because the latent confounders are structurally associated with both the observed WSI features and the diagnostic outcome, corresponding to the open backdoor path νβνβν. 2.3 The basic process of front door intervention. Therefore, this paper introduces causal intervention to reduce the confounding bias revealed in Equation (9). In Pearlβs causal framework, the conventional conditional probability ν(ν Μ β£ β£ ν) describes a passive observation of ν, whose prediction tends to absorb spurious correlations induced by the unobserved confounder ν. In contrast, ν(ν Μ β£νν(ν))describes an active intervention on ν, where the natural image-generation mechanism is replaced by a controlled feature-generation mechanism. νν(ν) is introduced to optimize ν into a purified feature distribution. This purified distribution is expected to be decoupled from the visual appearance of the original WSI and the acquisition- related confounding factors, while preserving diagnostically meaningful morphological features. In the SCM, this operation aims to cut off all incoming edges to ν, namely blocking the spurious backdoor path νβν, thereby making νstrictly independent of the unobserved confounder ν. The mediator variable ν is introduced as the concrete realization of νν(ν), where νdenotes the purified feature representation generated after intervening on ν. Since ν cannot be directly measured in practical WSI diagnosis, directly estimating ν(ν Μ β£νν(ν))through back-door adjustment is infeasible. In the presence of unobserved confounders, front-door intervention provides a more feasible solution. As shown in Figure 1.c, this method identifies the causal effect through an observable mediator variable ν, which transmits diagnostic information from νto ν. Under the deterministic mapping assumption of neural- network-based feature extraction, the standard front-door adjustment formula can be simplified as an expectation over the global input prior distribution: ν ( ν Μ |νν ( ν ) ) =νΌ ν β² βΌβ ( ν ) [ν(ν Μ β£ β£ ν ( ν ) ,ν β² )] ( 10 ) 2.4 Mechanism of Reducing Prediction Bias in Dual-Pathway Model After obtaining the computable front-door intervention distribution in Eq. (10), we further abstract a common computational form shared by front-door intervention-based MIL methods. We argue that the practical essence of the front-door intervention operator in MIL can be reduced to a dual-channel processing mechanism. Specifically, the same WSI-level input ν is processed from two different perspectives: ν 1 = Ξ¦ 1 ( ν ) ,ν 2 = Ξ¦ 2 ( ν ) , ( 11 ) where Ξ¦ 1 (β ) and Ξ¦ 2 (β ) denote two channel-specific transformation functions. Here, ν 1 and ν 2 are the channel-level representations of the same WSI bag ν. The two channels are not required to be individually causal; they may both be observational feature channels. The subsequent fusion and disentanglement of these two representations can induce the causal effect of νν(ν), yielding ν(ν Μ β£νν(ν)). Thus, the mediator variable in the front-door formula can be written as: ν=Ξ¨(ν 1 ,ν 2 ),(12) where Ξ¨(β ,β ) denotes a learnable operator for constructing the interventional mediator representation. Under the ideal information- preserving scenario, ν retains the diagnosis-relevant information contained in the joint representation ( ν 1 , ν 2 ) . The information obtained by the two channels under different processing mechanisms should provide different diagnostic perspectives for the same WSI input, emphasizing different aspects of pathological morphology, such as spatial organization, texture patterns, frequency-domain variations, or contextual dependencies. The information contained in these different perspectives should also be non-overlapping diagnostic information. In this paper, conditional Shannon entropy ν»(β β£β ) is introduced to formalize this point. To illustrate why such a dual-channel structure can outperform a single-channel MIL model, we take ν 1 as a representative single channel in the following discussion. Specifically, ν»(νβ£ν 1 ) measures the remaining uncertainty of the ideal WSI-level diagnostic outcome ν after observing the first- channel representation ν 1 , while ν»(νβ£ν 1 ,ν 2 ) measures the remaining diagnostic uncertainty after jointly observing both channel representations. The diagnosis-relevant non-redundancy condition is formulated as ν»(νβ£ν 1 )βν»(νβ£ν 1 ,ν 2 )>0.(13) Equation (13) states that, after ν 1 is known, the addition of ν 2 reduces the diagnostic uncertainty, i.e., the second channel still provides additional information about the ideal WSI-level diagnostic outcome ν. To demonstrate that the causal intervention is effective, we aim to show that the theoretically optimal front-door interventional discrepancy is smaller than the theoretically optimal observational discrepancy.The KL divergence is introduced as a distributional discrepancy measure ν· νΎνΏ (β ||β ). |ν ( ν Μ |νν ( ν ) ) βν(ν|ν)|β|ν(ν Μ |ν)βν(ν|ν)| =νΌ ν β² βΌν(ν) [ ν· KL ( ν(νβ£ν β² ) β₯ ν(ν Μ β£νν(ν β² )) )] βνΌ ν β² βΌν ( ν ) [ν· KL ( ν ( ν β£ β£ ν β² ) β₯ ν ( ν Μ β£ β£ ν β² ) ) ] =ν» ( ν | ν 1 ,ν 2 ) βν» ( ν | ν 1 ) <0 ( 14 ) Equation (14) shows that, under the diagnostic complementarity condition and the ideal information-preserving mediator assumption, the KL divergence between the theoretically optimal dual-channel interventional prediction and the ideal unconfounded causal predictive distribution is smaller than that between the theoretically optimal single-channel observational prediction and the ideal distribution. That is, the interventional prediction ν(ν Μ β£νν(ν)) is expected to be closer to the ideal unconfounded causal predictive distribution ν(νβ£ν) than the original observational prediction ν(ν Μ β£ν). This agrees with the optimization direction of causal MIL, namely reducing the discrepancy between the model prediction and the ideal causal target through feature-space intervention. Therefore, under the stated assumptions, the dual- channel front-door intervention provides a computable mechanism for mitigating confounding-induced prediction bias. 2.5 Additional conditional assumptions under non ideal conditions The analysis in Section 2.4 describes an ideal setting in which the dual-channel representation contains diagnosis-relevant non- redundant information and the mediator ν΄=ν³(νΏ ν ,νΏ ν ) preserves such information during fusion. However, in practical MIL models, these two requirements are not automatically satisfied. Therefore, we propose that abstracting the front-door intervention method into a dual-channel representation entails two additional conditions. First, each channel should possess independent and sufficient diagnostic capability. This condition requires that neither channel degenerates into purely noisy or non-diagnostic transformations. The remaining diagnostic uncertainty of a single channel, ν―(νβ£ νΏ ν ), should be sufficiently small, such that each channel, when used alone, can support baseline-level WSI classification. This requirement ensures that both νΏ ν and νΏ ν retain sufficient pathological evidence to predict the ideal WSI-level diagnostic outcome ν. Second, under the premise of independent diagnostic capability, a larger inter-channel feature discrepancy can serve as an indicator of stronger potential diagnostic complementarity. A larger discrepancy between the two channels indicates that they process the same WSI input from more distinct feature perspectives. For example, image-domain and frequency-domain transformations are unlikely to form strict subset relations, because they emphasize different aspects of pathological morphology. When both channels possess independent and sufficient diagnostic capability, non- diagnostic information is less likely to dominate either channelβ otherwise, excessive non-diagnostic information would impair the single-channel diagnostic performance. Therefore, a larger inter- channel discrepancy is more likely to correspond to non-overlapping diagnostic information, and the diagnostic complementarity between the two channels is positively correlated with ν―(νΏ ν β£νΏ ν ). 2.6 Network Instantiation of the Additional Assumptions Given a WSI bag ν=ν₯ ν ν=1 ν , two feature pathways extract the patch-level representations β 1ν =Ξ¦ 1 ( ν₯ ν ) ,β 2ν =Ξ¦ 2 ( ν₯ ν ) ( 15 ) Each pathway follows the general MIL classification structure. For the ν-th pathway, νβ1,2, the attention score and attention weight of the ν-th patch are calculated as ν νν = ν ν β€ β νν β ν ,ν νν = exp(ν νν ) βexp( ν ν=1 ν νν ) ( 16 ) where ν ν is a learnable attention vector and ν is the feature dimension. The corresponding WSI-level representation is ν§ ν =βν νν ν ν=1 β νν ( 17 ) Thus, both pathways contain patch-level feature extraction, attention-based aggregation, and WSI-level representation learning, and each can independently support slide-level classification when it retains sufficient pathological evidence. This satisfies the first assumption that both pathways possess independent diagnostic capability. The two pathways are then fused through a difference-driven attention mechanism. Let β Λ νν = β νν β₯β νν β₯ 2 ( 18 ) denote the normalized patch representation, and let the normalized first-pathway WSI representation be expressed as ν§Λ 1 =βν 1ν ν ν=1 β Λ 1ν ( 19 ) For each second-pathway patch feature, its discrepancy from the first-pathway WSI-level representation is defined via the inner product: ν ν =1ββ Λ 2ν β€ ν§Λ 1 ( 20 ) A larger ν ν indicates that the second-pathway patch feature is more different from the dominant WSI-level representation captured by the first pathway. Since ν§Λ 1 is an attention-weighted combination of the first-pathway patch features, this discrepancy can be expanded as ν ν =1ββ Λ 2ν β€ βν 1ν ν ν=1 β Λ 1ν =βν 1ν ν ν=1 (1ββ Λ 2ν β€ β Λ 1ν ) ( 21 ) Thus, the discrepancy between a second-pathway patch feature and the first-pathway WSI-level representation is exactly the attention- weighted average discrepancy between that feature and the first- pathway patch-level representations. Hence, a larger distance from β Λ 2ν to ν§Λ 1 indicates a larger average feature difference from the diagnostically emphasized patch distribution of the first pathway. The discrepancy scores are converted into fusion weights through softmax normalization: νΌ ν = exp(ν ν ) βexp( ν ν=1 ν ν ) , ( 22 ) so that second-pathway features that differ more from the first- pathway representation receive larger fusion weights. The attention- weighted second-pathway representation is ν 2 =βνΌ ν ν ν=1 β 2ν , ( 23 ) and the interventional mediator is constructed by fusing this representation with the first-pathway WSI-level feature: ν=ν§ 1 +ν ν ν 2 , ( 24 ) where ν ν is a learnable feature transformation. This thus satisfies the second assumption: when both pathways independently preserve diagnostic information, a larger feature difference indicates that the second pathway is more likely to provide pathological evidence not already emphasized by the first pathway. The difference-driven attention mechanism assigns greater contribution to these non-redundant features and incorporates them into the mediator ν. 3. REFERENCES [1] T. Nan et al., "Deep learning quantifies pathologistsβ visual patterns for whole slide image diagnosis," Nature Communications, vol. 16, no. 1, p. 5493, 2025/07/01 2025, doi: 10.1038/s41467-025- 60307-1. [2] M. Lu, D. Williamson, T. Chen, R. Chen, M. Barbieri, and F. Mahmood, "Data-efficient and weakly supervised computational pathology on whole-slide images," (in English), Nat. Biomed. Eng, vol. 5, no. 6, p. 555β+, Jun 2021, doi: 10.1038/s41551-020-00682- w. [3] S. Hajdu, "A Note from History: Microscopic Contributions of Pioneer Pathologists," (in English), Ann. Clin. Lab. Sci., vol. 41, no. 2, p. 201β206, Spr 2011. [Online]. Available: <Go to ISI>://WOS:000294371100015. [4] Z. Zhou, "A brief introduction to weakly supervised learning," (in English), Natl. Sci. Rev., Review vol. 5, no. 1, p. 44β53, Jan 2018, doi: 10.1093/nsr/nwx106. [5] G. Campanella et al., "Clinical-grade computational pathology using weakly supervised deep learning on whole slide images," Nature medicine, vol. 25, no. 8, p. 1301β1309, 2019. [6] M. Ilse, J. Tomczak, and M. Welling, "Attention-based Deep Multiple Instance Learning," presented at the Proceedings of the 35th International Conference on Machine Learning, Proceedings of Machine Learning Research, 2018. [Online]. Available: https://proceedings.mlr.press/v80/ilse18a.html. [7] Z. Shao et al., "TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification," in 35th Conference on Neural Information Processing Systems (NeurIPS), Electr Network, Dec 06β14 2021, LA JOLLA: Neural Information Processing Systems (Nips), in Advances in Neural Information Processing Systems, 2021. [Online]. Available: <Go to ISI>://WOS:000925183303031. [Online]. Available: <Go to ISI>://WOS:000925183303031 [8] H. Li et al., "Task-specific Fine-tuning via Variational Information Bottleneck for Weakly-supervised Pathology Whole Slide Image Classification," in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, CANADA, Jun 17β24 2023, LOS ALAMITOS: Ieee Computer Soc, in IEEE Conference on Computer Vision and Pattern Recognition, 2023, p. 7454β7463, doi: 10.1109/cvpr52729.2023.00720. [Online]. Available: <Go to ISI>://WOS:001058542607078 [9] J. Yao, X. Zhu, J. Jonnagaddala, N. Hawkins, and J. Huang, "Whole slide images based cancer survival prediction using attention guided deep multiple instance learning networks," Medical Image Analysis, vol. 65, p. 101789, 2020/10/01/ 2020, doi: https://doi.org/10.1016/j.media.2020.101789. [10] M. PavloviΔ et al., "Improving generalization of machine learning-identified biomarkers using causal modelling with examples from immune receptor diagnostics," Nature Machine Intelligence, vol. 6, no. 1, p. 15β24, 2024/01/01 2024, doi: 10.1038/s42256-023-00781-8. [11] J. Pearl, Causality. Cambridge university press, 2009. [12] I. R. Fulcher, I. Shpitser, S. Marealle, and E. J. Tchetgen Tchetgen, "Robust inference on population indirect causal effects: the generalized front door criterion," Journal of the Royal Statistical Society Series B: Statistical Methodology, vol. 82, no. 1, p. 199β 214, 2020. [13] A. N. Glynn and K. Kashin, "Front-door versus back-door adjustment with unmeasured confounding: Bias formulas for front- door and hybrid adjustments with application to a job training program," Journal of the American Statistical Association, vol. 113, no. 523, p. 1040β1049, 2018. [14] L. Jiao et al., "Causal Inference Meets Deep Learning: A Comprehensive Survey," Research, vol. 7, p. 0467, 2024, doi: doi:10.34133/research.0467. [15] H. Jin and C. X. Ling, "Using AUC and accuracy in evaluating learning algorithms," IEEE Transactions on Knowledge and Data Engineering, vol. 17, no. 3, p. 299β310, 2005, doi: 10.1109/TKDE.2005.50. [16] G. Naidu, T. Zuva, and E. M. Sibanda, "A review of evaluation metrics in machine learning algorithms," in Computer science on- line conference, 2023: Springer, p. 15β25. [17] T. Nan, Y. Ding, H. Quan, D. Li, M. Zou, and X. Cui, "Establishing truly causal relationship between whole slide image predictions and diagnostic evidence subregions in deep learning," arXiv e-prints, p. arXiv: 2407.17157, 2024. [18] K. Chen, S. Sun, and J. Zhao, "CaMIL: Causal Multiple Instance Learning for Whole Slide Image Classification," in 38th AAAI Conference on Artificial Intelligence (AAAI) / 36th Conference on Innovative Applications of Artificial Intelligence / 14th Symposium on Educational Advances in Artificial Intelligence, Vancouver, CANADA, Feb 20β27 2024, PALO ALTO: Assoc Advancement Artificial Intelligence, in AAAI Conference on Artificial Intelligence, 2024, p. 1120β1128. [Online]. Available: <Go to ISI>://WOS:001239880400050. [Online]. Available: <Go to ISI>://WOS:001239880400050 [19] X. Cui, W. Chen, and J. Su, "A Multiscale Frequency Domain Causal Framework for Enhanced Pathologic al Analysis," 2025. [Online]. Available: https://openreview.net/forum?id=6xrDPHhwD3. [Online]. Available: https://openreview.net/forum?id=6xrDPHhwD3 [20] A. Vaswani et al., "Attention Is All You Need," in 31st Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, CA, Dec 04β09 2017, vol. 30, LA JOLLA: Neural Information Processing Systems (Nips), in Advances in Neural Information Processing Systems, 2017. [Online]. Available: <Go to ISI>://WOS:000452649406008. [Online]. Available: <Go to ISI>://WOS:000452649406008 [21] K. He, X. Zhang, S. Ren, J. Sun, and IEEE, "Deep Residual Learning for Image Recognition," in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, Jun 27β30 2016, NEW YORK: Ieee, in IEEE Conference on Computer Vision and Pattern Recognition, 2016, p. 770β778, doi: 10.1109/cvpr.2016.90. [Online]. Available: <Go to ISI>://WOS:000400012300083 [22] O. Russakovsky et al., "ImageNet Large Scale Visual Recognition Challenge," (in English), Int. J. Comput. Vis., vol. 115, no. 3, p. 211β252, Dec 2015, doi: 10.1007/s11263-015-0816-y. [23] T. Lin, Z. Yu, H. Hu, Y. Xu, C. Chen, and IEEE, "Interventional Bag Multi-Instance Learning On Whole-Slide Pathological Images," in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, CANADA, Jun 17β24 2023, LOS ALAMITOS: Ieee Computer Soc, in IEEE Conference on Computer Vision and Pattern Recognition, 2023, p. 19830β19839, doi: 10.1109/cvpr52729.2023.01899. [Online]. Available: <Go to ISI>://WOS:001062531304015 [24] H. Zhang et al., "DTFD-MIL: Double-Tier Feature Distillation Multiple Instance Learning for Histopathology Whole Slide Image Classification," in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, Jun 18β24 2022, LOS ALAMITOS: Ieee Computer Soc, in IEEE Conference on Computer Vision and Pattern Recognition, 2022, p. 18780β18790, doi: 10.1109/cvpr52688.2022.01824. [Online]. Available: <Go to ISI>://WOS:000870783004059