Paper deep dive
Evaluating Causal Discovery Algorithms for Path-Specific Fairness and Utility in Healthcare
Nitish Nagesh, Elahe Khatibi, Thomas Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/22/2026, 5:27:07 AM
Summary
This paper evaluates causal discovery algorithms (PC, GES, FCI, NOTEARS, DAGMA, DAG-GNN) on their ability to recover ground-truth causal structures and support path-specific fairness analysis in healthcare datasets (Alzheimer's and Heart Failure). The study introduces a framework to measure the Causal Fairness-Utility Ratio (CFUR), finding that structural recovery significantly impacts the decomposition of direct, indirect, and spurious fairness effects, with FCI performing best on real-world clinical data.
Entities (6)
Relation Signals (3)
CFUR ā quantifiestradeoffbetween ā Fairness and Utility
confidence 98% Ā· We use the causal fairness utility ratio to quantify the trade-off between fairness gain and accuracy loss
Peter-Clark (PC) ā achievedbeststructuralrecoveryon ā Alzheimer's Disease Dataset
confidence 95% Ā· On synthetic data, Peter-Clark achieved the best structural recovery.
Fast Causal Inference (FCI) ā achievedhighestutilityon ā Heart Failure Clinical Records Dataset
confidence 95% Ā· On heart failure data, Fast Causal Inference achieved the highest utility.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Causal discovery in health data faces evaluation challenges when ground truth is unknown. We address this by collaborating with experts to construct proxy ground-truth graphs, establishing benchmarks for synthetic Alzheimer's disease and heart failure clinical records data. We evaluate the Peter-Clark, Greedy Equivalence Search, and Fast Causal Inference algorithms on structural recovery and path-specific fairness decomposition, going beyond composite fairness scores. On synthetic data, Peter-Clark achieved the best structural recovery. On heart failure data, Fast Causal Inference achieved the highest utility. For path-specific effects, ejection fraction contributed 3.37 percentage points to the indirect effect in the ground truth. These differences drove variations in the fairness-utility ratio across algorithms. Our results highlight the need for graph-aware fairness evaluation and fine-grained path-specific analysis when deploying causal discovery in clinical applications.
Tags
Links
- Source: https://arxiv.org/abs/2603.15926v1
- Canonical: https://arxiv.org/abs/2603.15926v1
Trouble viewing inline? Open PDF directly ā
Full Text
28,531 characters extracted from source content.
Expand or collapse full text
Evaluating Causal Discovery Algorithms for Path-Specific Fairness and Utility in Healthcare Nitish Nagesh1, Elahe Khatibi1, Thomas Hughes1, Mahdi Bagheri1, Pratik Gajane2, Amir M. Rahmani1 1 University of California Irvine 2 University of OrlĆ©ans, France Abstract Causal discovery in health data faces evaluation challenges when ground truth is unknown. We address this by collaborating with experts to construct proxy ground-truth graphs, establishing benchmarks for synthetic Alzheimerās disease and heart failure clinical records data. We evaluate the Peter-Clark, Greedy Equivalence Search, and Fast Causal Inference algorithms on structural recovery and path-specific fairness decomposition, going beyond composite fairness scores. On synthetic data, Peter-Clark achieved the best structural recovery. On heart failure data, Fast Causal Inference achieved the highest utility. For path-specific effects, ejection fraction contributed 3.37 percentage points to the indirect effect in the ground truth. These differences drove variations in the fairness-utility ratio across algorithms. Our results highlight the need for graph-aware fairness evaluation and fine-grained path-specific analysis when deploying causal discovery in clinical applications. 1 Introduction Determining which causal pathways drive disparity in health outcomes is a critical task in informatics for goals including targeting interventions, assessing comparative fairness of prediction models, and deciding which effects are legally or ethically permissible to adjust [5]. Disparities can arise through direct effects of protected attributes on outcomes, indirect effects mediated by clinical variables, or spurious effects from confounders. Causal fairness frameworks decompose total variation into direct, indirect, and spurious components, enabling a nuanced understanding of which pathways contribute to disparity [6, 14]. The trade-off between fairness and predictive utility [15, 16] quantifies the cost of blocking each pathway. One challenge is that a known causal graph is required to identify which variables act as mediators and which as confounders. Causal discovery algorithms learn graph structure from observational data [19, 9, 12], yet evaluating whether discovered graphs support reliable path-specific fairness analysis in clinical settings remains an open question. We address this gap by establishing expert-defined benchmarks and evaluating causal discovery for path-specific fairness on both synthetic and real-world clinical data. Our work makes the following contributions: ⢠We establish a causal graph benchmark for a real-world clinical dataset and ground our evaluation in a synthetic clinical benchmark. ⢠We map each discovered graph to a fairness model and apply causal discovery algorithms, evaluating both structural recovery and path-specific fairness decomposition. ⢠We examine the trade-off between fairness and utility by evaluating the causal fairness utility ratio per path, enabling fine-grained analysis of which pathways offer the best fairness gain per unit accuracy cost. Figure 1: Proposed Discovery and Evaluation Framework 2 Methods In this section, we describe our proposed architecture, datasets, ground truth graphs, causal discovery algorithms, evaluation metrics, and experimental setup. 2.1 Proposed Architecture We develop a pipeline to evaluate causal discovery algorithms for utility and fairness. Using a combination of domain knowledge and expert-driven inputs, we establish the ground truth causal graph. We then discover the underlying causal graph by running causal discovery algorithms. We evaluate the discovered graph on utility and causal fairness [14] metrics accounting for disentangled effects of individual mediators and confounders on outcome. Finally, we examine tensions between fairness and utility through the causal fairness utility ratio [16]. Our framework is outlined in Figure 1. 2.2 Datasets To evaluate the utility and fairness metrics, we leverage two datasets. We first use a synthetically generated Alzheimerās disease dataset that has a well defined ground truth graph [1]. Then we consider another real-world dataset related to heart failure [28] where we work with an expert to establish the causal graph. We generate synthetic data based on the known structural causal model [1] for the Alzheimerās disease dataset. The parameters under consideration are sex, education, age, apoe4, moca, av45, tau, brain volume and ventricular volume. Sex is the protected attribute. Age is as the name suggests. Education refers to the number of years of education. APOE4 is a genetic risk factor. Tau is a biomarker that suggests cognitive decline. MOCA is the Montreal Cognitive Assessment Score. Brain volume refers to the total brain matter volume. Ventricular volume refers to the total ventricular volume. We omit slice number and Brain MRI variables from our analysis since those details are references to the raw Alzheimerās disease dataset in the repo. We initially begin with experimental setup for generating linear dataset and then add the effect of unobserved confounders through a latent variable as is commonplace [23]. We generate 1,000 samples. Figure 2: Ground truth causal graph derived for Alzheimerās Disease dataset [1]. We use the heart failure clinical records dataset [7] for evaluating causal discovery algorithms and path-specific effects. The dataset comprises 299 participants with a combination of demographic and clinical features and is open source, making it well suited for analysis. The features include demographic variables (age, gender), comorbidities (anaemia, diabetes, hypertension status, smoking status), physiological measurements (serum creatinine, serum sodium, ejection fraction, platelets, creatinine phosphokinase), time from admission to death, and the mortality outcome. 2.3 Ground Truth Causal Graph We use the Alzheimerās disease benchmark dataset based on the paper in the causal modeling agents paper [1]. The graph is as shown in Fig 2. We selected a graph from a single expert to demonstrate the effect of protected attributes on the outcome along with the mediators and confounders. Figure 3: Benchmark causal graph for Heart Failure Clinical Record Dataset [7]. We establish the causal graph benchmark shown in Figure 3 by working closely with a domain expert. Higher blood pressure and elevated serum creatinine indicate presence of chronic kidney disease causing electrolyte imbalance leading to lower serum sodium [21]. Further, advanced heart failure increases likelihood of cardiac injury leading to elevated creatine phosphokinase (CPK) levels [2]. Lifestyle and metabolic factors act as critical precursors in this network. Smoking serves as a significant accelerator by causing elevated blood pressure and doubling the risk of developing heart failure [10]. Similarly, untreated hypertension eventually causes heart failure and a subsequent reduction in ejection fraction [13]. Diabetes acts as a central node in this causal structure, directly leading to kidney disease and frequently co-existing with high blood pressure to exacerbate heart failure progression [25]. Anemia often emerges as a consequence of CKD, forcing the heart to work harder to pump oxygen-deficient blood, thereby worsening heart failure outcomes [18]. Serum creatinine levels may rise due to direct damage from diabetes or as a secondary effect of low ejection fraction, leading to fluid retention and decreased serum sodium [18]. Finally, a lack of consistent follow-up, particularly in patients with chronic conditions like diabetes, leads to poor disease management and an increased probability of a death event [11]. 2.4 Causal Discovery We perform causal discovery using constraint-based, score-based, and continuous-optimization methods. We chose this mix to compare how different algorithm families recover structure under mixed-type data and latent confounders. For the Alzheimerās disease dataset, we ran five algorithms: PC [19, 9], GES [8], NOTEARS [26, 27], DAGMA [3], and DAG-GNN [22]. For the heart failure clinical records dataset, we report PC, GES, and FCI (Fast Causal Inference). We included FCI because it extends PC to handle latent confounders and selection bias, which are common in observational clinical data. PC and GES were implemented via causal-learn [24], using the Fisher-Z conditional independence test at α=0.05α=0.05 for PC and BIC scoring for GES. NOTEARS used the original Zheng et al. implementation [26]. We tuned Ī»ā0.001,0.01,0.1Ī»ā\0.001,0.01,0.1\ and threshold ā0,0.1,0.3ā\0,0.1,0.3\, selecting Ī»=0.1Ī»=0.1 and threshold 0 by F1. DAGMA and DAG-GNN were implemented via gcastle. For DAGMA we tuned Ī»ā0.001,0.01,0.02,0.05Ī»ā\0.001,0.01,0.02,0.05\ and selected Ī»=0.05Ī»=0.05. For DAG-GNN we tuned threshold ā0.1,0.3ā\0.1,0.3\ and selected 0.30.3. This Alzheimerās disease setup replicates [20] and [1]. For the heart failure clinical records dataset, we omitted NOTEARS, DAGMA, and LiNGAM because they produced sparse graphs with poor structural recovery in preliminary runs. 2.5 Evaluation Utility. We use structural metrics to assess how well discovered graphs match the ground truth: F1 score, structural Hamming distance (SHD), false discovery rate (FDR), true positive rate (TPR), and false positive rate (FPR). These metrics are standard in causal discovery benchmarks and allow comparison across algorithms. Causal Fairness. Going beyond generic fairness, causal fairness [14] decomposes total variation (TV) into direct (Ctf-DE), indirect (Ctf-IE), and spurious (Ctf-SE) components. The decomposition captures the impact of mediators and confounders on outcome disparity. The distinction between these metrics is discussed in [17]. We apply the CFA decomposition using the Standard Fairness Model. For the Alzheimerās disease dataset, the protected attribute is sex, the outcome is ventricular volume, the mediators are Montreal Cognitive Assessment and brain volume, and the confounders are education, age, APOE4, av45, and tau. For the heart failure clinical records dataset, we derive mediators and confounders from each discovered graph: mediators are variables on directed paths from the protected attribute to the outcome, and confounders are ancestors of the outcome that are neither the protected attribute nor mediators. Causal Fairness Utility Ratio (CFUR). We use the causal fairness utility ratio [16] to quantify the trade-off between fairness gain and accuracy loss when blocking each path (direct, indirect, spurious). 2.6 Experimental Setup Experiments were run on Linux with an NVIDIA RTX 3090 GPU. We implemented causal discovery using the algorithms described above. For fairness, we applied the CFA decomposition for composite and individual path-specific effects [14] and the causal fairness utility ratio [16]. For the Alzheimerās disease dataset, we generated 1,000 synthetic samples. For fairness decomposition, we used 200 bootstrap samples for confidence intervals. For the heart failure clinical records dataset, we used 299 participants. Bootstrap sample sizes were 30 for discovery metrics and 200 for fairness decomposition. Code to reproduce all experiments is available at https://github.com/nitish-nagesh/causal-discovery-fairness. 3 Results We evaluated causal discovery algorithms on utility and path-specific fairness for the Alzheimerās disease and heart failure clinical records datasets. Each subsection presents one major finding with supporting tables and figures. Discovered Graph Structures. To assess structural recovery, we ran five causal discovery algorithms (PC, GES, NOTEARS, DAGMA, DAG-GNN) on the Alzheimerās disease dataset against expert-defined ground truth. For the heart failure clinical records dataset, we ran PC, GES, and FCI. 3.1 Alzheimerās Disease Dataset Utility. To answer how well each algorithm recovered the Alzheimerās disease ground truth structure, we computed F1, SHD, FDR, TPR, and FPR. Table 1 reports the results. PC achieved the best F1 score of 0.50 and lowest structural Hamming distance of 13. GES was second-best with F1 score of 0.42 and structural Hamming distance of 16. Continuous-optimization methods NOTEARS, DAGMA, and DAG-GNN performed poorly on this mixed-type dataset. DAG-GNN yielded the lowest F1 score of 0.13. Having established structural recovery, we next evaluated fairness decomposition on the same graphs. Table 1: Causal discovery utility on synthetic Alzheimerās disease dataset. Algorithm F1 SHD FDR TPR FPR PC (Fisher-Z) 0.50 13 0.27 0.52 0.12 GES (BIC) 0.42 16 0.53 0.38 0.26 NOTEARS 0.42 18 0.40 0.43 0.53 DAGMA 0.26 18 0.60 0.19 0.18 DAG-GNN 0.13 20 0.82 0.10 0.26 Composite Causal Fairness. To decompose sex-based disparity on ventricular volume into direct, indirect, and spurious components, we applied the CFA decomposition (TV = Ctf-DE ā- Ctf-IE ā- Ctf-SE) [14]. Table 2 shows the results. TV was consistent across graphs at 0.105, reflecting data-driven disparity. The ground truth decomposed TV into direct effect of 0.108, indirect effect of ā-0.025, and spurious effect of 0.028. Discovered graphs PC and GES collapsed to direct effect only, with Ctf-IE and Ctf-SE equal to zero. This collapse reflects structural misspecification in the discovered graphs. We next examined variable-level contributions to the spurious component. Table 2: Composite causal fairness for Alzheimerās disease: decomposition by graph. Algorithm Ctf-DE Ctf-IE Ctf-SE Ground truth 0.108 ā-0.025 0.028 PC 0.105 0.000 0.000 GES 0.105 0.000 0.000 Individual Path-Specific Effects. To identify which confounders drove the spurious effect, we decomposed Ctf-SE by variable for the ground truth graph. Table 3 reports the contributions. Education contributed 2.31%, age 0.27%, and APOE4 3.97% positively. Av45 contributed ā-12.28% and tau ā-7.66% negatively. These variable-level contributions support clinical necessity analysis and intervention prioritization. We then evaluated the trade-off between fairness gain and accuracy loss per path. Table 3: Individual contributions to Ctf-SE for Alzheimerās disease ground truth, confounders. Variable Contribution (%) education ++2.31 age ++0.27 apoe4 ++3.97 av45 ā-12.28 tau ā-7.66 Causal FairnessāUtility Ratio (CFUR). To quantify the trade-off between fairness gain and accuracy loss per path, we computed CFUR for each algorithm [16]. Table 4 reports CFUR per path as mean ± SD. The direct effect (DE) had the highest CFUR across algorithms. Blocking the direct sex to ventricular volume path yielded the most fairness gain per unit accuracy cost. Spurious effect (SE) had negative CFUR for most algorithms, indicating that blocking confounder paths increased loss with little fairness benefit. The ground truth CFUR profile was more balanced than discovered graphs. These Alzheimerās disease results establish the utilityāfairness trade-off on a controlled benchmark. We next evaluate the same metrics on real-world heart failure clinical records data. Table 4: CFUR by path for Alzheimerās disease, mean ± SD. Algorithm CFUR DE CFUR IE CFUR SE Ground truth ++15.5 ± 9.1 ā-1.8 ± 8.0 ā-0.2 ± 0.1 PC ++48.5 ± 40.0 ++0.6 ± 1.6 ā-0.1 ± 0.1 GES ++342 ± 676 ++0.4 ± 1.6 ++1.7 ± 6.7 NOTEARS ++5.7 ± 4.4 ā-0.3 ± 0.3 ++8.6 ± 15.3 DAGMA ++3.0 ± 1.3 ā-0.0 ± 1.8 ā-1.0 ± 2.3 3.2 Heart Failure Clinical Records Dataset Utility. To evaluate structural recovery on HFCR, we compared PC, GES, and FCI against the expert-defined ground truth. Table 5 reports F1, SHD, FDR, and TPR. FCI achieved the best F1 score of 0.38 and lowest structural Hamming distance of 20, with true positive rate of 0.29. PC and GES had higher false discovery rates of 0.75 and 0.81 respectively and lower true positive rates of 0.14 each. FCI recovered more true edges with fewer false positives. We then applied the same fairness decomposition to the heart failure clinical records graphs. Table 5: Causal discovery utility on heart failure clinical records dataset. Algorithm F1 SHD FDR TPR FCI 0.38 20 0.45 0.29 PC (Fisher-Z) 0.18 24 0.75 0.14 GES (BIC) 0.16 25 0.81 0.14 Composite Causal Fairness. To decompose sex-based disparity on death event into direct, indirect, and spurious components, we applied the CFA decomposition on each HFCR graph. Table 6 reports the results. TV was consistent across graphs at approximately ā-0.4%, reflecting data-driven disparity. The ground truth decomposed TV into direct effect of ā-5.1%, indirect effect of 0.06%, and spurious effect of ā-4.8%. PC collapsed to direct effect only. GES and FCI recovered indirect and spurious components, with FCI showing the largest spurious contribution of ā-7.0%. We next examined variable-level contributions to Ctf-SE and Ctf-IE. Table 6: Composite causal fairness for heart failure clinical records: TV and decomposition by graph, percent. Graph TV Ctf-DE Ctf-IE Ctf-SE Ground truth ā-0.42 ā-5.11 0.06 ā-4.75 PC ā-0.42 ā-0.42 0.00 0.00 GES ā-0.42 ā-4.90 2.00 ā-6.48 FCI ā-0.42 ā-4.91 2.49 ā-6.97 Path-Specific Effects. To identify which variables drove the indirect and spurious effects, we decomposed Ctf-SE and Ctf-IE by variable for the heart failure clinical records ground truth graph. Table 7 reports the contributions. Ejection fraction contributed most to Ctf-IE at 3.37%. Age contributed 1.04%, platelets 0.39%, and serum sodium 0.38% to Ctf-SE. Table 7: Individual contributions to Ctf-SE and Ctf-IE for heart failure clinical records ground truth, percent. Effect Variable Contribution (%) Ctf-SE age ++1.04 Ctf-SE platelets ++0.39 Ctf-SE serum sodium ++0.38 Ctf-IE ejection fraction ++3.37 Ctf-IE cpk ++0.85 Ctf-IE high blood pressure ++0.29 Causal FairnessāUtility Ratio (HFCR). To quantify the fairnessāutility trade-off per path for the heart failure clinical records dataset, we computed CFUR for each graph. Table 8 reports the results. The ground truth showed positive CFUR for indirect effect at 10.0 ± 27.5 and spurious effect at 0.13 ± 0.26. Direct effect had negative CFUR of ā-3.8 ± 3.2. Discovered graphs yielded different profiles. FCI had the most negative direct-effect CFUR at ā-16.9 ± 40.1. GES preserved a positive indirect-effect CFUR of 6.3 ± 8.9. Table 8: CFUR by path for heart failure clinical records, mean ± SD. Graph CFUR DE CFUR IE CFUR SE Ground truth ā-3.8 ± 3.2 ++10.0 ± 27.5 ++0.13 ± 0.26 PC ā-4.0 ± 8.8 ā-0.9 ± 3.1 ++0.09 ± 0.30 GES ā-10.6 ± 23.1 ++6.3 ± 8.9 ++0.07 ± 0.25 FCI ā-16.9 ± 40.1 ++0.9 ± 1.2 ++0.22 ± 0.16 4 Discussion Prior work on causal discovery for fairness [4, 23] has focused on non-clinical datasets and composite fairness scores. Binkyte et al. [4] evaluate path-specific effects but assume linear relationships and do not fully incorporate spurious effects from confounders. Zanna et al. [23] use synthetic data with known structure but limit evaluation to composite fairness scores rather than variable-level path-specific decomposition. Neither addresses the utilityāfairness trade-off per path or provides benchmarks for real-world clinical data where expert-defined graphs can be established through domain collaboration. Our work extends causal fairness benchmarks [14, 16] by pairing structural discovery with path-specific decomposition on both synthetic and real-world clinical data. We established that graph choice drives fairness decomposition. Discovered graphs can collapse indirect and spurious components (e.g., PC on both datasets) or recover them with varying fidelity (GES, FCI). This graph-dependence has implications for how clinicians and policymakers interpret fairness analyses. Our results agree with the CFA framework in that total variation (TV) is data-driven and consistent across graphs, while the decomposition into direct, indirect, and spurious components depends on structure. Unlike the Alzheimerās disease dataset, where PC achieved best structural recovery, the heart failure clinical records dataset favored FCI. This finding is consistent with FCIās design for latent confounders and selection bias in observational clinical data. The CFUR profiles differed by graph and dataset. On the Alzheimerās disease dataset, blocking direct effects yielded the largest fairness gain per accuracy cost. On the heart failure clinical records dataset, indirect and spurious effects showed positive CFUR under the ground truth, suggesting that interventions on mediators and confounders may be viable in some settings. Our study has limitations. The Alzheimerās disease ground truth is expert-defined. The heart failure clinical records ground truth relies on expert consensus augmented by literature, and associations may not imply causation. Observational data may violate faithfulness and Markov equivalence. The HFCR sample size of 299 yields wide confidence intervals for path-specific effects. Doubly directed edges in expert graphs imply cyclic relationships that conflict with standard DAG assumptions. Sensitivity analysis is warranted. These limitations leave open questions about how to validate ground truth graphs and how to generalize path-specific fairness to larger, more diverse cohorts. Future work will expand the suite of causal discovery algorithms, datasets, and evaluation metrics. We will develop metrics for multiple protected attributes and extend the framework to broader health data domains. The pipeline we present can be applied to other clinical datasets where expert-defined graphs are available, enabling graph-aware fairness evaluation alongside structural discovery. 5 Conclusions We presented a pipeline for evaluating causal discovery algorithms on utility and path-specific fairness in healthcare. On synthetic Alzheimerās disease data, PC achieved the best structural recovery. On real-world heart failure clinical records data, FCI outperformed PC and GES. Composite fairness decomposition varied by graph and dataset. Discovered graphs often collapsed or distorted indirect and spurious components. CFUR quantified the trade-off between fairness gain and accuracy loss per path. Our results highlight the need for graph-aware fairness evaluation and fine-grained path-specific analysis when deploying causal discovery in clinical applications. References [1] A. Abdulaal, N. Montana-Brown, T. He, A. Ijishakin, I. Drobnjak, D. C. Castro, D. C. Alexander, et al. (2023) Causal modelling agents: causal graph discovery through synergising metadata-and data-driven reasoning. In The Twelfth International Conference on Learning Representations, Cited by: Figure 2, Figure 2, §2.2, §2.2, §2.3, §2.4. [2] A. S. Ali, B. A. Rybicki, M. Alam, N. Wulbrecht, K. Richer-Cornish, F. Khaja, H. N. Sabbah, and S. Goldstein (1999) Clinical predictors of heart failure in patients with first acute myocardial infarction. American heart journal 138 (6), p. 1133ā1139. Cited by: §2.3. [3] K. Bello, B. Aragam, and P. Ravikumar (2022) DAGMA: Learning DAGs via M-matrices and a Log-Determinant Acyclicity Characterization. In Advances in Neural Information Processing Systems, Cited by: §2.4. [4] R. BinkytÄ, K. Makhlouf, C. Pinzón, S. Zhioua, and C. Palamidessi (2023) Causal discovery for fairness. In Workshop on Algorithmic Fairness through the Lens of Causality and Privacy, p. 7ā22. Cited by: §4. [5] P. Brouillard, C. Squires, J. Wahl, K. P. Kording, K. Sachs, A. Drouin, and D. Sridhar (2025-06) The Landscape of Causal Discovery Data: Grounding Causal Discovery in Real-World Applications. arXiv. Note: arXiv:2412.01953 [cs] External Links: Link, Document Cited by: §1. [6] S. Chiappa (2019) Path-specific counterfactual fairness. In Proceedings of the AAAI conference on artificial intelligence, Vol. 33, p. 7801ā7808. Cited by: §1. [7] D. Chicco and G. Jurman (2020) Machine learning can predict survival of patients with heart failure from serum creatinine and ejection fraction alone. BMC medical informatics and decision making 20 (1), p. 16. Cited by: Figure 3, Figure 3, §2.2. [8] D. M. Chickering (2002) Optimal structure identification with greedy search. Journal of machine learning research 3 (Nov), p. 507ā554. Cited by: §2.4. [9] D. Colombo, M. H. Maathuis, et al. (2014) Order-independent constraint-based causal structure learning.. J. Mach. Learn. Res. 15 (1), p. 3741ā3782. Cited by: §1, §2.4. [10] N. Ding, A. M. Shah, M. J. Blaha, P. P. Chang, W. D. Rosamond, and K. Matsushita (2022) Cigarette smoking, cessation, and risk of heart failure with preserved and reduced ejection fraction. Journal of the American College of Cardiology 79 (23), p. 2298ā2305. Cited by: §2.3. [11] J. A. Ezekowitz, C. Van Walraven, F. A. McAlister, P. W. Armstrong, and P. Kaul (2005) Impact of specialist follow-up in outpatients with congestive heart failure. Cmaj 172 (2), p. 189ā194. Cited by: §2.3. [12] K. Makhlouf, S. Zhioua, and C. Palamidessi (2024) When causality meets fairness: a survey. Journal of Logical and Algebraic Methods in Programming 141, p. 101000. Cited by: §1. [13] S. P. Murphy, N. E. Ibrahim, and J. L. Januzzi Jr (2020) Heart failure with reduced ejection fraction: a review. Jama 324 (5), p. 488ā504. Cited by: §2.3. [14] D. Plecko and E. Bareinboim (2024) Causal fairness analysis: a causal toolkit for fair machine learning. Foundations and Trends in Machine Learning 17 (3), p. 304ā589. Cited by: §1, §2.1, §2.5, §2.6, §3.1, §4. [15] D. Plecko and E. Bareinboim (2024) Reconciling predictive and statistical parity: a causal approach. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 14625ā14632. Cited by: §1. [16] D. Plecko and E. Bareinboim (2025) Fairness-accuracy trade-offs: a causal perspective. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 26344ā26353. Cited by: §1, §2.1, §2.5, §2.6, §3.1, §4. [17] M. Schrƶder, D. Frauen, and S. Feuerriegel (2023) Causal fairness under unobserved confounding: a neural sensitivity framework. arXiv preprint arXiv:2311.18460. Cited by: §2.5. [18] D. S. Silverberg, D. Wexler, A. Iaina, S. Steinbruch, Y. Wollman, and D. Schwartz (2006) Anemia, chronic renal disease and congestive heart failureāthe cardio renal anemia syndrome: the need for cooperation between cardiologists and nephrologists. International urology and nephrology 38 (2), p. 295ā310. Cited by: §2.3. [19] P. Spirtes, C. N. Glymour, and R. Scheines (2000) Causation, prediction, and search. MIT press. Cited by: §1, §2.4. [20] A. Srivastava, L. Nagalapatti, G. Jajoo, A. Vashishtha, P. Krishnamurthy, and A. Sharma (2025) Realizing llmsā causal potential requires science-grounded, novel benchmarks. arXiv preprint arXiv:2510.16530. Cited by: §2.4. [21] W. W. Tang, M. A. Bakitas, X. S. Cheng, J. C. Fang, S. E. Fedson, A. G. Fiedler, P. Martens, W. I. Mccallum, M. O. Ogunniyi, J. Rangaswami, et al. (2024) Evaluation and management of kidney dysfunction in advanced heart failure: a scientific statement from the american heart association. Circulation 150 (16), p. e280āe295. Cited by: §2.3. [22] Y. Yu, J. Chen, T. Gao, and M. Yu (2019) DAG-gnn: dag structure learning with graph neural networks. In Proceedings of the 36th International Conference on Machine Learning, Cited by: §2.4. [23] K. Zanna and A. Sano (2025) Fairness-driven llm-based causal discovery with active learning and dynamic scoring. arXiv preprint arXiv:2503.17569. Cited by: §2.2, §4. [24] K. Zhang, S. Zhu, M. Kalander, I. Ng, J. Ye, Z. Chen, and L. Pan (2021) GCastle: a python toolbox for causal discovery. External Links: 2111.15155 Cited by: §2.4. [25] W. Zhao, P. T. Katzmarzyk, R. Horswell, W. Li, Y. Wang, J. Johnson, S. B. Heymsfield, W. T. Cefalu, D. H. Ryan, and G. Hu (2014) Blood pressure and heart failure risk among diabetic patients. International journal of cardiology 176 (1), p. 125ā132. Cited by: §2.3. [26] X. Zheng, B. Aragam, P. Ravikumar, and E. P. Xing (2018) DAGs with NO TEARS: Continuous Optimization for Structure Learning. In Advances in Neural Information Processing Systems, Cited by: §2.4, §2.4. [27] X. Zheng, C. Dan, B. Aragam, P. Ravikumar, and E. P. Xing (2020) Learning sparse nonparametric DAGs. In International Conference on Artificial Intelligence and Statistics, Cited by: §2.4. [28] Y. Zheng, B. Huang, W. Chen, J. Ramsey, M. Gong, R. Cai, S. Shimizu, P. Spirtes, and K. Zhang (2024) Causal-learn: causal discovery in python. Journal of Machine Learning Research 25 (60), p. 1ā8. Cited by: §2.2.