Paper deep dive
A feature-stable and explainable machine learning framework for trustworthy decision-making under incomplete clinical data
Justyna Andrys-Olek, Paulina Tworek, Luca Gherardini, Mark W. Ruddock, Mary Jo Kurt, Peter Fitzgerald, Jose Sousa
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 7/21/2026, 12:00:44 AM
Summary
The paper introduces CACTUS (Comprehensive Abstraction and Classification Tool for Uncovering Structures), an explainable machine learning framework designed to address robustness, interpretability, and feature instability in incomplete clinical datasets. Using a haematuria cohort of 568 patients evaluated for bladder cancer, the authors benchmark CACTUS against Random Forests and Gradient Boosting methods under controlled missing data scenarios (10%, 20%, 30%). Results demonstrate that CACTUS achieves competitive or superior predictive performance (balanced accuracy and recall) while maintaining significantly higher feature stability compared to classical ML models. The study highlights that feature stability is a critical, underutilized metric for assessing the trustworthiness of AI in biomedical decision-making, particularly for sex-stratified analysis.
Entities (9)
Relation Signals (5)
CACTUS ā appliedto ā Haematuria Biomarker Cohort
confidence 95% Ā· Using a real-world haematuria cohort comprising 568 patients evaluated for bladder cancer, we benchmark CACTUS...
Feature Stability ā assesses ā Trustworthiness
confidence 92% Ā· feature stability provides information complementary to conventional performance metrics and is essential for assessing the trustworthiness of machine learning models applied to biomedical data.
Bladder Cancer ā classifiedby ā CACTUS
confidence 90% Ā· The goal was to measure how well CACTUS classifies patients based on the full dataset and identify biomarkers specific to patients with BC
CACTUS ā outperforms ā Random Forest
confidence 88% Ā· CACTUS achieves competitive or superior predictive performance while maintaining markedly higher stability of top-ranked features... For the male and female data subsets, it shows clear superiority over classical ML methods
CACTUS ā outperforms ā Gradient Boosting
confidence 88% Ā· CACTUS achieves competitive or superior predictive performance... CACTUS produced better results regarding balanced accuracy and recall when compared with other classical ML models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Machine learning models are increasingly applied to biomedical data, yet their adoption in high stakes domains remains limited by poor robustness, limited interpretability, and instability of learned features under realistic data perturbations, such as missingness. In particular, models that achieve high predictive performance may still fail to inspire trust if their key features fluctuate when data completeness changes, undermining reproducibility and downstream decision-making. Here, we present CACTUS (Comprehensive Abstraction and Classification Tool for Uncovering Structures), an explainable machine learning framework explicitly designed to address these challenges in small, heterogeneous, and incomplete clinical datasets. CACTUS integrates feature abstraction, interpretable classification, and systematic feature stability analysis to quantify how consistently informative features are preserved as data quality degrades. Using a real-world haematuria cohort comprising 568 patients evaluated for bladder cancer, we benchmark CACTUS against widely used machine learning approaches, including random forests and gradient boosting methods, under controlled levels of randomly introduced missing data. We demonstrate that CACTUS achieves competitive or superior predictive performance while maintaining markedly higher stability of top-ranked features as missingness increases, including in sex-stratified analyses. Our results show that feature stability provides information complementary to conventional performance metrics and is essential for assessing the trustworthiness of machine learning models applied to biomedical data. By explicitly quantifying robustness to missing data and prioritising interpretable, stable features, CACTUS offers a generalizable framework for trustworthy data-driven decision support.
Tags
Links
- Source: https://arxiv.org/abs/2602.17364v1
- Canonical: https://arxiv.org/abs/2602.17364v1
Trouble viewing inline? Open PDF directly ā
Full Text
70,537 characters extracted from source content.
Expand or collapse full text
A feature-stable and explainable machine learning framework for trustworthy1 decision-making under incomplete clinical data2 Justyna Andrys-Olek 1 , Paulina Tworek 1 , Luca Gherardini 1 , Mark W. Ruddock 2 , Mary Jo Kurt 2 ,3 Peter Fitzgerald 2 , and Jose Sousa 1, 2 4 1 Computational Intelligence, Sano Centre for Computational Personalised Medicine, Krakow,5 Poland6 2 Clinical Studies Group, Randox Laboratories Ltd., Co., Antrim, United Kingdom7 3 Coimbra University, Multidisciplinary Institute of Ageing, MIA - Portugal, Coimbra, Portugal8 * Correspondence: j.andrys-olek@sanoscience.org9 ** Correspondence: p.tworek@sanoscience.org10 *** Correspondence: j.sousa@sanoscience.org11 SUMMARY12 Machine learning models are increasingly applied to biomedical data, yet their adoption in high-13 stakes domains remains limited by poor robustness, limited interpretability, and instability of14 learned features under realistic data perturbations, such as missingness. In particular, models15 that achieve high predictive performance may still fail to inspire trust if their key features fluc-16 tuate when data completeness changes, undermining reproducibility and downstream decision-17 making. Here, we present CACTUS (Comprehensive Abstraction and Classification Tool for18 Uncovering Structures), an explainable machine learning framework explicitly designed to ad-19 dress these challenges in small, heterogeneous, and incomplete clinical datasets. CACTUS20 integrates feature abstraction, interpretable classification, and systematic feature stability anal-21 ysis to quantify how consistently informative features are preserved as data quality degrades.22 Using a real-world haematuria cohort comprising 568 patients evaluated for bladder cancer,23 we benchmark CACTUS against widely used machine learning approaches, including random24 forests and gradient boosting methods, under controlled levels of randomly introduced missing25 data. We demonstrate that CACTUS achieves competitive or superior predictive performance26 while maintaining markedly higher stability of top-ranked features as missingness increases, in-27 cluding in sex-stratified analyses. Our results show that feature stability provides information28 complementary to conventional performance metrics and is essential for assessing the trust-29 worthiness of machine learning models applied to biomedical data. By explicitly quantifying30 robustness to missing data and prioritising interpretable, stable features, CACTUS offers a gen-31 eralizable framework for trustworthy data-driven decision support. Although demonstrated on32 a bladder cancer cohort, the proposed approach is broadly applicable to other domains where33 incomplete, high-dimensional data and model transparency are critical. These findings high-34 light feature stability as a critical yet underutilised dimension of model evaluation in data-driven35 biomedical research.36 KEYWORDS37 bladder cancer, biomarkers, decision-making, support system, machine learning, explainable AI,38 features stability39 1 arXiv:2602.17364v1 [cs.LG] 19 Feb 2026 INTRODUCTION40 Medical decision-making is a complex process that integrates medical knowledge and experi-41 ence to formulate a diagnosis or prepare a treatment plan, based on patient health data obtained42 from tests, medical examinations, and interviews, with the goal of maximising clinical benefit 1 .43 Despite rapid advances in machine learning for biomedical applications, model evaluation re-44 mains dominated by predictive performance metrics, often overlooking whether learned repre-45 sentations remain stable under realistic data perturbations. In practice, biomedical datasets are46 frequently incomplete, heterogeneous, and subject to acquisition biases, raising concerns about47 the robustness and reproducibility of model-derived insights. During diagnosis, most doctors48 use cognitive shortcuts to formulate a hypothesis by matching an individualās clinical features to49 typical symptoms of a condition and then confirming it with a series of diagnostic tests. A deci-50 sion made by clinicians must be grounded in the best available evidence, in accordance with an51 evidence-based medicine approach, which requires application of population-based data to the52 care of an individual patient (...) 2 . It is affected by biases and uncertainties inherent to medical53 reality, but a clinician must assess the situation and consider all available facts to reduce un-54 certainty 3 1 . Despite extensive training and years of experience, some of these challenges may55 persist, creating opportunities for frameworks that support decision-making to guide or improve56 the final evaluation. Hence, the search for reliable technological solutions that can serve as a57 decision aid is well justified.58 Applying computer logic in healthcare as a support in human decision-making is not a new59 concept, as the first serious attempts date back to the 1970s, when the first digital diagnostic60 assistance based on a decision tree (INTERNIST-1) was deployed in Massachusetts General61 Hospital. It was designed by the computer scientist Harry Pople at the University of Pittsburgh to62 encapsulate the expertise of internist Jack D. Myers and to provide insights for new diagnoses 4 .63 The performance of the tool was quickly found to be unsatisfactory due to multiple issues, among64 which were the attribution of findings to improper causes and the inability to explain its (model)65 thinking 4 . Years later, many new artificial intelligence (AI) tools remain unreliable or too complex66 to be trusted for use in medical settings 5,6 . One of the main issues is the black-box problem,67 where the internal decision-making processes of AI systems are not transparent or interpretable68 to users 5 . For example, the recent boom in Deep Neural Network (DNN) algorithms in healthcare69 has yielded new insights into how such technology can help. However, these models have sig-70 nificant drawbacks: they are highly complex, require substantial computational resources, and71 demand large datasets for training, which limits their practicality in real-world settings, especially72 in sensitive areas such as medicine 7 . DNNs comprise multiple layers and employ millions of73 parameters, such as filters and constraints, making them inherently complex and difficult to in-74 terpret 8 . Even the simplest deep models process data in non-trivial ways, making their decision-75 making process difficult to understand. As a result, there is a growing trend in the field to prioritise76 interpretable and explainable machine learning (ML) tools, mostly ones that are intended to be77 used in high-risk fields like health, banking or criminal justice, where understanding the modelās78 reasoning is crucial for trust and accountability 9 .79 One key factor in the adoption of AI-based decision-making support assistants by industry and80 individuals is trustworthiness. In medicine, where the tools clinicians use daily directly affect81 patientsā health and well-being, establishing trust is even more essential. Above all, the model82 must be explainable and provide clear reasoning behind the decision-making process to both83 doctors and patients 8 10 . A trustworthy AI model must address biases in patient data to ensure84 fairness while maintaining ethical standards, ensuring patient privacy, demonstrating clinical ef-85 fectiveness, and delivering reproducible results 11 .86 Another obstacle to digital transformation in healthcare is that medical datasets are difficult to87 2 use as input for ML models because they often contain substantial amounts of missing and noisy88 data, making it very challenging to collect complete information for every case 12 . There are sev-89 eral reasons for missing data, including incomplete medical records, incomplete surveys, and90 data loss due to lack of patient follow-up. They can lead to bias, loss of valuable information,91 reduced statistical power, and generalisability of the findings 13 . That is why any solution de-92 signed to analyse medical records should be robust to missing data and capable of handling it93 effectively.94 From a data science perspective, instability in feature importance with varying data complete-95 ness undermines both interpretability and trust, particularly in high-stakes domains. Yet, system-96 atic evaluation of feature stability remains rare in applied machine learning studies. Addressing97 this gap requires frameworks that explicitly quantify robustness to missingness while remaining98 interpretable and data-efficient. In this paper, we describe how a new ML framework, CACTUS99 (Comprehensive Abstraction and Classification Tool for Uncovering Structures), designed to han-100 dle unbalanced, incomplete, and biased datasets, addresses the aforementioned challenges,101 positioning it as a trustworthy AI-based framework 14 . Additionally, we show that CACTUS main-102 tains feature stability, even under conditions of randomly introduced missing data (MCAR). This103 resilience ensures that feature importance remains consistent, reinforcing trust and enhancing104 the explainability of predictions, which are critical requirements in healthcare and other high-105 stakes applications 15 .106 The Haematuria Biomarker (HaBio) cohort data were used to classify patients into two groups:107 those with bladder cancer (BC) and those without bladder cancer (non-BC) 16 . The goal was to108 measure how well CACTUS classifies patients based on the full dataset and identify biomark-109 ers specific to patients with BC for the whole population (both genders) and depending on sex110 (females, males). It was also investigated how introducing an increasing percentage of missing111 values influences classification and features (biomarkers) stability, as entering missing values112 simulates a real-world medical scenario, when not all patients underwent the same diagnostic113 tests or can arise from the loss of a sample batch. Different percentages of missing values (10%,114 20%, and 30%) were randomly introduced into each dataset, assuming that missingness affects115 any feature with the same probability and does not follow any specific pattern. In contrast to116 accuracy-centric evaluation, we argue that stability-aware analysis is essential for trustworthy117 pattern discovery in real-world data.118 RESULTS AND DISCUSSION119 Feature Stability Metrics as a Foundation for Trustworthy AI in Classifying120 BC and non-BC Patients121 Beyond predictive performance, we evaluate model behaviour through the lens of feature sta-122 bility, assessing whether the most informative features remain consistent as data completeness123 decreases.124 Features stability was previously introduced by K. CapaÅa, P. Tworek and J. Sousa ( 15 ), to125 assess how stable the features are when ranked by different ML models. In general, it is desir-126 able for trustworthy AI classification frameorks for medical diagnostics that the most important127 features for classification remain stable across datasets as the number of randomly introduced128 missing values increases.129 The stability of a given feature can be measured either by directly inspecting its significance130 value across datasets or by conducting a broader analysis using comparison-based metrics.131 Here, we considered the average relative change in feature importance for the 10 most important132 3 features for classification across all used datasets (differing in the number of missing values) for133 each tested model. The equation, firstly introduced by CapaÅa et al. ( 15 ), is reported in Equation134 7 in the supplementary materials.135 The smaller the average relative change, the more stable the feature produced by a given136 model, as it indicates that feature importance remains mostly unaffected with a gradual transi-137 tion from a complete dataset to its version with 30% of values removed. The standard deviation138 is used to quantify how widely feature importance varies across datasets with varying amounts139 of missing data. Hence, low average change and low standard deviation are the ideal character-140 istics of models that are resilient to missing data.141 4 Figure 1: Stability of features across ML methods. Average relative change in feature importance calculated for the 10 most important features for classification obtained by each method for: the total dataset (top), the male subjects in the dataset (middle), and the female subjects in the dataset (bottom). 5 Figure 1 shows that CACTUS achieves low average change in feature importance values142 in general. For the male and female data subsets, it shows clear superiority over classical ML143 methods, while for the entire dataset (total), only Random Forest performs better. Furthermore,144 CACTUS exhibits one of the smallest standard deviation errors, confirming the consistency of145 the results.146 Another Way of Feature Stability Analysis in Classifying BC and non-BC147 Patients: The Overlap of Top 10 Ranked Features148 Feature overlap among top-ranked variables provides an intuitive, practitioner-oriented measure149 of robustness, complementing numerical stability metrics by revealing whether models rely on150 consistent signals under data perturbations. The percentage of overlapping features among the151 10 most important ones for classification across datasets with different proportions of missing152 values is another measure of feature stability. It is desirable that the top n features are stable.153 Observing the same features consistently appearing in the top 10 ranks for each dataset variant154 (complete and with increasing percentage of removed values up to 30%) can provide confidence155 to healthcare professionals that these features are indeed important in the classification, and156 consequently in the diagnostic process. This aspect is particularly valuable from the medical157 point of view, as a clinician might want to focus primarily on the n top-ranked features without158 inspecting the stability values in detail. That way, an ML model can provide direct support in a159 decision-making process. Although individual features may not always remain stable (Fig. 1),160 CACTUS gives more consistent and repetitive results across all three datasets (total, males,161 and females) when considered as a group, as illustrated in Figure 2. Classical methods cannot162 produce such stability for all three subsets. RF and CatBoost are the second best after CACTUS163 only for the total population and the male datasets, but CatBoost fails for females, where the164 second best is RF and LGBM. Features for CACTUS and LGBM are additionally shown in tables165 1-6, while results for the rest of the tested methods are available in Supplementary Information166 (Tab. S1-S12).167 6 Figure 2: Overlapping features. Graphs plotted for each subset from the total population, males and females, presenting per- centage of overlapping features from the top 10 most important features for classification with increasing number of missing values in the dataset. 7 Table 1: The most important 10 features for classification obtained through CACTUS analysis for the combined dataset (total popula- tion). The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1HAEMATURIA0.304CLUSTERIN0.317CLUSTERIN0.253CLUSTERIN0.220 2CLUSTERIN0.297NSE0.295FAS0.219CYSTATIN-B0.220 3NSE0.279HAEMATURIA0.257TAR EXPOS.0.215NSE0.210 4TAR EXPOS.0.252MICROALBUMIN0.248NSE0.214VEGF0.208 5IL-1α0.242(S)PAI-1/TPA0.241(S)PAI-1/TPA0.213HAEMATURIA0.206 6MICROALBUMIN0.242CYSTATIN-B0.230MICROALBUMIN0.199IL-1α0.197 7(S)PAI-1/TPA0.239IL-1α0.229HAEMATURIA0.186FAS0.190 8FAS0.230BTA0.221TPA0.186(S)PAI-1/TPA0.190 9BTA0.229FAS0.219CYSTATIN-B0.185TAR EXPOS.0.187 10CYSTATIN-B0.229TPA0.218IL-80.172BTA0.180 Table 2: The most important 10 features for classification obtained through LGBM analysis for the combined dataset (total population). The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1NSE390.4NSE309.3NSE251.5PSA/TPSA214.9 2PSA/TPSA192.0PSA/TPSA260.5TAR EXPOS.155.2YRS OF SMOKING172.9 3YRS OF SMOKING178.7(S)PAI-1/TPA169.8CLUSTERIN140.8(S)PAI-1/TPA171.4 4HAEMATURIA169.2HAEMATURIA163.4HAEMATURIA115.7NSE152.6 5(S)EGF123.1(S)EGF100.4PERK108.3MIDKINE129.7 6(S)PAI-1/TPA100.5IL-1899.7MMP9105.8HAEMATURIA121.7 7RECURRENT UTI92.2PROLACTIN90.6PSA/TPSA101.6IL-1α117.6 8EGF79.6ACR88.1(S)EGF98.6(S)EGF104.1 9CLUSTERIN77.9CLUSTERIN87.6YRS OF SMOKING88.7CYSTATIN-B76.6 10D-DIMER76.1FAS77.5(S)PAI-1/TPA75.6ACR67.4 Table 3: The most important 10 features for classification obtained through CACTUS analysis for the males dataset. The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1NSE0.298NSE0.281FAS0.224(S)PSA/TPSA0.207 2TAR EXPOS.0.281VEGF0.260NSE0.223TAR EXPOS.0.202 3CLUSTERIN0.273CLUSTERIN0.245CLUSTERIN0.216NSE0.198 4FAS0.266FAS0.244PROLACTIN0.213FAS0.185 5PROLACTIN0.262CYSTATIN-B0.231CYSTATIN-B0.212PROLACTIN0.181 6VEGF0.258TAR EXPOS.0.227IL-1α0.202VEGF0.174 7CYSTATIN-B0.253PROLACTIN0.223GRO0.194MIDKINE0.171 8MIDKINE0.237(S)PSA/TPSA0.220TAR EXPOS.0.194CYSTATIN-B0.170 9(S)PSA/TPSA0.227TPA0.210VEGF0.188IL-70.166 10TPA0.223IL-1α0.1993CXCL160.186TPA0.165 8 Table 4: The most important 10 features for classification obtained through LGBM analysis for the male dataset. The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1NSE318.2NSE228.3NSE166.7NSE207.9 2YRS OF SMOKING172.5PSA/TPSA151.2PSA/TPSA127.4PSA/TPSA136.2 3PROLACTIN127.3TNF-α115.2YRS OF SMOKING102.5PROLACTIN119.9 4PSA/TPSA123.3PROLACTIN114.7CLUSTERIN91.8TAR EXPOS.111.6 5EGF98.3(S)PAI-1/TPA107.4IL-888.9TGF-β 189.8 6NGAL96.6HAEMATURIA103.3TGF-β 180.9FAS84.1 7(S)PAI-1/TPA85.8YRS OF SMOKING99.8PROLACTIN78.9HAEMATURIA82.0 8HAEMATURIA81.8(S)CRP88.5(S)PAI-1/TPA73.4YRS OF SMOKING76.8 9(S)EGF68.4TGF-β 182.0HAEMATURIA70.9(S)EGF75.1 10LASP-167.1CRP73.3(S)VEGF64.4NGAL67.9 Table 5: The most important 10 features for classification obtained through CACTUS analysis for the females dataset. The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1HAEMATURIA0.555HAEMATURIA0.497HAEMATURIA0.512MICROALBUMIN0.364 2MICROALBUMIN0.437MICROALBUMIN0.453CLUSTERIN0.429HAEMATURIA0.356 3IL-1α0.418BTA0.417IL-1α0.422IL-80.328 4CLUSTERIN0.418CLUSTERIN0.414MICROALBUMIN0.421ACR0.310 5BTA0.404IL-1α0.408CXCL160.401BTA0.304 6MCP-10.401IL-80.406BTA0.392D-DIMER0.300 7IL-80.399MCP-10.381MCP-10.361IL-70.294 8VEGF0.393PROTEIN0.381ACR0.351IL-1α0.287 9IL-130.387NSE0.376IL-80.332IL-130.285 10(S)PAI-1/TPA0.385(S)PAI-1/TPA0.365D-DIMER0.329CLUSTERIN0.283 Table 6: The most important 10 features for classification obtained through LGBM analysis for the females dataset. The prefix (s) indicates that the biomarker is measured in blood rather than in urine (default). RANKCOMPLETERANKCOMPLETE +10%RANKCOMPLETE +20%RANKCOMPLETE +30%RANK NUMBERVALUEMISSING VALUESVALUEMISSING VALUESVALUEMISSING VALUESVALUE 1HAEMATURIA122.8HAEMATURIA85.9HAEMATURIA159.9HAEMATURIA80.0 2MICROALBUMIN80.5MICROALBUMIN75.3CLUSTERIN92.3MICROALBUMIN65.7 3IL-1348.5IL-1α67.7IL-1α61.9MCP-143.4 4(S)PAI-1/TPA37.9YRS OF SMOKING52.3GRO41.1IL-842.4 5CLUSTERIN35.0BTA51.6MICROALBUMIN37.9CLUSTERIN21.3 6BTA30.2CLUSTERIN41.8YRS OF SMOKING35.9YRS OF SMOKING20.8 7YRS OF SMOKING30.1PROGRANULIN35.6(S)CYSTATIN-C29.8BTA17.2 8MCP-129.1FABP-A27.0ACR24.1IL-12P7014.8 9RECURRENT UTI25.1NSE25.9CXCL1623.2CXCL1614.4 10TRIGLICERIDES23.0IL-1325.5(S)PAI-1/TPA20.0CYSTATIN-C13.0 Performance of ML Methods in BC and non-BC Patients Classification168 The gold standard for ML models evaluation in comparison to the metrics obtained by them.169 Specific metrics should be selected depending on the problem being solved. In the BC and non-170 BC patients classification task, we primarily focus on two metrics: balanced accuracy and recall171 (sensitivity). Balanced accuracy is a measure of how frequently an ML model makes correct172 predictions when dealing with imbalanced datasets, while recall is the ability of the model to173 spot positive cases, here BC patients. Recall (sensitivity) is particularly important in medical174 applications, where we aim to identify as many positive cases as possible, i.e., people who are175 sick or at risk of becoming sick (here, BC patients). Additional metrics: F1 score, accuracy and176 precision are calculated and presented in Supplementary Materials (Figures S2-S4).177 CACTUS produced better results regarding balanced accuracy and recall when compared178 with other classical ML models (Fig. 3-4). CACTUS achieved higher balanced accuracy, espe-179 9 cially for females and males subsets (Fig. 3, bottom and middle), whereas for recall, it is superior180 for all three datasets and all levels of introduced missing values (Fig. 4). Because the model181 performance is just one of the metrics available for ML model evaluation, other information, like182 feature stability, should be included in the analysis. The high stability of the features, together183 with high performance measured with metrics, shows that CACTUS is effective for the analysis184 of small and incomplete medical datasets.185 186 10 Figure 3: Balanced Accuracy (BA). BA calculated for each subset (total, males and females) with increasing number of missing values in the dataset. 11 Figure 4: Recall (sensitivity). Sensitivity calculated for each subset (total, males and females) with increasing number of miss- ing values in the dataset. 12 The 10 most important biomarkers for classifying patients as BC or non-187 BC patients for both sexes identified by CACTUS188 As proven above, CACTUS achieves overall better results over classical ML models in the case189 of differentiating between BC and non-BC patients in all three datasets, with and without strati-190 fication by sex. A sex-based comparison of features is conducted, as performance metrics, es-191 pecially recall, show significant differences between males and females, suggesting sex-specific192 biomarkers in bladder cancer development. Descriptions of 10 of these biomarkers identified by193 CACTUS, together with heatmaps illustrating their discriminatory power between BC and non-BC194 cases depending on sex, are presented below. The analysis of the 10 most important features195 performed for the total population is provided in the supplementary material (Fig. S1).196 Figure 5: Heatmap for males subset. A heatmap showing the 10 most important features for BC/non-BC case classification in the male subset, with thresholds and significance values used to produce the ranks. Note: the thresholds shown on the graph represent the best value that separates the two classes (BC and non-BC cases) and do not correspond to thresholds set by medical institutions for diagnosis. 1. NSE (urine): neuron-specific enolase is a protein specific to neurons and peripheral en-197 docrine cells. It serves as a biomarker for small cell lung carcinoma, as increased body198 13 fluid levels of NSE may occur with malignant proliferation and thus can be of value in di-199 agnosis, staging and treatment of related neuroendocrine tumours (NETs) 17 . However, no200 link between NSE and bladder cancer has been established to this date.201 2. Tar exposure: it refers to the burnt matter left after smoking a cigarette. It forms a sticky202 layer on the respiratory mucosa, impairing its functions and leading to many lung dis-203 eases 18 . Various types of cigarettes leave different amounts of tar in the lungs, depending204 on how they are manufactured 18 . The link between tar exposure and upper airway cancers205 is well known. In general, smoking is an established risk factor for bladder cancer and tar206 plays an important role in carcinogenesis 19 .207 3. Clusterin (urine): Clusterin is a protein found widely in blood and tissues. It can exist in208 different forms and have distinct functions, depending on the location and expression levels.209 Although sCLU (secreted form of CLU) is known to play crucial roles in lipid transport210 and cellular lysis and adhesion under physiological conditions, elevated levels are often211 noted in several pathologies, like Alzheimerās disease, fibrosis, cardiovascular diseases or212 cancers 20 . Increased expression of clusterin is already linked with bladder cancer, making213 it a promising candidate biomarker for disease detection and prognosis 21 .214 4. FAS/CD95 (urine): FAS cell surface death receptor is a protein that starts the molecular215 cascade leading to apoptosis upon binding with its ligand (FASL). A soluble form of CD95216 (a product of alternative mRNA splicing) is, on the other hand, an antiapoptotic factor.217 Elevated levels of sFAS are linked with many cancers, including small cell lung, ovarian,218 endometrial, adrenocortical and colorectal cancers 22,23 . Moreover, sFAS has been iden-219 tified as a potential biomarker not only for predicting outcomes in patients with bladder220 cancer but also for detecting the disease itself 24,25 .221 5. Prolactin (serum): Prolactin is a hormone produced by the pituitary gland, which reg-222 ulates testosterone production and sexual functions in males. Elevated prolactin levels223 are observed in both nodular hyperplasia and prostate cancer patients 26 , and our results,224 shown in Figure 5, indicate a negative correlation between prolactin and bladder cancer225 and a positive correlation between prolactin and non-bladder cancer patients.226 6. VEGF (urine): vascular endothelial growth factor (VEGF) is a cytokine secreted by os-227 teoblasts, macrophages, cancer cells, megakaryocytes and platelets. It plays an important228 role in angiogenesis by stimulating vascular permeability and vessel growth 27 . It has been229 proven that an increased level of VEGF in tumours is associated with more intensive cell230 proliferation and metastasis, which translates to poor prognosis in cancer patients 28,29 .231 Moreover, it has been recognised as a potential biomarker associated with bladder can-232 cer 30 .233 7. Cystatin B (urine): Cystatin B is a member of cystatin proteins which act as inhibitors234 of cathepsin proteases by binding to them, thus preventing cellular protein degradation,235 particularly from lysosomal enzymes 31 . Increased urinary levels of cystatin B have been236 observed in patients with bladder cancer and show a strong association with tumour grade237 and disease stage 32 .238 8. Midkine (urine): Midkine, also known as neurite growth-promoting factor 2, is a small pro-239 tein that takes part in angiogenesis, fibrinolysis, cell division and migration. Under phys-240 iological conditions, its activity is restricted to specific tissue types and remains generally241 low 33 . However, its overexpression has been linked with inflammatory responses and car-242 cinogenesis in at least 20 distinct types of cancer, including bladder 34,35 .243 14 9. PSA/tPSA (serum): it refers to the measurement of the ratio of free PSA (prostate-specific244 serum antigen) to total PSA (tPSA), which enables more precise assessment of prostate245 cancer risk. Low PSA/tPSA ratio indicates high risk of prostate cancer , while a high246 PSA/tPSA ratio indicates benign changes in the gland. PSA is a biomarker of prostate247 cancer, but no record of a link between PSA and bladder cancer has been found to date.248 10. tPA: (tissue plasminogen activator) facilitates the breakdown of clots. The level of free249 tPA can be an informative biomarker in breast cancer, when it forms a complex with PAI-1250 protein 36 . However, there are no studies to date to show the link between standalone free251 tPA and bladder cancer.252 15 Figure 6: Heatmap for females subset. A heatmap showing the 10 most important features for BC/non-BC case classification in the male subset, with thresholds and significance values used to produce the ranks. Note: the thresholds shown on the graph represent the best value that separates the two classes (BC and non-BC cases) and do not correspond to thresholds set by medical institutions for diagnosis. To avoid repetition, the descriptions of overlapping features with MALES heatmap (Figure 5),253 Clusterin (urine) and VEGF (urine), are omitted here.254 1. Haematuria: A presence of red blood cells (RBCs) in the urine. When RBCs are visible to255 the naked eye, the condition is known as macrohaematuria. If they are visible only under a256 microscope, the condition is called microhaematuria. Haematuria indicates that pathologi-257 cal processes may take place within the urinary tract. These can be benign in nature, e.g.258 kidney stones, bacterial infection, menstruation (females) or prostate enlargement (males),259 but may also signal the presence of malignancies like bladder cancer 37,38 .260 2. Microalbuminuria: A condition defined by the excretion of 30ā300 mg of albumin in urine261 over a 24-hour period, confirmed in at least two out of three urine samples. It is a crucial262 indicator of kidney damage, as normally, proteins are retained in the blood by the renal263 glomeruli and not secreted in the urine. Two cohort studies have established a correlation264 16 between higher urinary albumin levels and an increased risk of developing cancers (urinary265 tract, lung, haematological), independently of kidney function 39,40 .266 3. BTA (urine): The bladder tumor antigen (BTA) test is designed to detect human comple-267 ment factor-H related protein (hCFHrp). It has been found that bladder cancer cell lines can268 produce and release hCFHrp, whereas normal epithelial cells do not. This suggests that269 hCFHrp secretion is associated with malignant transformation 41,42 . However, the presence270 of haematuria in urine samples can interfere with test results, leading to false positives.271 Hence, blood presence should be considered a confounding factor in diagnostic assess-272 ments with use of a BTA test 43 .273 4. IL-1α (urine): Interleukin-1α is a pro-inflammatory protein that promotes fever and sepsis,274 typically released as a result of cell death or tissue damage 44 . It is a key factor in the patho-275 genesis of multiple inflammation-related conditions, including various types of cancers. In276 the context of HaBio cohort, it is worth noting that elevated IL-1α levels were associated277 with malignant progression of bladder cancer 44,45 .278 5. MPC-1 (urine): MCP-1 stands for monocyte chemoattractant protein 1. It is a small cy-279 tokine that plays an important role in immunoregulation and inflammation 46 . It has been280 shown that MCP-1 is associated with infections, bowel disease, diabetes, and several types281 of cancer 47,48 . Furthermore, urinary MCP-1 concentration has been proposed as a poten-282 tial prognostic indicator of bladder cancer progression 48 .283 6. IL-8 (urine): Another protein from the interleukin group. IL-8 is a chemokine that plays284 an important role in the immune response to infection through chemotaxis 49 . It also pro-285 motes cell migration, angiogenesis, and metastasis, and is known as a pro-cancerogenic286 factor across various cancer subtypes. Strong correlation between elevated IL8 levels and287 bladder cancer made it recognised as a potential biomarker for the disease ( 50 ), along with288 VEGF and PAI-1 51 .289 7. IL-13 (urine): IL-13 is a cytokine acting as an immunomodulator. It is an important mediator290 in responses to parasitic infections, as well as in inflammatory processes, allergies, and291 cancers 52 . The link between bladder cancer and increased levels of IL-13 in blood was292 already established by Szymanska et. al 53 .293 8. PAI-1/tPA complex (serum): PAI-1 is an inhibitor of tPA (tissue plasminogen activator).294 Together they form a complex which regulates fibrin clot degradation through hydrolysis.295 High concentrations of PAI-1 increase the risk of blood clots, while low concentrations of296 PAI-1 are linked with an elevated risk of haemorrhages. Both tPA and PAI-1 proteins are297 recognised as factors involved in carcinogenesis 36,54 . Elevated PAI-1 levels have been as-298 sociated with various types of cancers, as the protein can protect tumour cells from apop-299 tosis and confer resistance to therapeutic treatments 55,56 . In breast cancer, the prognostic300 value of the PAI-1/tPA complex has been shown to be comparable to that of total PAI-1,301 and it serves as a predictor of poor clinical outcome 55 . What is the most important in the302 context of the HaBio cohort? PAI-1 was identified as a potential biomarker for detecting303 bladder cancer ( 57,58 ), especially when measured with IL8, VEGF and APOE in a panel 51 .304 Most of the 10 most relevant features shown in the heatmaps (Fig. 5-6, Fig. S1) are associ-305 ated with bladder cancer or with carcinogenesis more broadly. The detection of many of them in306 urine samples strongly suggests that malignant processes are occurring within the urinary tract.307 What is also notable is that there are sex-specific differences (only two of the top 10 biomarkers308 17 are common to both sexes), not only between features but also in threshold values for biomark-309 ers that overlap across datasets, thereby confirming the previously proposed hypothesis that310 biomarkers characteristic of bladder cancer development differ by sex. Hence, the results can311 be treated as a basis for a screening panel with sex-specific biomarkers for bladder cancer. In-312 stead of evaluating individual features separately, it would be more informative to interpret them313 as a set of diagnostic rules. Because the established biomarker for bladder cancer, BTA, may314 be influenced by haematuria ( 43 ), a panel of additional biomarkers would be desirable to confirm315 or exclude the presence of bladder cancer. From a clinical perspective, taking urinary samples316 rather than blood is more comfortable for patients. In addition, screening using urine samples317 does not require medical staff to draw blood from patients, thereby reducing the cost of this type318 of medical examination and the burden on the healthcare system. While biomarkers are impor-319 tant factors to assess the risk of BC for each patient, it is also important to take a closer look at320 the history of smoking, which translates to harmful tar exposure. It is the only feature that is not321 a biomarker and ranks highly. The risk for BC associated with tobacco smoking is established322 as ground truth and cannot be overlooked. In the male group, a potential confounding factor is323 comorbidity, as patients may present both with bladder cancer and prostate issues (e.g. BPE)324 at once, which could affect the classification results. Therefore, analysing a set of biomarkers325 measured at the same time point provides a more reliable basis for decision-making than relying326 on a single biomarker (e.g., BTA).327 18 DATA AND METHODS328 Data collection and curation329 Rather than treating missing values solely as a nuisance, we explicitly model increasing levels of330 missingness as a controlled stress test to assess model robustness and feature stability under331 realistic data degradation scenarios. First, the data were cleaned, and EDA (exploratory data332 analysis) was performed on the HaBio dataset to prepare it for analysis. The HaBio cohort333 dataset includes 675 patients with haematuria, categorised into two groups: control and bladder334 cancer. To qualify as a control patient in the HaBio cohort study, individuals had to meet specific335 criteria, including no prior history of cancer, positive haematuria (past or present), a negative336 cystoscopy within the last 3 months, and no prior chemotherapy or radiotherapy. Additionally,337 control patients were selected to match bladder cancer patients already enrolled in the study in338 terms of sex, approximate age range, and smoking status. On the other hand, patients in the339 bladder cancer group were required to have a history of positive haematuria (past or present)340 and have undergone cystoscopy within the last six months. They also could not have a history341 of cancer other than bladder cancer and needed to have either a confirmed bladder cancer342 diagnosis or a suspicion of the disease.343 During the final diagnosis, 107 patients were labelled undiagnosed, while 9 patients were344 initially misclassified as having bladder cancer but were later found to have an infection. To345 eliminate this ambiguity, the 107 undiagnosed entries were removed, resulting in a final dataset346 of 568 patients, comprising 201 bladder cancer cases and 367 non-bladder cancer cases (in-347 cluding 9 entries initially classified as bladder cancer), which served as criterion/target class for348 classification. The dataset was notably unbalanced by gender, with 130 female patients (23%)349 and 438 male patients (77%) (Figure 7).350 Figure 7: Distribution of the HaBio cohort patients regarding sex. A bar plot illustrating the distribution of bladder cancer (BC) and bladder cancer free (non-BC) patients by sex. 19 HaBio Cohort dataset includes multi-dimensional patient data collected at various stages of351 the study conducted between 17/10/2012 and 21/02/2020. At the time of recruitment, a nurse352 or HaBio clinician collected information on the patientās age, sex, current medications, medical353 history, exposure risks, lifestyle choices, stimulant use, occupation, cystoscopy date, and bladder354 cancer diagnosis date, if applicable. It was also reported whether haematuria was visible with355 the naked eye. Urine and blood samples were analysed to quantify 79 biomarkers. During the356 follow-up period, the clinician conducted a detailed pathological review, including assessment of357 the conditions identified as the underlying causes of hematuria.358 Given that the original dataset contained a wide range of patient-related information, careful359 feature selection was necessary to retain only the most relevant variables for meaningful anal-360 ysis. To prevent misleading conclusions caused by extreme data scarcity, features (columns)361 related to exposure data and medication details were excluded from the final dataset. The re-362 sulting dataset contained a well-balanced selection of features, providing a holistic view of each363 patientās characteristics. It included demographic variables, lifestyle choices, general health in-364 formation, occupational risk assessed by the Office of National Statistics and 79 biomarkers365 measured in the HaBio study. Subsequently, for a more detailed analysis, the dataset was strat-366 ified by sex, yielding 3 final cohorts: a combined dataset for both sexes (here named TOTAL), a367 subset of female patients (FEMALES), and a subset of male patients (MALES).368 The final dataset consisted of both continuous and categorical columns: āBlaCA_noBlaCAā369 (target column storing whether the final diagnosis for a patient was diagnosed with bladder can-370 cer (1) or not (0)), Haem_Macro_Micro (haematuria: macro or micro), Ģ s80HdG (urine, ng/ml),371 Albumin - Creatinine Ratio ((ACR), urine), BTA (urine, Uml), CD44 (serum, ng/ml), CEA (serum,372 ng/ml), CK20 (urine, ng/ml), Clusterin (urine, ng/ml), Creatinine (urine,μmol/L), CRP (urine,373 mg/ml), CRP (serum, mg/ml), CXCL16 (urine, ng/ml), Cystatin-B (urine, ng/ml), Cystatin-C374 (urine, ng/ml), Cystatin-C (serum, ng/ml), D-dimer (urine, ng/ml), EGF (urine, pg/ml), EGF375 (serum, pg/ml), FABP-A (serum, ng/ml), FAS (urine, pg/ml), GRO (serum, pg/ml), HAD (serum,376 U/l), IFN-γ (urine, pg/ml), IFN-γ (serum, pg/ml), IL-1α (urine, pg/ml), IL-1α (serum, pg/ml), IL-1β377 (urine, pg/ml), IL-1β (serum, pg/ml), IL2 (urine, pg/ml), IL2 (serum, pg/ml), IL3 (urine, pg/ml),378 IL4 (urine, pg/ml), IL4 (serum, pg/ml), IL6 (urine, pg/ml), IL6 (serum, pg/ml), IL7 (urine, pg/ml),379 IL8 (urine, pg/ml), IL8 (serum, pg/ml), IL10 (urine, pg/ml), IL10 (serum, pg/ml), IL12p70 (urine,380 pg/ml), IL13 (urine, pg/ml), IL18 (urine, pg/ml), IL23 (urine, pg/ml), LASP-1 (serum, pg/ml), M30381 (serum, U/l), M2PK (serum, ng/ml), MCP-1 (urine, pg/ml), MCP-1 (serum, pg/ml), Microalbu-382 min (urine, mg/l), Midkine (urine, pg/ml), MMP9 (urine, ng/ml), MMP9/NGAL (urine, ng/ml),383 MMP9/TIMP (urine, ng/ml), NGAL (urine, ng/ml), NSE (urine, ng/ml), osmolality (urine, mOsm),384 PAI-1/tPA (serum, ng/ml), pERK (urine, pg/ml), Progranulin (urine, ng/ml), Prolactin (serum,385 μlU/l), Protein (urine, mg/ml), PSA-TPSA (serum, ng/ml), S100A4 (serum, ng/ml), sIL2Rα (serum, 386 ng/ml), sIL6R (urine, ng/ml), TGF-β1 (urine, pg/ml), Thrombomodulin (urine, ng/ml), TNFα (urine,387 pg/ml), TNFα (serum, pg/ml), sTNFR1 (urine, ng/ml), sTNFR2 (urine, ng/ml), tPA (urine, ng/ml),388 VEGF (urine, pg/ml), VEGF (serum, pg/ml), HDL (serum,μmol/L), LDL (serum,μmol/L), Triglyc-389 erides (serum,μmol/L) and Cholesterol (serum,μmol/L). The biomarker columns in the original390 dataset have a small proportion of missing values, ranging from 3 to 24 per cent, with a mean391 of 6 and a median of 3 across features. Additionally, datasets contained: sex and age columns,392 weekly alcohol intake (ml), daily fluid intake (ml), smoking status (past, current, never), smok-393 ing history (in years) and exposure to tar (mg), diabetes status, number of different medications394 taken daily, frequency of the urinary tract infections, and ONS ranking (assessment of workplace395 hazards). In total, the dataset comprises 89 features.396 20 CACTUS397 CACTUS (Comprehensive Abstraction and Classification Tool for Uncovering Structures) is based398 on a modified Naive Bayes classification algorithm 14 . An innovative data processing component399 of CACTUS is an abstraction (discretisation) module that transforms feature values into 2 cat-400 egories (Up and Down) using the ROC (Receiver Operating Characteristic) curve. The module401 enables effective interpretability of the results and ensures anonymisation. Following abstraction,402 classification is performed by matching each record (patientās features) to each class and assign-403 ing the class that is most similar. Each feature (column) is ranked by its discriminative strength,404 thereby highlighting its contribution to the classification process. Features with similar distri-405 butions across target classes (e.g., healthy vs. sick) are considered less informative, whereas406 those with distinct patterns across classes are more informative, as they better differentiate the407 classes. Feature ranking provides clinicians with interpretable insights into the modelās reason-408 ing by identifying features that enable the record to be classified into the appropriate class, and409 it helps recognise biases and faulty behaviour, thereby improving the model and increasing trust410 in its outputs.411 CLASSIC EXPLAINABLE MACHINE LEARNING APPROACHES412 1. Random Forest (RF) is a classification method based on a decision tree ensemble 59 . It413 creates multiple random subsets of the data, passes them to multiple decision trees, and414 then combines their outputs by majority voting. In majority voting, the final class label for a415 given sample is determined by the class most frequently predicted by the individual trees.416 Extracting the most important features from the RF model involves measuring the reduc-417 tion in Gini impurity when a particular feature is used for splitting; the greater the impurity418 reduction, the more important the feature is to the model. As a result of feature importance419 calculations, one can obtain the features that are most informative for the classification420 task 60 .421 2. Selected Gradient Boosting Methods (AdaBoost, XGB, LGBM, CatBoost): Gradient422 boosting (GB) is an ML approach that, similarly to RF, is based on an ensemble of decision423 trees. However, in GB, weak learners are trained sequentially to minimise the errors made424 by their predecessors, yielding a strong learner that makes accurate predictions of the tar-425 get variable. AdaBoost was the first practical application of the GB algorithm that uses426 ensembles of decision trees 61,62 . It assigns higher weights to misclassified samples, mak-427 ing them more relevant for subsequent trees. However, AdaBoost is not robust with large428 datasets and does not support missing values. To apply this algorithm to datasets with429 missing values, an imputation algorithm is required to impute the missing values with the430 most likely numerical value. In this study, mean imputation was applied, replacing missing431 values with the mean of non-missing values for each feature. Feature importance in the432 AdaBoost method is computed by first obtaining impurity-based importance values from433 each weak learner and then averaging them using the learnersā weights, which reflect their434 predictive performance. EXtreme Gradient Boosting (XGB) is another gradient boosting435 algorithm, designed to handle sparse datasets 63 . Its prediction is computed as the sum436 of the outputs from all trees once the stopping criterion (e.g., a sufficiently small error)437 is met. Unlike AdaBoost, XGB supports missing values through its Sparsity Aware Split438 Finding algorithm. Feature importance in XGB is derived similarly to random forests, but439 instead of using Gini impurity, it is based on the average reduction in the modelās overall440 classification loss. Both methods quantify how much a feature helps reduce the loss, but441 in random forests, this is measured locally via impurity reduction at each split, whereas in442 21 XGB, it reflects the contribution to the overall loss reduction (global model improvement).443 Light Gradient Boosting Machine (LGBM), similarly to XGB, is an alternative GB method444 based on decision trees ensemble 64 . It is optimised for large, high-dimensional datasets445 and can handle missing values natively. It can serve as an alternative to XGB, as fea-446 ture importances can be calculated in the same way. Lastly, CatBoost, also a gradient447 boosting algorithm 65 . It was developed to efficiently handle categorical features, which are448 often problematic for other GB methods. In addition, it supports sparse datasets (missing449 values) and provides fast execution. In CB, feature importance is computed to quantify the450 increase in the overall loss function when a feature is excluded. Compared to XGB and451 LGBM, which report gain as an averaged loss reduction per split in the decision tree, CB452 returns the total (raw) loss reduction contributed by each feature.453 22 CONCLUSIONS454 We performed experiments using the HaBio dataset to classify patient cases into bladder can-455 cer and non-bladder cancer groups. CACTUS achieved the best overall performance in terms456 of balanced accuracy and sensitivity across all 3 subgroups (total, males, and females) com-457 pared with classical ML methods. What stands out is CACTUS showing superior feature sta-458 bility, which is important in the case of applying ML models to medical scenarios: the most459 important features for the model remained consistent across datasets with varying proportions460 of missing data and showed the highest degree of overlap even when missingness reached461 30%, confirming the robustness of the method. Although bladder cancer serves as a concrete462 and clinically relevant case study, the primary contribution of this work lies in its data-centric463 perspective on trustworthy machine learning. The challenges addressed hereāmissing data,464 limited sample sizes, and the need for interpretable and stable feature selectionāare common465 across biomedical and real-world datasets. These findings are particularly relevant from a clin-466 ical perspective, as medical datasets are often incomplete. Moreover, we have demonstrated467 that the set of relevant biomarkers yields greater information than any individual biomarker. On468 average, CACTUS demonstrated the best feature stability for males and females and ranked469 second after Random Forest for the total population. This difference between men and women470 may be attributed to distinct biomarker patterns, as the most important features (biomarkers)471 vary between these subgroups. Our results demonstrate that CACTUS is a promising candi-472 date as a trustworthy AI system for medical decision-making, reducing uncertainty associated473 with complex datasets and enabling the identification of key patterns that could support the de-474 velopment of novel bladder cancer screening assays. CACTUS is effective and stable not only475 with complete data but also with many missing fields, and it offers the possibility of integrating476 multi-dimensional data of various types (continuous, categorical). Our findings demonstrate that477 predictive performance alone is insufficient to characterise model reliability. Models with com-478 parable accuracy can exhibit markedly different behaviour in terms of feature stability, with im-479 plications for reproducibility, interpretability, and downstream decision-making. Feature stability,480 therefore, provides complementary information that should be considered alongside traditional481 evaluation metrics. Future work could apply our approach by dividing the datasets into biomarker482 and lifestyle/demographic subsets. This distinction could help clarify the relative contributions of483 each type of data, given that biomarkers appear to be more important for classification and may484 overshadow lifestyle/demographic factors, while also providing insight into which factors patients485 can control to reduce their risk of developing bladder cancer. Moreover, larger datasets should486 be used to confirm our initial findings based on the presented cohort. It is also important to487 note that the thresholds calculated by CACTUS to abstract features are not the same as those488 set by diagnostic companies. Thresholds represent the best value that separates two classes489 (here, BC and non-BC cases), and thus do not carry the same information as thresholds set by490 medical institutions for diagnosis, which result from a combination of features used in the model491 rather than from the influence of a single characteristic. Additionally, we should consider the492 differential impacts of each feature across populations. Lastly, another benefit of using CACTUS493 is that it does not require prior domain knowledge in the field of ML model building, conversely to494 when classic ML models are built (multiple hyperparameters need to be established before con-495 structing the model), as a result, CACTUS not only provides more reliable results explainability,496 which is important in the healthcare sector, but is also easy to use by medical staff without the497 need for lengthy and costly trainings. From a data science standpoint, this work emphasises the498 importance of evaluating how models respond to data perturbations rather than static datasets499 alone. Incorporating stability-oriented analyses enables more reliable knowledge extraction and500 supports the development of trustworthy machine learning systems in domains where data im-501 23 perfections are the norm rather than the exception. All the arguments presented here emphasise502 that CACTUS can successfully serve as a trustworthy AI framework for medical applications.503 By prioritising interpretability and feature stability under incomplete data, CACTUS contributes504 a generalizable framework for robust pattern discovery in complex datasets. This perspective505 complements performance-driven evaluation and supports more trustworthy data-driven insights506 across application domains, and suggests that stability-aware evaluation should be considered507 a standard component of pattern discovery pipelines operating on imperfect real-world data.508 RESOURCE AVAILABILITY509 Lead contact510 Requests for further information and resources should be directed to and will be fulfilled by the511 lead contact, Paulina Tworek (p.tworek@sanoscience.org) or Jose Sousa (j.sousa@sanoscience.org).512 Data and code availability513 ⢠The description of HaBio Clinical Study is presented at: https://w.isrctn.com/ISRCTN25823942.514 ⢠Access to the CACTUS tool is provided upon request by Jose Sousa (j.sousa@sanoscience.org).515 ⢠Any additional information required to reanalyze the data reported in this paper is available516 from the lead contact upon request.517 ACKNOWLEDGMENTS518 This project has received funding from the European Unionās Horizon 2020 research and in-519 novation programme under grant agreement No 857533 and from the International Research520 Agendas Programme of the Foundation for Polish Science No MAB PLUS/2019/13. The publi-521 cation was created within the project of the Minister of Science and Higher Education "Support522 for the activity of Centers of Excellence established in Poland under Horizon 2020" on the basis523 of the contract number MEiN/2023/DIR/3796. This project has received funding from the Eu-524 ropean Unionās Horizon 2020 research and innovation programme under grant agreement No525 857524.526 AUTHOR CONTRIBUTIONS527 Conceptualization, S.A.O. and P.T.; methodology, S.A.O. and P.T.; investigation, S.A.O. and P.T.;528 writing-āoriginal draft, S.A.O., P.T., L.G., M.W.R., M.J.K., P.F. and J.S.; funding acquisition, J.S.529 and M.W.R..; resources, S.A.O., P.T., L.G., M.W.R., M.J.K., P.F. and J.S.; supervision, J.S. and530 M.W.R.531 DECLARATION OF INTERESTS532 Mark W. Ruddock, Mary Jo Kurt and Peter Fitzgerald are employees of Randox Laboratories Ltd.533 24 DECLARATION OF GENERATIVE AI AND AI-ASSISTED TECH-534 NOLOGIES535 During the preparation of this work, the authors didnāt use any AI-assisted technologies, the536 authors reviewed and edited the content as needed and take full responsibility for the content of537 the publication.538 SUPPLEMENTAL INFORMATION INDEX539 Figure S1-S4 and Tables S1-S12 in a Supplementary Materials PDF.540 25 References541 1. Whang, W. (2013). Medical Decision-Making. Encyclopedia of Behavioral Medicine..542 Springer. doi: 10.1097/ACM.0000000000002902.543 2. Katz, D. (2001). Clinical Epidemiology and Evidence-Based Medicine: Fundamental544 Principles of Clinical Reasoning and Research. SAGE Publications.doi: 10.4135/545 9781452232638.546 3. Turner, J. (2013). Clinical Decision-Making. Encyclopedia of Behavioral Medicine.. Springer.547 doi: 10.1007/978-1-4419-1005-9_996.548 4. Miller, R., Pople, H.J., and Myers, J. (1982). Internist-1, an experimental computer-based549 diagnostic consultant for general internal medicine. The New England Journal of Medicine550 307. doi: 10.1056/NEJM198208193070803.551 5. Kaur, D., Uslu, S., Rittichier, K.J., and Durresi, A. (2022). Trustworthy artificial intelligence:552 A review. ACM Computing Surveys 55. doi: 10.1145/3491209.553 6. Beger, J. (2025). Not someone, but something: Rethinking trust in the age of medical ai.554 European Journal of Radiology Artificial Intelligence 3. doi: 10.1016/j.ejrai.2025.100038.555 7. Piccialli, F., Somma, V.D., Giampaolo, F., Cuomo, S., and Fortino, G. (2021). A survey on556 deep learning in medicine: Why, how and when? Information Fusion 66, 111ā137. doi:557 10.1016/j.inffus.2020.09.006.558 8. Ali, S., Abuhmed, T., El-Sappagh, S., Muhammad, K., Alonso-Moral, J.M., Confalonieri, R.,559 Guidotti, R., Del Ser, J., DĆaz-RodrĆguez, N., and Herrera, F. (2023). Explainable Artificial560 Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence.561 Information Fusion 99, 101805. doi: 10.1016/j.inffus.2023.101805.562 9. Rudin, C. (2019). Stop explaining black box machine learning models for high stakes563 decisions and use interpretable models instead.Nature Machine Intelligence 1. doi:564 10.1038/s42256-019-0048-x.565 10. Kaur, D., Uslu, S., Rittichier, K.J., and Durresi, A. (2022). Trustworthy Artificial Intelligence:566 A Review. ACM Computing Surveys 55. doi: 10.1145/3491209.567 11. Li, Y.H., Li, Y.L., Wei, M.Y., and Li, G.Y. (2024). Innovation and challenges of artificial568 intelligence technology in personalized healthcare. Scientific Reports 14. doi: 10.1038/569 s41598-024-70073-7.570 12. Heymans, M.W., and Twisk, J.W. (2022). Handling missing data in clinical research. Journal571 of Clinical Epidemiology 151, 185ā188. doi: 10.1016/j.jclinepi.2022.08.016.572 13. Marino, M., Lucas, J., Latour, E., and Heintzman, J.D. (2021). Missing data in primary573 care research: importance, implications and approaches. Family Practice 38. URL: https:574 //doi.org/10.1093/fampra/cmaa134. doi: 10.1093/fampra/cmaa134.575 14. Gherardini, L., Varma, V.R., CapaÅa, K., Woods, R., and Sousa, J. (2024). CACTUS: A576 Comprehensive Abstraction and Classification Tool for Uncovering Structures. ACM Trans.577 Intell. Syst. Technol. 15. doi: 10.1145/3649459.578 26 15. Capala, K., Tworek, P., and Sousa, J. (2025). Stability of Machine Learning Predictive579 Features Under Limited Data . IEEE Transactions on Knowledge & Data Engineering 37.580 doi: 10.1109/TKDE.2025.3580671.581 16. Williamson, K. (2019). The Haematuria Biomarker Study (HaBio). ISRCTN is the UKās582 clinical study registry. doi: 10.1186/ISRCTN25823942 available at: https://w.isrctn.583 com/ISRCTN25823942.584 17. Isgrò, M.A., Bottoni, P., and Scatena, R. (2015). Neuron-specific enolase as a biomarker:585 Biochemical and clinical aspects. In Advances in Cancer Biomarkers: From biochemistry586 to clinic for a critical revision. Springer Netherlands. doi: 10.1007/978-94-017-7215-0_9.587 18. Jebet, A., Kibet, J.K., Kinyanjui, T., and Nyamori, V.O. (2018). Environmental inhalants from588 tobacco burning: Tar and particulate emissions. Scientific African 1. doi: https://doi.org/589 10.1016/j.sciaf.2018.e00004.590 19. Freedman, N.D., Silverman, D.T., Hollenbeck, A.R., Schatzkin, A., and Abnet, C.C. (2011).591 Association between smoking and risk of bladder cancer among men and women. JAMA592 306. doi: 10.1001/jama.2011.1142.593 20. Xing Du, W.S., Zhongyao Chen (2025). Clusterin: structure, function and roles in disease.594 International Journal of Medical Sciences 22. doi: 10.7150/ijms.107159.595 21. Hazzaa, S.M. (2010). Clusterin as a diagnostic and prognostic marker for transitional596 cell carcinoma of the bladder. Pathology oncology research : POR 16. doi: 10.1007/597 s12253-009-9196-3.598 22. Shimizu, M., Kondo, M., Ito, Y., Kume, H., Suzuki, R., and Yamaki, K. (2005). Soluble fas599 and fas ligand provide new information on metastasis and response to chemotherapy in sclc600 patients. Cancer Detection and Prevention 29. doi: 10.1016/j.cdp.2004.09.001.601 23. Abbasova, S.G., Vysotskii, M.M., Ovchinnikova, L.K., Obusheva, M.N., Digaeva, M.A.,602 Britvin, T.A., Bahoeva, K.A., Karabekova, Z.K., Kazantzeva, I.A., Mamedov, U.R., Manuchin,603 I.B., and Davidov, M.I. (2009). Cancer and soluble fas. Bulletin of Experimental Biology and604 Medicine 148.605 24. Ahn, J.H., Kang, C.K., Kim, E.M., Kim, A.R., and Kim, A. (2022). Proteomics for early606 detection of non-muscle-invasive bladder cancer: Clinically useful urine protein biomarkers.607 Life 12. doi: 10.3390/life12030395.608 25. YOUICHI, M., OSAMU, Y., and BENJAMIN, B. (1998). Prognostic significacne of soluble609 fas in the serum of patients with bladder cancer. Journal of Urology 160. doi: 10.1016/610 S0022-5347(01)62960-4.611 26. Goffin, V., Hoang, D.T., Bogorad, R.L., and Nevalainen, M.T. (2011). Prolactin regulation612 of the prostate gland: a female player in a male game. Nature Reviews Urology 8. doi:613 10.1038/nrurol.2011.143.614 27. Ferrara, N. (1999). Role of vascular endothelial growth factor in the regulation of angiogen-615 esis. Kidney International 56. doi: 10.1046/j.1523-1755.1999.00610.x.616 28. Ghalehbandi, S., Yuzugulen, J., Pranjol, M.Z.I., and Pourgholami, M.H. (2023). The role of617 vegf in cancer-induced angiogenesis and research progress of drugs targeting vegf. Euro-618 pean Journal of Pharmacology 949. doi: 10.1016/j.ejphar.2023.175586.619 27 29. Elayat, G., Punev, I., and Selim, A. (2023). An overview of angiogenesis in bladder cancer.620 Current Oncology Reports 25. doi: 10.1007/s11912-023-01421-5.621 30. Rosser, C.J., Chang, M., Dai, Y., Ross, S., Mengual, L., Alcaraz, A., and Goodison, S.622 (2014). Urinary protein biomarker panel for the detection of recurrent bladder cancer. Can-623 cer epidemiology, biomarkers and prevention : a publication of the American Association for624 Cancer Research, cosponsored by the American Society of Preventive Oncology 23. doi:625 10.1158/1055-9965.EPI-14-0035.626 31. Guicciardi, M.E., Leist, M., and Gores, G.J. (2004). Lysosomes in cell death. Oncogene 23.627 doi: 10.1038/sj.onc.1207512.628 32. Feldman, A.S., Banyard, J., Wu, C.L., McDougal, W.S., and Zetter, B.R. (2009). Cystatin b629 as a tissue and urinary biomarker of bladder cancer recurrence and disease progression.630 Clinical Cancer Research 15. doi: 10.1158/1078-0432.CCR-08-1143.631 33. Yıldırım, B., Kulak, K., and Bilir, A. (). European journal of breast health 20. doi: 10.4274/632 ejbh.galenos.2024.2024-4-7.633 34. Zhou, L., Jiang, J., Fu, Y., Zhang, D., Li, T., Fu, Q., Yan, C., Zhong, Y., Dionigi, G., Liang,634 N., and Sun, H. (2021). Diagnostic performance of midkine ratios in fine-needle aspirates635 for evaluation of cytologically indeterminate thyroid nodules. Diagnostic Pathology 16. doi:636 10.1186/s13000-021-01150-y.637 35. Jones, D.R. (2014). Measuring midkine: the utility of midkine as a biomarker in cancer and638 other diseases. British Journal of Pharmacology 171. doi: 10.1111/bph.12601.639 36. de Witte, J.H., Sweep, C.G.J., Klijn, J.G.M., Grebenschikov, N., Peters, H.A., Look, M.P.,640 van Tienoven, T., Heuvel, J.J.T.M., Vries, J.B.D., Benraad, T., and Foekens, J.A. (1999).641 Prognostic value of tissue-type plasminogen activator (tpa) and its complex with the type-642 1 inhibitor (pai-1) in breast cancer. British Journal of Cancer 80. doi: 10.1038/sj.bjc.643 6690353.644 37. Sutton, J.M. (1990). Evaluation of hematuria in adults. JAMA 263. doi: 10.1001/jama.1990.645 03440180081037.646 38. Ingelfinger, J.R. (2021). Hematuria in adults. New England Journal of Medicine 385. doi:647 10.1056/NEJMra1604481.648 39. Luo, L., Yang, Y., Kieneker, L.M., Janse, R.J., Bosi, A., Mazhar, F., de Boer, R.A., de Bock,649 G.H., Gansevoort, R.T., and Carrero, J.J. (2023). Albuminuria and the risk of cancer: the650 stockholm creatinine measurements (scream) project. Clinical Kidney Journal 16. doi: 10.651 1093/ckj/sfad145.652 40. Luo, L., Kieneker, L.M., van der Vegt, B., Bakker, S.J.L., Gruppen, E.G., Casteleijn, N.F.,653 de Boer, R.A., Suthahar, N., de Bock, G.H., Aboumsallem, J.P., Vart, P., and Gansevoort,654 R.T. (2023). Urinary albumin excretion and cancer risk: the prevend cohort study. Nephrol-655 ogy Dialysis Transplantation 38. doi: 10.1093/ndt/gfad107.656 41. Cheng, Z.Z., Corey, M.J., Parepalo, M., Majno, S., Hellwage, J., Zipfel, P.F., Kinders, R.J.,657 Raitanen, M., Meri, S., and Jokiranta, T.S. (2005). Complement factor h as a marker for658 detection of bladder cancer. Clinical Chemistry 51. doi: 10.1373/clinchem.2004.042192.659 28 42. Kinders, R., Jones, T., Root, R., Bruce, C., Murchison, H., Corey, M., Williams, L., Enfield,660 D., and Hass, G.M. (1998). Complement factor h or a related protein is a marker for transi-661 tional cell cancer of the bladder. Clinical cancer research : an official journal of the American662 Association for Cancer Research 4.663 43. Miyake, M., Goodison, S., Rizwani, W., Ross, S., Grossman, H.B., and Rosser, C.J. (2012).664 Urinary bta: indicator of bladder cancer or of hematuria. World Journal of Urology 30. doi:665 10.1007/s00345-012-0935-9.666 44. Cavalli, G., Colafrancesco, S., Emmi, G., Imazio, M., Lopalco, G., Maggio, M.C., Sota,667 J., and Dinarello, C.A. (2021). Interleukin 1alpha: a comprehensive review on the role of668 il-1alpha in the pathogenesis and treatment of autoimmune and inflammatory diseases.669 Autoimmunity Reviews 20. doi: 10.1016/j.autrev.2021.102763.670 45. Yao, S.J., Ma, H.S., Liu, G.M., Gao, Y., and Wang, W. (2023). Increased il-1alpha expression671 is correlated with bladder cancer malignant progression. Archives of Medical Science 19.672 doi: 10.5114/aoms.2020.100677.673 46. Xu, J., Fu, S., Peng, W., and Rao, Z. (2012). Mcp-1-induced protein-1, an immune regulator.674 Protein & Cell 3. doi: 10.1007/s13238-012-2075-9.675 47. Singh, S., Anshita, D., and Ravichandiran, V. (2021). Mcp-1: Function, regulation, and676 involvement in disease. International Immunopharmacology 101. doi: 10.1016/j.intimp.677 2021.107598.678 48. Amann, Perabo, Wirger, Hugenschmidt, and Schultze-Seemann (1998). Urinary levels of679 monocyte chemo-attractant protein-1 correlate with tumour stage and grade in patients with680 bladder cancer. British Journal of Urology 82. doi: 10.1046/j.1464-410x.1998.00675.x.681 49. Meier, C., and Brieger, A. (2025). The role of il-8 in cancer development and its impact on682 immunotherapy resistance. European Journal of Cancer 218. doi: 10.1016/j.ejca.2025.683 115267.684 50. VandenBussche, C.J., Heaney, C.D., Kates, M., Hooks, J.J., Baloga, K., Sokoll, L., Rosen-685 thal, D., and Detrick, B. (2024). Urinary il-6 and il-8 as predictive markers in bladder urothe-686 lial carcinoma: A pilot study. Cancer Cytopathology 132. doi: https://doi.org/10.1002/687 cncy.22767.688 51. Bardowska, K., Krajewski, W., KoÅodziej, A., Ko Ģ scielska-Kasprzak, K., Bartoszek, D.,689 Ģ Zabi Ģ nska, M., Chorbi Ģ nska, J., Kubacki, F., Królicki, T., Krajewska, M., SzydeÅko, T., and690 Kami Ģ nska, D. (2025). Evaluation of six novel biomarkers for predicting recurrence of non-691 muscle invasive bladder cancer after endoscopic resectionā a prospective observational692 study. World Journal of Urology 43. doi: 10.1007/s00345-025-05485-9.693 52. Shi, J., Song, X., Traub, B., Luxenhofer, M., and Kornmann, M. (2021). Involvement of il-4,694 il-13 and their receptors in pancreatic cancer. International Journal of Molecular Sciences695 22. doi: 10.3390/ijms22062998.696 53. Szymanska, B., Sawicka, E., Jurkowska, K., Matuszewski, M., Dembowski, J., and Piwowar,697 A. (2021). The relationship between interleukin-13 and angiogenin in patients with bladder698 cancer. Journal of physiology and pharmacology : an official journal of the Polish Physio-699 logical Society 72. doi: 10.26402/jpp.2021.4.13.700 29 54. Placencio, V.R., and DeClerck, Y.A. (2015). Plasminogen activator inhibitor-1 in cancer:701 Rationale and insight for future therapeutic testing. Cancer Research 75. doi: 10.1158/702 0008-5472.CAN-15-0876.703 55. A, I., T, F., T, U., T, T., K-i, H., H, H., A, O., T, M., K, A., and T, Y. (2024). Plasminogen704 activator inhibitor-1 promotes immune evasion in tumors by facilitating the expression of705 programmed cell death-ligand 1. Frontiers in Immunology 15. doi: 10.3389/fimmu.2024.706 1365894.707 56. Kumara, H.S., Addison, P., Gamage, D.N., Pettke, E., Shah, A., Yan, X., Cekic, V., and708 Whelan, R.L. (2022). Sustained postoperative plasma elevations of plasminogen activator709 inhibitor-1 following minimally invasive colorectal cancer resection. Molecular and Clinical710 Oncology 16.711 57. Furuya, H., Sasaki, Y., Chen, R., Peres, R., Hokutan, K., Murakami, K., Kim, N., Chan,712 O.T.M., Pagano, I., DyrskjĆøt, L., Jensen, J.B., Malmstrom, P.U., Segersten, U., Sun, Y., Arab,713 A., Goodarzi, H., Goodison, S., and Rosser, C.J. (2022). Pai-1 is a potential transcriptional714 silencer that supports bladder cancer cell activity. Scientific Reports 12. doi: 10.1038/715 s41598-022-16518-3.716 58. Chan, O.T., Furuya, H., Pagano, I., Shimizu, Y., Hokutan, K., DyrskjĆøt, L.,717 Bjerggaard Jensen, J., Malmstrom, P.U., Segersten, U., Janku, F., and Rosser, C.J. (2017).718 Association of mmp-2, rb and pai-1 with decreased recurrence-free survival and over-719 all survival in bladder cancer patients. Oncotarget 8. doi: https://doi.org/10.18632/720 oncotarget.20686.721 59. Breiman, L. (2001). Random forests. Machine Learning 45. doi: 10.1023/A:1010933404324.722 60. Louppe, G., Wehenkel, L., Sutera, A., and Geurts, P. (2013). Understanding variable impor-723 tances in forests of randomized trees. In Proceedings of the 27th International Conference724 on Neural Information Processing Systems - Volume 1. Curran Associates Inc. p. 431ā439.725 61. Schapire, R.E. (2013). Explaining adaboost.In Empirical Inference: Festschrift in726 Honor of Vladimir N. Vapnik. Springer Berlin Heidelberg p. 37ā52.doi: 10.1007/727 978-3-642-41136-6_5.728 62. Schapire, R.E. (1999). A brief introduction to boosting. In Proceedings of the 16th Inter-729 national Joint Conference on Artificial Intelligence - Volume 2. San Francisco, CA, USA:730 Morgan Kaufmann Publishers Inc. p. 1401ā1406.731 63. Chen, T., and Guestrin, C. (2016). Xgboost: A scalable tree boosting system. In Proceed-732 ings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data733 Mining. Association for Computing Machinery. doi: 10.1145/2939672.2939785.734 64. Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.Y. (2017).735 Lightgbm: a highly efficient gradient boosting decision tree. In Proceedings of the 31st736 International Conference on Neural Information Processing Systems. Curran Associates737 Inc. p. 3149ā3157.738 65. Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A.V., and Gulin, A. (2018). Catboost:739 unbiased boosting with categorical features. In Proceedings of the 32nd International Con-740 ference on Neural Information Processing Systems. Curran Associates Inc. p. 6639ā6649.741 30