Paper deep dive
CEL: Comprehensive Counterfactual Explanations Library and Benchmark
Oleksii Furman, Łukasz Lenkiewicz, Marcel Musiałek, Maciej Zięba
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/2/2026, 12:55:38 PM
Summary
The paper introduces CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations in explainable AI. It addresses the lack of standardized evaluation by providing 18 datasets and 14 counterfactual methods (local, global, and group-wise) with a consistent evaluation protocol measuring validity, proximity, sparsity, and plausibility.
Entities (13)
Relation Signals (11)
CEL → includes → 18 Datasets
confidence 95% · CEL includes 18 datasets of varying size and complexity
CEL → includes → 14 Methods
confidence 95% · provides implementations or reimplementations of 14 widely used counterfactual methods
Wrocław University of Science and Technology → affiliatedwith → CEL
confidence 90% · Authors are from Wrocław University of Science and Technology ... we introduce CEL
CEL → evaluatesusing → Proximity
confidence 90% · The evaluation protocol incorporates multiple complementary metrics capturing ... proximity
CEL → evaluatesusing → Sparsity
confidence 90% · The evaluation protocol incorporates multiple complementary metrics capturing ... sparsity
CEL → evaluatesusing → Plausibility
confidence 90% · The evaluation protocol incorporates multiple complementary metrics capturing ... distributional plausibility
CEL → evaluatesusing → Validity
confidence 90% · The evaluation protocol incorporates multiple complementary metrics capturing validity
AReS → ispartof → CEL
confidence 90% · CEL implements AReS
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model's prediction to a desired outcome. While early methods primarily focused on minimal feature changes, recent work incorporates additional properties such as sparsity, actionability and plausibility. Despite this progress, fair and systematic evaluation remains challenging. Existing studies often rely on different data splits, predictive models, and evaluation metrics, which limits objective comparison across methods. To fill this gap, we introduce CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations designed to support consistent implementation and evaluation. CEL includes 18 datasets of varying size and complexity and provides implementations or reimplementations of 14 widely used counterfactual methods. Using this standardized setup, we conduct a comprehensive quantitative comparison across a variety of methods on datasets that differ in size, number, and types of attributes. The evaluation protocol incorporates multiple complementary metrics capturing validity, coverage, sparsity, proximity, and distributional plausibility, including density- and outlier-based measures to assess the realism of generated counterfactuals. To the best of our knowledge, this is the first comprehensive benchmark that systematically evaluates recent counterfactual explanation methods within a unified and reproducible framework. While prior libraries and benchmarking efforts exist in the literature, many are outdated, limited in scope, or lack consistent evaluation protocols. The proposed benchmark aims to improve reproducibility, enable fair comparison, and establish a workbench for the development of future counterfactual explanation methods.
Tags
Links
- Source: https://arxiv.org/abs/2607.22045v1
- Canonical: https://arxiv.org/abs/2607.22045v1
Trouble viewing inline? Open PDF directly →
Full Text
37,188 characters extracted from source content.
Expand or collapse full text
CEL: Comprehensive Counterfactual Explanations Library and Benchmark Oleksii Furman ⋆ , Łukasz Lenkiewicz ∗ , Marcel Musiałek, and Maciej Zięba Wrocław University of Science and Technology, Wrocław, Poland oleksii.furman, lukasz.lenkiewicz, maciej.zieba@pwr.edu.pl, 279704@student.pwr.edu.pl Abstract. Counterfactual explanations are a prominent approach in explainable artificial intelligence (xAI), providing actionable guidance on what input changes would alter a model’s prediction to a desired outcome. While early methods primarily focused on minimal feature changes, recent work incorporates additional properties such as spar- sity, actionability and plausibility. Despite this progress, fair and sys- tematic evaluation remains challenging. Existing studies often rely on different data splits, predictive models, and evaluation metrics, which limits objective comparison across methods. To fill this gap, we intro- duce CEL (Counterfactual Explanations Library), a unified library and benchmark for counterfactual explanations designed to support consis- tent implementation and evaluation. CEL includes 18 datasets of varying size and complexity and provides implementations or reimplementations of 14 widely used counterfactual methods. Using this standardized setup, we conduct a comprehensive quantitative comparison across a variety of methods on datasets that differ in size, number, and types of attributes. The evaluation protocol incorporates multiple complementary metrics capturing validity, coverage, sparsity, proximity, and distributional plau- sibility, including density- and outlier-based measures to assess the real- ism of generated counterfactuals. To the best of our knowledge, this is the first comprehensive benchmark that systematically evaluates recent counterfactual explanation methods within a unified and reproducible framework. Whereas prior libraries are often outdated, limited in scope, or lack consistent protocols, CEL aims to improve reproducibility, enable fair comparison, and provide a workbench for developing future counter- factual explanation methods. Keywords: Counterfactual Explanations· Explainable Artificial Intel- ligence· Benchmarking· Machine Learning 1 Introduction Machine learning models are increasingly deployed in high-stakes applications, where transparency and accountability are critical. Counterfactual explanations ⋆ O. Furman and Ł. Lenkiewicz contributed equally to this research. arXiv:2607.22045v1 [cs.LG] 24 Jul 2026 2O. Furman et al. Data Module DatasetsPreprocessing Model Module Explanation Engine Local: CCHVAE, DICE, PPCEF, ... Global: AReS, GLOBE-CE Group-Wise: GLANCE TCREx Predictive Classifiers: LR, MLP Regressors: Linear, MLP Actionability Constraints Feature Bounds Probabilistic Density Estimation: KDE, GMM Normalizing Flows Metrics Orchestrator Coverage & Validity Proximity: L1, L2, MAD Sparsity Plausibility: Likelihood, LOF, IsoForest Counterfactual Reports & Visualisations Fig. 1. Overview of the CEL architecture. It consists of four interconnected modules: (i) a Data Module handling dataset loading, preprocessing, and constraint specifica- tion; (i) a Model Module providing predictive and probabilistic backbones; (i) an Explanation Engine supporting local, global, and group-wise counterfactual methods; and (iv) a Metrics Orchestrator computing validity, proximity, sparsity, and plausibility measures. Arrows indicate data flow between components. have emerged as a prominent approach in explainable AI, offering actionable insights by identifying minimal changes required to alter a model’s prediction. Early formulations, most notably by Wachter et al. [32], focused primarily on two criteria: validity (the counterfactual must achieve the desired prediction) and proximity (the modification from the original instance should be minimal). While intuitive, these constraints alone often lead to unrealistic or impractical explanations. Subsequent work has expanded the set of desirable properties for counterfac- tual explanations. In addition to validity and proximity, researchers have empha- sized plausibility, requiring counterfactuals to lie in high-density regions of the data distribution; sparsity, encouraging modifications to as few features as possi- ble; and actionability, ensuring that only features that can be feasibly changed in practice are modified [8]. As a result, modern counterfactual generation methods typically optimize for multiple, and sometimes competing, criteria. This grow- ing set of constraints has significantly increased the complexity of both method design and evaluation. Since their introduction, the evaluation of counterfactual methods has re- mained fundamentally under-specified. For a given instance, multiple valid coun- terfactuals may exist, each satisfying different subsets of desirable properties [20]. Modern methods therefore optimize trade-offs among criteria such as validity, proximity, plausibility, sparsity, robustness, and actionability, requiring a multi- dimensional assessment rather than a single scalar metric [8]. CEL: Counterfactual Explanations Library and Benchmark3 In practice, studies use inconsistent evaluation setups: metrics are defined differently, different subsets of properties are emphasized, and experimental pro- tocols vary in data splits, preprocessing, feature encodings, constraints, and the predictive models being explained [21, 25]. Because counterfactuals are inher- ently model-dependent, such variation can substantially affect feasibility and quality, so observed differences may reflect experimental configuration rather than methodological advances. This lack of standardization makes fair compari- son difficult, yields unstable method rankings, and hinders reliable measurement of progress. To address this fragmentation, we introduce CEL, a unified benchmark and reproducible evaluation framework. Rather than serving solely as an implemen- tation library, CEL establishes a controlled protocol that standardizes datasets, predictive backbones, preprocessing pipelines, feature constraints, and evalu- ation metrics, enabling systematic and fair comparison across counterfactual paradigms. Figure 1 overviews the CEL architecture, connecting data manage- ment, predictive and generative modeling, explanation generation, and standard- ized evaluation. CEL includes 18 curated datasets spanning diverse domains and 14 widely used counterfactual generation methods implemented or reimplemented within a common interface. All methods are evaluated under harmonized training splits, shared predictive models, and a consistent metric taxonomy. The library is pub- licly available via pip install ce-library and the source code is hosted on GitHub 1 . Our main contributions are: – A controlled evaluation protocol for counterfactual explanations that standardizes datasets, preprocessing, predictive backbones, constraint han- dling, and metric definitions, enabling fair and reproducible comparison. – A benchmark including 18 datasets and 14 counterfactual generation methods implemented within a unified framework, supporting evaluation across local, global, and group-wise paradigms. – An open-source programming library designed to facilitate future method integration, transparent reporting, and community-driven benchmarking. 2 Related Works Research on counterfactual explanations has led to several libraries and bench- marking efforts designed to support implementation and evaluation. However, despite these contributions, a fully controlled and protocol-standardized bench- mark for counterfactual explanations is still lacking. CARLA [23] is among the earliest and most widely adopted benchmarking libraries for counterfactual explanations and algorithmic recourse, providing mul- tiple generation methods, integrated evaluation measures, and several datasets 1 Repository URL omitted for double-blind review. 4O. Furman et al. within a unified Python framework. While it greatly improved accessibility and comparability, it does not enforce a fully fixed experimental protocol across pre- dictive backbones, preprocessing choices, and constraint configurations. More recently, RobustX [11] targets the robustness of counterfactual expla- nations under model and data perturbations, offering standardized tools for this property. While valuable for analyzing stability, its scope is narrow and does not cover multiple counterfactual paradigms under a unified experimental protocol. In contrast to prior libraries, CEL emphasizes protocol-level control across datasets, predictive backbones, preprocessing steps, constraint handling, and metric definitions, enabling systematic and fair comparison across local, global, and group-wise counterfactual paradigms. 3 CEL: Counterfactual Explanations Library CEL is built from modular components that separate models, datasets, counter- factual methods, and evaluation metrics. This structure allows flexible compo- sition and simplifies extension with new methods, datasets, and metrics, while typed interfaces ensure consistent data handling. 3.1 Architecture and Design Principles CEL enforces a strict separation of concerns, so that data preprocessing, model training, explanation generation, and evaluation operate as decoupled yet in- teroperable components. It employs a configuration-driven workflow, letting researchers define experiments—model hyperparameters, dataset constraints, and metric selection—via hierarchical configuration files rather than hard-coded scripts. 3.2 Unified Method Interfaces A key challenge in benchmarking is the heterogeneity of counterfactual algo- rithms. CEL unifies them under a common abstraction, BaseCounterfactualMethod , which defines a standardized interface for explanation generation. All methods return a structuredExplanationResult , enabling consistent evaluation and di- rect comparison across approaches. To support diverse explanation paradigms, CEL adopts a Mixin-based design that unifies local, global, and group-wise methods under a common interface, enabling consistent integration and evaluation. 3.3 Data Management and Preprocessing Valid counterfactual evaluation requires rigorous data handling to ensuring that generated instances remain within valid feature domains. CEL implements a robust DatasetBase architecture that manages feature metadata, including ac- tionability constraints, mutability, and feature bounds. CEL: Counterfactual Explanations Library and Benchmark5 The library features a flexiblePreprocessingPipeline , allowing users to chain transformations such as scalers (Standard, MinMax, Robust) and encoders (OneHot, Ordinal). Crucially, the pipeline supports precise inverse transforma- tions, ensuring that counterfactuals generated in processed space can be accu- rately mapped back to the original input space for evaluation. For methods requiring continuous density estimation (e.g., normalizing flows), CEL includes a specialized Dequantization module. This module converts dis- crete categorical features into continuous values using variational or noise-based dequantizers, bridging the gap between discrete tabular data and continuous generative models. 3.4 Predictive and Generative Backbones CEL supports a wide range of model architectures via a unified PytorchBase interface, categorized into task-specific mixins. Discriminative Models. To ensure broad compatibility, the library im- plements ClassifierPytorchMixin andRegressionPytorchMixin . Implemented architectures range from standard Multi-Layer Perceptrons (MLP) and Logistic Regression to advanced Neural Oblivious Decision Ensembles (NODE), allowing researchers to test counterfactual generation against models of varying complex- ity and interpretability and easily extend their code with new implementations. Generative Models. Some of modern counterfactual methods and metrics rely on density estimation to ensure the plausibility of generated explanations. CEL integrates a GenerativePytorchMixin that standardizes likelihood estima- tion and sampling interfaces. The library provides implementations of state-of- the-art Normalizing Flows, including Masked Autoregressive Flow (MAF) [22], along with RealNVP and NICE. Additionally, it supports Kernel Density Esti- mation (KDE) and Gaussian Mixture Models (GMM) for baseline comparisons. These models allow users to estimate the likelihood of generated counterfactuals, a key component in assessing plausibility. 3.5 Evaluation Orchestration CEL introduces a MetricsOrchestrator which allows to use existing metrics or define custom ones. This component leverages a registry-based system to dynam- ically instantiate and compute metrics defined in the experiment configuration. The orchestrator handles input validation and compute a comprehensive suite of metrics covering validity, proximity, sparsity, and density-based plausibility scores. Metrics are detailed in the supplementary material. CEL modular workflow is demonstrated in Listing 1.1, which illustrates how the Dataset, Model, Method, and MetricsOrchestrator components interact. As shown, the framework abstracts the complexity of training and optimization, allowing users to generate and evaluate counterfactuals with minimal boilerplate code. 6O. Furman et al. 1 from cel.datasets import FileDataset 2 from cel.models.classifiers import MLPClassifier 3 from cel.cf_methods.local_methods import PPCEF 4 from cel.metrics import MetricsOrchestrator 5 6 # 1. Configuration -driven Data Loading 7 dataset = FileDataset(config_path="dataset.yaml") 8 9 # 2. Model Training (Unified PytorchBase Interface) 10 classifier = MLPClassifier (...) 11 classifier.fit(dataset.train_dataloader ()) 12 13 # 3. Counterfactual Generation (LocalMixin) 14 method = PPCEF(classifier , gen_model=flow_model , ...) 15 result = method.explain_dataloader(dataset.test_dataloader ()) 16 17 # 4. Standardized Evaluation 18 metrics = MetricsOrchestrator(conf_path="default.yaml") 19 scores = metrics.compute(result) Listing 1.1. Code snippet showing CEL modular workflow. 4 Benchmark A good counterfactual benchmark rests on three principles [23, 21]. First, a stan- dardized evaluation protocol—a fixed suite of diverse datasets, pre-defined predictive models, and consistent preprocessing pipelines—eliminates experi- mental variation as a confounding factor, a common issue in prior studies [25]. Second, comprehensiveness: broad coverage of methods spanning local, group- wise, and global paradigms, evaluated with multi-faceted metrics that capture validity, proximity, sparsity, and plausibility [8]. Third, reproducibility and extensibility: an open-source design that lets researchers replicate results and integrate their own methods, datasets, and metrics. CEL is built around these principles. 4.1 Datasets CEL includes 18 pre-configured datasets covering classification (13) and regres- sion (5) tasks, spanning domains such as finance and social sciences to test meth- ods across different distributions and complexities. Table 1 summarizes their characteristics. Consistent preprocessing is applied across all datasets. Continuous features are scaled to the [0, 1] range using Min-Max normalization. Categorical features are handled based on the method’s requirements For discriminative models and CF methods, categorical features are maintained in their original encoded for- mat. to facilitate training of density estimator we apply variational dequanti- zation to discrete features, mapping them into a continuous latent space while maintaining the discrete structure during inverse transformations. CEL: Counterfactual Explanations Library and Benchmark7 Table 1. Datasets included in CEL with feature types, distribution and references. DatasetTaskN Cat NumC Label Distribution Adult Census [2]Clf. 32,561 8420: 75.9%, 1: 24.1% Audit [29]Clf.775 0 2320: 60.6%, 1: 39.4% Bank Marketing [19] Clf. 40,004 9720: 88.3%, 1: 11.7% BlobsClf. 1,500 0230, 1, 2: 33.3% each Credit Default [35] Clf. 30,000 9 1420: 77.9%, 1: 22.1% Digits [28]Clf. 1,797 0 6410 Balanced (∼10% each) German Credit [9] Clf. 1,000 11720: 70.0%, 1: 30.0% GMC [24]Clf. 16,714 3720: 50.0%, 1: 50.0% HELOC [6]Clf. 10,459 0 2320: 52.2%, 1: 47.8% Law [17]Clf. 2,216 2320: 50.0%, 1: 50.0% Lending Club [10] Clf. 93,888 4820: 29.6%, 1: 70.4% MoonsClf. 1,024 0220: 50.0%, 1: 50.0% Wine [27]Clf.178 0 133 0: 33.1, 1: 39.9, 2: 27% SyntheticRegr. 1,000 02Cont.[0, 1] Concrete [34]Regr. 1,030 08Cont.[0, 1] Diabetes [5]Regr. 442 0 10Cont.[0, 1] Yacht [7]Regr. 308 06Cont.[0, 1] SCM20D [31]Regr. 8,966 0 61 Cont. (16)[0, 1] 16 4.2 Models CEL integrates distinct types of predictive and generative models implemented in PyTorch, which are essential for generating and evaluating counterfactual expla- nations. To ensure rigorous evaluation, CEL implements several predictors that we used in our benchmark. Classification Models: The library includes imple- mentations of Multi-Layer Perceptrons (MLP), Logistic Regression. Regression Models: For continuous target variables, CEL provides Linear Regression and MLP Regressor models. Density Estimator: To support plausibility evaluation and methods that rely on density estimation we utilise Masked Autoregressive Flow, accordingly to Wielopolski et al. [33]. 4.3 Methods The library implements 14 counterfactual explanation methods, categorized into local, global, and group-wise approaches. All methods adhere to a common in- terface, facilitating direct comparison. Local Methods Local methods generate explanations for individual instances. CEL includes classic optimization-based approaches such as Wachter et al. [32] and DiCE [20], which solve optimization problems to generate counterfactu- als. The library also supports several methods that focus on plausibility via density-based constraints and generative models, including the method of Artelt 8O. Furman et al. Table 2. Counterfactual explanation methods included in the CEL benchmark. Category Method Citation Local Wachter[32] Artelt[1] DiCE[20] CCHVAE [12] PPCEF[33] CEM[4] CEGP[16] CADEX [18] SACE[14] CEARM [30] Global AReS[26] GLOBE-CE [15] Group-wise GLANCE [13] T-CREx [3] et al. [1] (Artelt); CCHVAE [12], which leverages variational autoencoders for manifold-constrained generation; PPCEF [33], a flow-based approach with an explicit probabilistic formulation; and CEGP [16], which guides perturbations toward prototype-based counterfactuals. CEM [4] uses autoencoders to verify the plausibility of perturbed instances. The case-based method SACE [14] and CADEX [18] (constrained adversarial examples) are also included. Additionally, CEL implements CEARM [30], a Bayesian optimization-based method designed for regression models with a globally convergent search algorithm. Global Methods Global methods provide dataset-level insights. CEL imple- ments AReS [26], which generates global recourse rules, and GLOBE-CE [15], which learns global translation directions. Additionally, we include GLANCE [13] in our comparison by configuring it with a single group, effectively adapting this group-wise method for a global evaluation. These methods are essential for understanding high-level model behavior and identifying systemic biases. Group-wise Methods Bridging the gap between local and global, group-wise methods generate explanations for subgroups of data. CEL includes GLANCE [13] and TCREX [3], which allow users to inspect counterfactuals for specific clusters or demographic groups. 4.4 Metrics To ensure a holistic evaluation, CEL provides a comprehensive suite of metrics covering key counterfactual properties: Validity, Proximity, Sparsity, Plausibility. To ensure rigor evaluation we selected such metrics to evaluate each property: CEL: Counterfactual Explanations Library and Benchmark9 – Validity: Measures the success rate of generated counterfactuals in achiev- ing the target prediction. For continuous targets, validity is measured as the mean absolute error (MAE) between the predicted value of the counterfac- tual and the target value. – Proximity: Quantifies the distance between the original instance and the counterfactual using Euclidean distance for continuous features. For mixed data we apply combined Euclidean-Hamming distance, weighted proportion- ally by number of continuous and discrete features. – Sparsity: Evaluates the fraction of modified features, encouraging simpler explanations. – Plausibility: Assesses how well counterfactuals fit the data distribution using Log-Likelihood from density estimation model. We use MAF [22] for this benchmark, due to its superiority verified by Wielopolski et al. [33]. This evaluation ensures that methods are not just optimized for a single criterion (e.g., validity) but produce high-quality, actionable, and realistic explanations. We include a detailed explanation of all metrics in the supplementary material. 4.5 Experimental Setup Our evaluation systematically compares counterfactual methods across local, global, and group-wise paradigms under identical conditions, isolating method- ological differences from confounding factors. We evaluate all methods on 18 datasets of varying size, dimensionality, and feature composition, using two pre- dictive backbones per task type, and assess validity, proximity, sparsity, and plausibility to reveal both absolute performance and the trade-offs and failure modes that emerge across diverse settings. We adopt a standardized protocol using 5-fold cross-validation. For each fold we train the discriminative model on the training split, then fit a generative density estimator on data labeled by that model’s predictions, so conditional density is estimated against the appropriate class distribution. We model class- conditional densities with masked autoregressive flows (MAF) [22], which flexibly estimate complex distributions while preserving tractable likelihood computa- tion. For each dataset we define a target class and explain test instances whose pre- diction differs from it. For multiclass datasets we restrict evaluation to a binary setting between a selected class pair, aligning generation and density estimation with a well-defined decision boundary and enabling consistent comparison across datasets. For regression tasks, we set the problem to increasing the predicted value. The desired target value is defined as the original prediction increased by 20% of the overall range of the target variable. We measure validity as the Mean Absolute Error (MAE) between the counterfactual’s prediction and this target value. 10O. Furman et al. Fig. 2. Performance of local counterfactual methods across seven classification datasets. Rows represent datasets and columns represent evaluation metrics (Valid- ity, Euclidean-Hamming Distance, Sparsity, Log-Density, Computation Time). Each boxplot shows the distribution of metric values across methods: Artelt (blue), DiCE (light blue), CCHVAE (orange), PPCEF (light orange), CEGP (green), CADEX (red), and SACE (pink). 4.6 Results In this section, we present representative results from our benchmark. The com- plete evaluation across all datasets and models is provided in the supplemen- tary material. Our benchmark reports multiple metrics capturing different prop- erties of counterfactual explanations, including validity, proximity, and plausi- bility. This multi-dimensional evaluation enables analysis of trade-offs between competing objectives. For presentation, we selected a representative subset of datasets that covers the key data characteristics encountered in our full bench- mark: purely numerical features (Digits, Wine, HELOC, Moons), mixed cat- egorical and numerical features (Give Me Some Credit, Credit Default, Adult Census), and regression targets (Concrete, Diabetes, Synthetic). All reported re- sults are averaged across both Logistic Regression and MLP classifiers to ensure robustness across different model architectures. CEL: Counterfactual Explanations Library and Benchmark11 Fig. 3. Performance of counterfactual methods on regression tasks. Rows represent regression datasets and columns represent metrics (MAE, Sparsity, Log-Density, Com- putation Time). Methods: CEARM (blue), WACH (orange). Local Methods Figure 2 presents the performance of local methods across seven datasets. Most of the methods achieve perfect validity across all datasets, while CADEX shows moderate performance for mixed data and Artelt demon- strates the highest variance. The consistency of CCHVAE and PPCEF makes them suitable for applications where validity is the primary concern. CADEX generates the closest counterfactuals, followed by PPCEF and CCHVAE. SACE produces significantly more distant counterfactuals, suggesting counterfactual explanations with higher cost. DICE achieves the lowest sparsity, followed by CEGP and CADEX. PPCEF and CCHVAE demonstrate the high log-density, in- dicating its counterfactuals align well with the data manifold. DICE and CADEX show moderate plausibility with lower densities, suggesting their counterfactu- als may be less realistic. CCHVAE is the fastest method (11.70 s on average), followed by CADEX, while CEGP is the slowest (474.4 s), making it impractical for large-scale applications. Regression Methods For regression tasks (Figure 3), where validity is measured by Mean Absolute Error (MAE), WACH consistently outperforms CEARM. While both methods achieve comparable accuracy (MAE∼0.1) and high spar- sity, WACH produces significantly more plausible counterfactuals (positive log- density of 8.50) and is nearly 9× faster (7.4s vs. 69.5s). This makes WACH the clear preferred method for regression. 12O. Furman et al. Fig. 4. Performance of global counterfactual methods. Rows represent datasets and columns represent metrics. Methods: AReS (pink), GLOBE-CE (dark brown), Glob- alGLANCE (brown). Global Methods Global methods provide dataset-level insights by learning transformations applicable to entire classes. Figure 4 compares AReS, GLOBE- CE, and GLANCE (with 1 group), aggregated across Logistic Regression and MLP models. GLOBE-CE and GLANCE achieve perfect or near-perfect validity while AReS shows only moderate success rates: GLOBE-CE and GLANCE ag- gregate explanations by direction with varying magnitude, whereas AReS seeks a single universal fixed change. AReS generates closer counterfactuals when suc- cessful but modifies more features, while GLANCE produces more distant coun- terfactuals with higher sparsity. AReS shows the lowest log-density, though all methods exhibit high variance, reflecting the difficulty of maintaining distribu- tional alignment across diverse instances with a single global transformation. CEL: Counterfactual Explanations Library and Benchmark13 Fig. 5. Performance of group-wise counterfactual methods. Rows represent datasets and columns represent metrics. Methods: GLANCE (dark pink), T-CREx (light pink). Group-wise Methods Group-wise methods bridge the gap between local and global approaches by generating explanations for subgroups. Figure 5 presents results for GLANCE and T-CREx, averaged across Logistic Regression and MLP classifiers. This comparison reveals how these methods balance the personaliza- tion of local approaches with the efficiency of global methods. GLANCE consistently achieves higher validity, but T-CREx, when applica- ble, produces closer and more plausible counterfactuals, however with very low success rates. This highlights a clear trade-off between generating effective and minimally disruptive explanations for subgroups. Compared to global methods like AReS, group-wise approaches mitigate va- lidity collapse. Global methods attempt to find a single recourse rule for the entire population, often resulting in poor generalization (e.g., AReS yields only 14O. Furman et al. 0.17 validity on Adult). By constraining the scope to subgroups, GLANCE recov- ers the high validity typically associated with local methods while maintaining the interpretability of broad rules. Furthermore, subgroup targeting allows for significantly higher sparsity than global baselines, as the recourse rule need not accommodate the variance of the entire dataset. Remarkably, GLANCE matches the high validity of local methods despite this harder constraint, though local methods retain superior manifold adherence (lower LOF): group-wise shifts inevitably push peripheral group members into lower-density regions to satisfy the validity constraint for the cluster center. Overall, group-wise counterfactuals offer an effective semi-global explanation, avoiding the noise of local methods and the ineffectiveness of purely global rules. 5 Conclusions We introduced CEL, a unified benchmark and open-source library for standard- ized evaluation of counterfactual explanation methods. By harmonizing datasets, predictive backbones, preprocessing, and metrics, CEL enables fair and repro- ducible comparison across local, global, and group-wise paradigms, a significant step beyond the fragmented practices common in the field. Our benchmark of 14 methods on 18 datasets reveals that no single method uniformly dominates; each excels on some quality dimensions at the expense of others. Among local methods, PPCEF and CCHVAE achieve near-perfect valid- ity with strong plausibility, while CADEX and Wachter produce closer counter- factuals that may fall in low-density regions; DiCE and CEGP yield the spars- est explanations, and SACE delivers the strongest manifold adherence at higher proximity, with computational cost ranging from under 1s (Wachter) to over 470s (CEGP). At the global level, GLOBE-CE achieves minimal, high-validity pertur- bations whereas AReS frequently fails. Group-wise GLANCE recovers local-level validity with interpretable subgroup rules, while TCREx favors minimal pertur- bation over effectiveness. For regression, Wachter dominates across all metrics. These findings underscore that method selection should be guided by application- specific requirements. CEL is designed to be extensible, and we encourage the community to contribute new methods, datasets, and metrics. Promising future directions include evaluating counterfactuals for fairness and robustness, devel- oping user-centric trade-off optimization, and extending to more complex data modalities. We believe CEL will help accelerate progress toward more reliable and trustworthy explainable AI. References 1. Artelt, A., Hammer, B.: Convex density constraints for computing plausible coun- terfactual explanations. In: International conference on artificial neural networks. p. 353–365. Springer (2020) 2. Becker, B., Kohavi, R.: Adult Census Income. UCI Machine Learning Repository (1996). https://doi.org/10.24432/C5XW20 CEL: Counterfactual Explanations Library and Benchmark15 3. Bewley, T., Amoukou, S.I., Mishra, S., Magazzeni, D., Veloso, M.: Counterfactual metarules for local and global recourse. In: Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net (2024), https://openreview.net/forum?id=Ad9msn1SKC 4. Dhurandhar, A., Chen, P., Luss, R., Tu, C., Ting, P., Shanmugam, K., Das, P.: Ex- planations based on the missing: Towards contrastive explanations with pertinent negatives. In: Advances in Neural Information Processing Systems 31: Annual Con- ference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada. p. 590–601 (2018) 5. Efron, B., Hastie, T., Johnstone, I., Tibshirani, R.: Least angle regression” (with discussions). The Annals of Statistics 32 (01 2004) 6. FICO: FICO Explainable Machine Learning Challenge. https://community.fico. com/s/explainable-machine-learning-challenge (2018) 7. Gerritsma, J., Onnink, R., Versluis, A.: Yacht Hydrodynamics. UCI Machine Learning Repository (1981), DOI: https://doi.org/10.24432/C5XG7R 8. Guidotti, R.: Counterfactual explanations and how to find them: literature re- view and benchmarking. Data Mining and Knowledge Discovery 38(5), 2770–2824 (2024) 9. Hofmann, H.: Statlog (German Credit Data). UCI Machine Learning Repository (1994), DOI: https://doi.org/10.24432/C5NC77 10. Jagtiani, J., Lemieux, C.: The roles of alternative data and machine learning in fintech lending: evidence from the lendingclub consumer platform. Financial Man- agement 48(4), 1009–1029 (2019) 11. Jiang, J., Marzari, L., Purohit, A., Leofante, F.: Robustx: Robust counterfactual explanations made easy. arXiv preprint arXiv:2502.13751 (2025) 12. Karimi, A.H., Barthe, G., Balle, B., Valera, I.: Model-agnostic counterfactual ex- planations for consequential decisions. In: International conference on artificial intelligence and statistics. p. 895–905. PMLR (2020) 13. Kavouras, L., Psaroudaki, E., Tsopelas, K., Rontogiannis, D., Theologitis, N., Sacharidis, D., Giannopoulos, G., Tomaras, D., Markou, K., Gunopulos, D., et al.: Glance: Global actions in a nutshell for counterfactual explainability. arXiv preprint arXiv:2405.18921 (2024) 14. Keane, M.T., Smyth, B.: Good counterfactuals and where to find them: A case- based technique for generating counterfactuals for explainable ai (xai). In: Inter- national Conference on Case-Based Reasoning. p. 163–178. Springer (2020) 15. Ley, D., Mishra, S., Magazzeni, D.: Globe-ce: A translation based approach for global counterfactual explanations. In: International conference on machine learn- ing. p. 19315–19342. PMLR (2023) 16. Looveren, A.V., Klaise, J.: Interpretable counterfactual explanations guided by prototypes. In: Machine Learning and Knowledge Discovery in Databases. Research Track - European Conference, ECML PKDD 2021, Bilbao, Spain, September 13- 17, 2021, Proceedings, Part I. Lecture Notes in Computer Science, vol. 12976, p. 650–665. Springer (2021) 17. (LSAC), L.S.A.C.: Law School Admissions Bar Passage. Kaggle, https://w. kaggle.com/datasets/danofer/law-school-admissions-bar-passage 18. Moore, J., Hammerla, N., Watkins, C.: Explaining deep learning models with con- strained adversarial examples. In: Pacific Rim international conference on artificial intelligence. p. 43–56. Springer (2019) 19. Moro, S., Cortez, P., Rita, P.: A data-driven approach to predict the success of bank telemarketing. Decision Support Systems 62, 22–31 (2014) 16O. Furman et al. 20. Mothilal, R.K., Sharma, A., Tan, C.: Explaining machine learning classifiers through diverse counterfactual explanations. In: Proceedings of the 2020 conference on fairness, accountability, and transparency. p. 607–617 (2020) 21. de Oliveira, R.M.B., Martens, D.: A framework and benchmarking study for coun- terfactual generating methods on tabular data. Applied Sciences 11(16), 7274 (2021) 22. Papamakarios, G., Pavlakou, T., Murray, I.: Masked autoregressive flow for density estimation. Advances in neural information processing systems 30 (2017) 23. Pawelczyk, M., Bielawski, S., Heuvel, J.v.d., Richter, T., Kasneci, G.: Carla: a python library to benchmark algorithmic recourse and counterfactual explanation algorithms. arXiv preprint arXiv:2108.00783 (2021) 24. Pawelczyk, M., Broelemann, K., Kasneci, G.: Learning model-agnostic counterfac- tual explanations for tabular data. In: Proceedings of the web conference 2020. p. 3126–3132 (2020) 25. Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d’Alché Buc, F., Fox, E., Larochelle, H.: Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program). Journal of ma- chine learning research 22(164), 1–20 (2021) 26. Rawal, K., Lakkaraju, H.: Beyond individualized recourse: Interpretable and in- teractive summaries of actionable recourses. Advances in Neural Information Pro- cessing Systems 33, 12187–12198 (2020) 27. Repository, U.M.L.: Wine Data Set. https://archive.ics.uci.edu/dataset/ 109/wine (1991) 28. Repository, U.M.L.: Optical Recognition of Handwritten Digits. https://archive. ics.uci.edu/dataset/80/optical+recognition+of+handwritten+digits (1998) 29. Repository, U.M.L.: Audit Data Set. https://archive.ics.uci.edu/dataset/ 475/audit+data (2016) 30. Spooner, T., Dervovic, D., Long, J., Shepard, J., Chen, J., Magazzeni, D.: Counterfactual explanations for arbitrary regression models. arXiv preprint arXiv:2106.15212 (2021) 31. Spyromitros-Xioufis, E., Tsoumakas, G., Groves, W., Vlahavas, I.: Multi-target regression via input space expansion: treating targets as inputs. Machine Learning 104(1), 55–98 (July 2016). https://doi.org/10.1007/s10994-016-5546-z, https:// doi.org/10.1007/s10994-016-5546-z 32. Wachter, S., Mittelstadt, B., Russell, C.: Counterfactual explanations without opening the black box: Automated decisions and the gdpr. Harv. JL & Tech. 31, 841 (2017) 33. Wielopolski, P., Furman, O., Stefanowski, J., Zięba, M.: Probabilistically plausible counterfactual explanations with normalizing flows. arXiv preprint arXiv:2405.17640 (2024) 34. Yeh, I.C.: Concrete Compressive Strength. UCI Machine Learning Repository (1998), DOI: https://doi.org/10.24432/C5PK67 35. Yeh, I.C.: Default of Credit Card Clients. UCI Machine Learning Repository (2009), DOI: https://doi.org/10.24432/C55S3H