Paper deep dive
Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification
Guillermo Rubiños Rodríguez, Martín Ottavianelli, Mateo Alonso, Gonzalo Blázquez Gil, Boris-Stephan Rauchmann, Pablo Díez-Valle, Sergio Altares-López
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/28/2026, 3:51:12 AM
Summary
This study investigates the effectiveness of Hybrid Quantum-inspired Convolutional Neural Networks (HQiCNNs) compared to classical CNNs for medical image classification. Using classically-emulated quantum circuits (EQCs) within a parameter-matched architecture, the authors evaluate performance on retinal OCT and brain MRI datasets. Results indicate that HQiCNNs perform competitively, particularly in intermediate-data regimes, while classical CNNs excel with larger datasets. The study also introduces SHAP-based explainability tools (|SHAP|IoU and EMD_pos) to demonstrate that both models attend to anatomically plausible regions.
Entities (9)
Relation Signals (8)
HQiCNN → comparedwith → CNN
confidence 98% · we present a systematic study of the effectiveness of a Hybrid Quantum-inspired Convolutional Neural Network (HQiCNN) compared with a parameter-matched classical Convolutional Neural Network (CNN)
HQiCNN → usescomponent → EQC
confidence 95% · The quantum-inspired branch replaces the dense neural layer of the common backbone with the parameterized emulated quantum circuit (EQC)
HQiCNN → evaluatedon → OCT
confidence 92% · We evaluate both architectures on two clinically relevant imaging tasks, retinal optical coherence tomography (OCT) classification
HQiCNN → evaluatedon → OASIS-1
confidence 92% · Alzheimer’s disease staging from structural MRI (OASIS-1)
EMD_pos → derivedfrom → SHAP
confidence 90% · the second metric,EMD pos , quantitatively compares the similarity between the positive SHAP distributions
|SHAP|IoU → derivedfrom → SHAP
confidence 90% · The first tool is|SHAP|IoU, which plots the intersection areas between the top 10% absolute SHAP maps of the models
SHAP → usedfor → Explainability
confidence 90% · We therefore adopt SHAP as the basis for our interpretability analysis
Barren Plateaus → affects →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, while the scalability of quantum circuits leads to trainability issues. In this work, we investigate whether small, classically-emulated quantum circuit components can play a meaningful role within complex models, offering an alternative to purely classical convolutional architectures. To this end, we present a systematic study of the effectiveness of a Hybrid Quantum-inspired Convolutional Neural Network (HQiCNN) compared with a parameter-matched classical Convolutional Neural Network (CNN) that differs only in an intermediate dense neural layer. Both models are evaluated on two real-world medical datasets while systematically varying the different hyperparameters, ensuring a fair model comparison that is both dataset and hyperparameter independent. The results show that no architecture consistently dominates the other: the HQiCNN achieves its largest gains in intermediate-data regimes, whereas the CNN reaches the highest accuracies for the largest training sets in both datasets. Furthermore, removing entanglement produces comparable performance while enabling substantially better scalability of quantum simulations, and richer observable sets become beneficial only when sufficient training data are available. Finally, we propose two SHAP-based explainability tools for comparing the predictions between both models, $|SHAP|$IoU and $EMD_{pos}$ metric, to demonstrate that both architectures consistently attend to anatomically plausible regions. Thus, we provide a comprehensive benchmark showing that, under certain conditions, hybrid quantum-inspired models are an alternative that can offer benefits in practical tasks such as medical image classification.
Tags
Links
- Source: https://arxiv.org/abs/2607.21186v1
- Canonical: https://arxiv.org/abs/2607.21186v1
Trouble viewing inline? Open PDF directly →
Full Text
70,137 characters extracted from source content.
Expand or collapse full text
Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification Guillermo Rubi ̃ nos Rodr ́ıguez 1,2 , Mart ́ın Ottavianelli 1 , Mateo Alonso 1 , Gonzalo Bl ́ azquez Gil 1 , Boris-Stephan Rauchmann 3,4 , Pablo D ́ıez-Valle 1,+ , and Sergio Altares-L ́ opez 3,4,+,* 1 Instituto Tecnol ́ ogico de Galicia, Cant ́ on Grande 9, Planta 3, 15003 A Coru ̃ na, Spain 2 Centro Singular de Investigaci ́ on en Tecnolox ́ ıas Intelixentes (CiTIUS), Universidade de Santiago de Compostela, Santiago de Compostela, 15782, Spain 3 Department of Neuroradiology, LMU University Hospital, Ludwig Maximilian University of Munich, Marchioninistraße 15, 81377 Munich, Germany 4 Department of Psychiatry and Psychotherapy, LMU University Hospital, LMU Munich, Nussbaumstraße 7, 80336 Munich, Germany + These authors contributed equally to this work * slopez@med.lmu.de ABSTRACT Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks (HQCNN) as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, such as measurement-based gradients, while the scalability of quantum circuits in the number of qubits leads to trainability issues such as barren plateaus. On top of that, the success of HQCNNs has been associated with quantum architectures that can be efficiently simulated on classical computers. In this work, we treat this last point not as an issue but as an opportunity, and investigate whether small, classically-emulated quantum circuit components can play a meaningful role within complex models, offering an alternative to purely classical convolutional architectures. To this end, we present a systematic study of the effectiveness of a Hybrid Quantum-inspired Convolutional Neural Network (HQiCNN) compared with a parameter-matched classical Convolutional Neural Network (CNN) that differs only in an intermediate dense neural layer. Both models are evaluated on real-world medical datasets (retinal OCT and brain MRI for dementia staging) while systematically varying the training set size, learning rate, batch size, circuit entanglement, and observable set; ensuring a fair model comparison that is both dataset and hyperparameter independent. The results show that no architecture consistently dominates the other: the HQiCNN achieves its largest gains in intermediate-data regimes, whereas the CNN reaches the highest accuracies for the largest training sets in both datasets. We further show that removing entanglement produces comparable performance while enabling substantially better scalability of quantum simulations, and that richer observable sets become beneficial only when sufficient training data are available. Finally, we propose two SHAP-based explainability tools for comparing the predictions between both models,|SHAP|IoU and EMD pos metric, to demonstrate that both architectures consistently attend to anatomically plausible regions, while the HQiCNN learns more localized disease-relevant features under limited-data conditions. Thus, we provide a comprehensive benchmark showing that, under certain conditions, hybrid quantum-inspired models are an alternative and can offer benefits in practical tasks such as medical image classification. 1 Introduction Quantum machine learning (QML) 1 has emerged as a promising paradigm at the intersection of quantum computing and artificial intelligence (AI), with growing interest in its application to domains where classical models face persistent limitations in accuracy, robustness, or interpretability. Among these domains, medical image diagnosis stands out as a particularly demanding use case: clinical decision support systems must not only achieve high predictive performance but also offer a degree of transparency that allows practitioners to trust and validate model outputs. Variational quantum algorithms (VQAs) 2 , and in particular parameterized quantum circuits (PQCs) 3 embedded within hybrid quantum-classical architectures, have been proposed as a means of introducing nonlinear, high-dimensional feature transformations that may complement existing machine learning (ML) approaches, including classical deep learning (DL) models 4–6 . Specifically in the context of QML for image classification, there have been several proposals such as introducing variational PQCs after RESNET18 frameworks for transfer learning 7 , replacing the classical convolutional layer by PQCs 6 and substituting arXiv:2607.21186v1 [quant-ph] 23 Jul 2026 the classical dense layers inside convolutional architectures by PQCs 8 . However, most of these studies are constrained either by effective classical simulability 9 or by well-documented challenges like trainability issues related to the limited qubit count and barren plateaus, which have been shown to severely limit the scalability of PQC-based models 10 . Furthermore, the results of the previous image-classification studies do not address the recently raised important questions regarding whether reported quantum advantages in QML are genuine or artifacts of insufficiently controlled experimental comparisons with classical baselines, and if quantum properties are even necessary for the reported advantages 11 . At the same time, the broader ML literature applied to medicine has increasingly emphasized the need for models that are not only accurate but also simple and interpretable. While architectures such as Vision Transformers (ViT) 12 have achieved state-of-the-art performance on many imaging tasks, their complexity and limited transparency make them less suitable for clinical contexts, where explainability tools such as SHAP (SHapley Additive exPlanations) are important for justifying model predictions and building trust among medical practitioners 13 . This motivates the exploration of hybrid quantum-classical convolutional neural networks (HQCNNs) 5 as an alternative that retains the explainable AI and data efficiency of classical convolutional backbones while introducing a compact, parameter-matched quantum layer whose contribution to performance can be isolated and rigorously assessed. In this context, we propose a new framework that integrates classical convolutional neural networks (CNNs) 14 with small- size classically emulated PQCs (EQCs), forming the Hybrid Quantum-inspired Convolutional Neural Network (HQiCNN). Our approach enables a fair, architecturally controlled comparison to investigate the conditions under which the integration of an emulated quantum layer provides measurable advantages 11 . To this end, we design two models that share an identical convolutional backbone and differ only in the intermediate dense representation-learning layer, thereby attributing any observed performance differences specifically to the quantum-inspired component rather than to differences in model capacity. We evaluate both architectures on two clinically relevant imaging tasks, retinal optical coherence tomography (OCT) 15 classification and Alzheimer’s disease staging from structural MRI (OASIS-1) 16 , across a broad hyperparameter space including training set size, learning rate, batch size, circuit entanglement, and observable set size. In addition, instead of directly comparing the model-dependent SHAP maps, we derive two SHAP-based methods that allow us to identify whether the models rely on comparable or divergent evidence when making predictions. The first tool is|SHAP|IoU, which plots the intersection areas between the top 10% absolute SHAP maps of the models, while the second metric,EMD pos , quantitatively compares the similarity between the positive SHAP distributions. The combination of both tools allows us to observe that both models rely on similar relevant areas for a given image, while also showing consistent positive SHAP distributions across the whole test set. The main contributions of this article are: (i) the exploration and proposal of emulated quantum circuits embedded within classical CNN architectures; (i) a fair and comprehensive hyperparameter analysis to ensure an unbiased comparison between classical and quantum-emulated models; (i) a systematic study of the impact of observables and entanglement on the performance of the quantum-emulated model; (iv) the evaluation of the robustness of this comparison across two challenging medical image classification datasets; and (v) the design and implementation of SHAP-based explainability methods to compare the models’ predictions and assess the anatomical plausibility of the identified regions, as a first step towards evaluating their potential clinical relevance. This study is organized as follows. Section 2 introduces the fundamentals of quantum computing and hybrid parametric quantum-classical circuits. Section 3 describes the proposed models to be compared and the experimental methodology, including the description of the experiments conducted and the SHAP-based explainability tools used to compare results across models. Section 4 presents the results and discussion, covering the effects of training set size, learning rate, batch size, entanglement, and observable set size, together with the explainability analysis based on SHAP attributions. Finally, Section 5 summarizes the main conclusions of this study and outlines directions for future work. 2 Fundamentals of Quantum Computing Quantum computing is based on the principles of quantum mechanics, with the qubit as its fundamental unit. A qubit is described by a vector in a two-dimensional complex Hilbert space and can exist in a linear combination of the basis states|0⟩ and|1⟩, a property known as superposition: |ψ⟩ = α|0⟩+ β|1⟩ = α β , whereαandβare complex amplitudes satisfying|α| 2 +|β| 2 = 1 17 . Entanglement is a quantum correlation between subsystems such that the joint state cannot be factorized into a tensor product of individual states. Entangled qubits allow quantum circuits to encode joint probability distributions and correlations that are difficult to represent classically. Quantum computations are implemented using quantum gates, which are unitary operators acting on one or more qubits. 2/17 Single-qubit rotation gates around the Bloch sphere axes are commonly used in hybrid models: R X (θ) = cos θ 2 −i sin θ 2 −i sin θ 2 cos θ 2 , R Y (θ) = cos θ 2 − sin θ 2 sin θ 2 cos θ 2 , R Z (θ) = e −iθ/2 0 0 e iθ/2 .(1) These gates rotate a qubit state around the corresponding axis of the Bloch sphere. Multi-qubit gates, such as the CNOT, create entanglement by flipping a target qubit conditioned on the control qubit. In digital quantum computing, the most common observables used are different tensor product combinations of Pauli matrices 17 : X = 0 1 1 0 ,Y = 0 −i i0 ,Z = 10 0 −1 .(2) The sequencing of single-qubit rotations with tunable parametersθdefines a PQC, as shown in Figure 1b. PQCs act as trainable nonlinear transformations of input data in hybrid quantum-classical models and form the basic building block of so-called quantum neural networks (QNNs) 18 . A high-dimensional representation of the input data is obtained after the parameterized evolution, from which classical features are extracted by estimating the expectation values of a selected set of observables through repeated sampling (shot-based measurement) of the quantum state; these classical outputs can then be processed classically to predict or classify. The parametersθare optimized using classical gradient-based methods, with the gradient conventionally computed via the parameter-shift rule, which yields exact analytical gradients by evaluating the circuit at shifted parameter values and is compatible with optimizers commonly used in deep learning. Notable limitations include barren plateaus 10 , which hinder training as gradients vanish exponentially with circuit size 19 , the noise inherent to current quantum hardware 20 , and the finite sampling limitation. 3 Methods In this section, we introduce the proposed models that are to be compared on health-related image classification tasks. Then, we describe the experimental setup that allows us to fairly compare the algorithms. Finally, we introduce the two SHAP-based tools derived to compare and explain the predictions provided by both models. 3.1 Quantum-inspired and classical arquitectures The models we compare in this article are schematized in Figure 1. Both models share a common classical backbone that comprises three main components. First, a convolutional block composed of three convolutional layers, each incorporating ReLU activation, a kernel size of three, batch normalization, and max pooling, transforms the three-dimensional RGB input into a 128-dimensional feature vector. This vector is then compressed by a linear reduction layer, a linear layer without an activation function that reduces the 128-dimensional vector into a 4-dimensional vector, which can be encoded into either the quantum circuit or the linear layer of the CNN branch. Finally, an output layer, also a linear layer without an activation function, takes the 12-dimensional vector from the previous layer and maps it to a vector whose size corresponds to the number of labels. Note that the proposed architectures differ only in a single layer, allowing for a fair and accurate comparison of the impact of the emulated quantum circuit relative to a classical ReLU-activated neural network. We describe the different components in the following subsections. 3.1.1 Dense Layer (CNN branch) The 4-dimensional feature vector produced by the Linear Reduction layer is subsequently processed by a dense layer with ReLU activation, which transforms it into a 12-dimensional feature representation. This combination of convolutional and linear layers follows the standard architecture of deep convolutional models, providing a simple and effective transition between learned feature extraction and subsequent processing stages. We opt for a convolutional architecture over more complex alternatives such as ViT for several reasons. CNN-based models remain simpler and more interpretable than transformer-based architectures, a property of particular relevance in medical applications, where model transparency and the ability to justify predictions are often as important as raw performance. They also typically require fewer data and computational resources to train, resulting in faster convergence, which is advantageous given the moderate size of the dataset used in this study. Most importantly, the CNN branch is designed to serve as a fair classical baseline against which the quantum branch can be meaningfully compared, since both architectures share the same backbone and a comparable number of trainable parameters, differing only in the layer under study. Introducing a substantially more complex model, such as a ViT, would compromise this parameter-matched, architecture-controlled comparison, making it difficult to attribute performance differences specifically to the quantum component rather than to differences in model capacity or design. 3/17 Figure 1. Schematic representation of the model to compare. Both share a common architecture and only differ in the layer right before the output. The HQiCNN uses a preprocessing scaled sigmoid transformation to map the input data into the quantum circuit. (b) Parameterized EQC used within the HQiCNN algorithm. 3.1.2 Emulated Quantum Circuit (HQiCNN branch) The quantum-inspired branch replaces the dense neural layer of the common backbone with the parameterized emulated quantum circuit (EQC) shown in Figure 1b. The distinctive feature of emulated quantum circuits is that, whether due to their low correlations or their small size, they can be efficiently computed on classical devices. Since we are considering quantum architectures with only four qubits, they can be computed easily by quantum hardware emulation within a short time. The four-dimensional feature vector produced by the Linear Reduction layer is first encoded into the quantum state through angle encoding, usingR Y rotations applied to each qubit. A variational layer then applies trainableR X andR Z rotations, followed by a linear entanglement layer implemented with CNOT gates. Finally, the three local Pauli observables (X,Y, andZ) are measured on each of the four qubits, yielding a 12-dimensional output vector that is passed to the output classical layer. Since parameterized quantum gates implement rotations and therefore require input angles bounded within[0, 2π), we investigate how the choice of preprocessing applied to the four-dimensional vector v prior to encoding affects downstream performance, comparing three distinct strategies. The first employs a scaled sigmoid transformation,2π· sigmoid(v), which smoothly maps arbitrary real-valued inputs into the required interval while compressing extreme values toward its boundaries. The second applies a modulo transformation,remainder(v, 2π), which wraps the input periodically into the valid range without altering its relative scale, thereby preserving the original distances between values more faithfully than the sigmoid mapping. The third condition serves as a baseline in which no transformation is applied and the raw vector v is encoded directly, allowing us to assess whether explicit preprocessing offers any measurable advantage over leaving the input unconstrained. A study comparing these three options showed that the scaled sigmoid transformation consistently outperforms the other ones, as a result, we select this transformation for the rest of the study. Finally, we highlight that the proposed quantum-inspired implementation allows the hybrid model to compute gradients analytically via backpropagation, circumventing the need to estimate the quantum gradient through the parameter shift rule 21 . The parameter shift rule requires two forward passes per parameter to estimate the gradient, resulting in a total of2pforward passes for a quantum circuit ofpparameters. This quantum-inspired architecture can implement backpropagation to calculate all gradients in parallel using one single backward pass, effectively reducing the number of passes required to calculate the gradient from O(p) to one. 4/17 3.2 Experimental setup In order to ensure a faithful comparison between the results obtained from both architectures and to examine how the quantum layer affects the results, we study the models’ hyperparameters as explained in Ref. 11 . We compare the test accuracy of the classical and quantum emulated models for the OASIS and OCT datasets (See Section 4.1), ensuring that the effects are not dataset-specific. Furthermore, the role of the hybrid quantum layer is studied in depth by testing if the entanglement enhances performance. These tests are repeated for 10 randomly selected seeds and for different training sizes, providing a statistically robust evaluation of the results for different training sizes. Table 1 displays the studied hyperparameters for the OCT dataset. Since the balanced OASIS dataset was considerably smaller compared with the OCT dataset, we used the OASIS dataset experiment to help us determine the robustness of the model, assessing whether the observed trends between both models generalize across different datasets. Note that this robustness test is realized under a simplified 2D setting (see Section 4.1) and it is not intended to provide clinical conclusions, but rather serves as a robustness check, evaluating whether the behavior of both architectures generalizes to a distinct imaging modality and classification task. Table 1. The hyperparameter search space used for both models evaluated on the OCT dataset is described. The impact of the learning rate and training size combination on model performance is studied, resulting in 270 executions. Subsequently, the best learning rate for each model at a training size of 1000 samples is used to investigate the effect of batch size, resulting in 30 additional executions. Finally, the quantum properties of the HQiCNN are further analyzed by evaluating the importance of entanglement and the observable set across different training sizes, adding a total of 340 executions. HyperparameterValues# Values Training Size200, 400, 800, 1000, 1200, 1600, 2000, 12000, 300009 Learning Rate10 −2 , 10 −3 , 10 −4 3 Batch Size8, 16, 323 Quantum CircuitEntangled, Not Entangled2 Observable SetO 1 ,O 2 ,O 3 3 Seeds10 randomly selected seeds10 Total configurations tested9× 3× 10+ 3× 10+ 2× 8× 10+ 3× 6× 10 = 640 The remaining hyperparameters, those not included in Table 1, related to the training of the models were kept constant across all realized experiments. We highlight the use of GradScaler, which enables the optimizer to handle the quantum and classical gradients in a unified manner. Given the multiclass nature of the dataset, cross-entropy is selected as the loss function. This choice also justifies the absence of a softmax activation in the output layer, since the cross-entropy loss operates directly on unnormalized logits. To constrain the training and make the hyperparameter exploration feasible, we define 100 epochs as the maximum number of training epochs in all cases, complemented by an early stopping criterion with a patience hyperparameter of eight epochs, which stops training when the validation loss does not decrease within this interval. Additionally, a learning rate scheduler is configured with a patience of three epochs and a reduction factor of 0.3, allowing the learning rate to adapt to training plateaus. Both models are implemented using PyTorch 22 , including the quantum component, which is simulated through native PyTorch functions rather than a dedicated quantum computing framework. 3.3 Explainability tools Interpreting the results obtained for ML algorithms comprised of hidden layers and large trainable parameter counts is a difficult task. Several families of explainability techniques have been proposed to address this challenge. Gradient-based visualization methods such as Grad-CAM 23 and its refinement Grad-CAM++ 24 , generate class-discriminative localization maps by leveraging the gradients flowing into the final convolutional layers, offering an intuitive visualization of the image regions driving a CNNs prediction; however, these approaches are architecture-dependent, typically restricted to convolutional backbones, and provide coarser, lower-resolution attributions than perturbation-based alternatives. Attention-based methods, applicable to transformer architectures, and causal interpretability approaches, which attempt to move beyond mere correlation by estimating the causal effect of specific input features or latent factors on the model output 25 , represent complementary directions, though they generally demand additional modeling assumptions or architecture-specific access. In contrast, perturbation-based, model-agnostic tools such as SHAP (SHapley Additive exPlanations) 13 or LIME (Local Interpretable Model-Agnostic Explanations) 26 are widely adopted in the field of AI applied to medicine, as they can be applied uniformly across the different model architectures compared in this study and provide evidence that ML models are attending to anatomically relevant areas, rather than relying on spurious or irrelevant features 27 . We therefore adopt SHAP as the basis 5/17 for our interpretability analysis and propose two complementary techniques built upon it: the first is a tool that compares, for a single image, the SHAP areas that contribute the most to the prediction of both models; the second is a metric that quantifies, across the whole test set, the similarity between the positive SHAP maps generated by both models. 3.3.1 Absolute SHAP Intersection of Union (|SHAP|IoU) In order to effectively compare and validate that both models study the medical areas of interest for our images, we derive the |SHAP|IoU. This tool takes the top10%absolute SHAP values of the image distribution pixels for each model, defining the corresponding regions of interest R A and R B . The extracted regions are then plotted over the original image, highlighting the intersection and single model interest areas. It allows us to identify which areas the models intersect, proving similarity between them, and in which areas they differ, understanding interpretable image-based medical explanations on why one model performs better than the other one. 3.3.2 Positive SHAP Earth Mover’s Distance (EMD pos ) The previous method serves as an interpretability tool for a single image at a time. With the aim of correctly comparing the similarity between the positive SHAP distributions 13 (areas that improve the correct prediction of the models) across several images, we introduce the metricEMD pos . This metric is a generalization to 2 dimensions of the Wasserstein metric for 1 dimension probabilistic distributions. The first step is to normalize the positive values of the SHAP for each considered image, obtaining a 2 dimensional probabilistic distribution. Then we numerically calculate the full pairwise Euclidean cost matrix over all pixel coordinates: EMD pos = min γ∈ ∏ (ω HQiCNN |ω CNN ) ∑ i,j γ i,j ||p i − p j || 2 .(3) Theω HQiCNN|CNN represents the normalized positive SHAP distribution for our models, andp i,j are the normalized coordinates of the pixelsp = y 64 , x 64 , andγ i,j specifies the amount amount of probability mass transported from location i to target location j. In our considered datasets, the images that we use are transformed to tensors of resolution64× 64pixels to input into the model, so we explore 4096 grid points. This proposed metric has a lower and an upper bound. The lower bound indicates identical distributions while the upper bound indicates non correlated or opposite pixel distribution. The lower bound can be derived directly as0, while the upper bound depends on the resolution of the image. We derive the maximum bound by considering two distributions, one centered in the corner p = (0, 0) and the other distribution in the opposite corner position p = (63/64, 63/64). D = s 63 64 2 + 63 64 2 = 63 64 √ 2≈ 1.3921.(4) Therefore, 0 ≤ EMD pos ≤ 63 64 √ 2≈ 1.3921.(5) 4 Results and Discussion This section collects the discussion over the obtained results for the considered experiments that allow us to fairly compare the proposed architectures. First, we explain the considered datasets and their characteristics. Then, we discuss the hyperparameter results for the more balanced OCT dataset. Afterwards, we conduct a robustness test using the dementia dataset. Finally, a comparative study of explainability is performed for the OCT dataset. 4.1 Datasets The dataset used in this study consists of retinal optical coherence tomography (OCT) images, a widely used imaging modality that provides high-resolution cross-sectional views of the retina and is routinely employed in ophthalmic diagnosis 28 . It includes 84,495 JPEG images organized into two standard splits (training and test) and four clinically relevant classes: NORMAL, CNV (choroidal neovascularization), DME (diabetic macular edema), and DRUSEN, as can be shown in Figures 2a, 2b, 2c, 2d. Each image is labeled according to its diagnostic category and associated with an anonymized patient identifier and acquisition index to ensure traceability without compromising patient privacy. Data were collected from multiple international clinical centers between 2013 and 2017 using Spectralis OCT devices (Heidelberg Engineering) 15 . To ensure label quality, a multi-stage review process was applied. Images first underwent basic quality control to remove those with severe artifacts, followed by independent grading by multiple ophthalmologists. Final labels were confirmed by senior retinal specialists with extensive 6/17 clinical experience. In addition, a subset of validation images was independently re-annotated to further assess and mitigate potential labeling inconsistencies. In order to test the robustness of the proposed frameworks across different datasets, we also apply it to a dementia staging dataset. We conducted experiments using the OASIS-1 MRI dataset, which contains T1-weighted structural brain images of 416subjects aged 18 to 96 years 16 . The original 3D image files were converted to Nifti format using the FSL tool, and later converted to images in the dataset preparation. Specifically, the images are first reoriented to a standard anatomical orientation and corrected for scanner-induced intensity inhomogeneities using a bias field correction algorithm. The field of view is then cropped to a standard brain size, and non-brain tissue (skull, scalp) is removed via skull-stripping. The resulting brain-extracted volumes are linearly registered to the MNI152 template, a standard stereotactic reference space of the human brain widely used in neuroimaging (Montreal Neurological Institute, 152-subject average) 29, 30 , at 2 m resolution, using a 12 degrees-of-freedom affine transformation with trilinear interpolation. Tissue segmentation is subsequently performed on the registered T1-weighted images, and a White Matter (WM) mask is obtained by thresholding the WM partial volume estimation map. This mask is applied to the original T1w image to isolate WM voxel intensities, and their mean value is computed. Each T1w volume is then intensity-normalized by dividing it by this WM mean value, following the Shinohara normalization 31 approach, in order to reduce inter-subject and inter-scanner intensity variability. Finally, the resulting normalized 3D volumes are converted into 2D representations by extracting axial slices from each subject’s registered brain volume, which are subsequently exported as images for use in the classification framework. Clinical Dementia Rating (CDR) scale is a clinical assessment tool used to quantify the severity of dementia based on cognitive and functional performance, with higher scores indicating greater impairment 32 . Specifically, the CDR is obtained through a semi-structured interview with the patient and a reliable informant, evaluating six cognitive and functional domains: memory, orientation, judgment and problem solving, community affairs, home and hobbies, and personal care. Each domain is rated independently, and the results are combined, typically following the Washington University algorithm, into a global CDR score ranging from0(no impairment) to3(severe dementia), with intermediate stages (0.5,1,2) reflecting questionable, mild, and moderate impairment, respectively 33 . These stages are widely used in clinical practice to characterize disease progression, from questionable cognitive decline (CDR 0.5) to mild (CDR 1), moderate (CDR 2), and severe dementia (CDR 3). The OASIS dataset we use provides four cognitive stage labels, from which three classes were defined for this study. Although these labels do not correspond to a formal clinical CDR assessment, they can be reasonably approximated to established CDR stages based on the severity descriptions reported for the dataset: Non Demented(approximately corresponding to CDR0, i.e., no cognitive impairment), Very Mild Dementia (approximately corresponding to CDR0.5, i.e., questionable to very mild cognitive decline), and Mild/Moderate Dementia (approximately corresponding to CDR≥ 1, encompassing mild and moderate stages of dementia, in which cognitive impairment begins to affect daily functioning). Figures 2e, 2f, and 2g show representative samples for each class, with the last one comprising samples from both the Mild Dementia and Moderate Dementia categories, which were merged due to the latter having fewer than 500 samples, in order to achieve a better balance across the dataset. This three-class formulation enables a more comprehensive evaluation of the proposed HQiCNN framework by considering different stages of cognitive impairment while maintaining a controlled classification problem. Additionally, the majority class was downsampled to match the size of the minority classes, ensuring a balanced distribution among the three selected categories and preventing the model from being biased towards the most represented class. All images were resized to64× 64pixels and normalized to the range[0, 1]. Spatial downscaling was deliberately performed to reduce computational complexity and memory requirements, particularly for quantum-inspired and hybrid quantum–classical models with strict input-size constraints. Moreover, using low-resolution inputs facilitates faster experimentation. It allows testing the feasibility of HQiCNN architectures on resource-limited quantum hardware simulators while preserving the coarse anatomical patterns needed for early-stage classification. 4.2 Hyperparameter and quantum properties analysis This section provides a systematic investigation of how training hyperparameters and quantum properties influence the performance of the proposed hybrid quantum-inspired model and its classical counterpart on the balanced OCT dataset. We analyze each factor independently to identify its individual contribution, followed by a comprehensive comparison to elucidate the conditions under which the quantum-inspired approach provides advantages over the classical baseline. 4.2.1 Training size The amount of labeled data available for training is one of the most influential factors governing the performance of machine learning and quantum models. Since real world data acquisition and labeling can be costly or limited in practical settings, it is important to characterize how classification performance scales with the size of the training set, and whether classical and hybrid quantum-classical models exhibit similar or divergent learning behavior as more data becomes available. To this end, we systematically vary the training set size and evaluate its effect on test accuracy for both architectures under study. 7/17 (a) Drusen class sample.(b) Healthy Control class sample. (c) Choroidal Neovascularization (CNV) class sample.(d) Diabetic Macular Edema (DME) class sample. (e) Non Demented class sample.(f) Very Mild Demential class sample.(g) Mild/Moderate Dementia class sample. Figure 2. Class samples of the considered datasets. Sub-figures a, b,c and d correspond to the retinal OCT images 28 . Sub-figures e, f and g correspond to the OASIS-1 Dementia dataset 16 . Table 2 reveals a clear relationship between classification performance, the amount of training data available, and learning rates. For both CNN and HQiCNN (See Section 3.1), test accuracy generally increases as the training set grows from 200 to 30,000 samples, confirming that additional training examples allow the models to learn more discriminative feature representations and improve their generalization capability. This trend is particularly evident for learning rates of 0.01 and 0.001, where performance improvements are observed almost monotonically across increasing dataset sizes. The CNN model reaches accuracies above 0.90 only when trained with the largest subsets, whereas considerably lower values are obtained when fewer than 1,000 samples are available. A similar pattern can be observed for HQiCNN, suggesting that both architectures benefit from the additional information contained in larger training sets. Interestingly, the magnitude of the improvement is not constant across the explored range. The largest gains occur when moving from very small datasets toward intermediate training sizes. Beyond approximately 12,000 samples, performance increases become more gradual, indicating that both models begin to approach a saturation regime in which additional data provide diminishing returns. This behavior suggests that the convolutional feature extractor is already capable of capturing most of the relevant information available in the dataset once a sufficient number of samples is provided. Although the overall trend is similar for both architectures, HQiCNN exhibits slightly higher performance in several intermediate training size configurations (see in Figure 3), suggesting that the quantum-enhanced representation may be particularly useful when the available amount of training data is sufficient to learn meaningful feature interactions but not large enough for the classical model to fully exploit the underlying structure of the dataset. 8/17 Table 2. Mean test accuracy over ten random seeds for both models exploring different learning rates and training sizes. CNN LRTraining size 20040080010001200160020001200030000 0.01 0.436 ±0.104 0.495 ±0.135 0.575 ±0.049 0.670 ±0.082 0.698 ±0.098 0.749 ±0.083 0.843 ±0.020 0.929 ±0.015 0.930 ±0.010 0.001 0.467 ±0.098 0.535 ±0.125 0.633 ±0.058 0.661 ±0.074 0.734 ±0.055 0.751 ±0.076 0.774 ±0.127 0.924 ±0.010 0.937 ±0.006 0.0001 0.311 ±0.069 0.358 ±0.078 0.550 ±0.092 0.518 ±0.132 0.628 ±0.068 0.643 ±0.049 0.672 ±0.042 0.885 ±0.011 0.919 ±0.006 HQiCNN LRTraining size 20040080010001200160020001200030000 0.01 0.325 ±0.094 0.468 ±0.105 0.551 ±0.099 0.638 ±0.092 0.714 ±0.078 0.800 ±0.032 0.818 ±0.066 0.930 ±0.009 0.929 ±0.012 0.001 0.448 ±0.082 0.559 ±0.076 0.688 ±0.051 0.745 ±0.074 0.769 ±0.028 0.795 ±0.055 0.828 ±0.034 0.920 ±0.008 0.931 ±0.008 0.0001 0.347 ±0.089 0.414 ±0.085 0.542 ±0.080 0.582 ±0.053 0.600 ±0.051 0.653 ±0.037 0.684 ±0.058 0.897 ±0.008 0.920 ±0.008 2004008001000120016002000 Training size 0.01 0.001 0.0001 Learning rate 0.100 0.075 0.050 0.025 0.000 0.025 0.050 0.075 0.100 Accuracy difference (HQiCNN - CNN) Figure 3. Mean test accuracy across ten different seeds with for the CNN and HQiCNN models exploring different learning rates and training sizes. The read areas indicate a better mean accuracy performance of HQiCNN model, while the blue area indicate that the CNN outperforms the hybrid model. The used metric values are discretized, however interpolation is used to make the figure smoother. 4.2.2 Learning rate The learning rate has a substantial impact on the optimization process and, consequently, on the final classification performance. Across both architectures, the experiments indicate that a learning rate of 0.001 consistently provides the most favorable balance between accuracy and optimization stability. When the learning rate increases to 0.01, the models generally achieve a better performance. However, larger fluctuations between training configurations can also be observed, indicating a greater sensitivity to the optimization trajectory. While some experimental settings produce excellent results, others exhibit reduced stability, as reflected by the reported standard deviations. In contrast, the smallest learning rate evaluated, 0.0001, systematically leads to lower accuracies. This behavior suggests that the optimization process becomes excessively conservative, preventing the models from reaching highly discriminative solutions within the allocated training budget. Although performance still improves with increasing training size, neither CNN nor HQiCNN achieves the accuracy levels observed for larger learning rates. These 9/17 findings indicate that the benefits of the hybrid quantum architecture are strongly linked to effective optimization. A poorly chosen learning rate can limit the ability of both classical and quantum parameters to adapt to the training data, thereby masking potential advantages associated with the quantum circuit. To better understand the conditions under which HQiCNN provides an advantage over CNN, Figure 3 presents a heatmap showing the difference in mean test accuracy between both architectures in small and intermediate training size regimes (200-2000 images). Positive values indicate regions where HQiCNN outperforms CNN, whereas negative values indicate the opposite behavior. The heatmap reveals that the performance difference is not determined solely by training size or learning rate. Instead, it emerges from the interaction between both variables. Distinct regions of the parameter space exhibit different behaviors, suggesting that the effectiveness of the quantum layer depends on how optimization dynamics interact with data availability. The most prominent positive region appears around intermediate training sizes and learning rates close to 0.001. In this area, HQiCNN consistently achieves higher accuracies than the classical baseline, indicating that the quantum circuit may be capturing useful feature interactions that remain inaccessible to the fully connected layer used in the CNN. Negative regions are primarily concentrated at the extremes of the parameter space, particularly for very small datasets and for configurations where optimization is too aggressive. Under these conditions, the additional complexity introduced by the quantum circuit does not translate into improved predictive performance. 4.2.3 Batch size Batch size is a key hyperparameter in the training of neural network models, as it directly affects gradient estimation, convergence stability, and the effective noise present in the optimization process. In traditional sample-based hybrid quantum-classical models, batch size additionally influences on measurement statistics and the reliability of expectation-value estimates, making it particularly relevant to assess whether performance trends are robust across different batch size choices. Figure 4a presents the performance of the models for a training size of 1000 samples across three different batch sizes. As has been extensively documented in the machine learning literature, hyperparameter optimization constitutes a critical factor in model performance. In this context, the aim of this analysis is to assess whether the observed difference in accuracy for the training size of 1000 remains consistent across variations in batch size. The quantum inspired hybrid model outperforms its classical counterpart both in mean accuracy and stability, consistently across the three considered batch sizes. This result is consistent with previously published results regarding the optimization of Quantum Neural Networks hyperparameters 34 , where they find that the batch size influences the runtime and memory but not the model performance. 4.2.4 Entanglement The quantum circuit used in the previous results is the one shown in Figure 1b, which includes an entanglement layer implemented via CNOT gates just before measurement. Note that the presence of entanglement is a widely used indicator of the need for a genuine quantum computer, as entangled circuits with non-Clifford gates are exponentially costly to simulate classically. In Figure 4b, we compare the results between the entangled and non-entangled versions of the circuit. On average, the non-entangled version yields better results; however, at training sizes1000and1200, the entanglement produces more accurate and stable results. We remark that our output observables are local and that they may not be able to fully capture the correlations generated by entanglement, potentially limiting the impact of entanglement in the classification results. Since entanglement is not necessary for strong performance in general, our model can incorporate fully separable quantum circuits. This is an important finding of our workflow, as emulating quantum circuits scales linearly with the number of qubits for separable circuits and exponentially for entangled circuits 35 , allowing us to scale to a larger number of qubits. 4.2.5 Observable set size A crucial design choice in any quantum machine learning algorithm is the selection of the observable set used to extract classical information from the quantum state via measurement. This choice directly determines the dimensionality and expressivity of the feature space in which the subsequent classical model operates, and therefore has a direct bearing on both predictive accuracy and computational cost. Since each additional non-commuting observable requires an independent expectation-value estimation on quantum hardware (or an additional shot budget in simulation), understanding how the size and structure of the observable set impacts performance is essential for balancing accuracy against resource requirements. To investigate this trade-off, we compare three nested observable sets of increasing cardinality, corresponding to single-qubit (local), two-qubit (pairwise correlation), and three-qubit (triple correlation) Pauli terms, respectively. In Figure 4c, we directly compare the accuracy averaged over10random seeds for3different observable sets in the algorithm. These observable sets are: •O 1 =Γ i for Γ∈X,Y,Z and i∈1, 4. Previous results were obtained using this observable set. 10/17 81632 Batch size 40% 60% 80% Test accuracy CNN HQiCNN (a) Accuracy comparison between the CNN and HQiCNN models for different batch sizes. 40080010001200160020001200030000 Training size 60% 80% 100% Test accuracy HQiCNN (non-entangled) HQiCNN (entangled) (b) Accuracy comparison for the HQiCNN model with and without entanglement. (c) Accuracy comparison between the CNN and HQiCNN models for different batch sizes. Figure 4. Batch size and quantum properties comparisons for the OCT dataset. In subfigures 4a and 4b, the black horizontal lines of the box plots represent the median, the box covers from the first to the third quartile and the whiskers represent the maximum and minimum points within a 1.5×IQR from the quartiles. •O 2 =O 1 ∪Γ i Γ j for Γ∈X,Y,Z and i, j∈1, 4 with i̸= j. •O 3 =O 1 ∪O 2 ∪Γ i Γ j Γ k for Γ∈X,Y,Z and i, j,k∈1, 4 with i̸= j̸= k. These sets are nested,O 1 ⊂O 2 ⊂O 3 , so that any performance differences can be attributed purely to the additional higher-order correlators rather than to a change in the underlying single-qubit information already captured byO 1 . We observe that the observable length does not influence the results for smaller and intermediate training sizes. This suggests that in the low-data regime the model is limited primarily by the amount of training data rather than by the expressivity of the feature space, and that the local single-qubit observables inO 1 already capture the dominant part of the signal relevant for classification. However, large training sizes (above12000) show a clear linear relationship between observable length and accuracy performance, where the length of the observable slightly improves accuracy results, as we can see for training sizes12000and30000in Figure 4c. This indicates that once sufficient training data is available to reliably exploit a higher-dimensional feature space, the additional higher-order correlators inO 2 andO 3 encode complementary discriminative information that is not accessible from local observables alone. Thus, these results suggest that the observable set size interacts with the training-set size in a manner reminiscent of a bias–variance trade-off: richer observable sets increase the expressivity of the model, but this added expressivity can only be leveraged once enough data is available to estimate the corresponding decision boundary reliably. From a practical standpoint, this would have direct implications for quantum hardware implementations, where each additional observable inO 2 orO 3 incurs additional measurement overhead; and for classically emulated implementations, where no measurement or sampling overhead is required. Our results imply that, for hardware implementations, this overhead is only justified in the large-data regime; in data-limited settings, the smaller observable setO 1 achieves comparable accuracy at a lower measurement cost, making it the preferable choice when quantum resources are constrained. In the emulated setting, bigger and richer observable sets such asO 2 andO 3 measured on classically simulated small quantum circuits can extract more information and correlations between qubits for large training sizes. 4.3 Comparative analysis between CNN and HQiCNN A direct comparison between the two architectures reveals that neither model uniformly dominates the other across all experimental conditions. Instead, the relative performance depends on the interaction between training size and learning rate. The CNN baseline demonstrates strong and consistent performance across the entire parameter space, achieving the highest 11/17 overall accuracy for the largest considered training size in the experiments. This result highlights the effectiveness of classical convolutional representations for the considered classification task. The classical CNN uses a dense layer to map the reduced latent representation to the output space, whereas the HQiCNN encodes the same latent features into a four-qubit parameterized EQC, generating a higher-order feature representation before the final classification layer (see Figure 1). Since both models share the same convolutional backbone, the HQiCNN’s slightly better performance under certain regimes can be attributed primarily to the replacement of the intermediate dense layer by the parameterized EQC. One possible explanation is that entanglement and parameterized quantum operations enable the model to capture complex feature correlations that are difficult to represent with a shallow classical layer of comparable dimensionality, suggesting that the quantum component introduces additional representational flexibility that may be beneficial under certain conditions. However, the limited magnitude of the observed gains also suggests that the shared convolutional backbone already extracts the majority of the discriminative information, so the quantum layer acts primarily as a refinement mechanism rather than as a complete replacement for classical feature learning. As a result, we find that the relatively small performance gap between the two approaches is itself an important finding: it indicates that hybrid quantum-inspired architectures can achieve results comparable to state-of-the-art classical methods, supporting the feasibility of integrating quantum layers into practical machine-learning pipelines. 4.4 Robustness: Dementia use case In previous subsections, we have studied the performance of both models under different hyperparameter configurations using the OCT dataset, a large and highly balanced dataset that enabled a comprehensive evaluation across a wide range of experimental settings. In this subsection, we transfer the best-performing hyperparameter configuration obtained from the OCT experiments to the OASIS dementia MRI dataset (See Section 4.1) to evaluate the robustness and generalization capability of the proposed framework in a different medical imaging scenario. To obtain a balanced and clinically meaningful classification problem, the original OASIS categories were reorganized into three classes: Non Demented, Very Mild Dementia, and Mild/Moderate Dementia (See Figure 5). This reformulation allows the model to distinguish between cognitively healthy subjects and different dementia stages, while reducing the class imbalance present in the original dataset. 999300012000 Training size 60.0% 65.0% 70.0% 75.0% 80.0% 85.0% 90.0% 95.0% 100.0% Test accuracy CNN HQiCNN Figure 5. Accuracy comparison between the HQiCNN and CNN models for the dementia dataset (See Section 4.1). Each model uses the best learning rate found for the most comparable training size in Table 2. We remark that the use of a training size of 999 is cause by ensuring a balanced dataset within the 3 classes. The black horizontal lines of the box plots represent the median, the box covers from the first to the third quartile and the whiskers represent the maximum and minimum points within a 1.5×IQR from the quartiles. Figure 5 shows the accuracy results for different training sizes. We find a clear correlation between the results in this dataset and the OCT’s results shown in Table 2; thus, the models perform similarly on two different health-related image classification tasks. The HQiCNN outperforms the CNN on smaller training sizes; however, this advantage narrows as the training size grows, eventually reversing so the CNN performs slightly better at larger sizes. 12/17 4.5 Explainability To further investigate the decision-making process of the proposed models, a SHAP 13 based tool analysis was performed on OCT images. Figure 6a illustrates the spatial overlap of the top 10% absolute SHAP values for HQiCNN and a conventional CNN. Red regions correspond to image areas exclusively identified as relevant by HQiCNN, blue regions indicate areas uniquely emphasized by the CNN, and green regions represent the regions jointly considered important by both models. For this explainability study, we select the seeds whose per-model accuracies most closely matched their respective mean accuracy across all seeds. The attribution maps show that both architectures focus primarily on the retinal regions affected by the three considered diseases, particularly around the retinal pigment epithelium (RPE) and the characteristic elevations produced by the deposits (see Section 4.1), confirming that their predictions are based on clinically meaningful anatomical structures. However, clear differences emerge when the amount of training data is limited. Under the low-data scenario (1,000 training samples), HQiCNN exhibits a more concentrated and anatomically coherent attention pattern than the conventional CNN, as shown in the left side of Figure 6a. Its saliency maps are predominantly localized around the possibly clinical important areas of the retinal layers, whereas the CNN presents a more fragmented distribution of salient regions. This behavior suggests that HQiCNN learns more discriminative and clinically meaningful representations with fewer training samples, enabling a more accurate localization of disease-related features while reducing attention to less informative image regions. Note that as the training dataset increases to 30,000 samples, the overlap between the two models becomes substantially larger as can be seen in the right side of Figure 6a, indicating that both architectures progressively converge toward similar disease-relevant biomarkers. Nevertheless, HQiCNN still preserves several exclusive attention regions around the RPE boundaries and subtle retinal deformations, suggesting that its hierarchical feature-extraction strategy captures complementary structural information even in large-data settings. In order to get an averaged metric of similarity between the SHAP distribution and to avoid relying exclusively on single image interpretations, we use the previously introducedEMD pos metric. This measure enable us to quantify the similarity between the positive SHAP distribution of the studied models over a set of considered images. We calculate it for the models trained with1000and30000samples obtaining0.066± 0.019and0.045± 0.018, respectively. These results are close to the lower bound, indicating a clear overall correlation between the classical and quantum-inspired models. As expected, increasing the training size reduces the differences between the HQiCNN and CNN positive SHAP distributions. We further validate the interpretability provided by the proposed metricEMD pos by comparing its distribution for the two considered training sizes with a cross-image reference distribution, whereEMD pos is computed between SHAP maps corresponding to different images (see Section 3.3.2). As shown in Figure 6b, both the1000- and30000-sample training configurations yield consistently lowerEMD pos values than the cross-image distribution, as expected. This result indicates that the positive SHAP regions associated with correct disease classification remain spatially consistent across samples, being mainly concentrated within the same RPE area. Therefore, despite the differences in training size, both models identify similar anatomical regions as relevant for the correct prediction output, suggesting that the learned explanations are not only accurate but also spatially stable. This combined analysis shows that HQiCNN consistently attends to anatomically plausible retinal structures in its predictions while also achieving superior data efficiency by learning where to focus with significantly fewer training samples. The improved localization of these regions under low-data conditions, together with the increased agreement both in single and averaged metrics between models as the dataset grows, supports the robustness, explainability, and spatial consistency of the attention patterns identified by the proposed architecture for automated disease classification in OCT images. Note that, while these SHAP-based explanations highlight anatomically coherent and reproducible regions of interest, they do not by themselves constitute clinically validated biomarkers, as such validation would require dedicated assessment by expert clinicians. Nonetheless, the ability of HQiCNN to provide consistent and interpretable visual explanations of its predictions represents a promising step towards more transparent AI-assisted diagnostic systems, and future work involving expert clinical review will be needed to further assess the diagnostic relevance of the identified regions. 5 Conclusions and future work This work presents small emulated quantum circuits as a tool for improving the performance of standard networks, and a systematic evaluation of the conditions under which a HQiCNN provides advantages over a parameter-matched classical CNN on real world medical image classification tasks. By keeping the convolutional backbone identical and replacing only an intermediate dense layer, the observed differences can be directly attributed to the quantum-inspired component. The experimental results show that neither architecture consistently outperforms the other. The classical CNN achieves the highest global accuracy on the OCT dataset (up to93.7%), whereas the HQiCNN consistently performs better in intermediate training size regimes, particularly around800and2000training samples, while maintaining competitive performance across all evaluated configurations. The same trend is reproduced on the OASIS dementia dataset, indicating that the observed behavior is not specific to a single medical imaging task. 13/17 NORMAL DME CNV DRUSEN Training size 1000 Training size 30000 (a)|SHAP|IoUacross the four labels. Red regions denote the areas of highest influence for the HQiCNN model’s classification decision, blue regions denote the corresponding areas for the CNN model, and green regions represent the intersection of influence between the two models. pos (b) Image count distribution against theEMD pos value for both traininig sizes on the same images and the combined different image setting. The dashed vertical lines indicate the mean for each distribution. Figure 6. Explainability plots for comparing both models predictions for the OCT dataset. 14/17 The systematic hyperparameter analysis further shows that the effectiveness of the HQiCNN depends on the interaction between data availability and optimization. Empirically, we show that a learning rate of10 −3 provides the most stable performance, the HQiCNN remains consistently robust across different batch sizes, and removing entanglement produces comparable, and often slightly better, classification accuracy. This finding is particularly relevant because separable circuits scale linearly during classical simulation, whereas entangled circuits exhibit exponential complexity. Similarly, extending the observable set beyond local Pauli measurements yields measurable improvements for large training datasets, suggesting that the additional measurement cost is justified only when sufficient data are available. From an interpretability perspective, the proposed|SHAP|IoU and EMD pos metrics show that both architectures focus on the same clinically relevant anatomical structures. The average EMD pos decreases from0.066± 0.019for1000training samples to0.045± 0.018for30000samples, indicating an increasing agreement between both models as more data become available. Moreover, the HQiCNN produces more localized and clinically coherent explanations under limited-data conditions, suggesting greater data efficiency without relying on spurious image regions. Overall, the results indicate that the advantage provided by the quantum-inspired layer is not a direct consequence of introducing quantum operations, but rather depends on the interaction between the amount of available training data, the optimization strategy, and the selected measurement scheme. This suggests that the effectiveness of hybrid quantum-inspired models is strongly conditioned by their design choices and experimental configuration. Note that the presented approach is based on emulated shallow circuits with four qubits, which allows fast running on classical computers and avoids the challenges associated with practical quantum computation. In particular, these models do not face any additional limitations imposed by the quantum hardware, such as noise accumulation, decoherence effects, statistical uncertainty from finite-shot measurements, and restrictions imposed by the physical connectivity of quantum processors, which affect the behavior and scalability of these models. Future work will focus on the clinical validation of the proposed framework in hospital environments, including evaluation on larger multicenter cohorts and under realistic diagnostic conditions. In particular, the current evaluation lacks external validation on independent clinical datasets acquired at different institutions, as well as reader studies involving clinicians to assess the diagnostic utility and interpretability of the model’s predictions. Additionally, the reliance on a simplified 2D slice-based representation of the MRI volumes, rather than the full 3D structural information, may limit the anatomical context available to the model, and this will be addressed in future extensions of this work. Additionally, although this work provides a comprehensive evaluation of predictive performance across multiple experimental settings, it does not analyze the computational cost associated with these approaches, including training time, inference efficiency, and memory requirements. Future research will, therefore, focus on extending the proposed framework to larger emulable quantum circuits and other quantum-inspired techniques such as tensor networks or Pauli propagation, together with a detailed analysis of computational efficiency to determine the practical trade-offs between predictive performance and resource requirements in hybrid quantum-inspired learning models. 15/17 Data availability The data supporting the experiments conducted in this article are the OCT scans, available in Reference 15 , and the OA- SIS1 images dataset 16 used for the technical robustness section, available athttps://w.kaggle.com/datasets/ ninadaithal/imagesoasis. Details related to the OASIS MRI dataset preprocessing can be found athttps://w. kaggle.com/datasets/ninadaithal/oasis-1-shinohara/. Code availability The code used in this study can be made available from the corresponding authors upon reasonable request. References 1. Biamonte, J. et al. Quantum machine learning. Nature 549, 195–202, DOI: 10.1038/nature23474 (2017). 2. Cerezo, M. et al. Variational quantum algorithms. Nat. Rev. Phys. 3, 625–644 (2021). 3. Benedetti, M., Lloyd, E., Sack, S. & Fiorentini, M. Parameterized quantum circuits as machine learning models. Quantum science technology 4, 043001 (2019). 4.Goto, T., Tran, Q. H. & Nakajima, K. Universal approximation property of quantum machine learning models in quantum-enhanced feature spaces. Phys. Rev. Lett. 127, 090506, DOI: 10.1103/PhysRevLett.127.090506 (2021). 5. Liu, J. et al. Hybrid quantum-classical convolutional neural networks. Sci. China Physics, Mech. & Astron. 64, 290311 (2021). 6.Long, C., Huang, M., Ye, X., Futamura, Y. & Sakurai, T. Hybrid quantum-classical-quantum convolutional neural networks. Sci. Reports 15, 31780 (2025). 7.Mari, A., Bromley, T. R., Izaac, J., Schuld, M. & Killoran, N. Transfer learning in hybrid classical-quantum neural networks. Quantum 4, 340, DOI: 10.22331/q-2020-10-09-340 (2020). 8.Senokosov, A., Sedykh, A., Sagingalieva, A., Kyriacou, B. & Melnikov, A. Quantum machine learning for image classification. Mach. Learn. Sci. Technol. 5, 015040, DOI: 10.1088/2632-2153/ad2aef (2024). 9.Bermejo, P. et al. Quantum convolutional neural networks are effectively classically simulable. PRX Quantum 7, 020304, DOI: 10.1103/8qt9-72ts (2026). 10. Cerezo, M., Sone, A., Volkoff, T., Cincio, L. & Coles, P. J. Cost function dependent barren plateaus in shallow parametrized quantum circuits. Nat. Commun. 12, 1791, DOI: 10.1038/s41467-021-21728-w (2021). 11.Bowles, J., Ahmed, S. & Schuld, M. Better than classical? the subtle art of benchmarking quantum machine learning models (2024). 2403.07059. 12.Parvaiz, A. et al. Vision transformers in medical computer vision—a contemplative retrospection. Eng. Appl. Artif. Intell. 122, 106126 (2023). 13.Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Adv. neural information processing systems 30 (2017). 14. LeCun, Y., Bengio, Y. & Hinton, G. Deep learning. nature 521, 436–444 (2015). 15.Kermany, D., Zhang, K. & Goldbaum, M. Labeled optical coherence tomography (oct) and chest x-ray images for classification, DOI: 10.17632/rscbjbr9sj.2 (2018). 16.Marcus, D. S. et al. Open access series of imaging studies (oasis): Cross-sectional mri data in young, middle aged, nondemented, and demented older adults. J. Cogn. Neurosci. 19, 1498–1507, DOI: 10.1162/jocn.2007.19.9.1498 (2007). 17. Nielsen, M. A. & Chuang, I. L. Quantum computation and quantum information (Cambridge university press, 2010). 18. Abbas, A. et al. The power of quantum neural networks. Nat. computational science 1, 403–409 (2021). 19. McClean, J. R., Boixo, S., Smelyanskiy, V. N., Babbush, R. & Neven, H. Barren plateaus in quantum neural network training landscapes. Nat. Commun. 9, 4812, DOI: 10.1038/s41467-018-07090-4 (2018). 20.Wang, S. et al. Noise-induced barren plateaus in variational quantum algorithms. Nat. Commun. 12, DOI: 10.1038/ s41467-021-27045-6 (2021). 21.Schuld, M., Bergholm, V., Gogolin, C., Izaac, J. & Killoran, N. Evaluating analytic gradients on quantum hardware. Phys. Rev. A 99, 032331, DOI: 10.1103/PhysRevA.99.032331 (2019). 16/17 22. Paszke, A. et al. Automatic differentiation in pytorch. OpenReview (2017). 23.Selvaraju, R. R. et al. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), 618–626 (2017). 24. Chattopadhay, A., Sarkar, A., Howlader, P. & Balasubramanian, V. N. Grad-cam++: Generalized gradient-based visual explanations for deep convolutional networks. In 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), 839–847 (IEEE, 2018). 25. Moraffah, R., Karami, M., Guo, R., Raglin, A. & Liu, H. Causal interpretability for machine learning – problems, methods and evaluation. ACM SIGKDD Explor. Newsl. 22, 18–33 (2020). 26. Ribeiro, M. T., Singh, S. & Guestrin, C. " why should i trust you?" explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining, 1135–1144 (2016). 27. Vimbi, V., Shaffi, N. & Mahmud, M. Interpreting artificial intelligence models: a systematic review on the application of lime and shap in alzheimer’s disease detection. Brain informatics 11, 10 (2024). 28.Kermany, D. S. et al. Identifying medical diagnoses and treatable diseases by image-based deep learning. cell 172, 1122–1131 (2018). 29.Fonov, V. S., Evans, A. C., McKinstry, R. C., Almli, C. R. & Collins, D. L. Unbiased nonlinear average age-appropriate brain templates from birth to adulthood. NeuroImage 47, S102, DOI: 10.1016/S1053-8119(09)70884-5 (2009). 30.Fonov, V. et al. Unbiased average age-appropriate atlases for pediatric studies. NeuroImage 54, 313–327, DOI: 10.1016/j. neuroimage.2010.07.033 (2011). 31.Shinohara, R. T. et al. Statistical normalization techniques for magnetic resonance imaging. NeuroImage: Clin. 6, 9–19 (2014). 32.Rauchmann, B.-S., Laib, J., Ercik, B., Perneczky, R. & Altares-López, S. Multimodal ordinal modeling of alzheimer’s disease severity using structural mri and clinical data. arXiv preprint arXiv:2606.11794 (2026). 33. Morris, J. C. The Clinical Dementia Rating (CDR): Current version and scoring rules. Neurology 43, 2412–2414, DOI: 10.1212/wnl.43.11.2412-a (1993). 34.Moussa, C., Patel, Y. J., Dunjko, V., Bäck, T. & van Rijn, J. N. Hyperparameter importance and optimization of quantum neural networks across small datasets. Mach. Learn. 113, 1941–1966, DOI: 10.1007/s10994-023-06389-8 (2024). 35.Vidal, G. Efficient classical simulation of slightly entangled quantum computations. Phys. Rev. Lett. 91, 147902, DOI: 10.1103/PhysRevLett.91.147902 (2003). Funding This article is funded by PESL-Stiftung-Alzheimer 2026 in Bayern, Germany, (registered project-80765114 Pesl-Alzheimer- Stift, Principal Investigator: Dr.-Ing. Sergio Altares-López), which is focused on research in Alzheimer’s disease. Furthermore, researchers G.R., M.O., M.A., G.B. and P.D. acknowledge the support provided by project ARQADE (CER-20251019), funded by the CERVERA Research Programme of CDTI (Centre for Technological Development and Innovation). Author contributions statement G.R., M.O., M.A., P.D., and S.A. conceived and planned the experiments. G.R. and M.O. carried out the experiments. G.B., S.A., and B.R. provided funding. G.R., M.O., M.A., P.D., and S.A. contributed to the interpretation of the results. G.R. and S.A. took the lead in writing the manuscript. All authors provided critical feedback and helped shape the research, analysis and manuscript. Additional information Competing financial interests: The authors declare no competing financial interests. 17/17