Paper deep dive
Following the Diagnostic Trace: Visual Cognition-guided Cooperative Network for Chest X-Ray Diagnosis
Shaoxuan Wu, Jingkun Chen, Chong Ma, Cong Shen, Xiao Zhang, Jun Feng
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/20/2026, 12:26:48 PM
Summary
The paper introduces VCC-Net, a visual cognition-guided collaborative network for chest X-ray diagnosis that integrates radiologists' visual search traces (via eye-tracking or mouse) with AI inference. It employs a Visual Attention Generator (VAG) to learn hierarchical visual search strategies and a Visual Cognition-guided Classifier (VCC) with a cognition-graph co-editing module to align model representations with radiologist attention, improving diagnostic accuracy and interpretability on datasets like SIIM-ACR and EGD-CXR.
Entities (10)
Relation Signals (8)
VCC-Net ā evaluatedon ā EGD-CXR
confidence 98% Ā· Experiments on the public datasets... EGD-CXR... achieved classification accuracies of 85.05%
VCC-Net ā evaluatedon ā SIIM-ACR
confidence 98% Ā· Experiments on the public datasets SIIM-ACR... achieved classification accuracies of 88.40%
VCC-Net ā contains ā Visual Attention Generator
confidence 95% Ā· VCC-Net comprises two principal modules: The visual attention generator (VAG)...
VCC-Net ā contains ā Visual Cognition-guided Classifier
confidence 95% Ā· The visual cognition-guided classifier (VCC) employs a cognitionāgraph co-editing module...
VCC-Net ā evaluatedon ā TB-Mouse
confidence 95% Ā· Experiments on the... self-constructed TB-Mouse dataset achieved classification accuracies of 92.41%
Visual Cognition-guided Classifier ā contains ā Cognition-Graph Co-editing Module
confidence 92% Ā· The VCC employs a cognitionāgraph co-editing module to integrate VC for constructing a disease-aware graph.
Cognition-Graph Co-editing Module ā integrates ā Visual Cognition
confidence 90% Ā· A cognition-graph co-editing module subsequently integrates radiologist VC with model inference
Visual Attention Generator ā uses ā
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Computer-aided diagnosis (CAD) has significantly advanced automated chest X-ray diagnosis but remains isolated from clinical workflows and lacks reliable decision support and interpretability. Human-AI collaboration seeks to enhance the reliability of diagnostic models by integrating the behaviors of controllable radiologists. However, the absence of interactive tools seamlessly embedded within diagnostic routines impedes collaboration, while the semantic gap between radiologists' decision-making patterns and model representations further limits clinical adoption. To overcome these limitations, we propose a visual cognition-guided collaborative network (VCC-Net) to achieve the cooperative diagnostic paradigm. VCC-Net centers on visual cognition (VC) and employs clinically compatible interfaces, such as eye-tracking or the mouse, to capture radiologists' visual search traces and attention patterns during diagnosis. VCC-Net employs VC as a spatial cognition guide, learning hierarchical visual search strategies to localize diagnostically key regions. A cognition-graph co-editing module subsequently integrates radiologist VC with model inference to construct a disease-aware graph. The module captures dependencies among anatomical regions and aligns model representations with VC-driven features, mitigating radiologist bias and facilitating complementary, transparent decision-making. Experiments on the public datasets SIIM-ACR, EGD-CXR, and self-constructed TB-Mouse dataset achieved classification accuracies of 88.40%, 85.05%, and 92.41%, respectively. The attention maps produced by VCC-Net exhibit strong concordance with radiologists' gaze distributions, demonstrating a mutual reinforcement of radiologist and model inference. The code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2602.21657v1
- Canonical: https://arxiv.org/abs/2602.21657v1
Trouble viewing inline? Open PDF directly ā
Full Text
53,871 characters extracted from source content.
Expand or collapse full text
IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 20201 Following the Diagnostic Trace: Visual Cognition-guided Cooperative Network for Chest X-Ray Diagnosis Shaoxuan Wu, Jingkun Chen, Chong Ma, Cong Shen, Xiao Zhang, Jun Feng Abstract ā Computer-aided diagnosis (CAD) has signifi- cantly advanced automated chest X-ray diagnosis but re- mains isolated from clinical workflows and lacks reliable decision support and interpretability. Human-AI collabora- tion seeks to enhance the reliability of diagnostic models by integrating the behaviors of controllable radiologists. However, the absence of interactive tools seamlessly em- bedded within diagnostic routines impedes collaboration, while the semantic gap between radiologistsā decision- making patterns and model representations further limits clinical adoption. To overcome these limitations, we pro- pose a visual cognition-guided collaborative network (VCC- Net) to achieve the cooperative diagnostic paradigm. VCC- Net centers on visual cognition (VC) and employs clinically compatible interfaces, such as eye-tracking or the mouse, to capture radiologistsā visual search traces and atten- tion patterns during diagnosis. VCC-Net employs VC as a spatial cognition guide, learning hierarchical visual search strategies to localize diagnostically key regions. A cogni- tionāgraph co-editing module subsequently integrates ra- diologist VC with model inference to construct a disease- aware graph. The module captures dependencies among anatomical regions and aligns model representations with VC-driven features, mitigating radiologist bias and facili- tating complementary, transparent decision-making. Exper- iments on the public datasets SIIM-ACR, EGD-CXR, and self-constructed TB-Mouse dataset achieved classification accuracies of 88.40%, 85.05%, and 92.41%, respectively. The attention maps produced by VCC-Net exhibit strong concordance with radiologistsā gaze distributions, demon- strating a mutual reinforcement of radiologist and model inference. The code is available at https://github.com/IPMI- NWU/VCC-Net. Index Termsā Visual Cognition, Computer-aided Diagno- sis, Eye Tracking, Mouse Trajectory This work was supported in part by the National Natural Science Foundation of China under Grant 62403380 and the Shaanxi Province Postdoctoral Science Foundation under Grant 2024BSHSDZZ042. (Co- first authors: Shaoxuan Wu, Jingkun Chen; Corresponding authors: Xiao Zhang, Jun Feng.) Shaoxuan Wu, Xiao Zhang, and Jun Feng are with the College of Computer Science, Northwest University, Xiāan, China, 710127 (e-mail: wushaoxuan@stumail.nwu.edu.cn; xiaozhang, fengjun@nwu.edu.cn). Jingkun Chen is with the Department of Engineering Science, University of Oxford, Oxford, OX3 7DQ, United Kingdom (e-mail: jingkun.chen@eng.ox.ac.uk). Chong Ma is with the School of Computing and Artificial Intelligence, Southwest Jiaotong University, Chengdu 611756, China (e-mail: ma- chong@swjtu.edu.cn). Cong Shen is with the Department of PET/CT, The First Affiliated Hospital of Xiāan Jiaotong University, Xiāan, 710061, China. (shen- cong100217@fh.xjtu.edu.cn). I. INTRODUCTION Our Collaborative Paradigm Visual Attention Generator Diagnostic Trajectory x y t Human Visual Attention Diagnosis Result Visual Cognition- guided Classifier Network Visual Attention Visual Cognition Perception Preliminary Screening Correction Conventional Paradigm CAD System Fig. 1.The collaborative paradigm (right) leverages radiologistsā visual cognition to bridge radiologistsā cognition and model inference. Compared with conventional CAD (left), it enhances the consistency and reliability in clinical decisions. L UNG diseases such as pneumothorax, pneumonia, and tuberculosis impose a substantial global health burden, with tuberculosis alone accounting for more than ten million newly reported cases each year [43]. Chest X-ray remains a clinically diagnostic modality that radiologists rely on to guide therapeutic decisions. However, the growing volume of imag- ing examinations places considerable pressure on radiology services, potentially reducing diagnostic accuracy and slowing clinical decision-making. The rapid advancement of computer-aided diagnosis (CAD) has substantially improved automated chest X-ray diagnosis, offering radiologists effective support for rapid screening [36], [39]. Despite achievements, most models still rely on end-to- end, data-driven paradigms and function as isolated computa- tional models with limited integration into clinical workflows [37]. In addition, the models are susceptible to non-clinical factors [33] and demonstrate limited interpretability, factors that weaken radiologistsā trust [9]. These limitations hinder models from becoming trustworthy decision-making partners and restrict their widespread adoption in automated chest X- ray diagnosis. In recent years, the reassessment of the relationship between arXiv:2602.21657v1 [cs.CV] 25 Feb 2026 2IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2020 computer-aided diagnosis (CAD) and radiologists has emerged as a key research direction in chest X-ray diagnosis. Rather than functioning solely as a decision-support role, the new generation of models promotes collaboration, enabling radi- ologists and models to participate in the diagnostic process jointly [38]. By introducing radiologist decision-making pat- terns, human-AI collaboration enhances the transparency and reliability of diagnostic results while maintaining automation efficiency [41]. Nevertheless, the lack of interactive tools that integrate seamlessly into routine clinical workflows constrains effective collaboration, and the semantic discrepancy between radiologistsā decision-making patterns and model representa- tions continues to impede broader clinical adoption. Radiologistsā visual cognition (VC) during image reading represents a natural bridge between radiologist reasoning and model inference. VC captures the decision-making process of radiologists and can be captured through a clinically compatible mannerāsuch as eye-tracking or mouseāwithout interrupting standard diagnostic routines [8], [11]. The cogni- tive process typically follows a hierarchical search strategy, beginning with global structural scanning and progressing toward localized inspection of suspicious lesions [19]. Fine- grained abnormalities such as nodules often elude models, while radiologistsā attention patterns compensate for these gaps. Whereas model-generated cues reduce biases caused by fatigue or subjectivity. The bidirectional synergy between the two establishes a solid foundation for developing an effective collaborative diagnostic paradigm. Building upon these insights, a visual cognition-guided cooperative network (VCC-Net) is proposed to implement an efficient collaborative diagnostic paradigm. VCC-Net com- prises two principal modules: The visual attention generator (VAG) learns radiologistsā hierarchical visual search strategies, progressively refining focus from global to local regions to dynamically localize clinically relevant areas. The visual cognition-guided classifier (VCC) employs a cognitionāgraph co-editing module to integrate radiologistsā VC patterns with attention maps produced by the model. The resulting graph structure encodes pathological semantics and captures in- terdependencies among anatomical regions, enabling more comprehensive diagnostic reasoning. As illustrated in Fig.1, radiologistsā fixation provides spatial constraints that guide the model to concentrate on clinically pertinent regions, while the modelās attention maps assist radiologists in detecting sub- tle findings or correcting subjective deviations. Such mutual reinforcement facilitates complementary decision-making and cognitive alignment throughout the diagnostic process. Exper- imental evaluations confirm that VCC-Net delivers consistent improvements in both diagnostic accuracy and reliability. The main contributions of this paper are as follows: ⢠We propose VCC-Net, which learns and integrates radiol- ogistsā visual search traces and attention patterns through the VAG and VCC modules, respectively, to achieve a collaborative and complementary diagnostic paradigm. ⢠The VAG designed to emulate the hierarchical search strategy characteristic of radiologists, combines the global contextual modeling capabilities with the localized fea- ture extraction strengths. By leveraging VC, it captures radiologistsā visual search behavior and generates atten- tion maps that emphasize clinically significant regions. ⢠The VCC employs a cognitionāgraph co-editing module to integrate VC for constructing a disease-aware graph. The module captures inter-regional dependencies among anatomical regions and aligns model representations with VC features, mitigating radiologist bias and facilitating transparent diagnostic decisions. ⢠Extensive experiments on two public gaze datasets, SIIM- ACR and EGD-CXR, along with the self-constructed TB- Mouse dataset, demonstrate that the VCC-Net outper- forms current state-of-the-art methods in both diagnostic accuracy and clinical interpretability. I. RELATED WORK A. Learning Radiologistsā Attention for Medical Imaging Analysis In recent years, attention mechanisms have gained widespread adoption in medical image analysis. Inspired by radiologistsā VC, attention mechanisms guide models to pri- oritize key regions in images [14]. Pioneering studies aug- mented convolutional neural network (CNN) architectures with attention modules, sharpening their focus on diagnostically pertinent regions [15], [16]. These methods have integrated spatial and channel attention-through techniques, such as adap- tive aggregation and collective spatial attention more precisely highlight abnormal regions. Self-attention is inspired by the cognitive processes of the brain and aims to emulate the ability to selectively concentrate on salient information through dynamic weight allocation. It has gained significant attention due to its effectiveness in handling long-range dependencies. Recent methods have improved contextual representation by introducing global-local feature interaction mechanisms [17] or by integrating the local feature extraction capabilities of CNN with the long- range dependency modeling strengths of Transformers [35]. In addition, deformable Transformers have been employed to further refine attention to clinically significant regions within medical images [18]. The key advantage of attention mechanisms is their ability to simulate attention by selectively focusing on critical informa- tion, thus enhancing feature recognition. However, it remains uncertain whether the attention learned by models accurately reflects attention. VC, which encompasses attentional behavior during image interpretation, offers a promising avenue for guiding feature selection and improving model interpretabil- ityāan essential factor in fostering clinical trust in automated diagnostic models [9], [14]. VC is expected to mitigate the mismatch between model focus and the clinical judgment process by aligning attention mechanisms with radiologists- driven visual cues. B. Integrating Visual Cognition into Collaborative Medical Image Analysis Radiologistsā VC is reflected in the distribution of their attention across medical images and radiologistsā decision- making processes [11], [19], [40]. VC can be effectively recorded by tracking radiologistsā gaze or mouse trajectories. AUTHOR et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING3 Visual Cognition-guided Cooperative Network for Medical Image Diagnosis Stem Visual Attention Generator (VAG) Hard Head Diagnose Head Soft Head Stem Diagnose Head Visual Cognition-guided Classifier (VCC) FC Graph Conv FC Graph ķ¢ FC FC Feature Map f Conv BlockConv Block Visual Distance ķ a Feature Map f Graph ķ¢ Edge e Feature Distance ķ f Visual Attention Ģ ķ soft Feature Aggregation Feature Transform Distance Fusion Distance ķ Top K Encoder E(Ā·) Decoder D soft (Ā·) Guide GNN BlockCNN BlockCognition-Graph Co-editing Module (CGCM) CNN Block GNN Block Graph Construction with Feature Down-Sampling Pointwise Addition Up-Sampling CGCM ... ķ ķµ ļæ½ķ 1 ķ ķ ļæ½ķ 1 ķ ķ ļæ½ķ ķ ķ ķ ļæ½ķ ķµ ... ... ... ... ķ ķµ ļæ½ķ ķµ ķ ķ ķ ķ ... ķ ķ ķ ķ ... ķ ķµ ķ ķµ Feature Distance ķ f Feature Map f Graph ķ¢ Top K Decoder D hard (Ā·) Fig. 2.The proposed VCC-Net comprises two main components: (1) The VAG employs GNN and CNN to model both global and local visual search patterns of radiologists, and supervises learning through three pathways. The VAG generates radiologist-like visual attention based on the input medical images. (2) The VCC leverages radiologistsā VC to construct a graph structure and align visual and feature differences across regions, ultimately producing diagnostic outputs. Gaze record visual fixation points, highlighting regions most relevant to diagnosis [20], while mouse trajectories track user activity and stay points at varying granularities during interac- tion [21], [22]. This data collection framework preserves the integrity of clinical workflows while providing comprehensive, multidimensional data support for CAD. Recent studies have integrated radiologistsā VC into CAD. For example, GA-Net was proposed by Wang et al. [8], where gaze maps and attention consistency were used to guide the networkās focus, thereby improving diagnostic accuracy from knee X-rays. EG-ViT [9] uses gaze information to mask irrelevant background features, reducing shortcut learning. GazeGNN [10] embeds gaze into the graph neural network (GNN) input to enhance inference robustness, also improv- ing interpretability by constructing graph nodes from image patches. Xie et al. [23] introduced a medical image segmen- tation framework that integrates eye-tracking information as weak supervision, improving performance under limited data conditions. Additionally, Gajos et al. [24] developed Hevelius, a rapid mouse trajectory test, to support objective assessments of movement disorders in clinical trials. Li et al. [25] proposed AG-CNN, an attention-based CNN that addresses redundancy in retinal fundus images, improving glaucoma detection. Ad- ditionally, Hevelius was developed by Gajos et al. [24] as a rapid mouse trajectory test, designed to facilitate objective movement disorder assessments in clinical trials. AG-CNN was proposed by Li et al. [25], an attention-based CNN that addresses redundancy in retinal fundus images, thereby enhancing glaucoma detection. Existing studies have validated the potential of VC to enhance diagnostic performance in CAD [42]. Current main- stream methods primarily guide models to focus on spe- cific regions by embedding or masking background infor- mation, but have yet to explore the visual search patterns and anatomical associations inherent in VC. Graph-based methods offer a promising strategy for modeling spatial and semantic relationships in medical images. By leveraging VC- informed relational modeling, it is possible to capture inter- regional correlations. Aligning graph structures with cognitive strategies enables the learning of clinical decision-making processes, thereby contributing to improved performance and interpretability. I. METHOD The VCC-Net aims to improve chest X-ray diagnosis per- formance and transparency by leveraging radiologistsā visual cognition (VC), i.e., gaze or mouse trajectory, during film reading. As shown in Fig.2, the proposed VCC-Net consists of the VAG and the VCC. The VAG learns visual patterns from radiologists and generates corresponding visual attention from medical images. The VCC leverages visual attention to construct graph structures and align feature differences and visual differences, guiding the model to focus on disease- related regions, mirroring radiologistsā diagnostic processes. The following sections describe the data collection and label generation in Section I-A, the network architecture of the VAG in Section I-B, and the architecture of the VCC in Section I-C. A. Data Collection and Label Generation The eye tracker automatically collects the gaze data and can be seamlessly integrated into the radiologistās workflow. For the gaze points in the gaze datasets SIIM-ACR [26] and 4IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2020 Cognition-Graph Co-editing Module (CGCM) Reshape Feature Map f ķ» ķ Ćķ ķ Ćķ¶ Visual Attention Ģ ķ soft ķ» ķ Ćķ ķ DistanceFusion ķ ķ ļæ½ķ ķµ ... ... ......... ķ ķµ ļæ½ķ ķ ... ķ ķ ļæ½ķ 1 ... ķ ķ ļæ½ķ ķ ķ ķ ļæ½ķ ķ ...... ķ ķµ ļæ½ķ ķµ ķ ķ ķ ķ ķ ķ ķ ķµ ķ ķ ķ ķ ķ ķµ ķ ķ ... ... ķ ķ ļæ½ķ ķµ ... ... ......... ķ ķµ ļæ½ķ ķ ... ķ ķ ļæ½ķ 1 ... ķ ķ ļæ½ķ ķ ķ ķ ļæ½ķ ķ ...... ķ ķµ ļæ½ķ ķµ ķ ķ ķ ķ ķ ķ ķ ķµ ķ ķ ķ ķ ķ ķµ ķ ķ ... ... ķĆķ¶ Reshape ķ 1 ķ 2 ķ 3 ķ ķ ... ķ Feature Distance ķ f ķĆķ Visual Distanceķ a ķĆķ ķ ķ ļæ½ķ ķµ ... ... ......... ķ ķµ ļæ½ķ ķ ... ķ ķ ļæ½ķ 1 ... ķ ķ ļæ½ķ ķ ķ ķ ļæ½ķ ķ ...... ķ ķµ ļæ½ķ ķµ Normalization ķ ķ ļæ½ķ ķµ ... ... ......... ķ ķµ ļæ½ķ ķ ... ķ ķ ļæ½ķ 1 ... ķ ķ ļæ½ķ ķ ķ ķ ļæ½ķ ķ ...... ķ ķµ ļæ½ķ ķµ ... ... ......... ... ... ...... Distance ķ ķĆķ Normalization L align Top K Edge Index ķ f ķĆķ¾ Graph ķ¢ f ķ... ķ...ķ ķķ ...... ķµķķ...ķ ķ ķķ ķ ... ...... Pointwise Addition Euclidean Distance ļæ½ ķĆķ¶ Nodes ķ± f Nodes ķ± a ļæ½ ķ f ļæ½ ķ a Ćα ķ£ 1 ķ£ 2 ķ£ 3 ķ£ ķ ... Fig. 3. The CGCM aligns feature and visual distances for each image node by incorporating radiologistsā VC, optimizing the feature space structure. With the addition of soft visual attention. The CGCM also guides the network in constructing a graph structure concentrated on high-attention regions, establishing a robust model aligned with radiologistsā cognition. EGD-CXR [28], [29], we process gaze maps based on the data preprocessing steps from [9], which are referred to as soft visual attention y soft in this paper. These reflect the areas of attention of the radiologists. Compared to gaze data, mouse trajectory data are easier to collect and can be obtained using standard input devices, such as a mouse. To this end, we developed a system that automatically captures mouse trajectories during reading. As the radiologist interacts with the images, the system tracks the mouse position, recording a set of points M points = p 1 ,p 2 ,Ā· ,p m , where each point p represents the mouse position at a particular moment. After data collection, two key post-processing steps are performed. First, similar to the fixation extraction process used with gaze, we apply the I2MC [31] algorithm to identify stay points M stay = p 1 ,p 2 ,Ā· ,p n , which represent regions of prolonged attention during reading. Next, a 2D Gaussian kernel with a 150-pixel radius and a 25-pixel sigma is applied to the stay points, generating a soft visual attention y soft ā R HĆW , where H and W are the image height and width, respectively. The visual attention y soft reflects the radiologistās focus, with higher values indicating areas that attracted more attention. The same method is applied to gaze datasets to generate corresponding attention. B. Visual Attention Generator The VAG aims to learn radiologistsā VC from visual atten- tion generated through gaze or mouse trajectory. By combining the global modeling capability of GNNs and the local feature extraction capability of CNNs, the VAG generates visual attention corresponding to an image. This structure mimics the radiologistās diagnostic process, where they first identify suspicious regions globally and then focus on local regions to identify anomalies like small nodules or exudation. VAG consists of an encoder E(Ā·) and two decoders, D soft (Ā·) and D hard (Ā·). The encoder includes a feature-based graph construction layer and a GNN block for extracting global features. Specifically, the input image xā R HĆWĆ1 is first downsampled through a stem layer to H 4 Ć W 4 ĆC , where C represents the feature dimension. The graph construction layer transforms features at each position into a node v ā R C , forming a node-setV =v 1 ,v 2 ,Ā· ,v N , where N = H 4 ā W 4 . Based on theV , the distance matrixD f ā R NĆN is calculated as follows: D i,j f =ā„ v i ā v j ā„ 2 ,(1) where i and j denote the index of nodes. Using the distance matrix D f , the nearest neighbors of each node are identified, forming a graph structure G = V,E, where E denotes the edges between nodes. The graph can be represented by the feature vector X ā R NĆC . The GNN block performs feature aggregation and transformation operations, which can be formulated as: X 1 = FC 2 (GC(FC 1 (X))) + X,(2) Y = FC 4 (FC 3 (X 1 )) + X 1 ,(3) where FC(Ā·) denotes fully connected layers and GC(Ā·) indi- cates max-relative graph convolution [27]. Both decoders consist of four CNN blocks for local feature extraction. The CNN block is defined as: Z = Conv 2 (Conv 1 (Y )) + Y,(4) where Conv(Ā·) include 3Ć 3 convolution layers, batch nor- malization, and ReLU activation. Skip connections between encoder and decoder blocks at the same scale allow integration of global and local information. After through D soft (Ā·) and the soft head, the soft visual attention p soft is generated. Similarly, the hard visual attention p hard , consisting of binary values indicating regions of high attention, is generated after passing through D hard (Ā·) and the hard head. Finally, the predicted category p aux is obtained via the diagnose head. Both p aux and p hard as complementary information to improve p soft generation quality. The loss function for VAG is defined as: L V AG =L soft +L hard +L aux =L MSE (p soft ,y soft ) +L Dice (p hard ,y hard ) +L CE (p aux ,y cls ), (5) AUTHOR et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING5 where y soft is the ground truth soft visual attention, y cls is the category label, and y hard is a hard label obtained by thresholding y soft (i.e., y hard = I(y soft > threshold)). The indicator function I(Ā·) evaluates to 1 when the probability exceeds the threshold and 0 otherwise. The L MSE , L Dice , and L CE correspond to mean squared error loss, Dice loss, and cross-entropy loss, respectively. C. Visual Cognition-guided Classifier The VCC integrates visual cognition through the cognition- graph co-editing module (CGCM) layer to guide the modelās focus on task-relevant areas. By leveraging radiologistsā at- tention as supervisory signals, VCC aligns visual and feature differences across regions, enabling the model to learn fea- ture representations consistent with radiologistsā cognition. As shown in Fig.2, VCC consists of the CGCM and GNN blocks, which share a similar structure to the VAG encoder, but differ in graph construction. For a feature map f ā R H f ĆW f ĆC , it is treated as a node V f =v 1 ,v 2 ,...,v N , where N = H f ā W f . The prediction of soft visual attention map p soft is downsampled to the feature map size, resulting in Ģp soft ā R H f ĆW f , which is also converted into a point set V a = a 1 ,a 2 ,...,a N . Based on the f and Ģp soft , we compute two distance matrices, D f and D a , which are normalized via min-max scaling to obtain Ė D f and Ė D a . The Ė D f (i,j) and Ė D a (i,j) represent the differences between the i-th and j-th nodes in feature and visual attention space, respectively. By aligning Ė D f and Ė D a , we help the model learn feature representations that align with the radiologistās attention. The alignment loss is defined as: L align =L MSE ( Ė D f , Ė D a ).(6) Additionally, the distance matrices Ė D f and Ė D a are fused to form the final distance matrix D = Ė D f + α Ė D a . Using D, we identify the k-nearest neighbors for each node, forming an edge set E f and constructing the graph structure G f = V f ,E f . During the fusion stage, feature and visual distances provide complementary information to build a more robust graph structure. As shown in Fig.9, distance fusion effectively eliminates connections to areas unrelated to the disease. The loss function for VCC is: L V C =L CE (p cls ,y cls ) + Ī» align L align ,(7) where L align is a balancing coefficient for the alignment loss. p cls represents the diagnostic prediction results of VCC, which are obtained through several CGCM and GNN blocks, down- sampling, and the diagnose head. The total loss for VCC-Net is the weighted sum of VAG and VCC losses: L =L V C + Ī» V AG L V AG ,(8) where Ī» V AG is the balancing coefficient for the VAG loss. IV. EXPERIMENTS A. Dataset and Evaluation Metrics To evaluate the performance of VCC-Net, we used three datasets: two public gaze datasets, SIIM-ACR [26] and EGD- CXR [28], [29], and the self-constructed TB-Mouse dataset. TABLE I QUANTITATIVE COMPARISON BETWEEN OUR METHOD AND OTHER APPROACHES ON THE SIIM-ACR DATASET. BOLD VALUES REPRESENT THE BEST RESULTS. MethodAccāAUCā F1ā ā ResNet18 [1]83.2082.3585.25 ā ResNet50 [1] 84.0085.8183.63 ā ResNet101 [1]84.4086.1381.08 ā ViT [2]83.6084.1683.77 ā SwinT [3]84.4083.3183.69 ā ViG [4]83.2084.8982.81 ⦠M-SEN [5]84.8085.9384.03 ⦠EML-Net [6]85.2083.6585.25 ⦠DeepGaze [7]84.4085.8984.84 ⦠GA-Net18 [8]84.8071.2683.71 ⦠GA-Net50 [8]83.2070.2582.35 ⦠GA-Net101 [8] 84.8072.6884.03 ⦠EG-ViT [9]85.6075.3085.14 ⦠Ours88.4086.1288.08 ā GazeGNN [10]85.6085.1685.60 TABLE I QUANTITATIVE COMPARISON BETWEEN OUR METHOD AND OTHER APPROACHES ON THE EGD-CXR DATASET. BOLD VALUES REPRESENT THE BEST RESULTS. MethodAccāAUCā F1ā ā ResNet18 [1]71.9685.0272.17 ā ResNet50 [1]72.9086.4372.45 ā ResNet101 [1]74.7785.6374.90 ā ViT [2] 70.0985.4369.19 ā SwinT [3]71.9686.4474.39 ā ViG [4]75.7085.7175.62 ⦠M-SEN [5]78.5084.2777.45 ⦠EML-Net [6]77.5787.3375.47 ⦠DeepGaze [7]74.7787.8769.40 ⦠GazeMTL [11]78.5088.7077.90 ⦠IAA [12] 78.5090.0077.60 ⦠EffNet+G [13]77.5788.8077.00 ⦠GA-Net18 [8]77.5786.1377.47 ⦠GA-Net50 [8]78.5086.4377.91 ⦠GA-Net101 [8]79.4486.4279.20 ⦠EG-ViT [9]77.5785.5377.42 ⦠Ours 85.0591.5284.87 ā GazeGNN [10]83.1892.3082.30 The SIIM-ACR dataset includes 1,170 chest X-ray images, with corresponding gaze data; 268 of these images are pneu- mothorax. The EGD-CXR dataset comprises 1,083 chest X-ray images sourced from the MIMIC-CXR dataset [29], catego- rized into three classes: normal, congestive heart failure, and pneumonia. Each image includes gaze data. The TB-Mouse dataset contains 1,000 training images and corresponding mouse trajectories, collected as radiologists read the images, along with 2,200 test images. These images are classified into two categories: tuberculosis and normal. An experienced ra- diologist collected the mouse trajectory data using the system described in Section I-A. To evaluate the performance in diagnosis, we employed accuracy (Acc), area under the receiver operating characteristic curve (AUC), and F1 score (F1) as evaluation metrics. All experiments were conducted on an NVIDIA 3080 Ti GPU (12GB) using PyTorch. We initialized the network with pre- trained model weights from ImageNet [32] and trained it using the Adam optimizer. The training process consisted of 200 6IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2020 1.2.276.0.7230010.3.1.4.8323329.475.1517875163.123986 1.2.276.0.7230010.3.1.4.8323329.1056.1517875166.87722 1.2.276.0.7230010.3.1.4.8323329.1821.1517875169.736963 GA-NetOursImageGazeGNNResNetViGViTSwinT pneumothorax pneumothorax Fig. 4. Visualization of attention maps, generated with Grad-CAM, comparing different methods on the SIIM-ACR dataset. In the first column, the pneumothorax region is highlighted in yellow. The red areas in the attention maps reflect the focus regions of each network. The second column shows that VCC-Net achieves superior anomaly localization by accurately focusing on the pneumothorax region in comparison to other models. epochs, with a learning rate of 2Ć 10 ā4 and a batch size of 8. In Eq.7 and Eq.8, the values of Ī» align = 0.5 and Ī» V AG = 0.5 were used. Also in distance fusion, the hyperparameter α is set to 2.0. B. Results on Gaze Dataset In our experiments on the gaze datasets, we categorized methods into three groups: 1) Methods without VCā : ResNet [1], Vision Transformer [2], Swin Transformer [3], Vision GNN [4]; 2) Methods with VC during training ⦠: M-SEN [5], EML-Net [6], DeepGaze [7], GA-Net [8], EG-ViT [9]; 3) Methods with VC during inference ā : GazeGNN [10]. Table.I compares our method quantitatively with others on the SIIM-ACR dataset. VCC-Net significantly outperforms competing approaches in Acc, AUC, and F1. Specifically, compared to the best-performing EG-ViT, our model im- proves Acc by 2.80% (88.40% vs. 85.60%). Furthermore, we observe a substantial improvement in the F1, where our method outperforms GazeGNN (which utilizes VC during inference) by 2.48% (88.08% vs. 85.60%). Notably, unlike GazeGNN, our approach generates visual attention through the VAG, enhancing the modelās usability. Table.I presents a quantitative comparison of the EGD-CXR dataset. VCC- Net outperforms the first two categories of methods in Acc, AUC, and F1. When compared to the best-performing GA- Net, our method improves Acc by 5.61% (85.05% vs. 79.44%) and F1 by 5.67% (84.87% vs. 79.20%). These results con- firm that VCC-Net effectively leverages VC to guide and supervise the network. When compared to GazeGNN, our method achieves slightly lower AUC (91.52% vs. 92.30%) but outperforms GazeGNN in Acc (85.05% vs. 83.18%). This discrepancy may be attributed to GazeGNNās use of real VC during inference, which enhances sample differentiation by incorporating radiologistsā semantic information, whereas our method autonomously generates VC. Nevertheless, our AUC remains the highest among methods in the first two categories. We employed Grad-CAM [30] to visualize attention maps of different methods on the SIIM-ACR dataset, as shown in TABLE I QUANTITATIVE COMPARISON BETWEEN OUR METHOD AND OTHER METHODS ON THE TB-MOUSE DATASET. BOLD VALUES REPRESENT THE BEST RESULTS. MethodAccāAUCā F1ā ā ResNet18 [1]83.7392.2383.75 ā ResNet50 [1]88.9595.8788.96 ā ResNet101 [1]87.8696.0687.80 ā ViT [2]88.2395.6488.20 ā SwinT [3] 88.1894.9588.05 ā ViG [4] 88.4095.1788.33 ⦠M-SEN [5]89.8696.5589.85 ⦠EML-Net [6] 90.6497.2890.25 ⦠DeepGaze [7]90.4596.6890.42 ⦠GA-Net18 [8]87.1894.6587.15 ⦠GA-Net50 [8]91.0596.7690.98 ⦠GA-Net101 [8] 90.5996.9990.48 ⦠EG-ViT [9] 89.5596.6189.56 ⦠Ours 92.4197.8492.41 TABLE IV QUANTITATIVE RESULTS FROM ABLATION STUDIES ON VCC-NET COMPONENTS FOR THE SIIM-ACR AND TB-MOUSE DATASETS. BOLD VALUES REPRESENT THE BEST RESULTS. SIIM-ACR L soft L hard L aux L align Accā AUCā F1ā ā85.2085.5384.66 ā ā 86.4085.9086.58 ā86.4085.8885.97 ā ā ā86.8086.9186.32 ā87.20 87.5987.37 ā ā ā ā 88.4086.12 88.08 TB-Mouse L soft L hard L aux L align Accā AUCā F1ā ā89.8295.7289.71 ā ā 90.6897.7790.55 ā90.3297.4290.19 ā ā ā 91.0597.1790.93 ā 91.5097.0291.41 ā ā ā ā92.41 97.84 92.41 Fig.4. The first column shows the original image with the yellow mark indicating the pneumothorax region, and the red area highlighted in the other columns shows the areas AUTHOR et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING7 20150401002057 20130126000505 20121101000410 20150126001550 nodule exudation exudation exudation 20120326001486 GA-NetOursImageEG-ViTResNetViGViTSwinT 0.18 fibreexudation fibreexudation äŗę„ę§č”č”ęę£åčŗē»ę ø äøčŗéäøŗäø»ē大å°äøēćåÆåŗ¦äøåå ååøäøååēē²ē²ē¶ęē»čē¶é“å½±ļ¼ ę°é²ęøåŗäøéę§ē”¬ē»åéåē ē¶å ±å X线蔨ē°äøļ¼ę°ēęøåŗå¢ę®ē ē¶å¤§é½ ä½äŗäøę¹ 20180524003240 Fig. 5. Visualization of attention maps obtained from Grad-CAM comparing methods on the TB-Mouse dataset. In the first column, regions with exudation and nodules are marked with bounding boxes. Red areas indicate the focused regions in each networkās attention map. Compared to other methods, VCC-Net accurately localizes lesion areas. OursGround TruthImageUNetResUNetTransUNetMedNeXtU 2 Net pneumothorax pneumothorax pneumothorax Fig. 6. Comparison of visual attention generated by various methods on the SIIM-ACR dataset. Compared to others, the visual attention generated by our method closely matches that of real radiologists. The yellow circle highlights an area missed by our method that is non-abnormal on the X-ray, underscoring that VCC-Net predicted visual attention is highly aligned with actual abnormalities. the model focuses on. Methods without VC tend to focus on irrelevant regions, and even methods with VC, such as GazeGNN and GA-Net, struggle to localize abnormal areas accurately. Our model mitigates this issue, focusing more precisely on the pneumothorax region, thereby enhancing interpretability. In normal chest X-rays, VCC-Net correctly identifies the costophrenic angle, a critical anatomical land- mark. Fig.6 compares the visual attention with that from other methods. Our visual attention aligns well with the ground truth, although some regions (highlighted by the yellow circle) are missed. These areas do not show abnormalities in the image. Results suggest that VAG generates high-quality visual attention, further improving model transparency. C. Results on Mouse Dataset We compared our method against two categories of ap- proaches on the TB-Mouse dataset: methods without VC ā and with VC during training ā¦. Since GazeGNN requires real VC, which is rarely available in practice, we did not compare our method to it. Table.I shows the quantitative comparison on the TB-Mouse dataset, where our approach outperforms all others in Acc, AUC, and F1. Specifically, compared to the best-performing GA-Net50, our method improves Acc by 1.36% (92.41% vs. 91.05%) and F1 score by 1.43% (92.41% vs. 90.98%). Fig.5 illustrates the attention maps on the TB- Mouse dataset generated by Grad-CAM for different methods. The first column shows the original image with tuberculosis- related regions, such as exudation and nodules, marked with the bounding box. Areas highlighted by the network are shown in red. Compared to other methods, ours more effectively identifies the abnormal regions. Although ResNet detects some exudation areas, its localization is imprecise. In contrast, our method accurately focuses on tuberculosis-related abnormali- ties, highlighting its enhanced interpretability. D. Ablation Study 1) Effectiveness of Each Module of VCC-Net: We conducted ablation studies on the SIIM-ACR and TB-Mouse datasets to evaluate the contribution of each component in VCC-Net, as shown in Table.IV. When only the soft visual attention loss 8IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2020 Fig. 7. Ablation study on the hyperparameters Ī» align and Ī» V AG , which adjust feature alignment strength and the contribution of visual attention generation, respectively. Acc, F1, and AUC for the SIIM-ACR data are shown across different parameter settings. Optimal performance is achieved when both parameters are set to 0.5, with consistency observed across various settings, underscoring model stability. TABLE V QUANTITATIVE ABLATION RESULTS ON VAG STRUCTURE FOR THE SIIM-ACR AND TB-MOUSE DATASETS. BOLD VALUES REPRESENT THE BEST RESULTS SIIM-ACR MethodAccā AUCāF1ā VCC83.2084.8982.81 VAG (CNN) +VCC 85.6082.9485.14 VAG (GNN) +VCC87.2085.0786.14 VAG (GNN+CNN) +VCC88.40 86.12 88.08 TB-Mouse MethodAccā AUCāF1ā VCC88.4095.1788.33 VAG (CNN) +VCC 90.4596.6690.39 VAG (GNN) +VCC91.4197.2691.33 VAG (GNN+CNN) +VCC92.41 97.84 92.41 L soft was used, the Acc was 85.20% and 89.82%, respectively. IntroducingL hard , which guides the model to focus on regions emphasized by radiologists, alongside the auxiliary classifica- tion loss L aux , led to an improvement in Acc. Incorporating the alignment loss L align , which aligns feature differences with attention differences, resulted in a further Acc increase of 2.00% (87.20% vs. 85.20%) and 1.68% (91.50% vs. 89.82%) on the two datasets, respectively. This demonstrates that L align helps the model learn more effective representations by integrating cognition. Finally,L align was incorporated into both L hard and L aux , leading to improvements in accuracy, which reached 88.40% and 92.41%, respectively. A slight decrease in AUC was observed on the SIIM-ACR dataset. This may result from the hard visual attention, which tends to focus excessively on specific regions, potentially overlooking areas adjacent to anomalies. While this enhances the modelās ability to discriminate key features, it may reduce overall discrimination performance. 2) Effectiveness of VAGās Architecture: To investigate the effectiveness of the VAG architecture, we performed ablation experiments on the SIIM-ACR and TB-Mouse datasets, as shown in Table.V. Using only the original classifier, the Acc on the SIIM-ACR dataset was 83.20%. After generating visual attention with CNNs, Acc improved to 85.60%, showing that VCC can effectively incorporate VC to enhance model perfor- mance. When visual attention was generated using GNNs, Acc further increased to 87.20%. Combining GNN and CNN for attention map generation yielded the highest Acc of 88.40%, outperforming all other strategies. Paired t-tests confirmed the TABLE VI COMPARISON OF PERFORMANCE USING DIFFERENT TYPES OF VISUAL ATTENTION. BOLD INDICATES THE BEST RESULTS. SIIM-ACR MethodAccā AUCāF1ā Random82.4082.6482.64 Radiologistsā VC86.0084.8485.08 VAG 88.4086.1288.08 VAG+Radiologistsā VC88.80 88.56 88.54 TB-Mouse MethodAccā AUCāF1ā Random90.8595.5890.97 Radiologistsā VC 91.3297.7791.33 VAG 92.41 97.8492.41 VAG+Radiologistsā VC92.6497.79 92.64 statistical significance of these improvements, with p-values for Acc and F1 lower than 0.05. Experiment results support the effectiveness of VAG in integrating both global and local information for more refined attention map generation. 3) Ablation of Hyperparameters: To evaluate the impact of the hyperparameters Ī» align , Ī» V AG , and α on model per- formance, an ablation study was conducted on the SIIM- ACR dataset. Ī» align and Ī» V AG control the contributions of aligning radiologistsā visual attention with feature dif- ferences and generating visual attention maps, respectively. Fig.7 presents performance metrics for various parameter settings (0.1, 0.3, 0.5, 0.7, 1.0). The model exhibited opti- mal performance when both parameters were set to 0.5, and showed strong consistency and stability across different configurations. We analyzed the effect of varying α values (0, 0.1, 0.2, 0.5, 1.0, 2.0, 5.0, 10.0) on model performance dur- ing graph construction. When α was set too low (e.g., 0), the model depended heavily on learned features, yielding a relatively low accuracy of 83.2%. As α increased, performance improved progressively, reaching its highest accuracy of 88.4% at α = 2. Beyond that point, larger α values diminished the contribution of feature information, causing a decline in accuracy to 85.6% at α = 10. 4) The ablation of different types of visual attention: Table.VI compares the performance of using different types of vi- sual attention. Random refers to randomly sampled attention, whereas Radiologistsā VC denotes visual attention derived from radiologists. The third type represents visual attention generated by the proposed VAG, and the fourth combines the generated and radiologistsā VC through additive fusion. AUTHOR et al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING9 VAGHuman VCImage pneumothorax pneumothorax Fig. 8.Visualization reveals instances where radiologists, due to subjectivity or fatigue, misinterpret non-lesion structures as suspicious areas or overlook pneumothorax regions. The yellow-circled areas ex- emplify such occurrences. The attention produced by VAG achieved higher accuracy than VC on both datasets (88.40% vs. 86.00%, 92.41% vs. 91.32%). Such differences arise from the subjectivity and individual variability inherent in radiologistsā vision, as ra- diologists occasionally fixate on non-lesion regions perceived as suspicious. Fig.8 highlights this phenomenon, with yellow- circled areas illustrating examples of such attention. In con- trast, VAG focuses more consistently on pathology-relevant regions, yielding modest yet measurable gains in diagnostic performance. Integrating the model-generated visual attention with radiologistsā VC further improved accuracy relative to using predictions alone (88.80% vs. 88.40%, 92.64% vs. 92.41%). The results underscore a complementary relationship between radiologist and model. Radiologistsā attention pro- vides clinically spatial constraints, while the model reduces subjective bias and optimizes focus distribution, enhancing diagnostic consistency. Overall, the findings reveal the promise of collaborative mechanisms within a humanāAI collabora- tions and point toward the development of more reliable and efficient diagnostic models. E. Graph Structure and Distance Map Visualization Fig.9 demonstrates the complementary effect of attention distance and feature distance during graph construction. Two cases come from the SIIM-ACR and TB-Mouse datasets, respectively. Each image is based on a distinct distance metric, with the central node indicated by the red circle and its neighbors indicated by the yellow circle. The numbers in Fig.9 indicate the actual distance between the central node and that node. The yellow areas represent abnormal regions. Initially, the central node, influenced by both feature and visual dis- tances, often connects to irrelevant regions (e.g., background areas of chest X-rays). However, after distance fusion, these extraneous connections are effectively removed. As a result, the neighboring nodes of the central node focus more on the lung field, concentrating on potentially pathological areas. This confirms the complementarity of the two distance metrics, with the fusion strategy facilitating the construction of a graph structure rich in foreground nodes while reducing interference from irrelevant regions, thereby enhancing the transparency and interpretability of the model. 1.2.276.0.7230010.3.1.4.8323329.506.1517875163.231648.png 0 1 2 3 4 5 6 7 8 9 10 11 13 14 15 16 17 18 20 21 22 23 24 25 27 28 29 30 31 32 34 35 36 37 38 39 41 42 43 44 45 46 48 12 19 26 33 40 47 20120321001890 (a) Feature Distance (b) Visual Distance (c) Fusion Distance 0.270.170.180.130.25 0.310.25 0.32 0.400.030.370.16 0.47 0.42 0.11 0.43 0.02 0.06 0 0.09 0 0.0200.06 0.50 0.29 0.37 0.610.42 0.78 0.04 0.74 0.01 0.01 0.04 0.030.01 0 0.02 0.04 0.07 0.43 0.40 0.40 0.520.33 0.800.30 0.42 pneu motho rax Fig. 9. Visualization of graph structures under different distance met- rics. Examples from the SIIM-ACR and TB-Mouse datasets are shown. The red circle represents the central node with yellow neighboring nodes. Yellow areas highlight abnormal regions. 1.2.276.0.7230010.3.1.4.8323329.506.1517875163.231648 11 1.2.276.0.7230010.3.1.4.8323329.2174.1517875171.489135 26 20110930000084 11 20121212000521 16 0 1 2 3 4 5 6 7 8 9 10 11 13 14 15 16 17 18 20 21 22 23 24 25 27 28 29 30 31 32 34 35 36 37 38 39 41 42 43 44 45 46 48 12 19 26 33 40 47 exudation nodule exudation exudation exudation 1.2.276.0.7230010.3.1.4.8323329.1501.1517875168.127863 11 Attention Distance Feature Distance Abnormality Fusion Distance č·ē¦»čåęęę“å认ē„äøčÆä¹äæ”ęÆ ęå©äŗę“ē²¾åēčÆę ā¢čē¹č·ē¦»åę ā¢ę°čøčē¹äøę°čøåŗåč·ē¦»å° ā¢čŗéØčē¹äøčŗéåŗåč·ē¦»å° ā¢čęÆčē¹äøčŗé ā¢č·ē¦»ē»å使äøēøå ³åŗåé“ę“å čæē¦» å¾16 å¾ååé“č·ē¦»åÆč§åļ¼åØå¾ē»ęäøéę©äøäøŖčē¹ åÆč§åäøå ¶ä»čē¹ēč·ē¦»ćé¢č²č¶äŗ®č”Øē¤ŗč·ē¦»č¶å¤§ exudation exudation (b) Visual Distance(a) Feature DistanceAbnormality(c) Fusion Distance pneumothorax pneumothorax nodule Fig. 10.Visualization of distance maps based on various distance metrics. Examples from the SIIM-ACR and TB-Mouse datasets are shown. Darker colors indicate proximity to the central white point. Fig.10 further illustrates the distance between the central white point and other areas in the image, with darker shades indicating shorter distances. The two cases from the SIIM- ACR and TB-Mouse datasets display the chest X-ray along with abnormal pathological regions. Subsequent columns show distance maps for different distance metrics. In the first case, the central point (representing pneumothorax) is closer to other regions of the lung field in feature space, and the visual distance map successfully reduces the proximity to another pneumothorax area. In the second case, the central point of the exudation region is closer to other exudation and nodules in feature space, but its boundary is blurred, erroneously bringing it closer to background areas. The visual distance map increases the distance to unrelated regions, improving the construction of a more accurate graph structure. V. DISCUSSION AND CONCLUSION This paper proposes the visual cognition-guided cooper- ative network (VCC-Net) to enhance the performance and interpretability of computer-aided diagnosis. VCC-Net com- prises two components: visual attention generator (VAG) and visual cognition-guided classifier (VCC). VAG replicates the hierarchical visual search strategy of radiologists, generating attention maps to highlight critical regions. VCC constructs semantic associations between image regions and models inter-regional relationships using a graph structure, directing 10IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. X, NO. X, X 2020 the modelās focus toward key lesion areas and aligning its decision-making with the radiologistsā visual diagnostic pro- cess. Experiments on public gaze datasets (SIIM-ACR and EGD-CXR) and the self-developed TB-Mouse dataset, based on mouse trajectory, demonstrate that VCC-Net enhances diagnostic accuracy and interpretability significantly. Future work will explore the generalization of VCC-Net to additional imaging modalities and disease contexts to improve its adapt- ability. Another promising direction involves incorporating a human-in-the-loop mechanism to enable real-time radiologist feedback during inference, thereby refining decision-making and increasing clinical applicability. REFERENCES [1] K. He et al, āDeep residual learning for image recognition,ā in 2016 IEEE Conference on Computer Vision and Pattern Recognition, 2016, p. 770ā778. [2] A. Dosovitskiy et al, āAn image is worth 16x16 words: Transformers for image recognition at scale,ā in 9th International Conference on Learning Representations, ICLR, 2021. [3] Z. Liu et al, āSwin transformer: Hierarchical vision transformer using shifted windows,ā in Proceedings of the IEEE/CVF International Con- ference on Computer Vision, 2021, p. 10 012ā10 022. [4] K. Han et al, āVision GNN: An image is worth graph of nodes,ā Advances in Neural Information Processing Systems, vol. 35, p. 8291ā 8303, 2022. [5] Y. Cai et al, āMulti-task SonoEyeNet: Detection of fetal standardized planes assisted by generated sonographer attention maps,ā in Medical Image Computing and Computer Assisted Intervention. Springer, 2018, p. 871ā879. [6] S. Jia and N. D. Bruce, āEML-NET: An expandable multi-layer network for saliency prediction,ā Image and Vision Computing, vol. 95, p. 103887, 2020. [7] A. Linardos et al, āDeepGaze IIE: Calibrated prediction in and out- of-domain for state-of-the-art saliency modeling,ā in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, p. 12 919ā12 928. [8] S. Wang et al, āFollow my eye: Using gaze to supervise computer-aided diagnosis,ā IEEE Transactions on Medical Imaging, vol. 41, no. 7, p. 1688ā1698, 2022. [9] C. Ma et al, āEye-gaze-guided vision transformer for rectifying shortcut learning,ā IEEE Transactions on Medical Imaging, vol. 42, no. 11, p. 3384ā3394, 2023. [10] B. Wang et al, āGazeGNN: A gaze-guided graph neural network for chest X-ray classification,ā in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, p. 2194ā2203. [11] K. Saab et al, āObservational supervision for medical image classifi- cation using gaze data,ā in Medical Image Computing and Computer Assisted Intervention. Springer, 2021, p. 603ā614. [12] Y. Gao et al, āAligning eyes between humans and deep neural network through interactive attention alignment,ā Proceedings of the ACM on Human-Computer Interaction, vol. 6, no. CSCW2, p. 1ā28, 2022. [13] H. Zhu, S. Salcudean, and R. Rohling, āGaze-guided class activation mapping: Leverage human visual attention for network attention in chest X-rays classification,ā in Proceedings of the 15th International Sympo- sium on Visual Information Communication and Interaction, 2022, p. 1ā8. [14] X. Bai et al, āLearning from human attention for attribute-assisted visual recognition,ā IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, p. 11 152ā11 167, 2024. [15] X. Guo and Y. Yuan, āSemi-supervised WCE image classification with adaptive aggregated attention,ā Medical Image Analysis, vol. 64, p. 101733, 2020. [16] R. Gu et al, āCA-Net: Comprehensive attention convolutional neural net- works for explainable medical image segmentation,ā IEEE Transactions on Medical Imaging, vol. 40, no. 2, p. 699ā711, 2021. [17] S. He, P. E. Grant, and Y. Ou, āGlobal-local transformer for brain age estimation,ā IEEE Transactions on Medical Imaging, vol. 41, no. 1, p. 213ā224, 2022. [18] Y. Xie et al, āCoTr: Efficiently bridging CNN and transformer for 3D medical image segmentation,ā in Medical Image Computing and Computer Assisted Intervention. Springer, 2021, p. 171ā180. [19] M. Bhattacharya, S. Jain, and P. Prasanna, āRadioTransformer: A cascaded global-focal transformer for visual attentionāguided disease classification,ā in European Conference on Computer Vision. Springer, 2022, p. 679ā698. [20] M. Alsharid et al, āGaze-assisted automatic captioning of fetal ul- trasound videos using three-way multi-modal deep neural networks,ā Medical Image Analysis, vol. 82, p. 102630, 2022. [21] T. Katerina and P. Nicolaos, āMouse behavioral patterns and keystroke dynamics in End-User Development: What can they tell us about usersā behavioral attributes?ā Computers in Human Behavior, vol. 83, p. 288ā 305, 2018. [22] R. Horwitz et al, āLearning from mouse movements: Improving ques- tionnaires and respondentsā user experience through passive data collec- tion,ā Advances in Questionnaire Design, Development, Evaluation and Testing, p. 403ā425, 2020. [23] J. Xie et al, āIntegrating eye tracking with grouped fusion networks for semantic segmentation on mammogram images,ā IEEE Transactions on Medical Imaging, vol. 44, no. 2, p. 868ā879, 2025. [24] K. Z. Gajos et al, āComputer mouse use captures ataxia and parkinson- ism, enabling accurate measurement and detection,ā Movement Disor- ders, vol. 35, no. 2, p. 354ā358, 2020. [25] L. Li et al, āAttention based glaucoma detection: A large-scale database and CNN model,ā in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, p. 10 571ā10 580. [26] A.Zawackietal,āSIIM-ACRpneumothoraxsegmenta- tion,ā 2019. [Online]. Available: https://kaggle.com/competitions/ siim-acr-pneumothorax-segmentation [27] G. Li et al, āDeepGCNs: Can GCNs go as deep as CNNs?ā in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, p. 9267ā9276. [28] A. Karargyris et al, āCreation and validation of a chest X-ray dataset with eye-tracking and report dictation for AI development,ā Scientific Data, vol. 8, no. 1, p. 92, 2021. [29] A. E. Johnson et al, āMIMIC-CXR, A de-identified publicly available database of chest radiographs with free-text reports,ā Scientific Data, vol. 6, no. 1, p. 317, 2019. [30] R. R. Selvaraju et al, āGrad-CAM: Visual explanations from deep networks via gradient-based localization,ā in Proceedings of the IEEE International Conference on Computer Vision, 2017, p. 618ā626. [31] M. Nystr Ģ om and K. Holmqvist, āAn adaptive algorithm for fixation, saccade, and glissade detection in eyetracking data,ā Behavior Research Methods, vol. 42, no. 1, p. 188ā204, 2010. [32] J. Deng et al, āImageNet: A large-scale hierarchical image database,ā in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2009, p. 248ā255. [33] R. G. et al, āShortcut learning in deep neural networks,ā Nature Machine Intelligence, vol. 2, no. 11, p. 665ā673, 2020. [34] B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, āLearning deep features for discriminative localization,ā in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, June 2016. [35] X. Guo et al, āUCTNet: Uncertainty-guided CNN-Transformer hybrid networks for medical image segmentation,ā Pattern Recognition, vol. 152, p. 110491, 2024. [36] X. Zhang et al, āAn anatomy- and topology-preserving framework for coronary artery segmentation,ā IEEE Transactions on Medical Imaging, vol. 43, no. 2, p. 723ā733, 2024. [37] G. Zhao et al, āDiagnose like a radiologist: Hybrid neuro-probabilistic reasoning for attribute-based medical image diagnosis,ā IEEE Transac- tions on Pattern Analysis and Machine Intelligence, vol. 44, no. 11, p. 7400ā7416, 2022. [38] I. Pershin et al, āArtificial intelligence for the analysis of workload- related changes in radiologistsā gaze patterns,ā IEEE Journal of Biomed- ical and Health Informatics, vol. 26, no. 9, p. 4541ā4550, 2022. [39] Y. Cai et al, āSpatio-temporal visual attention modelling of standard biometry plane-finding navigation,ā Medical Image Analysis, vol. 65, p. 101762, 2020. [40] S. Wang et al, āImproving self-supervised medical image pre-training by early alignment with human eye gaze information,ā IEEE Transactions on Medical Imaging, p. 1ā1, 2025. [41] ā, āInteractive computer-aided diagnosis on medical image using large language models,ā Communications Engineering, vol. 3, no. 1, p. 133, 2024. [42] S. Wu et al, āGaze-Directed Vision GNN for mitigating shortcut learning in medical image,ā in Medical Image Computing and Computer Assisted Intervention. Springer, 2024, p. 514ā524. [43] A. Kalyanpur and N. Mathur, āApplications of artificial intelligence in thoracic imaging: a review,ā Academia Medicine, vol. 2, no. 1, 2025.