Paper deep dive
Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis
Pu Wang, Zhixuan Mao, Jialu Li, Zhuoran Zheng, Dianjie Lu, Youshan Zhang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/10/2026, 3:14:58 AM
Summary
The paper introduces a novel diagnostic framework for canine pneumothorax that combines VLM-guided Flow Matching for precise lesion segmentation with Random Matrix Theory (RMT) for spectral anomaly detection. By reframing diagnosis as a synergistic process of signal localization and statistical analysis, the model achieves superior performance on a new public pixel-level annotated dataset, effectively addressing data scarcity and the 'black box' interpretability issues of traditional deep learning models.
Entities (5)
Relation Signals (3)
VLM-FlowMatch → diagnoses → Canine pneumothorax
confidence 95% · VLM-FlowMatch reframes canine pneumothorax diagnosis as a unified signal localization and spectral analysis process.
VLM-FlowMatch → utilizes → Vision-Language Model
confidence 95% · our method employs a Vision-Language Model (VLM) to guide an iterative Flow Matching process
Random Matrix Theory → performs → Spectral Anomaly Detection
confidence 90% · We then apply Random Matrix Theory (RMT)... to analyze these features... for Spectral Anomaly Detection.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Automatic diagnosis of canine pneumothorax is challenged by data scarcity and the need for trustworthy models. To address this, we first introduce a public, pixel-level annotated dataset to facilitate research. We then propose a novel diagnostic paradigm that reframes the task as a synergistic process of signal localization and spectral detection. For localization, our method employs a Vision-Language Model (VLM) to guide an iterative Flow Matching process, which progressively refines segmentation masks to achieve superior boundary accuracy. For detection, the segmented mask is used to isolate features from the suspected lesion. We then apply Random Matrix Theory (RMT), a departure from traditional classifiers, to analyze these features. This approach models healthy tissue as predictable random noise and identifies pneumothorax by detecting statistically significant outlier eigenvalues that represent a non-random pathological signal. The high-fidelity localization from Flow Matching is crucial for purifying the signal, thus maximizing the sensitivity of our RMT detector. This synergy of generative segmentation and first-principles statistical analysis yields a highly accurate and interpretable diagnostic system (source code is available at: this https URL).
Tags
Links
- Source: https://arxiv.org/abs/2604.05482v1
- Canonical: https://arxiv.org/abs/2604.05482v1
Trouble viewing inline? Open PDF directly →
Full Text
37,203 characters extracted from source content.
Expand or collapse full text
Unifying VLM-Guided Flow Matching and Spectral Anomaly Detection for Interpretable Veterinary Diagnosis Pu Wang1,2, Zhixuan Mao1, Jialu Li3, Zhuoran Zheng4, Dianjie Lu5, Youshan Zhang6,∗ 1School of Mathematics, Shandong University, 2Shenzhen Loop Area Institute, 3Yeshiva University, 4Qilu University of Technology, 5Shandong Normal University, 6Chuzhou University, wangou@mail.sdu.edu.cn, youshan_zhang@chzu.edu.cn Abstract Automatic diagnosis of canine pneumothorax is challenged by data scarcity and the need for trustworthy models. To address this, we first introduce a public, pixel-level annotated dataset to facilitate research. We then propose a novel diagnostic paradigm that reframes the task as a synergistic process of signal localization and spectral detection. For localization, our method employs a Vision-Language Model (VLM) to guide an iterative Flow Matching process, which progressively refines segmentation masks to achieve superior boundary accuracy. For detection, the segmented mask is used to isolate features from the suspected lesion. We then apply Random Matrix Theory (RMT), a departure from traditional classifiers, to analyze these features. This approach models healthy tissue as predictable random noise and identifies pneumothorax by detecting statistically significant outlier eigenvalues that represent a non-random pathological signal. The high-fidelity localization from Flow Matching is crucial for purifying the signal, thus maximizing the sensitivity of our RMT detector. This synergy of generative segmentation and first-principles statistical analysis yields a highly accurate and interpretable diagnostic system (source code is available at: https://github.com/Pu-Wang-alt/Canine-pneumothorax). †footnotetext: * Corresponding author. This research was funded by the research project of Chuzhou University (Grant No. 2025qd36). I Introduction Canine pneumothorax is a common and potentially life-threatening emergency in veterinary clinical practice characterized by abnormal accumulation of gas in the pleural space between the lungs and the chest wall, resulting in lung collapse and severe respiratory distress [25]. Timely and accurate diagnosis is essential to guide emergency treatment and improve prognosis. At present, chest X-ray radiography is a common method for the diagnosis of canine pneumothorax. However, the interpretation of radiological images is highly dependent on the expertise and clinical experience of veterinarians. In some subtle or atypical cases, manual interpretation may be subjective, and in emergency situations, it is challenging to quickly and accurately delineate the extent of collapse for assessing the severity of the disease and making treatment plans (such as thoracocentesis). Therefore, it is of great clinical application value to develop an intelligent tool that can assist veterinarians in rapid, objective, and accurate diagnosis. Figure 1: Comparison of diagnostic approaches for canine pneumothorax. (a) The clinical challenge of subtle features. (b) The interpretability issue of ”black box” AI. (c) Our proposed framework. Recently, artificial intelligence technology represented by deep learning has made breakthroughs in the field of medical image analysis, and shows great potential, especially in lesion segmentation and classification tasks [2, 40]. In veterinary radiology, AI algorithms have been initially applied to tasks such as assessment of canine hip dysplasia [15], heart size measurement [32], and identification of certain skeletal abnormalities [27], showing great potential for improving diagnostic objectivity and efficiency. However, these traditional AI methods face two major bottlenecks. One is the extreme scarcity of large-scale, high-quality labeled data. There is a serious lack of standardized public datasets with high-quality expert annotations in the field of veterinary imaging. The construction of such a dataset is not only costly, but also requires the time of a large number of veterinary radiology experts. The second is the lack of interpretability. As illustrated in Figure 1, traditional models often function as “black boxes” that usually only provide numerical results for segmentation or classification and are unable to explain their diagnostic rationale, which limits their application in clinical decision making where a high degree of trust is required. Furthermore, distinct from general human medical imaging, veterinary radiology faces the unique challenge of extreme interspecific and interbreed anatomical variance (e.g., the thoracic cavity shape differences between a Chihuahua and a Great Dane). This high variance makes it difficult for standard supervised methods to abstract a unified “normal” representation, often leading to poor generalization. In contrast, our proposed framework addresses this by modeling the statistical properties of the signal rather than memorizing anatomical shapes, providing a transparent and trustworthy alternative by combining precise lesion localization with a quantitative anomaly score. With the development of large-scale pre-trained Foundation Models, especially large language models (LLMS) and Vision-language models (VLMS) [21], these models have gained unprecedented world knowledge and powerful zero-shot/few-shot inference capabilities through pre-training on massive multi-modal data [28, 29]. Their unique ability to understand and generate natural language opens up entirely new possibilities for building trustworthy human-computer interactive diagnostic systems [17]. Although LLM has shown great potential in the field of general human medicine, there is still a huge research gap in the highly specialized field of veterinary radiology. To address the data scarcity problem, we begin by constructing and releasing the first publicly available radiological image dataset with pixel-level expert annotations for canine pneumothorax. Based on this foundation, we propose an innovative VLM-FlowMatch segmentation framework, semantically guided lesion localization by iteratively refining an initial segmentation mask with a VLM-guided vector field. Finally, for diagnostic task, we introduce a paradigm based on RMT for anomaly detection, which quantifies the statistical perturbation from pathological signals within the focused lesion area to provide a robust Spectral Anomaly Score (SAS). I Related Work Canine medical image segmentation. Medical image segmentation is the cornerstone of computer-aided diagnosis, which aims to accurately identify anatomical structures and lesion regions at the pixel level [9]. Fully supervised deep learning models, represented by U-Net and its variants, have achieved outstanding achievements in numerous segmentation tasks and become the gold standard in this field [34]. However, the success of these models is premised on large-scale, high-quality pixel-level labeled data. In specialized fields such as veterinary radiology, the cost of obtaining such data is extremely high, severely limiting the application of fully supervised methods [45]. To address this challenge, the research community has explored a variety of data-efficient learning strategies, aiming to learn more robust features from limited labeled data. These methods include transfer learning [26], weakly supervised learning [33], and advanced techniques based on feature matching and distribution alignment [24]. Although these data-efficient methods effectively alleviate the problem of data dependence, the trained models still face two major limitations: (1) an accuracy bottleneck in fuzzy and subtle boundaries; (2) a lack of explanations for the diagnosis. Applications of Large Language Models in Medical Imaging. In recent years, large language models (LLMS) and vision-language models (VLM) have brought advances to the field of medical image analysis [42]. Although traditional deep learning models perform well on tasks such as classification or segmentation, their nature of not being able to communicate effectively with clinicians has been a major obstacle in their clinical translation. LLM has advanced logical reasoning and natural language interaction capabilities, which can transform complex pixel information into language that human doctors can understand and verify [14]. In Visual Question answering (VQA) and diagnostic AIDS, models are able to respond to natural language questions (such as Are there abnormalities in the image? ) to answer the specific content of the image, and even directly give preliminary diagnosis and classification recommendations [5]. In the automatic generation of radiology reports, the model automatically analyzes the input medical images and generates a structured and standardized diagnostic report, which can reduce the work burden of radiologists and standardize the quality of the report [1]. However, the reliability and factual accuracy of the model are still huge challenges, and sometimes it will produce plausible but inconsistent illusion [18]. Most studies use LLM as an isolated, end-of-process module that lacks intervention and insight into upstream image processing steps such as segmentation. Figure 2: Overview of our proposed synergistic framework for canine pneumothorax diagnosis. I Method VLM-FlowMatch reframes canine pneumothorax diagnosis as a unified signal localization and spectral analysis process. In Figure 2, the pipeline first employs a VLM-Infused U-Net and an Attentional Flow Matching module to generate a high-precision segmentation mask M M. This mask acts as a spatial filter to isolate the region of interest (ROI). Finally, features from the focused region are analyzed via a Random Matrix Theory (RMT)-based classifier to render a final diagnosis. I-A ViT-UNet for Initial Mask Generation Our architecture features a U-Net with a pre-trained Vision Transformer (ViT) backbone to leverage its global feature extraction for semantic understanding. For an input image X, the ViT encoder generates a visual feature map FimgF_img. To incorporate semantic guidance, we perform channel-wise multiplication between FimgF_img and a projected text feature vector FtxtF_txt (derived from prompt TpromptT_prompt), using spatial broadcasting for dimension alignment. We construct the U-Net’s skip connections by progressively upsampling this final text-fused feature map via bilinear interpolation. This forms a feature pyramid that matches the resolution of each decoder stage. The decoder then reconstructs the initial segmentation mask M(0)M^(0): M(0),Fimg,Ftxt=Ψ(X,Tprompt;θvlm-unet)M^(0),F_img,F_txt= (X,T_prompt; _vlm-unet) (1) I-B Iterative Refinement via VLM-Guided Flow Matching To refine M(0)M^(0), we learn a vector field v through flow matching. We model the refinement as a discretization of an ordinary differential equation (ODE): dxt=v(xt,t,Fcond)dtdx_t=v(x_t,t,F_cond)dt. At each timestep t, the network predicts the velocity vtv_t using a cross-attention mechanism. Cross-attention injects multimodal context to predict boundary-correcting velocities: vt=Φflow(CrossAttention(Q=f(xt),K=[Ftxt;Fimg],V=[Ftxt;Fimg])) splitv_t= _flow(CrossAttention(&Q=f(x_t),\\ K=[F_txt;F_img],V=[F_txt;F_img])) split (2) Here, the conditioning term FcondF_cond is implicitly handled by the K,VK,V projection. Training and Inference: We construct a probability path xt=(1−t)x0+tx1x_t=(1-t)x_0+tx_1 between the coarse mask x0=M(0)x_0=M^(0) and the ground truth x1=Mgtx_1=M_gt. The network vθv_θ is trained with Conditional Flow Matching (CFM) loss to regress the target velocity ut=x1−x0u_t=x_1-x_0. During inference, we solve the ODE using an Euler solver with T=10T=10 steps (dt=0.1dt=0.1). Starting from x0=M(0)x_0=M^(0), we iterate to obtain x1x_1, which is then binarized to yield the final mask M M.The comprehensive workflow of this guidance process is outlined in Algorithm 1. I-C Spectral Anomaly Detection for Diagnostic Classification We utilize Random Matrix Theory (RMT) to detect pathological signals. We isolate the ROI by Xfocus=X⊙M^X_focus=X M. The features extracted from XfocusX_focus, denoted as Fp∈ℝN×pF_p ^N× p, are standardized to zero mean and unit variance. Hypothesis Testing with RMT: Under the null hypothesis (H0H_0), we assume the standardized features of healthy tissue approximate a high-dimensional random noise system. According to the Marchenko-Pastur (MP) law, as N,p→∞N,p→∞ with aspect ratio p/N→yp/N→ y, the eigenvalues of the sample covariance matrix S=1NFpTFpS= 1NF_p^TF_p should fall within the support [λ−,λ+][ _-, _+], where λ±=(1±y)2 _±=(1± y)^2. Under the alternative hypothesis (H1H_1), pneumothorax introduces a low-rank structural signal U, modeling the covariance as a ”spiked” model: Fp=W+UF_p=W+U. This causes outlier eigenvalues to separate from the MP bulk spectrum (λi>λ+ _i> _+). We quantify this using the Spectral Anomaly Score (SAS): SAS(Xfocus)=∑λi>λ+(λi−λ+)SAS(X_focus)= _ _i> _+( _i- _+) (3) Algorithm 1 VLM Guided Flow Matching Refinement 0: Image X, Text Prompt TpromptT_prompt, Initial Mask M(0)M^(0), Steps NstepsN_steps 0: Refined Mask M M 1: Feature Extraction: 2: Fimg,Ftxt←VLM_Encoder(X,Tprompt)F_img,F_txt \_Encoder(X,T_prompt) 3: Initialization: 4: x0←M(0)x_0← M^(0) Start from coarse mask 5: dt←1/Nstepsdt← 1/N_steps 6: for k=0k=0 to Nsteps−1N_steps-1 do 7: t←k×dt← k× dt 8: Construct Query: 9: Q←FeatureExtract(xt)Q (x_t) 10: VLM Guidance (Cross Attention): 11: vt←Φflow(CrossAttn(Q,K=[Ftxt;Fimg],V=[Ftxt;Fimg]))v_t← _flow(CrossAttn(Q,K=[F_txt;F_img],V=[F_txt;F_img])) 12: ODE Solver Step (Euler): 13: xt+dt←xt+vt×dtx_t+dt← x_t+v_t× dt 14: xt+1←xt+dtx_t+1← x_t+dt 15: end for 16: M^←Binarize(x1) M (x_1) 17: return M M I-D Optimization Objective We employ a staged training strategy. The segmentation network is optimized using a hybrid loss ℒseg=ℒDice+λbceℒBCEL_seg=L_Dice+ _bceL_BCE. For diagnosis, the scalar SAS is fed into a logistic regression classifier Ψclf _clf. To address class imbalance, Ψclf _clf is trained using Focal Loss: ℒFocal(pt)=−αt(1−pt)γlog(pt)L_Focal(p_t)=- _t(1-p_t)^γ (p_t) (4) IV Experiments IV-A Dataset The dataset used in our study was sourced from the public Canine Thoracic Radiograph collection available on the Korean AI-Hub platform (https://aihub.or.kr/). To guarantee the reproducibility of our research, we partitioned this curated dataset into fixed training, validation, and test sets. Specifically, the training set comprises 8641 images, the validation set comprises 2468 images, and the test set comprises 1236 images. All model training and evaluation reported in this paper were conducted on this fixed partition to ensure fair and comparable results. The dataset exhibits a characteristic class imbalance consistent with medical screening scenarios. Specifically, the distribution follows an approximate 8:2 (4:1) ratio, where healthy cases constitute roughly 80% of the data and pneumothorax cases account for the remaining 20%. This uneven distribution necessitates the use of the F1-score as the primary metric for robust evaluation and motivates our adoption of the Focal Loss to prevent the model from biasing towards the majority class. To ensure the high quality and clinical relevance of the annotations, we established a rigorous consensus protocol. The raw data were annotated by three board-certified veterinary radiologists. In cases of disagreement regarding lesion boundaries, a senior radiologist reviewed and adjudicated the final ground truth. IV-B Implementation Details Our framework was implemented using PyTorch. The UnetFlowMatch model was trained on our training set using the Adam optimizer with an initial learning rate of 1e-4. We utilized OpenClip as the evaluation and feedback model. The prompts for the VLM were carefully designed to elicit structured refinement instructions. We utilized OpenCLIP (ViT-B/32). The visual encoder was frozen to preserve pre-trained knowledge, while the text encoder generated embeddings for the prompt: “A canine chest X-ray showing [pulmonary markings/pneumothorax]”. The VLM feature dimension is p=512p=512. The sample size N corresponds to the number of patches in the focused region (typically N≈196N≈ 196 for 224×224224× 224 input). The same VLM was used for diagnostic classification. All experiments were conducted using two NVIDIA RTX 4090 GPUs (24GB each). IV-C Quantitative results IV-C1 Comparison on Segmentation Performance Table I presents a comprehensive performance comparison between our method and a variety of state-of-the-art segmentation models. Our model consistently ranks first across all metrics on both the validation and test sets. Note that models like SAM and Swin-Transformer perform poorly (low mIoU). This is likely due to the significant domain shift between their pre-training data (natural images) and veterinary X-rays, causing them to segment the entire lung field rather than the specific pathological air pocket. Specifically, on the test set, our method achieves a top mDice of 0.8953 and mIoU of 0.8114. This performance not only surpasses classic U-Net-based architectures like PolypFlow (0.8019 mIoU) and powerful Transformer-based models like DeepLabv3+ (0.7733 mIoU), but also significantly outperforms other recent Mamba-based approaches such as Swin-UMamba (0.7820 mIoU). The consistent lead on both validation and test sets also suggests a strong generalization ability of our model. These results robustly validate the superiority of our proposed framework with its VLM-guided module. TABLE I: Performance comparison with state-of-the-art segmentation methods. Category Year Model Test Set mDice↑ mIoU↑ Unet-based 2015 Unet [34] 0.8774 0.7878 2018 Unet++ [22] 0.8712 0.7788 2025 PolypFlow [41] 0.8869 0.8019 2020 U2U^2Net [16] 0.8834 0.7965 2022 Swin-UNet [6] 0.8462 0.7424 2025 H-vmunet [44] 0.8281 0.7147 Others 2020 HRNet [19] 0.8780 0.7883 2017 SegNet [3] 0.8777 0.7880 2020 ResUnet [11] 0.8670 0.7727 2023 SAM [13] 0.6731 0.5277 Transformer-based 2018 DeepLabv3+ [7] 0.8681 0.7733 2021 Swin-Transformer [31] 0.5763 0.4165 Mamba-based 2024 Mamba-UNet [43] 0.8506 0.7481 2024 Swin-UMamba [30] 0.8733 0.7820 Ours 0.8953 0.8114 IV-C2 Comparison on Diagnostic Classification Performance We evaluated our framework on the diagnostic classification task against a comprehensive suite of baseline models [37, 35, 12, 23, 38, 8, 36, 46, 39, 10, 20, 4], as summarized in Table I. To ensure a strictly fair comparison, all baseline models were trained and evaluated on the same focused (masked) input data as our method. Specifically, the regions of interest were isolated using the segmentation masks, and the background was zeroed out for all classifiers. A key challenge of our dataset is class imbalance, making the F1-score the primary metric for robust evaluation. The impact of this imbalance is evident in baselines such as VGG16, which exhibits 100% Recall but extremely low Precision (0.2451), indicating convergence to a trivial solution of over-predicting the positive class. In contrast, the results clearly highlight the superiority of our proposed method: it avoids this failure mode, achieving the highest accuracy of 0.9032 and, more importantly, a balanced top-ranking F1-score of 0.7962. TABLE I: Performance comparison on the validation and test sets. Model Test Set Acc.↑ Prec.↑ Rec.↑ F1↑ GoogleNet [37] 0.8270 0.6787 0.5587 0.6129 VGG16 [35] 0.2451 0.2451 1.0000 0.3938 ResNet50 [12] 0.5908 0.3494 0.7769 0.4821 DenseNet201 [23] 0.8379 0.7263 0.5438 0.6219 Inceptionv3 [38] 0.8485 0.7680 0.5471 0.6390 Xception [8] 0.8582 0.8312 0.5289 0.6465 InceptionResnetV2 [36] 0.8687 0.8914 0.5289 0.6639 NasnetLarge [46] 0.8424 0.7109 0.6017 0.6517 EfficientNetB7 [39] 0.8233 0.6294 0.6793 0.6534 Vision Transformer [10] 0.6370 0.3441 0.5306 0.4174 CONVT [20] 0.4457 0.2858 0.8413 0.4267 Beit_large [4] 0.6191 0.3603 0.7140 0.4789 Ours 0.9032 0.8222 0.7719 0.7962 VGG16 and CONVT show high recall but suffer from very low precision, indicating a tendency to over-predict the positive class. Conversely, models like InceptionResNetV2 achieve high precision (0.8914) but at the expense of lower recall (0.5289). Our method, however, attains a strong balance, achieving a high precision of 0.8222 while maintaining a competitive recall of 0.7719. Figure 3: (a) This chart displays the Receiver Operating Characteristic curves. The proposed model achieves the best performance with a leading Area Under the Curve (AUC) score of 0.939. This is notably higher than other models. (b) This chart displays the Precision-Recall curves. The proposed model again shows superior performance, attaining the highest Average Precision score of 0.885. TABLE I: Ablation study of our proposed framework. Setting Model Components Seg. Performance Class. Performance Text Guidance Flow Matching RMT Input (Purification) mDice ↑ mIoU ↑ AUC ↑ F1-Score ↑ Experiment 1: Ablation on Segmentation Components (a) × × – 0.8830 0.7949 – – (b) ✓ × – 0.8736 0.7051 – – (c) ✓ ✓ – 0.8953 0.8114 – – Experiment 2: Ablation on Classification Synergy (d) ✓ ✓ Full Image – – 0.9054 0.7209 (e) ✓ ✓ Focused Image – – 0.9390 0.7962 To further assess the model’s performance across all classification thresholds, we plotted the ROC and Precision-Recall (P-R) curves, as shown in Figure 3. In the ROC analysis (Figure 3a), our model achieves a superior AUC of 0.939, indicating its strong overall discriminative ability. More importantly, given the class imbalance of our dataset, the P-R curve (Figure 3b) provides a more insightful evaluation. Our model again leads with the highest Average Precision of 0.885. Its P-R curve is positioned consistently above all others, demonstrating a robust ability to maintain high precision even as recall increases. Both metrics confirm the comprehensive superiority of our proposed framework. High-resolution images are shown in the GitHub link. Figure 4: Qualitative Comparison of Unet-based Segmentation Results. This figure presents a visual comparison of the segmentation performance of our proposed model against five Unet-based methods. IV-C3 Validation of RMT Assumptions: Empirical Spectral Distribution Analysis To empirically validate the core premise of our method—specifically that healthy tissue features follow the Marchenko-Pastur (MP) law while pneumothorax introduces outlier “spikes”. We visualized the Empirical Spectral Distribution (ESD) of the feature covariance matrices extracted from our test set. IV-D Qualitative Results To visually substantiate our quantitative findings, we provide qualitative comparisons of the segmentation results. As shown in Figure 4, while U-Net-based models can capture the general shape of the target, our method produces cleaner boundaries and more accurate contours. In contrast, our model robustly and accurately segments the target structure in all cases. These visualizations are in strong agreement with our superior quantitative metrics and demonstrate the practical effectiveness of our approach. IV-E Ablation Study We conducted a series of ablation studies to validate the effectiveness of our framework’s key components, with the results presented in Table I. Our segmentation ablation reveals that adding only VLM text guidance (b) degrades the baseline (a) performance. However, the Flow Matching module (c) is crucial for refining this raw guidance, creating a synergistic effect that significantly surpasses the baseline with an mIoU of 0.8114. The value of segmentation for classification is clear: focusing the input on the segmented lesion improved the F1-score from 0.7209 (full image) to 0.7962, confirming a strong synergistic benefit. V Conclusion In this work, we introduce a novel, interpretable framework for canine pneumothorax diagnosis and release the first accompanying public, pixel-level annotated dataset. Our method uniquely unifies VLM-guided Flow Matching for precise lesion localization with Random Matrix Theory (RMT) for diagnosis, reframing the task as the detection of statistical anomalies in purified pathological signals. This synergistic paradigm is proven to significantly outperform state-of-the-art models, offering a new path for developing trustworthy medical AI in data-scarce environments. References [1] O. Alfarghaly, R. Khaled, A. Elkorany, M. Helal, and A. Fahmy (2021) Automated radiology report generation using conditioned transformers. Informatics in Medicine Unlocked 24, p. 100557. Cited by: §I. [2] S. Asgari Taghanaki, K. Abhishek, J. P. Cohen, J. Cohen-Adad, and G. Hamarneh (2021) Deep semantic segmentation of natural and medical images: a review. Artificial intelligence review 54 (1), p. 137–178. Cited by: §I. [3] V. Badrinarayanan, A. Kendall, and R. Cipolla (2017) Segnet: a deep convolutional encoder-decoder architecture for image segmentation. IEEE transactions on pattern analysis and machine intelligence 39 (12), p. 2481–2495. Cited by: TABLE I. [4] H. Bao, L. Dong, S. Piao, and F. Wei (2021) Beit: bert pre-training of image transformers. arXiv preprint arXiv:2106.08254. Cited by: §IV-C2, TABLE I. [5] Y. Bazi, M. M. A. Rahhal, L. Bashmal, and M. Zuair (2023) Vision–language model for visual question answering in medical imagery. Bioengineering 10 (3), p. 380. Cited by: §I. [6] H. Cao et al. (2022) Swin-unet: unet-like pure transformer for medical image segmentation. In European conference on computer vision, p. 205–218. Cited by: TABLE I. [7] L. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam (2018) Encoder-decoder with atrous separable convolution for semantic image segmentation. In ECCV, p. 801–818. Cited by: TABLE I. [8] F. Chollet (2017) Xception: deep learning with depthwise separable convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1251–1258. Cited by: §IV-C2, TABLE I. [9] H. Cui, L. Hu, and L. Chi (2023) Advances in computer-aided medical image processing. Applied Sciences 13 (12), p. 7079. Cited by: §I. [10] A. Dosovitskiy, L. Beyer, A. Kolesnikov, et al. (2020) An image is worth 16x16 words: transformers for image recognition at scale. arXiv preprint arXiv:2010.11929. Cited by: §IV-C2, TABLE I. [11] D. et al. (2020) ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data. ISPRS Journal of Photogrammetry and Remote Sensing 162, p. 94–114. Cited by: TABLE I. [12] H. et al. (2016) Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 770–778. Cited by: §IV-C2, TABLE I. [13] K. et al. (2023) Segment anything. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4015–4026. Cited by: TABLE I. [14] L. et al. (2024) Snapkv: llm knows what you are looking for before generation. Advances in Neural Information Processing Systems 37, p. 22947–22970. Cited by: §I. [15] L. et al. (2025) Deep learning-based automated assessment of canine hip dysplasia. Multimedia Tools and Applications 84 (19), p. 21571–21587. Cited by: §I. [16] Q. et al. (2020) U2-net: going deeper with nested u-structure for salient object detection. Pattern recognition 106, p. 107404. Cited by: TABLE I. [17] R. et al. (2023) Contribution and performance of chatgpt and other large language models (llm) for scientific and research advancements: a double-edged sword. International Research Journal of Modernization in Engineering Technology and Science 5 (10), p. 875–899. Cited by: §I. [18] S. et al. (2024) Truth or mirage? towards end-to-end factuality evaluation with llm-oasis. arXiv preprint arXiv:2411.19655. Cited by: §I. [19] W. et al. (2020) Deep high-resolution representation learning for visual recognition. IEEE transactions on pattern analysis and machine intelligence 43 (10), p. 3349–3364. Cited by: TABLE I. [20] W. et al. (2021) Cvt: introducing convolutions to vision transformers. In Proceedings of the IEEE/CVF international conference on computer vision, p. 22–31. Cited by: §IV-C2, TABLE I. [21] Z. et al. (2024) Vision-language models for vision tasks: a survey. IEEE transactions on pattern analysis and machine intelligence 46 (8), p. 5625–5644. Cited by: §I. [22] Z. et al. (2018) Unet++: a nested u-net architecture for medical image segmentation. In International workshop on deep learning in medical image analysis, p. 3–11. Cited by: TABLE I. [23] G. Huang et al. (2017) Densely connected convolutional networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 4700–4708. Cited by: §IV-C2, TABLE I. [24] Q. Huang, X. Guo, Y. Wang, H. Sun, and L. Yang (2024) A survey of feature matching methods. IET Image Processing 18 (6), p. 1385–1410. Cited by: §I. [25] L. Jobson (2016) Nursing a canine patient with a pneumothorax—a patient care report. The Veterinary Nurse 7 (4), p. 240–244. Cited by: §I. [26] H. E. Kim, A. Cosa-Linan, N. Santhanam, M. Jannesari, M. E. Maros, and T. Ganslandt (2022) Transfer learning for medical image classification: a literature review. BMC medical imaging 22 (1), p. 69. Cited by: §I. [27] E. Kostenko, J. Šengaut, and A. Maknickas (2024) Machine learning in assessing canine bone fracture risk: a retrospective and predictive approach. Applied Sciences 14 (11), p. 4867. Cited by: §I. [28] W. Li et al. (2025) VT-fsl: bridging vision and text with llms for few-shot learning. NeurIPS. Cited by: §I. [29] W. Li et al. (2026) DVLA-rl: dual-level vision-language alignment with reinforcement learning gating for few-shot learning. ICLR. Cited by: §I. [30] J. Liu et al. (2024) Swin-umamba: mamba-based unet with imagenet-based pretraining. In International conference on medical image computing and computer-assisted intervention, p. 615–625. Cited by: TABLE I. [31] Z. Liu, Y. Lin, et al. (2021) Swin transformer: hierarchical vision transformer using shifted windows. In Proceedings of the IEEE/CVF international conference on computer vision, p. 10012–10022. Cited by: TABLE I. [32] L. P. Ramisetty (2024) Precision veterinary imaging: vertebral heart size measurement in dog x-rays with efficientnet-b7 and self-attention mechanisms. Unpublished manuscript] 2. Cited by: §I. [33] Z. Ren, S. Wang, and Y. Zhang (2023) Weakly supervised machine learning. CAAI Transactions on Intelligence Technology 8 (3), p. 549–580. Cited by: §I. [34] O. Ronneberger, P. Fischer, and T. Brox (2015) U-net: convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention, p. 234–241. Cited by: §I, TABLE I. [35] K. Simonyan and A. Zisserman (2014) Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556. Cited by: §IV-C2, TABLE I. [36] C. Szegedy, S. Ioffe, V. Vanhoucke, and A. Alemi (2017) Inception-v4, inception-resnet and the impact of residual connections on learning. In Proceedings of the AAAI conference on artificial intelligence, Vol. 31. Cited by: §IV-C2, TABLE I. [37] C. Szegedy et al. (2015) Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 1–9. Cited by: §IV-C2, TABLE I. [38] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna (2016) Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 2818–2826. Cited by: §IV-C2, TABLE I. [39] M. Tan and Q. Le (2019) Efficientnet: rethinking model scaling for convolutional neural networks. In International conference on machine learning, p. 6105–6114. Cited by: §IV-C2, TABLE I. [40] F. Wang, P. Wang, M. Zhao, C. Shan, and Z. Yang (2026) The power of modality: improving polyp segmentation with multimodal information. IET Image Processing 20 (1), p. e70305. Cited by: §I. [41] P. Wang, H. Ma, Z. Zhang, and Z. Zheng (2025) PolypFlow: reinforcing polyp segmentation with flow-driven dynamics. arXiv preprint arXiv:2502.19037. Cited by: TABLE I. [42] P. Wang et al. (2025) AgentPolyp: accurate polyp segmentation via image enhancement agent. IEEE Signal Processing Letters 32, p. 3062–3066. Cited by: §I. [43] Z. Wang, J. Zheng, Y. Zhang, G. Cui, and L. Li (2024) Mamba-unet: unet-like pure visual mamba for medical image segmentation. arXiv preprint arXiv:2402.05079. Cited by: TABLE I. [44] R. Wu, Y. Liu, P. Liang, and Q. Chang (2025) H-vmunet: high-order vision mamba unet for medical image segmentation. Neurocomputing 624, p. 129447. Cited by: TABLE I. [45] S. Xiao, N. K. Dhand, Z. Wang, K. Hu, P. C. Thomson, J. K. House, and M. S. Khatkar (2025) Review of applications of deep learning in veterinary diagnostics and animal health. Frontiers in Veterinary Science 12, p. 1511522. Cited by: §I. [46] B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le (2018) Learning transferable architectures for scalable image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, p. 8697–8710. Cited by: §IV-C2, TABLE I. Supplementary Material To offer a more granular analysis, Figure S1 displays the confusion matrices for all compared methods. The heatmap for our model (bottom right) provides a clear visualization of its balanced performance. It correctly identified 1762 negative cases (TN) and 467 positive cases (TP). More importantly, the number of misdiagnoses (False Positives, FP=101) and missed diagnoses (False Negatives, FN=138) are both effectively suppressed. This contrasts sharply with models like InceptionResnetV2, which, despite having very few FPs (FP=59, indicating a low rate of misdiagnosing healthy cases), missed a significant number of positive cases (FN=283), posing a high risk of missed diagnosis. Our framework’s ability to minimize both FN and FP demonstrates its robustness and clinical potential in handling imbalanced diagnostic data, achieving an optimal balance between identifying patients and avoiding false alarms. Figure S1: This figure presents a comparative analysis of the confusion matrices for the proposed model and twelve other models. Each matrix displays the counts for True Negatives (TN), False Positives (FP), False Negatives (FN), and True Positives (TP). The results highlight the superior performance of our model, which achieves a strong balance in correctly identifying both positive and negative instances while maintaining low error rates compared to the other methods. Figure S2: Qualitative Comparison of Segmentation Results. This figure presents a visual comparison of the segmentation performance of our proposed model against four methods. Figure S3: Qualitative Comparison of Transformer-based and Mamba-based Segmentation Results. The overall procedure, from signal purification to the final spectral anomaly scoring, is summarized in Algorithm S1. More significant performance gaps are observed against other architectural families. For instance, the general-purpose model SAM (Figure S2) and Transformer-based models like Swin-Transformer (Figure S3) largely fail on this task, producing severely fragmented or noisy results. Algorithm S1 Spectral Anomaly Detection and Diagnosis 0: Original Image X, Refined Mask M M, Theoretical Limit λ+ _+ 0: Diagnosis Label Y Y 1: Signal Purification: 2: Xfocus←X⊙M^X_focus← X M 3: Statistical Modelling: 4: Fp←VLM_Project(Xfocus)F_p \_Project(X_focus) 5: S←1NFpTFpS← 1NF_p^TF_p 6: λii=1p←EigenDecomposition(S)\ _i\_i=1^p (S) 7: Anomaly Scoring (SAS): 8: Scoresas←0Score_sas← 0 9: for each eigenvalue λi _i do 10: if λi>λ+ _i> _+ then 11: Scoresas←Scoresas+(λi−λ+)Score_sas← Score_sas+( _i- _+) 12: end if 13: end for 14: Final Diagnosis: 15: Y^←Ψclf(Scoresas) Y← _clf(Score_sas) 16: return Y Y