Paper deep dive
YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception
Marios Impraimakis, Daniel Vazquez, Feiyu Zhou
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 1:43:21 AM
Summary
This paper introduces a novel framework for interpretable object detection by coupling YOLOv10 with a Kolmogorov-Arnold Network (KAN) as a post-hoc surrogate model. The KAN uses seven geometric and semantic features to model and visualize the trustworthiness of YOLOv10 detections, providing transparency in ambiguous or degraded visual conditions. The system is further augmented with a BLIP vision-language model to provide natural language scene descriptions, creating a multimodal, interpretable perception pipeline.
Entities (4)
Relation Signals (3)
Kolmogorov-Arnold Network → modelstrustworthinessof → YOLOv10
confidence 98% · a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections
BLIP → generatescaptionsfor → Scene
confidence 95% · a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene
Kolmogorov-Arnold Network → usesfeatures → Geometric and Semantic Features
confidence 95% · model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their confidence scores in visually degraded or ambiguous scenes. To address this limitation, a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features. The additive spline-based structure of the Kolmogorov-Arnold network enables direct visualisation of each feature's influence. This produces smooth and transparent functional mappings that reveal when the model's confidence is well supported and when it is unreliable. Experiments on both Common Objects in Context (COCO), and images from the University of Bath campus demonstrate that the framework accurately identifies low-trust predictions under blur, occlusion, or low texture. This provides actionable insights for filtering, review, or downstream risk mitigation. Furthermore, a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene. This tool enables a lightweight multimodal interface without affecting the interpretability layer. The resulting system delivers interpretable object detection with trustworthy confidence estimates. It offers a powerful tool for transparent and practical perception component for autonomous and multimodal artificial intelligence applications.
Tags
Links
- Source: https://arxiv.org/abs/2603.23037v1
- Canonical: https://arxiv.org/abs/2603.23037v1
Trouble viewing inline? Open PDF directly →
Full Text
57,349 characters extracted from source content.
Expand or collapse full text
YOLOv10 with Kolmogorov-Arnold networks and vision-language foundation models for interpretable object detection and trustworthy multimodal AI in computer vision perception Marios Impraimakis Corresponding author: Marios Impraimakis (mi595@bath.ac.uk) University of Bath, Bath BA2 7AY, UK Daniel Vazquez University of Bath, Bath BA2 7AY, UK Feiyu Zhou Zhejiang University, Hangzhou 310027, China Abstract The interpretable object detection capabilities of a novel Kolmogorov-Arnold network framework are examined here. The approach refers to a key limitation in computer vision for autonomous vehicles perception, and beyond. These systems offer limited transparency regarding the reliability of their confidence scores in visually degraded or ambiguous scenes. To address this limitation, a Kolmogorov-Arnold network is employed as an interpretable post-hoc surrogate to model the trustworthiness of the You Only Look Once (Yolov10) detections using seven geometric and semantic features. The additive spline-based structure of the Kolmogorov-Arnold network enables direct visualisation of each feature’s influence. This produces smooth and transparent functional mappings that reveal when the model’s confidence is well supported and when it is unreliable. Experiments on both Common Objects in Context (COCO), and images from the University of Bath campus demonstrate that the framework accurately identifies low-trust predictions under blur, occlusion, or low texture. This provides actionable insights for filtering, review, or downstream risk mitigation. Furthermore, a bootstrapped language-image (BLIP) foundation model generates descriptive captions of each scene. This tool enables a lightweight multimodal interface without affecting the interpretability layer. The resulting system delivers interpretable object detection with trustworthy confidence estimates. It offers a powerful tool for transparent and practical perception component for autonomous and multimodal artificial intelligence applications. Keywords: Autonomous vehicles, Interpretable object detection, Kolmogorov-Arnold Network (KAN), You Only Look Once (YOLOv10), Trustworthy Explainable AI (XAI) 1 Introduction Object detection systems have attracted interest through the development of the one-stage You Only Look Once models. These detectors achieve real-time performance, making them attractive for deployment in visual perception tasks in autonomous vehicles [31, 89, 70, 67] , traffic [87], and pedestrian detection [48, 40]. To this end, Wang et al. [72] proposed to train the You Only Look Once models without a filtering step for faster and more accurate behavior. Dong et al. [19] developed a method that uses wavelet-based processing, improved feature pooling, and a combined loss design. Luz et al. [14] showed that combining Internet of Things devices, edge computing, and deep learning enables vehicle detection across hardware with widely varying processing speeds. Wang et al. [73] improved the detection of tiny objects by mixing information from different scales and surrounding context [32]. Haoyan et al. [23] proposed a lightweight model for steel‑surface defect detection by replacing standard layers with an adaptive downsampling structure [3, 7]. Chhimpa et al. [11], later, demonstrate that fine-tuning an object detection model can effectively identify brain tumors in medical images. An et al. [4] introduced also a small-target model which improves crack detection using a new head structure, a detail-enhancing pyramid, and a loss function. Zhang et al. [88] proposed an improved network specifically for aerial vehicle detection, Ding et al. [18] developed a model that uses attention and multi-scale processing to detect small objects from drone imagery, while Chen et al. [10] added an edge-focused feature module, improved multi-scale processing, and pruning and distillation steps. Sun et al. [65] also combined a transformer-based backbone with an attention-driven feature-fusion method to capture tiny details. Srinivasu et al. [63] studied how adjusting model settings and expanding the training data affects the detection performance. Mahapadi et al. [43] later optimized two nature-inspired search algorithms. Along these lines, Qin et al. [51] proposed a classroom-behavior detection system which improves recognition under crowded and multi-scale conditions. Furthermore, Xie et al. [80] described a model for autonomous-driving detection [76, 37, 20, 75, 35, 46, 30] that uses a new gating module, an improved feature pyramid, a redesigned multi-scale network, and a double-distillation training strategy. Liu et al. [39] introduced an enhanced model designed to locate defects in aircraft-engine components. Finally, Meng et al. [44] combine ideas from pose-estimation models for agricultural automation. Toward Kolmogorov-Arnold networks proposed by Liu et al. [41], Shuai et al. [59] introduced later a physics-informed framework that embeds scientific laws into the learning process [1, 33, 69, 68]. Wu et al. [79] presented Graph Kolmogorov-Arnold networks [36], while Howard et al. [25] developed multifidelity Kolmogorov-Arnold networks that use inexpensive low-quality data. Hasan et al. [24] combined sequence-learning networks to analyze brain-signal time series, while Baravsin et al. [6] conducted a study across over one hundred time-series datasets for real-world signals. Furthermore, Wang et al. [74] proposed convolutional Kolmogorov-Arnold networks, which aim to create smaller and more interpretable intrusion-detection models. Saravani et al. [57] compared Kolmogorov–Arnold networks with machine-learning approaches, Ansar et al. [5] modified the loss functions of Kolmogorov-Arnold networks incorporating a correlation measure to improve performance, while Ta et al. [66] introduced fully-composed Kolmogorov-Arnold networks that mix mathematical basis functions to learn complex relationships from low-dimensional inputs. Finally, Panahi et al. [49] proposed a general discovery framework using Kolmogorov-Arnold networks that can uncover governing equations from data. To the vision-language modeling end, Xue et al. [83] introduced a curated dataset, training procedures, and a family of large models that handle both images and text [8]. Yang et al. [84] presented VisionZip, a method that keeps only the most useful visual tokens, Xu et al. [82] proposed a model with chain-of-thought that performs its own multi-stage reasoning steps, while Vasu et al. [71] introduced a new vision encoder that produces fewer visual tokens to greatly speed up processing. To this end, Deng et al. [17] studied how vision-language models balance visual and textual information when they are given images and different kinds of text inputs, Zhou et al. [90] introduced the first benchmark fot both space and time in videos, while Shu et al. [58] proposed video-extra-large which summarizes the content of each time segment. Deitke et al. [16] presented some newly created datasets, Zhan et al. [86] also curated a large remote-sensing instruction dataset [38], and Hu et al. [26] created a high-quality remote-sensing image-captioning dataset for aerial images. Furthermore, Wu et al. [78] explored segmenting farmland applications, Dai et al. [15] investigated the interactions between people and their environments, and Xin et al. [81] examined a decomposed three-dimensional encoder, an improved image-text alignment method, and a dual-stream projector for medical images with text descriptions. Related to interpretable artificial intelligence in engineering [42, 29, 28, 47, 34], Singh et al. [60] introduced DiaXplain, a diabetes-diagnosis tool to deliver understandable results, Roy et al. [55] developed a system that emphasizes transparency for clinicians, while Yao et al. [85] proposed a method that detects interpretable discharge states. Chudasama et al. [12] introduced TrustKG, a knowledge-graph-based approach that improves the interpretability and reliability of hybrid medical systems. Gliner et al. [21] demonstrated a tool for electrocardiogram images, Song et al. [62] analyzed climate and lightning data to understand which conditions increase the likelihood of forest fire-related lightning [52]. Rashid et al. [53] presented an ensemble-based disease‑detection method, while Salvi et al. [56] combined uncertainty estimation with interpretability methods. Finally, Agrawal et al. [2] explored explanation techniques to improve the clarity of model reasoning. Related to trustworthy artificial intelligence, Chander et al. [9] described a range of modern relevant techniques with Wirz et al. [77] argued that trust in outputs depends both on how users perceive the system and on the specific decision situation. Cousineau et al. [13] considered findings from interviews with developers showing that creating trustworthy artificial-intelligence systems is difficult due to immature standards, inconsistent regulations, unclear definitions, and limited practical tools. Stamboliev et al. [64] analyzed European Union policy discourse and concluded that promoting trustworthy artificial intelligence serves as a shared idea that unites diverse political and industry stakeholders. Reinhardt et al. [54] examined the philosophical views of trust, Goisauf et al. [22] provided an interdisciplinary analysis of trust, while Jannel et al. [27] proposed the concept of trustability as a threshold that distinguishes appropriate trust in these systems. Parra et al. [50] introduced the REASON framework, which aims to integrate trustworthiness into the full life-cycle of future communication networks, and Moreno et al. [45] proposed a design approach intended to help developers embed trustworthy principles into medical systems. However, the internal behavior of such models remains difficult to interpret: the learned decision surfaces are highly nonlinear, their dependence on geometric and semantic features is opaque, and their confidence outputs offer limited insight into why detections are accepted, rejected, or assigned particular confidence levels. To this end, a framework that couples You Only Look Once with a Kolmogorov-Arnold networks is examined to model and interpret the internal confidence predictions of the detector. The focus is on the structured seven-dimensional set of interpretable features available at the output of the detection head: normalized bounding-box position, normalized size, predicted confidence, discrete class index, and relative image scale. Additionally, natural language explanations also offer complementary insight by aligning model outputs with human linguistic reasoning. The combination of these models results in a unified interpretability pipeline: You Only Look Once performs object detection, Kolmogorov-Arnold network provides a transparent surrogate that exposes the structure of the model’s confidence function, and the visual-language model produces natural language descriptions. This multimodal design [61] brings together visual, numerical, and linguistic interpretability within a single coherent framework. The remainder of this paper is organized as follows. Section 2 presents the You Only Look Once detector that provides the structured inputs for surrogate modeling. Section 3 introduces the Kolmogorov-Arnold network surrogate, detailing its mathematical foundation, spline-based formulation, and its role in approximating the model’s confidence function using interpretable univariate transformations. Section 4 describes the vision-language foundation model component, explaining how natural-language explanations are generated to complement the numerical interpretability offered by the surrogate. Section 5 presents the experimental results, including feature-level analyses, partial-dependence behavior, hidden-unit specialization, edge-importance structure, fidelity evaluation, and qualitative examples demonstrating the interpretability pipeline on the COCO and the University of Bath campus’ photos dataset. Section 6 provides a discussion of the interpretability findings, while Section 7 concludes the work. 2 Modeling confidence by You Only Look Once You Only Look Once is a one-stage object detector designed to perform real-time bounding-box prediction that directly regresses object locations and class probabilities seen in Fig. 1. Given an input image I∈ℝH×W×3I ^H× W× 3, the model computes a hierarchy of feature maps through a backbone network, producing multiscale feature tensors [F1,F2,F3][F_1,F_2,F_3] that encode semantic and spatial information. At each spatial location on each feature map, the model predicts a bounding box parameterized by its center and dimensions: b=(x,y,w,h)b=(x,\,y,\,w,\,h) (1) together with a score o∈[0,1]o∈[0,1] and a distribution over classes p(c|I)p(c|I). The final confidence for a detection is computed as: conf=o⋅p(c|I)conf=o· p(c|I) (2) which serves as the scalar quantity for the Kolmogorov-Arnold Network, Section 3. The network receives an input image of size 640×640×3640× 640× 3, which is processed by the backbone network to extract hierarchical visual features. The backbone primarily consists of convolutional layers and coarse-to-fine modules that progressively learn spatial and semantic representations from the input image. These extracted features are then passed to the neck component to aggregate multi-scale contextual information and enhance feature fusion. The refined feature maps are subsequently forwarded to the detection head, before the model produces a set of dense predictions corresponding to potential object locations and class probabilities across the image, resulting in approximately 8400 candidate detections. Figure 1: Architectural framework of You Only Look Once. C2f stands for coarse-to-fine, SPPF stands for spatial pyramid pooling fast, and PSA stands for partial self-attention. Let (x^,y^,w^,h^)( x, y, w, h) denote the network’s predicted box and (x∗,y∗,w∗,h∗)(x^*,y^*,w^*,h^*) denote the ground truth box assigned to that feature-map location, the model employs a differentiable intersection over union (IoU) metric-based loss function as: Lbox=1−IoU((x^,y^,w^,h^),(x∗,y∗,w∗,h∗))L_box=1-IoU\! (( x, y, w, h),\,(x^*,y^*,w^*,h^*) ) (3) The training objective is expressed as: L=λboxLbox+λobjLobj+λclsLclsL= _boxL_box+ _objL_obj+ _clsL_cls (4) where λbox,λobj,λcls _box, _obj, _cls are balancing coefficients. Each detection consists of the tuple (x,y,w,h,conf,c)(x,y,w,h,conf,c), from which seven numerical features are used by the Kolmogorov-Arnold Network of Section 3, seen as a framework in Fig. 2. These are the normalized spatial position (x,y)(x,y), the normalized size (w,h)(w,h), the predicted confidence confconf, the class index c, and the relative image scale s=wh6402s= wh640^2. Figure 2: Framework for interpretable learning in trustworthy detection perception. 3 Surrogate Kolmogorov-Arnold network modeling The Kolmogorov-Arnold network seen in Fig. 3 is inspired by the Kolmogorov-Arnold representation theorem, which states that any multivariate continuous function f:ℝn→ℝf:R^n can be expressed as a finite superposition of univariate continuous functions. Kolmogorov proved that there exist continuous functions ϕq _q and ψp,q _p,q such that: f(x1,…,xn)=∑q=12n+1ϕq(∑p=1nψp,q(xp))f(x_1,…,x_n)= _q=1^2n+1 _q\! ( _p=1^n _p,q(x_p) ) (5) Kolmogorov–Arnold networks operationalize this idea by replacing the learned ψp,q _p,q functions with trainable spline functions sj,k(⋅)s_j,k(·) to every connection between input feature k and hidden unit j, and replacing ϕq _q with linear combinations of hidden-unit activations as: hj(x)=∑k=1dsj,k(xk)h_j(x)\;=\; _k=1^ds_j,k(x_k) (6) where d is the input dimensionality. The model output is then obtained as: y(x)=∑j=1mwjhj(x)y(x)\;=\; _j=1^mw_j\,h_j(x) (7) where m is the number of hidden units and wjw_j are learned output weights. The model input signals are fed into a hidden layer in which each neuron applies a learnable univariate spline function, shown in the diagram as a white spline curve inside each hidden unit. The outputs of the hidden spline neurons are subsequently combined by a standard output neuron to produce the final prediction. Figure 3: Architecture of a Kolmogorov-Arnold network. In this work, the surrogate confidence model is a Kolmogorov-Arnold network with architecture width=[7,16,1], grid=5 and spline order=3 (cubic b-splines), meaning that the network receives seven interpretable numerical features derived from detections. These features consist of normalized bounding–box geometry (x,y,w,h)(x,y,w,h), predicted confidence, discrete class index, and relative image scale (wh/6402)(wh/640^2). Finally,, the model learning form is: c^(x)=∑j=116wj(∑k=17sj,k(xk)) c(x)= _j=1^16w_j ( _k=1^7s_j,k(x_k) ) (8) where x∈ℝ7x ^7 is the vector of You Only Look Once-derived features. This architecture directly enables the interpretability results. Partial-dependence plots arise naturally from evaluating a single spline while holding other dimensions fixed. Specifcically, monotonicity emerges from the smoothness properties of the fitted splines, hidden-unit roles can be inferred from correlations between spline outputs and input features, and input-hidden edge importance corresponds to norms of learned spline coefficients. Importanlty, each input–hidden connection uses a learnable spline made from B‑spline basis functions with grid = 5 and spline order = 3, each with its own trainable coefficient. During training, these coefficients adjust the curve shape, enabling smooth and interpretable nonlinear mappings while keeping the model simple and stable, seen in Fig. 3. 4 Text augmentation by vision-language foundation modeling Large vision-language foundation models aim to learn a joint representation between images and natural language, seen in Fig. 4. An input image I and a textual sequence T=(t1,…,tL)T=(t_1,…,t_L) are embedded into a unified representation followed by a joint conditioning module: H=FVL(Zimg,Ztxt)H=F_VL (Z_img,Z_txt ) (9) where, FVLF_VL consists of stacked transformer blocks, and ZimgZ_img and ZtxtZ_txt are the unified representations of the input. The generative language model then predicts a probability distribution over the vocabulary conditioned jointly on vision and text as: p(ti+1|t1,…,ti,I)=Softmax(Whi)p(t_i+1\,|\,t_1,…,t_i,I)=Softmax\! (W\,h_i ) (10) where, hih_i is the decoder hidden state produced from the fused representation H. The input image is first processed by the processor, which performs the required tokenisation, normalisation, and embedding transformations. These processed embeddings are then passed to a vision encoder and a text decoder, which jointly generate a natural‑language caption describing the visual content of the image. Figure 4: Architectural framework of visual-language models. BLIP stands for Bootstrapped Language-Image Pretraining. Bootstrapped language-image pretraining optimizes a contrastive objective of the form: LITC=−logexp(sim(Zimg,Ztxt)/τ)∑T′exp(sim(Zimg,Ztxt′)/τ)L_ITC=- \! (sim (Z_img,Z_txt )/τ ) _T \! (sim (Z_img,Z_txt )/τ ) (11) where sim(⋅,⋅)sim(·,·) is a cosine-similarity operator and τ is a temperature parameter. The model functions as a language-based interpretability layer that complements the numerical explanations provided by the Kolmogorov-Arnold network of Section 3. For each detection image I, BLIP receives the visual input and produces a caption T T by sampling from the conditional distribution p(T|I)p(T|I) given by the decoder. 5 Application to vehicle detection tasks The modeling framework is examined on the Common Objects in COntext dataset focused on vehicle detection tasks. The goal is to characterize the models internal confidence behavior using the Kolmogorov-Arnold network, and to demonstrate how the spline-based structure provides transparent explanations of the detector’s geometric and semantic sensitivities. The model, shown in Fig. 5, receives seven input features derived from You Only Look Once detections; bounding-box geometry, predicted confidence, class index, and image scale, and outputs a smooth approximation of the model’s confidence. The darker edges in the schematic indicate stronger learned importance. The mean absolute spline activation per input feature is shown in Fig. 6; here, the class index and confidence values dominate the activation magnitudes, while geometric terms (x,y,w,h)(x,y,w,h) exhibit weaker but consistent influence. Figures 7-10 present the partial dependence plots for individual features, demonstrating how the Kolmogorov-Arnold network output varies when only one input is changed. The dependence on bounding-box width (Fig. 7) shows a smooth monotonic increase, indicating that the model assigns higher confidence to wider objects. Height (Fig. 8) yields a similarly smooth but slightly weaker trend. The dependence on vertical position y (Fig. 9) is shallow with mild curvature, consistent with weak geometric sensitivity. Image scale (Fig. 10) exhibits only a small positive effect, confirming that detection confidence is not strongly biased across image resolutions. Across all plots, the curves are smooth and stable, illustrating that the spline components of the surrogate generalize without overfitting. Figures 11 and 12 show the feature-unit correlation bars. They provide fine-grained interpretability of the internal structure of the surrogate. The full input–hidden edge importance heatmap is also shown in Fig. 13, revealing that class and confidence supply strong signals to specific neurons, while geometric inputs connect more weakly. Figures 14-16 show the collection of spline functions learned by the Kolmogorov-Arnold network for all feature–unit pairs. These visualizations demonstrate the smoothness and diversity of the learned transformations. Each figure corresponds to one group of spline subplots. Finally, Figures 17–23 illustrate the application of the full system to real scenes. You Only Look Once detections are shown alongside Kolmogorov-Arnold network-predicted confidence estimates and vision-language captions. In high-quality images such as the truck (Fig. 17), police car (Fig. 18), and sand-car scenes (Fig. 19), the surrogate and detector agree closely. For more challenging examples featuring occlusions, blur, or clutter (Figs. 20–23), the Kolmogorov-Arnold network highlights regions where confidence becomes less predictable, demonstrating the surrogate’s ability to expose when detection reliability decreases. The numerical interpretability complements the linguistic explanations generated by the vision-language model. Figure 5: Kolmogorov-Arnold network model architecture with seven interpretable inputs and one output. Figure 6: Mean absolute spline activation per input feature. Figure 7: Partial dependence of the Kolmogorov-Arnold network on bounding-box width. Figure 8: Partial dependence on bounding-box height. Figure 9: Partial dependence on vertical position y. Figure 10: Partial dependence on image scale. Figure 11: Feature influence on a representative hidden unit. Figure 12: Feature influence on a different hidden unit. Figure 13: Input–hidden edge importance heatmap. Figure 14: Spline functions for all feature - unit pairs (Group 1, horizontal-axis: feature value, vertical‑axis: spline activation). Figure 15: Spline functions for all feature - unit pairs (Group 2). Figure 16: Spline functions for all feature - unit pairs (Group 3). Figure 17: Application to a truck scene. Figure 18: Application to a police car. Figure 19: Application to a vehicle in sand. Figure 20: Application to a van example. Figure 21: Application to a challenging scene with partial occlusion and blur. Figure 22: Application to a cluttered scene with multiple objects. Figure 23: Application to a scene with a tree stump and surrounding clutter. Table 1: Feature-level statistics from the network. Feature SplineActivation Saliency PD Delta x 0.107281804 1.43×10−6× 10^-6 −0.007746-0.007746 y 0.106276410 2.64×10−6× 10^-6 −0.025138-0.025138 w 0.061890736 4.17×10−6× 10^-6 0.0569680.056968 h 0.068972364 2.35×10−6× 10^-6 −0.045587-0.045587 conf 0.191350000 1.06×10−4× 10^-4 0.7069390.706939 cls 3.731932000 1.67×10−6× 10^-6 0.0297240.029724 scale 0.082707120 2.77×10−6× 10^-6 −0.0269816-0.0269816 Table 2: Hidden-node statistics from the network. Node Activation Importance Feature Correlation n0 4.908808 2.2996383 cls 0.999684334 n1 3.8972838 0.19116642 cls -0.999700487 n2 2.4302435 0.10419141 cls -0.996881783 n3 6.031918 3.575461 cls 0.999689877 n4 5.654485 3.7371411 cls 0.999682188 n5 6.2070827 5.189023 cls 0.999588966 n6 5.103179 0.22900744 cls -0.999275625 n7 0.48421517 0.27280924 cls -0.947130859 n8 1.4790272 0.48364034 cls -0.993801713 n9 0.6606282 0.2996985 conf -0.958717942 n10 5.8533688 2.5290031 cls 0.999058664 n11 9.327825 0.38113362 cls -0.999662876 n12 2.3340948 0.15058956 cls -0.997766674 n13 0.4784846 0.076826476 cls 0.985139787 n14 2.45556 0.08944808 cls -0.999845147 n15 2.6013832 1.7512177 cls 0.997587681 Table 3: Input-hidden feature influence values for each Kolmogorov-Arnold network hidden unit. Node x y w h conf cls scale n0 0.029854016 0.015423812 0.002957478 0.014497110 0.046715240 2.295937500 0.005541976 n1 0.000579606 0.002248535 0.001634980 0.001924950 0.003963163 0.191089600 0.000456766 n2 0.000505880 0.002597063 0.002021160 0.002741036 0.005081795 0.103393555 0.000257480 n3 0.011441349 0.020771712 0.031916957 0.032681603 0.085905920 3.580443100 0.004258949 n4 0.023302530 0.028058290 0.026650915 0.024421094 0.055909153 3.746629200 0.022205167 n5 0.039337154 0.018811513 0.039905798 0.024799882 0.136980850 5.175523300 0.015163474 n6 0.003420651 0.000805297 0.001067619 0.001793040 0.006243926 0.227985070 0.002086347 n7 0.012705687 0.014479717 0.004435883 0.003144299 0.086433806 0.245138540 0.001393769 n8 0.006420781 0.012792237 0.009923225 0.006143517 0.056687247 0.489789370 0.000277586 n9 0.045870207 0.021616014 0.054148242 0.003510625 0.305599600 0.025778690 0.019549502 n10 0.021460332 0.017167496 0.017201278 0.029365148 0.095861495 2.544660300 0.007906381 n11 0.000112305 0.002232117 0.001292159 0.003263123 0.007255736 0.379950200 0.000122972 n12 0.002334908 0.001076487 0.000819806 0.001019987 0.009017214 0.148895730 0.000324138 n13 0.009948854 0.005338817 0.003623644 0.006526013 0.003195604 0.074584790 0.000295031 n14 0.000178230 0.001063958 0.000743822 0.000746206 0.000768538 0.089053730 0.000452965 n15 0.019759016 0.017285729 0.023938041 0.010737362 0.115140710 1.733505800 0.001825250 Table 4: Feature influence metrics combining spline activation, saliency, partial dependence delta, and edge importance. Feature SplineAct Saliency PDP_Delta EdgeImportance Influence x 0.10728 2.86×10−6× 10^-6 −0.0077-0.0077 0.22723 0.01739 y 0.10628 5.29×10−6× 10^-6 −0.0251-0.0251 0.18177 0.01392 w 0.06189 8.34×10−6× 10^-6 0.05697 0.22228 0.04231 h 0.06897 4.70×10−6× 10^-6 −0.0456-0.0456 0.16731 0.00371 conf 0.19135 2.12×10−4× 10^-4 0.70694 1.02076 0.52001 cls 3.73193 3.33×10−6× 10^-6 0.02972 21.05236 0.52559 scale 0.08271 5.55×10−6× 10^-6 −0.02698-0.02698 0.08212 0.01082 Table 5: Kolmogorov-Arnold network fidelity across overall data and per-feature quantile bins. Scope Feature BinIndex BinRange N R2 MAE RMSE Overall 9098 0.99568467 0.010857625 0.01510572 Per-Feature x 0 [0.0078, 0.2581] 1820 0.995543874 0.010738888 0.014922408 Per-Feature x 1 [0.2581, 0.4369] 1819 0.995584850 0.011135795 0.015436348 Per-Feature x 2 [0.4369, 0.5569] 1820 0.995243065 0.011685996 0.016162014 Per-Feature x 3 [0.5569, 0.7306] 1819 0.995711381 0.010715223 0.015072694 Per-Feature x 4 [0.7306, 0.9956] 1820 0.996162495 0.010012295 0.013840625 Per-Feature y 0 [0.0090, 0.3730] 1820 0.993737923 0.012518246 0.017162563 Per-Feature y 1 [0.3730, 0.4995] 1819 0.996216900 0.010224011 0.014291467 Per-Feature y 2 [0.4995, 0.5889] 1820 0.996546371 0.009600522 0.013792516 Per-Feature y 3 [0.5889, 0.7164] 1819 0.995977398 0.010252580 0.014582687 Per-Feature y 4 [0.7164, 0.9839] 1820 0.995056666 0.011692083 0.015466231 Per-Feature w 0 [0.0078, 0.0633] 1820 0.994522331 0.009916774 0.013265161 Per-Feature w 1 [0.0633, 0.1175] 1819 0.996003601 0.009420515 0.013088631 Per-Feature w 2 [0.1175, 0.1942] 1820 0.996316109 0.009662537 0.013505891 Per-Feature w 3 [0.1942, 0.3645] 1819 0.995564989 0.010814356 0.015195336 Per-Feature w 4 [0.3645, 1.0000] 1820 0.992965135 0.014473127 0.019501280 Per-Feature h 0 [0.0134, 0.0996] 1820 0.993620586 0.010833444 0.014169815 Per-Feature h 1 [0.0996, 0.1767] 1819 0.994757525 0.011014964 0.014902389 Per-Feature h 2 [0.1767, 0.2982] 1820 0.994644570 0.011212999 0.015743511 Per-Feature h 3 [0.2982, 0.5189] 1819 0.994923902 0.011166411 0.015544278 Per-Feature h 4 [0.5189, 1.0000] 1820 0.995195339 0.010060562 0.015118543 Per-Feature conf 0 [0.2501, 0.3561] 1820 0.646103634 0.013054217 0.018254275 Per-Feature conf 1 [0.3561, 0.5248] 1819 0.908017482 0.010815061 0.015045941 Per-Feature conf 2 [0.5248, 0.7179] 1820 0.953721345 0.009243234 0.012182652 Per-Feature conf 3 [0.7179, 0.8677] 1819 0.880090188 0.011291197 0.014700041 Per-Feature conf 4 [0.8677, 0.9854] 1820 0.684149733 0.009884628 0.014724099 Per-Feature cls 1 [0.0000, 3.0000] 3558 0.999025715 0.005320589 0.007107449 Per-Feature cls 2 [3.0000, 29.0000] 1878 0.992935624 0.014938466 0.019866999 Per-Feature cls 3 [29.0000, 52.0000] 1833 0.992111145 0.014669338 0.018721599 Per-Feature cls 4 [52.0000, 79.0000] 1829 0.994633838 0.013618740 0.016740484 Per-Feature scale 0 [0.0745, 0.6250] 1800 0.995370806 0.011333028 0.015715895 Per-Feature scale 1 [0.6250, 0.6672] 1303 0.995758532 0.010512075 0.014713371 Per-Feature scale 2 [0.6672, 0.7375] 2343 0.995833673 0.010658501 0.014828051 Per-Feature scale 3 [0.7375, 0.7500] 225 0.995384340 0.011505265 0.015818063 Per-Feature scale 4 [0.7500, 1.0000] 3427 0.995732202 0.010832923 0.015066167 Table 6: Monotonicity analysis of each input feature. Feature MonotonicityScore Direction Strength x 0.021862028 Flat/Weak Weak y 0.034401234 Flat/Weak Weak w 0.376517522 Positive Moderate h 0.449024012 Positive Moderate conf 0.996658272 Positive Strong cls -0.146067093 Negative Weak scale 0.022657916 Flat/Weak Weak Table 1 summarizes the per‑feature statistics extracted from the Kolmogorov-Arnold network. The mean spline activation identifies class and confidence as the dominant drivers of the model’s confidence behavior, while the partial dependence deltas quantify each feature’s global effect across its domain. Geometric inputs play a secondary role, and the uniformly small saliency values demonstrate that the network remains smooth and stable, without sensitivity spikes or overfitting artefacts. Table 2 provides a detailed description of hidden-unit behavior within the Kolmogorov-Arnold network. Strong correlations between hidden-unit activations and the class feature reveal a consistent specialization pattern, with a few neurons carrying disproportionately large influence over the output. This structure supports a transparent interpretation of the model, where individual units serve identifiable semantic functions rather than mixing signals in an opaque manner. Table 3 presents the complete input–hidden edge importance matrix computed from spline coefficients. The results show strong connections from class and, to a lesser extent, confidence, feed into specific hidden units, while geometric features contribute weaker. Table 4 consolidates all interpretability metrics into a unified feature-influence score. The strong agreement between normalized statistics confirms that class and confidence dominate the model’s confidence landscape from multiple perspectives; activation strength, global effect, and structural edge contribution. This agreement reinforces the internal coherence of the netowork and provides a robust grounding for the qualitative partial dependence plots analyses. Table 5 evaluates the Kolmogorov-Arnold network fidelity across quantile bins for each feature. The model maintains stable fidelity across most of the feature space, demonstrating that the additive spline formulation accurately captures the functional relationship between the model’s confidence and its input parameters. The slight reductions in fidelity for the lowest and highest confidence bins align with the expected difficulty of modeling edge-case detections. Finally, Table 6 shows the monotonicity scores, offering a global characterization of the directionality of the model’s confidence behavior. Confidence emerges as an almost perfectly monotonic driver, providing clear interpretability, while geometric dimensions exhibit moderate monotonic trends. Class shows a weak negative correlation due to index encoding, but combined with the other metrics, its importance remains clearly established. 6 Discussion The interpretability analysis produced by the Kolmogorov-Arnold network reveals several consistent patterns. First, while the full suite of interpretability tables provides detailed quantitative insights, not all tables contribute unique information. The unified feature influence table is the most essential as it captures both the spline-activation behaviour and the structural edge importance values, effectively integrating the core components of two other tables. In contrast, the hidden-unit role analysis is crucial because it exposes how individual units specialize in response to class, confidence, or weaker geometric signals, providing a clear view of how the surrogate decomposes the model’s confidence function. The monotonicity scores likewise add interpretive value, showing that the most influential features also tend to exhibit strong or moderate monotonic relationships with the surrogate output. Together, these analyses confirm that the Kolmogorov-Arnold network approximation is not only accurate but also structurally coherent, with meaningful contributions distributed across identifiable units and feature interactions. More specifically, the conceptual diagrams show that the surrogate processes each input feature through dedicated spline transformations, highlighting the interpretability advantage of additive univariate nonlinearities. The feature explanations clarify the semantic meaning of the inputs and confirm that the surrogate assigns the strongest influence to class and confidence, as expected given the nature of the object detection task. 7 Conclusion This work presented a self-awareness framework that enhances the transparency and trustworthiness of You Only Look Once detections through an interpretable Kolmogorov-Arnold network. By modeling confidence as a structured sum of univariate spline functions, the network provides direct insight into how geometric and semantic features shape the reliability of each prediction. The visual-language captioning component assists the perception pipeline with lightweight multimodal descriptions. Overall, the framework provides an interpretable approach with: 1. Numerical insight into whether You Only Look Once’s confidence values are trustworthy. 2. Real-time and practical provision of confidence across the full feature domain, while simultaneously identifying unreliable detections in visually ambiguous scenes where the detector may fail. 3. Analysis of hidden units and monotonicity patterns for a coherent internal structure, with specialised neurons that consistently encode class information, confidence trends, and interpretable geometric influences. 4. A multimodal extension using visua-language descriptive scene-level context without altering the underlying modeling. Importantly, the findings show that the proposed framework enables interpretable detection with reliable confidence estimation, strengthening the transparency and robustness of modern automated systems. Acknowledgements The authors gratefully acknowledge the Microsoft Research (MSR) team for the dataset, while additional data were collected from the University of Bath campus in UK. References [1] K. M. Adnan, T. M. Ghazal, M. Saleem, M. S. Farooq, C. Y. Yeun, M. Ahmad, and S. Lee (2025) Deep learning driven interpretable and informed decision making model for brain tumour prediction using explainable ai. Scientific Reports 15 (1), p. 19223. Cited by: §1. [2] R. Agrawal, T. Gupta, S. Gupta, S. Chauhan, P. Patel, and S. Hamdare (2025) Fostering trust and interpretability: integrating explainable ai (xai) with machine learning for enhanced disease prediction and decision transparency. Diagnostic Pathology 20 (1), p. 105. Cited by: §1. [3] V. Akhil, G. Raghav, N. Arunachalam, and D. Srinivas (2020) Image data-based surface texture characterization and prediction using machine learning approaches for additive manufacturing. Journal of Computing and Information Science in Engineering 20 (2), p. 021010. Cited by: §1. [4] J. An, S. Dong, X. Wang, C. Li, and W. Zhao (2025) Research on uav aerial imagery detection algorithm for mining-induced surface cracks based on improved yolov10. Scientific Reports 15 (1), p. 30101. Cited by: §1. [5] T. Ansar and W. M. Ashraf (2025) Comparison of kolmogorov–arnold networks and multi-layer perceptron for modelling and optimisation analysis of energy systems. Energy and AI 20, p. 100473. Cited by: §1. [6] I. Barašin, B. Bertalanič, M. Mohorčič, and C. Fortuna (2025) Exploring kolmogorov–arnold networks for interpretable time series classification. International Journal of Intelligent Systems 2025 (1), p. 9553189. Cited by: §1. [7] P. M. Bhatt, R. K. Malhan, P. Rajendran, B. C. Shah, S. Thakar, Y. J. Yoon, and S. K. Gupta (2021) Image-based surface defect detection using deep learning: a review. Journal of Computing and Information Science in Engineering 21 (4), p. 040801. Cited by: §1. [8] M. B. Bowen, L. A. Smith, C. L. Carroll, M. Rahmanpour, T. Pan, and B. Morkos (2026) Navigating standards in engineering design through latent textual topology and llms. Journal of Computing and Information Science in Engineering, p. 1–14. Cited by: §1. [9] B. Chander, C. John, L. Warrier, and K. Gopalakrishnan (2025) Toward trustworthy artificial intelligence (tai) in the context of explainability and robustness. ACM Computing Surveys 57 (6), p. 1–49. Cited by: §1. [10] W. Chen, X. Ke, and S. Meng (2025) Small defect detection in printed circuit boards based on the multiscale edge strengthening and an improved yolov10. Scientific Reports 15 (1), p. 36445. Cited by: §1. [11] G. R. Chhimpa, S. Awasthi, N. Bhati, P. Yadav, and N. A. Wani (2025) A transfer learning-driven fine-tuning of yolov10 for improved brain tumor detection in mri images. Scientific Reports. Cited by: §1. [12] Y. Chudasama, H. Huang, D. Purohit, and M. Vidal (2025) Toward interpretable hybrid ai: integrating knowledge graphs and symbolic reasoning in medicine. IEEE Access 13, p. 39489–39509. Cited by: §1. [13] C. Cousineau, R. Dara, and A. Chowdhury (2025) Trustworthy ai: ai developers’ lens to implementation challenges and opportunities. Data and Information Management 9 (2), p. 100082. Cited by: §1. [14] G. P. Da Luza, G. M. Satoa, L. F. G. Gonzaleza, and J. F. Borin (2025) Smart parking with pixel-wise roi selection for vehicle detection using yolov8, yolov9, yolov10, and yolov11. Internet of Things, p. 101858. Cited by: §1. [15] D. Dai, L. Xu, Y. Li, Y. Zhang, and S. Xia (2025) Humanvlm: foundation for human-scene vision-language model. Information Fusion 123, p. 103271. Cited by: §1. [16] M. Deitke, C. Clark, S. Lee, R. Tripathi, Y. Yang, J. S. Park, M. Salehi, N. Muennighoff, K. Lo, L. Soldaini, et al. (2025) Molmo and pixmo: open weights and open data for state-of-the-art vision-language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 91–104. Cited by: §1. [17] A. Deng, T. Cao, Z. Chen, and B. Hooi (2025) Words or vision: do vision-language models have blind faith in text?. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 3867–3876. Cited by: §1. [18] R. Ding, C. Wang, X. Wu, G. Chen, J. Yang, H. Jia, S. Zhong, and R. Gu (2026) SAA-yolo: a scale-aware attention enhanced yolov10 for uav-based detection of medicinal plant phlomoides rotata in alpine grasslands. Computers and Electronics in Agriculture 245, p. 111570. Cited by: §1. [19] H. Dong, Y. Wang, and D. Miao (2025) Improved yolov10-based real-time helmet detection algorithm for complex scenarios. Journal of Real-Time Image Processing 22 (6), p. 197. Cited by: §1. [20] Q. Fan, Y. Li, M. Deveci, K. Zhong, and S. Kadry (2025) LUD-yolo: a novel lightweight object detection network for unmanned aerial vehicle. Information Sciences 686, p. 121366. Cited by: §1. [21] V. Gliner, I. Levy, K. Tsutsui, M. R. Acha, J. Schliamser, A. Schuster, and Y. Yaniv (2025) Clinically meaningful interpretability of an ai model for ecg classification. NPJ Digital Medicine 8 (1), p. 109. Cited by: §1. [22] M. Goisauf, M. Cano Abadia, K. Akyuz, M. Bobowicz, A. Buyx, I. Colussi, M. Fritzsche, K. Lekadir, P. Marttinen, M. T. Mayrhofer, et al. (2025) Trust, trustworthiness, and the future of medical ai: outcomes of an interdisciplinary expert workshop. Journal of Medical Internet Research 27, p. e71236. Cited by: §1. [23] H. Haoyan, T. Jinwu, W. Haibin, and L. Xinyun (2025) Ead-yolov10: lightweight steel surface defect detection algorithm research based on yolov10 improvement. IEEE Access. Cited by: §1. [24] M. Hasan, X. Zhao, W. Wu, J. Dai, X. Gu, and A. Noreen (2025) Long short-term memory and kolmogorov arnold network theorem for epileptic seizure prediction. Engineering Applications of Artificial Intelligence 154, p. 110757. Cited by: §1. [25] A. A. Howard, B. Jacob, and P. Stinis (2025) Multifidelity kolmogorov–arnold networks. Machine Learning: Science and Technology 6 (3), p. 035038. Cited by: §1. [26] Y. Hu, J. Yuan, C. Wen, X. Lu, Y. Liu, and X. Li (2025) Rsgpt: a remote sensing vision language model and benchmark. ISPRS Journal of Photogrammetry and Remote Sensing 224, p. 272–286. Cited by: §1. [27] R. Jannel and J. Tallant (2026) Trustability and trustworthiness: conceptual foundations and the case of ai. AI and Ethics 6 (1), p. 13. Cited by: §1. [28] J. Jiménez-Luna, F. Grisoni, and G. Schneider (2020) Drug discovery with explainable artificial intelligence. Nature Machine Intelligence 2 (10), p. 573–584. Cited by: §1. [29] D. W. Joyce, A. Kormilitzin, K. A. Smith, and A. Cipriani (2023) Explainable artificial intelligence for mental health through transparency and interpretability for understandability. npj digital medicine 6 (1), p. 6. Cited by: §1. [30] H. Jung, D. Paek, and S. Kong (2025) Open-source autonomous driving software platforms: comparison of autoware and apollo. arXiv preprint arXiv:2501.18942. Cited by: §1. [31] S. KC et al. (2022) Enhanced pothole detection system using yolox algorithm. Autonomous Intelligent Systems 2 (1), p. 22. Cited by: §1. [32] B. Khalili and A. W. Smyth (2024) SOD-yolov8—enhancing yolov8 for small object detection in aerial imagery and traffic scenes. Sensors 24 (19), p. 6209. Cited by: §1. [33] E. Kiyani, K. Shukla, J. F. Urbán, J. Darbon, and G. E. Karniadakis (2025) Optimizing the optimizer for physics-informed neural networks and kolmogorov-arnold networks. Computer Methods in Applied Mechanics and Engineering 446, p. 118308. Cited by: §1. [34] S. M. Lauritsen, M. Kristensen, M. V. Olsen, M. S. Larsen, K. M. Lauritsen, M. J. Jørgensen, J. Lange, and B. Thiesson (2020) Explainable artificial intelligence model to predict acute critical illness from electronic health records. Nature communications 11 (1), p. 3852. Cited by: §1. [35] D. Li, Z. Yang, W. Nai, Y. Xing, and Z. Chen (2025) A road lane detection approach based on reformer model. Egyptian Informatics Journal 29, p. 100625. Cited by: §1. [36] L. Li, Y. Zhang, G. Wang, and K. Xia (2025) Kolmogorov–arnold graph neural networks for molecular property prediction. Nature Machine Intelligence 7 (8), p. 1346–1354. Cited by: §1. [37] T. Li, J. Ruan, and K. Zhang (2025) The investigation of reinforcement learning-based end-to-end decision-making algorithms for autonomous driving on the road with consecutive sharp turns. Green Energy and Intelligent Transportation 4 (3), p. 100288. Cited by: §1. [38] Z. Li, D. Muhtar, F. Gu, Y. He, X. Zhang, P. Xiao, G. He, and X. Zhu (2025) Lhrs-bot-nova: improved multimodal large language model for remote sensing vision-language interpretation. ISPRS Journal of Photogrammetry and Remote Sensing 227, p. 539–550. Cited by: §1. [39] J. Liu, Y. Wei, R. Sun, and Y. Yue (2025) LMG-yolov10: an efficient defect detection model for aero-engine components based on improved yolov10. Aerospace Science and Technology, p. 110788. Cited by: §1. [40] W. Liu, X. Qiao, C. Zhao, T. Deng, and F. Yan (2025) VP-yolo: a human visual perception-inspired robust vehicle-pedestrian detection model for complex traffic scenarios. Expert Systems with Applications 274, p. 126837. Cited by: §1. [41] Z. Liu, Y. Wang, S. Vaidya, F. Ruehle, J. Halverson, M. Soljačić, T. Y. Hou, and M. Tegmark (2024) Kan: kolmogorov-arnold networks. arXiv preprint arXiv:2404.19756. Cited by: §1. [42] S. M. Lundberg, G. Erion, H. Chen, A. DeGrave, J. M. Prutkin, B. Nair, R. Katz, J. Himmelfarb, N. Bansal, and S. Lee (2020) From local explanations to global understanding with explainable ai for trees. Nature machine intelligence 2 (1), p. 56–67. Cited by: §1. [43] A. A. Mahapadi, V. Shirsath, and A. Pundge (2025) Real-time diabetic retinopathy detection using yolo-v10 with nature-inspired optimization. Biomedical Materials & Devices, p. 1–23. Cited by: §1. [44] Z. Meng, X. Du, R. Sapkota, Z. Ma, and H. Cheng (2025) YOLOv10-pose and yolov9-pose: real-time strawberry stalk pose detection models. Computers in Industry 165, p. 104231. Cited by: §1. [45] P. A. Moreno-SÃ, J. Del Ser, M. Van Gils, J. Hernesniemi, et al. (2025) A design framework for operationalizing trustworthy artificial intelligence in healthcare: requirements, tradeoffs and challenges for its clinical adoption. Information Fusion, p. 103812. Cited by: §1. [46] X. Nie, L. Zhu, Z. He, A. Cheng, S. Zhong, and E. Li (2025) Investigating 3d object detection using stereo camera and lidar fusion with bird’s-eye view representation. Neurocomputing 620, p. 129144. Cited by: §1. [47] G. Novakovsky, N. Dexter, M. W. Libbrecht, W. W. Wasserman, and S. Mostafavi (2023) Obtaining genetics insights from deep learning via explainable artificial intelligence. Nature Reviews Genetics 24 (2), p. 125–137. Cited by: §1. [48] M. Oussouaddi, O. Bouazizi, Z. e. A. A. Ismaili, Y. Attaoui, M. Chentouf, et al. (2025) DSR-yolo: a lightweight and efficient yolov8 model for enhanced pedestrian detection. Cognitive Robotics 5, p. 152–165. Cited by: §1. [49] S. Panahi, M. Moradi, E. M. Bollt, and Y. Lai (2025) Data-driven model discovery with kolmogorov-arnold networks. Physical Review Research 7 (2), p. 023037. Cited by: §1. [50] J. Parra-Ullauri, X. Zhou, S. Moazzeni, R. Hussain, X. Vasilakos, Y. Wu, R. Baby, M. H. Mahmud, G. Incorvaia, D. Hond, et al. (2025) Lifecycle management of trustworthy ai models in 6g networks: the reason approach. IEEE Wireless Communications 32 (2), p. 42–51. Cited by: §1. [51] B. Qin, H. Hu, and S. Du (2025) ACM-yolov10: research on classroom learning behavior recognition algorithm based on improved yolov10. IEEE Access. Cited by: §1. [52] L. T. Ramos, E. Casas, C. Romero, F. Rivas-Echeverría, and E. Bendek (2025) A study of yolo architectures for wildfire and smoke detection in ground and aerial imagery. Results in Engineering 26, p. 104869. Cited by: §1. [53] M. R. A. Rashid, M. A. E. Korim, M. Hasan, M. S. Ali, M. M. Islam, T. Jabid, R. U. Islam, and M. Islam (2025) An ensemble learning framework with explainable ai for interpretable leaf disease detection. Array 26, p. 100386. Cited by: §1. [54] K. Reinhardt (2023) Trust and trustworthiness in ai ethics. AI and Ethics 3 (3), p. 735–744. Cited by: §1. [55] P. Roy, M. Hasan, M. R. Islam, and M. P. Uddin (2025) Interpretable artificial intelligence (ai) for cervical cancer risk analysis leveraging stacking ensemble and expert knowledge. Digital health 11, p. 20552076251327945. Cited by: §1. [56] M. Salvi, S. Seoni, A. Campagner, A. Gertych, U. R. Acharya, F. Molinari, and F. Cabitza (2025) Explainability and uncertainty: two sides of the same coin for enhancing the interpretability of deep learning models in healthcare. International Journal of Medical Informatics 197, p. 105846. Cited by: §1. [57] M. J. Saravani, R. Noori, C. Jun, D. Kim, S. M. Bateni, P. Kianmehr, and R. I. Woolway (2025) Predicting chlorophyll-a concentrations in the world’s largest lakes using kolmogorov-arnold networks. Environmental Science & Technology 59 (3), p. 1801–1810. Cited by: §1. [58] Y. Shu, Z. Liu, P. Zhang, M. Qin, J. Zhou, Z. Liang, T. Huang, and B. Zhao (2025) Video-xl: extra-long vision language model for hour-scale video understanding. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 26160–26169. Cited by: §1. [59] H. Shuai and F. Li (2025) Physics-informed kolmogorov-arnold networks for power system dynamics. IEEE Open Access Journal of Power and Energy 12, p. 46–58. Cited by: §1. [60] S. Singh, N. A. Wani, R. Kumar, and J. Bedi (2025) DiaXplain: a transparent and interpretable artificial intelligence approach for type-2 diabetes diagnosis through deep learning. Computers and Electrical Engineering 126, p. 110470. Cited by: §1. [61] B. Song, R. Zhou, and F. Ahmed (2024) Multi-modal machine learning in engineering design: a review and future directions. Journal of Computing and Information Science in Engineering 24 (1), p. 010801. Cited by: §1. [62] S. Song, X. Zhou, S. Yuan, P. Cheng, and X. Liu (2025) Interpretable artificial intelligence models for predicting lightning prone to inducing forest fires. Journal of Atmospheric and Solar-Terrestrial Physics 267, p. 106408. Cited by: §1. [63] P. N. Srinivasu, G. L. A. Kumari, S. C. Narahari, S. Ahmed, and A. Alhumam (2025) Exploring the impact of hyperparameter and data augmentation in yolo v10 for accurate bone fracture detection from x-ray images. Scientific Reports 15 (1), p. 9828. Cited by: §1. [64] E. Stamboliev and T. Christiaens (2025) How empty is trustworthy ai? a discourse analysis of the ethics guidelines of trustworthy ai. Critical Policy Studies 19 (1), p. 39–56. Cited by: §1. [65] H. Sun, G. Yao, S. Zhu, L. Zhang, H. Xu, and J. Kong (2025) SOD-yolov10: small object detection in remote sensing images based on yolov10. IEEE Geoscience and Remote Sensing Letters 22, p. 1–5. Cited by: §1. [66] H. Ta, D. Thai, A. B. S. Rahman, G. Sidorov, and A. Gelbukh (2026) Fc-kan: function combinations in kolmogorov-arnold networks. Information Sciences, p. 123103. Cited by: §1. [67] B. Tang, J. Zhou, C. Zhao, Y. Pan, Y. Lu, C. Liu, K. Ma, X. Sun, R. Zhang, and X. Gu (2025) Using uav-based multispectral images and cgs-yolo algorithm to distinguish maize seeding from weed. Artificial Intelligence in Agriculture 15 (2), p. 162–181. Cited by: §1. [68] J. D. Toscano, T. Käufer, Z. Wang, M. Maxey, C. Cierpka, and G. E. Karniadakis (2025) AIVT: inference of turbulent thermal convection from measured 3d velocity data by physics-informed kolmogorov-arnold networks. Science advances 11 (19), p. eads5236. Cited by: §1. [69] J. D. Toscano, L. Wang, and G. E. Karniadakis (2025) KKANs: kurkova-kolmogorov-arnold networks and their learning dynamics. Neural Networks 191, p. 107831. Cited by: §1. [70] M. Usama, H. Anwar, and S. Anwar (2025) Vehicle and license plate recognition with novel dataset for toll collection. Pattern Analysis and Applications 28 (2), p. 57. Cited by: §1. [71] P. K. A. Vasu, F. Faghri, C. Li, C. Koc, N. True, A. Antony, G. Santhanam, J. Gabriel, P. Grasch, O. Tuzel, et al. (2025) Fastvlm: efficient vision encoding for vision language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, p. 19769–19780. Cited by: §1. [72] A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han, and G. Ding (2024) Yolov10: real-time end-to-end object detection. Advances in neural information processing systems 37, p. 107984–108011. Cited by: §1. [73] J. Wang, J. Su, Z. Wen, and Y. Sun (2025) Enhanced yolov10 for small object detection with context-aware and adaptive modules. International Journal of Multimedia Information Retrieval 14 (3), p. 22. Cited by: §1. [74] Z. Wang, A. Zainal, M. M. Siraj, F. A. Ghaleb, X. Hao, and S. Han (2025) An intrusion detection model based on convolutional kolmogorov-arnold networks. Scientific Reports 15 (1), p. 1917. Cited by: §1. [75] F. Wei and W. Wang (2025) Scca-yolo: a spatial and channel collaborative attention enhanced yolo network for highway autonomous driving perception system. Scientific Reports 15 (1), p. 6459. Cited by: §1. [76] J. Wei, A. As’arry, K. A. M. Rezali, M. Z. M. Yusoff, H. Ma, and K. Zhang (2025) A review of yolo algorithm and its applications in autonomous driving object detection. IEEE Access. Cited by: §1. [77] C. D. Wirz, J. L. Demuth, A. Bostrom, M. G. Cains, I. Ebert-Uphoff, D. J. Gagne I, A. Schumacher, A. McGovern, and D. Madlambayan (2025) (Re) conceptualizing trustworthy ai: a foundation for change. Artificial Intelligence 342, p. 104309. Cited by: §1. [78] H. Wu, Z. Du, D. Zhong, Y. Wang, and C. Tao (2025) FSVLM: a vision-language model for remote sensing farmland segmentation. IEEE Transactions on Geoscience and Remote Sensing 63, p. 1–13. Cited by: §1. [79] Y. Wu, Z. Zang, X. Zou, W. Luo, N. Bai, Y. Xiang, W. Li, and W. Dong (2025) Graph attention and kolmogorov–arnold network based smart grids intrusion detection. Scientific Reports 15 (1), p. 8648. Cited by: §1. [80] Y. Xie, D. Du, and M. Bi (2025) YOLO-ace: a vehicle and pedestrian detection algorithm for autonomous driving scenarios based on knowledge distillation of yolov10. IEEE Internet of Things Journal. Cited by: §1. [81] Y. Xin, G. C. Ates, K. Gong, and W. Shao (2025) Med3dvlm: an efficient vision-language model for 3d medical image analysis. IEEE Journal of Biomedical and Health Informatics. Cited by: §1. [82] G. Xu, P. Jin, Z. Wu, H. Li, Y. Song, L. Sun, and L. Yuan (2025) Llava-cot: let vision language models reason step-by-step. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 2087–2098. Cited by: §1. [83] L. Xue, M. Shu, A. Awadalla, J. Wang, A. Yan, S. Purushwalkam, H. Zhou, V. Prabhu, Y. Dai, M. S. Ryoo, et al. (2025) Blip-3: a family of open large multimodal models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 6124–6135. Cited by: §1. [84] S. Yang, Y. Chen, Z. Tian, C. Wang, J. Li, B. Yu, and J. Jia (2025) Visionzip: longer is better but not necessary in vision language models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 19792–19802. Cited by: §1. [85] Z. Yao, M. Wu, J. Qian, and D. Reynaerts (2025) Intelligent discharge state detection in micro-edm process with cost-effective radio frequency (rf) radiation: integrating machine learning and interpretable ai. Expert Systems with Applications 291, p. 128607. Cited by: §1. [86] Y. Zhan, Z. Xiong, and Y. Yuan (2025) Skyeyegpt: unifying remote sensing vision-language tasks via instruction tuning with large language model. ISPRS Journal of Photogrammetry and Remote Sensing 221, p. 64–77. Cited by: §1. [87] J. Zhang, Z. Wang, Y. Yi, L. Kuang, and J. Zhang (2025) PSFE-yolo: a traffic sign detection algorithm with pixel-wise spatial feature enhancement. Pattern Analysis and Applications 28 (1), p. 24. Cited by: §1. [88] Y. Zhang, X. Chen, S. Sun, H. You, Y. Wang, J. Lin, and J. Wang (2025) Vehicle detection in drone aerial views based on lightweight osd-yolov10. Scientific Reports 15 (1), p. 25155. Cited by: §1. [89] J. Zhong, D. Kong, Y. Wei, and B. Pan (2025) YOLOv8 and point cloud fusion for enhanced road pothole detection and quantification. Scientific Reports 15 (1), p. 11260. Cited by: §1. [90] S. Zhou, A. Vilesov, X. He, Z. Wan, S. Zhang, A. Nagachandra, D. Chang, D. Chen, X. E. Wang, and A. Kadambi (2025) Vlm4d: towards spatiotemporal awareness in vision language models. In Proceedings of the IEEE/CVF international conference on computer vision, p. 8600–8612. Cited by: §1.