Paper deep dive
Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry
Reza Ahmari, Ahmad Mohammadi, Vahid Hemmati, Nicholas Edmond, Hossein Z. Saghazadeh, Olusola Odeyomi, Parham Kebria, Abdollah Homaifar
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/9/2026, 2:56:55 AM
Summary
This paper presents an interpretable fuzzy inference framework for estimating continuous yaw commands for UAVs tracking UGVs using low-dimensional features from YOLO bounding boxes. The framework compares a Mamdani fuzzy system baseline with a first-order Takagi-Sugeno model, utilizing features like target centroid location, area, and aspect ratio. Evaluated on 6,169 samples from a VICON motion-capture environment, the Takagi-Sugeno model achieved high accuracy (MAE 0.140°) and demonstrated that the approach is transparent, data-efficient, and suitable for real-time deployment on resource-constrained platforms.
Entities (8)
Relation Signals (5)
Fuzzy Inference Framework ā estimates ā Yaw Estimation
confidence 95% Ā· presents an interpretable fuzzy-inference framework that generates continuous yaw commands
YOLO ā providesfeaturesfor ā Fuzzy Inference Framework
confidence 95% Ā· low-dimensional features extracted from YOLO boxes: target centroid location, area, and aspect ratio
Takagi-Sugeno Model ā achieveshigheraccuracythan ā Mamdani Fuzzy System
confidence 90% Ā· The Takagi-Sugeno model achieves a test-set mean absolute error of 0.140... These results show that the framework is transparent... suitable for real-time vision-based UAV guidance
Bounding-Box Geometry ā inputto ā Takagi-Sugeno Model
confidence 90% Ā· first-order Takagi--Sugeno model... derived from training-set quantiles... antecedent space is formed from the same three YOLO-derived bounding-box descriptors
Vicon ā providesgroundtruthfor ā Yaw Estimation
confidence 90% Ā· Ground-truth yaw labels are generated from a VICON motion-capture system
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) supports cooperative aerial--ground robotics, but reliable continuous yaw estimation from onboard vision remains challenging because of sensing uncertainty, limited computation, and the need for interpretable control. Existing deep-learning and geometric-reconstruction approaches often require large datasets, external localization, or complex modeling assumptions, reducing transparency and deployment suitability on resource-constrained platforms. We present an interpretable fuzzy-inference framework that generates continuous yaw commands from low-dimensional features extracted from YOLO boxes: target centroid location, area, and aspect ratio. No explicit geometric modeling is required. A Mamdani fuzzy system serves as an interpretable baseline using a shoulder--triangle--shoulder input partition. It is followed by a first-order Takagi--Sugeno model with three antecedent membership terms per input, whose parameters are derived from training-set quantiles, yielding a compact 27-rule structure. Evaluation uses 6{,}169 labeled samples from a VICON motion-capture environment. Across five randomized train--test splits, the Takagi--Sugeno model achieves a test-set mean absolute error of $0.140^\circ \pm 0.003^\circ$, a root mean squared error of $0.200^\circ \pm 0.008^\circ$, and a maximum absolute error of $1.254^\circ \pm 0.121^\circ$. Within-threshold accuracies are $99.676% \pm 0.270%$ for $\pm1^\circ$ and $100.000% \pm 0.000%$ for both $\pm3^\circ$ and $\pm5^\circ$. Directional consistency between image-plane horizontal displacement and predicted yaw sign reaches $90.254% \pm 0.612%$. These results show that the framework is transparent, data-efficient, computationally lightweight, and suitable for real-time vision-based UAV guidance toward mobile ground targets.
Tags
Links
- Source: https://arxiv.org/abs/2608.04121v1
- Canonical: https://arxiv.org/abs/2608.04121v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
70,480 characters extracted from source content.
Expand or collapse full text
[2] . Abdollah [1] of Computer Science, Carolina A&T State University, 1601 E Market St, , 27411, , 2] of Electrical and Computer Engineering, Carolina A&T State University, 1601 E Market St, , 27411, , Interpretable Fuzzy Inference for UAV Target Tracking Using Bounding-Box Geometry rahmari@aggies.ncat.edu amohammadi@aggies.ncat.edu . Vahid vhemmati@ncat.edu nedmond@aggies.ncat.edu . Saghazadeh hzamanisaghazadeh@aggies.ncat.edu . Olusola otodeyomi@ncat.edu . Parham pmkebria@ncat.edu homaifar@ncat.edu * [ Abstract Vision-based guidance of unmanned aerial vehicles (UAVs) toward unmanned ground vehicles (UGVs) is a key capability for cooperative aerialāground robotic systems; however, estimating a reliable continuous yaw command from onboard visual perception alone remains challenging due to sensing uncertainty, limited onboard computational resources, and the need for interpretable control strategies. Many existing approaches rely on deep learning or geometric reconstruction methods that require large datasets, external localization, or complex modeling assumptions, reducing transparency and suitability for real-time deployment on resource-constrained platforms. This paper presents an interpretable fuzzy-inference-based framework for generating continuous yaw commands using only low-dimensional vision features extracted from YOLO bounding boxes, including target centroid location, area, and aspect ratio, without explicit geometric modeling. A Mamdani fuzzy system is developed as an interpretable baseline using a shoulderātriangleāshoulder input partition, followed by a first-order TakagiāSugeno model with three antecedent membership terms per input, whose parameters are derived from training-set quantiles, yielding a compact 27-rule structure. The approach is evaluated on 6,169 labeled samples collected in a VICON motion-capture environment. Across five randomized trainātest splits, the TakagiāSugeno model achieves a test-set mean absolute error of 0.140ā±0.003ā0.140 ± 0.003 , a root mean squared error of 0.200ā±0.008ā0.200 ± 0.008 , and a maximum absolute error of 1.254ā±0.121ā1.254 ± 0.121 . The within-threshold accuracies are 99.676%±0.270%99.676\%± 0.270\% for ±1ā± 1 and 100.000%±0.000%100.000\%± 0.000\% for both ±3ā± 3 and ±5ā± 5 . Directional consistency between image-plane horizontal displacement and predicted yaw sign reaches 90.254%±0.612%90.254\%± 0.612\%, demonstrating that the proposed framework provides a transparent, data-efficient, and computationally lightweight solution for real-time vision-based UAV guidance toward mobile ground targets. keywords: Guidance Navigation Control (GNC), Fuzzy inference, UAV-UGV integration, Vision-Based navigation, Yaw estimation 1 Introduction Vision-based perception plays a central role in autonomous aerial and ground robotic systems, particularly in scenarios where precise relative positioning and alignment are required for navigation, tracking, or cooperative tasks [hutchinson2002tutorial, chaumette2006visual]. Unmanned aerial vehicles (UAVs) are increasingly deployed across a wide range of applications including surveillance, infrastructure inspection, logistics, and cooperative aerialāground robotics, highlighting the growing importance of reliable onboard perception and navigation capabilities [mohsan2023unmanned]. In such settings, compact geometric descriptors derived from object detections, such as bounding-box location, scale, and shape, provide an attractive alternative to dense image representations due to their low dimensionality and robustness to appearance variations. A recurring challenge in vision-to-control pipelines is the reliable mapping from noisy, perspective-dependent visual features to continuous control commands. Recent data-driven approaches, including deep neural networks and attention-based models, have demonstrated strong representational power for vision-based perception and regression tasks [bojarski2016end, muhammad2020deep, liu2024explainable, ming2024not]. While such methods can achieve high accuracy, they typically require large annotated datasets, introduce a high number of learned parameters, and rely on complex feature interactions that are difficult to interpret and validate within safety-critical control loops. In addition, their performance can be sensitive to distributional shifts in visual appearance, scale, or viewpoint, which are common in aerial robotics scenarios. In contrast, rule-based fuzzy inference systems explicitly encode the mapping between visual descriptors and control actions through transparent linguistic rules, enabling predictable behavior, easier validation, and graceful degradation under uncertainty [mamdani1975experiment, passino1998fuzzy]. These properties make fuzzy logic particularly well suited for low-dimensional, structured inputs such as bounding-box geometry, where interpretability, robustness, and controlled generalization are often more critical than maximizing perceptual expressiveness. Classical Mamdani fuzzy systems are widely used in control applications due to their transparency and intuitive linguistic structure. However, when applied to continuous regression tasks, Mamdani inference is known to exhibit limited numerical precision and output saturation effects, particularly as the dimensionality of the input space increases [passino1998fuzzy]. TakagiāSugeno fuzzy systems address these limitations by replacing fuzzy consequents with rule-local linear models, yielding improved approximation capability while preserving the interpretability of fuzzy antecedents [takagi1985fuzzy, sugeno1985industrial]. This paper presents a systematic and reproducible fuzzy-logic framework for estimating a continuous yaw correction command from YOLO bounding-box descriptors [redmon2016you]. The objective is not to introduce a new fuzzy-inference family or a new visual descriptor, but to formulate and validate a leakage-controlled, low-dimensional vision-to-yaw mapping in which interpretable fuzzy reasoning is applied directly to detector-produced visual perception outputs. Two fuzzy inference formulations are evaluated under a shared three-input antecedent rule structure. First, a Mamdani fuzzy controller is constructed as an interpretable baseline using robust percentile-based normalization, explicit shoulderātriangleāshoulder membership functions, and a deterministic same-side rule design. Second, a first-order TakagiāSugeno model is constructed using train-derived quantile membership parameters and rule-local linear consequents identified by least-squares regression. In both cases, the antecedent space is formed from the same three YOLO-derived bounding-box descriptors: horizontal center, normalized area, and aspect ratio. The technical contribution lies in the controlled formulation of a compact fuzzy yaw-estimation pipeline that bridges detector-based visual perception and fuzzy inference. The framework combines: (i) training-only feature and membership-function parameterization to avoid test-data leakage, (i) coverage-safe shoulderātriangleāshoulder antecedent partitions to reduce boundary activation ambiguity, (i) a shared interpretable antecedent rule structure for Mamdani and TakagiāSugeno inference, and (iv) multi-metric evaluation against VICON-derived yaw labels and standard low-dimensional regression baselines. This design clarifies the trade-off between transparent fuzzy directional reasoning and high-precision continuous yaw regression. The contributions supported by this paper are as follows: 1. A reproducible low-dimensional vision-to-yaw estimation framework that maps YOLO bounding-box geometry to a continuous yaw correction variable using only horizontal center, normalized area, and aspect ratio. 2. A leakage-controlled membership-function design in which feature normalization and fuzzy antecedent parameters are derived from the training split only and applied unchanged during test-time inference. 3. A coverage-safe shoulderātriangleāshoulder fuzzy partition for the antecedent variables, reducing boundary activation ambiguity while preserving a compact 27-rule fuzzy structure. 4. A controlled comparison between Mamdani and first-order TakagiāSugeno inference under a shared antecedent rule structure, clarifying the trade-off between interpretable directional reasoning and high-precision continuous yaw estimation. 5. A comparative evaluation against standard low-dimensional regression baselines and computational-volume measurements, enabling assessment of accuracy, worst-case error, side-consistency, interpretability, and inference cost. 2 Related Work Vision-based heading and yaw estimation has been studied extensively in the contexts of autonomous navigation, visual servoing, and cooperative aerialāground systems. Existing approaches can be broadly categorized into geometry-based methods, data-driven regression models, and fuzzy or rule-based controllers [hutchinson2002tutorial, chaumette2006visual, wu2022survey]. This section reviews prior work along these axes, with emphasis on methods that leverage bounding-box or image-plane features and those that balance interpretability with predictive accuracy. 2.1 Vision-Based Heading and Yaw Estimation Early work on vision-based yaw and heading estimation largely relied on explicit geometric reasoning and visual servoing principles. Classical image-based visual servoing (IBVS) formulations estimate orientation errors directly from image features such as centroids, edges, or point correspondences, and generate control commands through analytical Jacobians [hutchinson2002tutorial, chaumette2006visual]. While these methods provide theoretical guarantees under idealized conditions, their performance degrades in the presence of detection noise, partial occlusion, and imperfect feature association, which are common in real-world robotic deployments. More recent studies have leveraged object detection outputs, particularly bounding boxes, as compact representations of target pose in the image plane. Bounding-box center displacement provides a direct image-plane cue for lateral misalignment, while box area and scale variation act as depth-related cues in monocular settings through perspective projection and region-feature geometry [hutchinson2002tutorial, chaumette2006visual, chaumette2004image]. These approaches reduce reliance on precise feature tracking and enable integration with modern deep detectors, but they require robust mappings from low-dimensional visual descriptors to continuous control signals [wu2022survey]. 2.2 Data-Driven Regression Approaches With the increasing availability of labeled data, fully data-driven models have been proposed to regress yaw or steering commands directly from visual inputs or intermediate features. Convolutional neural networks have been used to predict steering angles or yaw corrections end-to-end from raw images, as demonstrated in learning-based driving and aerial navigation systems [bojarski2016end]. Although such models can achieve high numerical accuracy, they typically require large training datasets, exhibit limited interpretability, and provide little insight into failure modes when operating outside the training distribution [muhammad2020deep]. Hybrid approaches have also been explored, where deep networks extract intermediate features that are subsequently mapped to control variables using shallow regressors or linear models. While these methods reduce model complexity relative to end-to-end learning, they still rely on opaque feature representations and often lack explicit mechanisms for incorporating domain knowledge or enforcing monotonic relationships between visual cues and control outputs. 2.3 Fuzzy Logic Controllers in Robotics Fuzzy logic control has a long history in mobile robotics and autonomous systems, particularly in scenarios involving uncertainty, imprecise sensing, or heuristic reasoning. Mamdani-type fuzzy controllers have been widely applied to navigation, obstacle avoidance, and alignment tasks due to their linguistic rule structure and intuitive interpretability [mamdani1975experiment, passino1998fuzzy]. In visual navigation contexts, fuzzy rules such as āif target is left, turn leftā provide transparent decision logic that can be inspected and modified by human designers. Prior work has shown that Mamdani-type fuzzy inference systems, while highly interpretable and well suited for linguistic control design, can exhibit limitations when applied to continuous regression and estimation problems. Because the output of a Mamdani system is synthesized through fixed output membership functions and defuzzification, the resulting mapping may suffer from limited numerical resolution and sensitivity to the chosen output partition, particularly when modeling smooth inputāoutput relationships. Such characteristics can lead to coarse control signals and reduced approximation accuracy in strongly continuous domains [mendel2002type, riza2015frbs]. In contrast, fuzzy rule-based systems with functional consequents, including TakagiāSugeno formulations, are explicitly designed to improve regression performance by representing local inputāoutput relationships more precisely. TakagiāSugeno fuzzy systems have been proposed as a remedy to these limitations by replacing fuzzy-set consequents with parametric functions, typically linear in the inputs. In robotics applications, first-order Sugeno models have demonstrated improved approximation capability while retaining the structured partitioning of the input space provided by fuzzy membership functions [takagi1985fuzzy, sugeno1985industrial]. More broadly, robust yaw regulation remains an important control problem across autonomous robotic platforms. For example, Karade et al. [karade2025robustyaw] investigated robust yaw angle control for autonomous underwater vehicles using a dynamic surface-based optimized second-order sliding-mode framework under disturbances and model uncertainty. Although that work focuses on model-based control rather than vision-based fuzzy inference, it further highlights the importance of accurate and robust yaw-related decision making in autonomous robotic systems. The consequent parameters can be identified using least-squares techniques, enabling data-driven optimization without sacrificing the interpretability of the rule antecedents. Hybrid fuzzyāneural approaches have also been explored to enhance robustness and adaptability in uncertain control environments [kebria2019type, kebria2019adaptive]. 2.4 Bounding-Box Features and Normalization Several studies have emphasized the importance of feature normalization when using bounding-box descriptors for control or estimation. Bounding-box area and aspect ratio are particularly sensitive to distance changes, partial visibility, and detector noise. Percentile-based or quantile-based normalization schemes have been employed in related contexts to mitigate the influence of outliers and to stabilize downstream learning algorithms [hastie2009elements, kebria2019fuzzy, loh2025theoretical]. Despite their practical importance, preprocessing and membership-function parameterization strategies are often under-specified in fuzzy perception pipelines, leading to ambiguity regarding trainingātest leakage and reproducibility. In the present framework, all feature-scaling and membership-function parameters are derived from the training split only and applied unchanged during test-time inference. This ensures that the reported Mamdani and TakagiāSugeno results are not affected by test-set information leakage. 2.5 Positioning of the Present Work The primary contribution of this work is therefore not a new fuzzy-inference formulation itself, but a framework-level integration in which interpretable fuzzy reasoning is applied directly to detector-produced visual perception outputs rather than to direct physical state measurements. Specifically, YOLO-derived bounding-box descriptors are used as compact geometric inputs for monocular UAVāUGV relative yaw estimation, establishing a lightweight perception-level fuzzy inference framework that bridges object-detection-based visual perception and fuzzy reasoning. Unlike prior fuzzy-control studies that emphasize heuristic tuning or simulation-only validation, the proposed approach is evaluated on a real-world dataset with high-fidelity VICON-based ground-truth yaw measurements. By evaluating Mamdani inference, TakagiāSugeno inference, standard regression baselines, and computational cost under the same bounding-box feature setting, the study clarifies the trade-offs among interpretability, numerical accuracy, worst-case error, side-consistency, and inference efficiency in vision-based UAV yaw estimation. 3 Methodology The objective of this section is to develop a reproducible pipeline for predicting a continuous yaw correction command, Īø (degrees), from vision-based features. Each data sample corresponds to a single image frame containing a YOLO detected UGV bounding box, from which geometric descriptors are extracted. The corresponding ground-truth yaw reference is obtained from a VICON motion-capture system. Two fuzzy inference systems are evaluated using an identical antecedent structure with three input variables and three linguistic terms per input (27 rules total). First, a Mamdani baseline designed for interpretability, and secondly, a first-order TakagiāSugeno model, in which the consequent parameters are identified by least squares on the training split. 3.1 Notation and Problem Definition Each sample corresponds to one image frame and one detected target bounding box. Let the image resolution be WĆHWĆ H pixels and the detected bounding box be ā¬=(x1,y1,x2,y2), =(x_1,y_1,x_2,y_2), (1) 0ā¤x1<x2ā¤W, 0ā¤y1<y2ā¤H. 0⤠x_1<x_2⤠W,0⤠y_1<y_2⤠H. where (x1,y1)(x_1,y_1), and (x2,y2)(x_2,y_2) denote the pixel coordinates of the top-left and bottom-right corners of the bounding box, respectively. The goal is to estimate a yaw Īø^āā Īø from a three-dimensional feature vector =[cx,a,r]āā3x=[c_x,\;a,\;r] T ^3 (2) via a mapping Īø^=fā() Īø=f(x). Where, (cx,a,r)(c_x,\;a,\;r), denote the horizontal center, normalized area, and aspect ratio features extracted from the bounding box. Depending on the model formulation, these features may be used directly or after normalization. These parameters are defined in the following section. In this paper we evaluate the perception-to-yaw mapping offline to isolate inference accuracy and generalization. In deployment, Īø Īø is intended to serve as a yaw correction reference that can be fed to a low-level yaw-rate or attitude controller within the UAVās onboard control loop. 3.2 Ground-Truth Yaw from VICON Ground-truth yaw labels are generated from a VICON motion-capture system providing millimeter-level positional accuracy [merriaux2017study]. Let u=[pu,x,pu,y,pu,z]p_u=[p_u,x,p_u,y,p_u,z] T and g=[pg,x,pg,y,pg,z]p_g=[p_g,x,p_g,y,p_g,z] T denote the UAV and UGV positions in the VICON frame. The horizontal relative position of UGV with respect to the position of UAV is Īāxāy=[pg,xāpu,xpg,yāpu,y]. _xy= bmatrixp_g,x-p_u,x\\ p_g,y-p_u,y bmatrix. (3) The ground-truth yaw label Īø is defined as the signed line-of-sight bearing from the UAV to the UGV: Īø=atan2ā”(pg,yāpu,y,pg,xāpu,x),Īø=atan2\! (p_g,y-p_u,y,\;p_g,x-p_u,x ), (4) For reporting and evaluation, Īø is expressed in degrees. Negative values correspond to leftward bearings and positive values correspond to rightward bearings under the adopted coordinate convention. 3.3 Feature Definitions from Bounding Boxes The dataset includes the raw bounding-box coordinates (x1,y1,x2,y2)(x_1,y_1,x_2,y_2) and the derived features (cx,cy,a,r)(c_x,c_y,a,r). We use cxc_x, a, and r as inputs. For completeness, the feature definitions are listed below. Normalized Horizontal Center The bounding-box center is xc=x1+x22,yc=y1+y22,x_c= x_1+x_22, y_c= y_1+y_22, (5) and the normalized horizontal center is cx=xcWā[0,1].c_x= x_cWā[0,1]. (6) Although cy=yc/Hc_y=y_c/H is available, it is not used in the three-input models reported here. Normalized Area Let w=x2āx1w=x_2-x_1 and h=y2āy1h=y_2-y_1 denote the bounding-box width and height (pixels). The normalized area is a=wāhWāHā[0,1].a= whWHā[0,1]. (7) Aspect Ratio In our experiments, bounding-box aspect ratio r is used as the third input to both fuzzy models and matches the dataset column aspect_ratio. It is defined as r=hw,r= hw, (8) With this definition, larger r corresponds to a taller or narrower box, while smaller r corresponds to a wider box. This convention yields a positive feature that captures perspective-induced shape distortion and remains numerically stable under aerial viewpoints. Let =(n,Īøn)n=1ND=\(x_n, _n)\_n=1^N with N=6169N=6169. Indices are shuffled and split into training and test sets. To account for split variability, all experiments are repeated over five randomized splits with different fixed seeds: ātraināŖātest=1,ā¦,N,ātrainā©ātest=ā ,I_train _test=\1,ā¦,N\, _train _test= , (9) with |ātrain|=ā0.7āNā|I_train|= 0.7N and |ātest|=Nā|ātrain||I_test|=N-|I_train|. To reduce sensitivity to outliers and to stabilize membership-function design, the Mamdani baseline uses robust percentile-based normalization. For each feature xācx,a,rxā\c_x,a,r\, compute on the training set only: q0.05ā(x)=Q0.05ā(xn:nāātrain), q_0.05(x)=Q_0.05(\x_n:n _train\), (10) q0.95ā(x)=Q0.95ā(xn:nāātrain), q_0.95(x)=Q_0.95(\x_n:n _train\), where Qpā(ā )Q_p(Ā·) denotes the empirical p-quantile. Each sample (train or test) is then clipped and mapped to [0,1][0,1]: xclip x_clip =minā”(maxā”(x,q0.05ā(x)),q0.95ā(x)), = ( (x,\;q_0.05(x)),\;q_0.95(x) ), (11) xnorm x_norm =xclipāq0.05ā(x)q0.95ā(x)āq0.05ā(x)+εā[0,1]. = x_clip-q_0.05(x)q_0.95(x)-q_0.05(x)+ ā[0,1]. (12) This yields norm=[cx,norm,anorm,rnorm]x_norm=[c_x,norm,a_norm,r_norm] T for the Mamdani controller. For the TakagiāSugeno model, membership parameters are derived from training quantiles directly in the original feature scale (Section 3.5), and no additional robust mināmax normalization is applied. The fuzzy partitions used in this work combine triangular and shoulder membership functions. Central linguistic terms are represented using triangular membership functions, while boundary terms are represented using left- and right-shoulder functions to ensure nonzero coverage at the extremes. For a central triangular term with parameters (α,β,γ)(α,β,γ), where αā¤Ī²ā¤Ī³Ī±ā¤Ī²ā¤Ī³, we define μtriā(x;α,β,γ)=maxā”(0,minā”(xāαβāα+ε,γāxγāβ+ε)), _tri(x;α,β,γ)= \! (0,\; \! ( x-αβ-α+ ,\; γ-xγ-β+ ) ), (13) which is bounded in [0,1][0,1] and numerically stable due to ε>0 >0. 3.4 Mamdani Fuzzy Inference System The Mamdani controller operates on the normalized inputs normx_norm and assigns three linguistic terms to each input [mamdani1975experiment, passino1998fuzzy]. For the lateral position feature cxc_x, the terms are Left, Center, and Right. For the area feature a (a distance proxy, since larger image area indicates a closer target), the terms are Far, Mid, and Near. For the aspect ratio feature r, the terms are Wide, Normal, and Tall. All terms are implemented using the same three-term shoulderātriangleāshoulder partition on [0,1][0,1]; only the linguistic names differ by input. To ensure nonzero membership at the endpoints, the Left and Right terms are implemented as shoulder functions, while the Center term is implemented as a triangular function: μLā(x)=1,xā¤0,0.5āx0.5+ε,0<x<0.5,0,xā„0.5,μCā(x)=μtriā(x;0,0.5,1),μRā(x)=0,xā¤0.5,xā0.50.5+ε,0.5<x<1,1,xā„1. gathered _L(x)= cases1,&x⤠0,\\[3.0pt] 0.5-x0.5+ ,&0<x<0.5,\\[6.0pt] 0,&xā„ 0.5, cases _C(x)= _tri(x;0,0.5,1), \\ _R(x)= cases0,&x⤠0.5,\\[3.0pt] x-0.50.5+ ,&0.5<x<1,\\[6.0pt] 1,&xā„ 1. cases gathered (14) When applied to anorma_norm, the sets μL,μC,μR\ _L, _C, _R\ correspond to Far, Mid, Near; when applied to rnormr_norm, they correspond to Wide, Normal, Tall. With three inputs and three terms per input, the rule base contains Nr=33=27N_r=3^3=27 rules, each rule i has the form: āi:If ācxā is āAiā and āaā is āBiā and ārā is āCi,then āĪøā is āDi,R_i:\ If c_x is A_i and a is B_i and r is C_i,\ then Īø is D_i, (15) where Ai,Bi,CiāL,C,RA_i,B_i,C_iā\L,C,R\ and DiD_i is one of five output labels. Rule Firing Strength Given normx_norm, rule i fires with strength (minimum t-norm) wi=minā”(μAiā(cx,norm),μBiā(anorm),μCiā(rnorm)).w_i= \! ( _A_i(c_x,norm),\; _B_i(a_norm),\; _C_i(r_norm) ). (16) Output Sets and Defuzzification Let Īømin _ and Īømax _ denote the minimum and maximum training labels, respectively. Five triangular output membership functions are constructed from [Īømin,Īømax][ _ , _ ] using a padding padĪø=0.05ā(ĪømaxāĪømin)pad_Īø=0.05( _ - _ ) and uniform breakpoints over [ā,h]=[ĪømināpadĪø,Īømax+padĪø][ ,h]=[ _ -pad_Īø,\ _ +pad_Īø]: μSharpLeftā(Īø) _SharpLeft(Īø) =μtriā(Īø;ā,ā,ā+0.25ā(hāā)), = _tri(Īø; , , +25(h- )), (17) μLeftā(Īø) _Left(Īø) =μtriā(Īø;ā,ā+0.25ā(hāā),ā+0.50ā(hāā)), = _tri(Īø; , +25(h- ), +50(h- )), μZeroā(Īø) _Zero(Īø) =μtriā(Īø;ā+0.25ā(hāā),ā+0.50ā(hāā),ā+0.75ā(hāā)), = _tri(Īø; +25(h- ), +50(h- ), +75(h- )), μRightā(Īø) _Right(Īø) =μtriā(Īø;ā+0.50ā(hāā),ā+0.75ā(hāā),h), = _tri(Īø; +50(h- ), +75(h- ),h), μSharpRightā(Īø) _SharpRight(Īø) =μtriā(Īø;ā+0.75ā(hāā),h,h). = _tri(Īø; +75(h- ),h,h). Mamdani implication uses clipping: μiā(Īø)=minā”(wi,μDiā(Īø)), _i(Īø)= (w_i, _D_i(Īø) ), (18) and aggregation uses the maximum operator: μoutā(Īø)=maxi=1,ā¦,27ā”μiā(Īø). _out(Īø)= _i=1,ā¦,27 _i(Īø). (19) The crisp output is computed by centroid defuzzification over Ī=[Īøminā5ā,Īømax+5ā] =[ _ -5 , _ +5 ]: Īø^=ā«ĪĪøāμoutā(Īø)āĪøā«Īμoutā(Īø)āĪø. Īø= _ Īø\, _out(Īø)\,dĪø _ _out(Īø)\,dĪø. (20) In implementation, the integrals are approximated on a uniform grid with M=1201M=1201 points. 3.4.1 Same-Side Rule Logic The rule generator is designed to promote sign consistency between the lateral displacement of the target in the image plane and the predicted yaw command. The normalized horizontal center feature cxc_x determines the turn direction, where Left corresponds to a negative yaw correction, Right corresponds to a positive yaw correction, and Center corresponds to a near-zero yaw command. The normalized area feature anorma_norm and the aspect ratio feature rnormr_norm modulate the aggressiveness of the yaw response. As an example, turns are sharpened when the target is categorized as Near (large anorma_norm), indicating close proximity, and/or Tall (large rnormr_norm), corresponding to vertically elongated or narrow bounding boxes. These conditions empirically correlate with increased sensitivity to lateral misalignment under the adopted aerial viewing geometry. This same-side mapping encourages the direction of the yaw command to remain consistent with the visual lateral offset of the target, while allowing the magnitude of the response to adapt smoothly based on target scale and shape. Rule Base Specification The complete 27-rule base for both the Mamdani and TakagiāSugeno systems is deterministically generated from the Cartesian product of the three linguistic terms for each input: Left,Center,Right\Left,Center,Right\ for cxc_x, Far,Mid,Near\Far,Mid,Near\ for a, and Wide,Normal,Tall\Wide,Normal,Tall\ for r. In the Mamdani controller, each antecedent combination is assigned a linguistic consequent label DiD_i according to the same-side mapping described above. In the TakagiāSugeno model, the same antecedent grid is retained, but each rule uses a learned first-order consequent ziā()=piācx+qiāa+siār+tiz_i(x)=p_ic_x+q_ia+s_ir+t_i instead of a fixed linguistic consequent label. Because the antecedent grid is generated algorithmically and contains 27 rules, the full rule table is provided in Appendix 6 for completeness and reproducibility. 3.5 First-Order TakagiāSugeno Fuzzy Inference System The TakagiāSugeno model uses the same 27-rule antecedent structure but replaces fuzzy consequents with rule-local linear functions [takagi1985fuzzy, sugeno1985industrial]. For readability, we denote the three fuzzy sets for cxc_x as Left, Center, Right; for a as Far, Mid, Near; and for r as Wide, Normal, Tall, although the underlying three-term parameterization follows the same left-shoulder, center-triangle, right-shoulder construction in each case. Unless otherwise stated, this model operates on the original feature scale (cx,a,r)(c_x,a,r). Quantile-Derived Membership Parameters (Coverage-Guaranteed). For each input variable xācx,a,rxā\c_x,a,r\, training-set quantiles q0.05ā(x)q_0.05(x), q0.50ā(x)q_0.50(x), and q0.95ā(x)q_0.95(x) are computed using only samples in ātrainI_train. To guarantee nonzero fuzzy coverage and prevent cases where all rule firing strengths become zero, the TakagiāSugeno model uses a three-term partition consisting of a left shoulder function, a central triangular function, and a right shoulder function: μLeftā(x) _Left(x) =1,xā¤q0.05ā(x),q0.50ā(x)āxq0.50ā(x)āq0.05ā(x)+ε,q0.05ā(x)<x<q0.50ā(x),0,xā„q0.50ā(x), = cases1,&x⤠q_0.05(x),\\[3.0pt] q_0.50(x)-xq_0.50(x)-q_0.05(x)+ ,&q_0.05(x)<x<q_0.50(x),\\[6.0pt] 0,&xā„ q_0.50(x), cases (21) μCenterā(x) _Center(x) =μtriā(x;q0.05ā(x),q0.50ā(x),q0.95ā(x)), = _tri\! (x;q_0.05(x),\,q_0.50(x),\,q_0.95(x) ), (22) μRightā(x) _Right(x) =0,xā¤q0.50ā(x),xāq0.50ā(x)q0.95ā(x)āq0.50ā(x)+ε,q0.50ā(x)<x<q0.95ā(x),1,xā„q0.95ā(x). = cases0,&x⤠q_0.50(x),\\[3.0pt] x-q_0.50(x)q_0.95(x)-q_0.50(x)+ ,&q_0.50(x)<x<q_0.95(x),\\[6.0pt] 1,&xā„ q_0.95(x). cases (23) All quantiles are computed on ātrainI_train only and held fixed for test-time inference. Firing Strengths, Normalization, and Output For each rule i, the firing strength is computed using the product t-norm: wi=μAiā(cx)āμBiā(a)āμCiā(r),w_i= _A_i(c_x)\, _B_i(a)\, _C_i(r), (24) and normalized weights are wĀÆi=wiāk=127wk+ε. w_i= w_i _k=1^27w_k+ . (25) Each rule has a first-order consequent ziā()=piācx+qiāa+siār+ti,z_i(x)=p_ic_x+q_ia+s_ir+t_i, (26) and the Sugeno output is Īø^=āi=127wĀÆiāziā(). Īø= _i=1^27 w_i\,z_i(x). (27) Least-Squares Identification Let βāā108β ^108 collect all consequent coefficients pi,qi,si,tii=127\p_i,q_i,s_i,t_i\_i=1^27. For each training sample n, construct the regression row Ļ(n)āā108 Ļ^(n) ^108 by stacking wĀÆi(n)ācx(n),wĀÆi(n)āa(n),wĀÆi(n)ār(n),wĀÆi(n)i=127\ w_i^(n)c_x^(n), w_i^(n)a^(n), w_i^(n)r^(n), w_i^(n)\_i=1^27. Stacking all training rows yields ΦāāNtrainĆ108 ^N_trainĆ 108 and labels āāNtrainy ^N_train. The parameters are estimated by least squares: β^=argā”minβā”āΦāβāā22, β= _β\| β-y\|_2^2, (28) which is computed using numerically stable solvers (SVD/QR). At test time, the train-derived membership parameters and the learned β β are fixed. For each test sample, membership degrees, firing strengths, normalized weights, and Īø Īø are computed without re-fitting. 3.6 Evaluation Metrics Let ei=Īø^iāĪøie_i= Īø_i- _i denote the test error for sample i. We report mean absolute error (MAE), root mean square error (RMSE), maximum absolute error (MAXAE), and within-threshold accuracy for Ļā1ā,3ā,5āĻā\1 ,3 ,5 \: MAE =1Ntestāāi=1Ntest|ei|, = 1N_test _i=1^N_test|e_i|, (29) RMSE =1Ntestāāi=1Ntestei2, = 1N_test _i=1^N_teste_i^2, (30) MAXAE =maxi=1,ā¦,Ntestā”|ei|, = _i=1,ā¦,N_test|e_i|, (31) AccĀ±Ļ _Ā±Ļ =1Ntestāāi=1Ntestā(|ei|ā¤Ļ)Ć100%. = 1N_test _i=1^N_testI (|e_i|ā¤Ļ )Ć 100\%. (32) Within-threshold accuracy at tighter tolerances (±1ā± 1 and ±3ā± 3 ) is included to quantify fine-grained yaw estimation performance near zero, which is particularly relevant for stable closed-loop heading regulation and avoidance of control chatter. Directional (Side) Consistency (SC) In addition to numerical error metrics, we quantify whether the predicted yaw command preserves the intended turning direction implied by the image-plane lateral displacement. Let mcxm_c_x denote the median of cxc_x computed on the training set only. We define a three-class side label from the horizontal target location: scxā(cx)=Left,cx<mcx,Right,cx>mcx,Center,otherwise.s_c_x(c_x)= casesLeft,&c_x<m_c_x,\\ Right,&c_x>m_c_x,\\ Center,&otherwise. cases (33) Similarly, we map the predicted yaw to a three-class directional label using a small deadband Ī“ (degrees) to avoid counting near-zero corrections as left/right: sĪøā(Īø^)=Left,Īø^<āĪ“,Center,|Īø^|ā¤Ī“,Right,Īø^>Ī“.s_Īø( Īø)= casesLeft,& Īø<-Ī“,\\ Center,&| Īø|ā¤Ī“,\\ Right,& Īø>Ī“. cases (34) Directional (Side) consistency (SC) is then computed on the test set as the percentage of samples whose side labels agree: SideCons=1Ntestāāi=1Ntestā(scxā(cx,i)=sĪøā(Īø^i))Ć100%.SideCons= 1N_test _i=1^N_testI\! (s_c_x(c_x,i)=s_Īø( Īø_i) )Ć 100\%. (35) In our experiments, Ī“=3āĪ“=3 and mcxm_c_x is computed from the training split only and held fixed for test evaluation. Side-consistency is used solely as an evaluation metric and is not explicitly enforced during fuzzy inference or parameter estimation. 4 Dataset Collection and Processing The dataset used in this study was collected and fully documented in our prior work [ahmarijournal, ahmariconf]. In this paper, we reuse the same experimental dataset and ground-truth generation procedure to enable a controlled evaluation of interpretable fuzzy inference systems for vision-based yaw estimation. For completeness and reproducibility, we summarize the key aspects of the hardware setup, data acquisition procedure, and labeling pipeline here, while referring the reader to [ahmarijournal] for full implementation details, system diagrams, and extended experimental context. 4.1 Experimental Setup Data were collected in an indoor motion-capture arena equipped with a VICON system capable of full six degrees-of-freedom (6DoF) tracking for both the UAV and UGV, as illustrated in Figure 1. The VICON infrastructure provides millimeter-level positional accuracy, enabling high-fidelity relative pose annotation synchronized with image frames [merriaux2017study]. During data collection, the UAV remained stationary while a single UGV traversed the scene under diverse headings and positions. This controlled setup was intentionally used to isolate the relationship between detector-derived bounding-box geometry and VICON-derived yaw correction without introducing additional compensation variables associated with moving-camera effects. Because the current framework operates on frame-level bounding-box descriptors, each image frame is treated as an independent observation of the relative target configuration in the UAV camera view [ahmariconf, ahmarijournal]. Figure 1: Top view of the laboratory scene showing VICON camera placement and UGV/UAV layout. The UAV platform is equipped with two monocular cameras: a forward-facing camera (C1) used to observe the UGV for detection and heading estimation, and a downward-facing camera (C2) intended for landing verification once alignment is achieved. The present study uses only the forward-facing stream (C1), consistent with the goal of mapping image-plane bounding-box geometry to a yaw correction command. The data collection pipeline integrates VICON tracking outputs with the ROS-based robotics stack using a hybrid WindowsāLinux environment, as described in [ahmariconf, ahmarijournal]. A custom Python interface is used to synchronize VICON logs and camera timestamps, producing frame-aligned records that associate each captured image with the corresponding ground-truth 6DoF state of the UAV and UGV. This synchronization is essential for producing reliable yaw supervision and for evaluating vision-based estimators under realistic detection noise and viewpoint variation. Each image frame was annotated with a bounding box around the UGV to train a YOLO object detector tailored to the experimental environment [ahmarijournal, ahmariconf]. As part of the detector-development procedure reported in a prior work, augmentation operations such as image flipping, rotation, brightness variation, and small translations were used to improve the robustness of UGV bounding-box detection. Once trained, the detector produces bounding boxes parameterized by pixel coordinates (x1,y1,x2,y2)(x_1,y_1,x_2,y_2) per frame. These coordinates form the basis for the low-dimensional geometric features used in this paper. While the original dataset contains additional metadata (including temporal alignment and 3D pose logs), the focus of the present work is on the mapping from bounding-box descriptors to yaw correction. Ground-truth yaw labels are derived from VICON-based relative pose measurements as explained in equations 3 and 4. 5 Results and Discussion This section reports an experimental evaluation of the proposed three-input fuzzy yaw controllers using five randomized trainātest splits to quantify run-to-run variability. The evaluation is based on VICON-labeled image frames collected with a stationary UAV and a moving UGV; therefore, the reported results quantify yaw-estimation performance under controlled perception conditions. We first analyze the Mamdani controller as an interpretable baseline and then report the performance of the first-order TakagiāSugeno controller with consequent parameters identified by least squares on the training split. All quantitative results are reported as mean ± standard deviation across runs. 5.1 Mamdani Results The Mamdani fuzzy controller is designed as an interpretable baseline that promotes sign consistency between the lateral displacement of the target and the commanded yaw correction. Across five randomized trainātest splits, the three-input Mamdani controller achieves a test-set MAE of 4.041ā±0.054ā4.041 ± 0.054 , RMSE of 4.866ā±0.052ā4.866 ± 0.052 , and MAXAE of 9.332ā±0.091ā9.332 ± 0.091 . Within-threshold accuracies are 17.299%±1.658%17.299\%± 1.658\%, 41.005%±1.036%41.005\%± 1.036\%, and 64.279%±1.042%64.279\%± 1.042\% for tolerance bands of ±1ā± 1 , ±3ā± 3 , and ±5ā± 5 , respectively. These results indicate that, while the Mamdani controller reliably preserves the dominant turning direction, its fixed output membership functions and centroid defuzzification lead to limited numerical precision, particularly for small yaw corrections. Beyond numerical accuracy, we evaluate SC between the image-plane lateral displacement and the predicted yaw sign using the same criterion across all runs. Averaged over five splits, the Mamdani controller achieves 91.183%±0.712%91.183\%± 0.712\% directional agreement between the sign implied by cxc_x and the sign of the predicted yaw. For completeness, Table 1 presents a comparative summary of the test-set performance of the Mamdani and TakagiāSugeno models across all reported metrics. Figure 2: Mamdani controller predictions versus ground-truth yaw on the test set. The solid black line denotes ideal prediction Īø^=Īø Īø=Īø, while the dashed and dotted gray lines indicate ±1ā± 1 and ±3ā± 3 error bands, respectively. Points are colored by absolute prediction error |Īø^āĪø|| Īø-Īø|. The representative split yields MAE =4.086ā=4.086 , RMSE =4.912ā=4.912 , MAXAE =9.195ā=9.195 , and Acc±1ā=18.3%± 1 =18.3\%. Figure 2 shows the predicted yaw angle as a function of the ground-truth yaw on the test set. The predictions exhibit a clear monotonic relationship with the ground truth, indicating that the controller captures the dominant dependency between image-plane lateral displacement and yaw correction. Most predictions follow the identity-line trend, confirming that the same-side rule logic provides directionally meaningful behavior in the majority of cases. However, structured deviations can be observed at larger yaw magnitudes, where the predicted responses bend away from the ideal line. This behavior is consistent with the fixed fuzzy output partitions used by Mamdani inference. These effects are inherent to the Mamdani formulation, which relies on fixed output membership functions and maxāmin inference. The corresponding error histogram is shown in Figure 3. The error distribution spans both negative and positive values and remains relatively broad, which is consistent with the moderate RMSE and the limited percentage of predictions within the tighter ±1ā± 1 and ±3ā± 3 bands. Figure 3: Test-set error histogram for the Mamdani controller on a representative split. The solid black vertical line denotes zero error, and the dashed gray lines denote ± . The representative split yields MAE =4.086ā=4.086 , Acc±3ā=40.2%± 3 =40.2\%, and Acc±5ā=64.0%± 5 =64.0\%. This behavior is consistent with the limited expressiveness of fixed fuzzy consequents and the absence of data-driven tuning in the Mamdani framework. Nevertheless, the Mamdani controller provides an interpretable baseline whose behavior is predictable and stable, making it suitable for safety-critical or explainability-focused applications. These limitations, however, motivate the adoption of a TakagiāSugeno fuzzy inference system, which replaces fixed fuzzy consequents with data-driven local linear models to improve regression fidelity while retaining an interpretable rule structure. 5.2 TakagiāSugeno Results The first-order TakagiāSugeno controller retains the same antecedent structure and rule base as the Mamdani controller but replaces fuzzy output sets with rule-local linear consequents identified via least-squares regression. This formulation preserves the interpretability of the fuzzy antecedents while substantially increasing expressive power for continuous regression. Table 1 provides a consolidated comparison of the test-set performance of the Mamdani and TakagiāSugeno controllers across all reported metrics, including numerical accuracy and directional side-consistency. Relative to the Mamdani baseline, the TakagiāSugeno model achieves a marked reduction in both MAE and RMSE, along with a substantial improvement in within-threshold accuracy across all tolerance bands. Directional side-consistency remains high for both models. The Mamdani controller achieves slightly higher SC, consistent with its sign-oriented rule design, while the TakagiāSugeno model remains comparably consistent despite learning its consequents from data. Table 1: Test-set performance comparison of Mamdani and TakagiāSugeno controllers (3 inputs), reported as mean ± standard deviation over five randomized splits. Model MAE RMSE MAXAE ±1ā± 1 ±3ā± 3 ±5ā± 5 SC Mamdani 4.04±0.054.04±0.05 4.87±0.054.87±0.05 9.33±0.099.33±0.09 17.30±1.6617.30±1.66 41.01±1.0441.01±1.04 64.28±1.0464.28±1.04 91.18±0.7191.18 0.71 Sugeno 0.14±0.000.14 0.00 0.20±0.010.20 0.01 1.25±0.121.25 0.12 99.68±0.2799.68 0.27 100.00±0.00100.00 0.00 100.00±0.00100.00 0.00 90.25±0.6190.25±0.61 Figure 4 illustrates the predicted yaw versus ground truth on the test set. In contrast to the Mamdani controller, the Sugeno predictions track the identity line more closely than the Mamdani controller across most of the yaw range. The point cloud remains tightly concentrated around the ideal line, and nearly all predictions fall within the ±1ā± 1 band, indicating that the learned rule-local linear consequents substantially improve numerical resolution without saturation. Figure 4: TakagiāSugeno controller predictions versus ground truth on the test set. The solid black line denotes ideal prediction Īø^=Īø Īø=Īø, and the dashed gray lines indicate the ±1ā± 1 error band. Points are colored by absolute prediction error |Īø^āĪø|| Īø-Īø|. The representative split yields MAE =0.141ā=0.141 , RMSE =0.197ā=0.197 , MAXAE =1.124ā=1.124 , and Acc±1ā=99.9%± 1 =99.9\%. Notably, the Sugeno controller maintains high accuracy near zero yaw while simultaneously preserving fidelity at large magnitudes, a regime where the Mamdani controller exhibits increased variance. The error histogram in Figure 5 is sharply peaked around zero, with a near-zero mean error, indicating negligible systematic bias and substantially lower dispersion than the Mamdani baseline. Across five randomized splits, the TakagiāSugeno model achieves a test-set MAE of 0.140ā±0.003ā0.140 ± 0.003 , RMSE of 0.200ā±0.008ā0.200 ± 0.008 , and MAXAE of 1.254ā±0.121ā1.254 ± 0.121 (Table 1). The model also achieves 99.676%±0.270%99.676\%± 0.270\% accuracy within ±1ā± 1 and 100.000%±0.000%100.000\%± 0.000\% accuracy within both ±3ā± 3 and ±5ā± 5 . Compared to the Mamdani controller, the Sugeno model substantially reduces the frequency and magnitude of large errors and exhibits a more compact error distribution around zero. Figure 5: Test-set error histogram for the TakagiāSugeno controller. The solid black vertical line denotes zero error, and the dashed gray lines denote ± . The representative split yields MAE =0.141ā=0.141 , Acc±3ā=100.0%± 3 =100.0\%, and Acc±5ā=100.0%± 5 =100.0\%. To complement the quantitative evaluation, Figure 6 provides a frame-level qualitative comparison on representative test samples, showing the detected bounding boxes and the corresponding yaw predictions relative to the VICON ground truth. Figure 6: Qualitative comparison on representative frames (left, center, and right yaw cases). Each cell shows the same detected bounding box with overlayed yaw value: Predicted yaw Īø Īø from the corresponding controller. The TakagiāSugeno model tracks the ground truth more closely across the full yaw range, while the Mamdani baseline exhibits larger deviations in magnitude under challenging cases. The displayed frames are drawn from the held-out test set to represent left, center, and right yaw cases using the same instances for both controllers, ensuring a consistent qualitative comparison. Table 2 reports the absolute prediction error for the representative frames shown in Figure 6. While both controllers preserve the correct turning direction, the Mamdani controller exhibits larger magnitude errors, particularly in Frame 3, where output saturation effects become evident. In contrast, the TakagiāSugeno model consistently achieves lower absolute error across all frames, including near-zero and large-magnitude yaw cases. This frame-level comparison reinforces the quantitative results and highlights the improved numerical fidelity of the Sugeno formulation. Table 2: Absolute prediction error corresponding to the qualitative examples shown in Figure 6. Frame |Ī|| | Mamdani (deg) |Ī|| | Sugeno (deg) Frame 1 0.43 0.11 Frame 2 0.27 0.05 Frame 3 5.25 0.02 Using the same side-consistency criterion adopted for the Mamdani evaluation, the Sugeno controller achieves 90.254%±0.612%90.254\%± 0.612\% SC on the test set, indicating that the learned consequents largely preserve the intended same-side behavior while allowing deviations in visually ambiguous cases where minimizing regression error is favored. The comparison between the Mamdani and TakagiāSugeno controllers highlights a clear trade-off between interpretability and predictive accuracy. The Mamdani controller offers transparent reasoning through fixed linguistic consequents and is designed to promote sign-consistent behavior through its same-side rule mapping, but its performance is limited by coarse output resolution and centroid defuzzification. In contrast, the Sugeno controller preserves the same interpretable antecedent structure while significantly improving accuracy through least-squares identification of rule consequents. Importantly, both controllers operate on the same three input features derived from YOLO bounding boxes. Therefore, the observed performance improvement can be attributed to the inference and consequent formulation, rather than changes in feature representation. These results demonstrate that the proposed Sugeno formulation provides a strong balance between interpretability and performance, making it well suited for vision-based yaw control in UAVāUGV interaction scenarios and motivating a broader comparison with standard low-dimensional regression baselines operating on the same features. 6 Comparative Evaluation This section evaluates the proposed fuzzy inference models from three complementary perspectives. First, the models are compared with standard regression baselines operating on the same three bounding-box descriptors. Second, the computational volume and inference time of each method are reported. Third, the proposed yaw-estimation formulation is contrasted with image-based fuzzy visual servoing to clarify the difference between predicting an angular correction from bounding-box geometry and directly regulating image-plane error. 6.1 Comparison with Standard Regression Baselines To evaluate the fuzzy inference models against conventional low-dimensional regression approaches, five standard regressors are considered using the same three bounding-box descriptors. The compared regressors include linear regression, ridge regression, support vector regression (SVR), random forest regression, and a shallow multilayer perceptron (MLP). All models are evaluated using the same five-run 70/30 trainātest protocol and the same performance metrics used for the fuzzy models. Table 3: Comparison with standard regression baselines using the same three bounding-box features. Results are reported as mean ± standard deviation over five randomized trainātest splits. Model MAE (ā) RMSE (ā) MAXAE (ā) Acc±1ā± 1 (%) SC (%) Linear Regression 0.668±0.0090.668± 0.009 0.887±0.0080.887± 0.008 5.443±0.2425.443± 0.242 78.271±0.94878.271± 0.948 90.124±0.55590.124± 0.555 Ridge Regression 0.658±0.0090.658± 0.009 0.889±0.0100.889± 0.010 5.641±0.2655.641± 0.265 79.082±0.89579.082± 0.895 90.113±0.54290.113± 0.542 SVR 0.135±0.0030.135± 0.003 0.224±0.0350.224± 0.035 3.045±3.3593.045± 3.359 99.427±0.20499.427± 0.204 90.189±0.54990.189± 0.549 Random Forest 0.124±0.0030.124± 0.003 0.204±0.0060.204± 0.006 1.404±0.1811.404± 0.181 99.630±0.12799.630± 0.127 90.232±0.55390.232± 0.553 Shallow MLP 0.491±0.0120.491± 0.012 0.632±0.0250.632± 0.025 2.329±0.3052.329± 0.305 87.618±2.55687.618± 2.556 90.773±0.62090.773± 0.620 Mamdani 4.041±0.0544.041± 0.054 4.866±0.0524.866± 0.052 9.332±0.0919.332± 0.091 17.299±1.65817.299± 1.658 91.183±0.71291.183± 0.712 TakagiāSugeno 0.140±0.0030.140± 0.003 0.200±0.0080.200± 0.008 1.254±0.1211.254± 0.121 99.676±0.27099.676± 0.270 90.254±0.61290.254± 0.612 The comparison shows that the three bounding-box descriptors contain strong predictive information for yaw estimation. Linear regression and ridge regression capture the dominant geometric trend, but their MAE, RMSE, and MAXAE remain substantially higher than those of the nonlinear regressors and the TakagiāSugeno model. This indicates that a single global linear mapping is insufficient to fully represent the feature-to-yaw relationship in the present dataset. Among the conventional regression baselines, random forest achieves the lowest average MAE, while SVR also provides competitive average error. However, the TakagiāSugeno model achieves the lowest RMSE, lowest MAXAE, and highest Acc±1ā± 1 among the evaluated models. The difference between random forest and TakagiāSugeno in MAE is small, whereas the TakagiāSugeno model provides a lower RMSE and a lower MAXAE, indicating a tighter overall and worst-case error profile under the reported test protocol. These results show that the TakagiāSugeno formulation is competitive with standard low-dimensional regression baselines while preserving an interpretable fuzzy antecedent-rule structure. The Mamdani controller is less accurate as a continuous regressor, but it achieves the highest side-consistency, confirming its role as an interpretable directional baseline. 6.2 Computational-Volume and Inference-Time Comparison For practical UAV applications, computational cost is an important consideration in addition to prediction accuracy. The fuzzy models and standard regression baselines are therefore compared in terms of general model structure, fitting time, and inference time per test sample. Timing values are measured on the same hardware and averaged over the test-set prediction procedure. Table 4: Computational-volume and inference-time comparison. Timing values are reported as mean ± standard deviation over five randomized trainātest splits. Model Main structure Fit time (s) Inference time/sample (ms) Linear Regression Linear mapping 0.001±0.0010.001± 0.001 0.000053±0.0000020.000053± 0.000002 Ridge Regression Regularized linear mapping 0.474±1.0280.474± 1.028 0.000070±0.0000180.000070± 0.000018 SVR RBF-kernel regression 2.977±0.5552.977± 0.555 0.084327±0.0011270.084327± 0.001127 Random Forest Tree ensemble 1.519±0.0701.519± 0.070 0.013159±0.0028360.013159± 0.002836 Shallow MLP One-hidden-layer neural regressor 0.493±0.0310.493± 0.031 0.000124±0.0000160.000124± 0.000016 Mamdani Fuzzy rules + centroid defuzzification 0.001±0.0000.001± 0.000 0.090467±0.0005710.090467± 0.000571 TakagiāSugeno Fuzzy rules + weighted consequent average 0.020±0.0010.020± 0.001 0.000538±0.0000200.000538± 0.000020 The computational comparison shows that all evaluated methods operate on compact three-dimensional feature vectors rather than high-dimensional image inputs. Linear and ridge regression provide the lowest inference cost, but their prediction errors are substantially larger than those of the strongest nonlinear methods. The shallow MLP also has low inference time, but its numerical accuracy is lower than that of the TakagiāSugeno model, SVR, and random forest. SVR provides competitive average error but has substantially higher measured inference time than the TakagiāSugeno model. Random forest achieves the lowest MAE among the evaluated models, but it relies on an ensemble structure and has a higher measured inference time than the TakagiāSugeno model. The TakagiāSugeno model provides a favorable compromise between accuracy, interpretability, and computational cost. Its inference is based on fuzzy rule activation, rule-local consequent evaluation, and normalized weighted averaging. Under the measured test protocol, it achieves substantially lower inference time than SVR, random forest, and Mamdani inference while maintaining competitive accuracy and the lowest reported MAXAE. The Mamdani model remains useful as an interpretable directional baseline, but its centroid-defuzzification step is computationally heavier in the present implementation. 6.3 Comparison with Image-Based Fuzzy Visual Servoing A closely related vision-based fuzzy control framework is presented in [visualServoFuzzyUAV], where fuzzy inference is applied to reactive visual servoing of UAV yaw. In that approach, the controller inputs are the horizontal pixel residual exā(k)=xtargetā(k)āxcenter,e_x(k)=x_target(k)-x_center, and its discrete-time derivative Īāexā(k)=exā(k)āexā(kā1), e_x(k)=e_x(k)-e_x(k-1), while the output is a yaw-related control command that drives the target toward image centering. The framework is validated through laboratory experiments and UAV flight tests, where performance is evaluated primarily using image-plane tracking errors and controller response trajectories. Although both approaches employ fuzzy inference for vision-driven yaw behavior, they differ in representation, objective, and evaluation domain. The method in [visualServoFuzzyUAV] operates directly in pixel coordinates. Under the standard pinhole projection model, pixel displacement depends on focal length and viewing geometry, implying that the magnitude of exā(k)e_x(k) varies with camera intrinsic parameters and target distance. As a result, membership function ranges and rule sensitivities are tied to the sensing configuration used during controller design, and transferring the controller across different camera setups may require recalibration. The controller also incorporates the discrete derivative Īāexā(k) e_x(k) to improve responsiveness. However, discrete differentiation can amplify high-frequency measurement noise, meaning that tracker jitter, partial occlusions, or bounding-box fluctuations may introduce transient spikes in the derivative term unless additional filtering is applied. This reflects a reactive visual servoing design in which instantaneous perception residuals are directly coupled to control actions. The methodological differences between the image-based fuzzy servoing approach and the proposed framework are summarized in Table 5. Table 5: Methodological comparison with the image-based fuzzy visual servoing framework in [visualServoFuzzyUAV]. Aspect Image-Based Fuzzy Servoing [visualServoFuzzyUAV] Proposed Framework Objective Image-plane yaw centering Angular yaw estimation Inputs Pixel residual exe_x and derivative Īāex e_x Normalized geometric features Camera Dependence Residual magnitude tied to camera parameters Reduced pixel-scale dependence Temporal Handling Discrete derivative term No derivative input Noise Sensitivity Derivative amplifies tracking jitter Reduced sensitivity to measurement noise Output Yaw control command Yaw correction angle Evaluation Image-plane tracking error Angular error (VICON ground truth) Inference Study Single fuzzy formulation Mamdani vs. Sugeno The proposed framework instead formulates yaw alignment as an angular correction problem. Visual observations are mapped to resolution-normalized geometric features derived from detected bounding boxes, and fuzzy inference predicts a continuous yaw correction variable that is evaluated directly against independent ground-truth orientation measurements. By avoiding discrete error differentiation as a primary input, the proposed approach reduces sensitivity to derivative-induced noise amplification. Performance is therefore characterized statistically in the angular domain, enabling direct assessment of heading estimation accuracy under controlled experimental conditions. 7 Conclusion This paper introduced a structured and reproducible fuzzy-logic framework for estimating continuous yaw correction commands from compact vision-based descriptors derived from YOLO bounding boxes. By operating directly on low-dimensional geometric features, the proposed approach enables efficient and interpretable vision-to-control mapping without reliance on external localization at inference time or high-dimensional image representations. Two fuzzy inference formulations were evaluated under identical conditions: a Mamdani controller serving as an interpretable baseline and a first-order TakagiāSugeno model with consequents identified via least-squares regression. While the Mamdani system demonstrated reliable directional behavior and consistent sign alignment between lateral displacement and yaw output, its performance was constrained by the inherent discretization and saturation effects of fixed output membership functions. In contrast, the TakagiāSugeno formulation preserved the same fuzzy antecedent structure while achieving substantially higher numerical accuracy, attaining a test-set MAE of 0.140ā±0.003ā0.140 ± 0.003 , RMSE of 0.200ā±0.008ā0.200 ± 0.008 , and MAXAE of 1.254ā±0.121ā1.254 ± 0.121 across five randomized splits, with 99.676%±0.270%99.676\%± 0.270\%, 100.000%±0.000%100.000\%± 0.000\%, and 100.000%±0.000%100.000\%± 0.000\% of predictions within ±1ā± 1 , ±3ā± 3 , and ±5ā± 5 , respectively. Directional agreement between image-plane lateral displacement and predicted yaw sign was maintained at 90.254%±0.612%90.254\%± 0.612\%, indicating that the learned consequents preserve the intended turning behavior in most cases. A central contribution of this work lies in the controlled comparison of Mamdani and TakagiāSugeno inference under a shared antecedent rule structure, training-derived preprocessing, and training-derived membership-function parameterization. This design clarifies how the consequent formulation affects vision-based yaw-estimation performance while preserving interpretability at the fuzzy-antecedent level. The comparative results show that the Mamdani controller remains useful for explainable directional reasoning, whereas the TakagiāSugeno model provides substantially higher numerical precision and a favorable accuracyāinterpretabilityācomputational-cost trade-off relative to standard low-dimensional regression baselines. Importantly, all models operate on compact geometric features derived from detected bounding boxes and avoid reliance on high-dimensional image representations. Overall, the proposed framework offers a practical, data-efficient, and computationally lightweight approach for vision-based UAV yaw guidance and establishes a clear methodological foundation for integrating fuzzy inference into perception-driven robotic control pipelines. 8 Future Work This work establishes a reproducible framework for mapping vision-derived bounding-box geometry to continuous yaw commands using interpretable fuzzy inference. Several extensions can further expand the capabilities of the proposed approach. Future work will evaluate the framework under dynamic UAV ego-motion and closed-loop flight conditions, where camera motion, detector jitter, and control latency may affect the observed bounding-box descriptors. First, the current formulation operates on frame-level geometric descriptors. Incorporating temporal information, such as short-term motion cues or sequential feature aggregation, may improve robustness under rapid target motion or intermittent detections. Second, the present design employs a compact three-term fuzzy partition per input to preserve interpretability and computational efficiency; future work may explore adaptive or hierarchical fuzzy structures that allow richer partitions while maintaining transparent rule representations. Finally, extending evaluation beyond the controlled indoor VICON environment to more diverse operating conditions, including outdoor scenarios, illumination changes, background variability, different camera configurations, and alternative detector architectures, will provide further insight into the robustness of bounding-box-based fuzzy inference for perception-driven UAV guidance. Acknowledgements This work was supported in part by the National Science Foundation (NSF) under Grant No. 2301553; in part by the National Aeronautics and Space Administration (NASA) University Leadership Initiative (ULI) under Grant Nos. 80NSSC20M0161 and 80NSSC25M7098; in part by the U.S. Department of Transportation (USDOT) under Grant No. 69A3552348327; and in part by the North Carolina Department of Transportation (NCDOT) under Grant No. RP2025-43. References Appendix A Fuzzy Rule Base This appendix lists the complete 27-rule antecedent grid used in both fuzzy systems. The linguistic terms are defined as follows: (i) cxc_x: Left, Center, Right; (i) a: Far, Mid, Near (larger a implies a closer target); (i) r: Wide, Normal, Tall (larger r implies a taller or narrower box). The Mamdani consequent label DiāSharpLeft,Left,Zero,Right,SharpRightD_iā\SharpLeft,Left,Zero,Right,SharpRight\ is assigned by the same-side mapping implemented in our rule-generation function. For the TakagiāSugeno model, the antecedents are identical, but each rule uses a first-order consequent ziā()=piācx+qiāa+siār+tiz_i(x)=p_ic_x+q_ia+s_ir+t_i, where (pi,qi,si,ti)(p_i,q_i,s_i,t_i) are learned by least squares on the training set. Table 6: Complete 27-rule base with requested linguistic terms. Sugeno uses the same antecedents with rule-local linear consequents ziā()=piācx+qiāa+siār+tiz_i(x)=p_ic_x+q_ia+s_ir+t_i. Rule i cxc_x term a term r term Mamdani consequent DiD_i Sugeno consequent ziā()z_i(x) 1 Left Far Wide Left p1ācx+q1āa+s1ār+t1p_1c_x+q_1a+s_1r+t_1 2 Left Far Normal Left p2ācx+q2āa+s2ār+t2p_2c_x+q_2a+s_2r+t_2 3 Left Far Tall SharpLeft p3ācx+q3āa+s3ār+t3p_3c_x+q_3a+s_3r+t_3 4 Left Mid Wide Left p4ācx+q4āa+s4ār+t4p_4c_x+q_4a+s_4r+t_4 5 Left Mid Normal Left p5ācx+q5āa+s5ār+t5p_5c_x+q_5a+s_5r+t_5 6 Left Mid Tall SharpLeft p6ācx+q6āa+s6ār+t6p_6c_x+q_6a+s_6r+t_6 7 Left Near Wide SharpLeft p7ācx+q7āa+s7ār+t7p_7c_x+q_7a+s_7r+t_7 8 Left Near Normal SharpLeft p8ācx+q8āa+s8ār+t8p_8c_x+q_8a+s_8r+t_8 9 Left Near Tall SharpLeft p9ācx+q9āa+s9ār+t9p_9c_x+q_9a+s_9r+t_9 10 Center Far Wide Zero p10ācx+q10āa+s10ār+t10p_10c_x+q_10a+s_10r+t_10 11 Center Far Normal Zero p11ācx+q11āa+s11ār+t11p_11c_x+q_11a+s_11r+t_11 12 Center Far Tall Zero p12ācx+q12āa+s12ār+t12p_12c_x+q_12a+s_12r+t_12 13 Center Mid Wide Zero p13ācx+q13āa+s13ār+t13p_13c_x+q_13a+s_13r+t_13 14 Center Mid Normal Zero p14ācx+q14āa+s14ār+t14p_14c_x+q_14a+s_14r+t_14 15 Center Mid Tall Zero p15ācx+q15āa+s15ār+t15p_15c_x+q_15a+s_15r+t_15 16 Center Near Wide Zero p16ācx+q16āa+s16ār+t16p_16c_x+q_16a+s_16r+t_16 17 Center Near Normal Zero p17ācx+q17āa+s17ār+t17p_17c_x+q_17a+s_17r+t_17 18 Center Near Tall Zero p18ācx+q18āa+s18ār+t18p_18c_x+q_18a+s_18r+t_18 19 Right Far Wide Right p19ācx+q19āa+s19ār+t19p_19c_x+q_19a+s_19r+t_19 20 Right Far Normal Right p20ācx+q20āa+s20ār+t20p_20c_x+q_20a+s_20r+t_20 21 Right Far Tall SharpRight p21ācx+q21āa+s21ār+t21p_21c_x+q_21a+s_21r+t_21 22 Right Mid Wide Right p22ācx+q22āa+s22ār+t22p_22c_x+q_22a+s_22r+t_22 23 Right Mid Normal Right p23ācx+q23āa+s23ār+t23p_23c_x+q_23a+s_23r+t_23 24 Right Mid Tall SharpRight p24ācx+q24āa+s24ār+t24p_24c_x+q_24a+s_24r+t_24 25 Right Near Wide SharpRight p25ācx+q25āa+s25ār+t25p_25c_x+q_25a+s_25r+t_25 26 Right Near Normal SharpRight p26ācx+q26āa+s26ār+t26p_26c_x+q_26a+s_26r+t_26 27 Right Near Tall SharpRight p27ācx+q27āa+s27ār+t27p_27c_x+q_27a+s_27r+t_27