Paper deep dive
DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation
Zihan Yang, Yang Guo, Hongxing Zhang, Dan Lu, Siyuan Yao
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cues. Previous works typically attempt to construct discriminative representation of the small objects. However, the wide range frequency domain noises and label ambiguities have been greatly overlooked, which significantly hinders the accurate localization. To address these issues, we propose a novel small object detection (SOD) detector termed DyFrDet, which is able to precisely localize the small object by dynamically suppressing the background distractions in frequency domain. Specifically, we propose a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) to adaptively suppress low-frequency redundancy and excessive high-frequency noises. The DyFrFPN transforms the hierarchical features into frequency domain representation, and introduces a Dynamic Band Predictor (DBP) to preserve the discriminative components for small object identification. Afterwards, we present a novel Label Disambiguation Module (LDM), which leverages probabilistic distributions to explicitly model and alleviate the inherent ambiguity of target labels, yielding efficient improvement in localization precision of the small objects with low-resolution. Extensive experiments demonstrate that DyFrDet achieves state-of-the-art performance across multiple benchmarks, indicating its effectiveness and robustness in various challenging scenarios. Our code is available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.02495v1
- Canonical: https://arxiv.org/abs/2608.02495v1
Trouble viewing inline? Open PDF directly →
Full Text
56,866 characters extracted from source content.
Expand or collapse full text
by DyFrDet: Towards Accurate Small Object Detection via Dynamic Frequency Suppression with Label Disambiguation Zihan Yang zihanyang@buaa.edu.cn 0009-0003-8818-5360 Hangzhou International Innovation Institute of Beihang UniversityHangzhouChina , Yang Guo guoyang4409@gmail.com 0009-0000-0455-3217 Beijing University of Posts and TelecommunicationsBeijingChina , Hongxing Zhang hongxingzhang@bupt.edu.cn 0000-0001-6637-5655 Beijing University of Posts and TelecommunicationsBeijingChina , Dan Lu danlu@buaa.edu.cn 0000-0002-9274-9398 Hangzhou International Innovation Institute of Beihang UniversityHangzhouChina and Siyuan Yao yaosiyuan04@gmail.com 0000-0002-8479-5124 Shenzhen Campus of Sun Yat-sen UniversityShenzhenChina (2026) Abstract. Despite the remarkable progress over the past decades, accurately identifying small objects remains challenging because of their insufficient visual cues. Previous works typically attempt to construct discriminative representation of the small objects. However, the wide range frequency domain noises and label ambiguities have been greatly overlooked, which significantly hinders the accurate localization. To address these issues, we propose a novel small object detection (SOD) detector termed DyFrDet, which is able to precisely localize the small object by dynamically suppressing the background distractions in frequency domain. Specifically, we propose a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) to adaptively suppress low-frequency redundancy and excessive high-frequency noises. The DyFrFPN transforms the hierarchical features into frequency domain representation, and introduces a Dynamic Band Predictor (DBP) to preserve the discriminative components for small object identification. Afterwards, we present a novel Label Disambiguation Module (LDM), which leverages probabilistic distributions to explicitly model and alleviate the inherent ambiguity of target labels, yielding efficient improvement in localization precision of the small objects with low-resolution. Extensive experiments demonstrate that DyFrDet achieves state-of-the-art performance across multiple benchmarks, indicating its effectiveness and robustness in various challenging scenarios. Our code is available at https://github.com/ManOfStory/DyFrDet. Small Object Detection, Frequency-aware Suppression, Label Disambiguation †journalyear: 2026†copyright: c†conference: Proceedings of the 34th ACM International Conference on Multimedia; November 10–14, 2026; Rio de Janeiro, Brazil†booktitle: Proceedings of the 34th ACM International Conference on Multimedia (M ’26), November 10–14, 2026, Rio de Janeiro, Brazil†doi: 10.1145/3767308.3835215†isbn: 979-8-4007-2213-4/2026/11†ccs: Computing methodologies Object detection 1. Introduction Small Object Detection (SOD), as a critical subtask of generic object detection, aims to accurately localize and recognize objects with limited size. SOD plays an essential role in a wide range of real-world applications, including unmanned aerial vehicle (UAV) surveillance, remote sensing localization, scene monitoring, obstacle avoidance, and autonomous driving. However, due to the lack of visual cues and complex scenarios, directly applying generic object detectors (Girshick et al., 2014; Ren et al., 2016; Liu et al., 2016; Lin et al., 2017b; Carion et al., 2020; Zhu et al., 2021) to small object detection often leads to significant performance degradation. Figure 1. (a) Visualization of the heatmap using different frequency suppression settings. “GT” denotes ground-truth annotations. “Encoder Feature” is extracted from CFINet (Yuan et al., 2023). “Remove Low” applies the low-frequency suppression strategy from HS-FPN (Shi et al., 2025). “Remove Low &\& High” further suppresses both low and high frequency components. (b) Label ambiguities in small targets. The green boxes indicate the ground truth annotations, whereas the red boxes are closer to the actual object location. Please zoom in for details. From a technical viewpoint, contemporary small object detection frameworks (Tan et al., 2020; Yang et al., 2022; Yuan et al., 2023; Liu et al., 2024; Sun et al., 2025) predominantly make efforts on constructing semantically discriminative features to alleviate the deficiency of feature representation inherent to small objects. Although these approaches yield significant improvements in general small object detection, they still struggle to maintain robustness in highly cluttered complex scenarios. Primarily, due to the diminutive spatial footprint, small objects are highly susceptible to interference from complex background textures and fine-grained noise, which often leads to false positives or missed detections. Some advanced approaches like HS-FPN (Shi et al., 2025) introduce a Discrete Cosine Transform (DCT) based filtering strategy to remove the low-frequency components, aiming to encourage the model to focus on the attentive small objects rather than large-scale distractors. However, the high-frequency distractors may still hamper the localization capability. Besides, the low resolution of small objects often leads to blurred boundaries, which inevitably introduces label ambiguity to supervise the model training, making it challenging to achieve accurate object localization. To address these issues, in this paper, we propose a novel dynamic frequency suppressive detector termed DyFrDet, which is able to precisely localize the small object by dynamically suppressing the background distractions in frequency domain. As shown in Fig. 1(a), the encoder features are typically noisy in complex local regions. If the low-frequency redundancies and high-frequency noises can be efficiently suppressed, the detection performance can be greatly boosted. To achieve this, DyFrDet introduces a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) to transform the hierarchical pyramid features into different frequency components. A Dynamic Band Predictor (DBP) is designed to control the frequency representation, allowing the detector to suppress both low-frequency redundancies and high-frequency noises simultaneously. Meanwhile, as the low-resolution small objects often introduce label ambiguity, we propose a Label Disambiguation Module (LDM) to alleviate the performance degradation due to the unreliable model training under ambiguous annotations. The small object’s coordinate regression is tackled as a position-aware distributional prediction task. By paying more attention to the high confidence samples and restraining the ambiguous samples, DyFrFPN can predict the small object’s localization more precisely. Experimental results on several popular benchmarks validate the effectiveness of our proposed method. In conclusion, the main contributions of this paper are summarized as follows: • We propose DyFrDet, which introduces Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) to decompose the hierarchical pyramid features into different frequency components, and suppresses both low-frequency redundancies and high-frequency noises simultaneously for small object detection. • We introduce a Label Disambiguation Module (LDM), which pays more attention to the high confidence samples and restrains the ambiguous samples, leading to more accurate bounding box estimation. • Extensive experiments on several popular datasets verify the effectiveness and robustness of DyFrDet in handling extremely small objects under both densely and sparsely distributed scenarios. 2. Related Work 2.1. Small Object Detection Recent advances in small object detection can be broadly categorized into three directions: data-oriented, architecture-oriented, and feature-oriented methods. For data-oriented approaches, (Kisantal et al., 2019; Chen et al., 2019) improve performance through data augmentation, while (Xu et al., 2021, 2022a, 2022b; Yuan et al., 2023) optimize the label assignment process during training, alleviating the scarcity of positive samples for small objects. These methods enhance supervision from the training perspective, thereby improving detection performance. For architecture-oriented methods, FPN (Lin et al., 2017a) introduces a feature pyramid that leverages top-down aggregated multi-scale features for detection. QueryDet (Yang et al., 2022) proposes a cascade architecture and designs a Cascade Sparse Query (CSQ) mechanism to further boost performance. Other works (Li et al., 2019; Ghiasi et al., 2019; Tan et al., 2020; Qiao et al., 2021; Yao et al., 2026, 2024) also explore advanced architectural designs, achieving improved performance through better multi-scale representation and feature utilization. For feature-oriented methods, (Wu et al., 2020; Kim et al., 2021) adopt similarity learning to enhance representations by leveraging similar small objects. HS-FPN (Shi et al., 2025) introduces a high-pass filter to suppress low-frequency components, encouraging the network to focus more on small objects. In addition, Sun et al. (Sun et al., 2025) observe that high-frequency noise can degrade small object features and propose Spectral Enhancement (SET), which employs a heterogeneous architecture for foreground and background feature refinement. Different from previous works, our method dynamically attenuates both low-frequency redundancy and high-frequency noise in a channel-wise manner, enabling more flexible and adaptive frequency modulation. 2.2. Frequency Representation As an important tool in image processing, frequency domain analysis is widely employed in deep learning research. For feature modeling, Chi et al. (Chi et al., 2020) propose Fast Fourier Convolution (FFC), which extends conventional convolution by incorporating frequency-domain modeling, enabling non-local receptive fields and cross-scale feature interactions. Chen et al. (Chen et al., 2025) introduce FDAM, which integrates dynamic high-pass and low-pass filters into the network architecture to address a key limitation of Vision Transformers (ViTs), namely their inherent low-pass filtering effect that leads to frequency attenuation and loss of fine-grained details. Qin et al. (Qin et al., 2021) show that global average pooling (GAP) is a special case of the discrete cosine transform (DCT), and further generalize channel attention mechanisms in the frequency domain via the proposed FcaNet, achieving strong performance on benchmarks such as ImageNet (Deng et al., 2009) and COCO (Lin et al., 2014). SpectFormer (Patro et al., 2025) introduces a hybrid architecture that combines spectral layers with multi-head self-attention, enabling the model to jointly capture global frequency representations and spatial dependencies. FSEL (Sun et al., 2024) incorporates frequency-domain transformations to alleviate the sensitivity and locality limitations of spatial features in camouflaged object detection. In this paper, we jointly leverage frequency-domain and spatial features to better distinguish small objects from the background, dynamically suppressing irrelevant distractions and noise in feature representations. Figure 2. The overall architecture of DyFrDet consists of two main components: Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) and Label Disambiguation Module (LDM). DyFrFPN integrates a Dynamic Band Predictor (DBP), which adaptively predicts two channel-wise frequency thresholds to suppress irrelevant components. The masked features are then passed to the LDM to enhance the final localization performance. 2.3. Label Disambiguation Label ambiguity is often caused by human annotation bias or low-quality images. This issue becomes especially severe in small object detection, where the distinction between object and background is often unclear. Most existing works (Yao et al., 2021; Ye et al., 2022; Zhang et al., 2025; Huang et al., 2025) treat bounding box regression as a deterministic prediction problem. However, such approaches exhibit significant limitations when faced with imprecise or ambiguous object boundaries. Recently, modeling bounding boxes as probability distributions rather than fixed coordinates has attracted increasing attention across various domains, including visual tracking under uncertain or adverse conditions. For instance, UAST (Zhang et al., 2022) discretizes localization outputs into probability distributions, while UMDATrack (Yao et al., 2025b) and UncTrack (Yao et al., 2025a) introduce uncertainty-aware tracking frameworks to improve localization reliability under challenging conditions. This probabilistic formulation has demonstrated strong capabilities in resolving label ambiguity and enhancing localization robustness under complex or uncertain conditions. However, to the best of our knowledge, the distributional bounding box regression for small object detection has not been carefully explored. In this work, we demonstrate the potential of our approach to mitigate label ambiguity and better handle complex scenarios in small object detection(SOD). 3. Methodology In this section, we present the overall architecture of the proposed DyFrDet. As illustrated in Fig. 2, DyFrDet consists of a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) and a Label Disambiguation Module(LDM). The DyFrFPN transforms pyramid features into frequency domain and employs a Dynamic Band Predictor (DBP) to adaptively suppress irrelevant frequency components. The LDM enforces the detector to focus on the high confidence samples and restrain the ambiguous samples to facilitate small object localization. 3.1. Introduction to DyFrFPN Given an input image tI_t, we first pass it to a ResNet-50 backbone to extract multi-scale feature maps 2,3,4,5\c_2,c_3,c_4,c_5\, whose spatial resolutions are reduced by the factor of 4,8,16,32\4,8,16,32\ and are passed through the proposed DyFrFPN. Specifically, the backbone feature maps are processed by a standard FPN to obtain the top-down aggregated multi-scale features 2,3,4,5\p_2,p_3,p_4,p_5\. These features are then transformed into the frequency domain using the Fast Fourier Transform (FFT), resulting in frequency representations 2,3,4,5\P_2,P_3,P_4,P_5\. The FFT is computed as follows: (1) Pi(u,v)=∑x,ypi(x,y)⋅e−j2π(xHu+yWv),j2=−1.P_i(u,v)= _x,yp_i(x,y)· e^-j2π ( xHu+ yWv ), j^2=-1. where i∈2,3,4,5i∈\2,3,4,5\ and pi(x,y)p_i(x,y) denotes the spatial domain feature at location (x,y)(x,y) in level i, Pi(u,v)P_i(u,v) is its corresponding frequency representation at coordinate (u,v)(u,v). This transformation enables the decomposition of features into different frequency components: the low-frequency signals in the top-left region of the spectrum and the high-frequency ones in the bottom-right. Since the multi-scale features differ only in spatial resolution while sharing the same processing operations, we omit the scale index i in the following discussion for brevity. Dynamic Band Predictor. Instead of applying uniform suppression across all channels, we predict the dynamic band range for each channel to conduct channel-wise precise denoising. To achieve this, we send frequency-domain representation ∈ℝC×H×WP ^C× H× W and the corresponding spatial-domain feature ∈ℝC×H×Wp ^C× H× W into DBP to estimate the suppression thresholds [α1,α2]∈ℝC[ _1, _2] ^C. Particularly, the complex-valued frequency feature P is first decomposed into the amplitude spectrum A and phase spectrum Φ , which can be computed as follows: (2) (u,v) (u,v) =Re(P(u,v))2+Im(P(u,v))2, = Re(P(u,v))^2+Im(P(u,v))^2, Φ(u,v) (u,v) =arctan(Im(P(u,v))Re(P(u,v))), = ( Im(P(u,v))Re(P(u,v)) ), where Re(⋅)Re(·) and Im(⋅)Im(·) denote the real and imaginary parts of the complex-valued feature, respectively. After obtaining the amplitude spectrum A and phase spectrum Φ , we pass them through a convolutional block to aggregate the frequency representations. The aggregated components are then transformed back into the spatial domain via the Inverse Fast Fourier Transform (IFFT). The complete process is formulated as: (3) ′(u,v) (u,v) =ConvBlock((u,v)), =ConvBlock(A(u,v)), Φ′(u,v) (u,v) =ConvBlock(Φ(u,v)), =ConvBlock( (u,v)), =ℱ−1 =F^-1 [′(u,v)⋅exp(jΦ′(u,v))], [A (u,v)· (j (u,v) )], where ′∈ℝC×H×WA ^C× H× W and Φ′∈ℝC×H×W ^C× H× W denote the amplitude and phase components, respectively. Afterwards, we compress the spatial information of the feature maps using global pooling operations. We apply Global Max Pooling (GMP) and Global Average Pooling (GAP) to obtain compact feature representations, which can be given by: (4) ′ =Concat(GMP(),GAP()), =Concat(GMP(f),\ GAP(f)), ′ =Concat(GMP(),GAP()), =Concat(GMP(p),\ GAP(p)), where ′∈ℝ2C×H′W′f ^2C× H W and ′∈ℝ2C×H′W′p ^2C× H W , we further project them into query, key, and value embeddings for attention computation as follows: (5) Q,K,V=[Wq,Wk,Wv]T⋅[′,′,′],Q,\ K,\ V=[W_q,W_k,W_v]^T·[p ,\ f ,\ f ], where WqW_q, WkW_k, Wv∈ℝ2C×dW_v ^2C× d are learnable linear projection matrices. Q,K,V∈ℝd×H′W′Q,K,V ^d× H W is the query, key and value embeddings. To adaptively suppress the irrelevant noises, we predict the suppression thresholds α1,α2∈RC _1, _2∈R^C through an attention mechanism followed by a feed-forward network (FFN). To ensure positivity and enable flexible scaling, we adopt an exponential parameterization as follows: (6) dα dα =FFN(′+softmax(QK⊤d)V), =FFN (p +softmax ( QK d )V ), α1=αl⋅edα1,α2=αh⋅edα2, _1= _l· e^d _1, _2= _h· e^d _2, where α1,α2∈ℝC _1, _2 ^C denote the predicted lower and upper suppression thresholds, respectively. αl,αh∈ℝ _l, _h denote the predefined suppression thresholds. Based on the predicted α1 _1 and α2 _2, we define a spatial mask as follows: (7) ℳ(x,y,j)=0,if x<α1j⋅W and y<α1j⋅H0,if x>α2j⋅W and y>α2j⋅H1,otherwise,M(x,y,j)= cases0,&if x< _1^j· W and y< _1^j· H\\ 0,&if x> _2^j· W and y> _2^j· H\\ 1,&otherwise, cases where α1j _1^j and α2j _2^j represent the lower and upper suppression thresholds of the j-th channel. The pixels within the top-left region (i.e., low-frequency redundant areas) and the bottom-right region (i.e., high-frequency noise areas) are masked out (set to 0), while the remaining spatial locations are retained (set to 1). By applying the predicted ℳM to the frequency feature map P, we obtain the noise suppressed feature ′P . Which can be given by: (8) ′=−β⋅(⊙(1−ℳ)),P =P-β· (P (1-M) ), where ⊙ denotes element-wise multiplication, and β is a hyperparameter that controls the suppression rate. Finally, we convert the suppressed frequency-domain features back to the spatial domain using the Inverse Fast Fourier Transform (IFFT), yielding the enhanced multi-scale features ∗,∗,∗,∗\p^*_2,p^*_3,p^*_4,p^*_5\. 3.2. Label Disambiguation Module To precisely localize targets and mitigate the negative impact of label ambiguities, we introduce a Label Disambiguation Module (LDM). Let P=(Px,Py,Pw,Ph)P=(P_x,P_y,P_w,P_h) denote the proposal box, where (Px,Py)(P_x,P_y) represents the center coordinates and (Pw,Ph)(P_w,P_h) denotes the size of the proposal, respectively. Similarly, G=(Gx,Gy,Gw,Gh)G=(G_x,G_y,G_w,G_h) represent the ground-truth bounding box. The bounding box regression process can be formulated as: (9) tx t_x =(Gx−Px)/Pw, =(G_x-P_x)/P_w, G^x G_x =Pwdx+Px, =P_wd_x+P_x, ty t_y =(Gy−Py)/Ph, =(G_y-P_y)/P_h, G^y G_y =Phdy+Py, =P_hd_y+P_y, tw t_w =log(Gw/Pw), = (G_w/P_w), G^w G_w =Pwexp(dw), =P_w (d_w), th t_h =log(Gh/Ph), = (G_h/P_h), G^h G_h =Phexp(dh), =P_h (d_h), where =dx,dy,dw,dhD=\d_x,d_y,d_w,d_h\ denotes the predicted offsets used to transform the proposal P to the estimated bounding box G G. =tx,ty,tw,thT=\t_x,t_y,t_w,t_h\ represents the targets offsets corresponding to the ground truth box G and anchor P. In the following discussion, we denote D and T by x and gtx_gt for simplicity. Different from the traditional deterministic coordinate regression methods (Ren et al., 2016; Cai and Vasconcelos, 2018; Yao et al., 2025b) that minimize L1 loss between the predicted offsets x and the ground truth offset gtx_gt, the proposed DyFrDet models them as a pair of Gaussian distribution and Dirac delta distribution, represented as: (10) ()=(,),gt()=δ(−gt), _ θ(x)=N( μ_ θ, σ), _gt(x)=δ(x-x_gt), where ()P_ θ(x) denotes the predicted localization distribution, and gt()Q_gt(x) represents the ground-truth distribution. Here μ is the mean of the Gaussian distribution, and σ is the predicted covariance that determines the quality of the localization. Then the label quality can be reflected by the maximum value of σ, which is denoted as σm _m. As shown in Fig. 3, the object boundaries become increasingly blurred as σm _m increases from top to bottom. Figure 3. Visualization of different small objects under varying σm _m values. The green boxes indicate ground-truth annotations, and blue boxes denote predicted bounding boxes. From top to bottom, the value of σm _m gradually increases. The full names of class abbreviations are as follows: t-light (traffic-light), t-camera (traffic-camera), w-cone (warning cone). During training, we minimize the Kullback–Leibler (KL) divergence between the predicted distribution ()P_ θ(x) and the ground-truth distribution gt()Q_gt(x). The loss function is defined as: (11) ℒLDM _LDM =KL(gt()∥()) =KL(Q_gt(x) _ θ(x)) ∝−∼gt[log()] -E_x _gt [ _ θ(x) ] ∝(gt−)(gt−)⊤(⊤)+log(⊤), (x_gt- μ)(x_gt- μ) ( σ σ)+ ( σ σ), To mitigate the negative impact of these samples with large label ambiguities, we design a weighted function as follows: (12) ω(σm)=1if σm≤ρε+(1−ε)[1−(σm−ρ1−ρ)2]if ρ<σm<1,ω( _m)= \ array[]l1&if _m≤ρ\\ +(1- ) [1- ( _m-ρ1-ρ )^2 ]&if ρ< _m<1, array . where ε and ρ are hyperparameters to control the penalty range. When σm≤ρ _m≤ρ, ω(σm)ω( _m) remains 1, while when σm→1 _m→ 1, ω(σm)ω( _m) becomes closer to ε . By reweighting the loss term, we are able to reduce the detrimental effects of label ambiguity in training. Finally, the overall loss function is defined as: (13) ℒ=γℒIoU+ℒCLS+ω(σm)∗ℒLDM,L= _IoU+L_CLS+ω( _m)*L_LDM, where ℒIoUL_IoU denotes the IoU loss. ℒCLSL_CLS is the cross-entropy classification loss for candidate proposals. The coefficient γ is a hyperparameter that balances the loss contribution. 4. Experiment In this section, we first provide the implementation details of our method and comparisons with other state-of-the-art detectors on two popular benchmarks: AI-TOD (Wang et al., 2021) and SODA (Cheng et al., 2023). SODA includes two subsets, SODA-D and SODA-A. Next, we conduct ablation studies to evaluate the effectiveness of our approach. Finally, we present qualitative visualizations to analyze the impact on feature representations. Implementation Details All experiments are conducted on a single NVIDIA RTX 3090 GPU. For AI-TOD, the patch size is fixed to 800×800800× 800, while for SODA-D and SODA-A, the patches are resized to 1200×12001200× 1200. The hyperparameters αl _l and αh _h are set to 0.05 and 0.95. The suppression rate β is set to 0.5. The weight function parameters ε and ρ are set to 0.5, 0.8, respectively. The loss weights γ are set to 0.9. The model is trained for 36 epochs on AI-TOD and for 12 epochs on SODA-D and SODA-A. To ensure that the suppressed frequency components correspond to meaningful distractors, the Dynamic Frequency Suppression Strategy is activated from the 24th epoch on AI-TOD and the 8th epoch on both SODA-D and SODA-A. 4.1. Comparison with state-of-the-art AI-TOD. The AI-TOD dataset contains 700,621 instances in 8 categories in 28,036 aerial images. Compared to existing object detection datasets in aerial images, the mean size of objects in AI-TOD is about 12.8 pixels, which is much smaller than others. We report the results of our proposed DyFrDet on AI-TOD test set. As shown in Table 1, DyFrDet achieves the best performance across all metrics. Compared with HS-FPN (Shi et al., 2025), it improves (AP, AP50, AP75, APvt_vt, APt_t, APs_s, APm_m) by (3.6%, 1.5%, 4.6%, 3.1%, 3.8%, 3.4%, 2.7%), respectively. Table 1. Comparison with state-of-the-art detection methods on AI-TOD test set. All metrics follow the COCO-style evaluation protocol, including overall AP and performance across different scene types. ∗ indicates methods using ResNet-50 as the backbone. Bold numbers indicate the best results. Method Source Backbone AP AP50 AP75 APvt_vt APt_t APs_s APm_m Query-based Detectors DAB-DETR(Liu et al., 2022) ICLR2022 R50 4.9 16.0 1.7 1.7 3.6 7.0 18.0 DAB-Deformable-DETR(Liu et al., 2022) ICLR2022 R50 16.5 42.6 9.9 7.9 15.2 23.8 31.9 DINO-Defomable-DETR(Zhang et al., 2023) ICLR2023 R50 23.2 56.6 15.4 9.9 23.1 29.3 37.6 DINO-5scale w/SET (Sun et al., 2025) CVPR2025 R50 26.6 57.1 20.8 13.2 27.1 31.5 – Multi-Stage Detectors Faster R-CNN (Ren et al., 2016) TPAMI2017 R50+FPN 11.1 26.3 7.6 0.0 7.2 23.3 33.6 Cascade R-CNN (Cai and Vasconcelos, 2018) CVPR2018 R50+FPN 13.8 30.8 10.5 0.0 10.5 25.5 36.6 DetectorRS(Qiao et al., 2021) CVPR2021 R50+FPN 14.8 32.8 11.4 0.0 10.8 28.3 38.0 QueryDet(Yang et al., 2022) CVPR2022 R50+FPN 12.2 29.3 7.9 2.4 10.5 18.5 26.3 CFINet(Yuan et al., 2023) ICCV2023 R50+FPN 24.7 53.9 18.6 11.7 26.4 28.1 32.2 KLDet(Zhou and Zhu, 2024) TGRS2024 R50+FPN 19.6 46.4 13.7 8.4 20.6 22.7 26.4 RFLA(Xu et al., 2022b) ECCV2022 R50 w/SAC+FPN 24.8 55.2 18.5 9.3 24.8 30.3 38.2 DNTR(Liu et al., 2024) TGRS2024 R50 w/SAC+FPN 26.2 56.7 20.2 12.8 26.4 31.0 37.0 SimD(Shi et al., 2024) IROS2024 R50 w/SAC+FPN 26.6 55.9 21.2 13.4 27.5 30.9 37.8 DetectorRS w/FIDP (Bian et al., 2025) CVPR2025 R50 w/SAC+FPN 24.3 54.4 18.3 8.5 24.9 29.8 – HS-FPN(Shi et al., 2025) AAAI2025 R50 w/SAC+FPN 25.1 55.7 19.1 12.1 25.3 29.9 36.9 DyFrDet∗ – R50+DyFrFPN 26.0 55.3 19.6 13.6 27.3 29.0 34.5 DyFrDet – R50 w/SAC+DyFrFPN 28.7 57.2 23.7 15.2 29.1 33.3 39.6 Table 2. Ablation study on AI-TOD test set. When DyFrFPN is disabled, it is replaced with a standard FPN. When LDM is removed, an FFN with L1 loss is used instead. DyFrFPN LDM AP APvt_vt APt_t APs_s APm_m × × 23.6 11.9 24.8 27.0 30.7 ✓ × 24.9 12.3 26.2 28.7 33.3 × ✓ 24.4 12.0 25.6 27.3 32.5 ✓ ✓ 26.0 13.6 27.3 29.0 34.5 Table 3. Comparison of the static and dynamic frequency suppression strategies. “Static” denotes suppressing frequency components outside the fixed thresholds [α1 _1, α2 _2] during the FPN stage. “Dynamic” refers to the proposed DyFrFPN, which dynamically predicts α1 _1 and α2 _2 to suppress noise. All experiments are conducted on AI-TOD test set. Strategy α1 _1 α2 _2 AP APvt_vt APt_t APs_s APm_m None – – 24.4 12.0 25.6 27.3 32.5 Static 0.05 1.00 24.9 12.1 26.4 28.3 32.1 Static 0.05 0.95 25.3 13.1 26.7 28.8 32.8 Static 0.10 0.90 24.8 12.3 26.1 27.7 33.3 Static 0.15 0.85 24.7 11.8 26.0 28.0 32.5 Static 0.20 0.80 25.1 11.9 26.6 28.5 32.4 Static 0.25 0.75 24.7 12.8 25.8 28.8 32.6 Static 0.30 0.70 25.0 11.5 26.2 28.3 32.3 Dynamic – – 26.0 13.6 27.3 29.0 34.5 SODA-A. The SODA-A dataset consists of 2,513 aerial images with 872,069 oriented bounding box annotations, focusing on small object detection. SODA-A has an average resolution of 4761 × 2777 pixels and contains approximately 347 instances per image, presenting a highly dense object distribution that challenges existing detection models in clustered scenes. Table 4. Comparison with oriented object detection methods on SODA-A test set. Metrics follow COCO-style evaluation. ∗ indicates methods using ResNet-50 as the backbone. Bold numbers indicate the best results. Method Source Backbone AP AP50 AP75 APes_es APrs_rs APgs_gs APN One-stage Detectors Rotated RetinaNet(Lin et al., 2017b) ICCV2017 R50 + FPN 26.8 63.4 16.2 9.1 22.0 35.4 28.2 Oriented RepPoints(Li et al., 2022) CVPR2022 R50 + FPN 26.3 58.8 19.0 9.4 22.6 32.4 28.5 DHRec(Nie and Huang, 2022) TPAMI2022 R50 + FPN 30.1 68.8 19.8 10.6 24.6 40.3 34.6 CFPT(Du et al., 2025) TGRS2025 R50+CFPT 25.9 63.3 14.3 9.0 21.2 34.6 27.9 LEGNet(Lu et al., 2025) ICCVW2025 LWGNet + FPN 29.6 58.7 26.4 10.4 26.0 39.6 32.0 Two-stage Detectors Rotated Faster R-CNN(Ren et al., 2016) TPAMI2017 R50 + FPN 32.5 70.1 24.3 11.9 27.3 42.2 34.4 Gliding Vertex(Xu et al., 2020) TPAMI2020 R50 + FPN 31.7 70.8 22.6 11.7 27.0 41.1 33.8 Oriented R-CNN(Xie et al., 2021) ICCV2021 R50 + FPN 34.4 70.7 28.6 12.5 28.6 44.5 36.7 DODet(Cheng et al., 2022) TGRS2022 R50 + FPN 31.6 68.1 23.4 11.3 26.3 41.0 33.5 CFINet(Yuan et al., 2023) ICCV2023 R50 + FPN 34.4 73.1 26.1 13.5 29.3 44.0 35.9 DecoupleNet(Lu et al., 2024) TGRS2024 DecoupleNet + FPN 36.6 71.3 33.3 12.2 31.0 47.7 40.2 STD(Yu et al., 2024) AAAI2024 R50+FPN 30.3 67.4 21.2 12.4 26.6 36.7 30.7 SR-TOD(Cao et al., 2024) ECCV2024 R50+FPN 32.7 70.5 24.1 12.2 27.5 42.6 34.5 GauCho(Marques et al., 2025) CVPR2025 R50 + FPN 33.2 70.1 25.0 9.9 27.8 44.9 36.4 Unc-SOD(Yuan et al., 2026) TIP2026 R50+FPN 34.8 73.6 26.4 13.8 29.7 44.7 36.5 DyFrDet∗ – R50 + DyFrFPN 36.0 73.3 30.1 13.8 30.6 46.0 38.0 DyFrDet – DecoupleNet + DyFrFPN 37.8 73.4 34.3 12.9 31.9 49.5 41.0 Table 4 summarizes the performance on SODA-A benchmark. DyFrDet achieves state-of-the-art results with (37.8%, 73.4%, 34.3%, 31.9%, 49.5%, 41.0%) on (AP, AP50, AP75, APrs_rs, APgs_gs, APN_N), surpassing state-of-the-art GauCho (Marques et al., 2025) by (4.6%, 3.3%, 9.3%, 3.0%, 4.1%, 4.6%, 4.6%) across all metrics. Table 5. Ablation study of different compositional inputs of DBP. “F” denotes the method only using the frequency feature, “SF” only uses the spatial feature, and “F+SF” uses both. Results are reported in terms of Average Precision (AP) on various object sizes. Strategy AP APvt_vt APt_t APs_s APm_m F 24.8 11.9 26.2 28.2 32.7 SF 24.6 12.2 26.0 28.2 32.7 F+SF 26.0 13.6 27.3 29.0 34.5 Table 6. Ablation study on different suppression rates β in Eq. 8. β=0.0β=0.0 indicates no suppression, while β=1.0β=1.0 denotes direct filtering. β AP APvt_vt APt_t APs_s APm_m 0.00 24.4 12.0 25.6 27.3 32.5 0.25 25.0 12.5 26.1 29.0 32.9 0.50 26.0 13.6 27.3 29.0 34.5 0.75 24.6 11.9 25.9 27.5 32.3 1.00 25.3 12.7 26.5 28.6 33.2 SODA-D. The SODA-D dataset contains 24,828 high-resolution images with 278,433 annotated instances spanning 9 categories: people, rider, bicycle, motor, vehicle, traffic sign, traffic light, traffic camera, and warning cone. It exhibits rich diversity in terms of locations, weather conditions, period, camera viewpoints, and traffic scenarios. With an average image resolution of 3407 × 2470, the dataset is particularly well-suited for detecting small and tiny objects in complex environments. We report the performance of DyFrDet on SODA-D test set, as summarized in Table 7. Our DyFrDet achieves state-of-the-art performance across all metrics, with an average AP of 31.3%. Compared with HS-FPN(Shi et al., 2025), DyFrDet yields consistent improvements of (1.7%, 5.3%, 0.1%, 1.5%, 1.4%, 2.0%, 0.9%) in terms of (AP, AP50, AP75, APes_es, APrs_rs, APgs_gs, APN_N). Table 7. Comparison with state-of-the-art detection approaches on SODA-D test set. All metrics follow COCO-style evaluation, including overall AP and performance across different scene types. Bold numbers indicate the best results. Method Source Backbone AP AP50 AP75 APes_es APrs_rs APgs_gs APN One-stage Detectors RetinaNet(Lin et al., 2017b) ICCV2017 R50+FPN 28.2 57.6 23.7 11.9 25.2 34.1 44.2 FCOS(Tian et al., 2019) ICCV2019 R50+FPN 23.9 49.5 19.9 6.9 19.4 30.9 40.9 ATSS(Zhang et al., 2020) CVPR2020 R50+FPN 26.8 55.6 22.1 11.7 23.9 32.2 41.3 DyHead(Dai et al., 2021) CVPR2021 R50+FPN 27.5 56.1 23.2 12.4 24.4 33.0 41.9 KLDet(Zhou and Zhu, 2024) TGRS2024 R50+FPN 25.9 53.8 21.4 10.7 22.2 31.9 41.6 CFPT(Du et al., 2025) TGRS2025 R50+CFPT 27.5 54.5 23.8 7.1 22.4 36.2 45.9 Two-stage Detectors Faster R-CNN(Ren et al., 2016) TPAMI2017 R50+FPN 28.9 59.7 24.2 13.9 25.6 34.3 43.2 Cascade RPN(Vu et al., 2019) NIPS2019 R50+FPN 29.1 56.5 25.9 12.5 25.5 35.4 44.7 KL(He et al., 2019) CVPR2019 R50+FPN 29.4 59.2 24.7 14.0 26.3 36.0 44.2 RFLA(Xu et al., 2022b) ECCV2022 R50+FPN 29.7 60.2 25.2 13.2 26.9 35.4 44.6 CFINet(Yuan et al., 2023) ICCV2023 R50+FPN 30.7 60.8 26.7 14.7 27.8 36.4 44.6 DNTR(Liu et al., 2024) TGRS2024 R50 w/SAC+RFP 29.6 57.8 26.5 13.1 26.7 35.5 43.4 SR-TOD(Cao et al., 2024) ECCV2024 R50+FPN 29.3 60.0 24.5 13.8 26.0 35.6 43.4 HS-FPN(Shi et al., 2025) AAAI2025 R50 + FPN 29.6 56.8 26.7 13.6 26.4 35.3 45.3 DyFrDet∗ – R50+DyFrFPN 31.3 62.1 26.8 15.1 27.8 37.3 46.2 4.2. Ablation Study Ablation of Different Variations. To evaluate the effectiveness of each component, we selectively ablate DyFrFPN and LDM from DyFrDet. When DyFrFPN is removed, it is replaced with a standard FPN. Similarly, when LDM is removed, it is replaced with an FFN head trained using the L1 loss. As shown in Table 2, the baseline detector without DyFrFPN and LDM achieves (23.6%, 11.9%, 24.8%, 27.0%, 30.7%) in terms of (AP, APvt_vt, APt_t, APs_s, APm_m), respectively. When DyFrFPN is introduced, the performance improves to (24.9%, 12.3%, 26.2%, 28.7%, 33.3%), demonstrating its ability to effectively suppress low-frequency redundancies and high-frequency noises. Additionally, incorporating LDM into the baseline also leads to performance gains, achieving (24.4%, 12.0%, 25.6%, 27.3%, 32.5%) for (AP, APvt_vt, APt_t, APs_s, APm_m), respectively. Furthermore, combining both DyFrFPN and LDM yields the best performance, with improvements of (2.4%, 1.7%, 2.5%, 2.0%, 3.8%) over the baseline. These results validate the effectiveness of the proposed DyFrFPN and LDM modules. Ablation of Different Suppression Strategies. To verify the effectiveness of the proposed Dynamic Frequency Suppression Strategy, we analyze the performance gains under different frequency suppression strategies. In our experiments, “None” denotes that no frequency components are suppressed. “Static” with various settings of α1 _1 and α2 _2 refers to the strategy that statically suppresses frequency components outside the range [α1,α2][ _1, _2]. As shown in Table 3, when following the setting of HS-FPN (Shi et al., 2025), where only low-frequency components are suppressed by setting α1=0.05 _1=0.05 and α2=1.00 _2=1.00, the performance improves (0.5%, 0.1%, 0.8%, 1.0%) in terms of (AP, APvt_vt, APt_t, APs_s), respectively. When high-frequency suppression is also introduced (e.g., setting α1=0.05 _1=0.05 and α2=0.95 _2=0.95), the performance further improves (0.4%, 1.0%, 0.3%, 0.5%, 0.7%) over the previous setting in terms of (AP, APvt_vt, APt_t, APs_s, APm_m). These experiments validate the importance of suppressing high-frequency noise. Moreover, the dynamic strategy yields the best overall performance, reaching (26.0%, 13.6%, 27.3%, 29.0%, 34.5%) on (AP, APvt_vt, APt_t, APs_s, APm_m), which demonstrates the effectiveness of our proposed suppression strategy. Band Predictor Settings. We perform an ablation study to analyze the impact of different input feature compositions for the Dynamic Band Predictor (DBP), as reported in Table 5. Specifically, we investigate three variants that utilize frequency features (F) alone, spatial features (SF) alone, and the combination of both (F+SF). When using either individual branch, F and SF achieve comparable performance, obtaining overall AP scores of 24.8 and 24.6, respectively, which indicates that both frequency-domain representations and spatial cues provide valuable information for band prediction. By jointly incorporating F and SF, DBP achieves consistent improvements across all object scales and obtains the highest overall AP of 26.0, validating the necessity of each branch and the complementary roles of frequency and spatial information. Figure 4. Visualization of features map at the P2 level on SODA-D test set. The first row shows results from the baseline, while the second row presents results from DyFrDet. The green, blue, and red boxes denote true positives (TP), false positives (FP), and false negatives (FN), respectively. Please zoom in for details. Suppression Rate. We conduct an ablation study on different suppression rates, as shown in Table 6. When β=0.00β=0.00, no suppression is applied. When β=1.00β=1.00, it corresponds to full filtering. The best performance is achieved when β=0.50β=0.50, indicating that moderate suppression yields an optimal trade-off between information retention and noise reduction. No suppression allows distractors to interfere with the target, while full suppression may mistakenly remove useful target features. 4.3. Qualitative Visualization To qualitatively assess the effectiveness of our method, we visualize the feature heatmaps at the P2 level on SODA-D dataset, as shown in Fig. 4. The heatmaps produced by CFINet exhibit cluttered and dispersed activations, which often lead to false positives and missed detections. In contrast, DyFrDet generates more concentrated activations on foreground regions, with suppressed responses in the background, resulting in more reliable detection performance, demonstrating its stronger ability to distinguish small objects from background clutter. 5. Conclusion In this paper, we propose DyFrDet, a two-stage detector composed of a Dynamic Frequency-aware Feature Pyramid Network (DyFrFPN) and a Label Disambiguation Module (LDM). The former leverages Dynamic Band Predictor (DBP) to adaptively control the frequency representation, allowing the detector to suppress both low-frequency redundancies and high-frequency noises. The latter achieves accurate bounding box estimation by paying more attention to the high confidence samples and restraining the ambiguous samples. Experimental results demonstrate that our method achieves state-of-the-art performance on several widely used small object detection benchmarks. 6. Acknowledgement This work was supported by Project No. 25KYHX131B: Construction of Physical Neural Networks for Cognitive Countermeasures. References (1) Bian et al. (2025) Jinghao Bian, Mingtao Feng, Weisheng Dong, Fangfang Wu, Jianqiao Luo, Yaonan Wang, and Guangming Shi. 2025. Feature Information Driven Position Gaussian Distribution Estimation for Tiny Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition. 30376–30386. Cai and Vasconcelos (2018) Zhaowei Cai and Nuno Vasconcelos. 2018. Cascade r-cnn: Delving into high quality object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 6154–6162. Cao et al. (2024) Bing Cao, Haiyu Yao, Pengfei Zhu, and Qinghua Hu. 2024. Visible and clear: Finding tiny objects in difference map. In European Conference on Computer Vision. 1–18. Carion et al. (2020) Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-end object detection with transformers. In European Conference on Computer Vision. 213–229. Chen et al. (2019) Changrui Chen, Yu Zhang, Qingxuan Lv, Shuo Wei, Xiaorui Wang, Xin Sun, and Junyu Dong. 2019. Rrnet: A hybrid detector for object detection in drone-captured images. In Proceedings of the IEEE/CVF international conference on computer vision workshops. 0–0. Chen et al. (2025) Linwei Chen, Lin Gu, and Ying Fu. 2025. Frequency-dynamic attention modulation for dense prediction. In IEEE International Conference on Computer Vision. 22620–22632. Cheng et al. (2022) Gong Cheng, Yanqing Yao, Shengyang Li, Ke Li, Xingxing Xie, Jiabao Wang, Xiwen Yao, and Junwei Han. 2022. Dual-aligned oriented detector. IEEE Transactions on Geoscience and Remote Sensing 60 (2022), 1–11. Cheng et al. (2023) Gong Cheng, Xiang Yuan, Xiwen Yao, Kebing Yan, Qinghua Zeng, Xingxing Xie, and Junwei Han. 2023. Towards large-scale small object detection: Survey and benchmarks. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 11 (2023), 13467–13488. Chi et al. (2020) Lu Chi, Borui Jiang, and Yadong Mu. 2020. Fast fourier convolution. Advances in Neural Information Processing Systems 33 (2020), 4479–4488. Dai et al. (2021) Xiyang Dai, Yinpeng Chen, Bin Xiao, Dongdong Chen, Mengchen Liu, Lu Yuan, and Lei Zhang. 2021. Dynamic head: Unifying object detection heads with attentions. In IEEE Conference on Computer Vision and Pattern Recognition. 7373–7382. Deng et al. (2009) Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In IEEE Conference on Computer Vision and Pattern Recognition. 248–255. Du et al. (2025) Zewen Du, Zhenjiang Hu, Guiyu Zhao, Ying Jin, and Hongbin Ma. 2025. Cross-Layer Feature Pyramid Transformer for Small Object Detection in Aerial Images. IEEE Transactions on Geoscience and Remote Sensing 63 (2025), 1–14. Ghiasi et al. (2019) Golnaz Ghiasi, Tsung-Yi Lin, and Quoc V Le. 2019. Nas-fpn: Learning scalable feature pyramid architecture for object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 7036–7045. Girshick et al. (2014) Ross Girshick, Jeff Donahue, Trevor Darrell, and Jitendra Malik. 2014. Rich feature hierarchies for accurate object detection and semantic segmentation. In IEEE Conference on Computer Vision and Pattern Recognition. 580–587. He et al. (2019) Yihui He, Chenchen Zhu, Jianren Wang, Marios Savvides, and Xiangyu Zhang. 2019. Bounding box regression with uncertainty for accurate object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 2888–2897. Huang et al. (2025) Shihua Huang, Zhichao Lu, Xiaodong Cun, Yongjun Yu, Xiao Zhou, and Xi Shen. 2025. DEIM: DETR with Improved Matching for Fast Convergence. 15162-15171 pages. Kim et al. (2021) Jung Uk Kim, Sungjune Park, and Yong Man Ro. 2021. Robust small-scale pedestrian detection with cued recall via memory learning. In IEEE Conference on Computer Vision and Pattern Recognition. 3050–3059. Kisantal et al. (2019) Mate Kisantal, Zbigniew Wojna, Jakub Murawski, Jacek Naruniec, and Kyunghyun Cho. 2019. Augmentation for small object detection. arXiv preprint arXiv:1902.07296 (2019). Li et al. (2022) Wentong Li, Yijie Chen, Kaixuan Hu, and Jianke Zhu. 2022. Oriented reppoints for aerial object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 1829–1838. Li et al. (2019) Yanghao Li, Yuntao Chen, Naiyan Wang, and Zhaoxiang Zhang. 2019. Scale-aware trident networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 6054–6063. Lin et al. (2017a) Tsung-Yi Lin, Piotr Dollár, Ross Girshick, Kaiming He, Bharath Hariharan, and Serge Belongie. 2017a. Feature pyramid networks for object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 2117–2125. Lin et al. (2017b) Tsung-Yi Lin, Priya Goyal, Ross Girshick, Kaiming He, and Piotr Dollár. 2017b. Focal loss for dense object detection. In IEEE International Conference on Computer Vision. 2980–2988. Lin et al. (2014) Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014. Microsoft coco: Common objects in context. In European Conference on Computer Vision. 740–755. Liu et al. (2024) Hou-I Liu, Yu-Wen Tseng, Kai-Cheng Chang, Pin-Jyun Wang, Hong-Han Shuai, and Wen-Huang Cheng. 2024. A DeNoising FPN With Transformer R-CNN for Tiny Object Detection. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–15. Liu et al. (2022) Shilong Liu, Feng Li, Hao Zhang, Xiao Yang, Xianbiao Qi, Hang Su, Jun Zhu, and Lei Zhang. 2022. DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR. In International Conference on Learning Representations. Liu et al. (2016) Wei Liu, Dragomir Anguelov, Dumitru Erhan, Christian Szegedy, Scott Reed, Cheng-Yang Fu, and Alexander C Berg. 2016. Ssd: Single shot multibox detector. In European Conference on Computer Vision. 21–37. Lu et al. (2025) Wei Lu, Si-Bao Chen, Hui-Dong Li, Qing-Ling Shu, Chris HQ Ding, Jin Tang, and Bin Luo. 2025. Legnet: Lightweight edge-Gaussian driven network for low-quality remote sensing image object detection. arXiv preprint arXiv:2503.14012 (2025). Lu et al. (2024) Wei Lu, Si-Bao Chen, Qing-Ling Shu, Jin Tang, and Bin Luo. 2024. DecoupleNet: A Lightweight Backbone Network with Efficient Feature Decoupling for Remote Sensing Visual Tasks. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–13. Marques et al. (2025) José Henrique Lima Marques, Jeffri Murrugarra-Llerena, and Claudio R. Jung. 2025. GauCho: Gaussian Distributions with Cholesky Decomposition for Oriented Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition. 3593–3602. Nie and Huang (2022) Guangtao Nie and Hua Huang. 2022. Multi-oriented object detection in aerial images with double horizontal rectangles. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 4 (2022), 4932–4944. Patro et al. (2025) Badri N Patro, Vinay P Namboodiri, and Vijay S Agneeswaran. 2025. Spectformer: Frequency and attention is what you need in a vision transformer. In Proceedings of the IEEE/CVF winter conference on applications of computer vision. 9543–9554. Qiao et al. (2021) Siyuan Qiao, Liang-Chieh Chen, and Alan Yuille. 2021. Detectors: Detecting objects with recursive feature pyramid and switchable atrous convolution. In IEEE Conference on Computer Vision and Pattern Recognition. 10213–10224. Qin et al. (2021) Zequn Qin, Pengyi Zhang, Fei Wu, and Xi Li. 2021. Fcanet: Frequency channel attention networks. In IEEE International Conference on Computer Vision. 783–792. Ren et al. (2016) Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2016. Faster R-CNN: Towards real-time object detection with region proposal networks. IEEE Transactions on Pattern Analysis and Machine Intelligence 39, 6 (2016), 1137–1149. Shi et al. (2024) Shuohao Shi, Qiang Fang, Xin Xu, and Tong Zhao. 2024. Similarity distance-based label assignment for tiny object detection. In IEEE/RSJ International Conference on Intelligent Robots and Systems. 13711–13718. Shi et al. (2025) Zican Shi, Jing Hu, Jie Ren, Hengkang Ye, Xuyang Yuan, Yan Ouyang, Jia He, Bo Ji, and Junyu Guo. 2025. HS-FPN: High frequency and spatial perception FPN for tiny object detection. In AAAI Conference on Artificial Intelligence. 6896–6904. Sun et al. (2025) Huixin Sun, Runqi Wang, Yanjing Li, Linlin Yang, Shaohui Lin, Xianbin Cao, and Baochang Zhang. 2025. SET: Spectral Enhancement for Tiny Object Detection. In IEEE Conference on Computer Vision and Pattern Recognition. 4713–4723. Sun et al. (2024) Yanguang Sun, Chunyan Xu, Jian Yang, Hanyu Xuan, and Lei Luo. 2024. Frequency-spatial entanglement learning for camouflaged object detection. In European Conference on Computer Vision. 343–360. Tan et al. (2020) Mingxing Tan, Ruoming Pang, and Quoc V Le. 2020. Efficientdet: Scalable and efficient object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 10781–10790. Tian et al. (2019) Zhi Tian, Chunhua Shen, Hao Chen, and Tong He. 2019. Fcos: Fully convolutional one-stage object detection. In IEEE International Conference on Computer Vision. 9627–9636. Vu et al. (2019) Thang Vu, Hyunjun Jang, Trung X Pham, and Chang Yoo. 2019. Cascade RPN: Delving into high-quality region proposal network with adaptive convolution. Advances in Neural Information Processing Systems 32 (2019). Wang et al. (2021) Jinwang Wang, Wen Yang, Haowen Guo, Ruixiang Zhang, and Gui-Song Xia. 2021. Tiny Object Detection in Aerial Images. In IEEE International Conference on Pattern Recognition. 3791–3798. Wu et al. (2020) Jialian Wu, Chunluan Zhou, Qian Zhang, Ming Yang, and Junsong Yuan. 2020. Self-mimic learning for small-scale pedestrian detection. In Proceedings of the ACM International Conference on Multimedia. 2012–2020. Xie et al. (2021) Xingxing Xie, Gong Cheng, Jiabao Wang, Xiwen Yao, and Junwei Han. 2021. Oriented R-CNN for object detection. In IEEE International Conference on Computer Vision. 3520–3529. Xu et al. (2022a) Chang Xu, Jinwang Wang, Wen Yang, Huai Yu, Lei Yu, and Gui-Song Xia. 2022a. Detecting tiny objects in aerial images: A normalized Wasserstein distance and a new benchmark. ISPRS Journal of Photogrammetry and Remote Sensing 190 (2022), 79–93. Xu et al. (2022b) Chang Xu, Jinwang Wang, Wen Yang, Huai Yu, Lei Yu, and Gui-Song Xia. 2022b. RFLA: Gaussian receptive field based label assignment for tiny object detection. In European Conference on Computer Vision. 526–543. Xu et al. (2021) Chang Xu, Jinwang Wang, Wen Yang, and Lei Yu. 2021. Dot distance for tiny object detection in aerial images. In IEEE Conference on Computer Vision and Pattern Recognition. 1192–1201. Xu et al. (2020) Yongchao Xu, Mingtao Fu, Qimeng Wang, Yukang Wang, Kai Chen, Gui-Song Xia, and Xiang Bai. 2020. Gliding vertex on the horizontal bounding box for multi-oriented object detection. IEEE Transactions on Pattern Analysis and Machine Intelligence 43, 4 (2020), 1452–1459. Yang et al. (2022) Chenhongyi Yang, Zehao Huang, and Naiyan Wang. 2022. QueryDet: Cascaded sparse query for accelerating high-resolution small object detection. In IEEE Conference on Computer Vision and Pattern Recognition. 13668–13677. Yao et al. (2025a) Siyuan Yao, Yang Guo, Yanyang Yan, Wenqi Ren, and Xiaochun Cao. 2025a. UncTrack: Reliable Visual Object Tracking With Uncertainty-Aware Prototype Memory Network. IEEE Transactions on Image Processing 34 (2025), 3533–3546. Yao et al. (2021) Siyuan Yao, Xiaoguang Han, Hua Zhang, Xiao Wang, and Xiaochun Cao. 2021. Learning Deep Lucas-Kanade Siamese Network for Visual Tracking. IEEE Transactions on Image Processing 30 (2021), 4814–4827. Yao et al. (2026) Siyuan Yao, Dongxiu Liu, Taotao Li, Shengjie Li, Wenqi Ren, and Xiaochun Cao. 2026. UAGLNet: Uncertainty-Aggregated Global–Local Fusion Network With Cooperative CNN–Transformer for Building Extraction. IEEE Transactions on Geoscience and Remote Sensing 64 (2026), 1–14. Yao et al. (2024) Siyuan Yao, Hao Sun, Tian-Zhu Xiang, Xiao Wang, and Xiaochun Cao. 2024. Hierarchical graph interaction transformer with dynamic token clustering for camouflaged object detection. IEEE Transactions on Image Processing 33 (2024), 5936–5948. Yao et al. (2025b) Siyuan Yao, Rui Zhu, Ziqi Wang, Wenqi Ren, Yanyang Yan, and Xiaochun Cao. 2025b. UMDATrack: Unified multi-domain adaptive tracking under adverse weather conditions. In IEEE International Conference on Computer Vision. 6466–6475. Ye et al. (2022) Botao Ye, Hong Chang, Bingpeng Ma, Shiguang Shan, and Xilin Chen. 2022. Joint feature learning and relation modeling for tracking: A one-stream framework. In European Conference on Computer Vision. 341–357. Yu et al. (2024) Hongtian Yu, Yunjie Tian, Qixiang Ye, and Yunfan Liu. 2024. Spatial transform decoupling for oriented object detection. In AAAI Conference on Artificial Intelligence, Vol. 38. 6782–6790. Yuan et al. (2026) Xiang Yuan, Gong Cheng, Jiacheng Cheng, Ruixiang Yao, and Junwei Han. 2026. Unc-SOD: An Uncertainty Learning Framework for Small Object Detection. IEEE Transactions on Image Processing 35 (2026), 1127–1142. Yuan et al. (2023) Xiang Yuan, Gong Cheng, Kebing Yan, Qinghua Zeng, and Junwei Han. 2023. Small Object Detection via Coarse-to-fine Proposal Generation and Imitation Learning. In IEEE International Conference on Computer Vision. 6317–6327. Zhang et al. (2025) Chang-Bin Zhang, Yujie Zhong, and Kai Han. 2025. Mr. DETR: Instructive Multi-Route Training for Detection Transformers. In IEEE Conference on Computer Vision and Pattern Recognition. 9933–9943. Zhang et al. (2022) Dawei Zhang, Yanwei Fu, and Zhonglong Zheng. 2022. UAST: Uncertainty-aware siamese tracking. In International Conference on Machine Learning. 26161–26175. Zhang et al. (2023) Hao Zhang, Feng Li, Shilong Liu, Lei Zhang, Hang Su, Jun Zhu, Lionel M. Ni, and Heung-Yeung Shum. 2023. DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection. In International Conference on Learning Representations. Zhang et al. (2020) Shifeng Zhang, Cheng Chi, Yongqiang Yao, Zhen Lei, and Stan Z Li. 2020. Bridging the gap between anchor-based and anchor-free detection via adaptive training sample selection. In IEEE Conference on Computer Vision and Pattern Recognition. 9759–9768. Zhou and Zhu (2024) Zhuangzhuang Zhou and Yingying Zhu. 2024. KLDet: Detecting Tiny Objects in Remote Sensing Images via Kullback–Leibler Divergence. IEEE Transactions on Geoscience and Remote Sensing 62 (2024), 1–16. Zhu et al. (2021) Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai. 2021. Deformable DETR: Deformable Transformers for End-to-End Object Detection. In International Conference on Learning Representations.