Paper deep dive
CSI-tuples-based 3D Channel Fingerprints Construction Assisted by MultiModal Learning
Chenjie Xie, Li You, Ruirong Chen, Gaoning He, Xiqi Gao
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/27/2026, 1:35:14 AM
Summary
The paper proposes a modularized multimodal framework for constructing 3D Channel Fingerprints (3D-CF) in low-altitude communication systems. By modeling 3D-CF as a collection of CSI-tuples (LAV positions and statistical CSI) rather than discretized grids, the authors formulate the construction as a multimodal regression task. The framework integrates geographic environment maps, communication measurements, and LAV coordinates using three modules: Correlation-based Multimodal Fusion (Corr-MMF), Multimodal Representation (MMR), and CSI Regression (CSI-R), achieving 27.5% higher accuracy than state-of-the-art methods.
Entities (6)
Relation Signals (4)
Multimodal Framework → contains → Corr-MMF
confidence 98% · includes a correlation-based multimodal fusion (Corr-MMF) module
Multimodal Framework → contains → MMR
confidence 98% · a multimodal representation (MMR) module
Multimodal Framework → contains → CSI-R
confidence 98% · a CSI regression (CSI-R) module
3D-CF → iscomposedof → CSI-tuples
confidence 95% · establish the 3D-CF model as a collection of CSI-tuples
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Low-altitude communications can promote the integration of aerial and terrestrial wireless resources, expand network coverage, and enhance transmission quality, thereby empowering the development of sixth-generation (6G) mobile communications. As an enabler for low-altitude transmission, 3D channel fingerprints (3D-CF), also referred to as the 3D radio map or 3D channel knowledge map, are expected to enhance the understanding of communication environments and assist in the acquisition of channel state information (CSI), thereby avoiding repeated estimations and reducing computational complexity. In this paper, we propose a modularized multimodal framework to construct 3D-CF. Specifically, we first establish the 3D-CF model as a collection of CSI-tuples based on Rician fading channels, with each tuple comprising the low-altitude vehicle's (LAV) positions and its corresponding statistical CSI. In consideration of the heterogeneous structures of different prior data, we formulate the 3D-CF construction problem as a multimodal regression task, where the target channel information in the CSI-tuple can be estimated directly by its corresponding LAV positions, together with communication measurements and geographic environment maps. Then, a high-efficiency multimodal framework is proposed accordingly, which includes a correlation-based multimodal fusion (Corr-MMF) module, a multimodal representation (MMR) module, and a CSI regression (CSI-R) module. Numerical results show that our proposed framework can efficiently construct 3D-CF and achieve at least 27.5% higher accuracy than the state-of-the-art algorithms under different communication scenarios, demonstrating its competitive performance and excellent generalization ability. We also analyze the computational complexity and illustrate its superiority in terms of the inference time.
Tags
Links
- Source: https://arxiv.org/abs/2603.25288v1
- Canonical: https://arxiv.org/abs/2603.25288v1
Trouble viewing inline? Open PDF directly →
Full Text
86,246 characters extracted from source content.
Expand or collapse full text
CSI-tuples-based 3D Channel Fingerprints Construction Assisted by MultiModal Learning Chenjie Xie, Li You, Ruirong Chen, Gaoning He, Xiqi Gao Part of this work was accepted by IEEE WCNC 2026 [58]. Chenjie Xie, Li You, and Xiqi Gao are with the National Mobile Communications Research Laboratory, Southeast University, Nanjing 210096, China, and also with the Purple Mountain Laboratories, Nanjing 211111, China (e-mail: cjxie@seu.edu.cn, lyou@seu.edu.cn, xqgao@seu.edu.cn). Ruirong Chen and Gaoning He are with the Huawei Technologies Co., Ltd., Shenzhen 518129, China (e-mail: ruirongchen@huawei.com, hegaoning@huawei.com). Abstract Low-altitude communications can promote the integration of aerial and terrestrial wireless resources, expand network coverage, and enhance transmission quality, thereby empowering the development of sixth-generation (6G) mobile communications. As an enabler for low-altitude transmission, 3D channel fingerprints (3D-CF), also referred to as the 3D radio map or 3D channel knowledge map, are expected to enhance the understanding of communication environments and assist in the acquisition of channel state information (CSI), thereby avoiding repeated estimations and reducing computational complexity. In this paper, we propose a modularized multimodal framework to construct 3D-CF. Specifically, we first establish the 3D-CF model as a collection of CSI-tuples based on Rician fading channels, with each tuple comprising the low-altitude vehicle’s (LAV) positions and its corresponding statistical CSI. In consideration of the heterogeneous structures of different prior data, we formulate the 3D-CF construction problem as a multimodal regression task, where the target channel information in the CSI-tuple can be estimated directly by its corresponding LAV positions, together with communication measurements and geographic environment maps. Then, a high-efficiency multimodal framework is proposed accordingly, which includes a correlation-based multimodal fusion (Corr-MMF) module, a multimodal representation (MMR) module, and a CSI regression (CSI-R) module. Numerical results show that our proposed framework can efficiently construct 3D-CF and achieve at least 27.5% higher accuracy than the state-of-the-art algorithms under different communication scenarios, demonstrating its competitive performance and excellent generalization ability. We also analyze the computational complexity and illustrate its superiority in terms of the inference time. I Introduction As the sixth-generation (6G) mobile communications continue to evolve, massive demands for low-altitude applications have arisen across transportation, agriculture, and emergency services, catalyzing the vigorous development of low-altitude communications [70, 61, 11, 22, 63]. Currently, low-altitude networks strive to provide real-time communication and navigation services for low-altitude vehicles (LAVs), facilitating the interoperability and coordination with terrestrial networks, delivering communication support for specific regions, and enhancing both network coverage and transmission quality [70, 63]. In the foreseeable future, low-altitude communications will further promote the synergistic integration of aerial and terrestrial resources, empowering a new paradigm for the development of mobile communications. However, acquiring high-quality channel state information (CSI) in low-altitude communication systems remains an unresolved challenge, which significantly impacts the performance of low-altitude wireless transmission. Due to the spatially non-uniformly distributed scatterers [40, 33], the radio environments in low-altitude scenarios become more complicated, and the signal propagation characteristics vary significantly across different altitudes, rendering the acquisition of accurate CSI considerably challenging. On the other hand, LAVs in low-altitude airspace exhibit high mobility, constrained power consumption, limited payload, and restricted computing power [40, 32], making the acquisition of real-time and efficient CSI more difficult. Fortunately, channel fingerprints (CF), also referred to as the channel knowledge map (CKM) [69] or the radio environment map (REM) [65], have emerged as a novel approach to address the aforementioned challenge. By definition, CF is a site-specific database storing the user-location-related CSI. In terrestrial communication systems, 2D-CF has demonstrated its effectiveness in assisting the acquisition of channel information, which supports the resource management [5], beam selection [55], wireless positioning [72], and other applications without repetitive estimations. Comparably, in low-altitude communication systems, 3D-CF will also directly provide the corresponding CSI based on the positions or trajectories of LAVs, thereby reducing pilot overhead, conserving wireless resources, and improving transmission efficiency. It is foreseeable that the introduction of 3D-CF will certainly bring a new perspective to the development of low-altitude communications. Currently, several studies have commenced to explore approaches for high-efficiency 3D-CF construction. For example, the authors in [49] detected the number of radiation sources based on the path loss (PL) model and constructed a 3D spectrum map accordingly. Nevertheless, adopting such a one-size-fits-all PL model inevitably leads to substantial errors for 3D-CF construction due to its considerable variability across different altitudes. To address the limitations of model-based approaches, data-driven methods have been extensively investigated. For instance, the authors in [14] leveraged LAV-based measurements to evaluate the interpolation algorithms, including nearest neighbor (N), linear, inverse distance weighting (IDW), and ordinary Kriging, validating their feasibility for 3D-CF reconstruction. Gaussian process regression (GPR), sparse Bayesian learning, and compressed sensing (CS) were also developed to reduce the required number of samples in [7, 50, 41, 42, 56]. Additionally, t-singular value decomposition (t-SVD), fiber sampling tensor decomposition (FSTD), block-term tensor decomposition (BTD), and other tensor-based algorithms were adopted to further enhance the computational efficiency by leveraging the smoothness prior of measurements in [67, 46, 28, 45]. However, these pure data-driven methods are entirely environment-blind. In practice, wireless channels are profoundly influenced by the geographic environment through a complex interplay of signal reflection, diffraction, and scattering mechanisms [35, 27], especially in low-altitude communication systems where line-of-sight (LOS) paths are predominant. To empower the geographic environment-assisted 3D-CF construction, researchers turn to machine learning (ML) for feasible solutions. For instance, the authors in [68] designed two deep neural networks (DNNs) to jointly reconstruct the 3D REM and its communication environment. Moreover, generative artificial intelligence (GenAI) was adopted for high-quality 3D-CF generation, supporting both radiation-aware and radiation-unaware scenarios with sparse spatial observations based on generative adversarial network (GAN) or diffusion model (DM) [71, 13, 51]. However, these computer vision-based GenAI algorithms need to model 3D-CF as images, requiring a discretization for the target region where LAVs in the same grid (pixel) share the identical CSI (pixel values) [59, 19]. This assumption introduces some critical limitations. Firstly, the uniform discretization reduces the flexibility of 3D-CF. Unlike terrestrial networks, LAV trajectories in low-altitude airspace are highly communication-demand-driven [57]. Indiscriminate gridding merely results in a mismatch between 3D-CF resolution and communication demands while wasting computational resources. Secondly, the intra-grid CSI sharing induces significant errors for 3D-CF applications. Particularly when buildings exist within the grid, PL and channel shadowing for LAVs on opposite sides may differ substantially and cannot be represented by the same CSI. Thirdly, the grid-based 3D model triggers an exponential increase in data volume, resulting in high computational complexity and being unsuitable for practical deployment. Multimodal learning (MML), a computer agent with intelligent capabilities such as understanding, reasoning, and learning, enables the integration of diverse data modalities for predictive tasks and offers a promising avenue to resolve the aforementioned issues of GenAI by transcending the image-only processing constraint. Currently, [24, 43, 60, 4, 17, 23, 1, 62, 20, 52, 16] have demonstrated enormous application potentials of MML-based CSI prediction in wireless communications. Particularly for 3D-CF construction in low-altitude systems, which involves more complex data modalities, including geographic environment maps, sparse communication measurements, and LAV coordinates, MML is expected to process and integrate the underlying information across these modalities, thereby offering novel solutions for 3D-CF construction and enhancing environmental awareness in low-altitude communication systems. Motivated by the above discussions, we investigate the 3D-CF construction based on MML for low-altitude communication systems. Specifically, we encapsulate the LAV’s coordinates and its corresponding channel information into a CSI-tuple, bypassing the discretization operation and directly fitting the mapping relationship between tuple elements to minimize the 3D-CF errors. Notably, the CSI stored in 3D-CF can be flexibly defined according to practical transmission requirements, such as the received signal strength (RSS), coverage, LOS probability, or even the channel covariance matrix and power angle spectrum (PAS). Then, we sufficiently consider the prior knowledge to assist the construction of 3D-CF, including geographic environment maps, sparse communication measurements, and LAV positions. Due to their heterogeneous data structures that can not be processed by a single-type network, we regard them as multimodal data and transform the 3D-CF construction problem into a multimodal regression task, thereby developing a modularized 3D-CF multimodal framework. The main contributions of this paper can be summarized as follows: • Based on the ground-to-LAV Rician fading channel model, we propose a 3D-CF model that is more suitable for low-altitude communication systems. Specifically, the 3D-CF is conceptualized as a collection of CSI-tuples with each tuple comprising the LAV position and its corresponding channel information. This model enables the adaptive adjustments of CSI-tuples according to practical communication demands, thus making the flexible 3D-CF construction possible. • Given the heterogeneous structures of different prior data, we formulate the 3D-CF construction problem as a multimodal regression task, where the target channel information in the CSI-tuple can be estimated directly by its corresponding LAV location, geographic environment maps, and measurement data. • Based on the structural characteristics and internal relations of prior data, we propose a highly efficient modularized multimodal framework for 3D-CF construction. During the data processing stage, a correlation-based multimodal fusion (Corr-MMF) module and a multimodal representation (MMR) module are designed based on the relativity between communication environments and measurement data, which extract and learn features of the CSI distribution in horizontal and vertical directions, respectively. In the CSI estimation phase, we align these different features via embedding operations and design the channel state information regression (CSI-R) module to estimate CSI by leveraging LAV positions as conditional inputs, thereby recovering CSI-tuples and accomplishing the 3D-CF construction. • We present numerical results to show that the proposed modularized multimodal framework achieves at least 27.5% higher accuracy than state-of-the-art algorithms in 3D-CF construction. Experimental results also demonstrate its competitive performance in generalization capability and computational complexity. The rest of this paper is organized as follows. In Section I, we establish the ground-to-LAV channel model and 3D-CF model in low-altitude communication scenarios, and formulate the construction problem accordingly. Section I elaborates on the structure of our proposed multimodal framework. Numerical results are presented in Section IV. Finally, we conclude the paper in Section V. Notations: ȷ=−1 = -1 denotes the imaginary unit. T a^T represents the transpose of vector a and ‖A‖F||A||_F is the Frobenius norm for matrix A. ℂM×N×KC^M× N× K denotes the M×N×KM× N× K dimensional complex tensor space and ℝ3R^3 represents the three-dimensional real space. ⋅E\·\ denotes the expectation operation. (,B)CN( a,B) represents the complex Gaussian distribution with mean a and covariance B. The notation ≜ is used for definitions. I System Model In this section, we introduce the channel model for ground-to-LAV links in low-altitude airspace. Then, a CSI-tuples-based 3D-CF model is established accordingly, which is highly flexible and does not rely on the spatial gridding. Based on the characteristics of 3D-CF, we formulate the construction problem as a multimodal regression task with the assistance of prior data. I-A Channel Model As illustrated in Fig. 1, we consider a ground-to-LAV system in low-altitude airspace, where the base station (BS), positioned at (x,y,0)(x,y,0), is equipped with a uniform linear array (ULA) comprising NBSN_ BS antenna elements [34]. For simplicity, each LAV in the target region employs a single antenna and moves in a constant velocity in a time interval of interest. To better characterize the ground-to-LAV links in target areas, we assume that each channel includes a line-of-sight (LOS) path and several reflected paths, both contributing to the received signal for a specific LAV [31]. By adopting the correlated Rician fading channel, the downlink (DL) channel between the BS and the m-th LAV over the n-th symbol can be modeled as [31, 18] m[n]=βm(¯m[n]+~m[n]), h_m[n]= _m ( h_m[n]+ h_m[n] ), (1) where βm _m represents the large-scale channel fading coefficient, ¯m[n] h_m[n] and ~m[n] h_m[n] denote the LOS component and NLOS component, respectively. Figure 1: A typical ground-to-LAV communication scenario, where all possible channel components include one LOS path and several reflected paths. Its propagation geometry takes the BS as the origin. For the LOS component ¯m[n] h_m[n], define K as the Rician factor and we have [18, 29, 48, 74] ¯m[n]=K+1(ϕm,0,θm,0)eȷ(2πνmm,0Tsn+φm,0), h_m[n]= KK+1 α( _m,0, _m,0)e (2π _m ξ_m,0T_sn+ _m,0), (2) where νm _m is the Doppler shift, TsT_s is the system sampling duration, φm,0 _m,0 is the phase shift for LOS component, and m,0≜mm,0T ξ_m,0 v_m k_m,0^T. m v_m is the unit velocity vector with an elevation angle ϕm,v _m,v and an azimuth angle θm,v _m,v, which is given by m=[cos(θm,v)sin(ϕm,v),sin(θm,v)sin(ϕm,v),cos(ϕm,v)]. v_m=[ ( _m,v) ( _m,v), ( _m,v) ( _m,v), ( _m,v)]. (3) m,0 k_m,0 is the unit wave vector with an elevation angle ϕm,0 _m,0 and an azimuth angle θm,0 _m,0, which is given by [37] m,0=[cos(θm,0)sin(ϕm,0),sin(θm,0)sin(ϕm,0),cos(ϕm,0)]. k_m,0=[ ( _m,0) ( _m,0), ( _m,0) ( _m,0), ( _m,0)]. (4) (ϕm,0,θm,0) α( _m,0, _m,0) is the steering vector with an elevation angle ϕm,0 _m,0 and an azimuth angle θm,0 _m,0, which can be expressed as [64] (ϕm,0,θm,0)=[1,eȷ2πdmλcζm,0,…,eȷ2π(NBS−1)dmλcζm,0], α( _m,0, _m,0)=[1,e 2π d_m _c _m,0,…,e 2π(N_ BS-1) d_m _c _m,0], (5) where dmd_m is the inter-antenna spacing, λc _c is the wavelength, and ζm,0≜cos(θm,0)sin(ϕm,0) _m,0 ( _m,0) ( _m,0). For the NLOS component ~m[n] h_m[n], define L as the number of NLOS paths, we have [29, 30, 3] ~m[n]=1K+1∑l=1L(ϕm,l,θm,l)Leȷ(2πνmm,lTsn+φm,l), h_m[n]= 1K+1 _l=1^L α( _m,l, _m,l) Le (2π _m ξ_m,lT_sn+ _m,l), (6) where m,l≜mm,lT ξ_m,l v_m k_m,l^T with m v_m and m,l k_m,l similar to (3) and (4), respectively. Assume that ϕm,ll=1L\ _m,l\_l=1^L, θm,ll=1L\ _m,l\_l=1^L, and φm,ll=1L\ _m,l\_l=1^L are independent random variables, then, according to the central limit theorem [2], when L tends to infinity, ~m[n] h_m[n] will approximate a zero-mean complex Gaussian random process, i.e., ~m[n]∼(0,m) h_m[n] (0, _m), where m _m represents the positive semi-definite spatial covariance matrix of the NLoS components for the m-th LAV [31, 18, 29, 30, 3]. Based on the analysis of (2) and (6), the channel model between the BS and the m-th LAV over the n-th symbol can be expressed as m[n]∼(H,R) h_m[n] (H,R) [31, 18, 29], where mean H=βm¯m[n]H= _m h_m[n] and covariance R=βmmR= _m _m. In accordance with this channel model, the RSS can be written by gm=PBS‖m[n]‖F2,g_m=P_ BS|| h_m[n]||_F^2, (7) where PBSP_ BS denotes the transmit power. I-B 3D-CF Model and Problem Formulation Based on the channel model presented in Section I-A, we next introduce the 3D-CF model for LAVs in low-altitude airspace and then formulate the construction problem. I-B1 3D-CF Model Traditional 3D-CF model typically partitions the target area into grids, where all receivers within the same grid share the identical channel information, thereby converting CF into an image [19, 59]. In contrast, we model 3D-CF as a collection of CSI-tuples (,Ω)\(X, )\, where X represents the LAV coordinates array and Ω denotes its associated channel information. On one hand, for any LAV in low-altitude airspace, we can always locate its corresponding CSI-tuple in 3D-CF and obtain the accurate channel information, thereby reducing errors induced by spatial gridding. On the other hand, the collection of CSI-tuples can be dynamically adjusted or reconstructed according to terminal density, low-altitude traffic load, and spatial utilization, maximizing the 3D-CF flexibility while minimizing unnecessary computational overhead in practical applications. In accordance with the channel model in Section I-A, we define the RSS in (7) as the target channel information stored in 3D-CF, i.e., Ω=gm =g_m. Therefore, the collection of CSI-tuples is ultimately expressed as =(m,Ψ(m))|Ψ:m∈ℝ3→gm,G=\(X_m, (X_m))| :X_m ^3→ g_m\, (8) which is our proposed 3D-CF model. Note that we define Ω as RSS solely for the convenience of elucidating the model, task, and methodology. In practice, Ω can be defined as different channel information according to practical communication requirements, such as PL, delay, Doppler shift, or even the channel covariance matrix, and the optimal beam indices, to match diverse applications. I-B2 Problem Formulation Under the definition of (8), the problem of constructing 3D-CF is transformed into the task of exploring function Ψ , that is, finding a mapping relationship from the LAV location to its corresponding RSS. However, directly fitting Ψ is extremely challenging due to the absence of distinct correlations between LAV location and its RSS. Consequently, we need to seek some prior information to facilitate the construction of Ψ . On one hand, the low-altitude geographic environment, which can be viewed as an image and conveniently captured by tools like RGB cameras, exerts a significant influence on the distribution of RSS. In terrestrial networks, geographic information has been proven to be effective and crucial in assisting the reconstruction of 2D-CF [19, 47, 6, 26]. In low-altitude airspace, environmental effects on RSS are more pronounced due to the dominance of LOS paths in ground-to-LAV channels. Therefore, the low-altitude geographic environment, denoted as ℰE, is one of the essential prior information to facilitate the construction of Ψ . On the other hand, measurable CF sampling data near the ground, denoted as tensor groG_ gro, can also serve as the prior information, as they partly reflect the signal propagation characteristics and reveal the underlying relationship between terminal locations and RSS. Based on ℰE and groG_ gro, the mapping relationship Ψ can be rewritten as Ψ:(m,ℰ,gro)→gm,m∈ℝ3. :(X_m,E,G_ gro)→ g_m,X_m ^3. (9) In practice, it is challenging to derive a feasible analytical solution by traditional interpolation methods, hence, we employ a deep neural network Ψ′ to fit Ψ in (9). In particular, the network Ψ′ involves three modal variables as inputs: the low-altitude communication environment ℰE, which encompasses both horizontal and vertical information of all buildings, vegetation, and other structures in the target low-altitude airspace; CF sampling data groG_ gro, which indicate the signal propagation characteristics; and the LAV position mX_m. Therefore, we formulate the problem of fitting Ψ by the network Ψ′ , i.e., the construction of 3D-CF, as a multimodal regression task, which is given by argminΘ _ \ ‖Ψ′[(m,ℰ,gro);Θ]−gm‖F2 \|| [(X_m,E,G_ gro); ]-g_m||_F^2 \ (10) s.t. s.t.\ gm=Ψ(m,ℰ,gro), g_m= (X_m,E,G_ gro), (10a) m∈ℝ3, _m ^3, (10b) where Θ is the trainable parameters for the network Ψ′ . I MultiModal Framework for 3D-CF Construction As analyzed in Section I-B, mutual relationships among mX_m, ℰE, and groG_ gro reveal the underlying patterns of spatial RSS distribution, hence, in this section, we develop a modularized multimodal framework to fit the mapping relationship Ψ and construct 3D-CF. As shown in Fig. 2, our proposed 3D-CF Multimodal framework includes three essential modules: the Corr-MMF module, the MMR module, and the CSI-R module, where the first two modules are designed to extract features of CSI distribution and geographic environments in horizontal and vertical directions, respectively, and the third module is responsible for spacial CSI prediction and 3D-CF reconstruction based on these features. Next, we will introduce them respectively. Figure 2: Diagram of the proposed 3D-CF MultiModal framework. This scheme includes three essential modules: the Corr-MMF module, the MMR module, and the CSI-R module. The first two modules are designed to extract features of CSI distribution and geographic environments in horizontal and vertical directions, respectively, and the third module is responsible for spacial CSI prediction and 3D-CF reconstruction. I-A The Correlation-based MultiModal Fusion (Corr-MMF) Module In conventional multimodal learning tasks, data from two or more media often exhibit strong correlations. Since their structural characteristics are heterogeneous, it is necessary to conduct a unified encoding and combination, known as MultiModal Fusion (MMF) [38]. By definition, MMF is the process of extracting and integrating features from two or more media to perform the subsequent regression or classification [15]. It leverages the correlation and complementarity among different data to keep critical features and remove redundant ones, integrating various information into a stable multimodal representation [21]. In our 3D-CF construction task, available CF measurements groG_ gro near the ground exhibit a strong correlation with the horizontal geographic environment information ℰhE_h [66, 25, 54]: on one hand, RSS in groG_ gro reveal the possible distribution of buildings, vegetation, and other structures in the target area [44]; on the other hand, ℰhE_h indicates the potential reflection, diffraction, scattering, and obstruction during signal propagation, thus affecting the distribution of RSS [36]. Consequently, exploring and fusing the correlated characteristics between groG_ gro and ℰhE_h are crucial for the multimodal framework to learn 3D-CF patterns in the horizontal direction, which motivates our design of the Corr-MMF module. In particular, the low-dimensional feature representations extracted and fused by the Corr-MMF module must satisfy the following two criteria: Criterion 1: The low-dimensional feature representations should maximally preserve the critical information inherent to both groG_ gro and ℰhE_h; Criterion 2: The low-dimensional feature representations should effectively preserve the correlated information among groG_ gro and ℰhE_h. (a) Diagram of the proposed Corr-MMF module. The network includes a two-stages encoder, a latent space, and a virtual decoder. The terminal attention mechanism (TAM) and channel attention mechanism (CAM) are introduced in this two-stages encoder as well. (b) Diagram of the channel attention mechanism (CAM). Figure 3: Diagram of the Corr-MMF module and its CAM. I-A1 Network Design for Criterion 1 For Criterion 1, a workable structure is the feature extractor (encoder) used to implement the key information extraction for both groG_ gro and ℰhE_h. As shown in Fig. 3(a), the encoder includes two stages: feature extraction and feature fusion. During the feature extraction stage, two sub-encoders, E1E_1 and E2E_2, are employed to extract features of groG_ gro and ℰhE_h, respectively, eliminating data redundancy and achieving dimensionality reduction. Specifically, each of E1E_1 and E2E_2 consists of several convolutional layers, each of which is accompanied by a rectified linear unit (ReLU) to conduct the downsampling operation. To further speed up the convergences and simultaneously enhance the generalization performance of the network, we incorporate the batch normalization (BN) after each convolutional layer. Furthermore, we observe that RSS of the specific LAV in 3D-CF exhibits a strong correlation with the channel information and geographical environment in its vicinity [8]. Consequently, a terminal attention mechanism (TAM) is designed for E1E_1 and E2E_2 to focus on data blocks closer to the LAV. Specifically, we construct two Gaussian masks M_G and MℰM_E, with their dimensions identical to groG_ gro and ℰhE_h, respectively. For the (m,nm,n)-th element in M_G and MℰM_E, we have M(m,n)=Mℰ(m,n)=e−dm,n22σ2,M_G(m,n)=M_E(m,n)=e^- d_m,n^22σ^2, (11) where dm,nd_m,n denotes the distance between this (m,nm,n)-th element and LAV position, and σ2σ^2 the variance of Gaussian distribution for M_G and MℰM_E. It is worth noting here that M_G and MℰM_E in TAM will assign different weights to groG_ gro and ℰhE_h at different spatial positions, thereby optimizing the process of feature extraction. Meanwhile, the Gaussian-distributed weights inherently maintain continuity and differentiability, facilitating the backward propagation computations for each sub-encoder. Overall, outputs of the sub-encoders E1E_1 and E2E_2 can be written as E1 _ E_1 =BN(ReLU(Conv(M⊙gro)))∈ℂh×w×c, = BN( ReLU( Conv(M_G _ gro))) ^h× w× c, (12) E2 _ E_2 =BN(ReLU(Conv(Mℰ⊙ℰh)))∈ℂh×w×c, = BN( ReLU( Conv(M_E _h))) ^h× w× c, (13) where ⊙ represents the Hadamard product, h,w,ch,w,c are heights, widths, and channels of E1O_ E_1 and E2O_ E_2, respectively. During the feature fusion stage, we adopt an add layer to integrate E1O_ E_1 and E2O_ E_2. Then, the added feature F=E1+E2∈ℂh×w×cO_ F=O_ E_1+O_ E_2 ^h× w× c is fed into a sub-encoder E3E_3 for further extraction of critical information, thereby completing the feature fusion. Note that the add layer here does not expand the number of channels in FO_ F, but rather exponentially enhances the feature informativeness they contain. Therefore, the subsequent sub-encoder E3E_3 must be capable of adaptively evaluating the importance of each channel to discern crucial features and emphasize them. To this end, we introduce a channel attention mechanism (CAM) to assign varying weights to different channels [53], as depicted in Fig. 3(b). Specifically, CAM employs an average pooling and a max pooling to separately obtain the global statistical information of each channel in mixed features FO_ F, denoted as avgpool(F)∈ℂ1×1×c avgpool(O_ F) ^1× 1× c and maxpool(F)∈ℂ1×1×c maxpool(O_ F) ^1× 1× c, respectively. Subsequently, shared convolutional layers with a kernel size of 11 are utilized to convert avgpool(F) avgpool(O_ F) and maxpool(F) maxpool(O_ F) into two sets of preliminary weight vectors. The summation of these two weights is then normalized via the Sigmoid function to generate the ultimate channel attention vector Vc∈ℂ1×1×cV_c ^1× 1× c, which is given by Vc=SigmoidConv[avgpool(F)]+Conv[maxpool(F)].V_c= Sigmoid\ Conv[ avgpool(O_ F)]+ Conv[ maxpool(O_ F)]\. (14) By broadcasting and multiplying VcV_c with FO_ F, the sub-encoder E3E_3 will focus more on the weighted crucial features, thereby optimizing the entire performance. Overall, assume that the sub-encoder E3E_3 has the same structure as E1E_1 and E2E_2, its outputs can be expressed as E3=BN(Relu(Conv(MC⊙F)))∈,O_ E_3= BN( Relu( Conv(M_C _ F))) , (15) where MC∈ℂh×w×cM_C ^h× w× c is the channel attention tensor by broadcasting VcV_c, with dimensions identical to those of FO_ F. As shown in Fig. 3(a), E3O_ E_3 represents the data manifold in the latent space Z, which is precisely the output of the Corr-MMF module. For ease of understanding, we denote E3O_ E_3 as CorrMMFO_ CorrMMF. Based on the above two-stage encoder, a virtual decoder D is adopted accordingly to further ensure the maximal preservation of key features described in Criterion 1. Note that the decoder D is termed as “virtual” because the Corr-MMF module exclusively requires the output CorrMMFO_ CorrMMF, while D solely serves as an optimization feedback. I-A2 Correlation Evaluation for Criterion 2 For Criterion 2, we have to develop additional constraints for Corr-MMF module to effectively preserve the correlated information among groG_ gro and ℰhE_h. For ease of derivation and analysis, denote the two-stage encoder in Fig. 3(a) as function f(⋅)f(·), where f(Z)=CorrMMFf(Z)=O_ CorrMMF with the two-view input Z=(gro,ℰh)Z=(G_ gro,E_h). Consequently, we can separately obtain the key features from groG_ gro and ℰhE_h by setting ℰh=0E_h=0 and gro=0G_ gro=0, denoted as f(Zℰ=0)f(Z_E=0) and f(Z=0)f(Z_G=0), respectively. Let fi(Zℰ=0)f_i(Z_E=0) and fi(Z=0)f_i(Z_G=0) represent the i-th elements of f(Zℰ=0)f(Z_E=0) and f(Z=0)f(Z_G=0), respectively, by introducing the adjusted cosine similarity, the correlation can be expressed as corr[f(Zℰ=0),f(Z=0)]=∑i=1NΔg,iΔe,i∑i=1NΔg,i2∑i=1NΔe,i2,corr[f(Z_E=0),f(Z_G=0)]= _i=1^N _ g,i _ e,i _i=1^N _ g,i^2 _i=1^N _ e,i^2, (16) where Δg,i=fi(Zℰ=0)−f(Zℰ=0)¯ _ g,i=f_i(Z_E=0)- f(Z_E=0), Δe,i=fi(Z=0)−f(Z=0)¯ _ e,i=f_i(Z_G=0)- f(Z_G=0), f(Zℰ=0)¯ f(Z_E=0) and f(Z=0)¯ f(Z_G=0) are the mean values for the key features from groG_ gro and ℰhE_h, respectively. Note that maximizing (16) empowers the Corr-MMF module to effectively preserve the correlation between features extracted from groG_ gro and ℰhE_h, which fulfills the requirements of Criterion 2. I-A3 Objective Function for Corr-MMF Module Based on the whole network in the above 1) and the correlation evaluation in the above 2), we then develop a matching objective function to train the Corr-MMF module so that it can uniformly satisfy the requirements of Criteria 1 and 2. Denote the network in Fig. 3(a) as ℱF, with its trainable parameters being ϑ . Note that ℱF is formed by cascading the encoder f(⋅)f(·) with a virtual decoder D. First, we introduce a fusion-reconstruction loss ℒfusion(ϑ)L_ fusion( ) to ensure that the fused features f(Z)f(Z) can be restored to the original two-view input Z=(gro,ℰh)Z=(G_ gro,E_h), which is given by ℒfusion(ϑ)=‖ℱ(Z;ϑ)−Z‖F.L_ fusion( )=||F(Z; )-Z||_F. (17) Next, a correlation loss ℒcorr(ϑ)L_ corr( ) is adopted to ensure that the features extracted and fused by Corr-MMF module effectively preserve the correlation between groG_ gro and ℰhE_h. According to (16), ℒcorr(ϑ)L_ corr( ) is designed as ℒcorr(ϑ)=1−corr[f(Zℰ=0),f(Z=0)],L_ corr( )=1-corr[f(Z_E=0),f(Z_G=0)], (18) where ℒcorr(ϑ)∈[0,2]L_ corr( )∈[0,2]. Additionally, we also introduce a cross-reconstruction loss ℒcross(ϑ)L_ cross( ) to assist the process of model training, given by ℒcross(ϑ)=‖ℱ(Z=0;ϑ)−gro‖F+‖ℱ(Zℰ=0;ϑ)−ℰh‖F.L_ cross( )=||F(Z_G=0; )-G_ gro||_F+||F(Z_E=0; )-E_ h||_F. (19) It is worth noting here that the cross-reconstruction loss ℒcross(ϑ)L_ cross( ) carries the physical meaning in practice: Due to the relativity between ℰhE_h and groG_ gro, it is theoretically possible to recover ℰhE_h from groG_ gro and vice versa [19]. This process emphasizes not only the reconstruction of ℰhE_h and groG_ gro but also their correlation, serving as a further enhancement for both ℒcorr(ϑ)L_ corr( ) and ℒfusion(ϑ)L_ fusion( ). Taking all the losses ℒfusion(ϑ)L_ fusion( ), ℒcorr(ϑ)L_ corr( ), and ℒcross(ϑ)L_ cross( ) into consideration, we present the objective function for Corr-MMF module as follows: ℒobj(ϑ)=ℒfusion(ϑ)+ℒcross(ϑ)+λℒcorr(ϑ),L_ obj( )=L_ fusion( )+L_ cross( )+ _ corr( ), (20) where λ is employed to control the proportion among different losses. When λ approaches 0, Corr-MMF module disregards the cross-modal correlation, failing Criterion 2. Conversely, when λ tends to infinity, it neglects the preservation of key features and the reconstruction of distinct modalities, violating Criterion 1. Therefore, λ requires prudent selection to achieve an optimal balance in Corr-MMF module performance. I-B The MultiModal Representation (MMR) Module In conventional multimodal learning tasks, raw multimedia data cannot be directly processed by machines, so it is necessary to conduct the unified description and processing, known as the MultiModal Representation (MMR) [38, 15]. By definition, MMR refers to the process of representing information from several media in a tensor or vector form [38]. While similar to MMF, it places greater emphasis on the uniformity across data representations rather than the feature fusion among different modalities. In our 3D-CF construction task, the horizontal geographic information ℰhE_ h and CF measurements groG_ gro have already been fused by Corr-MMF module (f(Z)=CorrMMFf(Z)=O_ CorrMMF), hence, we need to further process the vertical geographic information ℰvE_v to align with the data representation of CorrMMFO_ CorrMMF, which motivates our design of the MMR module. (a) Diagram of the MMR module. It employs an auto-encoder as the core architecture, including an input layer, a downsampling layer, a latent space, an upsampling layer, and an output layer. The spatial attention mechanism (SAM) is introduced as well. (b) Diagram of the spatial attention mechanism (SAM). Figure 4: Diagram of the MMR module and its SAM. Specifically, the MMR module takes ℰvE_v as the input and employs an auto-encoder as its core architecture, which comprises an input layer, a downsampling layer, a latent space, an upsampling layer, and an output layer, as shown in Fig. 4(b). The input layer is primarily utilized for data reshaping, transforming ℰvE_v into an easily processable tensor ℰv∈ℂh×w×cO_E_v ^h× w× c. The following downsampling layer and upsampling layer are similar to the encoder E3E_3 and virtual decoder D in Fig. 3(a), respectively. It is worth noting here that this structural similarity ensures the uniformity between CorrMMFO_ CorrMMF and ℰvO_E_v, which is the core of the MMR module. The output layer, corresponding to the input layer, is finally appended to reconstruct data back to ℰvE_v. Additionally, through a meticulous analysis of the data structure of ℰvO_E_v, we observe that pixels in feature matrices on each channel imply the specific characteristics of spatial communication environment at different positions. Therefore, the MMR module is expected to adaptively evaluate the importance of each pixel, thus highlighting critical environmental features at key locations. To this end, we introduce a spatial attention mechanism (SAM) to assign varying weights to different pixels in each channel of ℰvO_E_v [53], as depicted in Fig. 4(b). Specifically, SAM employs a global max pooling and a global average pooling to separately obtain feature maps of each pixel fibers in ℰvO_E_v, denoted as maxpool(ℰv)∈ℂh×w×1 maxpool(O_E_v) ^h× w× 1 and avgpool(ℰv)∈ℂh×w×1 avgpool(O_E_v) ^h× w× 1, respectively. Subsequently, we construct the integrated statistical information by concatenating maxpool(ℰv) maxpool(O_E_v) and avgpool(ℰv) avgpool(O_E_v), and feed it into a convolutional layer to derive the preliminary weight matrix. Afterwards, the Sigmoid function is introduced for normalization to acquire the ultimate spatial attention matrix Vs∈ℂh×w×1V_s ^h× w× 1, which is given by Vs V_s =SigmoidConv[Concat(avgpool(ℰv),maxpool(ℰv))]. = Sigmoid\ Conv[ Concat( avgpool(O_E_v), maxpool(O_E_v))]\. (21) By broadcasting and multiplying VsV_s with ℰvO_E_v, the downsampling process will focus more on the weighted crucial environmental features at key locations, thereby improving the performance of the MMR module. Overall, the output of the MMR module, namely the low-dimensional representation of ℰvE_v, can be expressed as MMR=BN(Relu(Conv(MS⊙ℰv))),O_ MMR= BN( Relu( Conv(M_S _E_v))), (22) where MS∈ℂh×w×cM_S ^h× w× c is the spatial attention tensor by broadcasting VsV_s, with dimensions identical to those of ℰvO_E_v. I-C The Channel State Information Regression (CSI-R) Module In the 3D-CF construction task, ℰE and groG_ gro have been, respectively, transformed into low-dimensional representations CorrMMFO_ CorrMMF and MMRO_ MMR with identical data structures, where CorrMMFO_ CorrMMF embodies the 3D-CF patterns in the horizontal direction and MMRO_ MMR embodies the environment patterns in the vertical direction. Based on this, we next design the CSI-R module to predict the channel information at a specific LAV location X in low-altitude airspace, thereby enabling the flexible 3D-CF construction free from grid constraints. Figure 5: Diagram of the proposed CSI-R module. The network includes a feature embedding layer and a fully connected regression layer. As illustrated in Fig. 5, The CSI-R module uses CorrMMFO_ CorrMMF and MMRO_ MMR as inputs and consists of a feature embedding layer and a fully connected (FC) regression layer. Regarding the feature embedding layer, we first partition the feature maps of CorrMMFO_ CorrMMF and MMRO_ MMR into several equally-sized patches by a sliding window. Then, a convolutional layer with its kernel size matching the patch size is followed to generate a sequence of embedding vectors. Flattening these sequences yields a one-dimensional embedding output, which will be fed into the subsequent FC regression layer. Note that the entire feature embedding process captures elemental similarities in the feature maps of CorrMMFO_ CorrMMF and MMRO_ MMR, respectively, positioning elements with higher similarity closer in the embedding space, thus accelerating the convergence and enhance the accuracy. The FC regression consists of several dense layers, each followed by a dropout layer to prevent severe overfitting. During the regression process, 3D coordinates of the LAV are vectorized and incorporated as conditional information, prompting the prediction of location-specific channel characteristics for 3D-CF construction. In summary, the mapping relationship Ψ can be effectively approximated through the cascade of Corr-MMF module, MMR module, and CSI-R module, thereby enabling the construction of CSI-tuples in (8) and achieving the configuration of 3D-CF according to practical communication requirements. This certainly provides novel insights into the low-altitude communications. IV Numerical Results In this section, numerical results are provided to evaluate the performance of the proposed 3D-CF multimodal framework. First, we introduce the generation of our Sionna-based datasets and detail the experiment setup. Then, we explore the impact of λ on the accuracy of 3D-CF construction. By employing Kriging interpolation, GPR, GAN, and FL as benchmarks, we next compare the 3D-CF construction performance across different scenarios, demonstrating the prediction accuracy and generalization capability of the proposed 3D-CF multimodal framework. Finally, we analyze the computational complexity. IV-A Datasets All datasets are generated by NVIDIA Sionna, an open-source library for research on wireless communication systems [12]. Specifically, we first download the geographic environment maps of Nanjing, China, from OpenStreetMap (OSM) [10], a collaborative mapping project maintained by a global community of volunteers who continuously update and validate geographic information. These maps are then imported into Blender to construct communication scenarios, which are subsequently transferred to Sionna for dataset generation via ray tracing (RT). Note that all selected areas represent typical urban macro-cell or micro-cell scenarios with varying building shapes, quantities, and distributions, thereby ensuring strong data diversity to support model training and facilitate the evaluation of generalization capability. To facilitate our experiments, the BS is randomly deployed in the target region with NBS=64N_ BS=64, and the LAV flight altitudes are confined to 25−8025-80 m. CF measurement data groG_ gro are uniformly sampled at a height of 1.51.5 m with a resolution of 11 m. Note that there exists an inherent resolution-accuracy-complexity trade-off: enhanced near-ground sampling density improves the accuracy of 3D-CF construction, but increases the computational complexity. Hence, the resolution of groG_ gro should be determined based on practical limitations in real-world scenarios. Based on the specific interface for wireless communication simulation in Sionna, we conduct ray tracing to obtain all possible LOS and NLOS paths between BS and LAVs, establish the channel model in accordance with (1), compute RSS according to (7), and construct 3D-CF by (8). To further enhance the stability of network training, we adopt the “max-min” linear normalization to scale the raw RSS in 3D-CF into [0,1][0,1], which is given by gm′=maxgm−gthr(gm)max−gthr,0,g_m = \ g_m-g_ thr(g_m)_ max-g_ thr,0 \, (23) where gthrg_ thr is the RSS threshold since signals below gthrg_ thr are practically undetectable by LAV in real-world scenarios [26]. Additional parameters used to generate the datasets are displayed in Table I. TABLE I: Parameters used to generate datasets in the platform of Sionna. Parameters Value Communication scenario urban macro-cell / micro-cell Size of the target region 256×256256× 256 Number of antenna elements for BS 6464 BS height 2525 m LAV flight altitudes 25−8025-80 m Sampling height of groG_ gro 1.51.5 m Carrier frequency 3.53.5 GHz Maximum number of paths 55 RSS threshold −147-147 dB Transmit power 2323 dBm IV-B Experiment Setup In correspondence with the 3D-CF multimodal framework in Section I, we configure all hyper-parameters of the three modules in Table I. Specifically, regarding the Corr-MMF module, the batch size is set to be 128128 and the epochs are 6060, with the learning rate programmed to be 0.00010.0001 for the first 3535 epochs and then linearly decaying to zero. For the MMR module, the training epochs are set to be 6060, and the learning rate is initialized at 0.0010.001 for the first 3535 epochs and subsequently reduced linearly to zero. The training process of the CSI-R module requires 1515 epochs with a reduced batch size of 3232. The learning rate is set to be 0.00010.0001 for the first 1010 epochs, followed by a linear decay to zero. All the simulations are implemented by TensorFlow, with the computer equipped with an Intel(R) Core(TM) i7-12700 and a GeForce GTX 4090. To evaluate the performance of the proposed multimodal framework in 3D-CF construction, we employ the mean absolute error (MAE) and root mean square error (RMSE) as metrics, which can be expressed as MAE MAE =1n∑i=1n|g~m′−gm′|, = 1n _i=1^n| g_m -g_m |, (24) RMSE RMSE =1n∑i=1n(g~m′−gm′)2, = 1n _i=1^n( g_m -g_m )^2, (25) where g~m′ g_m is the predicted RSS in CSI-tuples in 3D-CF. Since MAE provides equal weights to all errors, it can intuitively reflect the overall 3D-CF construction performance. Conversely, RMSE exhibits greater sensitivity to outliers, facilitating the detection of extreme prediction deviations in our proposed 3D-CF multimodal framework. The combination of these two metrics enables a more comprehensive evaluation of the experimental results. TABLE I: Hyper-parameters of Corr-MMF, MMR, and CSI-R modules in 3D-CF MultiModal framework. Parameter Module Corr-MMF MMR CSI-R Epochs 60 50 15 Delay Epochs 35 35 10 Learning Rate 0.005 0.001 0.0001 Batch Size 128 128 32 Optimizer Adam IV-C Influence of λ in the Corr-MMF module In Corr-MMF module, λ would influence the network performance by controlling the proportions among different losses. Therefore, in this section, we examine its impact on the 3D-CF construction error and analyze its effect on the convergence behavior of the Corr-MMF module. Figure 6: MAE and RMSE performance of the 3D-CF multimodal framework under different λ. As shown in Fig. 6, the proposed multimodal framework can achieve the minimal 3D-CF construction error at λ=1λ=1, with MAE=0.029 MAE=0.029 and RMSE=0.060 RMSE=0.060, capturing approximately 89.689.6% of the correlation between ℰhE_h and groG_ gro. As λ gradually decreases, the weight of ℒcorr(ϑ)L_ corr( ) reduces and the 3D-CF construction error increases progressively. Notably at λ=0λ=0, the performance of MAE and RMSE deteriorates by 24.124.1% and 1515% respectively, with captured correlation between ℰhE_h and groG_ gro dropping to 19.819.8%. In practice, an excessively small λ will prevent the Corr-MMF module from extracting correlated features between ℰhE_h and groG_ gro, thereby disabling the operation of feature fusion and degrading the construction performance of the proposed 3D-CF multimodal framework. Similarly, as λ increases, the weight of ℒcorr(ϑ)L_ corr( ) rises and the performance of 3D-CF construction declines. Especially at λ=10λ=10, where λℒcorr(ϑ) _ corr( ) can be considered infinite compared with the magnitude of other losses, the performance of MAE and RMSE deteriorates by 34.434.4% and 13.313.3%, respectively. Here, an excessively large λ will unreasonably overemphasize the correlation between ℰhE_h and groG_ gro (capturing as much as 99.9%), completely suppressing their distinctive features and leading to a collapse of the Corr-MMF module. To conclude, the determination of λ must strike an effective balance between distinctive feature extraction and correlated feature fusion, thus optimizing the performance of Corr-MMF module and achieving accurate 3D-CF construction. Figure 7: The convergence behavior of the validation loss for the Corr-MMF module under different λ. Fig. 7 presents the convergence behavior of the validation loss for the Corr-MMF module under different λ. A fundamental observation is that larger λ corresponds to greater λℒcorr(ϑ) _ corr( ), consequently resulting in higher values of the objective function ℒobj(ϑ)L_ obj( ). Nevertheless, the validation loss is lower when λ=1λ=1 compared to λ=0.5λ=0.5, which demonstrates that λ=1λ=1 can optimize the network performance to the greatest extent. This result is consistent with the conclusion obtained in Fig. 6. Furthermore, smaller λ leads to smoother loss curves and contributes to a more stable network training, whereas larger λ results in greater fluctuations. This occurs because smaller λ balances the magnitudes of different losses, enabling steadier execution of the gradient descent algorithm. Conversely, larger λ causes ℒcorr(ϑ)L_ corr( ) to be more dominant, making the network significantly more susceptible to its oscillation and hence resulting in fluctuations across different epochs. To conclude, variations in λ substantially impact the stability of network training. IV-D Evaluation of the 3D-CF Construction Performance In this section, we compare the performance of the proposed 3D-CF multimodal framework with four benchmarks and evaluate the generalization capability under different scenarios. IV-D1 Benchmarks To evaluate the performance of our proposed 3D-CF multimodal framework, the Kriging interpolation [14], GPR [7], GAN [13], and FL [9] are adopted as benchmarks. • Kriging interpolation [14]: Kriging is a classical interpolation method for constructing 3D-CF. Based on the prior sampling data, it achieves RSS estimation at arbitrary locations by incorporating distances and modeling 3D spatial correlation through different variograms. The Kriging interpolation method does not require a grid-based model for CF, thus making the flexible 3D-CF construction possible. • GPR [7]: Gaussian process regression is a widely used statistical non-parametric model for 3D-CF construction. It can construct the optimal approximator of RSS distribution by designing specific kernel functions based on signal propagation characteristics. This GPR-based method does not require the grid-based model as well, therefore making the 3D-CF construction more flexible. • GAN [13]: Generative adversarial network is a significant deep generative model employed for 3D-CF construction. Unlike the Kriging-based method and GPR-based method, GAN perceives the communication environment and utilizes it as conditional information to generate 3D-CF. However, it requires the grid-based model to represent CF as a multi-channel image, which necessitates a uniform partitioning of the target area and assigns the same RSS for all LAVs within the same grid, resulting in lower accuracy and reduced flexibility. • FL [9]: Federated learning represents a state-of the-art framework to recover 3D-CF. Embedded with deep neural networks, it enables the collaborative utilization of data from multiple LAVs to establish the global 3D-CF. Since the regression-based FL framework requires no special assumptions about the CF model, it balances both accuracy and flexibility in the process of 3D-CF construction. The sampling rate for Kriging interpolation and GPR is set to be 55%, while the training, validation, and test datasets are identical for all other ML-based methods. (a) 3D-CF for scenario 1. (b) 3D-CF for scenario 2. (c) 3D-CF for scenario 3. (d) 3D-CF for scenario 4. Figure 8: Illustrations of 3D-CF under four randomly selected communication scenarios. The RSS distribution at four horizontal planes, specifically at heights of 10 m, 20 m, 30 m, and 40 m, are illustrated along the Z axis. TABLE I: Comparison of the 3D-CF multimodal framework with state-of-the-art approaches under different communication scenarios regarding the RMSE and MAE performance. Scenario 1 Scenario 2 Scenario 3 Scenario 4 RMSE MAE RMSE MAE RMSE MAE RMSE MAE Kriging [14] 0.522 0.467 0.348 0.287 0.340 0.289 0.459 0.384 GPR [7] 0.343 0.223 0.291 0.201 0.304 0.161 0.325 0.194 GAN [13] 0.234 0.195 0.208 0.156 0.269 0.221 0.266 0.223 Federated Learning [9] 0.088 0.067 0.041 0.034 0.035 0.029 0.056 0.039 Multimodal framework (ours) 0.069↓ 0.044 ↓ 0.026 ↓ 0.019 ↓ 0.024 ↓ 0.021 ↓ 0.045 ↓ 0.027 ↓ IV-D2 Comparison to Different Benchmarks Table I presents the comparison of 3D-CF construction performance between the proposed multimodal framework and the four other baselines. Compared to the non-AI methods like Kriging interpolation [14] and GPR [7], the 3D-CF multimodal framework demonstrates a reduction in RMSE by factors of 7.57.5 and 4.94.9, respectively, and a decrease in MAE by factors of 10.610.6 and 5.15.1, respectively, significantly enhancing the accuracy of 3D-CF construction. Fundamentally, the Kriging-based method and GPR merely fit the data itself without exploring the impact of communication environments on channel characteristics. Particularly in low-altitude airspace, where the physical environments become more complex, pure data fitting is no longer sufficient to accurately predict the distribution of channel information. Secondly, compared to the classical generative model GAN [13], the proposed multimodal framework can achieve 3.43.4-fold and 4.44.4-fold performance advantages in RMSE and MAE metrics, respectively. One point should be noted that in low-altitude airspace, the physical environment comprises different data modalities in horizontal and vertical dimensions, with LAV coordinates evolving into ternary arrays. Since GAN in [13] is limited to processing merely single-modal data, the 3D-CF construction accuracy is inevitably compromised. Moreover, the adoption of GAN needs to model 3D-CF as images, resulting in a rigid and inflexible construction and cannot achieve the non-uniform density adaptation according to practical communication requirements. Our proposed multimodal framework effectively addresses these limitations, thereby significantly reducing 3D-CF construction errors and enhancing flexibility. Finally, compared to the state-of-the-art method, the proposed multimodal framework still outperforms the FL-based approach [9] by 27.527.5% and 52.252.2% in terms of the RMSE and MAE, respectively. This stems from the operation of feature extraction and feature fusion for diverse data modalities via Corr-MMF module and MMR module, which enables the final regression network to comprehensively learn the mapping relationship between LAV’s positions and its RSS. To conclude, owing to the CSI-tuples-based model and the module-based design, the proposed multimodal framework exhibits superior performance and high flexibility in 3D-CF construction. IV-D3 Comparison Under Different Scenarios Fig. 8 presents the 3D-CF across four randomly selected communication scenarios with different urban structures or building densities. For illustrative purposes only, the RSS distribution across four horizontal planes at 10 m, 20 m, 30 m, and 40 m is given along the Z‑axis. Table I provides the comparison of RMSE and MAE performance across these scenarios. It can be observed that, regardless of the scenario, our proposed multimodal framework can always achieve 3D-CF construction with smaller errors than the baselines, demonstrating its generalization ability with RMSE and MAE standard deviations of 0.00120.0012 and 0.00140.0014, respectively. IV-E Ablation Experiments In this subsection, we conduct two ablation experiments to respectively demonstrate the contributions of different input modalities and attention mechanisms to the proposed 3D-CF multimodal framework. IV-E1 Input Modality Table IV presents the contribution of each input modality to the proposed 3D-CF multimodal framework. Firstly, when ℰvE_v is missing, the model fails to learn the vertical signal-propagation characteristics, leading to a rapid performance deterioration. This outcome indicates that the construction of 3D-CF must account for the three-dimensional nature of the low-altitude environment. If existing 2D-CF construction methods are directly applied to low-altitude scenarios that only consider the horizontal building distribution, severe model mismatch will arise. Secondly, when both ℰhE_h and groG_ gro are missing, the model cannot learn the horizontal characteristics of the CSI distribution, leading to significant degradation in model performance as well. However, when only one of them is missing, the 3D-CF constructed by the proposed architecture experiences only a 4.4% drop in performance. In practice, both horizontal environmental information and near-ground measurements can reflect the signal-propagation characteristics in the horizontal direction, and the absence of either alone does not cause model collapse. TABLE IV: Contribution of each input modality to the 3D-CF multimodal framework. Geographic environments ℰE Measurements groG_ gro Evaluation metrics ℰhE_h ℰvE_v RMSE MAE ✓ ✓ ✓ 0.045 0.027 ✓ ✓ ✗ 0.047 0.028 ✗ ✓ ✓ 0.047 0.029 ✓ ✗ ✓ 0.687 0.675 ✗ ✗ ✓ 0.700 0.692 In summary, the ablation study on input modalities demonstrates the effectiveness and necessity of each type of prior information, thereby validating the rationality of our proposed 3D-CF multimodal framework. IV-E2 Attention Mechanisms Table V presents the impact of three different attention mechanisms, including TAM, CAM, and SAM, on the 3D-CF construction. Firstly, the model performance deteriorates severely when the TAM is removed, whereas the degradation is less pronounced when the CAM or SAM is removed. In principle, the TAM operates directly on the inputs, primarily strengthening the near-ground measurement data near the LAV projection. In contrast, the CAM and SAM function internally within the network, emphasizing key features through self-learning. Once the TAM is absent, the data itself lacks essential weighting, leading to severe performance degradation regardless of the internal network design. Secondly, the performance decline in the absence of CAM is less severe than that when the SAM is missing. Since both the CAM and TAM are mechanisms within the Corr-MMF module, even if CAM is removed, the TAM mechanism still enables this module to achieve a relatively effective training result. TABLE V: Ablation study on different attention mechanisms for 3D-CF construction. Corr-MMF module MMR module Evaluation metrics CAM TAM SAM RMSE MAE ✓ ✓ ✓ 0.045 0.027 ✗ ✓ ✓ 0.065 0.052 ✓ ✗ ✓ 0.546 0.455 ✓ ✓ ✗ 0.290 0.282 In summary, the ablation study on attention mechanisms demonstrates that all TAM, CAM, and SAM play indispensable roles in the proposed multimodal framework, enabling the accurate and efficient 3D-CF construction. IV-F Complexity Comparison In this subsection, we analyze the complexity of the proposed 3D-CF multimodal framework versus benchmarks and present comparative results of their inference times. Figure 9: Comparison of the 3D-CF inference time between the proposed multimodal framework and state-of-the-art approaches. Define N as the amount of training data. For the Kriging-based method, we must solve the Kriging system of equations with a size of N×N× N to complete the construction of 3D-CF, the complexity of which remains (N3)O(N^3) even when employing the Cholesky decomposition [39]. For the GPR-based method, the complexity of covariance matrix inversion is (N2)O(N^2), and the complexity of marginal likelihood approximation in the process of solving posterior probability is (N)O(N). Consequently, the overall complexity for the GPR-based 3D-CF construction is (N3)O(N^3) [73]. For the GAN-based approach, its complexity can be expressed as ∑ℓ=1LGANℓ(NCℓ−1HℓWℓCℓKℓ2) _ =1^L_ GANO_ (NC_ -1H_ W_ C_ K^2_ ), where LGANL_ GAN is the number of convolution layers, KℓK_ is size of the convolution kernel, and Hℓ,WℓH_ ,W_ , and CℓC_ are heights, widths, and channels of the ℓ -th layer’s output, respectively. For the FL-based approach which is mainly composed of the multilayer perceptron, its complexity is given by ∑p=1PFL(Ndphp) _p=1^P_ FLO(Nd_ph_p), where PFLP_ FL is the number of hidden layers, dpd_p is the input dimension and hph_p represents the number of neurons in the p-th hidden layer. Regarding our proposed 3D-CF multimodal framework, its complexity is the sum of complexities from all three modules, which is given by ∑ℓ=1LCorrMMF+LMMR(NCℓ−1HℓWℓCℓKℓ2)+∑p=1PCSIR(Ndphp) _ =1^L_ CorrMMF+L_ MMRO(NC_ -1H_ W_ C_ K^2_ )+ _p=1^P_ CSIRO(Nd_ph_p), where the first term contains both the Corr-MMF module and MMR module due to their similar network structures. Fig. 9 further presents the inference time of 3D-CF construction for different methods. As observed, the proposed multimodal framework significantly reduces the inference time compared to the Kriging-based and GPR-based methods. It also achieves lower construction errors and enhanced flexibility with comparable time complexity compared with the GAN-based approach. Relative to the state-of-the-art FL algorithm, the proposed scheme nearly halves the computational time. It should be emphasized that the proposed 3D-CF is designed to be trained, inferred, and deployed at the BS side. Given the abundant computational resources available at the BS, this framework is practical. To conclude, our multimodal framework exhibits competitive advantages in computational complexity, indicating great potential for practical applications. V Conclusion In this paper, we proposed a modularized multimodal framework to construct 3D-CF for low-altitude communications. Firstly, we established the 3D-CF model based on the ground-to-LAV Rician fading channels, which were defined as a collection of CSI-tuples with each tuple composed of LAV’s positions and its corresponding channel information. Due to the heterogeneous structures of different prior data such as LAV coordinates, geographic environment maps, and sampling data, we transformed the 3D-CF construction problem into a multimodal regression task and proposed a high-efficiency modularized multimodal framework accordingly, where the Corr-MMF module and the MMR module were designed to extract features of CSI distribution in horizontal and vertical directions, and the CSI-R module was developed to estimate the target CSI and reconstruct the 3D-CF. Numerical results demonstrated the competitive performance and generalization ability of our proposed 3D-CF multimodal framework, which attains an accuracy improvement of at least 27.5% over the benchmarks under different communication scenarios. We also analyzed the computational complexity and illustrated its superiority in terms of the inference time. In the future, it will be interesting to quantify the impact of measurement data resolution on the constructed 3D-CF accuracy and its computational complexity. Furthermore, considering air-to-air channels and dynamic environmental factors also constitutes a highly promising research direction for low-altitude 3D-CF techniques. References [1] Y. Ahn, J. Kim, S. Kim, K. Shim, J. Kim, S. Kim, and B. Shim (Oct. 2023) Toward intelligent millimeter and terahertz communication for 6G: Computer vision-aided beamforming. IEEE Wireless Commun. 30 (5), p. 179–186. External Links: Document Cited by: §I. [2] M. Cardone, A. Dytso, and C. Rush (Apr. 2023) Entropic central limit theorem for order statistics. IEEE Trans. Inf. Theory 69 (4), p. 2193–2205. External Links: Document Cited by: §I-A. [3] H. Chang, J. Bian, C. Wang, Z. Bai, W. Zhou, and e. M. Aggoune (May 2019) A 3D non-stationary wideband GBSM for low-altitude UAV-to-ground V2V MIMO channels. IEEE Access 7 (), p. 70719–70732. External Links: Document Cited by: §I-A, §I-A. [4] G. Charan, T. Osman, A. Hredzak, N. Thawdar, and A. Alkhateeb (Austin, TX, USA, Apr. 2022) Vision-position multi-modal beam prediction using real millimeter wave datasets. In in Proc. IEEE WCNC 2022, Vol. , p. 2727–2731. External Links: Document Cited by: §I. [5] H. Che, L. You, J. Wang, Z. Jin, C. Xie, and X. Gao (2025 (early access)) Channel charting-assisted non-orthogonal pilot allocation for uplink XL-MIMO transmission. Chin. J. Electron.. Cited by: §I. [6] G. Chen, Y. Liu, T. Zhang, J. Zhang, X. Guo, and J. Yang (May 2023) A graph neural network based radio map construction method for urban environment. IEEE Commun. Lett. 27 (5), p. 1327–1331. External Links: Document Cited by: §I-B2. [7] X. Chen, X. Zhong, Z. Zhang, L. Dai, and S. Zhou (Oct. 2025) High-efficiency urban 3D radio map estimation based on sparse measurements. IEEE Trans. Veh. Technol. 74 (10), p. 16488–16493. External Links: Document Cited by: §I, 2nd item, §IV-D1, §IV-D2, TABLE I. [8] Z. Cui, C. Briso-Rodríguez, K. Guan, l. Güvenç, and Z. Zhong (Sep. 2020) Wideband air-to-ground channel characterization for multiple propagation environments. IEEE Antennas Wirel. Propag. Lett. 19 (9), p. 1634–1638. External Links: Document Cited by: §I-A1. [9] Q. Gong, F. Wu, D. Yang, L. Xiao, and Z. Liu (Dec. 2023) 3D radio map reconstruction and trajectory optimization for cellular-connected UAVs. J. Commun. Inf. Networks 8 (4), p. 357–368. External Links: Document Cited by: 4th item, §IV-D1, §IV-D2, TABLE I. [10] M. Haklay and P. Weber (Dec. 2008) OpenStreetMap: User-generated street maps. IEEE Pervasive Comput. 7 (4), p. 12–18. External Links: Document Cited by: §IV-A. [11] N. Hossein Motlagh, T. Taleb, and O. Arouk (Dec. 2016) Low-altitude unmanned aerial vehicles-based internet of things services: Comprehensive survey and future perspectives. IEEE Internet Things J. 3 (6), p. 899–922. External Links: Document Cited by: §I. [12] J. Hoydis, F. A. Aoudia, S. Cammerer, M. Nimier-David, N. Binder, G. Marcus, and A. Keller (Kuala Lumpur, Malaysia, Dec. 2023) Sionna RT: Differentiable ray tracing for radio propagation modeling. In Proc. IEEE Globecom Workshops (GC Wkshps), Vol. , p. 317–321. External Links: Document Cited by: §IV-A. [13] T. Hu, Y. Huang, J. Chen, Q. Wu, and Z. Gong (Jun. 2023) 3D radio map reconstruction based on generative adversarial networks under constrained aircraft trajectories. IEEE Trans. Veh. Technol. 72 (6), p. 8250–8255. External Links: Document Cited by: §I, 3rd item, §IV-D1, §IV-D2, TABLE I. [14] A. Ivanov, K. Tonchev, V. Poulkov, A. Manolova, and A. Vlahov (Tampa, FL, United states, Nov. 2023) Interpolation accuracy evaluation for 3D radio environment maps construction. In IEEE Int. Symp. Wireless Pers. Multimedia Commun. (WPMC), Vol. , p. 1–7. External Links: Document Cited by: §I, 1st item, §IV-D1, §IV-D2, TABLE I. [15] S. Jabeen, X. Li, M. S. Amin, O. E. F. Bourahla, S. Li, and A. Jabbar (Feb. 2022) A review on methods and applications in multimodal deep learning. ACM Trans. Multimedia Comput. Commun. Appl. 19, p. 1– 41. Cited by: §I-A, §I-B. [16] F. Jiang, L. Dong, Y. Peng, K. Wang, K. Yang, C. Pan, and X. You (Jan. 2025) Large AI model empowered multimodal semantic communications. IEEE Commun. Mag. 63 (1), p. 76–82. External Links: Document Cited by: §I. [17] F. Jiang, Y. Peng, L. Dong, K. Wang, K. Yang, C. Pan, D. Niyato, and O. A. Dobre (Dec. 2024) Large language model enhanced multi-agent systems for 6G communications. IEEE Wireless Commun. 31 (6), p. 48–55. External Links: Document Cited by: §I. [18] F. Jiang and A. L. Swindlehurst (Jun. 2012) Optimization of UAV heading for the ground-to-air uplink. IEEE J. Sel. Areas Commun. 30 (5), p. 993–1005. External Links: Document Cited by: §I-A, §I-A, §I-A, §I-A. [19] Z. Jin, L. You, J. Wang, X. Xia, and X. Gao (Feb. 2025) An I2I inpainting approach for efficient channel knowledge map construction. IEEE Trans. Wireless Commun. 24 (2), p. 1415–1429. External Links: Document Cited by: §I, §I-B1, §I-B2, §I-A3. [20] Z. Jin, L. You, D. Wing Kwan Ng, X. Xia, and X. Gao (May 2025) Near-field channel estimation for XL-MIMO: A deep generative model guided by side information. IEEE Trans. Cognit. Commun. Networking 12 (), p. 628–643. External Links: Document Cited by: §I. [21] G. Joshi, R. Walambe, and K. Kotecha (Mar. 2021) A review on explainability in multimodal deep neural nets. IEEE Access 9 (), p. 59800–59821. External Links: Document Cited by: §I-A. [22] H. Kang, J. Joung, J. Kim, J. Kang, and Y. S. Cho (Sep. 2020) Protect your sky: A survey of counter unmanned aerial vehicle systems. IEEE Access 8 (), p. 168671–168710. External Links: Document Cited by: §I. [23] H. Kim, T. Roh, and B. Shim (Washington, DC, USA, Oct. 2024) Multi-modal sensing-aided beam management for 6G communication systems. In in Proc. IEEE VTC2024-Fall, Vol. , p. 1–5. External Links: Document Cited by: §I. [24] S. Kim, S. Jeong, J. Wu, B. Shim, and M. Z. Win (Dec. 2025) Large multimodal model-based environment-aware channel estimation. IEEE J. Sel. Areas Commun. 43 (12), p. 4059–4075. External Links: Document Cited by: §I. [25] H. Lei, Y. Yan, J. Liu, Q. Han, and Z. Li (Oct. 2024) Hierarchical multi-UAV path planning for urban low altitude environments. IEEE Access 12 (), p. 162109–162121. External Links: Document Cited by: §I-A. [26] R. Levie, Ç. Yapar, G. Kutyniok, and G. Caire (Jun. 2021) RadioUNet: fast radio map estimation with convolutional neural networks. IEEE Trans. Wireless Commun. 20 (6), p. 4001–4015. External Links: Document Cited by: §I-B2, §IV-A. [27] B. Li and J. Chen (Nov. 2024) Radio map-assisted approach for interference-aware predictive UAV communications. IEEE Trans. Wireless Commun. 23 (11), p. 16725–16741. External Links: Document Cited by: §I. [28] C. Li, Z. Dou, and Y. Lin (Dec. 2024) Fast 3-D radio map reconstruction via cross tensor approximation. IEEE Internet Things J. 11 (24), p. 40619–40633. External Links: Document Cited by: §I. [29] H. Li, L. Ding, Y. Wang, and Z. Wang (Aug. 2023) Air-to-ground channel modeling and performance analysis for cellular-connected UAV swarm. IEEE Commun. Lett. 27 (8), p. 2172–2176. External Links: Document Cited by: §I-A, §I-A, §I-A, §I-A. [30] H. Li, Y. Wang, C. Sun, and Z. Wang (Mar. 2024) User-centric cell-free massive MIMO for IoT in highly dynamic environments. IEEE Internet Things J. 11 (5), p. 8658–8675. External Links: Document Cited by: §I-A, §I-A. [31] J. Li, Q. Pan, Z. Wan, P. Zhu, D. Wang, M. Lou, J. Jin, F. Liu, and X. You (Dec. 2023) Low altitude 3-D coverage performance analysis of cell-free RAN for 6G systems. IEEE Trans. Veh. Technol. 72 (12), p. 16163–16176. External Links: Document Cited by: §I-A, §I-A, §I-A. [32] J. Li, C. Zhou, J. Liu, M. Sheng, N. Zhao, and Y. Su (Feb. 2024) Reinforcement learning-based resource allocation for coverage continuity in high dynamic UAV communication networks. IEEE Trans. Wireless Commun. 23 (2), p. 848–860. External Links: Document Cited by: §I. [33] J. Li, L. Yang, W. Hao, I. Ahmad, H. Liu, F. Shu, and D. Niyato (Aug. 2025) Multi-layer transmitting RIS-aided receiver for collaborative jamming and anti-jamming networks. IEEE Trans. Wireless Commun. 24 (8), p. 6518–6534. External Links: Document Cited by: §I. [34] J. Li, L. Yang, C. You, I. Ahmad, P. S. Bithas, M. Di Renzo, and D. Niyato (Dec. 2025) Absorptive RIS-assisted near-field covert communication with fluid antenna systems. IEEE J. Sel. Areas Commun. 44 (), p. 2052–2070. External Links: Document Cited by: §I-A. [35] W. Liu and J. Chen (Sep. 2023) UAV-aided radio map construction exploiting environment semantics. IEEE Trans. Wireless Commun. 22 (9), p. 6341–6355. External Links: Document Cited by: §I. [36] W. Liu and J. Chen (Sep. 2023) UAV-aided radio map construction exploiting environment semantics. IEEE Trans. Wireless Commun. 22 (9), p. 6341–6355. External Links: Document Cited by: §I-A. [37] M. Qian, L. You, X. Xia, and X. Gao (Oct. 2024) On the spectral efficiency of multi-user holographic MIMO uplink transmission. IEEE Trans. Wireless Commun. 23 (10), p. 15421–15434. External Links: Document Cited by: §I-A. [38] D. Ramachandram and G. W. Taylor (Nov. 2017) Deep multimodal learning: A survey on recent advances and trends. IEEE Signal Process Mag. 34 (6), p. 96–108. External Links: Document Cited by: §I-A, §I-B. [39] K. Sato and T. Fujii (Mar. 2017) Kriging-based interference power constraint: Integrated design of the radio environment map and transmission power. IEEE Trans. Cognit. Commun. Networking 3 (1), p. 13–25. External Links: Document Cited by: §IV-F. [40] S. Shao, W. Zhu, and Y. Li (Chongqing, China, Jun. 2022) Radar detection of low-slow-small UAVs in complex environments. In in Proc. IEEE ITAIC 2022, Vol. 10, p. 1153–1157. External Links: Document Cited by: §I. [41] F. Shen, G. Ding, Q. Wu, and Z. Wang (Apr. 2023) Compressed wideband spectrum mapping in 3D spectrum-heterogeneous environment. IEEE Trans. Veh. Technol. 72 (4), p. 4875–4886. External Links: Document Cited by: §I. [42] F. Shen, Z. Wang, G. Ding, K. Li, and Q. Wu (Jan. 2022) 3D compressed spectrum mapping with sampling locations optimization in spectrum-heterogeneous environment. IEEE Trans. Wireless Commun. 21 (1), p. 326–338. External Links: Document Cited by: §I. [43] H. Shimomura, Y. Koda, T. Kanda, K. Yamamoto, T. Nishio, and A. Taya (Las Vegas, NV, USA, Jan. 2023) Vision-aided frame-capture-based CSI recomposition for WiFi sensing: a multimodal approach. In in Proc. IEEE CCNC 2023, Vol. , p. 913–914. External Links: Document Cited by: §I. [44] J. Song, R. He, Z. Zhang, M. Yang, B. Ai, H. Zhang, and R. Chen (Xi’an, China, Jul. 2024) 3D environment reconstruction based on ISAC channels. In Proc. IEEE Int. Conf. Ubiquitous Commun. (Ucom), Vol. , p. 487–491. External Links: Document Cited by: §I-A. [45] H. Sun and J. Chen (Aug. 2024) Integrated interpolation and block-term tensor decomposition for spectrum map construction. IEEE Trans. Signal Process. 72 (), p. 3896–3911. External Links: Document Cited by: §I. [46] H. Sun and J. Chen (Jun. 2024) Energy-modified leverage sampling for radio map construction via matrix completion. IEEE Signal Process Lett. 31 (), p. 1780–1784. External Links: Document Cited by: §I. [47] K. Suto, S. Bannai, K. Sato, K. Inage, K. Adachi, and T. Fujii (Jun. 2021) Image-driven spatial interpolation with deep learning for radio map construction. IEEE Wireless Commun. Lett. 10 (6), p. 1222–1226. External Links: Document Cited by: §I-B2. [48] J. Tang, X. Gao, L. You, D. Shi, J. Yang, X. Xia, X. Zhao, and P. Jiang (Jun. 2025) Massive MIMO-OFDM channel acquisition with time-frequency phase-shifted pilots. IEEE Trans. Commun. 73 (6), p. 4520–4535. External Links: Document Cited by: §I-A. [49] J. Wang, Z. Lin, Q. Zhu, Q. Wu, T. Lan, Y. Zhao, Y. Bai, and W. Zhong (Mar. 2024) 3D spectrum mapping and reconstruction under multi-radiation source scenarios. China Commun. 23 (2), p. 20–34. Cited by: §I. [50] J. Wang, Q. Zhu, Z. Lin, J. Chen, G. Ding, Q. Wu, G. Gu, and Q. Gao (Oct. 2024) Sparse Bayesian learning-based hierarchical construction for 3D radio environment maps incorporating channel shadowing. IEEE Trans. Wireless Commun. 23 (10), p. 14560–14574. External Links: Document Cited by: §I. [51] X. Wang, Q. Zhang, N. Cheng, J. Chen, Z. Zhang, Z. Li, S. Cui, and X. Shen (2025) RadioDiff-3D: A 3D× 3D radio map dataset and generative diffusion based benchmark for 6G environment-aware communication. IEEE Trans. Network Sci. Eng. (), p. 1–18. External Links: Document Cited by: §I. [52] Z. Wen, G. Li, Z. Liu, Y. Li, and S. Han (Chengdu, China, Dec. 2024) A multi-modal learning framework for MIMO channel information acquisition. In in Proc. IEEE ICCC 2024, Vol. , p. 2393–2399. External Links: Document Cited by: §I. [53] S. Woo, J. Park, J. Lee, and I. S. Kweon (Munich, Germany, Sep. 2018) CBAM: Convolutional block attention module. In Proc. Eur. Conf. Comput. Vis. (ECCV), p. 3–19. Cited by: §I-A1, §I-B. [54] D. Wu, Y. Qiu, Y. Zeng, and F. Wen (Dec. 2024) Environment-aware channel estimation via integrating channel knowledge map and dynamic sensing information. IEEE Wireless Commun. Lett. 13 (12), p. 3608–3612. External Links: Document Cited by: §I-A. [55] D. Wu, Y. Zeng, S. Jin, and R. Zhang (May 2024) Environment-aware hybrid beamforming by leveraging channel knowledge map. IEEE Trans. Wireless Commun. 23 (5), p. 4990–5005. External Links: Document Cited by: §I. [56] Q. Wu, F. Shen, Z. Wang, and G. Ding (Oct. 2020) 3D spectrum mapping based on ROI-driven UAV deployment. IEEE Network 34 (5), p. 24–31. External Links: Document Cited by: §I. [57] T. Wu, J. Liu, J. Liu, Z. Huang, H. Wu, C. Zhang, B. Bai, and G. Zhang (Apr. 2022) A novel AI-based framework for AoI-optimal trajectory planning in UAV-assisted wireless sensor networks. IEEE Trans. Wireless Commun. 21 (4), p. 2462–2475. External Links: Document Cited by: §I. [58] C. Xie, L. You, R. Chen, G. He, and X. Gao (Kuala Lumpur, Malaysia, Apr. 2026) MML-based 3D channel fingerprints construction for low-altitude communications. In in Proc. IEEE WCNC 2026, p. 1–6. Cited by: CSI-tuples-based 3D Channel Fingerprints Construction Assisted by MultiModal Learning. [59] C. Xie, L. You, Z. Jin, J. Tang, X. Gao, and X. Xia (Nov. 2025) CF-CGN: channel fingerprints extrapolation for multi-band massive MIMO transmission based on cycle-consistent generative networks. IEEE J. Sel. Areas Commun. 43 (11), p. 3722 – 3736. External Links: Document Cited by: §I, §I-B1. [60] Z. Xin, Y. Liu, J. Xing, J. Huang, J. Bian, Z. Bai, and C. Wang (Feb. 2026) Multimodal fusion-based channel prediction and characterization for mmWave UAV A2G communications. IEEE Trans. Commun. 74 (), p. 5089–5104. External Links: Document Cited by: §I. [61] C. Xu, X. Liao, J. Tan, H. Ye, and H. Lu (Apr. 2020) Recent research progress of unmanned aerial vehicle regulation policies and technologies in urban low altitude. IEEE Access 8 (), p. 74175–74194. External Links: Document Cited by: §I. [62] Y. Yang, F. Gao, C. Xing, J. An, and A. Alkhateeb (Jul. 2021) Deep multimodal learning: Merging sensory data for massive MIMO channel prediction. IEEE J. Sel. Areas Commun. 39 (7), p. 1885–1898. External Links: Document Cited by: §I. [63] X. Ye, Y. Mao, X. Yu, S. Sun, L. Fu, and J. Xu (2025) Integrated sensing and communications for low-altitude economy: A deep reinforcement learning approach. IEEE Trans. Wireless Commun. (), p. 1–1. External Links: Document Cited by: §I. [64] Y. Ye, L. You, J. Wang, H. Xu, K. Wong, and X. Gao (Jan. 2024) Fluid antenna-assisted MIMO transmission exploiting statistical CSI. IEEE Commun. Lett. 28 (1), p. 223–227. External Links: Document Cited by: §I-A. [65] H. B. Yilmaz, T. Tugcu, F. Alagöz, and S. Bayhan (Dec. 2013) Radio environment map as enabler for practical cognitive radio networks. IEEE Commun. Mag. 51 (12), p. 162–169. External Links: Document Cited by: §I. [66] C. Yin, Z. Xiao, X. Cao, X. Xi, P. Yang, and D. Wu (Apr. 2018) Offline and online search: UAV multiobjective path planning under dynamic urban environment. IEEE Internet Things J. 5 (2), p. 546–558. External Links: Document Cited by: §I-A. [67] K. Yin, S. Fang, F. Chu, and Y. Fan (Dec. 2024) Compressed tensor completion: Approach for UAV-aided 3-D radio map construction. IEEE Internet Things J. 11 (24), p. 40516–40531. External Links: Document Cited by: §I. [68] P. Zeng and J. Chen (Seoul, Korea, Republic of, May 2022) UAV-aided joint radio map and 3D environment reconstruction using deep learning approaches. In in Proc. IEEE ICC 2022, Vol. , p. 5341–5346. External Links: Document Cited by: §I. [69] Y. Zeng and X. Xu (Jun. 2021) Toward environment-aware 6G communications via channel knowledge map. IEEE Wireless Commun. 28 (3), p. 84–91. External Links: Document Cited by: §I. [70] H. Zhang, Z. Han, G. C. Alexandropoulos, and N. H. Tran (Apr. 2022) Special issue on aerial access networks for 6G. Journal of Communications and Networks 24 (2), p. 121–124. External Links: Document Cited by: §I. [71] S. Zhang, S. Jiang, W. Lin, Z. Fang, K. Liu, H. Zhang, and K. Chen (Apr. 2025) Generative AI on SpectrumNet: An open benchmark of multiband 3-D radio maps. IEEE Trans. Cognit. Commun. Networking 11 (2), p. 886–901. External Links: Document Cited by: §I. [72] L. Zhao, Z. Fei, X. Wang, J. Huang, Y. Li, and Y. Zhang (Mar. 2025) IMNet: interference-aware channel knowledge map construction and localization. IEEE Wireless Commun. Lett. 14 (3), p. 856–860. External Links: Document Cited by: §I. [73] P. Zhen, B. Zhang, Y. Xu, Z. Chen, H. Wang, and D. Guo (Aug. 2022) Radio environment map construction based on Gaussian process with positional uncertainty. IEEE Wireless Commun. Lett. 11 (8), p. 1639–1643. External Links: Document Cited by: §IV-F. [74] Y. Zhu, L. You, Q. Kong, G. Seco-Granados, and X. Gao (Feb. 2025) Robust precoding for massive MIMO LEO satellite localization systems. IEEE Trans. Veh. Technol. 74 (2), p. 3434–3438. External Links: Document Cited by: §I-A.