Paper deep dive
Large Multimodal Model-Based Environment-Aware Mobility Management
Seokhyun Jeong, Sangmok Shin, Seungnyun Kim, Jiao Wu, Byonghyo Shim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/14/2026, 3:20:08 AM
Summary
The paper introduces LMM-EMM, a large multimodal model-based environment-aware mobility management framework for ultra-dense networks. It leverages LMMs to process RGB-D images and wireless measurements, extracting environmental context to predict channel capacities and make proactive handover decisions, significantly outperforming conventional deep learning and 5G NR approaches.
Entities (7)
Relation Signals (6)
LMM-EMM → uses → Large Multimodal Models
confidence 95% · The proposed framework leverages LMMs to process multimodal sensing data and extract rich contextual information on the surrounding environment.
LMM-EMM → outperforms → Conventional Deep Learning Approaches
confidence 93% · Simulation results demonstrate that the proposed scheme achieves substantial channel capacity improvements over conventional deep learning (DL)-based approaches, achieving about 45% improvement compared to the 5G NR mobility management scheme.
LMM-EMM → learns → Channel Capacity Map
confidence 92% · Using the extracted environmental information, the proposed scheme learns the intrinsic mapping from UE and SBS positions to channel capacity, referred to as channel capacity map (CCM).
LMM-EMM → optimizes → Proactive Handover
confidence 91% · Based on the predicted channel capacities, we determine proactive handover decisions maximizing the cumulative channel capacities.
Large Multimodal Models → processes → RGB-D Images
confidence 90% · By leveraging LMMs, the proposed scheme extracts contextual information on the surrounding environments from RGB-D images to capture user equipment (UE) mobility patterns.
Channel Capacity Map → maps → User Equipment to Small Base Stations
confidence 88% · The CCM characterizes the achievable channel capacity as a function of the UE position, the SBS position, and environmental factors (e.g., reflectors).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recently, large language models (LLMs) have been successfully adopted in various fields, including wireless communications, robotics, and autonomous vehicles, owing to their outstanding adaptability and reasoning abilities. Despite their huge potential, the application of LLMs for mobility management is relatively scarce since it requires not only analyzing wireless measurements but also predicting dynamic user trajectories and making real-time handover decisions across densely deployed small base stations (SBSs). In this paper, we propose an environment-aware mobility management scheme based on large multimodal models (LMMs), which extend capabilities of LLMs to process multimodal sensing data. By leveraging LMMs, the proposed scheme extracts contextual information on the surrounding environments from RGB-D images to capture user equipment (UE) mobility patterns and identify signal reflections and blockages caused by static reflectors and dynamic obstacles. Using the extracted environmental information, the proposed scheme learns the intrinsic mapping from UE and SBS positions to channel capacity, referred to as channel capacity map (CCM), from which future channel capacities along UE trajectories are predicted. Based on the predicted channel capacities, we determine proactive handover decisions maximizing the cumulative channel capacities. Simulation results demonstrate that the proposed scheme achieves substantial channel capacity improvements over conventional deep learning (DL)-based approaches.
Tags
Links
- Source: https://arxiv.org/abs/2607.09795v1
- Canonical: https://arxiv.org/abs/2607.09795v1
Trouble viewing inline? Open PDF directly →
Full Text
101,372 characters extracted from source content.
Expand or collapse full text
Large Multimodal Model-Based Environment-Aware Mobility Management Seokhyun Jeong, , Sangmok Shin, , Seungnyun Kim, , Jiao Wu, , and Byonghyo Shim Received 23 September, 2025; revised 13 March, 2026 and 17 May, 2026; accepted 3 July, 2026. This work was supported in part by the National Research Foundation (NRF) of Korea under Grant RS-2022-NR070834 and 2022M3C1A3099336, the Institute of Information & Communications Technology Planning & Evaluation(IITP)-ITRC(Information Technology Research Center) grant funded by the Korea government(MSIT)(IITP-2026-2021-0-02048) and the National Research Foundation, Singapore and Infocomm Media Development Authority under its Communications and Connectivity Bridging Funding Initiative. An earlier version of this paper was presented in part at the IEEE Global Communications Conference (GLOBECOM), Taipei, Taiwan, December 2025. (Corresponding author: Byonghyo Shim.) Seokhyun Jeong, Sangmok Shin, and Byonghyo Shim are with the Department of Electrical and Computer Engineering and the Institute of New Media and Communications, Seoul National University, Seoul 08826 Republic of Korea (e-mail: shjeong@islab.snu.ac.kr; smshin@islab.snu.ac.kr; bshim@snu.ac.kr). Seungnyun Kim is with the Information Systems Technology and Design (ISTD) pillar, Singapore University of Technology and Design, Singapore 487372 (e-mail: snkim94@mit.edu). Jiao Wu is with the Computer, Electrical and Mathematical Sciences and Engineering Division, King Abdullah University of Science and Technology, Thuwal 23955-6900 Saudi Arabia (e-mail: jiao.wu@kaust.edu.sa). Abstract Recently, large language models have been successfully adopted in various fields, including wireless communications, robotics, and autonomous vehicles, owing to their outstanding adaptability and reasoning abilities. Despite their huge potential, the application of LLMs for mobility management is relatively scarce since it requires not only analyzing wireless measurements but also predicting dynamic user trajectories and making real-time handover decisions across densely deployed small base stations. In this paper, we propose an environment-aware mobility management scheme based on large multimodal models, which extend capabilities of LLMs to process multimodal sensing data. By leveraging LMMs, the proposed scheme extracts contextual information on the surrounding environments from RGB-D images to capture user equipment (UE) mobility patterns and identify signal reflections and blockages caused by static reflectors and dynamic obstacles. Using the extracted environmental information, the proposed scheme learns the intrinsic mapping from UE and SBS positions to channel capacity, referred to as channel capacity map (CCM), from which future channel capacities along UE trajectories are predicted. Based on the predicted channel capacities, we determine proactive handover decisions maximizing the cumulative channel capacities. Simulation results demonstrate that the proposed scheme achieves substantial channel capacity improvements over conventional deep learning (DL)-based approaches. I Introduction LLM, such as ChatGPT and Gemini, are attracting great attention these days for their powerful capabilities to solve unlimited complex tasks [1]. Indeed, owing to the vast network size and extensive pretraining on diverse datasets, LLMs can solve a bewildering variety of problems with minimal examples (i.e., few-shot learning), or even without any example at all (i.e., zero-shot learning). Benefited by advanced natural language processing (NLP) capabilities, LLMs can handle task instructions in the form of text prompts. Recently, functionalities of LLMs have been extended to simultaneously process multimodal data, including images, audio, and video [2]. This extended version, called LMM, can extract rich contextual information on the surrounding environment and thus can handle complex real-world tasks. For this reason, LMMs can be used for a variety of applications, including smart factories, robotics, autonomous driving, and personal assistance [3]. While there are recent research efforts integrating LMM into wireless systems, papers addressing LMM-based mobility management are relatively scarce [4, 5]. This is because mobility management involves far more diversified functionalities than those required by resource allocation, such as predicting the dynamic movements of UEs and solving intricate cell association problems. In fact, the primary goal of mobility management is to associate each UE with an appropriate base station (BS) to ensure a reliable connection (i.e., cell association) [6]. In particular, in ultra-dense networks, wireless channels between a UE and a SBS fluctuate rapidly as the UE moves through densely deployed SBSs (e.g., pico and femto cells) and thus a delay in the cell association process will affect the reliable connection considerably (see Fig. 1) [7, 8, 9]. Figure 1: Visualization of mobility management in UDN systems. To guarantee seamless connectivity in UDNs, a handover technique transferring the ongoing cell association to a nearby SBS is needed. In 4G Long-Term Evolution (LTE) and 5G New Radio (NR), handovers are triggered based on UE-reported reference signal received power (RSRP) measurements when the target SBS (T-SBS) provides a higher RSRP than the serving SBS (S-SBS) over four consecutive reports [10]. While this approach has been used for many years, it is not quite effective for 5G and upcoming 6G networks, since it induces a considerable handover delay [6]111Note that the S-SBS can initiate the handover only after receiving a sequence of RSRP reports from the UE. In 5G NR, the time interval between adjacent RSRP reports is 4040-120120 ms, and thus the total handover latency can reach up to 360 ms. Clearly, this time far exceeds the channel coherence time of 5G NR. For example, in 5G NR with carrier frequency of fc=3.5f_c=3.5\,GHz and UE speed of v=10v=10\,m/s, the large-scale fading channel coherence time is 40916πcvfc≈145.140 916π cvf_c≈ 145.1 ms [11].. To mitigate this issue, several approaches for reducing the handover delay by proactively performing handover have been suggested. In [12, 13, 14], proactive mobility management techniques that utilize machine learning (ML) and DL models have been studied. In [15, 16, 17], deep reinforcement learning (DRL)-based mobility management techniques that use policy optimization for the handover control have been proposed. The potential weakness of these techniques is that they rely exclusively on wireless measurements (e.g., RSRP), which hinders rapid adaptation to abrupt channel variations caused by the dynamic environment (e.g., human/vehicle movements and sudden environmental changes) [18]. Recently, integrated sensing and communications (ISAC)-based handover techniques that leverage the sensing data (e.g., images, radar signals, and LiDAR point clouds) to predict potential link blockages have gained attention [19, 20, 21, 22, 23]. While the ISAC-based approaches can incorporate environmental features to some extent, these techniques focus primarily on detecting line-of-sight (LoS) path and overlook the impact of multipath propagation (e.g., reflection and scattering), limiting their effectiveness in rich scattering environments [24]. An aim of this paper is to put forth an environment-aware mobility management framework empowered by the LMM technology. The proposed framework, referred to as large multimodal model-based environment-aware mobility management (LMM-EMM), preemptively associates the UE with a proper SBS by leveraging environmental context and UE mobility patterns extracted from multimodal data (e.g., RGB-D images and wireless measurements). Key ingredient in achieving this mission is LMM, a generative artificial intelligence (AI) model specialized in extracting correlated features across multiple modalities and performing high-level reasoning. In the proposed framework, LMM learns the UE movement pattern and the reflection geometry between the UE and SBSs by analyzing the correlations between the physical environments (e.g., buildings, road structures, and dynamic obstacles) obtained from bird’s eye-view (BEV) maps, SBS-view sensing images, and also the wireless measurements. These features are then exploited to construct a fundamental mapping between the UE position and the channel capacity, henceforth dubbed as channel capacity map (CCM). Using CCM along with the predicted UE trajectory, the channel capacity can be estimated without requiring any real-time measurements, thereby facilitating fast and accurate environment-aware handovers. The main contributions of this paper are summarized as follows: • We propose an LMM-based environment-aware mobility management framework that ensures fast and reliable connectivity. Unlike conventional schemes that rely primarily on wireless signal measurements, LMM-EMM exploits the multimodal reasoning capability of LMM to incorporate a richer contextual understanding of the surrounding environment, thereby supporting more predictive and robust mobility decisions. • We develop an environment-aware channel capacity estimation technique that captures the reflection geometry of the propagation channel. To this end, we construct the CCM, an end-to-end mapping from the UE position to channel capacity. A main step of the CCM construction is analyzing the environmental context and inferring the underlying propagation characteristics, and we use LMM for this purpose. By leveraging CCM, we can accurately predict the future channel capacity for each SBS and make proactive handover decisions that reduce latency and enhance link reliability. • From extensive simulations on realistic UDN environments, we show that the proposed large multimodal model-based environment-aware mobility management (LMM-EMM) achieves a significant capacity gain over the conventional mobility management techniques. Specifically, LMM-EMM achieves about 45% improvement in channel capacity compared to the 5G NR mobility management scheme. Even when compared with the long short-term memory (LSTM)-based and DRL-based approaches, LMM-EMM yields 21% and 15% capacity gains, respectively. The rest of this paper is organized as follows: In Section I, we explain the UDN system model and briefly introduce conventional mobility management techniques. In Section I, we present the overall architecture of LMM-EMM. In Section IV, we discuss practical issues. In Section V, we demonstrate the numerical results and conclude the paper in Section VI. Notations: Bold upper and lower case symbols denote matrices and vectors, respectively. Sets are denoted by calligraphic font, for example, X. Superscript (⋅)T(·)^T denotes the transpose. x(t)x^(t) and x(t1:t2)x^(t_1:t_2) denote the value of x at time slot t and x(t1),x(t1+1),⋯,x(t2)\x^(t_1),x^(t_1+1),·s,x^(t_2)\ where t1<t2t_1<t_2, respectively. 1⊗2X_1 _2 denotes the Kronecker product of matrices 1X_1 and 2X_2. ‖2||x||_2 and []n[x]_n denote the Euclidean norm and the nnth element of a vector x, respectively. [⋅]E[·] denotes an expectation of a random variable. |||X| denotes the number of elements in a set X. ⋃i _iX_i denotes the union of the sets iX_i over all i. I Ultra-dense Network System In this section, we provide a brief overview of UDN systems and discuss conventional mobility management techniques. I-A Downlink UDN System Model We consider a millimeter wave (mmWave) multiple-input single-output (MISO) UDN system consisting of M SBSs equipped with N=Nx×NyN=N_x× N_y uniform planar array (UPA) antennas and a UE equipped with a single antenna. The set of SBS indices is denoted as ℳ=1,2,…,MM=\1,2,…c,M\. The SBSs are connected to the macro base station (MBS) through backhaul links to share transmit data and control signals. We consider a global coordinate system (GCS) where the position vectors of the mmth SBS and the UE at the ttth time slot are sbs,m(t)=[xsbs,mysbs,mzsbs,m]Tp_sbs,m^(t)=[x_sbs,m\,\,y_sbs,m\,\,z_sbs,m]^T and ue(t)=[xue(t)yue(t)zue(t)]Tp_ue^(t)= [x_ue^(t)\,\,y_ue^(t)\,\,z_ue^(t) ]^T, respectively. We also consider an orthogonal frequency division multiplexing (OFDM) system with S subcarriers, a carrier frequency of fcf_c, and a system bandwidth of B. The frequency of the ssth subcarrier is fs=fc+(s−S2)BSf_s=f_c+(s- S2) BS for s=1,⋯,Ss=1,·s,S. Assuming the equal power allocation among the data streams, the MISO channel capacity R(t)(m)R^(t)(m) between the mmth SBS and the UE at the time slot t is R(t)(m) R^(t)(m) = = B∑s=1Slog2(1+PtNσn2‖m(t)[s]‖22) B _s=1^S _2 (1+ P_tN _n^2 \|h_m^(t)[s] \|_2^2 ) -1.99997pt (1) where m(t)[s]∈ℂNh_m^(t)[s] ^N is the downlink channel vector from the mmth SBS to the UE at the ssth subcarrier and the ttth time slot, PtP_t is the SBS transmission power, and σn2 _n^2 is the noise power [25]. Since R(t)(m)R^(t)(m) depends on m(t)[s]s=1S\h_m^(t)[s]\_s=1^S, it varies dynamically as the UE moves [26]. Thus, if the link quality of S-SBS is lower than the predefined threshold due to the movement of the UE, the UE must search for a T-SBS and initiate a handover process. In our work, we adopt a block-fading geometric multipath channel model where the channel remains constant within a time slot with duration τs _s. Each time slot is divided into the channel estimation (CE) period τce _ce, during which resource blocks are dedicated to transmitting reference signals (e.g., CSI-RS), and the data transmission period τdt _dt. Under this channel model, the downlink channel vector m(t)[s]h_m^(t)[s] is expressed as the sum of an LoS path (denoted by l=0l=0) and L−1L-1 non-line-of-sight (NLoS) paths: m(t)[s] _m^(t)[s] = = βm,0(t)αm,0(t)e−j2πfsτm,0(t)(θm,0(t),ϕm,0(t)) _m,0^(t) _m,0^(t)e^-j2π f_s _m,0^(t)a( _m,0^(t), _m,0^(t)) (2) +∑l=1L−1βm,l(t)αm,l(t)e−j2πfsτm,l(t)(θm,l(t),ϕm,l(t)) + _l=1^L-1 _m,l^(t) _m,l^(t)e^-j2π f_s _m,l^(t)a( _m,l^(t), _m,l^(t)) -1.99997pt where βm,l(t) _m,l^(t) is the large-scale fading coefficient (LSFC), αm,l(t)∼(0,1) _m,l^(t) (0,1) is the small-scale fading coefficient (SSFC), τm,l(t) _m,l^(t) is the time delay, and θm,l(t)∈[0,2π) _m,l^(t)∈[0,2π) and ϕm,l(t)∈[0,π] _m,l^(t)∈[0,π] are the azimuth angle of departure (AoD) and zenith angle of departure (ZoD) of the llth path between the mmth SBS and the UE, respectively. Also, (θm,l(t),ϕm,l(t))∈ℂNa ( _m,l^(t), _m,l^(t) ) ^N is the SBS array steering vector given by (θm,l(t),ϕm,l(t)) ( _m,l^(t), _m,l^(t) ) = = [1⋯ej2πdλ(Nx−1)cosθm,l(t)sinϕm,l(t)] [1\,·s\,e^j 2π dλ(N_x-1) _m,l^(t) _m,l^(t) ] (3) ⊗[1⋯ej2πdλ(Ny−1)sinθm,l(t)sinϕm,l(t)] \,\, [1\,·s\,e^j 2π dλ(N_y-1) _m,l^(t) _m,l^(t) ] where d and λ are the antenna spacing and signal wavelength, respectively. Specifically, βm,l[s] _m,l[s] is expressed as βm,l[s] _m,l[s] = = Γm,l4πfsτm,le−(8π2σm,l2fs2cos2(ψm,l)c2+cτm,lkabs(fs)2) _m,l4π f_s _m,le^- ( 8π^2 _m,l^2f_s^2 ^2( _m,l)c^2+ c _m,lk_abs(f_s)2 )\! (4) where kabs(fs)k_abs(f_s), c, Γm,l _m,l, σm,l _m,l, and ψm,l _m,l are the molecular absorption coefficient, the speed of light, the Fresnel coefficient, the roughness coefficient, and the angle of incidence, respectively [27]. These parameters are determined by the UE and SBS positions uep_ue and sbs,mp_sbs,m, as well as the reflection surface ℰm,l=∣m,lT=bm,lE_m,l=\x _m,l^Tx=b_m,l\ (‖m,l‖2=1\|n_m,l\|_2=1) [28]: τm,l _m,l = = 1c‖m−2(γm,l−bm,l)m,l‖2 1c\|d_m-2( _m,l-b_m,l)n_m,l\|_2 (5) ψm,l _m,l = = arccos(|γm,l+δm,l|‖m‖22+4γm,lδm,l) ( | _m,l+ _m,l| \|d_m\|_2^2+4 _m,l _m,l ) (6) Γm,l _m,l = = cosψm,l−ϵ−sin2ψm,lcosψm,l−ϵ+sin2ψm,l _m,l- ε- ^2 _m,l _m,l- ε+ ^2 _m,l -1.99997pt (7) where m=ue−sbs,md_m=p_ue-p_sbs,m, γm,l=m,lTue−bm,l _m,l=n_m,l^Tp_ue-b_m,l, δm,l=m,lTsbs,m−bm,l _m,l=n_m,l^Tp_sbs,m-b_m,l, and ϵε is the dielectric permittivity of the reflector. By concatenating m(t)[s]h_m^(t)[s] for all subcarriers, we obtain the frequency domain channel matrix m(t)∈ℂN×SH_m^(t) ^N× S: m(t) _m^(t) = = [m(t)[1]m(t)[2]⋯m(t)[S]] [h_m^(t)[1]\,h_m^(t)[2]\,·s\,h_m^(t)[S] ] (8) = = ∑l=0L−1βm,l(t)αm,l(t)(θm,l(t),ϕm,l(t))T(τm,l(t)) _l=0^L-1 _m,l^(t) _m,l^(t)a ( _m,l^(t), _m,l^(t) )b^T ( _m,l^(t) ) (9) where (τm,l(t))∈ℂSb( _m,l^(t)) ^S is the phase shift vector of the OFDM subcarriers given by (τm,l(t))=[e−j2πf1τm,l(t)e−j2πf2τm,l(t)⋯e−j2πfSτm,l(t)]T. ( _m,l^(t))= [e^-j2π f_1 _m,l^(t)\ e^-j2π f_2 _m,l^(t)\ ·s\ e^-j2π f_S _m,l^(t) ]^T. (10) For brevity, we define the collection of the geometric channel parameters m,l(t)G_m,l^(t) of the llth path between the mmth SBS and UE at the ttth time slot as m,l(t)=βm,l(t),αm,l(t),τm,l(t),ϑm,l(t),φm,l(t),θm,l(t),ϕm,l(t).G_m,l^(t)= \ _m,l^(t), _m,l^(t), _m,l^(t), _m,l^(t), _m,l^(t), _m,l^(t), _m,l^(t) \. (11) Then, we can express m(t)H_m^(t) as a function of m,l(t)l=1L\G_m,l^(t)\_l=1^L as m(t)=(m,0(t))+∑l=1L−1(m,l(t)) _m^(t)=P (G_m,0^(t) )+ _l=1^L-1P (G_m,l^(t) ) (12) where (m,l(t))∈ℂN×SP(G_m,l^(t)) ^N× S denotes the llth multipath component of m(t)H_m^(t) corresponding to m,l(t)G_m,l^(t): (m,l(t))=βm,l(t)αm,l(t)(θm,l(t),ϕm,l(t))T(τm,l(t)).P(G_m,l^(t))= _m,l^(t) _m,l^(t)a ( _m,l^(t), _m,l^(t) )b^T ( _m,l^(t) ). (13) Note that geometric channel parameters vary continuously when reflectors remain unchanged. However, changes in reflectors can cause abrupt transitions, which lead to discontinuities in channel capacity [29]. Remark 1. As shown in Fig. 2, the channel capacity R(t)(m)R^(t)(m) for each SBS is piecewise continuous over time. The discontinuity arises from the discontinuity of reflection points on different reflectors. Figure 2: Piecewise continuity of channel capacity with UE movement. I-B Channel Capacity Map In this subsection, we introduce the concept of the CCM, which characterizes the achievable channel capacity as a function of the UE position, the SBS position, and environmental factors (e.g., reflectors). For notational simplicity, the time index t is omitted without loss of generality. Lemma 1. When the number of transmit antennas N is sufficiently large, the channel capacity Rideal(m)R_ideal(m) can be expressed as a function of LSFCs βm,l[s]l=0L−1\ _m,l[s]\_l=0^L-1. That is, Rideal(m)=B∑s=1Slog2(1+Ptσn2(βm,0[s]+∑l=1L−1βm,l[s])).R_ideal(m)=B _s=1^S _2 (1+ P_t _n^2 ( _m,0[s]+ _l=1^L-1 _m,l[s] ) ). (14) Proof. See Appendix A. ∎ In practice, moving obstacles (e.g., vehicles and pedestrians) can intermittently block the LoS link between UE and SBS. Such dynamic blockages significantly impact the achievable channel capacity due to the pronounced power disparity between the LoS and NLoS components [27]. Remark 2. The achievable channel capacity is a function of static channel capacity and LoS indicator as R(m)=clos(m)Rideal(m)+(1−clos(m))Rnlos(m)R(m)=c_los(m)R_ideal(m)+(1-c_los(m))R_nlos(m) (15) where Rnlos(m)R_nlos(m) is the capacity of NLoS channel given by Rnlos(m)=B∑s=1Slog2(1+Ptσn2∑l=1L−1βm,l[s]).R_nlos(m)=B _s=1^S _2 (1+ P_t _n^2 _l=1^L-1 _m,l[s] ). (16) Also, clos(m)c_los(m) is the LoS indicator defined as clos(m)=1if the LoS for SBS m exists0otherwise.c_los(m)= cases1&\!\!if the LoS for SBS $m$ exists\\ 0&\!\!otherwise. cases (17) From Equation (14) and (16), one can see that Rideal(m)R_ideal(m) and Rnlos(m)R_nlos(m) are functions of LSFCs βm,l[s]l=0L−1\ _m,l[s]\_l=0^L-1. Moreover, as shown in Equation (4)-(7), the LSFC βm,l[s] _m,l[s] is determined by the UE position uep_ue, SBS position sbs,mp_sbs,m, and the reflection surface ℰm,l=∣m,lT=bm,lE_m,l=\x _m,l^Tx=b_m,l\. Therefore, Rideal(m)R_ideal(m) and Rnlos(m)R_nlos(m) can be directly obtained from uep_ue, sbs,mp_sbs,m, and ℰ=∪m,lℰm,lE= _m,l\,E_m,l. Remark 3. The static channel capacities Rideal(m)R_ideal(m) and Rnlos(m)R_nlos(m) are fully characterized by three components: the UE position uep_ue, the SBS position sbs,mp_sbs,m, and the reflector information ℰE, through CCM fccmf_ccm: (Rideal(m),Rnlos(m)) (R_ideal(m),R_nlos(m) ) = = fccm(ue,sbs,m,ℰ). f_ccm(p_ue,p_sbs,m,E). (18) Remarks 2 and 3 show that predicting R(m)R(m) requires the UE position uep_ue, the SBS positions sbs,m∣m∈ℳ\p_sbs,m m \, the reflector information ℰE, and the LoS indicators clos(m)∣m∈ℳ\c_los(m) m \. Figure 3: Illustration of service interruption during the handover. I-C Handover Process The handover process in 5G NR consists of three main phases: preparation, execution, and completion [30]. The preparation phase begins when the RSRP of T-SBS exceeds that of S-SBS by a configured offset (i.e., event A3 [30]). At this point, the UE initiates the handover process by sending the radio resource control (RRC) measurement report to S-SBS. If S-SBS continues to receive these reports throughout the designated period (i.e., time-to-trigger), it sends the handover request to T-SBS, followed by a handover acknowledgment from T-SBS to S-SBS. Next, in the execution phase, S-SBS transfers the UE identification and security information (e.g., encryption algorithms) to T-SBS, along with any downlink data that has not yet been delivered to the UE. Then, the UE establishes a new connection to T-SBS via the random access. Lastly, in the completion phase, the core network finalizes the handover process by redirecting the UE data path to T-SBS. Figure 4: Overall procedure of the proposed LMM-EMM. Note that during the handover process, data transmission is temporarily interrupted since the UE cannot receive data from T-SBS until UE authentication and synchronization are completed (see Fig. 3). Thus, the effective channel capacity Reff(t)(m(t−1:t))R_eff^(t)(m^(t-1:t)) at the ttth time slot is expressed as a weighted sum of channel capacities during and after the handover process (i.e., R(t)(m(t))(1−ho(m(t−1:t))R^(t)(m^(t))(1- 1_ho(m^(t-1:t)) and R(t)(m(t))R^(t)(m^(t))): Reff(t)(m(t−1:t)) R_eff^(t) (m^(t-1:t) ) = = 1τs(τhoR(t)(m(t))(1−ho(m(t−1:t))) 1 _s ( _hoR^(t) (m^(t) )(1- 1_ho(m^(t-1:t))) +(τs−τho−τce)R(t)(m(t))) +( _s- _ho- _ce)R^(t) (m^(t) ) ) = = τs−τceτs(1−μho(m(t−1:t)))R(t)(m(t)) _s- _ce _s(1-μ 1_ho(m^(t-1:t)))R^(t) (m^(t) ) where τho _ho is the handover process time, and μ is the handover cost coefficient defined as μ=τhoτs−τce.μ= _ho _s- _ce. (21) Also, ho(m(t−1:t))∈0,1 1_ho(m^(t-1:t))∈\0,1\ is the binary handover indicator at the ttth time slot defined as ho(m(t−1:t))=1if m(t−1)≠m(t)0otherwise. 1_ho(m^(t-1:t)) -1.99997pt= -1.99997pt cases1&if $m^(t-1)≠ m^(t)$\\ 0&otherwise. cases (22) Note that in UDN environments, UEs may undergo frequent handovers due to the dense deployment of SBSs, which consumes a considerable portion of the total transmission time and thus degrades throughput significantly. I-D Conventional Mobility Management Techniques A major drawback of the 5G NR handover is that the signal quality deterioration and decision latency are considerable since handover is triggered only after the RSRP drop has been reported. This issue is even more pronounced in UDN scenarios where handover occurs frequently due to the reduced cell coverage and increased number of SBSs. To overcome the limitations of the reactive handover mechanism, proactive handover techniques that trigger handover before the signal quality degradation have been proposed [14]. Essence of this approach is to forecast the future RSRP for all candidate SBSs using historical RSRP measurements and then pick the SBS that maximizes the throughput. For the RSRP prediction, recurrent neural network (RNN) architectures such as LSTM and gated recurrent unit (GRU) have been widely employed [13, 31]. While conventional proactive handover techniques are effective to some extent, they exclusively rely on radio-frequency (RF) signals so that they fall short in capturing abrupt changes in the channel (e.g., blockages caused by dynamic obstacles). Note that RF signals capture gradual variations in signal quality, so it is very difficult to distinguish whether signal quality variations are transient (e.g., sudden blockages) or persistent (e.g., continuous path loss changes due to UE movement) by merely checking the received signals [18]. To address this issue, approaches that leverage sensing data (e.g., RGB images, radar signals, and LiDAR point clouds) obtained from various sensors (e.g., camera, radar, and LiDAR) have been proposed recently [19, 20, 21, 22, 23]. These techniques focus primarily on identifying the LoS component of the channel so that they suffer a degradation of the handover quality in NLoS scenarios. I-E Proactive Handover Decision Problem Formulation Effective mobility management must ensure seamless execution of handovers while avoiding unnecessary ones (e.g., ping-pong effects) caused by channel variations [32]. To this end, we formulate a long-term proactive handover problem that not only responds to instantaneous degradations but also seeks to maximize channel capacity over extended periods. Essence of the proactive handover problem P is to determine the SBS indices m(T:T+Tp)m^(T:T+T_p) for the time slots T to T+TpT+T_p, with an objective to maximize the cumulative channel capacity: : P: maxm(T:T+Tp) -3.00003pt _m^(T:T+T_p) ∑t=T+TpReff(t)(m(t−1:t)) _t=T^T+T_pR_eff^(t)(m^(t-1:t)) (23a) s.t. Reff(t)(m(t−1:t))≥Rmin(t)∀t=T,⋯,T+Tp -11.99998ptR_eff^(t) (m^(t-1:t) )≥ R_min^(t) 5.0pt∀ t=T,·s,T+T_p (23b) where Rmin(t)R_min^(t) is the minimum capacity requirement to ensure the quality of service (QoS) of the UE at the ttth time slot. The salient feature of P is that the objective function accounts for the cumulative channel capacity over Tp+1T_p+1 future time slots, which is clearly distinct from conventional approaches maximizing the instantaneous capacity (i.e., R(T)(m(T))R^(T)(m^(T))) exclusively. Using the cumulative channel capacity as a performance metric, we can reduce redundant handovers caused by instantaneous channel fluctuations. To determine the optimal SBS indices for future time slots, we need to know the future channel capacities (i.e., R(T+t)(m(T+t))t=0Tp∣m∈ℳ \\R^(T+t)(m^(T+t))\_t=0^T_p m \). Obviously, due to the causality issue, future channel capacities cannot be directly observed and need to be predicted instead. Unfortunately, conventional approaches predicting future channel capacities based on temporal correlation (e.g., Kalman filter and RNNs) are not so effective, in particular for urban environments, due to the discontinuity of channel capacity (see Remark 1). I Large Multimodal Model-Based Environment-Aware Mobility Management The primary goal of LMM-EMM is to proactively associate the UE with proper SBSs that maximize channel capacity. To this end, we leverage the multimodal reasoning capabilities of LMM. Specifically, we first extract the UE mobility patterns and scattering geometry from the multimodal sensing data. Unlike conventional DL-based approaches that only capture spatial correlations of the channel without any contextual information, LMM-EMM exploits the road geometry and surrounding environment to infer the behavioral intention of the UE (e.g., turning or accelerating) and perceive signal reflection patterns. Next, using extracted features, we construct CCM, a spatial mapping between the UE position and the channel capacity, used for the estimation of the future channel capacity along the predicted UE trajectory. Furthermore, by analyzing RGB-D images with LMM, we predict potential path blockages and then refine the channel capacity, thereby facilitating optimal SBS selection even under high mobilities. A key feature of LMM-EMM is that LMM extracts environmental knowledge (e.g., the spatial layout of roads and buildings) from multimodal data and reuses this shared understanding across trajectory prediction, channel capacity estimation, and blockage prediction. Unlike conventional DL approaches that learn independent input-output mappings for each task without sharing physical knowledge, LMM-EMM leverages a common environmental interpretation to produce mutually coherent predictions, thereby achieving more reliable mobility management performance. The procedure of LMM-EMM consists of four steps (see Fig. 4). 1. UE trajectory prediction: We predict the future UE trajectory ue(T:T+Tp)p_ue^(T:T+T_p) using historical UE positions ue(T−Tw:T−1)p_ue^(T-T_w:T-1) and the BEV map ℐbevI_bev. 2. Static channel capacity estimation via CCM: By utilizing CCM, we estimate the static channel capacities Rideal(T:T+Tp)(m),Rnlos(T:T+Tp)(m)∣m∈ℳ \R_ideal^(T:T+T_p)(m),R_nlos^(T:T+T_p)(m) m \ from the predicted UE trajectories. 3. Blockage prediction and dynamic channel capacity refinement: We predict path blockages from dynamic obstacles by tracking obstacle states from SBS-view sensing images, and then derive the future achievable channel capacities R(T:T+Tp)(m)∣m∈ℳ \R^(T:T+T_p)(m) m \. 4. Proactive handover optimization: We determine future SBS indices m(T:T+Tp)m^(T:T+T_p) by solving P in (23) using dynamic programming (DP). R(T:T+Tp)(m)∣m∈ℳ \R^(T:T+T_p)(m) m \ obtained in step 3 is used in this process. I-A LMM-Based UE Trajectory Prediction In the UE trajectory prediction step, we estimate the future UE trajectory using historical UE positions and the BEV map:222UE position can be acquired from various positioning techniques leveraging global navigation satellite system (GNSS), sensing, and wireless measurements[33, 34]. ue(T:T+Tp) _ue^(T:T+T_p) = = ftraj(ue(T−Tw:T−1),ℐbev) f_traj (p_ue^(T-T_w:T-1),I_bev ) (24) where ftrajf_traj is the UE trajectory prediction function and TwT_w is the observation time window. Also, ℐbevI_bev is the BEV map image representing the static road topology, lane structures, and surrounding environments, which can be captured at the rooftop of a building or a satellite [35]. It is important to note that the UE movement is constrained by physical environments, such as road and building areas, which are denoted by RA_R and BA_B, respectively [36]. For example, vehicles move along lanes with their trajectories confined within road boundaries (i.e., Prob(ue(t)=∣∉R)=0Prob(p_ue^(t)=p _R)=0), and pedestrians cannot traverse obstacles like buildings and walls (i.e., Prob(ue(t)=∣∈B)=0Prob(p_ue^(t)=p _B)=0) (see Fig. 5). In LMM-EMM, we exploit BEV images to incorporate RA_R and BA_B in UE trajectory prediction. This approach filters out impossible position predictions and also provides contextual cues for recognizing mobility patterns (e.g., turning at intersections) [37]. Specifically, we learn the conditional probability of the future UE positions given past positions using the next token prediction capability of LMM in an autoregressive manner: Prob( ( ue(T:T+Tp)∣ue(T−Tw:T−1),R,B) _ue^(T:T+T_p) _ue^(T-T_w:T-1),A_R,A_B ) (25) =∏t=0TpProb(ue(T+t)∣ue(T−Tw:T+t−1),R,B). = _t=0^T_pProb (p_ue^(T+t) _ue^(T-T_w:T+t-1),A_R,A_B ). Note that LMM is trained with the autoregressive language modeling objective (i.e., next token prediction), where the model learns to predict the most likely subsequent token given all preceding tokens in a sequence [38]. This next token prediction mechanism of LMM is particularly beneficial in LMM-EMM since the next token prediction can be interpreted as predicting the future UE position based on the historical UE trajectory. An intriguing feature of LMM-EMM is the instruction learning strategy that trains the model using natural language instructions specifying not only the input, task, and desired output, but also the task-relevant context. In our case, the instruction prompts explicitly specify which environmental information should be emphasized in the multimodal data, allowing LMM to focus on the features most relevant to each task (see Fig. 6). For the UE trajectory prediction task, the instruction prompts are structured as follows: • Input description: “The input is a sequence of historical UE positions from (T−Tw)(T-T_w)th time slot to (T−1)(T-1)th time slot ue(T−Tw:T−1)p_ue^(T-T_w:T-1) and the BEV map image ℐbevI_bev.” • Environment description: “The UE movement is constrained by physical environments such as road boundaries RA_R and buildings BA_B. Ensure the UE trajectory is in feasible regions within the given environment.” • Task instruction: “Find the UE trajectory for the next Tp+1T_p+1 time slots ue(T:T+Tp)p_ue^(T:T+T_p).” Given the instruction prompts, the LMM generates the response prompt as “The predicted UE trajectory for the next Tp+1T_p+1 time slots is ue(T:T+Tp)p_ue^(T:T+T_p).” Figure 5: Illustration of LMM-based trajectory prediction using BEV images. I-B LMM-Based Static Channel Capacity Estimation via CCM In the channel capacity estimation step, we learn the CCM fccmf_ccm in (18), which maps the UE and SBS positions (ue(t),sbs,m)(p_ue^(t),p_sbs,m) to the static channel capacities (Rideal(t)(m),Rnlos(t)(m))(R_ideal^(t)(m),R_nlos^(t)(m)), parametrized by reflector information ℰE. Two main challenges in learning fccmf_ccm are: 1) the discontinuous variation of static channel capacities with UE mobility (see Remark 1), and 2) the difficulty of directly acquiring reflector information ℰE, as accurately modeling and extracting ℰE in real environments is highly complex. To address this issue, we learn a surrogate function f~ccm f_ccm that takes the BEV map image, which implicitly encodes ℰE: (Rideal(t)(m),Rnlos(t)(m)) (R_ideal^(t)(m),R_nlos^(t)(m) ) = = f~ccm(ue(t),sbs,m,ℐbev). f_ccm (p_ue^(t),p_sbs,m,I_bev ). (26) Furthermore, we exploit the multimodal reasoning capability of LMMs to interpret reflector information from the BEV map and to pinpoint locations where channel capacity discontinuities may occur. In doing so, LMM-EMM integrates environmental context into the surrogate model, facilitating reliable approximation of fccmf_ccm and precise estimation of static channel capacities at arbitrary UE and SBS locations. The instruction prompts for the LMM-based channel capacity estimation are structured as follows: • Input description: “The input is the future UE position ue(t)p_ue^(t) and the BEV map ℐbevI_bev.” • Environment description: “Note that the reflector positions in the provided image affect the channel capacity by reflecting or blocking the propagation paths.” • Task instruction: “Estimate the static channel capacities (Rideal(t)(m),Rnlos(t)(m))∣m∈ℳ\(R_ideal^(t)(m),R_nlos^(t)(m)) m \ from the UE position.” Given the instruction prompts, we obtain the estimated channel capacities from the LMM response: “The estimated static channel capacities are (Rideal(t)(m),Rnlos(t)(m))∣m∈ℳ \(R_ideal^(t)(m),R_nlos^(t)(m)) m \.” Figure 6: Examples of instruction prompt and LMM response. I-C LMM-Based Blockage Prediction and Dynamic Channel Capacity Refinement In the blockage prediction and dynamic channel capacity refinement step, we first estimate the LoS indicators in (17) clos(T:T+Tp)(m)∣m∈ℳ \c_los^(T:T+T_p)(m) m \ from SBS-view sensing images. Using the estimated LoS indicators, together with the static channel capacity estimates Rideal(T:T+Tp)(m),Rnlos(T:T+Tp)(m)∣m∈ℳ \R_ideal^(T:T+T_p)(m),R_nlos^(T:T+T_p)(m) m \, we then compute the achievable channel capacity R(T:T+Tp)(m)∣m∈ℳ \R^(T:T+T_p)(m) m \ according to (15). Specifically, let obs,m=1,…,Nobs,mN_obs,m=\1,…c,N_obs,m\ be the set of objects near the SBS m. For each object i∈obs,mi _obs,m, we denote the center position of the object i in the camera coordinate system as obs,i(t)=[xobs,i(t)yobs,i(t)dobs,i(t)]q_obs,i^(t)= [x_obs,i^(t)\,\,y_obs,i^(t)\,\,d_obs,i^(t) ] (27) where xobs,i(t)x_obs,i^(t), yobs,i(t)y_obs,i^(t), and dobs,i(t)d_obs,i^(t) are x-coordinate, y-coordinate, and depth of the object centroid, respectively. Then, the 2D rectangular region i(t)⊂ℝ2A_i^(t) ^2 in the image occupied by the object i at time t is given by i(t)= _i^(t)= [xobs,i(t)−12wobs,i(t),xobs,i(t)+12wobs,i(t)] [x_obs,i^(t)- 12w_obs,i^(t),x_obs,i^(t)+ 12w_obs,i^(t) ] (28) ×[yobs,i(t)−12hobs,i(t),yobs,i(t)+12hobs,i(t)] × [y_obs,i^(t)- 12h_obs,i^(t),y_obs,i^(t)+ 12h_obs,i^(t) ] where wobs,i(t)w_obs,i^(t) and hobs,i(t)h_obs,i^(t) are the width and height of the bounding box of the object i in the camera coordinate system, respectively. Using this, clos(t)(m)c_los^(t)(m) is defined as clos(t)(m)=][c]l′s0if(x¯_ue^(t), y¯_ue^(t)) ∈ A_i^(t)andd_obs, i^(t) ¡ d_ue^(t)∃i ∈N_obs,m1otherwisec_los^(t)(m) -1.99997pt= -3.00003pt \ IEEEeqnarraybox[][][c]l s0& -10.00002ptif -3.99994pt$( x_ue^(t), y_ue^(t)) -1.99997pt∈ -1.99997pt A_i^(t)$ -6.00006ptand -3.99994pt$d_obs, i^(t) -1.99997pt< -1.99997pt d_ue^(t)$ -5.0pt$∃ i -1.00006pt∈ -1.99997ptN_obs,m$\\ 1& -10.00002ptotherwise IEEEeqnarraybox . (29) where due(t)d_ue^(t)=\,= ‖ue(t)−sbs,m(t)‖2\|p_ue^(t)-p_sbs,m^(t)\|_2 and (x¯ue(t),y¯ue(t))( x_ue^(t), y_ue^(t)) is the x- and y-coordinates of UE in the camera coordinate system at time slot t (see Fig. 7). Note that (x¯ue(t),y¯ue(t))( x_ue^(t), y_ue^(t)) can be obtained from the GCS UE position ue(t)p_ue^(t) through rotation as (x¯ue(t),y¯ue(t))=(xr(t)zr(t),yr(t)zr(t))( x_ue^(t), y_ue^(t))= ( x_r^(t)z_r^(t), y_r^(t)z_r^(t) ) where [xr(t)yr(t)zr(t)]T [x_r^(t)\,\,y_r^(t)\,\,z_r^(t) ]^T = = (ue(t)−sbs,m). (p_ue^(t)-p_sbs,m ). (30) Here, R is the camera rotation matrix given by [6] =[sinθ−cosθ0−sinϕcosθ−sinϕsinθcosϕcosϕcosθcosϕsinθsinϕ]R= bmatrix θ&- θ&0\\ - φ θ&- φ θ& φ\\ φ θ& φ θ& φ bmatrix (31) where θ and ϕφ denote the yaw and pitch angles of the camera. To determine clos(t)(m)c_los^(t)(m), the SBS m needs to know the future positions of objects obs,i(T:T+Tp)q_obs,i^(T:T+T_p) for i∈obs,mi _obs,m. Analogous to UE trajectory prediction, these positions are inferred from historical positions obs,i(T−Tw:T−1)q_obs,i^(T-T_w:T-1). Unlike UE position data, which can be reported from UE via uplink channel, historical object positions are difficult to obtain since no communication links exist between the SBS and the objects [39]. As a remedy, we utilize object detection (OD) techniques to accurately extract obs,i(T−Tw:T−1)q_obs,i^(T-T_w:T-1) from SBS-view RGB-D images ℐsbs,mI_sbs,m, and then use LMM to predict the future object trajectories: obs,i(T:T+Tp) _obs,i^(T:T+T_p) = = fobs(obs,i(T−Tw:T−1),ℐsbs,m) f_obs (q_obs,i^(T-T_w:T-1),I_sbs,m ) (32) where fobsf_obs is the obstacle trajectory prediction function. The instruction prompts for the LMM-based obstacle position prediction are designed as follows: • Input description: “The input is a sequence of the historical obstacle positions in the camera coordinate system obs,i(T−Tw:T−1)q_obs,i^(T-T_w:T-1) and the SBS-view sensing image ℐsbsI_sbs.” • Environment description: “In the camera coordinate system, obstacle movement is constrained by environmental elements including road boundaries for vehicles and immobile structures such as buildings that define non-navigable regions observed in ℐsbsI_sbs.” • Task instruction: “Find the positions of dynamic obstacles for the next (Tp+1)(T_p+1) time slots obs,i(T:T+Tp)q_obs,i^(T:T+T_p).” The LMM response prompts generated from the instruction prompts are “The positions of dynamic obstacles for the next Tp+1T_p+1 time slots are obs,i(T:T+Tp)q_obs,i^(T:T+T_p).” Once we obtain the predicted static channel capacities and LoS indicators Rideal(T:T+Tp)(m),Rnlos(T:T+Tp)(m),clos(T:T+Tp)(m)∣m∈ℳ, \R_ideal^(T:T+T_p)(m),R_nlos^(T:T+T_p)(m),c_los^(T:T+T_p)(m) m \, for all SBSs, we acquire the achievable channel capacity R(T:T+Tp)(m)∣m∈ℳ \R^(T:T+T_p)(m) m \ using (15). Figure 7: Illustration of LMM-based dynamic blockage prediction with SBS-view sensing images. I-D Proactive Handover Optimization In the proactive handover optimization step, we find the optimal SBS indices m^(T:T+Tp) m^(T:T+T_p) in future time slots by solving the optimization problem P in (23) based on the estimated channel capacities R(T:T+Tp)(m)∣m∈ℳ \R^(T:T+T_p)(m) m \. One intuitive yet naive approach to solve P is an exhaustive search, which evaluates all possible sequences of SBSs. Although exhaustive search guarantees the optimal solution, it becomes computationally prohibitive in UDN because the number of candidates grows exponentially as MTp+1M^T_p+1. To obtain a tractable solution of P, we adopt a DP-based approach, which solves complex problems by decomposing them into a series of simpler subproblems [40]. Owing to the cumulative structure of the capacity, the optimal solution at each time slot can be recursively derived from the solutions of previous time slots. In fact, the computational complexity of the proposed DP-based method is TpM2T_pM^2, which is significantly lower than that of the exhaustive search. Specifically, let g(T+t,m(T+t))g(T+t,m^(T+t)) be the maximum cumulative capacity when the UE is connected to the SBS m(T+t)m^(T+t) at time slot T+tT+t: g(T+t,m(T+t))=maxm(T:T+t−1)[ g (T+t,m^(T+t) )= -10.20006pt _m^(T:T+t-1)\, [ ∑t′=0tReff(T+t′)(m(T+t′−1:T+t′)) _t =0^tR_eff^(T+t ) (m^(T+t -1:T+t ) ) (33) ×min(m(T+t′−1:T+t′))] × 1_min (m^(T+t -1:T+t ) ) ] where min(m(t−1:t)) 1_min(m^(t-1:t)) is the capacity compliance indicator defined as min(m(t−1:t))=1if Reff(t)(m(t−1:t))≥Rmin(t)−∞if Reff(t)(m(t−1:t))<Rmin(t). 1_min(m^(t-1:t))= cases1&if $R_ eff^(t) (m^(t-1:t) )≥ R_min^(t)$\\ -∞&if $R_ eff^(t) (m^(t-1:t) )<R_min^(t)$. cases (34) Note that g(T+t,m(T+t))g(T+t,m^(T+t)) can be decomposed as g(T+t,m(T+t))=maxm(T+t−1)[g(T+t−1,m(T+t−1)) g(T+t,m^(T+t))= -1.99997pt _m^(T+t-1) [g (T+t-1,m^(T+t-1) ) +Reff(T+t)(m(T+t−1:T+t))min(m(T+t−1:T+t))]. 10.00002pt+R_eff^(T+t) (m^(T+t-1:T+t) ) 1_min (m^(T+t-1:T+t) ) ]. (35) Based on this recurrence formula, starting from initial g(T,m(T))=Reff(T)(m(T−1:T))min(m(T−1:T))g(T,m^(T))=R_eff^(T)(m^(T-1:T)) 1_min(m^(T-1:T)), we sequentially compute g(T+t,m(T+t))t=1Tp∣m∈ℳ \\g(T+t,m^(T+t))\_t=1^T_p m \, which are subsequently stored in the DP table. After that, m^(T:T+Tp) m^(T:T+T_p) are retrieved in reverse chronological order from the computed DP table. First, m^(T+Tp) m^(T+T_p) is obtained as m^(T+Tp) m^(T+T_p) = = argmaxm(T+Tp)g(T+Tp,m(T+Tp)). argmax_m^(T+T_p)\,\,g (T+T_p,m^(T+T_p) ). (36) Then, m^(t) m^(t) are determined recursively for t=Tp−1,⋯,0t=T_p-1,·s,0: m^(T+t) m^(T+t) = = argmaxm(T+t)[g(T+t,m(T+t)) argmax_m^(T+t)\,\, [g(T+t,m^(T+t)) (37) +Reff(T+t+1)(m(T+t:T+t+1))min(m(T+t:T+t+1))]. -15.00002pt+R_eff^(T+t+1)(m^(T+t:T+t+1)) 1_min(m^(T+t:T+t+1)) ]. Finally, the SBSs execute handovers according to the determined indices m^(T:T+Tp) m^(T:T+T_p). IV Practical Issues In this section, we discuss several practical issues related to the implementation of LMM-EMM, focusing on end-to-end latency and challenges associated with its deployment. IV-A End-to-end Latency The key requirement for reliable handover decision-making is that the total end-to-end latency of LMM-EMM remain shorter than the timescale over which the channel varies sufficiently to affect the ergodic channel capacity. As shown in Equation (38), the ergodic channel capacity is determined by slow-varying channel parameters, such as angles and path losses, whose coherence time TcT_c (i.e., beam coherence time) is given by [41] Tc=Dvcosθcos−1(ψ2logζ+1)T_c= Dv θ ^-1(ψ^2 ζ+1) (38) where D is the communication distance, θ is the beam angle, v is the speed of the UE, ψ is the beamwidth, and ζ∈[0,1]ζ∈[0,1] is the threshold ratio of the received power to its peak value. For example, when a UE moves at 25km/h25\,km/h with D=70mD=70\,m, θ=60∘θ=60 , and ζ=0.8ζ=0.8, and the SBS is equipped with N=32N=32 antennas such that ψ≈4N=0.125radψ≈ 4N=0.125\,rad, then the beam coherence time is Tc≈840msT_c≈ 840\,ms. In our implementation, the latency of each LMM-EMM module on an NVIDIA L40S GPU is summarized as follows: • UE trajectory prediction: T1=215T_1=215\,ms (2020\,ms for UE location report and 195195\,ms for LMM inference) • Static channel capacity estimation: T2=195T_2=195\,ms (195195\,ms for LMM inference) • Blockage prediction: T3=220T_3=220\,ms (2525\,ms for multimodal image capture and processing, and 195195\,ms for LMM inference) • Channel capacity refinement, report, and handover: T4=55T_4=55\,ms (2525\,ms for channel capacity refinement and report, and 3030\,ms for DP-based handover decision) Since blockage prediction can be performed in parallel with UE trajectory prediction and static channel capacity estimation, as it does not depend on others’ outputs, these three modules are completed within max(T1+T2,T3)=410 (T_1+T_2,T_3)=410\,ms. Therefore, the total end-to-end latency of LMM-EMM is Ttot=max(T1+T2,T3)+T4=465T_tot= (T_1+T_2,T_3)+T_4=465\,ms, which is substantially shorter than the beam coherence time Tc≈840T_c≈ 840\,ms (see Fig. 8). Moreover, this latency is expected to decrease further as GPU hardware continues to evolve, leaving considerable room for further optimization in future implementations. Figure 8: End-to-end latency vs. beam coherence time. IV-B Deployment Issues IV-B1 Noisy Sensing In real-world settings, sensing images are often degraded by environmental factors such as rain, snow, fog, or dust. To demonstrate the robustness of the proposed LMM-EMM framework to imperfections in sensing data, we evaluate the handover performance of LMM-EMM under noisy sensing conditions caused by environmental factors. Under this setting, LMM-EMM achieves the average channel capacity of 4.5114.511\,bps/Hz, which corresponds to only a 4% reduction from the noise-free case. Furthermore, when the denoising and restoration techniques for adverse visual conditions are applied, LMM-EMM achieves the average channel capacity of 4.7244.724\,bps/Hz (see Fig. 9(a)) [42]. This corresponds to only a 1%1\% decrease compared to 4.7704.770\,bps/Hz in the noise-free scenario (see Fig. 18). IV-B2 FoV and Resolution Issues One practical issue of LMM-EMM is that the SBS-view camera has a limited field of view (FoV) and resolution, which may restrict its ability to capture the entire scene. While it is true that a single camera has a limited FoV, the entire coverage area can be monitored by deploying multiple cameras, each covering a different sectorized region, similar to antenna sectorization. For example, three RGB-D cameras, each with a FoV of 120∘120 , can cover the entire area surrounding the SBS. Regarding the image resolution issue, objects that contribute to LoS blockage (e.g., car, bus, and truck) can still be effectively identified even in low-resolution images because they are typically large enough to occupy multiple pixels in the captured images. Furthermore, in our work, we crop and resize the image region around the predicted UE position to magnify relevant objects and improve detection performance (see Fig. 9(b)). (a) Illustration of the noisy sensing image and object detection after applying the denoising network. (b) Illustration of the cropped image region around the predicted UE position. Figure 9: Examples of sensing images used for blockage prediction under noisy and low-resolution scenarios. IV-B3 Availability of Cameras and BEV Maps One might be concerned that LMM-EMM can operate only when RGB-D cameras are installed and BEV maps are available at the SBSs. However, with ISAC emerging as a key component of upcoming 6G, SBSs are expected to be equipped with diverse sensing modalities including RGB-D cameras. Moreover, BEV maps can be readily obtained from publicly available sources such as OpenStreetMap [43] and Google Maps. If the serving SBS cannot collect the sensing data due to the absence of sensors, it can request a neighboring SBS that can observe the UE and surrounding obstacles to estimate the blockage state. Specifically, the serving SBS transmits the estimated UE trajectory to the neighboring SBS. The neighboring SBS then constructs the LoS line segment between the serving SBS and the UE, and uses its own RGB-D camera to detect obstacles and represent them as 3D boxes. By checking whether the LoS line segment intersects the 3D boxes, the blockage state can be determined [39]. If no such neighbor exists, the blockage state can be inferred from past channel measurements (e.g., signal-to-interference-plus-noise ratio (SINR)), and the inferred past state can be used as a proxy for the future LoS state. Such temporal consistency is reasonable because obstacles that cause blockage (e.g., buses, vehicles) typically have non-negligible volume, so the blockage state does not change significantly over a short time interval. IV-B4 Robustness to Reflector Material Mismatch A possible issue in LMM-EMM is that incorrect knowledge about the material properties may introduce errors in the channel capacity estimation. However, since material properties generally do not change significantly over time, the associated errors can be controlled via fine-tuning with real datasets. Even if the reflector material is incorrectly characterized, LMM-EMM can still maintain reasonable performance because the propagation geometry (e.g., LoS paths, reflection angles) remains unchanged despite the material mismatch. To validate this, we conduct an ablation study in which the reflector material in the validation dataset is changed from glass to concrete. We observe a performance degradation of 5.2% compared with the original case, which demonstrates that the proposed technique still maintains reasonable performance under incorrect knowledge of reflector materials (see Fig. 12). IV-B5 Synchronization According to the IEEE 1588-v2 protocol, inter-cell time synchronization errors within 0.1μ0.1\, may exist [44]. To evaluate the impact of these synchronization errors, we measure the average channel capacity under such timing offsets. We observe that the average channel capacity remains 4.770bps/Hz4.770\,bps/Hz, identical to the perfectly synchronized case of 4.770bps/Hz4.770\,bps/Hz, indicating no performance degradation. This is because the synchronization error is sufficiently small that the UE displacement during this interval is negligible, and thus the corresponding channel can be considered as unchanged. Figure 10: Average channel capacity of LMM-EMM with reflector material mismatch. Figure 11: Average channel capacity vs. SBS position error. Figure 12: Cosine similarity between approximated channel and actual channel. IV-B6 SBS Position Error Inaccurate SBS positions may affect blockage detection performance because the estimated LoS path can deviate from the actual one. To evaluate the impact of SBS position error, we plot the average channel capacity of the proposed LMM-EMM as a function of the SBS position error. We observe that under an SBS position error of 11 m, the average channel capacity reaches 4.7584.758 bps/Hz, which is only 0.30.3% lower than that under perfect SBS position knowledge (see Fig. 12). Moreover, since SBS positions are typically fixed after deployment, they can be calibrated over time, so persistent location errors are unlikely in practical systems. IV-B7 Soft Blockage In practical environments, a soft blockage model may arise due to partial blockage and diffraction, which could introduce discrepancies in the channel capacity estimation. However, in the considered mmWave scenario, the impact of such effects is limited for two main reasons. First, diffraction and refraction are generally weak at mmWave frequencies due to the short wavelength and high penetration loss [45]. As a result, the corresponding path components contribute much less to the received power than the reflected paths. Their effect on channel capacity variation is thus limited, and neglecting them is a reasonable approximation in the considered mmWave scenario. Second, the impact of partial blockage on the channel capacity variation is expected to be limited in the considered mmWave scenario. This is because objects that induce partial blockage (e.g., pedestrians) typically yield a much smaller path gain βm,l[s] _m,l[s] than specular reflectors (e.g., buildings), as they usually have a significantly larger roughness coefficient σm,l _m,l (see Equation (4))333For instance, a path partially blocked by a pedestrian with a roughness coefficient of σm,l=3m _m,l=3\,m yields a βm,l[s] _m,l[s] that is approximately 1/4671/467 of that for a typical glass building with σm,l=1m _m,l=1\,m.. Consequently, the channel capacity variation caused by soft blockages is relatively minor compared with the dominant effects of LoS and reflected paths. To validate the adopted LoS/NLoS approximation combined with the reflection model, we computed the cosine similarity between the approximated and actual channels at various roughness coefficients of the objects. We observe that the cosine similarity exceeds 0.930.93 even when the roughness coefficient is 3m3\, m, which confirms that the approximated model is sufficiently accurate for channel capacity comparison and handover decision-making (see Fig. 12). V Experimental Results V-A Simulation Setup We consider the UDN system where M=14M=14 SBSs equipped with Nx×Ny=8×4N_x× N_y=8× 4 UPA antennas serve a UE equipped with a single antenna. The antenna orientations of SBS and UE are (0∘,−10∘,0∘)(0 ,-10 ,0 ) and (0∘,0∘,0∘)(0 ,0 ,0 ), respectively. The UE moves along the road at a speed of v=25v=25 km/h and it turns left, turns right, or continues straight ahead at each intersection with probabilities of 25%, 25%, and 50%, respectively. The carrier frequency and bandwidth are fc=28f_c=28 GHz and B=100B=100 MHz, respectively. The observation time window TwT_w and the prediction length TpT_p are set to 55. We consider the signal-to-noise ratio (SNR) of 1515 dB and the handover interruption time of τho=36 _ho=36 ms. Also, an RGB-D camera having a resolution of 1024×7681024× 768 and a FoV of 120∘120 is mounted on the top side of the SBS. For the pretrained LMM, we adopt the LLaVA-1.5-7B model [2], which is one of the most prevalent open-source LMMs. We conduct experiments on an Intel Xeon Gold 6326 CPU server equipped with a 16-core CPU and an NVIDIA L40S GPU with 48 GB of memory. We compare the mobility management performance of the proposed LMM-EMM with four conventional techniques and the ideal scheme with perfect channel state information (CSI): 1. Perfect CSI-based scheme: an upper bound of average capacity assuming all channel capacities and blockage indicators are perfectly known. 2. DRL-based scheme [17]: A proactive handover scheme that uses deep Q-learning (DQN) to optimize the handover parameters (i.e., handover offset and time-to-trigger (T)) with handover-type prediction. • State: RSRP and SINR from both the serving and target SBSs, and the UE’s distance to the target SBS, speed, and moving direction. • Action: handover margin and time-to-trigger. • Reward: negative weighted sum of the ping-pong, too-early, and too-late handover occurrence rates. 3. Vision-aided scheme [19]: A proactive handover scheme that employs SBS-view camera images and a YOLOv3 object detection model to predict the LoS path blockages and proactively trigger handover. 4. LSTM-based scheme [13]: A proactive handover scheme that uses LSTM to predict the future channel capacity in an autoregressive manner and performs handover based on the predicted capacity. 5. 5G NR handover [10]: A reactive handover scheme where the SBS determines the handover decisions based on consecutive RSRP reports from the UE. Figure 13: Visualizations of the simulation environment on real-world scenarios. Figure 14: Comparison of CCMs generated by LMM-EMM and the conventional DL-based scheme. V-B Multimodal Dataset Generation and LMM Fine-Tuning V-B1 Multimodal Dataset Generation To construct a realistic dataset for UDN scenarios, we develop a comprehensive pipeline that integrates 3D environmental modeling, wireless channel generation, and sensory data acquisition. First, we reconstructed a 3D urban environment resembling a real-world deployment using building layouts extracted from OpenStreetMap, which has been reported to provide sub-meter positional accuracy [43]. Based on this 3D environment, we performed ray tracing and wireless channel generation using NVIDIA Sionna RT [46], which provides physically consistent channel modeling based on ray tracing and 3GPP TR 38.901 [47]. Since 3GPP TR 38.901 is widely adopted to emulate realistic propagation characteristics (e.g., reflection and blockage), the generated wireless channels closely resemble practical deployments. Therefore, although the data are simulated, the dataset maintains high physical fidelity and supports reliable performance evaluation for practical deployment scenarios. To validate performance in general wireless environments, we generate two additional real-world scenarios: 1) an urban environment characterized by high building density and tall structures, and 2) a suburban environment with moderate building density and low heights (less than 1010 m) (see Fig. 13). V-B2 LMM Fine-Tuning For LMM fine-tuning, we employ supervised fine-tuning (SFT) in which the LMM is trained on input–output pairs generated by a real-world wireless simulator (NVIDIA Sionna RT [46]) to minimize the negative log-likelihood (NLL) loss J for the response tokens ∗ w^*: =−1Ntoken∑i=1Ntokenlog([i]wi∗).J=- 1N_token _i=1^N_token ([ π_i]_w_i^* ). (39) By adjusting the model parameters through SFT, LMM-EMM can capture fine-grained mobility behaviors and channel variations, thereby enhancing mobility management performance. One major issue in SFT is the excessive computational burden of updating a massive number of network parameters (e.g., 7 billion parameters in LLaVA-1.5-7B), which results in extended training times and considerable resource demands. We address this issue by adopting the low-rank adaptation (LoRA) technique, which adjusts only a small fraction of parameters relevant to the task [48]. Specifically, we select a subset of parameters ∈ℝX×YW ^X× Y and update it with the product of two low-rank matrices ∈ℝX×rA ^X× r and ∈ℝr×YB ^r× Y: ′=+W =W+AB -3.00003pt (40) where ′∈ℝX×YW ^X× Y and r denote the updated model parameters and the rank of low-rank matrices, respectively. We use a LoRA rank of r=16r=16 and a learning rate of 1×10−41× 10^-4, which is decayed by a factor of 0.1 every 10 steps. The generated dataset for fine-tuning contains Ntrain=16 000N_train=16\,000, Nval=2 000N_val=2\,000, and Ntest=2 000N_test=2\,000 samples for training, validation, and test, respectively. Figure 15: Visual comparison of predicted UE trajectories for LMM-EMM and the DL-based trajectory prediction. V-C Simulation Results In Fig. 14, we visualize CCM generated by LMM-EMM and the conventional DL-based scheme [49]. To assess the quality of the generated CCM, we evaluate the normalized mean square error (NMSE) against the ground-truth CCM as shown in Appendix B. The results indicate that the proposed LMM-EMM generates a more accurate CCM than the DL-based scheme, achieving a 6.36.3 dB improvement in NMSE. Additionally, in Fig. 15, we compare the predicted UE trajectory of LMM-EMM and the DL-based trajectory prediction. We observe that the trajectory predicted by LMM-EMM closely follows the ground-truth path and accurately captures direction changes and UE movement. Overall, these findings indicate that LMM-EMM understands environmental features well, including intersections and obstacle locations, so that it can facilitate reliable handover decisions. Figure 16: Average channel capacity as a function of SNR. Figure 17: Average channel capacity as a function of the UE speed. Figure 18: Average channel capacity as a function of the number of SBSs. In Fig. 18, we evaluate the channel capacity as a function of SNR. We see that the proposed LMM-EMM achieves a significant improvement in the channel capacity over the conventional mobility management techniques. For instance, at an SNR of 15 dB, LMM-EMM achieves more than 45% gain in channel capacity compared to the 5G NR handover scheme. Even when compared to the LSTM-based technique, LMM-EMM achieves a 21% gain in channel capacity. This is because the proposed LMM-EMM can properly comprehend the environmental information to predict the accurate channel capacity from multiple SBSs. In Fig. 18, we evaluate the handover performance as a function of the UE speed. Due to the accurate estimation of the future trajectory of the UE and corresponding channels from the environment, the proposed LMM-EMM outperforms conventional mobility management techniques in terms of channel capacity. For example, when the speed of the UE is 2525 km/h, LMM-EMM achieves more than 50% and 23% higher channel capacity compared to the 5G NR handover and LSTM-based technique, respectively. Since LMM-EMM intelligently analyzes road geometry to accurately estimate the UE trajectory and corresponding channel capacities, it achieves robust performance across diverse mobility scenarios. To evaluate the performance of LMM-EMM in various deployment scenarios, we plot the channel capacity of the UE as a function of the number of SBSs in Fig. 18. Since each UDN system contains a different number of SBSs, validation across diverse numbers of SBSs is crucial to demonstrate generalization performance. We see that the channel capacity of the proposed LMM-EMM is much higher than that of conventional techniques. For example, when the number of SBSs is 1212, LMM-EMM achieves more than 40% higher capacity compared to the 5G NR handover. By exploiting CCM, LMM-EMM can accurately estimate the channel capacity of each SBS and associate the appropriate SBS with the UE, regardless of the number of SBSs. In Fig. 21, we evaluate the handover performance with various handover interruption times. The results demonstrate that the proposed LMM-EMM significantly outperforms conventional mobility management techniques. For instance, when τho=54 _ho=54 ms, LMM-EMM achieves more than 38% capacity gain over the 5G NR handover. Even when compared to the vision-based technique, LMM-EMM achieves a 17% increase in channel capacity. Notably, LMM-EMM exhibits minimal degradation in average capacity as the handover interruption time increases. This indicates that LMM-EMM can avoid redundant handovers through accurate recognition of transient channel variations caused by dynamic obstacles. To assess the generalization performance of LMM-EMM, we evaluate the channel capacity across diverse scenarios, including urban and suburban areas, in Fig. 21. We observe that LMM-EMM exhibits superior performance against conventional techniques across all wireless scenarios. For example, LMM-EMM achieves 30% and 26% higher channel capacity than the DRL-based technique in urban and suburban settings, respectively. This generalization performance results from the multimodal reasoning capability of LMM, which accurately interprets the spatial distribution of reflectors and the UE mobility patterns in diverse environments. Figure 19: Average channel capacity as a function of handover interruption time. Figure 20: Average channel capacity as a function of SNR in various wireless scenarios. Figure 21: Average channel capacity as a function of antenna array size. To verify that LMM-EMM remains effective even with smaller antenna arrays, in Fig. 21, we evaluate the channel capacity under various antenna array sizes. We observe that although the instantaneous channel experiences fluctuations due to small-scale fading, these fluctuations rarely change the handover decision. For example, even when the antenna array size is 88, LMM-EMM achieves 4.0174.017\,bps/Hz, which corresponds to 90%90\% of the channel capacity achieved by the optimal handover with perfect CSI. These results indicate that the proposed LMM-EMM retains most of the ideal handover gain even when the number of antennas is small. In Fig. 24, we evaluate the average channel capacity performance of LMM-EMM across different architectural layouts, including highway and indoor scenarios. To this end, we perform few-shot fine-tuning using only 600600 samples collected from those environments. We observe that at an SNR of 1515\,dB, LMM-EMM still achieves 1010% and 77% channel capacity gain over the conventional DRL-based approach in highway and indoor scenarios, respectively (see Fig. 24). This strong adaptability of LMM-EMM stems from the ability of LMM to capture the scenario-invariant propagation mechanisms (e.g., reflection and blockage) that are shared across different environments. While architectural layouts vary across environments, these variations primarily affect site-specific geometry, not the fundamental propagation physics. This means that LMM can quickly understand the propagation characteristics in new environments and rapidly adapt to unseen scenarios. TABLE I: Impact of input UE trajectory errors on prediction accuracy and communication performance. Noise variance 00\,m 0.50.5\,m 11\,m 1.51.5\,m 22\,m !20!white UE trajectory prediction RMSE 0.2520.252\,m 0.4900.490\,m 0.6750.675\,m 0.8440.844\,m 0.9640.964\,m !0!white Static channel capacity estimation NMSE −9.627-9.627\,dB −9.327-9.327\,dB −8.549-8.549\,dB −7.542-7.542\,dB −6.379-6.379\,dB !20!white Blockage prediction accuracy 95.5%95.5\% 95.4%95.4\% 94.9%94.9\% 93.7%93.7\% 92.2%92.2\% !0!white Average channel capacity 4.7704.770\,bps/Hz 4.7294.729\,bps/Hz 4.6574.657\,bps/Hz 4.5564.556\,bps/Hz 4.3714.371\,bps/Hz !20!white 10% worst case channel capacity 2.4902.490\,bps/Hz 2.4792.479\,bps/Hz 2.4402.440\,bps/Hz 1.9871.987\,bps/Hz 1.7981.798\,bps/Hz !0!white 20% worst case channel capacity 2.7442.744\,bps/Hz 2.7362.736\,bps/Hz 2.6702.670\,bps/Hz 2.3632.363\,bps/Hz 2.1812.181\,bps/Hz !20!white Handover failure rate 4.894.89\,% 5.985.98\,% 6.526.52\,% 8.428.42\,% 9.249.24\,% !0!white Service interruption probability 6.616.61\,% 6.816.81\,% 6.906.90\,% 7.257.25\,% 7.827.82\,% TABLE I: Ablation study on the impact of UE trajectory prediction, CCM estimation, and blockage prediction error Error source Amount of error UE trajectory Channel capacity Average prediction RMSE estimation NMSE channel capacity !20!white No noise addition - 0.2520.252\,m −9.627-9.627\,dB 4.7704.770\,bps/Hz !0!white Input UE trajectory Noise variance = 11\,m 0.6750.675\,m −9.061-9.061\,dB 4.6594.659\,bps/Hz !20!white UE trajectory prediction Noise variance = 11\,m 1.0831.083\,m −5.853-5.853\,dB 4.6324.632\,bps/Hz !0!white CCM estimation Noise variance = 33\,bps/Hz - −4.680-4.680\,dB 4.5544.554\,bps/Hz !20!white Blockage prediction Error probability +10%p - - 4.7474.747\,bps/Hz Fig. 24 shows an ablation study where each module is replaced with a CNN+LSTM model. Replacing the trajectory prediction, channel capacity estimation, and blockage prediction modules with CNN+LSTM leads to data-rate degradations of 6.46.4%, 24.424.4%, and 6.16.1%, respectively, since a simple discriminative model lacks the multimodal reasoning provided by LMM. The degradation is particularly noticeable for the channel capacity estimation module, because the channel capacity is more sensitive to variations and piecewise discontinuities in the propagation environment than the UE and blockage trajectories. Also, in Fig. 24, we conduct an experiment where the backbone LMM is replaced with lightweight models: 1) MobileVLM-1.7B [50] and 2) TinyLLaVA-1.5B [51]. We observe that MobileVLM-1.7B and TinyLLaVA-1.5B achieve 4.2634.263 bps/Hz and 4.3054.305 bps/Hz, respectively. Despite reducing the parameter size by more than 75% compared with LLaVA-1.5-7B, the performance degradation remains marginal, which demonstrates the feasibility of lightweight deployment. To investigate how the accumulation of errors affects overall system performance, we conduct an error propagation experiment across the entire framework in Table I. Specifically, we first inject artificial noise into the input UE trajectory and then evaluate its impact on various metrics, including the UE trajectory prediction root mean square error (RMSE), static channel capacity estimation NMSE, average channel capacity, and the worst 10%10\% and 20%20\% case channel capacity. The definitions of RMSE and NMSE are provided in Appendix B. When noise with variance 11\,m is added to the input UE trajectory, the UE trajectory prediction RMSE is 0.6750.675\,m and the average channel capacity is 4.6574.657\,bps/Hz. Compared with the noise-free case, the channel capacity decreases by only 2.4%2.4\%, indicating that the performance degradation is limited. Table I presents the robustness of LMM-EMM to the error propagation in real-world scenarios. The error level for each module was chosen to reflect practically plausible conditions, motivated by prior studies [33, 52, 53]. As shown in Table I, the performance is most sensitive to errors in the CCM estimation module. When the noise with a variance of 33 bps/Hz is added to the CCM estimation, the cumulative capacity per unit bandwidth decreases to 4.5544.554 bps/Hz, which is the lowest among all error sources. This is because the noise in the CCM makes it difficult for LMM to learn the underlying physical relationship between the propagation environment (e.g., reflection) and estimate the resulting channel capacity. In real-world scenarios, small-scale elements (e.g., trees, small objects) can cause temporary blockages and scattering effects, leading to some minor errors in blockage detection and channel capacity estimation. To account for such real-world variability, in Table I, we evaluate the average channel capacity under additional blockage detection errors and CCM estimation noise. We observe that even with an additional blockage error of 20% and a CCM noise standard deviation of 22\,bps/Hz, LMM-EMM achieves an average channel capacity of 4.6594.659\,bps/Hz, corresponding to only a 2.4%2.4\% performance degradation. This robustness is because handover decisions are often dominated by user location and major blockage events caused by large reflectors or obstacles, rather than by minor dispersion effects from small-scale objects. Figure 22: Average channel capacity as a function of SNR in different architectural layouts after few-shot adaptation. Figure 23: Average channel capacity vs. SNR, where each module is replaced by a CNN+LSTM. Figure 24: Average channel capacity vs. SNR with the backbone replaced by a lightweight LMM. In Table IV, we present the memory and computational complexities of LMM-EMM compared to conventional mobility management techniques. Although LMM-EMM exhibits the highest memory usage and longest inference time, its inference time remains shorter than the maximum handover latency (i.e., 360360\,ms) specified in 5G NR. Furthermore, the results can be pre-computed in advance of actual handover events by predicting future channel capacities, which further relaxes the practical constraints on inference time. This indicates that LMM-EMM has a feasible time complexity and is therefore compatible with the 5G NR standard. Moreover, recent progress in lightweight LMM and GPU hardware is expected to further alleviate memory and time complexities, so these issues are unlikely to pose practical concerns [54]. VI Conclusion In this paper, we proposed an environment-aware mobility management scheme for mmWave UDN systems. The key idea of the proposed LMM-EMM scheme is to leverage the CCM, an intrinsic mapping from UE and SBS positions to channel capacity. By harnessing the multimodal reasoning capability of LMMs, LMM-EMM captures the mobility pattern of the UE as well as reflection geometry and transient blockages caused by obstacles. Using the trained CCM with the predicted UE trajectory and blockage indicators, LMM-EMM estimates future channel capacities and employs DP to proactively determine handover decisions that maximize cumulative channel capacity. From numerical evaluations on various environments, we demonstrated that LMM-EMM achieves substantial channel capacity gains over the conventional DL-based methods. In this paper, we restricted our attention to mobility management but we believe that there are many research extensions of LMM-EMM such as user scheduling and random access. TABLE I: Average channel capacity (bps/Hz) at SNR =15=15 dB, with blockage detection error rate pblockp_block and CCM noise standard deviation σn _n. σn _n p_block 10% 15% 20% 25% 30% !20!white 0.5 bps/Hz 4.7544.754 4.7544.754 4.7434.743 4.7384.738 4.7274.727 !0!white 1 bps/Hz 4.7364.736 4.7354.735 4.7264.726 4.7214.721 4.7064.706 !20!white 1.5 bps/Hz 4.7064.706 4.7044.704 4.6974.697 4.6954.695 4.6784.678 !0!white 2 bps/Hz 4.6654.665 4.6684.668 4.6594.659 4.6624.662 4.6414.641 !20!white 2.5 bps/Hz 4.6424.642 4.6474.647 4.6474.647 4.6534.653 4.6164.616 Appendix A Proof of Lemma 1 In MISO systems, the channel capacity is expressed as R(t)(m) R^(t)(m) = = B∑s=1Slog2(1+PtNσn2∑i=1N|hm(t)[s,i]|2) B _s=1^S _2 (1+ P_tN _n^2 _i=1^N|h_m^(t)[s,i]|^2 ) (41) where hm(t)[s,i]h_m^(t)[s,i] is the iith element of m(t)[s]h_m^(t)[s]. When N is sufficiently large, by applying the law of large numbers, we obtain 1N∑i=1N|hm(t)[s,i]|2 1N _i=1^N|h_m^(t)[s,i]|^2 ≈ ≈ [|hm(t)[s,i]|2] [|h_m^(t)[s,i]|^2 ] (43) = = ∑l=0L−1βm,l(t)[|αm,l(t)|2] _l=0^L-1 _m,l^(t)E [| _m,l^(t)|^2 ] +∑p≠lβm,p(t)βm,l(t)[αm,p(t)α¯m,l(t)]ejϑ + _p≠ l _m,p^(t) _m,l^(t)E [ _m,p^(t) α_m,l^(t) ]e^j = = βm,0(t)+∑l=1L−1βm,l(t). _m,0^(t)+ _l=1^L-1 _m,l^(t). (44) This is from [|αm,l(t)|2]E [| _m,l^(t)|^2 ]=\,=\,11 and [αm,p(t)α¯m,l(t)]E [ _m,p^(t) α_m,l^(t) ]=\,=\,0 for p≠lp≠ l, where α¯m,l(t) α_m,l^(t) is the conjugate of αm,l(t)α_m,l^(t) and ϑ is some real value between 0 and 2π2π. By plugging (44) into (41), we finally obtain R(t)(m)=B∑s=1Slog2(1+Ptσn2(βm,0(t)+∑l=1L−1βm,l(t))).R^(t)(m)=B _s=1^S _2 (1+ P_t _n^2 ( _m,0^(t)+ _l=1^L-1 _m,l^(t) ) ). (45) TABLE IV: Memory usage and inference time Methods Memory usage Inference time !20!white Proposed LMM-EMM 14.1614.16\,GB 195.72195.72\,ms !0!white DRL-based scheme 1.241.24\,GB 45.6845.68\,ms !20!white Vision-aided scheme 0.820.82\,GB 14.9114.91\,ms !0!white LSTM-based scheme 0.270.27\,GB 6.696.69\,ms Appendix B Evaluation Metrics This appendix defines the evaluation metrics used in this paper, namely the RMSE for UE trajectory prediction and the NMSE for channel capacity estimation. 1. UE trajectory prediction RMSE: The RMSE for UE trajectory prediction is defined as RMSE=[(^ue(T:T+Tp)−ue(T:T+Tp))2]RMSE= E [ ( p_ue^(T:T+T_p)-p_ue^(T:T+T_p) )^2 ] (46) where ^ue(T:T+Tp) p_ue^(T:T+T_p) is the estimated future UE trajectory over the prediction horizon from T to T+TpT+T_p. 2. Channel capacity estimation NMSE: The NMSE for channel capacity estimation is defined as NMSE=[(R^(t)(m)−R(t)(m))2R(t)(m)2]NMSE=E [ ( R^(t)(m)-R^(t)(m) )^2R^(t)(m)^2 ] (47) where R^(t)(m) R^(t)(m) is the estimated channel capacity at the ttth time slot for the mmth SBS. References [1] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al., “Language models are few-shot learners,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 33, 2020, p. 1877–1901. [2] H. Liu, C. Li, Q. Wu, and Y. J. Lee, “Visual instruction tuning,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 36, 2023, p. 34 892–34 916. [3] J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V. Le, D. Zhou et al., “Chain-of-thought prompting elicits reasoning in large language models,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 35, 2022, p. 24 824–24 837. [4] A. Maatouk, N. Piovesan, F. Ayed, A. De Domenico, and M. Debbah, “Large language models for telecom: Forthcoming impact on the industry,” IEEE Commun. Mag., vol. 63, no. 1, p. 62–68, 2025. [5] H. J. Yang, H. Kim, H. Noh, S. Kim, and B. Shim, “Large multimodal model-empowered task-oriented autonomous communications: Design methodology and implementation challenges,” IEEE Veh. Technol. Mag., 2026. [6] Y. Ahn, J. Kim, S. Kim, S. Kim, and B. Shim, “Sensing and computer vision-aided mobility management for 6G millimeter and terahertz communication systems,” IEEE Trans. Commun., vol. 72, no. 10, p. 6044–6058, 2024. [7] J. Moon, S. Kim, H. Ju, and B. Shim, “Energy-efficient user association in mmWave/THz ultra-dense network via multi-agent deep reinforcement learning,” IEEE Trans. Green Commun. Netw., vol. 7, no. 2, p. 692–706, 2023. [8] S. Kim, J. Wu, and B. Shim, “Efficient channel probing and phase shift control for mmWave reconfigurable intelligent surface-aided communications,” IEEE Trans. Wireless Commun., vol. 23, no. 1, p. 231–246, 2023. [9] M. Kamel, W. Hamouda, and A. Youssef, “Ultra-dense networks: A survey,” IEEE Commun. Surveys Tuts., vol. 18, no. 4, p. 2522–2545, 2016. [10] Radio Resource Control (RRC), 3rd Generation Partnership Project 3GPP™ TS 38.331 V18.4.0, Dec. 2024, Release 18. [11] R. H. Clarke, “A statistical theory of mobile-radio reception,” Bell Sys. Technol. J., vol. 47, no. 6, p. 957–1000, 1968. [12] L. Yan, H. Ding, L. Zhang, J. Liu, X. Fang, Y. Fang, M. Xiao, and X. Huang, “Machine learning-based handovers for sub-6 GHz and mmWave integrated vehicular networks,” IEEE Trans. Wireless Commun., vol. 18, no. 10, p. 4873–4885, 2019. [13] S. H. A. Shah and S. Rangan, “Multi-cell multi-beam prediction using auto-encoder LSTM for mmWave systems,” IEEE Trans. Wireless Commun., vol. 21, no. 12, p. 10 366–10 380, 2022. [14] H.-S. Park, H. Kim, C. Lee, and H. Lee, “Mobility management paradigm shift: from reactive to proactive handover using AI/ML,” IEEE Netw., vol. 38, no. 2, p. 18–25, 2024. [15] L. Jiao, P. Wang, A. Alipour-Fanid, H. Zeng, and K. Zeng, “Enabling efficient blockage-aware handover in RIS-assisted mmWave cellular networks,” IEEE Trans. Wireless Commun., vol. 21, no. 4, p. 2243–2257, 2021. [16] W. Huang, M. Wu, Z. Yang, K. Sun, H. Zhang, and A. Nallanathan, “Self-adapting handover parameters optimization for SDN-enabled UDN,” IEEE Trans. Wireless Commun., vol. 21, no. 8, p. 6434–6447, 2022. [17] K. Sun, Q. Han, Z. Yang, W. Huang, H. Zhang, and V. C. Leung, “Proactive handover type prediction and parameter optimization based on machine learning,” IEEE Trans. Wireless Commun., vol. 24, no. 4, p. 3515–3528, 2025. [18] S. Kim, J. Moon, J. Kim, Y. Ahn, D. Kim, S. Kim, K. Shim, and B. Shim, “Role of sensing and computer vision in 6G wireless communications,” IEEE Wireless Commun., vol. 31, no. 5, p. 264–271, 2024. [19] G. Charan, M. Alrabeiah, and A. Alkhateeb, “Vision-aided 6G wireless communications: Blockage prediction and proactive handoff,” IEEE Trans. Veh. Technol., vol. 70, no. 10, p. 10 193–10 208, 2021. [20] U. Demirhan and A. Alkhateeb, “Radar aided proactive blockage prediction in real-world millimeter wave systems,” in Proc. IEEE Int. Conf. Commun. (ICC), 2022, p. 4547–4552. [21] Y. Liu, J. Wu, S. Kim, and B. Shim, “Vision-aided blockage prediction and proactive handover for indoor mmWave and terahertz communications,” in Proc. IEEE Global Commun. Conf. (GLOBECOM), 2023, p. 7411–7416. [22] C. Chaccour, W. Saad, M. Debbah, and H. V. Poor, “Joint sensing, communication, and AI: A trifecta for resilient THz user experiences,” IEEE Trans. Wireless Commun., vol. 23, no. 9, p. 11 444–11 460, 2024. [23] Y. Feng, C. Zhao, H. Luo, F. Gao, F. Liu, and S. Jin, “Networked ISAC based UAV tracking and handover towards low-altitude economy,” IEEE Trans. Wireless Commun., vol. 24, no. 9, p. 7670–7685, 2025. [24] C. Shen, C. Tekin, and M. van der Schaar, “A non-stochastic learning approach to energy efficient mobility management,” IEEE J. Sel. Areas Commun., vol. 34, no. 12, p. 3854–3868, 2016. [25] A. F. Molisch, M. Steinbauer, M. Toeltsch, E. Bonek, and R. S. Thoma, “Capacity of MIMO systems based on measured wireless channels,” IEEE J. Sel. Areas Commun., vol. 20, no. 3, p. 561–569, 2002. [26] J. Tang, F. Tang, S. Long, M. Zhao, and N. Kato, “Utilizing large language models for advanced optimization and intelligent management in space-air-ground integrated networks,” IEEE Netw., vol. 39, no. 5, p. 173–181, 2024. [27] D. Solomitckii, Q. C. Li, T. Balercia, C. R. Da Silva, S. Talwar, S. Andreev, and Y. Koucheryavy, “Characterizing the impact of diffuse scattering in urban millimeter-wave deployments,” IEEE Wireless Commun. Lett., vol. 5, no. 4, p. 432–435, 2016. [28] S. Kim, S. Jeong, J. Wu, B. Shim, and M. Z. Win, “Large multimodal model-based environment-aware channel estimation,” IEEE J. Sel. Areas Commun., vol. 43, no. 12, p. 4059–4075, 2025. [29] H. Ju, S. Jeong, S. Kim, B. Lee, and B. Shim, “Transformer-assisted parametric CSI feedback for mmWave massive MIMO systems,” IEEE Trans. Wireless Commun., vol. 23, no. 12, p. 18 774–18 787, 2024. [30] Evolved Universal Terrestrial Radio Access Network (E-UTRAN); X2 Application Protocol (X2AP), 3rd Generation Partnership Project 3GPP™ TS 36.423 V18.4.0, Mar. 2025, Release 18. [31] X. Hu, Y. Huo, X. Dong, F.-Y. Wu, and A. Huang, “Channel prediction using adaptive bidirectional GRU for underwater MIMO communications,” IEEE Internet Things J., vol. 11, no. 2, p. 3250–3263, 2023. [32] Q. Liu, G. Chuai, J. Wang, and J. Pan, “Proactive mobility management with trajectory prediction based on virtual cells in ultra-dense networks,” IEEE Trans. Veh. Technol., vol. 69, no. 8, p. 8832–8842, 2020. [33] L. Italiano, B. C. Tedeschini, M. Brambilla, H. Huang, M. Nicoli, and H. Wymeersch, “A tutorial on 5G positioning,” IEEE Commun. Surveys Tuts., vol. 27, no. 3, p. 1488–1535, 2024. [34] S. Kim, J. Moon, J. Wu, B. Shim, and M. Z. Win, “Vision-aided positioning and beam focusing for 6G terahertz communications,” IEEE J. Sel. Areas Commun., vol. 42, no. 9, p. 2503–2519, 2024. [35] Y. Zeng, J. Chen, J. Xu, D. Wu, X. Xu, S. Jin, X. Gao, D. Gesbert, S. Cui, and R. Zhang, “A tutorial on environment-aware communications via channel knowledge map for 6G,” IEEE Commun. Surveys Tuts., vol. 26, no. 3, p. 1478–1519, 2024. [36] Z. Lan, L. Liu, B. Fan, Y. Lv, Y. Ren, and Z. Cui, “Traj-LLM: A new exploration for empowering trajectory prediction with pre-trained large language models,” IEEE Trans. Intell. Veh., vol. 10, no. 2, p. 794–807, 2025. [37] I. Keum, J. Son, H. Kim, J. Moon, and B. Shim, “Deep learning-based NLoS localization using geometric map image,” IEEE Trans. Veh. Technol., 2025. [38] Y. Nie, N. H. Nguyen, P. Sinthong, and J. Kalagnanam, “A time series is worth 64 words: Long-term forecasting with transformers,” in Proc. Int. Conf. Learn. Representations (ICLR), 2023. [39] S. Kim, S. Saha, S. Jeong, B. Shim, and M. Z. Win, “Large multimodal model-based environment-aware beam management,” IEEE J. Sel. Areas Commun., vol. 44, p. 991–1007, 2026. [40] R. Bellman, “Dynamic programming,” Science, vol. 153, no. 3731, p. 34–37, 1966. [41] V. Va, J. Choi, and R. W. Heath, “The impact of beamwidth on temporal channel variation in vehicular channels and its implications,” IEEE Trans. Veh. Technol., vol. 66, no. 6, p. 5014–5029, 2016. [42] P. W. Patil, S. Gupta, S. Rana, S. Venkatesh, and S. Murala, “Multi-weather image restoration via domain translation,” in Proc. Int. Conf. Comput. Vis. (ICCV), 2023, p. 21 696–21 705. [43] M. Haklay and P. Weber, “Openstreetmap: User-generated street maps,” IEEE Pervasive Comput., vol. 7, no. 4, p. 12–18, 2008. [44] SMPTE Engineering Guideline - SD-SDI and HD-SDI Standards Roadmap, IEEE Standard EG 2111-1:2021, 2021. [45] G. R. MacCartney, S. Deng, S. Sun, and T. S. Rappaport, “Millimeter-wave human blockage at 73 ghz with a simple double knife-edge diffraction model and extension for directional antennas,” in IEEE Veh. Technol. Conf. (VTC-Fall), 2016. [46] J. Hoydis, F. Aït Aoudia, S. Cammerer, M. Nimier-David, N. Binder, G. Marcus, and A. Keller, “Sionna RT: Differentiable ray tracing for radio propagation modeling,” in Proc. IEEE Global Commun. Conf. Workshops (GC Wkshps), 2023, p. 317–321. [47] Study on channel model for frequencies from 0.5 to 100 GHz, 3rd Generation Partnership Project 3GPP™ TR 38.901 V18.0.0, Sep. 2020, Release 18. [48] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen et al., “LoRA: Low-rank adaptation of large language models,” in Proc. Int. Conf. Learn. Representations (ICLR), 2022. [49] Y. Zeng and X. Xu, “Toward environment-aware 6G communications via channel knowledge map,” IEEE Wireless Commun., vol. 28, no. 3, p. 84–91, 2021. [50] X. Chu et al., “Mobilevlm: A fast, strong and open vision language assistant for mobile devices,” arXiv preprint arXiv:2312.16886, 2023. [51] B. Zhou et al., “Tinyllava: A framework of small-scale large multimodal models,” arXiv preprint arXiv:2402.14289, 2024. [52] T. Nishio et al., “Proactive received power prediction using machine learning and depth images for mmWave networks,” IEEE J. Sel. Areas Commun., vol. 37, no. 11, p. 2413–2427, 2019. [53] M. Alrabeiah and A. Alkhateeb, “Deep learning for mmWave beam and blockage prediction using sub-6 GHz channels,” IEEE Trans. Commun., vol. 68, no. 9, p. 5504–5518, 2020. [54] J. Chen, L. Ye, J. He, Z.-Y. Wang, D. Khashabi, and A. Yuille, “Efficient large multi-modal models via visual context compression,” in Proc. Adv. Neural Inf. Process. Syst. (NeurIPS), vol. 37, 2024, p. 73 986–74 007. Seokhyun Jeong (Student Member, IEEE) received the B.S. degree from the Department of Electrical and Computer Engineering, Seoul National University, Seoul, South Korea, in 2022, where he is currently pursuing the Ph.D. degree in electrical and computer engineering. His main areas of research are in artificial intelligence for wireless communications. Sangmok Shin (Student Member, IEEE) received the B.S. degree from Daegu Gyeongbuk Institute of Science and Technology (DGIST), Daegu, South Korea, in 2023. He is currently pursuing the Ph.D. degree in Electrical and Computer Engineering at Seoul National University, Seoul, South Korea. His research interests include AI-assisted physical-layer wireless communications, with particular emphasis on beam management, channel estimation, and channel prediction. Mr. Shin was a recipient of the Samsung Humantech Paper Award Bronze Prize and the Graduate Student of the Year Award from the Department of Electrical and Computer Engineering in 2025. Seungnyun Kim (Member, IEEE) received the B.S. (with honors) and Ph.D. degrees in electrical and computer engineering from Seoul National University (SNU), Seoul, South Korea, in 2016 and 2023, respectively. He is an Assistant Professor at the Singapore University of Technology and Design (SUTD). Prior to joining SUTD, he was a Postdoctoral Fellow with the Wireless Information and Network Sciences Laboratory, Massachusetts Institute of Technology, Cambridge, MA, USA. His research interests include information theory, optimization methods, and machine learning with applications to real-world problems, including wireless communications, network localization and navigation, and non-terrestrial networks. Dr. Kim was a recipient of the Sejong Science Fellowship from Korean Government in 2023, the Best Ph.D. Dissertation Award from SNU in 2023, the Qualcomm Innovation Fellowship Finalist in 2021, and the Samsung Humantech Paper Award Gold Prize in 2019. Jiao Wu (Member, IEEE) received the B.S. degree in communication engineering from North China Electric Power University (NCEPU), Beijing, China, in 2015, the M.S. degree in electronics and communication engineering from University of Electronic Science and Technology of China (UESTC), Chengdu, China, in 2018, and the Ph.D. degree in electrical and computer engineering from Seoul National University (SNU), Seoul, South Korea, in 2023. She is currently a Postdoctoral Fellow with the Computer, Electrical and Mathematical Sciences and Engineering Division (CEMSE), King Abdullah University of Science and Technology (KAUST), Thuwal, Saudi Arabia. Her research interests include signal processing, optimization techniques, and machine learning with the applications to reconfigurable intelligent surfaces-assisted communications, extremely large-scale antenna systems, integrated sensing and communications, and non-terrestrial networks. Byonghyo Shim (Fellow, IEEE) received the B.S. and M.S. degrees in Control and Instrumentation Engineering from Seoul National University, South Korea, in 1995 and 1997, respectively, and the M.S. degree in mathematics and the Ph.D. degree in Electrical and Computer Engineering from the University of Illinois at Urbana-Champaign (UIUC), Champaign, IL, USA, in 2004 and 2005, respectively. From 1997 to 2000, he was an Officer (First Lieutenant) and an Academic full-time Instructor in the Department of Electronics Engineering, Korean Air Force Academy. From 2005 to 2007, he was a Staff Engineer with Qualcomm Inc., San Diego, CA, USA. From 2007 to 2014, he was an Associate Professor with the School of Information and Communication, Korea University, Seoul. Since 2014, he has been with Seoul National University (SNU), where he is currently a Professor of the Department of Electrical and Computer Engineering and Vice Dean of Engineering College. His research interests include wireless communications, deep learning, and statistical signal processing. Dr. Shim was a recipient of the M. E. Van Valkenburg Research Award from the ECE Department, University of Illinois, in 2005, the Haedong Young Engineer Award from IEIE in 2010, the Irwin Jacobs Award from Qualcomm and KICS in 2016, the Shinyang Research Award from the Engineering College of SNU in 2017, the Okawa Foundation Research Award in 2020, the IEEE Comsoc AP Outstanding Paper Award in 2021, the JCN Best Paper Award in 2024, and the SNU Academic Research Award in 2025. Dr. Shim was an Elected Member of the Signal Processing for Communications and Networking (SPCOM) Technical Committee of the IEEE Signal Processing Society. He has served as an Associate Editor for IEEE Transactions on Wireless Communications (TWC), IEEE Transactions on Communications (TCOM), IEEE Transactions on Vehicular Technology (TVT), IEEE Transactions on Signal Processing (TSP), IEEE Wireless Communications Letters (WCL), and Journal of Communications and Networks (JCN) and a Guest Editor for IEEE Journal on Selected Areas in Communications (JSAC).