Paper deep dive
SonarLLM: A Native Sonar--Optical Multimodal Large Language Model for Underwater Perception
Cong Su, longxuan ma, Ling Dong, Guofeng Tang, Weijie Yin, Haohui Chen, Zhengtao Yu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/29/2026, 4:23:53 AM
Summary
The paper introduces SonarLLM, a multimodal large language model designed for underwater perception that treats imaging sonar as a native modality alongside optical data. It addresses the limitations of existing models by using a sonar-specific encoder (PSVT), physics-aware feature enhancement modules (Optical-VFE and Acoustic-VFE), and a reliability-aware hierarchical fusion mechanism (AGFM with DeepStack). The authors also propose SonarBench, a paired benchmark for evaluating sonar-optical complementarity under controlled optical degradation. SonarLLM significantly outperforms baselines in recognition, counting, and VQA tasks, demonstrating that robust heterogeneous perception requires representing and weighting sonar according to its specific sensing characteristics.
Entities (11)
Relation Signals (10)
SonarLLM ā usescomponent ā PSVT
confidence 95% Ā· SonarLLM ... introduces an independent sonar encoder ... We initialize a Polar-aware Sonar Vision Transformer (PSVT)
SonarLLM ā usescomponent ā Qwen3-VL-8B
confidence 95% Ā· SonarLLM retains the pretrained Qwen3-VL-8B optical encoder
SonarLLM ā usescomponent ā Optical-VFE
confidence 95% Ā· dedicated Visual Feature Enhancement (VFE) modules operate in each modality: Optical-VFE targets scattering-related corruption
SonarLLM ā usescomponent ā Acoustic-VFE
confidence 95% Ā· Acoustic-VFE targets reverberation-related components and range attenuation
SonarLLM ā usescomponent ā AGFM
confidence 95% Ā· AGFM predicts quality-aware modality weights
SonarLLM ā usescomponent ā DeepStack
confidence 95% Ā· dual-stream hierarchical DeepStack delivers reweighted optical and sonar features
SonarBench ā createdby ā SonarLLM
confidence 90% Ā· We also introduce SonarBench ... By fixing the scene and sonar observation ... SonarBench enables controlled measurement
SonarLLM ā developedby ā Kunming University of Science and Technology
confidence 90% Ā· Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct range-azimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonar-optical complementarity. We propose SonarLLM, a sonar-optical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasks: recognition, counting, visual question answering, and captioning; and, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics.
Tags
Links
- Source: https://arxiv.org/abs/2608.24325v1
- Canonical: https://arxiv.org/abs/2608.24325v1
Trouble viewing inline? Open PDF directly ā
Full Text
54,437 characters extracted from source content.
Expand or collapse full text
SonarLLM: A Native SonarāOptical Multimodal Large Language Model for Underwater Perception DOI: X.XXXXXXXConference: Make sure to enter the correct conference title from your rights confirmation email; June 03ā05, 2018; Woodstock, NYISBN: 978-1-4503-X-X/2018/06CCS: Computing methodologies Computer visionCCS: Computing methodologies Machine learningCCS: Computing methodologies Natural language processing Cong Su email: siliconevolution@gmail.com Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China , Longxuan Ma Note: Corresponding author. email: lxma@kust.edu.cn Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China , Ling Dong email: ling.dong@kust.edu.cn Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China , Guofeng Tang email: 920903629@q.com Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China , Weijie Yin email: 18102382274@163.com Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China , Haohui Chen email: 2954390791@q.com Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China and Zhengtao Yu email: ztyu@hotmail.com Affiliation: Faculty of Information Engineering and Automation, Kunming University of Science and Technology, Kunming, China Affiliation: Yunnan Key Laboratory of Artificial Intelligence, Kunming, China Received 5 June 2009 Abstract. Reliable underwater perception requires complementary sensing under variable visibility. Optical cameras capture appearance and semantics but degrade rapidly with turbidity, whereas imaging sonar preserves geometry while exhibiting distinct rangeāazimuth structure and acoustic artifacts. Existing MLLMs, built primarily on optical encoders, are therefore ill-suited to model sonar or adaptively exploit sonarāoptical complementarity. We propose SonarLLM, a sonarāoptical MLLM that treats sonar as a native perceptual modality. It combines a sonar-specific encoder, modality-specific physics-aware feature enhancement, and reliability-aware hierarchical fusion to align acoustic structure with optical semantics and dynamically adjust their contributions as sensing quality changes. We also introduce SonarBench, a paired benchmark that spans four tasksārecognition, counting, visual question answering, and captioningāand, across the benchmark, three input settings: sonar-only, optical-only, and fusion. By fixing the scene and sonar observation while varying optical degradation, SonarBench enables controlled measurement of cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points, and 68.7% under fusion, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain grows from 6.0 to 36.0 points as turbidity increases, indicating the increasing complementary value of sonar under controlled optical degradation. Together, these results show that robust heterogeneous perception depends not only on adding sonar, but on representing and weighting it according to its sensing characteristics. Keywords: Multimodal Large Language Model, Underwater Perception and Understanding, Sonar-Optical 1. Introduction Underwater environments impose severe and highly variable sensing conditions. Recent multimodal large language models (MLLMs) have advanced open-ended visual understanding, while underwater models such as MarineGPT (Zheng et al., 2023) and NAUTILUS (Xu et al., 2025) extend these capabilities to marine image captioning, question answering, and recognition. However, these models remain predominantly dependent on optical imagery. Absorption and scattering progressively erase color, texture, contrast, and object boundaries as turbidity increases (Li et al., 2020). Imaging sonar is largely independent of illumination and can preserve object contours, spatial structure, and range information under poor visibility (Neupane and Seok, 2020). Optical and sonar observations are therefore complementary: one provides appearance and semantic detail, whereas the other supplies more stable geometric evidence under optical degradation (Li et al., 2025), as illustrated in Fig. 1. Figure 1. Paired sonarāoptical observations under controlled optical degradation. Optical evidence weakens with turbidity, whereas sonar preserves structural cues for the same scene. Exploiting this complementarity requires more than adding sonar as another visual input. Existing sonarāoptical systems are primarily designed for task-specific detection or tracking (Li et al., 2025; Wu et al., 2026) and do not address open-ended language reasoning. Natural-image encoders are also poorly matched to sonar, whose rangeāazimuth geometry and artifacts include speckle, reverberation, acoustic shadows, and range-dependent propagation loss (Neupane and Seok, 2020; Steiniger et al., 2022). Moreover, modality reliability is observation-dependent: optical evidence may collapse under turbidity, while sonar can be corrupted by acoustic noise and artifacts, making fixed fusion vulnerable to an unreliable sensor (Park et al., 2025). Together, these limitations motivate a sonarāoptical MLLM that jointly addresses modality-specific representation, degradation-aware enhancement, and reliability-aware interaction across sensors and semantic levels. To address these key challenges, we propose SonarLLM, a sonarāoptical MLLM that treats imaging sonar as a native perceptual modality rather than an auxiliary image. First, for modality-specific representation, SonarLLM retains the pretrained Qwen3-VL-8B optical encoder (Bai et al., 2025) and introduces an independent sonar encoder equipped with a multi-scale Sonar Stem and rangeāazimuth positional encoding. Second, for modality-dependent degradation, dedicated Visual Feature Enhancement (VFE) modules operate in each modality: Optical-VFE targets scattering-related corruption, whereas Acoustic-VFE targets reverberation-related components and range attenuation. Third, AGFM predicts quality-aware modality weights, while dual-stream hierarchical DeepStack delivers reweighted optical and sonar features to multiple language-model layers. Finally, progressive training establishes sonar-domain representations, cross-modal correspondence, reliability learning, and instruction-following capability. We further introduce SonarBench, a paired benchmark for controlled evaluation of sonarāoptical understanding. It covers recognition, counting, visual question answering, and captioning and spans sonar-only, optical-only, and fusion settings across the benchmark as a whole. Fusion evaluation focuses on the three accuracy-based tasks, while captioning provides a separate diagnostic of open-ended generation. Optical observations are evaluated from clear to heavily turbid conditions. Unlike protocols that compare different samples across conditions, SonarBench fixes the underlying scene and sonar observation while varying only optical quality. This design separates gains from improved unimodal modeling from gains attributable to complementary sonar evidence as optical reliability deteriorates. Experiments demonstrate both strong sonar understanding and robust cross-modal complementarity. SonarLLM achieves 72.0% macro accuracy across sonar-only recognition, counting, and VQA, outperforming the strongest baseline by 34.4 percentage points; substantially larger 27Bā35B general-purpose MLLMs do not close this gap. Under fusion input, SonarLLM reaches 68.7%, exceeding the best baseline by 25.1 points. For recognition and counting, the fusion-over-optical gain increases from 6.0 points under clear conditions to 36.0 points under heavy turbidity, while fusion performance remains comparatively stable. SonarLLM achieves the best overall performance among the evaluated MLLMs on SonarBench. Our main contributions are summarized as follows: ⢠We formulate sonarāoptical language understanding as a heterogeneous sensing problem and propose SonarLLM, which unifies sonar-specific representation, modality-specific feature enhancement, reliability-aware gating, and hierarchical interaction within a shared language model. ⢠We introduce SonarBench, a paired benchmark whose controlled intervention keeps the scene and sonar observation fixed while varying optical quality, thereby isolating unimodal sonar capability from cross-modal complementarity. ⢠Through controlled comparisons, representation analysis, and structural ablations, we show that native sonar modeling provides capabilities not recovered by model scale or instruction tuning alone, while reliability-aware fusion becomes increasingly valuable as optical evidence deteriorates. 2. Related Work 2.1. Underwater Vision-Language and Sonar Perception Recent work has extended vision-language learning to marine and underwater environments. MarineGPT (Zheng et al., 2023) uses domain-specific imageātext data and instruction tuning for captioning, question answering, and recognition, while AquaticCLIP (Alawode et al., 2025) learns underwater visionālanguage representations through large-scale contrastive pretraining. NAUTILUS (Xu et al., 2025) incorporates physics-aware enhancement, and OceanGPT (Bi et al., 2024) targets broader ocean-domain knowledge. OceanPile (Xue et al., 2026) and OceanGym (Xue et al., 2025) provide ocean multimodal data and embodied evaluation, respectively, but do not address sonar-native open-ended reasoning under paired optical degradation. These efforts demonstrate the value of domain data and priors, yet do not study the heterogeneous language-level interaction considered here. Imaging sonar provides complementary structural observations independent of ambient illumination. Existing sonar research primarily addresses detection, classification, segmentation, and tracking (Neupane and Seok, 2020; Steiniger et al., 2022), supported by datasets such as UATD (Xie et al., 2022) and SCTD (Zhang et al., 2022). Paired datasets including RGBS50 (Li et al., 2025), UMOD (Wu et al., 2026), and SOVIS (Chen et al., 2026) further demonstrate the value of sonarāoptical fusion for specific discriminative tasks. However, these systems typically fuse task-dependent features or predictions and do not connect acoustic representations to open-ended language reasoning. SonarLLM instead models sonar as a native perceptual modality and enables continuous interaction among acoustic structure, optical semantics, and language representations. 2.2. Heterogeneous Multimodal Fusion and Evaluation General-purpose MLLMs employ diverse mechanisms to connect visual and linguistic representations. LLaVA (Liu et al., 2023b; Liu et al., 2023a) uses learned projection, BLIP-2 (Li et al., 2023) introduces a Q-Former, Flamingo (Alayrac et al., 2022) injects visual context through cross-layer attention, and Qwen3-VL (Bai et al., 2025) incorporates hierarchical visual features through DeepStack. Gated and reliability-aware fusion methods (Arevalo et al., 2017; Park et al., 2025) additionally adjust modality contributions according to sensor quality. Nevertheless, these approaches generally assume visually homogeneous inputs or do not explicitly account for sensors with different imaging geometry, degradation processes, and environment-dependent reliability. SonarLLM combines modality-specific representation and enhancement with reliability-aware hierarchical interaction to address these requirements jointly. Existing underwater evaluation resources are similarly divided by modality or task. UIEB (Li et al., 2020) and EUVP (Islam et al., 2020) focus on optical enhancement; UATD and SCTD support sonar perception; and RGBS50, UMOD, and SOVIS target specific paired-sensor tasks. Vision-language benchmarks such as NautData (Xu et al., 2025) and UWBench (Zhang et al., 2025) remain centered on optical observations. They therefore cannot isolate whether multimodal gains arise from stronger unimodal modeling or from genuinely complementary sonar evidence. SonarBench addresses this limitation through paired interventions that keep the scene and sonar observation fixed while varying only optical quality, enabling controlled evaluation of sonar representation, cross-modal complementarity, and reliability adaptation. 3. Method 3.1. Overall Architecture Given an optical image IoI_o, an imaging-sonar observation IsI_s, and a textual instruction x, SonarLLM extracts heterogeneous visual representations and performs joint reasoning within a shared language model. As shown in Fig. 2, it preserves the pretrained optical encoder ā°oE_o of Qwen3-VL-8B (Bai et al., 2025) and introduces an independent sonar encoder ā°sE_s, thereby treating sonar as a native perceptual modality rather than an auxiliary image input. Figure 2. Overall architecture of SonarLLM, comprising sonar-native representation, modality-specific feature enhancement, reliability-aware hierarchical fusion, and progressive training. Each visual encoder produces a final feature FmF_m and three intermediate features Fm(k)k=13\F_m^(k)\_k=1^3 extracted after Transformer blocks 8, 16, and 24, where māo,smā\o,s\. Final features are processed by modality-specific Visual Feature Enhancement (VFE) modules, projected into the language space, and reweighted by AGFM0. Intermediate features bypass the VFEs and are processed by AGFMk before being projected and injected into language-model layers 8, 16, and 24, respectively, through dual-stream DeepStack. This separation preserves modality-specific information while enabling cross-modal interaction at multiple semantic depths. 3.2. Sonar-Native Visual Representation Natural-image encoders are poorly matched to sonar statistics and rangeāazimuth geometry. We initialize a Polar-aware Sonar Vision Transformer (PSVT) from the Qwen3-VL visual tower to retain transferable visual priors, and adapt it through a multi-scale Sonar Stem and explicit geometric positional encoding. Multi-scale Sonar Stem. Before patch embedding, the Sonar Stem captures echoes, boundaries, and acoustic shadows at different receptive fields: (1) ĪāIs I_s =Conv1Ć1ā[f7ā(f5ā(f3ā(Is)))], =Conv_1Ć 1 [f_7 (f_5 (f_3(I_s) ) ) ], I~s I_s =Is+γāĪāIs, =I_s+γ I_s, where f3f_3, f5f_5, and f7f_7 denote convolutional transformations with different receptive fields. The zero-initialized coefficient γ gradually introduces sonar-specific structure without disrupting the transferred representation at initialization. RangeāAzimuth Positional Adaptation. For a sonar patch centered at range rir_i and azimuth Īøj _j, we augment the transferred positional representation as: (2) ziājs=ziājbase+Ī»pā[Ļrā(ri)+ĻĪøā(Īøj)],z_ij^s=z_ij^base+ _p [ _r(r_i)+ _Īø( _j) ], where Ļr _r and ĻĪø _Īø encode physical range and azimuth, respectively. PSVT can therefore retain pretrained visual priors while explicitly adapting to sonar imaging geometry. 3.3. Modality-Specific Feature Enhancement Optical and sonar observations undergo different physical degradations. A compact abstraction of their image-formation processes is: (3) Ioā(u) I_o(u) =Jā”(u)āeāβādā(u)+Bāā[1āeāβādā(u)], =J(u)e^-β d(u)+B_ā [1-e^-β d(u) ], Sā”(r,Īø) S(r,Īø) =S0āTāSā(Īø)r2āeā2āαphyār+Rā”(r,Īø), = S_0TS(Īø)r^2e^-2 _phyr+R(r,Īø), where dā”(u)d(u) is scene range, β is optical attenuation, Jā”(u)J(u) is undegraded scene radiance, and BāB_ā is backscattered light. In the sonar model, S0S_0 is a source-level constant, αphy _phy is acoustic attenuation, TāSā(Īø)TS(Īø) is target strength, and Rā”(r,Īø)R(r,Īø) is reverberation. Motivated by these distinct attenuation and interference terms, we apply separate Optical-VFE and Acoustic-VFE modules to the final high-level features, while intermediate DeepStack features bypass the VFEs. The modules perform feature-space correction rather than inversion of the raw physical image-formation processes. Optical-VFE. A learnable query ebe_b estimates feature corruption related to scattering in FoF_o, while a frozen DINOv2-L encoder (Oquab et al., 2023) provides a generic structural reference DoD_o: (4) Co C_o =MHAā”(eb,Fo,Fo), =MHA(e_b,F_o,F_o), Foc F_o^c =LNā”(FoāWbāCo), =LN (F_o-W_bC_o ), FĀÆo F_o =LNā”[Foc+ĪØoā(Do)]. =LN [F_o^c+ _o(D_o) ]. Here, WbW_b is a learned channel projection, and ĪØo _o maps DoD_o to the token shape of FocF_o^c. Optical-VFE performs feature correction rather than physical inversion. Acoustic-VFE. A learnable acoustic query ere_r estimates a feature component associated with reverberation, while token range parameterizes learned compensation for range-dependent attenuation: (5) Cs C_s =MHAā”(er,Fs,Fs), =MHA(e_r,F_s,F_s), Fscā[i] F_s^c[i] =Fsā[i]āĻā”(gr)āWsāCs, =F_s[i]-Ļ(g_r)W_sC_s, Gā”(ri) G(r_i) =expā”[2ālogā”(rirmin)+2āαcā(riārmin)], = \! [2 ( r_ir_ )+2 _c(r_i-r_ ) ], FĀÆsā[i] F_s[i] =LNā”(Gā”(ri)āFscā[i]). =LN (G(r_i)F_s^c[i] ). Here, WsW_s is a learned channel projection, Ļ is sigmoid, and grg_r is a scalar gate. The learned range-compensation coefficient αc=softplusā”(α^c) _c=softplus( α_c) is distinct from the physical αphy _phy in Eq. (3). Because rir_i is the physical range of the sonar token, the compensation retains the native rangeāazimuth geometry. 3.4. Reliability-Aware Hierarchical Fusion The reliability of optical and sonar observations varies across both environments and spatial regions. SonarLLM therefore combines the Adaptive Gated Fusion Module (AGFM) with dual-stream DeepStack to model global sensor reliability and local token quality at multiple visual levels. At level k, let Fm(k)=fm,i(k)i=1NmF_m^(k)=\f_m,i^(k)\_i=1^N_m. For k=0k=0, Fm(0)F_m^(0) denotes the enhanced and projected final representation; for kā1,2,3kā\1,2,3\, it denotes the feature extracted after visual block 8, 16, or 24. AGFM first estimates token quality and aggregates it into modality-level reliability: (6) qm,i(k) q_m,i^(k) =hm(k)(LN(fm,i(k))),qĀÆm(k)=Nmā1āj=1Nmqm,j(k), =h_m^(k) (LN (f_m,i^(k) ) ), q_m^(k)=N_m^-1 _j=1^N_mq_m,j^(k), ĀÆ(k) q^(k) =[qĀÆo(k),qĀÆs(k)],(k)=softmax(ĀÆ(k)/Ļk). = [ q_o^(k), q_s^(k) ], ^(k)=softmax ( q^(k)/ _k ). Here, hm(k):ādkāāh_m^(k):R^d_k\!ā\!R is a token-shared scalar scorer. The vector (k)=[go(k),gs(k)]g^(k)=[g_o^(k),g_s^(k)] contains relative modality weights used as controlled reliability proxies. Each Ļk _k is learnable, lower-bounded at 0.050.05, and initialized to 2.02.0. AGFM further captures spatially non-uniform quality through mean-preserving token modulation: (7) um,i(k) u_m,i^(k) =1+Ļā”(qm,i(k))2,uĀÆm(k)=Nmā1āj=1Nmum,j(k), = 1+Ļ (q_m,i^(k) )2, u_m^(k)=N_m^-1 _j=1^N_mu_m,j^(k), Ļm,i(k) _m,i^(k) =um,i(k)uĀÆm(k),f~m,i(k)=gm(k)Ļm,i(k)fm,i(k). = u_m,i^(k) u_m^(k), f_m,i^(k)=g_m^(k) _m,i^(k)f_m,i^(k). The global factor gm(k)g_m^(k) allocates reliability across sensors, while Ļm,i(k) _m,i^(k) redistributes importance among tokens without changing their mean scale. For single-modality input, AGFM reduces to an identity mapping. Crucially, AGFM reweights rather than merges the two modality sequences. At the final level, the reweighted optical and sonar tokens form the initial multimodal context. At the three intermediate levels, the reweighted features are projected by modality-specific DeepStack mergers and injected into language layers 8, 16, and 24, respectively: (8) H(āk) H^( _k) āH(āk)+o(k)ā[o(k)ā(F~o(k))] ā H^( _k)+J_o^(k) [P_o^(k) ( F_o^(k) ) ] +s(k)ā[s(k)ā(F~s(k))]. +J_s^(k) [P_s^(k) ( F_s^(k) ) ]. Equation (8) is applied at language layers 8, 16, and 24. Here, m(k)P_m^(k) is a modality-specific DeepStack merger that maps intermediate features to the language-model hidden space, and m(k)J_m^(k) aligns and adds them to the visual-token span of the corresponding modality. Maintaining separate optical and sonar streams before injection avoids premature compression of their heterogeneous representations. 3.5. Progressive Training Strategy We train SonarLLM in four stages that progressively establish sonar-domain representations, acoustic semantics, cross-modal reliability, and language-level reasoning. Direct joint optimization would require limited paired data to simultaneously resolve domain shift, semantic organization, sensor correspondence, gate calibration, and instruction following. The staged curriculum first stabilizes acoustic representations and then introduces cross-modal and language supervision, reducing interference among these heterogeneous objectives. Table 1 summarizes the resulting schedule. Table 1. Progressive training schedule. CE denotes category cross-entropy. Stage Data/objective Trainable modules I Unlabeled sonar / MAE PSVT, decoder I Labeled sonar / CE PSVT/Stem, classifier I Paired / āalignL_align Sonar path, VFEs, AGFM IV Instructions / āinstL_inst LoRA, interfaces Stage I: Sonar-Domain Adaptation. PSVT is first adapted using masked reconstruction on unlabeled sonar images. Because low-response background occupies a large portion of each frame, normalized masked patches xĀÆi x_i are weighted by their local variation wiw_i: (9) āadapt=āiāā³wiāāā”(zi)āxĀÆiā22āiāā³wi.L_adapt= _i w_i \|D(z_i)- x_i \|_2^2 _i w_i. Here, xĀÆi=(xiāμā”(xi))/(Ļā”(xi)+ε) x_i=(x_i-μ(x_i))/(Ļ(x_i)+ ) and wi=clipā”(Ļā”(xi),Ļmin,Ļmax)w_i=clip(Ļ(x_i), _ , _ ). Only the sonar pathway is optimized, and the decoder D is discarded afterward. Stage I: Acoustic Semantic Learning. Category supervision then organizes the adapted sonar representations according to acoustic semantics: (10) āsem _sem =ā1Bān=1Bāc=1Cyn,clogps,c(n), =- 1B _n=1^B _c=1^Cy_n,c p_s,c^(n), ps(n) p_s^(n) =softmaxā”(Wcāhs(n)+bc). =softmax (W_ch_s^(n)+b_c ). where hs(n)=Poolā”[ā°sā(I~s(n))]h_s^(n)=Pool[E_s( I_s^(n))]. The optical pathway, language model, and fusion modules remain frozen. Stage I: Cross-Modal Alignment and Reliability Learning. Synchronized sonarāoptical pairs are used with stochastic optical degradation across a continuous severity range, while clear pairs are retained as anchors. Training combines global contrastive alignment, hierarchical feature alignment, BCE-based gate supervision, and gate regularization: (11) āNCE _NCE =InfoNCEā”(zo,zs), =InfoNCE(z_o,z_s), āhier _hier =āk=13[1ācosā”(zo(k),zs(k))], = _k=1^3 [1- (z_o^(k),z_s^(k) ) ], āgate _gate =14āāk=03BCEā”(gs(k),Ļsā(Ī·)), = 14 _k=0^3BCE (g_s^(k), _s(Ī·) ), āent _ent =ā14āk=03āb(gs(k)), =- 14 _k=0^3H_b\! (g_s^(k) ), āalign _align =āNCE+0.5āāhier+0.2āāgate+0.1āāent. =L_NCE+0.5L_hier+0.2L_gate+0.1L_ent. Here, zmz_m and zm(k)z_m^(k) are pooled final and intermediate representations, respectively; InfoNCE uses temperature 0.070.07, and ābH_b denotes binary entropy. The BCE target Ļsā(Ī·)=clampā”(0.5+0.4āĪ·,0,1) _s(Ī·)=clamp(0.5+0.4Ī·,0,1) shifts supervision from balanced fusion toward sonar as optical degradation increases, making gsg_s a supervised degradation proxy rather than a general reliability estimate. The negative-entropy term discourages early gate collapse. Modality-missing examples preserve compatibility with sonar-only and optical-only inputs. Stage IV: Multimodal Instruction Tuning. Finally, the model is trained on sonar-only, optical-only, and paired sonarāoptical instructions. For visual input āāIs,Io,(Is,Io)Iā\I_s,I_o,(I_s,I_o)\, prompt x, and target answer y, we minimize the answer-only autoregressive objective: (12) āinst=ā1||ātālogpĪ(ytā£x,ā,y<t),L_inst=- 1|A| _t p_ (y_t x,I,y_<t ), where A contains answer-token positions and Ī includes the trainable LoRA parameters (Hu et al., 2021) and multimodal interfaces. This stage transfers the learned sonar representations and reliability-aware alignment to open-ended recognition, counting, question answering, and captioning. 4. Experiments Our evaluation follows the causal structure of SonarLLM. We first test whether a sonar-native pathway provides capabilities that cannot be recovered through model scale or instruction tuning alone. We then use paired optical degradation to isolate when sonar becomes complementary and whether AGFM responds to the resulting reliability shift. Finally, representation analysis and controlled ablations trace these gains to sonar-domain adaptation, cross-modal alignment, sonar geometry, modality-specific enhancement, and hierarchical interaction before we quantify their computational cost. Unless otherwise specified, all models use identical task prompts, input protocols, and deterministic decoding. 4.1. Experimental Setup Training Data. Stage I pools approximately 98K candidate sonar images from RGBS50, UATD, SCTD, DeeperSense, FLC+FLS, and OceanGym (Xue et al., 2025), together with sonar-filtered frames from OceanInstruct and OceanPileās OceanInstruction (Xue et al., 2026). After empty-frame removal, outlier filtering, byte-level deduplication, and brightness-balanced sampling, we retain 40K unlabeled images for sonar-domain adaptation. Stage I uses 23-class sonar object annotations aggregated from these sources. Stage I uses synchronized RGBS50 sonarāoptical pairs with online optical degradation. Stage IV uses a 635K-sample instruction mixture comprising 533K sonar-related and 102K optical samples, with sonar-only, optical-only, and paired inputs. SonarBench. SonarBench evaluates recognition, counting, VQA, and captioning. Recognition, counting, and VQA use sonar-only, optical-only, and fusion inputs; captioning uses sonar-only and optical-only inputs. Optical and fusion settings are evaluated under clear, turbid, and heavily turbid conditions. Across degradation levels, the scene, question, and sonar observation remain fixed, and only the paired RGB image is modified. This design is a controlled stress test of optical reliability rather than a simulation of the complete distribution of natural turbidity. Within this protocol, paired fusion analysis is defined over recognition, counting, and VQA, whereas captioning is reported separately as a diagnostic of open-ended semantic generation. Table 2. SonarBench evaluation structure. S/O/F denote sonar, optical, and fusion inputs; C/T/H denote clear, turbid, and heavy optical conditions. Task Input conditions Metric Subsets Recognition S; O/FĆC/T/H Sem. accuracy 7 Counting S; O/FĆC/T/H Exact accuracy 7 VQA S; O/FĆC/T/H Sem. accuracy 7 Captioning S; OĆC/T/H GOOD / METEOR 4 Total 25 The three optical conditions provide a controlled contrast in observation quality. Clear images remain unchanged; turbid images combine reduced brightness, additive white noise, colored veiling, and Gaussian blur; and heavy-turbid images retain the same photometric attenuation while replacing white noise with multi-scale correlated disturbances. Training degradation data and benchmark renderings use the same generator, so the comparison isolates response to a known reliability shift rather than out-of-family generalization to arbitrary natural turbidity. Evaluation samples are drawn from held-out RGBS50 and UMOD video sequences, forming 25 subsetsā4 sonar-only, 12 optical-only, and 9 fusionāwith 150 QA instances each. Each record contains one question, while an underlying sequenceāframe may recur across tasks, modalities, and degradation conditions to support paired comparisons. Source sequences are partitioned before all training stages, and a sequenceāframe audit confirms zero overlap between training and evaluation. Metrics and Judging. Recognition and VQA use Qwen3.8-Max (Qwen Team, 2026c) to judge semantic correctness against reference answers. Counting uses integer parsing and exact match; unparseable outputs (0.9%) are incorrect. Captioning reports semantic GOOD rate, with NLTK METEOR (Banerjee and Lavie, 2005) (Ć100Ć 100) used only as a lexical diagnostic. Macro averages cover recognition, counting, and VQA; captioning is reported separately. VQA equally macro-averages attribute, existence, and counting scores, so its values need not be integer multiples of 1/1501/150; reported improvements are computed from unrounded aggregate scores. To assess possible same-family judge bias, we independently re-evaluate a stratified set of 300 recognition/VQA outputs with Kimi-K3 (Kimi Team, 2026). The judges achieve 94.3% agreement and Cohenās Īŗ=0.87Īŗ=0.87, and all model rankings remain unchanged after manual inspection of disagreements. Baselines. We compare against Qwen3-VL-8B, InternVL3.5-8B (Wang et al., 2025), MiniCPM-V 4.5 (Yu et al., 2025), the larger Qwen3.6-35B-A3B and Qwen3.8-27B models (Qwen Team, 2026a; Qwen Team, 2026b), and the underwater-domain OceanGPT-o-7B and NAUTILUS-7B models. Qwen3-VL-8B+LoRA uses the same Stage-IV data and LoRA configuration as SonarLLM, providing a controlled test of whether instruction tuning alone explains the gains. For fusion input, all baselines use their native multi-image interface with a fixed sonarāoptical order and explicit modality labels. Implementation Details. SonarLLM uses Qwen3-VL-8B with 448Ć448448Ć 448 inputs. The language model is frozen in Stages IāI. Stage I uses a 0.60 masking ratio. Stage I optimizes PSVT and the Sonar Stem at a learning rate of 5Ć10ā55Ć 10^-5. Stage I degrades the optical input with probability 0.70.7, samples Ī·ā¼ā”(0.1,1.0)Ī· (0.1,1.0) for degraded examples, and jointly trains the sonar pathway, both VFEs, and AGFM at a learning rate of 10ā410^-4. Stage IV uses LoRA with r=128r=128, α=256α=256, and learning rate 10ā410^-4. The architecture and Stage-IV adapters add 882.4M and 349.2M parameters, respectively, resulting in 10.3B total parameters versus 8.77B for the backbone when the frozen DINOv2-L structural-prior encoder is included. 4.2. Main Results Table 3. SonarBench results (%; 25 subsets, n=150n=150 each). Recognition/VQA use semantic judging, counting uses exact match, and ā” denotes caption GOOD rate. Averages exclude captioning. Best per row in bold. Modality Task (Condition) Qwen3-VL -8B Qwen3-VL -8B+LoRA InternVL 3.5-8B MiniCPM-V 4.5 NAUTILUS -7B OceanGPT -o-7B Qwen3.6 -35B-A3B Qwen3.8 -27B Sonar LLM Sonar Recognition 25.3 28.7 29.3 23.3 14.0 2.0 27.3 32.0 64.7 Counting 40.7 51.3 33.3 47.3 32.7 14.0 35.3 32.7 84.7 VQA 30.2 32.6 27.5 21.7 25.8 20.8 24.2 22.5 66.5 Captionā” 8.0 4.0 7.3 4.7 12.0 5.3 10.0 4.7 33.3 Optical Recognition (Clear) 41.3 38.7 38.7 36.0 34.7 10.7 44.0 40.0 77.3 Recognition (Turbid) 37.3 36.0 36.0 35.3 36.0 12.7 40.7 38.7 63.3 Recognition (Heavy) 38.0 35.3 33.3 29.3 30.7 9.3 34.7 39.3 44.7 Counting (Clear) 71.3 69.3 60.7 66.0 59.3 40.7 73.3 72.7 66.0 Counting (Turbid) 37.3 39.3 26.0 32.0 24.0 22.7 29.3 22.7 37.3 Counting (Heavy) 8.0 16.0 6.7 10.0 6.0 10.0 10.0 6.0 36.0 VQA (Clear) 40.9 37.0 37.0 41.1 32.6 22.9 38.7 36.8 47.5 VQA (Turbid) 34.7 32.5 30.8 32.0 28.2 21.5 32.8 28.5 44.7 VQA (Heavy) 28.1 28.2 26.3 22.9 20.7 20.3 24.1 24.7 41.2 Caption (Clear)ā” 23.3 26.7 25.3 22.0 18.7 16.0 26.7 22.0 29.3 Caption (Turbid)ā” 24.7 26.0 24.7 25.3 22.7 12.0 24.7 21.3 28.7 Caption (Heavy)ā” 15.3 9.3 10.0 10.7 14.0 6.0 13.3 12.0 14.7 Fusion Recognition (Clear) 37.3 30.7 34.7 34.7 34.0 15.3 40.7 40.7 83.3 Recognition (Turbid) 36.0 30.0 36.0 39.3 33.3 15.3 38.7 40.0 76.7 Recognition (Heavy) 39.3 30.0 33.3 32.0 30.7 10.7 37.3 38.7 80.7 Counting (Clear) 66.7 36.0 65.3 62.0 62.7 52.7 23.3 63.3 72.0 Counting (Turbid) 44.7 36.0 59.3 51.3 63.3 52.0 25.3 61.3 70.7 Counting (Heavy) 47.3 36.0 48.7 36.0 62.0 55.3 24.0 40.0 72.0 VQA (Clear) 39.1 30.3 33.9 38.0 38.0 11.6 25.6 38.5 54.6 VQA (Turbid) 33.7 31.5 34.4 26.9 28.9 15.5 24.7 36.1 53.5 VQA (Heavy) 26.4 30.1 28.6 20.4 22.6 13.5 20.3 34.5 55.2 Sonar Average 32.1 37.5 30.0 30.8 24.2 12.3 28.9 29.1 72.0 Optical Average 37.4 36.9 32.8 33.8 30.2 19.0 36.4 34.4 50.9 Fusion Average 41.2 32.3 41.6 37.8 41.7 26.9 28.9 43.7 68.7 Native Sonar Understanding. SonarLLM obtains 72.0% sonar-only macro accuracy, exceeding the strongest baseline by 34.4 points. The gain is consistent across recognition (64.7%), counting (84.7%), and VQA (66.5%), and its caption GOOD rate reaches 33.3% versus 12.0% for the best baseline. Increasing general-purpose capacity does not close the gap: Qwen3.6-35B-A3B and Qwen3.8-27B achieve only 28.9% and 29.1%, respectively. Underwater-domain models are also substantially weaker, supporting the need for sonar-native representation rather than model scaling or domain knowledge alone. Optical and Fusion Performance. SonarLLM achieves 50.9% on optical input and ranks first on every recognition and VQA subset. Its main optical weakness is clear/turbid counting, where it trails the best baselines. With both sensors, however, SonarLLM reaches 68.7%, 25.1 points above the strongest baseline. Fusion recognition remains between 76.7% and 83.3% across all degradation levels. In contrast, heterogeneous multi-image input alone is unreliable: Qwen3-VL-8B+LoRA drops from 36.9% optical accuracy to 32.3% under fusion, and Qwen3.6-35B-A3B drops from 36.4% to 28.9%. Dedicated heterogeneous representation and interaction therefore provide substantial benefits when exploiting the second sensor. Controlling Instruction Tuning and Sampling Uncertainty. Qwen3-VL-8B+LoRA uses the same Stage-IV data and language-side LoRA as SonarLLM. Relative to this control, SonarLLM improves sonar, optical, and fusion accuracy by 34.4, 14.0, and 36.5 points, showing that instruction tuning alone does not reproduce the observed gains. A 10,000-replication scene-clustered bootstrap gives 95% confidence intervals of [28.1,40.5][28.1,40.5], [11.3,16.7][11.3,16.7], and [34.0,39.0][34.0,39.0] for these gains; all lower bounds remain positive. These intervals quantify test-set sampling uncertainty but not variation across independent training runs. 4.3. Controlled Degradation and Reliability-Aware Fusion Because optical and fusion evaluations share scenes, questions, and sonar observations, fusion minus optical accuracy provides a paired estimate of the benefit of adding sonar. Table 4 summarizes this controlled comparison. The fusion-over-optical gain expands from 6.0 to 36.0 points for recognition and counting, and from 7.1 to 14.0 points for VQA. Meanwhile, clear-to-heavy fusion changes are limited to ā2.6-2.6, 0.00.0, and +0.6+0.6 points, respectively, despite much larger optical declines. For recognition, fusion also exceeds the stronger individual sensor by 6.0, 12.0, and 16.0 points across clear, turbid, and heavy conditions, showing that sonar increasingly complements rather than merely replaces optical sensing. Table 4. Controlled degradation summary derived from Table 3 (percentage points). Ī (FāO) is fusion minus optical accuracy; C/T/H denote clear/turbid/heavy conditions. Task Ī (FāO) C/T/H Optical Cā Fusion Cā Recognition +6.0/+13.4/+36.0+6.0/+13.4/+36.0 ā32.6-32.6 ā2.6-2.6 Counting +6.0/+33.4/+36.0+6.0/+33.4/+36.0 ā30.0-30.0 0.00.0 VQA +7.1/+8.8/+14.0+7.1/+8.8/+14.0 ā6.3-6.3 +0.6+0.6 Figure 3. Reliability-aware fusion analysis on 150 held-out scenes. (a) Sonar weight increases with optical degradation. (b) Under the same checkpoint, AGFMās advantage over equal weighting grows from 0.0 to 2.7 points. We next evaluate whether AGFM exhibits the intended internal response on an independent probe set of 150 RGBS50 scenes, each rendered at five degradation levels (750750 observations). As shown in Fig. 3(a), mean sonar weight increases from 0.491 to 0.637, 0.717, 0.769, and 0.795. Severity and sonar weight have Spearman correlation Ļ=0.749Ļ=0.749 (scene-blocked 95% CI [0.68,0.81][0.68,0.81]), and 95.3% of scenes show non-decreasing sonar weight. Because Stage I explicitly supervises the gate using degradation strength, we interpret this trend as verification of AGFMās intended internal response, rather than as independent evidence of emergent reliability estimation. Using the same checkpoint, we replace AGFM by equal weighting at inference. The accuracy advantage is 0.0, 0.2, and 2.7 points under clear, turbid, and heavy conditions, respectively; clear-to-heavy degradation is 0.7 points with AGFM versus 3.4 points with equal weighting. This pattern clarifies that complementarity is not an unconditional advantage of using more sensors. Recognition benefits consistently because optical appearance and acoustic contour and range provide distinct evidence, whereas fusion counting and VQA do not always exceed sonar-only performance when the answer depends primarily on geometry or degraded optical cues become distracting. AGFM is therefore not an oracle that selects the best modality for every question; it acts mainly as a degradation buffer that prevents a deteriorating sensor from dominating the shared representation. This distinction separates representation from allocation: the sonar pathway determines which acoustic evidence is available to language reasoning, while AGFM controls its influence as observation reliability shifts. The following ablations examine these responsibilities at the representation and interaction levels. The controlled protocol does not cover non-uniform scattering, dynamic suspended particles, sonar-specific failures, or temporal misalignment; evaluation on naturally degraded paired observations remains necessary. 4.4. Ablation and Representation Analysis Ablations use a diagnostic split constructed from video sequences disjoint from both training and the main SonarBench test set, with 25 modalityātaskācondition combinations and 150 samples each (N=3,750N=3,750). Recognition and VQA use rule-based scoring rather than the semantic judge in Table 3; absolute values across the two tables are not directly comparable. Domain Adaptation and Representation Formation. Table 5. Sonar front-end adaptation on the rule-scored validation split. R/C avg. is the mean of recognition and counting accuracy; METEOR is reported in %. Variant R/C avg. Rec. Count METEOR Transplant (no MAE) 72.0 64.0 80.0 45.1 MAE-rā64r64 69.3 61.3 77.3 46.8 MAE-rā128r128 74.7 71.3 78.0 49.2 As shown in Table 5, Stage-I MAE with the same r=128r=128 language adaptation improves the sonar R/C average from 72.0% to 74.7%, recognition by 7.3 points, and METEOR by 4.1, although counting decreases by 2.0 points. MAE-rā64r64 reaches only 69.3%, indicating that adapted acoustic features also require sufficient language-side capacity. Mask ratios and Stage-I data scale are examined separately in Table 6. Table 6. Sensitivity of the sonar R/C average (%) to Stage-I masking and data scale. Factor Tested values R/C avg. Mask ratio 0.50 / 0.60 / 0.75 73.3 / 74.7 / 69.3 Unlabeled images 10K / 20K / 40K 70.0 / 72.7 / 74.7 An intermediate masking ratio performs best: excessive masking removes sparse acoustic structure, whereas insufficient masking weakens the adaptation signal. Increasing unlabeled data produces consistent but diminishing gains; we therefore use a 0.60 mask ratio and 40K images. Figure 4. Progressive representation formation. (a) Sonar frames processed by the pretrained optical tower. (b) Class structure after Stages IāI. (c) Similarity of matched and mismatched sonarāoptical pairs before and after alignment and instruction tuning. Panel (c) uses 150 matched same-scene pairs and 22,350 mismatched cross-scene pairs under the clear condition; intervals denote 95% confidence intervals for the mean-similarity margin. Figure 4 shows that the silhouette score of sonar categories increases from ā0.18-0.18 under the pretrained optical tower to 0.160.16 after Stages IāI. However, matched and mismatched sonarāoptical pairs remain nearly indistinguishable at this point (Īalign=0.001 _align=0.001, 95% CI [0.000,0.001][0.000,0.001]). Stage I raises the alignment margin to 0.660 ([0.638,0.671][0.638,0.671]), supporting the conclusion that cross-modal correspondence is learned beyond unimodal sonar structure. After Stage IV, the margin remains substantial at 0.423 ([0.404,0.435][0.404,0.435]), indicating that instruction tuning relaxes strict feature similarity while preserving task-relevant correspondence. Structural Components. Table 7. Structural ablation on the rule-scored validation split (%). Fusion, Optical, and Sonar Macro average recognition, counting, and VQA. Drop: clear-to-heavy fusion decline. ā : sonar front-end retraining required. Mean Fusion is an independently retrained equal-weight variant. Variant Fusion Drop Optical Sonar Macro SonarLLM 82.1 1.3 72.5 65.8 Mean Fusion 80.8 3.8 72.3 65.6 w/o gate loss 81.2 3.0 72.7 64.9 w/o both VFE 79.8 3.1 67.2 61.4 w/o Optical-VFE 80.7 2.6 68.8 65.4 w/o Acoustic-VFE 80.4 2.4 72.1 62.0 w/o DeepStack 78.9 2.8 71.8 59.6 w/o Sonar Stemā 73.9 5.3 71.6 58.3 w/o Polar PEā 68.6 7.1 71.4 55.7 Table 7 identifies sonar geometry as the dominant factor: removing Polar PE reduces fusion/sonar-macro accuracy by 13.5/10.1 points, and removing the Sonar Stem reduces them by 8.2/7.5 points. The VFEs exhibit the intended modality specificity: removing Optical-VFE mainly reduces optical accuracy (3.7 points), whereas removing Acoustic-VFE mainly reduces sonar-macro accuracy (3.8 points). Removing DeepStack causes a 6.2-point sonar-macro loss, showing the importance of intermediate acoustic structure. Removing both VFEs causes consistent losses of 2.3, 5.3, and 4.4 points on fusion, optical, and sonar inputs, respectively, confirming complementary rather than redundant modality-specific corrections. Quality-aware weighting contributes most strongly to robustness. Mean Fusion and w/o gate loss reduce static fusion accuracy by only 1.3 and 0.9 points, but increase the clear-to-heavy decline from 1.3 to 3.8 and 3.0 points. Together with the same-checkpoint comparison in Fig. 3, this consistently supports AGFM as a quality-aware adaptation mechanism. The Sonar Stem and Polar PE variants require front-end retraining and are therefore not strict inference-time ablations, but their consistent losses across sonar, fusion, and degradation sensitivity support their necessity. Together, the ablations reveal complementary roles: the Sonar Stem and Polar PE establish acoustic representations, the VFEs correct modality-specific corruption, DeepStack preserves intermediate structure, and AGFM stabilizes fusion under asymmetric observation quality. 4.5. Efficiency Table 8. Inference efficiency on one A100-80GB GPU. Results use BF16, FlashAttention-2, batch size 1, two 448Ć448448Ć 448 images, 512 visual tokens, and 128 generated tokens; all SonarLLM measurements and latency and throughput are mean± over 10 runs. Model Params Peak mem. Prefill (ms) Decode (tok/s) Qwen3-VL-8B 8.77B 17.7GB 77.2±0.277.2±0.2 38.4±0.138.4±0.1 SonarLLM 10.3B 21.6GB 137.5±1.1137.5±1.1 21.1±1.221.1±1.2 Table 8 shows the cost of native sonar processing: parameters and peak memory rise by 17.4% and 22.0%, prefill latency by 78.1%, and decoding throughput falls by 45.1%. Most added capacity is inherited: the 736.9M sonar tower is pretrained and frozen DINOv2-L contributes 0.30B parameters; only 25.1M non-LoRA adaptation and fusion parameters are randomly initialized. 5. Conclusion This work shows that robust sonarāoptical reasoning requires representations tailored to each modality and fusion conditioned on observation quality. SonarLLM combines a sonar-native pathway with SonarBenchās paired protocol, which varies optical quality while holding the scene and sonar observation fixed. It achieves 72.0% sonar-only macro accuracy and 68.7% fusion accuracy, 34.4 and 25.1 points above the strongest baselines. The fusion-over-optical gain rises from +6.0+6.0 to +36.0+36.0 points as optical quality deteriorates, showing that sonar becomes increasingly valuable as optical reliability falls. These results support representing heterogeneous sensors according to their observation processes and weighting them according to observation quality. The controlled protocol does not fully validate naturally occurring turbidity; future work will extend to real paired data, task-aware routing, and temporal sonarāoptical reasoning. References Alawode et al. (2025) B. O. Alawode, I. I. Ganapathi, S. Javed, N. Werghi, M. Bennamoun, and A. Mahmood AquaticCLIP: a vision-language foundation model for underwater scene analysis. External Links: 2502.01785 Cited by: §2.1. Alayrac et al. (2022) J. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Millican, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 35, p. 23716ā23736. Cited by: §2.2. Arevalo et al. (2017) J. Arevalo, T. Solorio, M. Montes-y-Gómez, and F. A. GonzĆ”lez Gated multimodal units for information fusion. In ICLR 2017 Workshop Track, Cited by: §2.2. Bai et al. (2025) S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, et al. Qwen3-vl technical report. External Links: 2511.21631 Cited by: §1, §2.2, §3.1. Banerjee and Lavie (2005) S. Banerjee and A. Lavie METEOR: an automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, Ann Arbor, Michigan, p. 65ā72. External Links: Link Cited by: §4.1. Bi et al. (2024) Z. Bi, N. Zhang, Y. Xue, Y. Ou, D. Ji, G. Zheng, and H. Chen OceanGPT: a large language model for ocean science tasks. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 3357ā3372. External Links: Document, Link Cited by: §2.1. Chen et al. (2026) W. Chen, P. Tinn, P. G. Auran, M. Ludvigsen, and P. H. Haro A sonar-visual dataset for cross-modal underwater robot perception. External Links: 2606.01398, Document Cited by: §2.1. Hu et al. (2021) E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen LoRA: low-rank adaptation of large language models. External Links: 2106.09685, Document Cited by: §3.5. Islam et al. (2020) M. J. Islam, Y. Xia, and J. Sattar Fast underwater image enhancement for improved visual perception. IEEE Robotics and Automation Letters 5 (2), p. 3227ā3234. Cited by: §2.2. Kimi Team (2026) Kimi Team Kimi K3: open frontier intelligence. External Links: 2607.24653 Cited by: §4.1. Li et al. (2020) C. Li, C. Guo, W. Ren, R. Cong, J. Hou, S. Kwong, and D. Tao An underwater image enhancement benchmark dataset and beyond. IEEE Transactions on Image Processing 29, p. 4376ā4389. External Links: Document Cited by: §1, §2.2. Li et al. (2023) J. Li, D. Li, S. Savarese, and S. Hoi BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the 40th International Conference on Machine Learning (ICML), p. 19730ā19742. Cited by: §2.2. Li et al. (2025) Y. Li, B. Wang, J. Sun, X. Wu, and Y. Li RGB-sonar tracking benchmark and spatial cross-attention transformer tracker. IEEE Transactions on Circuits and Systems for Video Technology 35 (3), p. 2260ā2275. External Links: Document Cited by: §1, §1, §2.1. Liu et al. (2023a) H. Liu, C. Li, Y. Li, and Y. J. Lee Improved baselines with visual instruction tuning. External Links: 2310.03744 Cited by: §2.2. Liu et al. (2023b) H. Liu, C. Li, Q. Wu, and Y. J. Lee Visual instruction tuning. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 36. Cited by: §2.2. Neupane and Seok (2020) D. Neupane and J. Seok A review on deep learning-based approaches for automatic sonar target recognition. Electronics 9 (11), p. 1972. External Links: Document, Link Cited by: §1, §1, §2.1. Oquab et al. (2023) M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P. Huang, S. Li, I. Misra, M. Rabbat, V. Sharma, G. Synnaeve, H. Xu, H. JĆ©gou, J. Mairal, P. Labatut, A. Joulin, and P. Bojanowski DINOv2: learning robust visual features without supervision. External Links: 2304.07193, Document Cited by: §3.3. Park et al. (2025) K. Park, Y. Kim, D. Kim, and J. W. Choi Resilient sensor fusion under adverse sensor failures via multi-modal expert fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), p. 6720ā6729. Cited by: §1, §2.2. Qwen Team (2026a) Qwen Team Qwen3.6-35B-A3B: agentic coding power, now open to all. Note: Official model releaseAccessed: 2026-08-24 External Links: Link Cited by: §4.1. Qwen Team (2026b) Qwen Team Qwen3.8-27B. Note: Official model cardAccessed: 2026-08-24 External Links: Link Cited by: §4.1. Qwen Team (2026c) Qwen Team Qwen3.8-Max: a new bar for coding and cowork. Note: Official model releaseAccessed: 2026-08-24 External Links: Link Cited by: §4.1. Steiniger et al. (2022) Y. Steiniger, D. Kraus, and T. Meisen Survey on deep learning based computer vision for sonar imagery. Engineering Applications of Artificial Intelligence 114, p. 105157. Cited by: §1, §2.1. Wang et al. (2025) W. Wang, Z. Gao, L. Gu, et al. InternVL3.5: advancing open-source multimodal models in versatility, reasoning, and efficiency. External Links: 2508.18265 Cited by: §4.1. Wu et al. (2026) Y. Wu, W. Wang, C. Lin, M. Hou, and M. Liu Towards multimodal underwater object detection: a bidirectional feature recomposition network and visual-sonar dataset. Expert Systems with Applications 316, p. 131710. External Links: Document Cited by: §1, §2.1. Xie et al. (2022) K. Xie, J. Yang, and K. Qiu A dataset with multibeam forward-looking sonar for underwater object detection. Scientific Data 9 (1), p. 739. External Links: Document, Link Cited by: §2.1. Xu et al. (2025) W. Xu, C. Wang, D. Liang, Z. Zhao, X. Jiang, P. Zhang, and X. Bai NAUTILUS: a large multimodal model for underwater scene understanding. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §1, §2.1, §2.2. Xue et al. (2025) Y. Xue, M. Mao, X. Ru, Y. Zhu, B. Ren, S. Qiao, M. Wang, S. Deng, X. An, N. Zhang, Y. Chen, and H. Chen OceanGym: a benchmark environment for underwater embodied agents. External Links: 2509.26536 Cited by: §2.1, §4.1. Xue et al. (2026) Y. Xue, N. Zhang, T. Wu, Z. Ma, D. Ji, Z. Wang, G. Zheng, and H. Chen OceanPile: a large-scale multimodal ocean corpus for foundation models. External Links: 2605.00877 Cited by: §2.1, §4.1. Yu et al. (2025) T. Yu, Z. Wang, C. Wang, et al. MiniCPM-V 4.5: cooking efficient MLLMs via architecture, data, and training recipe. External Links: 2509.18154 Cited by: §4.1. Zhang et al. (2025) D. Zhang, C. Rong, B. Li, F. Wang, Z. Zhao, J. Gao, and X. Li UWBench: a comprehensive vision-language benchmark for underwater understanding. External Links: 2510.18262 Cited by: §2.2. Zhang et al. (2022) P. Zhang, J. Tang, H. Zhong, M. Ning, D. Liu, and K. Wu Self-trained target detection of radar and sonar images using automatic deep learning. IEEE Transactions on Geoscience and Remote Sensing 60, p. 4701914. External Links: Document Cited by: §2.1. Zheng et al. (2023) Z. Zheng, J. Zhang, T. Vu, S. Diao, Yue Him Wong Tim, and S. Yeung MarineGPT: unlocking secrets of ocean to the public. External Links: 2310.13596 Cited by: §1, §2.1.