Paper deep dive
Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations
Axi Niu, Jieheng Li, Kang Zhang, Qingsen Yan, Jinqiu Sun, Yanning Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 89%
Last extracted: 8/4/2026, 4:35:05 AM
Summary
The paper introduces TGFusion, a text-guided latent-space flow matching framework for Infrared-Visible Image Fusion (IVIF) under complex degradations. It addresses the limitation of fixed global text representations by employing a Prompt-conditioned Multi-stream Joint Flow Transformer. This architecture treats text as an independent semantic stream that interacts bidirectionally with visual streams (visible, infrared, and fusion) via joint attention, allowing dynamic adaptation to spatially varying degradations. Experiments on public benchmarks (MSRS, LLVIP, M3FD, TNO) and the DDL-12 degradation benchmark demonstrate superior perceptual quality, structural preservation, and robustness compared to existing methods.
Entities (14)
Relation Signals (13)
TGFusion → solves → Infrared-Visible Image Fusion
confidence 95% · Infrared-visible image fusion under realistic degradation scenarios is a challenging task... we propose TGFusion
TGFusion → uses → Prompt-conditioned Multi-stream Joint Flow Transformer
confidence 95% · we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream
TGFusion → employs → Flow Matching
confidence 90% · TGFusion, a text-guided latent-space flow matching framework
Prompt-conditioned Multi-stream Joint Flow Transformer → enables → Joint Attention
confidence 90% · Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations
TGFusion → evaluatedon → LLVIP
confidence 90% · The evaluation covers four public IVIF benchmarks: ... LLVIP
TGFusion → evaluatedon → DDL-12
confidence 90% · Training is conducted on DDL-12... Extensive experiments on public benchmarks and complex degradation scenarios
TGFusion → evaluatedon → MSRS
confidence 90% · The evaluation covers four public IVIF benchmarks: MSRS
TGFusion → evaluatedon → M3FD
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Tags
Links
- Source: https://arxiv.org/abs/2608.00530v1
- Canonical: https://arxiv.org/abs/2608.00530v1
Trouble viewing inline? Open PDF directly →
Full Text
47,143 characters extracted from source content.
Expand or collapse full text
Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations Axi Niu1, Jieheng Li1 , Kang Zhang2, Qingsen Yan1, Jinqiu Sun3, Yanning Zhang1 Abstract Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations. 1 Introduction Infrared-visible image fusion (IVIF) aims to integrate complementary information from infrared and visible images into a unified representation (Liu et al. 2025). Infrared images emphasize thermally salient targets and remain reliable under illumination variations, whereas visible images provide fine textures and structural details essential for human and machine perception. By combining these modality-specific strengths, IVIF supports a broad range of applications, including surveillance (Zhang et al. 2018), autonomous driving (Bao et al. 2023), object perception (Jain et al. 2023), and downstream semantic understanding (Zhang et al. 2023). Deep fusion networks based on CNNs (Liang et al. 2022; Liu et al. 2022), autoencoders (Li and Wu 2018; Li et al. 2023), generative models (Ma et al. 2019), and Transformers (Ma et al. 2022; Zhang et al. 2022) have substantially improved feature extraction, cross-modal interaction, and reconstruction. However, most are developed under relatively ideal imaging conditions and implicitly assume that source images provide reliable modality cues (Zhao et al. 2023a, 2024). In real-world scenarios, infrared and visible images always undergo diverse modality-specific and composite degradations of varying severity, which existing fusion methods often fail to handle, thereby degrading fusion quality and limiting robustness (Yi et al. 2024; Tang et al. 2025). A straightforward strategy is to restore the degraded input from each modality before fusion. However, this cascaded pipeline incurs additional computational overhead and decouples restoration from fusion, hindering their end-to-end joint optimization (Cao et al. 2025). Recent degradation-aware and text-guided methods encode degradation states or fusion preferences in text and incorporate these semantics into fusion either by modulating visual features with fixed textual representations (Tang et al. 2025) or by converting them into object-level spatial priors for semantic-aware fusion (Zhang et al. 2025). Such textual cues provide information beyond corrupted visual observations and help identify reliable modality-specific evidence. Despite this benefit, language guidance is generally defined in advance and passed to the visual pathway in a largely one-way manner. Because it is not revised according to the current fusion state, it may not adequately reflect spatially nonuniform corruption or local content variations. Text therefore serves mainly as auxiliary conditioning instead of evolving together with the fused representation during generation. As illustrated in Figure 1, ControlFusion (Tang et al. 2025), which relies on fixed textual conditioning, and OmniFuse (Zhang et al. 2025), which translates textual semantics into object-level spatial priors, still leave residual degradation artifacts, obscure local structures, and attenuate infrared targets under degradations. These distortions degrade task-relevant visual cues and can propagate to downstream perception, resulting in erroneous predictions. Figure 1: Overview of TGFusion (top) and qualitative comparisons under degradation scenarios (bottom). Motivated by these limitations, TGFusion formulates IVIF with degraded inputs as conditional transport in a compact latent space, using structured textual descriptions as semantic priors that guide the generative trajectory from noise to the target fusion representation in a compact VAE latent space. Throughout this trajectory, linguistic guidance is continually adapted to the evolving latent state and modality-specific evidence. This progressive semantic–visual interaction links textual priors to scene structures and thermally salient regions, enabling reliable cue selection and fusion representation generation. As shown in Figure 1, TGFusion produces robust fusion results across diverse degraded conditions. In summary, our main contributions are as follows: • We present TGFusion, which formulates degradation-robust IVIF as text-guided latent-space flow matching and integrates corruption handling with fusion in a single generative process. • We develop a Prompt-conditioned Multi-stream Joint Flow Transformer that organizes text as an independent semantic stream alongside the fusion, visible, and infrared streams. Joint attention supports token-level interaction and layer-wise updating across the four streams, overcoming the limitations of fixed global textual conditioning. • TGFusion consistently outperforms representative two-stage restoration–fusion pipelines and recent all-in-one degradation-aware methods under various degradation conditions. Superior downstream semantic segmentation performance further validates its effectiveness in supporting high-level perception. 2 Related Work Figure 2: Overview of the proposed TGFusion framework Figure 3: Detailed architecture of the proposed Prompt-conditioned Multi-stream Joint Flow Transformer. Text-Guided Image Fusion Text provides high-level priors that complement visual observations and support degradation-aware or preference-controlled fusion. With vision–language encoders such as CLIP (Radford et al. 2021), existing approaches mainly follow two directions. The first describes imaging conditions to improve robustness. Text-IF (Yi et al. 2024) uses task and degradation descriptions to regulate restoration and fusion, while ControlFusion (Tang et al. 2025) introduces language–vision degradation prompts that characterize degradation categories and severities under compound corruptions. The second uses text to specify semantic preferences or regions of interest. Text-DiFuse (Zhang et al. 2024a) enhances text-relevant foreground regions through feature remodulation; TextFusion (Cheng et al. 2025) establishes text–region associations for preference-aware fusion; and OmniFuse (Zhang et al. 2025) combines linguistic semantics with spatial localization for compound-degradation scenarios. Collectively, these methods establish text as a flexible interface for describing degradation states and fusion preferences. However, its role is generally confined to auxiliary conditioning rather than participating in progressive multimodal integration. Diffusion and Flow Matching for Image Fusion Diffusion models formulate image fusion as conditional generation, providing strong priors for structural plausibility and perceptual quality. DDFM (Zhao et al. 2023b) combines diffusion posterior sampling with source-image constraints, while CCF (Cao et al. 2024) adaptively selects fusion constraints from a condition bank during sampling. For corrupted inputs, DRMF (Tang et al. 2024) composes modality-specific degradation priors, extending diffusion fusion to degradation-robust settings. Recent studies move generative fusion into compact latent spaces. OmniFuse (Zhang et al. 2025) combines latent diffusion with text-driven semantics for compound degradations. MMAIF (Cao et al. 2025) includes a latent-space flow-matching variant (Lipman et al. 2023) for text-guided fusion across multiple tasks and degradation types. Overall, generative fusion has progressed from diffusion-based constrained sampling to latent diffusion and flow matching, enabling more flexible modeling of multimodal fusion distributions. Building on this progression, TGFusion adopts latent flow matching as its generative foundation and integrates textual guidance throughout latent evolution. 3 Method Preliminaries Conditional flow matching. Conditional flow matching learns a time-dependent velocity field that transports a tractable noise distribution to a target data distribution (Lipman et al. 2023). Let Z0∼(0,)Z_0 (0,I) denote noise and (Z1,C)(Z_1,C) denote a target sample with condition C. For t∼[0,1]t [0,1], the linear probability path and its target velocity are Zt=(1−t)Z0+tZ1,ut=dZtdt=Z1−Z0.Z_t=(1-t)Z_0+tZ_1, u_t= dZ_tdt=Z_1-Z_0. (1) The learned conditional velocity field vθ(Zt,t,C)v_θ(Z_t,t,C) is optimized by ℒflow=t,Z0,Z1,C[‖vθ(Zt,t,C)−(Z1−Z0)‖22].L_flow=E_t,Z_0,Z_1,C [ \|v_θ(Z_t,t,C)-(Z_1-Z_0) \|_2^2 ]. (2) At inference, the learned ODE is integrated from t=0t=0 to t=1t=1 to transform Gaussian noise into a target sample. Latent-space generative modeling. For computational efficiency, a pretrained VAE encoder ℰE maps images into compact latent representations, where flow matching based generation is performed, and its decoder D reconstructs the generated latent into image space. Overview Given a degraded visible image IvdI_v^d, a degraded infrared image IrdI_r^d, and a structured prompt p, TGFusion θG_θ aims to generate the fused infrared-visible image I^f I_f I^f=θ(Z0,Ivd,Ird,p),whereZ0∼(0,). I_f=G_θ (Z_0,I_v^d,I_r^d,p ), ~Z_0 (0,I). (3) Since image fusion has no unique ground-truth target, we use a frozen Text-IF model (Yi et al. 2024) to generate a pseudo fusion target IfI_f from the corresponding clean source pair. As illustrated in Figure 2, a frozen VAE encodes IvdI_v^d, IrdI_r^d, and IfI_f to get the latent-space samples ZvZ_v, ZrZ_r, and Z1Z_1. Meanwhile, a frozen CLIP text encoder converts the prompt p into the semantic condition c. The Prompt-conditioned Multi-stream Joint Flow Transformer then predicts the conditional velocity field from ZtZ_t, conditioned on ZvZ_v, ZrZ_r, and c. Degradation Semantic Prior Modeling We construct each prompt from task, modality-specific degradation, and generation-goal descriptions. The degradation description specifies the types and severities of the source images. ChatGPT (OpenAI 2023) is used to generate diverse linguistic variants. The prompt is formulated as p=qtask⊕qdeg⊕qgoal,p=q_task q_deg q_goal, (4) where ⊕ denotes textual concatenation. During training, prompt variants consistent with the annotated degradation states are randomly sampled, while fixed variants are used during inference for reproducibility. A frozen CLIP (Radford et al. 2021) text encoder τ(⋅)τ(·) maps p into token-level representations: c=τ(p)=[c1,c2,…,cL],c=τ(p)=[c_1,c_2,…,c_L], (5) where L denotes the token-sequence length. The complete token sequence is preserved as the textual input to the subsequent flow Transformer. Prompt-conditioned Multi-stream Joint Flow Transformer The Prompt-conditioned Multi-stream Joint Flow Transformer parameterizes the flow dynamics. As shown in Figure 3, the proposed transformer preserves stream-specific representations, aggregates the timestep embedding and condition summaries into a global condition gtg_t, and progressively exchanges information across streams through joint attention, stream-wise AdaLN, and gated residual updates. The refined fusion stream is then projected to u^t u_t. Multi-stream input embedding and global-condition construction. Let ℳ=fuse,vis,ir,textM=\fuse,vis,ir,text\ denote the four streams. The visual latents ZtZ_t, ZvZ_v, and ZrZ_r are flattened into spatial token sequences, and c is the CLIP token sequence for global semantic prior. Stream-specific projections Πfuse ^fuse, Πvis ^vis, Πir ^ir, and Πtext ^text map them to the initial representations F0fuseF_0^fuse, F0visF_0^vis, F0irF_0^ir, and F0textF_0^text, respectively, within a shared hidden space. The fusion stream represents the generative state, while the remaining streams provide visual and semantic conditions. A global condition vector is constructed as gt=ψ(t)+MLP(∑m∈vis,ir,textPmρ(F0m)),g_t=ψ(t)+MLP ( _m∈\vis,ir,text\P^mρ(F_0^m) ), (6) where ψ(t)ψ(t) is the time embedding, ρ(⋅)ρ(·) is token-wise average pooling, and PmP^m is a stream-specific projection. Pooling is used only for AdaLN conditioning; the complete token sequences remain available to joint attention. Prompt-conditioned joint block. For each block l∈0,…,Lb−1l∈\0,…,L_b-1\, the global condition gtg_t is mapped to stream-specific AdaLN (Peebles and Xie 2023) parameters: (bl,am,sl,am,γl,am,bl,fm,sl,fm,γl,fm)=Γlm(gt),m∈ℳ.(b_l,a^m,s_l,a^m, _l,a^m,b_l,f^m,s_l,f^m, _l,f^m)= _l^m(g_t), m . (7) After AdaLN modulation, each stream produces queries, keys, and values. Two-dimensional rotary position encoding (Su et al. 2024) is applied to the fusion, visible, and infrared streams, while the text stream retains the positional information encoded by CLIP. The resulting QKV representations are concatenated along the token dimension as QlQ_l, KlK_l, and VlV_l, and joint attention is computed by Ol=Softmax(QlKl⊤dh)Vl,O_l=Softmax ( Q_lK_l d_h )V_l, (8) where dhd_h is the head dimension. The output OlO_l is subsequently split according to the original token ranges, and each stream is updated through its own output projection, gated residual connection, and FFN. This joint operation enables layer-wise bidirectional interaction among semantic and visual representations while preserving stream-specific transformation paths. Conditional velocity-field prediction. After LbL_b blocks, the fusion stream predicts the conditional velocity: u^t=Wo((1+so)⊙LN(FLbfuse)+bo), u_t=W_o ((1+s_o) (F_L_b^fuse)+b_o ), (9) where (bo,so)(b_o,s_o) are generated from gtg_t, and WoW_o restores the latent spatial layout. The remaining streams serve only as jointly updated conditional contexts and have no separate prediction heads. MSRS LLVIP M3FD Methods EN AG C CLIP-IQA TReS NIQE EN AG C CLIP-IQA TReS NIQE EN AG C CLIP-IQA TReS NIQE CDDFuse 6.70 3.73 0.60 0.14 31.24 3.92 7.35 4.30 0.69 0.37 57.05 4.11 6.90 4.86 0.54 0.29 61.16 4.22 EMMA 6.72 3.78 0.60 0.18 27.01 4.57 7.35 4.67 0.69 0.40 51.93 4.23 6.92 5.33 0.50 0.40 55.64 5.45 Text-IF 6.72 3.82 0.59 0.16 32.55 3.97 7.33 5.06 0.67 0.38 58.86 3.67 6.84 5.04 0.47 0.26 62.52 4.41 Text-DiFuse 7.23 4.15 0.59 0.13 33.26 5.17 7.67 5.36 0.66 0.33 61.19 4.38 7.02 4.68 0.48 0.31 63.28 4.68 DRMF 7.27 4.08 0.52 0.16 39.13 4.60 7.41 5.99 0.55 0.43 66.83 3.60 6.74 4.51 0.43 0.33 68.76 4.77 OmniFuse 7.06 3.36 0.54 0.20 32.52 6.28 7.26 3.72 0.67 0.42 49.02 5.50 6.88 4.78 0.47 0.30 49.61 5.95 ControlFusion 7.11 3.91 0.55 0.14 36.43 3.95 7.36 5.02 0.68 0.36 68.03 3.95 6.83 4.60 0.50 0.30 72.13 3.85 TGFusion 7.49 4.16 0.47 0.28 65.70 3.67 7.51 4.55 0.61 0.48 69.92 3.27 7.23 5.63 0.40 0.37 69.82 4.15 Table 1: Quantitative comparison on the MSRS, LLVIP, and M3FD datasets. Complete TNO results are reported in Appendix B. VI (Random noise, RN) VI (Over-exposure, OE) VI (Blur) VI (Low-light, L) Methods CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CDDFuse 0.27 45.27 61.34 7.38 0.15 41.24 44.29 7.52 0.14 34.71 33.03 7.34 0.11 42.33 42.16 6.80 EMMA 0.23 42.51 46.14 7.41 0.14 39.40 38.27 7.56 0.14 33.90 27.50 7.38 0.16 41.53 34.93 6.82 Text-IF 0.26 44.76 60.54 7.40 0.13 41.38 48.39 7.15 0.16 34.67 34.16 7.36 0.14 39.97 44.42 7.17 Text-DiFuse 0.22 42.01 50.99 7.51 0.12 37.97 39.75 7.45 0.15 35.01 35.14 7.51 0.16 42.36 41.57 7.22 DRMF 0.27 47.02 67.68 7.49 0.18 42.91 53.39 7.24 0.23 35.06 39.84 7.46 0.22 44.40 51.00 7.36 OmniFuse 0.12 38.09 38.11 7.27 0.25 37.70 41.08 7.34 0.24 32.58 32.44 7.25 0.24 39.75 37.05 7.03 ControlFusion 0.19 47.60 58.02 7.28 0.16 45.94 52.18 7.32 0.14 31.86 31.15 7.21 0.15 44.79 48.62 6.87 TGFusion 0.30 57.40 83.73 7.52 0.26 49.90 69.60 7.52 0.29 60.74 87.77 7.60 0.24 51.42 72.78 7.46 VI (Rain) VI (Rain and Haze, RH) IR (Random noise, RN) IR (Low-contrast, LC) Methods CLIP-IQA MUSIQ TReS SD CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CDDFuse 0.14 43.36 43.62 47.11 0.37 59.51 71.68 6.99 0.25 42.00 41.32 7.40 0.20 45.61 47.30 6.87 EMMA 0.15 41.90 36.25 50.50 0.46 58.31 58.98 7.07 0.21 39.00 35.00 7.40 0.18 44.84 41.28 7.20 Text-IF 0.17 43.86 44.65 46.41 0.34 60.91 71.99 6.81 0.20 31.44 45.21 7.35 0.15 42.39 47.81 7.15 Text-DiFuse 0.13 39.96 39.82 56.63 0.36 60.36 64.49 7.25 0.16 48.64 56.69 7.51 0.13 44.08 47.07 6.94 DRMF 0.20 44.60 48.86 52.03 0.43 65.82 76.29 7.15 0.20 45.07 53.71 7.40 0.20 44.34 52.76 7.46 OmniFuse 0.25 38.30 38.00 41.36 0.36 48.26 44.79 6.76 0.23 39.09 38.74 7.26 0.24 40.06 39.53 7.26 ControlFusion 0.16 44.83 47.88 41.94 0.36 63.20 83.11 6.84 0.16 44.73 49.08 7.30 0.16 46.92 59.33 7.18 TGFusion 0.23 44.70 56.75 58.06 0.41 64.75 82.05 7.35 0.25 52.07 76.43 7.56 0.25 51.47 76.21 7.53 IR (Stripe noise, SN) VI (L) and IR (SN) VI (RH) and IR (RN) VI (OE) and IR (LC) Methods CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CLIP-IQA MUSIQ TReS EN CDDFuse 0.20 44.36 45.78 7.34 0.28 53.55 62.97 6.93 0.52 53.64 65.89 7.05 0.35 59.75 76.27 7.51 EMMA 0.18 43.01 39.24 7.39 0.36 52.76 54.76 6.96 0.45 53.70 65.80 7.04 0.30 56.67 63.67 7.38 Text-IF 0.19 44.13 46.43 7.37 0.25 58.11 72.50 7.07 0.46 53.59 76.46 7.00 0.29 58.61 77.71 6.95 Text-DiFuse 0.15 42.15 43.69 7.50 0.26 58.69 63.23 7.44 0.45 60.85 76.41 7.29 0.30 56.37 72.60 7.21 DRMF 0.20 45.08 54.09 7.42 0.39 62.03 75.64 7.39 0.43 65.02 74.40 7.11 0.35 61.50 79.07 7.09 OmniFuse 0.24 39.47 39.32 7.26 0.32 49.76 49.32 7.05 0.35 44.61 43.94 6.76 0.32 50.91 56.57 7.03 ControlFusion 0.17 44.50 50.31 7.31 0.33 59.22 71.55 7.06 0.36 64.08 82.57 6.88 0.33 64.63 83.25 7.27 TGFusion 0.25 53.47 80.60 7.60 0.41 63.96 82.51 7.48 0.41 64.71 82.04 7.32 0.41 65.05 81.91 7.39 Table 2: Quantitative comparisons on the DDL-12 benchmark under complex degradation scenarios. Optimization and Inference TGFusion combines the latent-space flow matching objective in Eq. 2 with pixel-space fusion regularization. Given the current latent state ZtZ_t, the predicted velocity u^t=vθ(Zt,t,Zv,Zr,c) u_t=v_θ(Z_t,t,Z_v,Z_r,c) is extrapolated over the remaining interval 1−t1-t to estimate the final fusion latent at t=1t=1, which is then decoded into the fused image: Z^1=Zt+(1−t)u^t,I^f=(Z^1), Z_1=Z_t+(1-t) u_t, I_f=D( Z_1), (10) where Z^1 Z_1 denotes the estimated fusion latent at the endpoint of the flow trajectory, and (⋅)D(·) is the frozen VAE decoder. Following widely used fusion constraints (Tang et al. 2022; Yi et al. 2024), intensity and maximum-gradient terms are imposed on the luminance channel: ℒint=1HW‖(I^f)−max((Ivc),(Irc))‖1,L_int= 1HW \|Y( I_f)- \! (Y(I_v^c),Y(I_r^c) ) \|_1, (11) ℒgrad=1HW‖|∇(I^f)|−max(|∇(Ivc)|,|∇(Irc)|)‖1,L_grad= 1HW \|| ( I_f)|- \! (| (I_v^c)|,| (I_r^c)| ) \|_1, (12) where IvcI_v^c and IrcI_r^c are clean source images, (⋅)Y(·) extracts luminance, and ∇ is the Sobel operator. The objective is ℒpixel=ℒint+λgℒgrad,L_pixel=L_int+ _gL_grad, (13) ℒtotal=ℒflow+λpixℒpixel,L_total=L_flow+ _pixL_pixel, (14) where λg _g and λpix _pix are fixed weighting coefficients. At inference, starting from Z0∼(0,)Z_0 (0,I), the conditional ODE is solved with Euler updates Zi+1=Zi+(ti+1−ti)vθ(Zi,ti,Zv,Zr,c),Z_i+1=Z_i+(t_i+1-t_i)v_θ(Z_i,t_i,Z_v,Z_r,c), (15) and the final result is decoded as I^f=(ZN) I_f=D(Z_N). Figure 4: Qualitative comparison of fused results under complex degradation scenarios of the DDL-12 degradation benchmark. 4 Experiments Experimental Setup Training details. TGFusion adopts a frozen Stable Diffusion 3 VAE (Esser et al. 2024) and a frozen CLIP ViT-L/14 text encoder (Radford et al. 2021). Its backbone consists of 12 Prompt-conditioned Multi-stream Joint Flow Transformer blocks, with a hidden dimension of 768 and 12 attention heads. Training is conducted on DDL-12 (Tang et al. 2025) using AdamW (Loshchilov and Hutter 2019) with a peak learning rate of 1×10−41× 10^-4. The loss weights are set to λpix=4 _pix=4 and λg=3 _g=3. Further implementation and optimization details are provided in Appendix A. Evaluation datasets. The evaluation covers four public IVIF benchmarks: MSRS (Tang et al. 2022), LLVIP (Jia et al. 2021), M3FD (Liu et al. 2022), and TNO (Toet 2017). Robustness to input degradations is assessed on the DDL-12 benchmark (Tang et al. 2025). Comparison methods. Comparison methods include CDDFuse (Zhao et al. 2023a), EMMA (Zhao et al. 2024), Text-IF (Yi et al. 2024), Text-DiFuse (Zhang et al. 2024a), DRMF (Tang et al. 2024), OmniFuse (Zhang et al. 2025), and ControlFusion (Tang et al. 2025). For methods requiring pre-enhancement, we use SwinIR (Liang et al. 2021), AirNet (Li et al. 2022), and WDNN (Guan et al. 2019) for infrared noise, low contrast, and stripe noise, respectively; LMPEC (Afifi et al. 2021) for visible overexposure; and AdaIR (Cui et al. 2025) for the remaining visible degradations. Metrics. Since source-referenced metrics can be confounded by degradations in the input images (Zhang et al. 2024b), we combine non-reference measures of information and detail (EN, SD, and AG), perceptual-quality metrics (CLIP-IQA, MUSIQ, TReS, and NIQE), and C as a complementary measure of source fidelity. Fusion Performance on Public Benchmarks Table 1 presents quantitative comparisons on public benchmarks. TGFusion achieves the best or second-best results in nearly all non-reference evaluations, demonstrating strong information and detail preservation, perceptual quality, and naturalness. Despite these advantages, TGFusion obtains only moderate C scores, whereas CDDFuse and EMMA, which do not explicitly model input degradations, achieve higher C without consistent gains in non-reference metrics. Since public benchmarks contain modality-specific degradations (Zhang et al. 2025), source correlation reflects content retention but cannot distinguish informative structures from degraded content. The qualitative comparisons in Appendix C further show that higher-C results also retain visible degradations and artifacts. We therefore primarily use no-reference metrics in subsequent comparisons. Metric Infrared Visible CDDFuse EMMA Text-IF Text-DiFuse DRMF OmniFuse ControlFusion TGFusion mIoU 65.53 69.98 70.85 70.63 70.72 71.75 70.04 69.12 71.10 71.83 mAcc 75.14 80.30 79.17 78.82 80.25 80.84 79.82 78.49 81.08 81.47 Table 3: Semantic segmentation evaluation on MFNet. Figure 5: Visualization of semantic–visual interactions in TGFusion. Top: Stage-wise attention from the fusion stream to the VIS, IR, and Text streams. Bottom: Localization of the haze token and qualitative comparison with ControlFusion. Fusion Comparison under Degradation Scenarios We further evaluate robustness under representative single-modality and compound degradations. As summarized in Table 2, TGFusion consistently ranks first or second, with obvious advantages in CLIP-IQA, MUSIQ, and TReS while maintaining competitive EN and SD. Compared with standard benchmarks, its gains are more pronounced under corrupted observations, indicating robustness to changes in modality reliability rather than to a specific type of degradation. Figure 4 corroborates this trend. Despite degradation-specific pre-enhancement, competing methods still inherit restoration residuals. In particular, ControlFusion (Tang et al. 2025) retains veil-like haze and underexposure in the haze and low-light cases; under rain-haze and rain, visible rain streaks remain, and scene textures are degraded. TGFusion effectively suppresses these residual corruptions, including simultaneous low-light and stripe-noise interference, while recovering visible structures and preserving infrared-target saliency. The qualitative evidence agrees with the perceptual-metric gains, demonstrating more reliable aggregation of complementary information under complex degradations. Please refer to Appendix D for more visual results. Analysis of Semantic–Visual Interaction We visualize the semantic–visual interactions within TGFusion. The upper panel of Figure 5 presents stage-wise attention from the fusion stream to the VIS, IR, and Text streams. Early blocks show broad responses across the three conditions, supporting initial inter-modal interaction. As features evolve, VIS attention increasingly traces scene structures, while IR attention concentrates on thermally salient targets. In later blocks, the Text stream also develops content-aware spatial responses. This broad-to-selective evolution suggests that joint interaction aligns text guidance with the emerging fusion objective, allowing text to support local refinement rather than remain a fixed global prior. The lower panel of Figure 5 shows the response of the individual haze token. Its high-response regions generally follow the spatial distribution of haze shown in the corresponding haze-level maps, particularly in heavily degraded areas. Compared with ControlFusion, which retains noticeable haze residuals in these regions, TGFusion more effectively suppresses haze in the corresponding high-response areas while preserving scene structures and infrared target saliency. The consistency among the haze distribution, token response, and fusion outcome indicates that the text stream can localize corrupted content and provide region-sensitive guidance for degradation suppression and reliable information aggregation across modalities. Semantic Segmentation Evaluation We further evaluate downstream semantic utility on MFNet (Ha et al. 2017) by separately training SegNeXt (Guo et al. 2022) on infrared, visible, and each method’s fused images under the same protocol. As reported in Table 3, TGFusion achieves the highest mIoU and mAcc, reaching 71.83% and 81.47%, respectively. Qualitative comparisons in Appendix E further show that effective information restoration and integration produce predictions more consistent with the ground truth, notably suppressing the false responses observed with competing methods. Detailed per-class IoU and accuracy results are also provided in the appendix E. Ablation Study Figure 6: Qualitative comparison of network-component ablations on the standard MSRS dataset. Config. EN AG CLIP-IQA TReS NIQE I 7.39 3.84 0.21 46.22 4.46 I 7.15 3.71 0.17 36.73 5.01 I 7.18 3.83 0.25 64.41 5.15 IV 7.42 4.37 0.20 46.52 4.27 V 7.49 4.16 0.28 65.70 3.67 Table 4: Quantitative network-component ablation results on the standard MSRS dataset. As shown in Figure 6 and Table 4, replacing joint attention with independent stream-wise self-attention (I) reduces the relative saliency of infrared targets against bright background structures, with consistent degradation across all metrics. Removing the independent text stream (I) further weakens the emphasis on local thermal targets and produces the lowest EN, AG, CLIP-IQA, and TReS, demonstrating that global textual modulation alone cannot provide sufficient semantic–visual interaction. Performing flow matching directly in pixel space (I) causes severe brightness amplification, washed-out structures, and color artifacts. Although it achieves the second-best CLIP-IQA and TReS, its lower EN and AG together with the worst NIQE indicate inferior information preservation and visual naturalness. Removing the pixel-space fusion loss (IV) introduces brightness drift and over-accentuates local structures, such as pedestrian contours and road cracks. Its highest AG therefore reflects stronger local gradients rather than balanced perceptual quality, as evidenced by the substantially lower CLIP-IQA and TReS. In contrast, the full model (V) preserves natural colors and visible structures while maintaining clear thermal-target saliency, achieving the best EN, CLIP-IQA, TReS, and NIQE. These results validate the complementary contributions of joint multi-stream interaction, the independent text stream, latent-space flow matching, and pixel-space fusion supervision. 5 Conclusion This work demonstrates that textual priors can move beyond static global conditioning to actively support degradation-robust IVIF. By coupling progressive semantic–visual interaction with latent flow generation, TGFusion mitigates diverse degradations while retaining fine structural details and salient thermal targets. Results on standard benchmarks, single and compound degradations, and semantic segmentation confirm its robust fusion quality and downstream utility. These findings suggest a promising direction for reliable multimodal fusion under adverse imaging conditions. References M. Afifi, K. G. Derpanis, B. Ommer, and M. S. Brown (2021) Learning multi-scale photo exposure correction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 9157–9167. Cited by: §4. F. Bao, X. Wang, S. H. Sureshbabu, G. Sreekumar, L. Yang, V. Aggarwal, V. N. Boddeti, and Z. Jacob (2023) Heat-assisted detection and ranging. Nature 619 (7971), p. 743–748. Cited by: §1. B. Cao, X. Xu, P. Zhu, Q. Wang, and Q. Hu (2024) Conditional controllable image fusion. Advances in Neural Information Processing Systems 37, p. 120311–120335. Cited by: §2. Z. Cao, Y. Zhong, Z. Wang, and L. Deng (2025) MMAIF: multi-task and multi-degradation all-in-one for image fusion with language guidance. In Proceedings of the IEEE/CVF International Conference on Computer Vision, p. 11744–11754. Cited by: §1, §2. C. Cheng, T. Xu, X. Wu, H. Li, X. Li, Z. Tang, and J. Kittler (2025) TextFusion: unveiling the power of textual semantics for controllable image fusion. Information Fusion 117, p. 102790. Cited by: §2. Y. Cui, S. W. Zamir, S. Khan, A. Knoll, M. Shah, and F. Khan (2025) AdaIR: adaptive all-in-one image restoration via frequency mining and modulation. In The Thirteenth International Conference on Learning Representations, Cited by: §4. P. Esser, S. Kulal, A. Blattmann, R. Entezari, J. Müller, H. Saini, Y. Levi, D. Lorenz, A. Sauer, F. Boesel, et al. (2024) Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first international conference on machine learning, Cited by: §4. J. Guan, R. Lai, and A. Xiong (2019) Wavelet deep neural network for stripe noise removal. IEEE Access 7, p. 44544–44554. Cited by: §4. M. Guo, C. Lu, Q. Hou, Z. Liu, M. Cheng, and S. Hu (2022) Segnext: rethinking convolutional attention design for semantic segmentation. Advances in neural information processing systems 35, p. 1140–1156. Cited by: §4. Q. Ha, K. Watanabe, T. Karasawa, Y. Ushiku, and T. Harada (2017) MFNet: towards real-time semantic segmentation for autonomous vehicles with multi-spectral scenes. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), p. 5108–5115. Cited by: Figure 9, §4. D. K. Jain, X. Zhao, G. González-Almagro, C. Gan, and K. Kotecha (2023) Multimodal pedestrian detection using metaheuristics with deep convolutional neural network in crowded scenes. Information Fusion 95, p. 401–414. Cited by: §1. X. Jia, C. Zhu, M. Li, W. Tang, and W. Zhou (2021) LLVIP: a visible-infrared paired dataset for low-light vision. In Proceedings of the IEEE/CVF international conference on computer vision, p. 3496–3504. Cited by: §4. B. Li, X. Liu, P. Hu, Z. Wu, J. Lv, and X. Peng (2022) All-In-One Image Restoration for Unknown Corruption. In IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA. Cited by: §4. H. Li and X. Wu (2018) DenseFuse: a fusion approach to infrared and visible images. IEEE transactions on image processing 28 (5), p. 2614–2623. Cited by: §1. H. Li, T. Xu, X. Wu, J. Lu, and J. Kittler (2023) Lrrnet: a novel representation learning guided fusion network for infrared and visible images. IEEE transactions on pattern analysis and machine intelligence 45 (9), p. 11040–11052. Cited by: §1. J. Liang, J. Cao, G. Sun, K. Zhang, L. Van Gool, and R. Timofte (2021) Swinir: image restoration using swin transformer. In Proceedings of the IEEE/CVF international conference on computer vision, p. 1833–1844. Cited by: §4. P. Liang, J. Jiang, X. Liu, and J. Ma (2022) Fusion from decomposition: a self-supervised decomposition approach for image fusion. In European conference on computer vision, p. 719–735. Cited by: §1. Y. Lipman, R. T. Chen, H. Ben-Hamu, M. Nickel, and M. Le (2023) Flow matching for generative modeling. In International Conference on Learning Representations, Cited by: §2, §3. J. Liu, X. Fan, Z. Huang, G. Wu, R. Liu, W. Zhong, and Z. Luo (2022) Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5802–5811. Cited by: §1, §4. J. Liu, G. Wu, Z. Liu, D. Wang, Z. Jiang, L. Ma, W. Zhong, X. Fan, and R. Liu (2025) Infrared and visible image fusion: from data compatibility to task adaption. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 (4), p. 2349–2369. Cited by: §1. I. Loshchilov and F. Hutter (2019) Decoupled weight decay regularization. In International Conference on Learning Representations, External Links: Link Cited by: §4. J. Ma, L. Tang, F. Fan, J. Huang, X. Mei, and Y. Ma (2022) SwinFusion: cross-domain long-range learning for general image fusion via swin transformer. IEEE/CAA Journal of Automatica Sinica 9 (7), p. 1200–1217. Cited by: §1. J. Ma, W. Yu, P. Liang, C. Li, and J. Jiang (2019) FusionGAN: a generative adversarial network for infrared and visible image fusion. Information fusion 48, p. 11–26. Cited by: §1. OpenAI (2023) ChatGPT: large language model. Note: https://openai.comAccessed: 2025-02-20 Cited by: §3. W. Peebles and S. Xie (2023) Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, p. 4195–4205. Cited by: §3. A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In International conference on machine learning, p. 8748–8763. Cited by: §2, §3, §4. J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu (2024) Roformer: enhanced transformer with rotary position embedding. Neurocomputing 568, p. 127063. Cited by: §3. L. Tang, Y. Deng, X. Yi, Q. Yan, Y. Yuan, and J. Ma (2024) DRMF: degradation-robust multi-modal image fusion via composable diffusion prior. In Proceedings of the 32nd ACM International Conference on Multimedia, p. 8546–8555. Cited by: §2, §4. L. Tang, Y. Wang, Z. Cai, J. Jiang, and J. Ma (2025) ControlFusion: a controllable image fusion network with language-vision degradation prompts. Advances in Neural Information Processing Systems 38, p. 118304–118327. Cited by: Appendix A, Figure 8, §1, §1, §2, §4, §4, §4, §4. L. Tang, J. Yuan, H. Zhang, X. Jiang, and J. Ma (2022) PIAFusion: a progressive infrared and visible image fusion network based on illumination aware. Information Fusion 83, p. 79–92. Cited by: §3, §4. A. Toet (2017) The tno multiband image data collection. Data in brief 15, p. 249. Cited by: §4. X. Yi, H. Xu, H. Zhang, L. Tang, and J. Ma (2024) Text-IF: leveraging semantic text guidance for degradation-aware and interactive image fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 27026–27035. Cited by: §1, §2, §3, §3, §4. H. Zhang, L. Cao, and J. Ma (2024a) Text-DiFuse: an interactive multi-modal image fusion framework based on text-modulated diffusion model. Advances in Neural Information Processing Systems 37, p. 39552–39572. Cited by: §2, §4. H. Zhang, L. Cao, X. Zuo, Z. Shao, and J. Ma (2025) OmniFuse: composite degradation-robust image fusion with language-driven semantics. IEEE Transactions on Pattern Analysis and Machine Intelligence 47 (9), p. 7577–7595. Cited by: §1, §2, §2, §4, §4. H. Zhang, X. Zuo, J. Jiang, C. Guo, and J. Ma (2024b) MRFS: mutually reinforcing image fusion and segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, p. 26974–26983. Cited by: §4. J. Zhang, H. Liu, K. Yang, X. Hu, R. Liu, and R. Stiefelhagen (2023) CMX: cross-modal fusion for rgb-x semantic segmentation with transformers. IEEE Transactions on intelligent transportation systems 24 (12), p. 14679–14694. Cited by: §1. J. Zhang, A. Liu, D. Wang, Y. Liu, Z. J. Wang, and X. Chen (2022) Transformer-based end-to-end anatomical and functional image fusion. IEEE Transactions on Instrumentation and Measurement 71, p. 1–11. Cited by: §1. Y. Zhang, B. Song, X. Du, and M. Guizani (2018) Vehicle tracking using surveillance with multimodal data fusion. IEEE Transactions on Intelligent Transportation Systems 19 (7), p. 2353–2361. Cited by: §1. Z. Zhao, H. Bai, J. Zhang, Y. Zhang, S. Xu, Z. Lin, R. Timofte, and L. Van Gool (2023a) Cddfuse: correlation-driven dual-branch feature decomposition for multi-modality image fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 5906–5916. Cited by: §1, §4. Z. Zhao, H. Bai, J. Zhang, Y. Zhang, K. Zhang, S. Xu, D. Chen, R. Timofte, and L. Van Gool (2024) Equivariant multi-modality image fusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, p. 25912–25921. Cited by: §1, §4. Z. Zhao, H. Bai, Y. Zhu, J. Zhang, S. Xu, Y. Zhang, K. Zhang, D. Meng, R. Timofte, and L. Van Gool (2023b) DDFM: denoising diffusion model for multi-modality image fusion. In Proceedings of the IEEE/CVF international conference on computer vision, p. 8082–8093. Cited by: §2. Figure 7: Qualitative comparisons on public benchmarks. Supplementary Contents This supplementary material provides further support for the methodology and experimental findings presented in the main paper. Specifically, it includes: • Detailed training and inference settings. • Quantitative results on TNO. • Qualitative comparisons on public benchmarks. • Extended qualitative comparisons across diverse degradation scenarios. • Detailed semantic segmentation results on MFNet, including qualitative comparisons and per-class IoU and accuracy evaluation results. Appendix A Implementation Details TGFusion is trained on DDL-12 (Tang et al. 2025) for 120 epochs using bf16 precision. Paired infrared and visible images undergo synchronized random horizontal and vertical flipping, followed by aligned random cropping to 256×256256× 256. The trainable parameters are optimized using AdamW with β1=0.9 _1=0.9, β2=0.999 _2=0.999, and a weight decay of 0.010.01. The effective batch size is 32 across two GPUs. The learning rate is linearly warmed up to 1×10−41× 10^-4 during the first 5 epochs and then cosine-decayed to 1×10−61× 10^-6. Training is implemented in PyTorch 2.4.1 with CUDA 12.1 and conducted on two A100-SXM4-40GB GPUs for approximately three days. During inference, the fused latent representation is generated using a 25-step Euler ODE solver with a classifier-free guidance scale of 3.03.0. Appendix B Quantitative Comparison on TNO Table 5 reports the complete quantitative results on TNO. TGFusion achieves the best EN, AG, and CLIP-IQA scores and ranks second in both TReS and NIQE, although it does not attain a leading C score. This overall performance trend is consistent with the findings on the public benchmarks reported in the main paper, demonstrating TGFusion’s competitive image fusion capability. Appendix C Visual Comparison on Public Benchmarks Figure 7 presents qualitative comparisons on four public benchmarks. On MSRS, TGFusion preserves the salient thermal response of the pedestrian while maintaining illumination transitions and structural details in the surrounding dark regions. On LLVIP, EMMA exhibits regular grid artifacts despite its higher C, whereas OmniFuse introduces conspicuous blockwise color distortions. TGFusion avoids these artifacts while retaining clear pedestrian, vehicle, and road structures. On M3FD, it alleviates the veil-like, low-contrast appearance observed in competing results and preserves both the salient vehicle response and clearer window and balcony structures. On TNO, TGFusion maintains a prominent small thermal target together with better-defined branches and ground textures. These comparisons further show that stronger source correlation can accompany the retention of degradations or artifacts. Overall, TGFusion achieves a more favorable balance among informative source-content preservation, reliable cue selection, and perceptual quality. Appendix D Extended Qualitative Comparisons across Diverse Degradation Scenarios Method EN AG C CLIP-IQA TReS NIQE CDDFuse 7.120 4.896 0.513 0.262 27.718 4.346 EMMA 7.158 4.760 0.493 0.287 28.026 5.206 Text-IF 7.159 4.917 0.484 0.237 31.370 3.923 Text-DiFuse 7.145 4.115 0.472 0.247 36.851 4.481 DRMF 7.195 4.462 0.384 0.242 31.826 4.886 OmniFuse 7.043 3.564 0.468 0.221 36.401 6.189 ControlFusion 6.977 4.479 0.488 0.274 53.162 4.524 TGFusion 7.315 5.798 0.405 0.297 36.967 4.102 Table 5: Quantitative comparison on the TNO dataset. Figure 8: Additional qualitative comparisons under representative degradation scenarios from DDL-12 (Tang et al. 2025). Figure 9: Qualitative semantic-segmentation comparison on MFNet (Ha et al. 2017). Figure 8 provides additional qualitative comparisons across representative modality-specific and cross-modal compound degradations. Competing methods often retain corrupted source cues, resulting in residual noise, rain streaks, brightness distortions, or excessive smoothing that weakens thermal saliency and structural details. In contrast, TGFusion more effectively limits degradation propagation while preserving salient targets and recovering scene content. These results are consistent with the proposed design: structured degradation priors progressively interact with the visual streams through joint attention, facilitating the selection of reliable modality cues, while latent-space flow matching supports high-quality structure and texture recovery. Overall, TGFusion achieves a favorable balance among degradation mitigation, thermal-target preservation, and detail reconstruction across diverse and complex degradation scenarios. Class Infrared Visible CDDFuse EMMA Text-IF Text-DiFuse DRMF OmniFuse ControlFusion TGFusion Intersection over Union (IoU, %) car 87.60 89.56 89.41 89.16 90.16 89.96 89.56 89.86 89.99 89.94 person 70.54 64.00 72.79 73.57 75.11 74.37 73.34 72.21 74.53 72.56 bike 66.05 69.04 69.33 70.12 70.24 70.24 70.18 68.38 71.57 70.58 curve 52.41 56.50 58.91 59.98 59.34 59.56 58.08 55.78 59.61 58.54 car stop 66.87 73.74 73.09 73.56 74.22 73.52 71.55 69.53 72.85 76.95 guardrail 56.64 74.47 71.48 66.61 67.82 70.74 67.97 64.43 65.70 72.99 color cone 60.02 66.63 63.21 64.73 66.52 69.12 61.62 66.34 68.58 67.77 bump 64.14 65.90 68.60 67.32 62.33 66.45 68.04 66.42 66.00 65.28 mIoU↑ 65.53 69.98 70.85 70.63 70.72 71.75 70.04 69.12 71.10 71.83 Per-class Accuracy (%) car 92.29 93.83 92.99 92.92 94.35 94.09 93.98 93.69 93.85 94.32 person 81.61 74.00 83.48 83.62 86.45 84.67 83.06 82.33 85.28 82.93 bike 76.63 77.64 77.81 78.37 78.22 79.49 77.93 76.24 80.43 78.55 curve 64.61 71.35 71.47 75.20 75.12 74.25 72.01 67.84 74.37 73.07 car stop 77.42 81.50 81.57 79.79 80.35 80.75 77.99 75.87 80.78 85.16 guardrail 61.75 84.58 77.07 70.05 74.27 78.53 79.37 75.05 79.23 79.24 color cone 70.12 76.66 70.03 73.75 80.77 77.67 73.21 76.21 77.78 79.31 bump 76.66 82.82 78.97 76.85 72.50 77.24 80.97 80.65 76.89 79.15 mAcc↑ 75.14 80.30 79.17 78.82 80.25 80.84 79.82 78.49 81.08 81.47 Table 6: Per-class IoU and accuracy comparison for semantic segmentation on MFNet. Appendix E Detailed Semantic Segmentation Results Although the best per-class results are distributed among different methods, TGFusion achieves the highest mIoU and mAcc of 71.83 and 81.47, respectively. Its simultaneous advantages in region overlap and category-wise accuracy demonstrate balanced semantic utility rather than gains confined to isolated categories. As shown in Figure 9, most competing methods produce a conspicuous bike response adjacent to the pedestrian despite its absence from the ground truth. TGFusion substantially suppresses this false response while preserving the principal car, person, and road-structure regions. Together, the quantitative and qualitative results indicate that TGFusion retains discriminative semantic information while reducing misleading cues, thereby supporting more reliable downstream segmentation.