Paper deep dive
User-Aware Conditional Generative Total Correlation Learning for Multi-Modal Recommendation
Jing Du, Zesheng Ye, Congbo Ma, Feng Liu, Flora. D. Salim
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 4/10/2026, 2:04:35 AM
Summary
The paper introduces GTC, a conditional Generative Total Correlation learning framework for multi-modal recommendation (MMR). GTC addresses the limitations of existing disentanglement methods by employing an interaction-guided diffusion model for user-aware content feature filtering and optimizing a tractable lower bound of total correlation to capture higher-order cross-modal dependencies, achieving state-of-the-art performance.
Entities (5)
Relation Signals (3)
GTC → improves → Multi-modal Recommendation
confidence 98% · Experiments on standard MMR benchmarks show GTC consistently outperforms state-of-the-art
GTC → optimizes → Total Correlation
confidence 95% · it pursues holistic cross-modal alignment and maximizes the total correlation
GTC → utilizes → Diffusion Model
confidence 95% · We employ an interaction-guided diffusion model to perform user-aware content feature filtering
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-modal recommendation (MMR) enriches item representations by introducing item content, e.g., visual and textual descriptions, to improve upon interaction-only recommenders. The success of MMR hinges on aligning these content modalities with user preferences derived from interaction data, yet dominant practices based on disentangling modality-invariant preference-driving signals from modality-specific preference-irrelevant noises are flawed. First, they assume a one-size-fits-all relevance of item content to user preferences for all users, which contradicts the user-conditional fact of preferences. Second, they optimize pairwise contrastive losses separately toward cross-modal alignment, systematically ignoring higher-order dependencies inherent when multiple content modalities jointly influence user choices. In this paper, we introduce GTC, a conditional Generative Total Correlation learning framework. We employ an interaction-guided diffusion model to perform user-aware content feature filtering, preserving only personalized features relevant to each individual user. Furthermore, to capture complete cross-modal dependencies, we optimize a tractable lower bound of the total correlation of item representations across all modalities. Experiments on standard MMR benchmarks show GTC consistently outperforms state-of-the-art, with gains of up to 28.30% in NDCG@5. Ablation studies validate both conditional preference-driven feature filtering and total correlation optimization, confirming the ability of GTC to model user-conditional relationships in MMR tasks. The code is available at: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2604.03014v1
- Canonical: https://arxiv.org/abs/2604.03014v1
Trouble viewing inline? Open PDF directly →
Full Text
77,400 characters extracted from source content.
Expand or collapse full text
User-Aware Conditional Generative Total Correlation Learning for Multi-Modal Recommendation Jing Du The University of New South Wales Sydney, Australia jing.du2@unsw.edu.au Zesheng Ye University of Melbourne Melbourne, Australia zesheng.ye@unimelb.edu.au Congbo Ma New York University Abu Dhabi Abu Dhabi, UAE cm7196@nyu.edu Feng Liu University of Melbourne Melbourne, Australia fengliu.ml@gmail.com Flora Salim The University of New South Wales Sydney, Australia flora.salim@unsw.edu.au Abstract Multi-modal recommendation (MMR) enriches item representations by introducing item content, e.g., visual and textual descriptions, to improve upon interaction-only recommenders. The success of MMR hinges on aligning these content modalities with user preferences derived from interaction data, yet dominant practices based on disentangling modality-invariant preference-driving signals from modality-specific preference-irrelevant noises are flawed. First, they assume a one-size-fits-all relevance of item content to user preferences for all users, which contradicts the user-conditional fact of preferences. Second, they optimize pairwise contrastive losses separately toward cross-modal alignment, systematically ignoring higher-order dependencies inherent when multiple con- tent modalities jointly influence user choices. In this paper, we introduce GTC, a conditionalGenerativeTotalCorrelation learn- ing framework. We employ an interaction-guided diffusion model to perform user-aware content feature filtering, preserving only personalized features relevant to each individual user. Further- more, to capture complete cross-modal dependencies, we optimize a tractable lower bound of the total correlation of item representa- tions across all modalities. Experiments on standard MMR bench- marks show GTC consistently outperforms state-of-the-art, with gains of up to 28.30% in NDCG@5. Ablation studies validate both conditional preference-driven feature filtering and total correla- tion optimization, confirming the ability of GTC to model user- conditional relationships in MMR tasks. The code is available at: https://github.com/jingdu-cs/GTC. CCS Concepts • Information systems→Recommender systems;• Comput- ing methodologies→ Neural networks. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference’17, Washington, DC, USA © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-x-x-x-x/Y/M https://doi.org/10.1145/n.n Keywords Multi-modal Recommendation, Diffusion Models, Total Correlation, Cross-modal Alignment ACM Reference Format: Jing Du, Zesheng Ye, Congbo Ma, Feng Liu, and Flora Salim. 2026. User- Aware Conditional Generative Total Correlation Learning for Multi-Modal Recommendation. In . ACM, New York, NY, USA, 11 pages. https://doi.org/ 10.1145/n.n 1 Introduction Recommender systems have evolved beyond collaborative filtering to embrace multi-modal recommendation (MMR), where diverse and rich content information, such as product images and textual descriptions, is used to enrich item representations by providing an explicit view of item characteristics that complements user prefer- ences derived from sparse user-item interactions [12]. When these information-rich content modalities are effectively aligned with in- teraction data modality, MMR promises more comprehensive user preferences grounded in semantically-meaningful item attributes, thus improving recommendation performance [23, 27]. Yet, a tension exists inherently between these modalities. Fig. 1 reveals a striking empirical observation: representations learned from user interactions diverge substantially from those extracted from visual and textual content, implying that the rich information in content modalities does not always align with the behavioral sig- nals of user interactions holistically [10]. This divergence exposes a fundamental challenge in MMR. Only partial features drive a user’s decision to engaged with an item; while the rest, from this user’s perspective, are noise that do not help to reveal the true preference, reflected by user-item interactions. As such, effective MMR models must be able to distinguish “relevant” preference-driving signals from “irrelevant” noises, filtering item features based on their rele- vance to user preference to avoid performance degradation [34]. The predominant response is multi-modal Disentangled Repre- sentation Learning (DRL), which attempts to isolate a “modality- invariant” preference-driving component of the item representation from “modality-specific” ones [13,28], and optimize for pairwise contrastive losses to align modalities thereafter, assuming that core preference, e.g., a T-shirt’s appeals might be “comfortable”, remain consistent across modalities and users; whilst noisy attributes, e.g., the use of words like “must-have”, may vary instead [4, 36]. arXiv:2604.03014v1 [cs.IR] 3 Apr 2026 Conference’17, July 2017, Washington, DC, USAJing Du et al. interact-imginteract-txtimg-txt 0.000 0.005 0.010 0.015 0.020 0.025 0.030 Cosine Distance Cell Baby Sports interact-imginteract-txtimg-txt 0 2 4 6 8 10 Euclidean Distance Cell Baby Sports Figure 1: The representations from user-item interactions ex- hibit low cosine similarity (left) and high Euclidean distance (right) to representations from visual and textual content modalities, on Amazon Sports dataset. In contrast, visual and textual representations are more aligned, implying the gap between latent user preference and explicit item attributes. T-Shirt by AW Apparel Crafted from 100% premium cotton, it offers comfort. Its bright orange hue and classic cut blend contemporary style with timeless simplicity. T-Shirt by AW Apparel Crafted from 100% premium cotton, it offers comfort. Its bright orange hue and classic cut blend contemporary style with timeless simplicity. User AUser B brand style color fabric Figure 2: The user-conditional nature of “appealing” fea- ture relevance. The item on the right appeals to a user who likes “orange” and “cotton”, while a different user is more interested in the brand (“AW Apparel”) and style, for whom color and fabric are less important. Thus, what constitutes a preference-driving signal is not only determined by the item attributes themselves, but also by individual user preference. However, the prevailing DRL practices are both impractical and suboptimal (detailed in Sec. 2.3). One issue is that they enforce a one-size-fits-all separation of item representations over all users. This implicitly assumes the relevance of a feature to user prefer- ences is fixed and universal, when in fact it is inherently conditional on each user. A user may prioritize a T-shirt’s “orange” color and “cotton” fabric, while another cares only about the brand (Fig. 2). By doing so, DRL methods may discard features relevant to some users but not others, failing to capture the entangled nature of item attributes and user-specific behaviors [2]. Second, existing MMR practices optimize pairwise contrastive losses independently—aligning interaction-textual, interaction-visual, and visual-textual modalities in parallel [21,36]. This is incomplete when more than two modalities are involved, as it cannot model interdependencies that emerge when modalities interact collectively. Consider a user who might engage with a product only when both its visual appeal and textual description align with the interest, nei- ther modality alone would drive the interaction. Pairwise alignment treats visual and textual signals as independent factors, failing to model this joint dependency, resulting in suboptimal alignment. Such limitations, i.e., the flawed assumption of universal content feature relevance to user preferences and the incomplete modeling of cross-modal dependencies, call for an alternative MMR paradigm. In Sec. 3, we propose a new framework called GTC that learns toGenerate interaction-guided item representations for content modalities and maximize cross-modalTotalCorrelation. GTC is built upon two principles: (1) since user-item interactions directly reflects user preference towards items, it leverages a diffusion model, guided by individual user interaction histories, and performs user- conditional preference filtering to denoise visual and textual features and preserve only those relevant to that specific user interactions; (2) it pursues holistic cross-modal alignment and maximizes the total correlation [25], which captures the complete dependency structure, including higher-order interactions that pairwise methods miss. Concretely, following standard practices [27,34], GTC first en- codes user embedding and three modality-specific item embeddings from interaction, textual and visual contents, based on user-item bi- partite graphs. It then refines the visual and textual representations using an interaction-guided diffusion model, leveraging its strength in representation learning [29]. Next, it aligns the interaction with two user-aware content representations by maximizing a tractable lower bound on their total correlation. Lastly, it integrates the inter- action item embedding with aligned content features to form the final representation for recommendation. Our contributions are: • User-aware generative filtering that tailors content features to indivudal user preferences, eliminating impractical uni- versal feature relevance assumptions (Sec. 3.2). • Total correlation maximization, the first attempt to model higher-order cross-modal dependencies for MMR, mitigating issues of separate pairwise alignments (Sec. 3.3). •A new framework that integrates these two principles to achieve a new state-of-the-art in MMR. •Extensive validation across standard MMR benchmarks that confirm consistent effectiveness of GTC (Sec. 4). 2 Background In this section, we first formulate the problem of MMR and contex- tualize previous work within this formulation. We then pinpoint their limitations that motivate our approach. 2.1 Problem Formulation Recommendation. LetU= 푢 푚 |U| 푚=0 be the user set andI= 푖 푛 |I| 푛=0 be the item set. Given an interaction matrix R∈ 0,1 |U|×|I| , where푅 푚푛 =1 denotes an observed interaction exists between the푚-th user and the푛-th item. These interactions are naturally represented as a user-item bipartite graphG= ⟨ U, I,E ⟩ , where an edge(푢 푚 ,푖 푛 ) ∈ Eexists if and only if푅 푚푛 =1 [4,5]. The goal of recommendation is to learn a model parameterized byΘthat maps a user푢 푚 and an item푖 푛 to a shared latent representation space e 푚 = ℎ 푢 (R;Θ) ∈ R 푑 and e 푛 = 푔 푖 (R;Θ) ∈ R 푑 , by optimizingΘto minimize a ranking-based loss function, e.g., Bayesian Personalized Ranking (BPR) loss [18], defined over a scoring function, typically the dot product푠(푢 푚 ,푖 푛 )=e ⊤ 푚 e 푛 . Such that the scoring function 푠(푢 푚 ′ ,·) can rank unseen items for each user푢 푚 ′ . Multi-modal Recommendation. MMR extends this setup by in- corporating heterogeneous item content modalities. We consider each푖 푛 has accompanying item image and textual description, pre- processed via pre-trained encoders, e.g., a convolutional neural network [7] for visual data and a Transformer [17] for textual data, leading to visual featuresV= v 푖 ∈ R 푑 푣 푖∈I and textual features T=t 푖 ∈ R 푑 푡 푖∈I over all items. Notice that the item representation map becomes e 푛 = 푔 푖 (v 푖 ,t 푖 ,R;Θ)by integrating both visual and User-Aware Conditional Generative Total Correlation Learning for Multi-Modal RecommendationConference’17, July 2017, Washington, DC, USA textual features, in addition to the interaction data. The crucial design challenge in MMR, which has driven the evolution of the field, is how to design this function to effectively combine signals from the interaction modality with those from content modalities. This paper categorizes them into two paradigms. 2.2 Previous Paradigms in MMR Direct Fusion. Early methods focused on learning representations for each modality and fusing them. This began with direct concate- nation [27] and evolved to using attention that assign weights to each modality, recognizing their unequal contributions to user pref- erence [10,23]. A primary challenge with this paradigm is that content modalities are often noisy (i.e., features irrelevant to user preference) that can degrade performance when fused indiscrimi- nately. While later methods leveraged denoising steps like user-item interaction graph pruning [26] and spectral filtering [15, 35], they operate on a user-agnostic basis, either fusing all available informa- tion or filtering noise based on universal criteria, failing to model the fact that the feature relevance is often specific to individuals. Disentanglement. The challenge of cleanly removing preference- irrelevant noise from content features spurred a paradigm shift towards disentangled representation learning (DRL) [1,14,24,33]. DRL methods aim to isolate a modality-invariant component (the assumed core preference-driving signal), from modality-specific components (the assumed noise) [3,5]. This is often enforced via pairwise contrastive optimization, which aligns item representa- tions from different modalities in a shared space [21,36]. Later studies focused on prompting statistical independence between signals and noise to facilitate disentanglement [6,11], sticking on fixed interaction-content relevance. 2.3 Motivation of This Study Still, DRL inherits the user-agnostic flaw of direct fusion methods and even introduces new limitations. Limitation of User-Agnostic Disentanglement. The core premise of DRL is that features can be universally categorized as “relevant” or “irrelevant” to user preference [3]. However, this is misaligned with how preferences are formed. Feature relevance is inherently user-conditional, as two users can have distinct preference pat- terns: a fashion-conscious person for whom visual aesthetics drive purchasing decisions, and a functionality-focused user who priori- tizes textual technical specifications. The enforcement of universal decomposition inevitably discards certain features that might be vital to specific users, creating a bottleneck affecting users whose preferences deviate from population-level patterns. This highlights the need for user-aware preference filtering that can adaptively determine feature relevance of content, based on individual interac- tion patterns. Such methods should preserve the benefits of noise reduction while maintaining sensitivity to user-specific preference. Limitation of Pairwise Alignment. Also, the alignment strategy used in previous DRL methods is theoretically incomplete. A holistic alignment requires capturing the complete statistical dependence among all modalities, a quantity measured by total correlation [25]. Let S,V,T be the random variables for the interaction, visual, and textual representations of an item. Following [20], total correlation TC(S,V,T)is defined as the KL-divergence between their joint distribution and the product of their marginals, decomposed as: 3· TC(S, V, T) = 3· 퐷 KL (푝(S, V, T) ∥ 푝(S)푝(V)푝(T)) = 2· [ 퐼(S; V)+ 퐼(S; T)+ 퐼(V; T) ] | z pairwise alignment, captured by prior studies + 퐼(S; V|T)+ 퐼(S; T|V)+ 퐼(V; T|S) | z higher-order dependencies, missed by prior studies , (1) where퐼(·;·)is the mutual information between two random vari- ables. Unequivocally, pairwise contrastive alignment that maxi- mizes a sum of pairwise mutual information between every two modalities, implemented in existing studies, is an incomplete proxy for total correlation. They neglect higher-order dependencies, such as the information shared between interaction and text once the visual context is known, leading to a suboptimal alignment that fails to model the whole picture. To overcome the limitations, we propose the principled GTC that departs from universal disentanglement assumptions to perform user-conditional feature filtering, and replaces incomplete pairwise alignment with an optimization towards total correlation 1 . 3 Methodology 3.1 Overview We now introduce the proposed framework for MMR, which learns toGenerate user-aware item representations and maximizes cross- modalTotalCorrelation (GTC) with two design choices: (1) gener- ating user-aware content features via user-item interaction guided diffusion (Sec. 3.2); and (2) maximizing (a tractable lower bound of ) the total correlation for holistic cross-modal alignment (Sec. 3.3). These aligned representations are then fused to perform the recom- mendation task (Sec. 3.4). Fig. 3 overviews the overall framework. 3.2 User-Aware Content Feature Generation Initial Embeddings. We begin by obtaining initial embeddings for all nodes in the bipartite graphG. For the interaction modality, we randomly initialize a node embedding for each user r 푢 and item r 푒 , capturing a latent collaborative space. For content modalities, item features are captured by pre-trained encoders, denoted as r 푣 and r 푡 for visual and textual node features. Since users do not have contents in our setup, their representations in the content modalities are given by the same initialized embeddings r 푢 . We then obtain modality-specific embeddings by propagating corresponding features through three parallel LightGCNs [8] over the shared graph structure. Let R 푚 be the input feature matrix for modality푚and E 푚 be the output embedding matrix, we have E 푚 = LightGCN( ̄ A, R 푚 ), for푚 ∈ E,V,T,(2) whereE,V,Tare the interaction, visual and textual spaces. ̄ Ais the normalized adjacency matrix of the whole bipartite graphG. From these outputs, we define three modality-specific embedding matrices: user and item S : = E E from interaction signalsE, visual and textual V : = E V , T : = E T from respective modalities. 1 The primary obstacle to this is the intractability of optimizing total correlation directly, we will address this in Sec. 3.3. Conference’17, July 2017, Washington, DC, USAJing Du et al. cross-modal alignment loss Eq. (9) <latexit sha1_base64="Aymnjg0c4KMCsEGiz6DIOeTCkNE=">AAAB/nicbVDLSgMxFM34rPU1Kq7cBIvgqsxIUZfFblyIVrAP6AxDJk3b0CQzJBmhDAP+ihsXirj1O9z5N2baWWjrgcDhnHu5JyeMGVXacb6tpeWV1bX10kZ5c2t7Z9fe22+rKJGYtHDEItkNkSKMCtLSVDPSjSVBPGSkE44bud95JFLRSDzoSUx8joaCDihG2kiBfehxpEcYsfQmC1JPcti4u80Cu+JUnSngInELUgEFmoH95fUjnHAiNGZIqZ7rxNpPkdQUM5KVvUSRGOExGpKeoQJxovx0Gj+DJ0bpw0EkzRMaTtXfGyniSk14aCbzsGrey8X/vF6iB5d+SkWcaCLw7NAgYVBHMO8C9qkkWLOJIQhLarJCPEISYW0aK5sS3PkvL5L2WdU9r9bua5X6VVFHCRyBY3AKXHAB6uAaNEELYJCCZ/AK3qwn68V6tz5mo0tWsXMA/sD6/AEVpZWS</latexit> L CON User-item Bi-partite Graph Modality-specifc Encoder LightGCN Graph Encoder <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω <latexit sha1_base64="DAo/YvITrhTTyNtbCl4oVBOR3Mo=">AAAB/nicbVDLSsNAFL3xWesrKq7cBIvgqiRS1GXRjcsK9gFNKJPJpB06mQkzE6GEgr/ixoUibv0Od/6NkzYLbT0wzOGce5kzJ0wZVdp1v62V1bX1jc3KVnV7Z3dv3z447CiRSUzaWDAheyFShFFO2ppqRnqpJCgJGemG49vC7z4SqajgD3qSkiBBQ05jipE20sA+9kPBIjVJzJX7JFWUCT4d2DW37s7gLBOvJDUo0RrYX34kcJYQrjFDSvU9N9VBjqSmmJFp1c8USREeoyHpG8pRQlSQz+JPnTOjRE4spDlcOzP190aOElUkNJMJ0iO16BXif14/0/F1kFOeZppwPH8ozpijhVN04URUEqzZxBCEJTVZHTxCEmFtGquaErzFLy+TzkXdu6w37hu15k1ZRwVO4BTOwYMraMIdtKANGHJ4hld4s56sF+vd+piPrljlzhH8gfX5A1w3lmc=</latexit> ω Interaction-guided Diffusion <latexit sha1_base64="iq9Fhz6AfE+jVKti/h9pbu8wfz4=">AAAB/nicbVDLSgMxFM34rPU1Kq7cBIvgqsxIUZdFEV2IVLAP6AxDJk3b0CQzJBmhDAP+ihsXirj1O9z5N2baWWjrgcDhnHu5JyeMGVXacb6thcWl5ZXV0lp5fWNza9ve2W2pKJGYNHHEItkJkSKMCtLUVDPSiSVBPGSkHY4uc7/9SKSikXjQ45j4HA0E7VOMtJECe9/jSA8xYultFqSe5PD66i4L7IpTdSaA88QtSAUUaAT2l9eLcMKJ0JghpbquE2s/RVJTzEhW9hJFYoRHaEC6hgrEifLTSfwMHhmlB/uRNE9oOFF/b6SIKzXmoZnMw6pZLxf/87qJ7p/7KRVxoonA00P9hEEdwbwL2KOSYM3GhiAsqckK8RBJhLVprGxKcGe/PE9aJ1X3tFq7r1XqF0UdJXAADsExcMEZqIMb0ABNgEEKnsEreLOerBfr3fqYji5Yxc4e+APr8wcMhZWM</latexit> L GEN reconstruction loss Eq. (4) <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="ZfUXZicXga9/bva7s61oGdMl2oQ=">AAACBXicbVC7TsMwFHXKq5RXgBGGiAqJqUpQBYwVLIxFog+piSLHcVqrjm3ZDlIVdWHhV1gYQIiVf2Djb3DaDNByJMtH59yre++JBCVKu+63VVlZXVvfqG7WtrZ3dvfs/YOu4plEuIM45bIfQYUpYbijiaa4LySGaURxLxrfFH7vAUtFOLvXE4GDFA4ZSQiC2kihfexHnMZqkpov97FQhHI2DXNfjMg0tOtuw53BWSZeSeqgRDu0v/yYoyzFTCMKlRp4rtBBDqUmiOJpzc8UFhCN4RAPDGUwxSrIZ1dMnVOjxE7CpXlMOzP1d0cOU1UsaipTqEdq0SvE/7xBppOrICdMZBozNB+UZNTR3CkicWIiMdJ0YghEkphdHTSCEiJtgquZELzFk5dJ97zhXTSad81667qMowqOwAk4Ax64BC1wC9qgAxB4BM/gFbxZT9aL9W59zEsrVtlzCP7A+vwBIXmZoQ==</latexit> ω ω <latexit sha1_base64="iq9Fhz6AfE+jVKti/h9pbu8wfz4=">AAAB/nicbVDLSgMxFM34rPU1Kq7cBIvgqsxIUZdFEV2IVLAP6AxDJk3b0CQzJBmhDAP+ihsXirj1O9z5N2baWWjrgcDhnHu5JyeMGVXacb6thcWl5ZXV0lp5fWNza9ve2W2pKJGYNHHEItkJkSKMCtLUVDPSiSVBPGSkHY4uc7/9SKSikXjQ45j4HA0E7VOMtJECe9/jSA8xYultFqSe5PD66i4L7IpTdSaA88QtSAUUaAT2l9eLcMKJ0JghpbquE2s/RVJTzEhW9hJFYoRHaEC6hgrEifLTSfwMHhmlB/uRNE9oOFF/b6SIKzXmoZnMw6pZLxf/87qJ7p/7KRVxoonA00P9hEEdwbwL2KOSYM3GhiAsqckK8RBJhLVprGxKcGe/PE9aJ1X3tFq7r1XqF0UdJXAADsExcMEZqIMb0ABNgEEKnsEreLOerBfr3fqYji5Yxc4e+APr8wcMhZWM</latexit> L GEN reconstruction loss Eq. (4) Total Correlation Maximization Similarity-based Fusion & Prediction recommendation BPR loss Eq. (11) <latexit sha1_base64="teNzxgB7JDpWwTRrTyTr5DmzNvk=">AAAB/nicbVDLSgMxFM3UV62vUXHlJlgEV2VGirosdePCRRX7gM4wZNK0DU0yQ5IRyjDgr7hxoYhbv8Odf2OmnYW2HggczrmXe3LCmFGlHefbKq2srq1vlDcrW9s7u3v2/kFHRYnEpI0jFsleiBRhVJC2ppqRXiwJ4iEj3XBynfvdRyIVjcSDnsbE52gk6JBipI0U2EceR3qMEUtvsyD1JIfN1n0W2FWn5swAl4lbkCoo0ArsL28Q4YQToTFDSvVdJ9Z+iqSmmJGs4iWKxAhP0Ij0DRWIE+Wns/gZPDXKAA4jaZ7QcKb+3kgRV2rKQzOZh1WLXi7+5/UTPbzyUyriRBOB54eGCYM6gnkXcEAlwZpNDUFYUpMV4jGSCGvTWMWU4C5+eZl0zmvuRa1+V682mkUdZXAMTsAZcMElaIAb0AJtgEEKnsEreLOerBfr3fqYj5asYucQ/IH1+QMbuJWW</latexit> L BPR <latexit sha1_base64="Y2773C1yEPjrlV3cOtBqNAwi4Gk=">AAAB9XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7APasWTSTBuaSYYko5Rh/sONC0Xc+i/u/Bsz7Sy09UDgcM693JMTxJxp47rfTmlldW19o7xZ2dre2d2r7h+0tUwUoS0iuVTdAGvKmaAtwwyn3VhRHAWcdoLJTe53HqnSTIp7M42pH+GRYCEj2FjpoR9hMw7CVGWDNMkG1Zpbd2dAy8QrSA0KNAfVr/5QkiSiwhCOte55bmz8FCvDCKdZpZ9oGmMywSPas1TgiGo/naXO0IlVhiiUyj5h0Ez9vZHiSOtpFNjJPKVe9HLxP6+XmPDKT5mIE0MFmR8KE46MRHkFaMgUJYZPLcFEMZsVkTFWmBhbVMWW4C1+eZm0z+reRf387rzWuC7qKMMRHMMpeHAJDbiFJrSAgIJneIU358l5cd6dj/loySl2DuEPnM8fV8GTEw==</latexit> r u <latexit sha1_base64="IXx7hPDuMT2OBPYugGKqYkd0RmQ=">AAAB9XicbVDLSsNAFL2pr1pfVZduBovgqiQi6rLoxmUF+4A2lsl00g6dPJi5UUrIf7hxoYhb/8Wdf+OkzUJbDwwczrmXe+Z4sRQabfvbKq2srq1vlDcrW9s7u3vV/YO2jhLFeItFMlJdj2ouRchbKFDybqw4DTzJO97kJvc7j1xpEYX3OI25G9BRKHzBKBrpoR9QHHt+qrJBitmgWrPr9gxkmTgFqUGB5qD61R9GLAl4iExSrXuOHaObUoWCSZ5V+onmMWUTOuI9Q0MacO2ms9QZOTHKkPiRMi9EMlN/b6Q00HoaeGYyT6kXvVz8z+sl6F+5qQjjBHnI5of8RBKMSF4BGQrFGcqpIZQpYbISNqaKMjRFVUwJzuKXl0n7rO5c1M/vzmuN66KOMhzBMZyCA5fQgFtoQgsYKHiGV3iznqwX6936mI+WrGLnEP7A+vwBVjyTEg==</latexit> r t <latexit sha1_base64="45bpSYtgnREc5+2Q0FdL305HNhU=">AAAB9XicbVDLSsNAFL2pr1pfUZduBovgqiRS1GXRjcsK9gFtLJPppB06mYSZSaWE/IcbF4q49V/c+TdO2iy09cDA4Zx7uWeOH3OmtON8W6W19Y3NrfJ2ZWd3b//APjxqqyiRhLZIxCPZ9bGinAna0kxz2o0lxaHPacef3OZ+Z0qlYpF40LOYeiEeCRYwgrWRHvsh1mM/SGU2SKfZwK46NWcOtErcglShQHNgf/WHEUlCKjThWKme68TaS7HUjHCaVfqJojEmEzyiPUMFDqny0nnqDJ0ZZYiCSJonNJqrvzdSHCo1C30zmadUy14u/uf1Eh1ceykTcaKpIItDQcKRjlBeARoySYnmM0MwkcxkRWSMJSbaFFUxJbjLX14l7Yuae1mr39erjZuijjKcwCmcgwtX0IA7aEILCEh4hld4s56sF+vd+liMlqxi5xj+wPr8AVlGkxQ=</latexit> r v <latexit sha1_base64="Y2773C1yEPjrlV3cOtBqNAwi4Gk=">AAAB9XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7APasWTSTBuaSYYko5Rh/sONC0Xc+i/u/Bsz7Sy09UDgcM693JMTxJxp47rfTmlldW19o7xZ2dre2d2r7h+0tUwUoS0iuVTdAGvKmaAtwwyn3VhRHAWcdoLJTe53HqnSTIp7M42pH+GRYCEj2FjpoR9hMw7CVGWDNMkG1Zpbd2dAy8QrSA0KNAfVr/5QkiSiwhCOte55bmz8FCvDCKdZpZ9oGmMywSPas1TgiGo/naXO0IlVhiiUyj5h0Ez9vZHiSOtpFNjJPKVe9HLxP6+XmPDKT5mIE0MFmR8KE46MRHkFaMgUJYZPLcFEMZsVkTFWmBhbVMWW4C1+eZm0z+reRf387rzWuC7qKMMRHMMpeHAJDbiFJrSAgIJneIU358l5cd6dj/loySl2DuEPnM8fV8GTEw==</latexit> r u <latexit sha1_base64="Y2773C1yEPjrlV3cOtBqNAwi4Gk=">AAAB9XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7APasWTSTBuaSYYko5Rh/sONC0Xc+i/u/Bsz7Sy09UDgcM693JMTxJxp47rfTmlldW19o7xZ2dre2d2r7h+0tUwUoS0iuVTdAGvKmaAtwwyn3VhRHAWcdoLJTe53HqnSTIp7M42pH+GRYCEj2FjpoR9hMw7CVGWDNMkG1Zpbd2dAy8QrSA0KNAfVr/5QkiSiwhCOte55bmz8FCvDCKdZpZ9oGmMywSPas1TgiGo/naXO0IlVhiiUyj5h0Ez9vZHiSOtpFNjJPKVe9HLxP6+XmPDKT5mIE0MFmR8KE46MRHkFaMgUJYZPLcFEMZsVkTFWmBhbVMWW4C1+eZm0z+reRf387rzWuC7qKMMRHMMpeHAJDbiFJrSAgIJneIU358l5cd6dj/loySl2DuEPnM8fV8GTEw==</latexit> r u <latexit sha1_base64="3ix4ez9mkrq+xbqtUZ4SHORIKWs=">AAAB9XicbVDLSsNAFL2pr1pfVZduBovgqiQi6rLoxmUF+4A2lsn0ph06eTAzUUrIf7hxoYhb/8Wdf+OkzUJbDwwczrmXe+Z4seBK2/a3VVpZXVvfKG9WtrZ3dveq+wdtFSWSYYtFIpJdjyoUPMSW5lpgN5ZIA09gx5vc5H7nEaXiUXivpzG6AR2F3OeMaiM99AOqx56fymyQYjao1uy6PQNZJk5BalCgOah+9YcRSwIMNRNUqZ5jx9pNqdScCcwq/URhTNmEjrBnaEgDVG46S52RE6MMiR9J80JNZurvjZQGSk0Dz0zmKdWil4v/eb1E+1duysM40Riy+SE/EURHJK+ADLlEpsXUEMokN1kJG1NJmTZFVUwJzuKXl0n7rO5c1M/vzmuN66KOMhzBMZyCA5fQgFtoQgsYSHiGV3iznqwX6936mI+WrGLnEP7A+vwBP3GTAw==</latexit> r e <latexit sha1_base64="IBbgQLF+0E3r5TIHUgy/f18h/fY=">AAAB8XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV+sK2lEyaaUMzmSG5I5Shf+HGhSJu/Rt3/o2ZdhbaeiBwOOdecu7xYykMuu63U1hb39jcKm6Xdnb39g/Kh0ctEyWa8SaLZKQ7PjVcCsWbKFDyTqw5DX3J2/7kLvPbT1wbEakGTmPeD+lIiUAwilZ67IUUx36QNmaDcsWtunOQVeLlpAI56oPyV28YsSTkCpmkxnQ9N8Z+SjUKJvms1EsMjymb0BHvWqpoyE0/nSeekTOrDEkQafsUkrn6eyOloTHT0LeTWUKz7GXif143weCmnwoVJ8gVW3wUJJJgRLLzyVBozlBOLaFMC5uVsDHVlKEtqWRL8JZPXiWti6p3Vb18uKzUbvM6inACp3AOHlxDDe6hDk1goOAZXuHNMc6L8+58LEYLTr5zDH/gfP4AyXORAQ==</latexit> T <latexit sha1_base64="vvQbFvx+rCW34TIJa00M9KzySqc=">AAAB8XicbVDLSgMxFL1TX7W+qi7dBIvgqsxIUZdFNy4r2Ae2Q8mkd9rQTGZIMkIZ+hduXCji1r9x59+YtrPQ1gOBwzn3knNPkAiujet+O4W19Y3NreJ2aWd3b/+gfHjU0nGqGDZZLGLVCahGwSU2DTcCO4lCGgUC28H4dua3n1BpHssHM0nQj+hQ8pAzaqz02IuoGQVh1pr2yxW36s5BVomXkwrkaPTLX71BzNIIpWGCat313MT4GVWGM4HTUi/VmFA2pkPsWipphNrP5omn5MwqAxLGyj5pyFz9vZHRSOtJFNjJWUK97M3E/7xuasJrP+MySQ1KtvgoTAUxMZmdTwZcITNiYgllitushI2ooszYkkq2BG/55FXSuqh6l9Xafa1Sv8nrKMIJnMI5eHAFdbiDBjSBgYRneIU3RzsvzrvzsRgtOPnOMfyB8/kDzH2RAw==</latexit> V <latexit sha1_base64="1psKDpBYNiVFl1nO3qlIsGili+0=">AAAB8XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7QPbUjLpnTY0kxmSjFCG/oUbF4q49W/c+Tdm2llo64HA4Zx7ybnHjwXXxnW/ncLK6tr6RnGztLW9s7tX3j9o6ihRDBssEpFq+1Sj4BIbhhuB7VghDX2BLX98k/mtJ1SaR/LBTGLshXQoecAZNVZ67IbUjPwgvZ/2yxW36s5AlomXkwrkqPfLX91BxJIQpWGCat3x3Nj0UqoMZwKnpW6iMaZsTIfYsVTSEHUvnSWekhOrDEgQKfukITP190ZKQ60noW8ns4R60cvE/7xOYoKrXsplnBiUbP5RkAhiIpKdTwZcITNiYgllitushI2ooszYkkq2BG/x5GXSPKt6F9Xzu/NK7TqvowhHcAyn4MEl1OAW6tAABhKe4RXeHO28O/Ox3y04OQ7h/AHzucPx+6RAA==</latexit> S <latexit sha1_base64="1psKDpBYNiVFl1nO3qlIsGili+0=">AAAB8XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7QPbUjLpnTY0kxmSjFCG/oUbF4q49W/c+Tdm2llo64HA4Zx7ybnHjwXXxnW/ncLK6tr6RnGztLW9s7tX3j9o6ihRDBssEpFq+1Sj4BIbhhuB7VghDX2BLX98k/mtJ1SaR/LBTGLshXQoecAZNVZ67IbUjPwgvZ/2yxW36s5AlomXkwrkqPfLX91BxJIQpWGCat3x3Nj0UqoMZwKnpW6iMaZsTIfYsVTSEHUvnSWekhOrDEgQKfukITP190ZKQ60noW8ns4R60cvE/7xOYoKrXsplnBiUbP5RkAhiIpKdTwZcITNiYgllitushI2ooszYkkq2BG/x5GXSPKt6F9Xzu/NK7TqvowhHcAyn4MEl1OAW6tAABhKe4RXeHO28O/Ox3y04OQ7h/AHzucPx+6RAA==</latexit> S <latexit sha1_base64="Y2773C1yEPjrlV3cOtBqNAwi4Gk=">AAAB9XicbVDLSgMxFL1TX7W+qi7dBIvgqsyIqMuiG5cV7APasWTSTBuaSYYko5Rh/sONC0Xc+i/u/Bsz7Sy09UDgcM693JMTxJxp47rfTmlldW19o7xZ2dre2d2r7h+0tUwUoS0iuVTdAGvKmaAtwwyn3VhRHAWcdoLJTe53HqnSTIp7M42pH+GRYCEj2FjpoR9hMw7CVGWDNMkG1Zpbd2dAy8QrSA0KNAfVr/5QkiSiwhCOte55bmz8FCvDCKdZpZ9oGmMywSPas1TgiGo/naXO0IlVhiiUyj5h0Ez9vZHiSOtpFNjJPKVe9HLxP6+XmPDKT5mIE0MFmR8KE46MRHkFaMgUJYZPLcFEMZsVkTFWmBhbVMWW4C1+eZm0z+reRf387rzWuC7qKMMRHMMpeHAJDbiFJrSAgIJneIU358l5cd6dj/loySl2DuEPnM8fV8GTEw==</latexit> r u <latexit sha1_base64="h5usJAUa4jOuhhQxrNFNLdxWaxw=">AAAB+XicbVDLSsNAFL2pr1pfUZduBovgqiQi6rLoxmWFvqAJZTKdtEMnkzAzKZSQP3HjQhG3/ok7/8ZJm4W2Hhg4nHMv98wJEs6Udpxvq7KxubW9U92t7e0fHB7ZxyddFaeS0A6JeSz7AVaUM0E7mmlO+4mkOAo47QXTh8LvzahULBZtPU+oH+GxYCEjWBtpaNtehPUkCDMvwDJr5/nQrjsNZwG0TtyS1KFEa2h/eaOYpBEVmnCs1MB1Eu1nWGpGOM1rXqpogskUj+nAUIEjqvxskTxHF0YZoTCW5gmNFurvjQxHSs2jwEwWOdWqV4j/eYNUh3d+xkSSairI8lCYcqRjVNSARkxSovncEEwkM1kRmWCJiTZl1UwJ7uqX10n3quHeNK6fruvN+7KOKpzBOVyCC7fQhEdoQQcIzOAZXuHNyqwX6936WI5WrHLnFP7A+vwBFZyT9w==</latexit> ̄ T <latexit sha1_base64="q0Vi3jLcIp/1V7W7Z7S9aJMu3yw=">AAAB+XicbVDLSsNAFL2pr1pfUZduBovgqiQi6rLoxmVF+4AmlMl00g6dTMLMpFBC/sSNC0Xc+ifu/BsnbRbaemDgcM693DMnSDhT2nG+rcra+sbmVnW7trO7t39gHx51VJxKQtsk5rHsBVhRzgRta6Y57SWS4ijgtBtM7gq/O6VSsVg86VlC/QiPBAsZwdpIA9v2IqzHQZh5AZbZY54P7LrTcOZAq8QtSR1KtAb2lzeMSRpRoQnHSvVdJ9F+hqVmhNO85qWKJphM8Ij2DRU4osrP5slzdGaUIQpjaZ7QaK7+3shwpNQsCsxkkVMte4X4n9dPdXjjZ0wkqaaCLA6FKUc6RkUNaMgkJZrPDMFEMpMVkTGWmGhTVs2U4C5/eZV0LhruVePy4bLevC3rqMIJnMI5uHANTbiHFrSBwBSe4RXerMx6sd6tj8VoxSp3juEPrM8fFBaT9g==</latexit> ̄ S <latexit sha1_base64="LHHEcNf9s5zS56viB6bV+dcVaMs=">AAAB+XicbVDLSsNAFL2pr1pfUZduBovgqiRS1GXRjcsK9gFNKJPppB06mYSZSaGE/IkbF4q49U/c+TdO2iy09cDA4Zx7uWdOkHCmtON8W5WNza3tnepubW//4PDIPj7pqjiVhHZIzGPZD7CinAna0Uxz2k8kxVHAaS+Y3hd+b0alYrF40vOE+hEeCxYygrWRhrbtRVhPgjDzAiyzbp4P7brTcBZA68QtSR1KtIf2lzeKSRpRoQnHSg1cJ9F+hqVmhNO85qWKJphM8ZgODBU4osrPFslzdGGUEQpjaZ7QaKH+3shwpNQ8CsxkkVOteoX4nzdIdXjrZ0wkqaaCLA+FKUc6RkUNaMQkJZrPDcFEMpMVkQmWmGhTVs2U4K5+eZ10rxrudaP52Ky37so6qnAG53AJLtxACx6gDR0gMINneIU3K7NerHfrYzlascqdU/gD6/MHGKiT+Q==</latexit> ̄ V Figure 3: Overall illustration of the proposed GTC framework. Interaction-Guided Denoising. The first core step of GTC is to refine the generic content embeddings(V,T)to reflect individual user preferences. To achieve this, we regard the interaction embed- dings S as a reflection of user preference and treat it as a user-specific guidance to denoise the content features V and T. Concretely, we learn an interaction-guided diffusion model over the content em- bedding spaces [29], which we apply independently to each initial content embedding matrix C 0 ∈ V, T. The forward diffusion process defines a Markov chain that grad- ually injects Gaussian noise into the content embedding C 0 over푇 timesteps. It ends up with C 푇 being a pure Gaussian, from which the state C 푡 at any timestep 푡 can be sampled in a closed form: C 푡 = √ ̄ 훼 푡 C 0 + √ 1− ̄ 훼 푡 흐,(3) where흐 ∼ N(0,I), and ̄ 훼 푡 = Î 푡 푠=1 ( 1− 훽 푠 )is determined by a predefined linear variance schedule훽 푡 푇 푡=1 . The reverse process aims to invert Eq. (3) by modeling the con- ditional distribution푝 휃 (C 푡−1 |C 푡 ,S)parameterized by a neural net- work흐 휙 that predicts the noise흐added at timestep푡during the forward corruption process. This denoising network is conditioned on the interaction embeddings S, which encode user-specific in- teraction patterns. Intuitively, conditioning on S encourages흐 휙 to leverage user preference signals to determine which parts of the intermediate content signal are noise and should therefore be re- moved. We instantiate흐 휙 with a U-Net architecture [19]. To provide temporal context, we map timestep푡to a high-dimensional embed- dingPE(푡)using sinusoidal positional encoding [22], which is in- jected into the network. Following the simplified training objective [9], the network흐 휙 is trained by minimizing the noise-prediction loss over timesteps 푡= 1, . . .,푇 , L GEN = E 푡,C 0 ,흐 h 흐−흐 휙 C 푡 ,푡, S 2 i .(4) Note that Eq. (4) applies to both V and T simultaneously. Content Features Generation. After training, we generate the user- conditional content representations e C 0 ∈ e V 0 , e T 0 from the learned diffusion model. We first sample the terminal state from a standard Gaussian C 푇 ∼N(0,I), and then iteratively denoise C 푇 back to C 0 using the trained denoising network흐 휙 . Concretely, for timesteps 푡=푇, . . ., 1, we update e C 푡−1 = 1 √ 훼 푡 e C 푡 − 1− 훼 푡 √ 1− ̄ 훼 푡 ,흐 휙 ( e C 푡 ,푡, S) + 휎 푡 z,(5) where z∼ N(0,I)is sampled from standard Gaussian, and휎 푡 controls the stochasticity.흐 휙 is instantiated as a U-Net architecture following [19]. We omit further architectural details for brevity. 3.3 Modal Alignment via Total Correlation Having generated user-aware content representations, we align S with e V 0 and e T 0 by maximizing their total correlationTC(S, e V 0 , e T 0 ), which quantifies how strongly the triplet(S, e V 0 , e T 0 )are dependent. As Eq. (1) establishes, maximizingTC(S, e V 0 , e T 0 )is sufficient to si- multaneously optimize both the pairwise alignment (e.g.,퐼(S; e V 0 )) and the higher-order dependencies (e.g.,퐼(S; e V 0 | e T 0 )) that simpler pairwise objectives neglect [21, 36]. For notational convenience, we hereinafter denote ̄ V and ̄ T the generated content representations e V 0 and e T 0 . Defined as the KL divergence between the joint distribution and the product of marginals [25], TC(S, ̄ V, ̄ T) can be expressed as: TC(S, ̄ V, ̄ T)= E S, ̄ V, ̄ T log 푝(S, ̄ V, ̄ T) 푝(S)푝( ̄ V)푝( ̄ T) .(6) Since direct maximization of Eq. (6) is intractable, we follow [20] and optimize a tractable variational lower bound on Eq. (6) via the InfoNCE contrastive loss [16], aligning the modalities by distin- guishing samples from the joint distribution푝(S, ̄ V, ̄ T)against those sampled from the product of marginals. Multi-modal Contrastive Formulation. We thus formulate the alignment objective as a contrastive learning task over푀mini- batches of multi-modal samples from the underlying data distribu- tion. Sampling Strategy. For a batch of푁samples, we treat the matched triplet(S 푖 , ̄ V 푖 , ̄ T 푖 )as the positive anchor for푖=1, . . .,푁drawn from the joint distribution푝(S, ̄ V, ̄ T). To construct negative samples that approximate the product of marginals, we independently shuffle the User-Aware Conditional Generative Total Correlation Learning for Multi-Modal RecommendationConference’17, July 2017, Washington, DC, USA indices of the visual and textual representations within the batch, denoted by휋 푣 and휋 푡 . For each anchor푖, we thus construct푁 −1 negative samples against the matched positive triplet(S 푖 , ̄ V 푖 , ̄ T 푖 ). Objective. We define the contrastive log-likelihood for anchor S as: L S→ ̄ V, ̄ T CON = log 푁 ∑︁ 푖=1 expℎ(S 푖 , ̄ V 푖 , ̄ T 푖 ) Í 푗≠푖 expℎ(S 푖 , ̄ V휋 푣(푗) , ̄ T휋 푣(푗) ) ,(7) whereℎ(x,y,z)≜ Í 퐷 푑=1 x (푑) y (푑) z (푑) is a multilinear inner prod- uct. Maximizing this objective corresponds to maximizing a lower bound on TC(S, ̄ V, ̄ T) [20, 25]: TC(S, ̄ V, ̄ T) ≥ log푁 + E D [L S→ ̄ V, ̄ T CON ].(8) Here, the expectation is taken over푀mini-batches of samples. According to [16], this bound holds because the optimalℎ ∗ implicitly estimates the log-density ratio defined within Eq. (6). Symmetrization. AsTC(S, ̄ V, ̄ T)is a symmetric measure over three modalities, while the single-anchor lossL S→ ̄ V, ̄ T CON is not, we sym- metrize the objective to cover all modal perspectives, such that L CON =L S→ ̄ V, ̄ T CON +L ̄ V→S, ̄ T CON +L ̄ T→S, ̄ V CON .(9) By minimizing Eq. (9), we encourage the model to capture higher- order dependencies among S, ̄ V , and ̄ T that are inaccessible to purely pairwise contrastive objectives, leading to a more comprehensive multi-modal alignment. 3.4 Fused Representation for Recommendation Similarity-based Fusion. With user-aware and aligned represen- tations S, ̄ V, ̄ Tavailable, we now integrate them into a coherent rep- resentation. For each item (or user–item entity) with embeddings S 푖 , ̄ V 푖 , and ̄ T 푖 ∈ R 퐷 , we first construct a fused content represen- tation via element-wise interaction ̄ C 푓 = ̄ V⊙ ̄ T. We then use this representation to refine the original interaction embedding S using a similarity-based gating mechanism that dynamically controls the content contributions based on their relevance to the interaction patterns. This yields a content-aware update to S as S 푓 = 훼 · ̄ C 푓 = softmax S· ̄ C 푓 /휏 ||S||·|| ̄ C 푓 || · ̄ C 푓 ,(10) where휏is a learnable temperature parameter. To preserve the primary collaborative signal in S while allowing the content signal to enhance it in a similarity-adaptive manner, we obtain the final fused representation via a residual connection ̄ S= S+ S 푓 . Recommendation. The final user and item representations are used to compute the ranking score푠(푢,푖)and predict the user’s next interaction. We optimize the interacted items to be ranked higher than un-interacted ones by minimizing the BPR loss [18]: L BPR = ∑︁ (푢,푖 po ,푖 ne )∈D − log휎(푠(푢,푖 po )−푠(푢,푖 ne )),(11) whereDis the set of training triplets, and푖 po and푖 ne are interacted and un-interacted items for user푢. 휎(·) is the sigmoid function. 3.5 Training Objective The complete GTC framework is trained end-to-end by minimizing a composite objective function that combines recommendation loss (Eq. (11)), generation loss (Eq. (4)), and total correlation lower bound (Eq. (9)), with ℓ 2 regularization applied to parametersΘ, such that L GTC =L BPR + 휔 1 ·L GEN + 휔 2 ·L CON +||Θ|| 2 ,(12) where휔 1 and휔 2 are hyperparameters that control the relative importance of the three components. 3.6 Computational Complexity The computational complexity of GTC mainly contains three com- ponents: (1)Encoders. GTC employs three parallel LightGCNs, resulting inO(퐿|E|푑). (2)Interaction-guided Denoisingand Content Features Generation. Interaction-guided Denoising samples a single timestep푡, constructing noisy content representa- tion C 푡 from C 0 via the closed-form forward process. Ignoring the negligible O(퐵푑)noise sampling and element-wise operations, the per-iteration time complexity is O(푐표푠푡 휙 (퐵,푑)), where푐표푠푡 휙 (퐵,푑) denotes the complexity of a single-forward pass of the U-Net net- work흐 휙 on a mini-batch of size퐵with푑-dimensional content representations. In the Content Features Generation stage, we iter- atively apply the learned reverse diffusion from C 푇 to C 0 . Each of the푇reverse steps requires one forward pass of흐 휙 , yielding overall complexity O(푇 · 푐표푠푡 휙 (퐵,푑)). (2)Total Correlation.L CON is implemented as a multi-modal InfoNCE loss over푁triplets. We compute scores against all in-batch candidates for each anchor using multilinear functionℎ(·). This all-pairs construction yields O(푁 2 푑). By adding three such terms, which add only a constant factor, the overall complexity remains O(푁 2 푑). 4 Experiments In this section, we conduct a set of experiments to validate GTC, structured around six research questions (RQs): RQ1) Can GTC outperform selected MMR baselines? RQ2) How does each component contribute to performance? RQ3) What are the impacts of visual and textual modalities? RQ4) Does GTC effectively ensure cross-modal balance? RQ5) Does GTC guarantee user preference consistency? RQ6) How do hyperparameters affect the GTC’s performance? 4.1 Experimental Setup Datasets. Following standard practice [36], we evaluate GTC on 3 widely used public benchmarks from the Amazon Review Datasets: Sports and Outdoors (Sports), Baby (Baby), and Cellphone (Cell). We filter for users and items with at least five interactions and randomly split the data into 80% for training, 10% for validation, and 10% for testing. The dataset statistics are provided in Table. 2. Baselines. We compare GTC against 10 recent MMR methods, in- cluding fusion-based methods (FREEDOM [35], GRCN [26], SMORE [15], LATTICE [32], MGCN [30], PGL [31]) and distenglement-based methods (SLMRec [21], DRAGON [34], BM3 [36], LGMRec [6]). Implementation. For content modality encoders, ResNet-50 [7] and sentence-transformer [17] are used to extract visual features (푑 푣 =4096) and textual features (푑 푡 =384). All baselines and GTC are trained using the Adam optimizer with a learning rate of 0.001 and a representation dimension of 64. For GTC, the diffusion pro- cess uses푇=500 timesteps for Sports,푇=600 for Baby, and Conference’17, July 2017, Washington, DC, USAJing Du et al. Table 1: Overall Performance in Sports (up), Baby (middle), and Cell (down) dataset. For each comparison, the best-performing method is bolded. Improved indicates the percentage improvement of the proposed GTC over therunner-up model. * denote statistically significant improvements, validated by a paired t-test at a significance level of푝<0.05 against therunner-up model. ModelsFREEDOMSMOREGRCNLATTICEBM3PGLSLMRecLGMRecMGCNDRAGONGTCImproved(%) NDCG @50.01960.01810.02450.02620.02670.02940.02910.02910.03140.07490.096128.30% @100.02570.02260.03160.03350.03430.03760.03680.03750.03990.08800.110625.68% @200.03270.02740.03940.04200.04300.04720.04510.04680.04960.10050.124223.58% @500.04370.03360.05100.05440.05520.06110.05640.06010.06360.11750.141120.09% Recall @50.02900.02680.03730.03920.03970.04410.04340.04420.04750.10400.130725.67% @100.04770.04070.05900.06130.06280.06940.06680.06980.07340.14410.174921.37% @200.07480.05920.08880.09420.09650.10650.09890.10580.11150.19280.227518.00% @500.12890.08970.14610.15550.15660.17510.15450.17120.18060.27640.310812.45% MAP @50.01590.01470.01950.02130.02170.02380.02360.02340.02530.06400.083029.69% @100.01840.01660.02240.02420.02470.02710.02670.02680.02870.06930.088928.28% @200.02020.01790.02440.02650.02710.02960.02890.02930.03130.07270.092627.37% @500.02190.01880.02630.02840.02900.03180.03070.03140.03350.07540.095326.39% ModelsFREEDOMSMOREGRCNLATTICEBM3PGLSLMRecLGMRecMGCNDRAGONGTCImproved(%) NDCG @50.02670.01940.02200.02250.02190.02610.02340.02610.02630.02630.02837.60% @100.03350.02430.02860.02880.02870.03360.02960.03430.03370.03440.03563.49% @200.04120.02940.03630.03680.03740.04260.03610.04340.04320.04340.04421.84% @500.05170.03530.04830.04910.05030.05640.04600.05580.05300.05420.05762.13% Recall @50.03900.02960.03260.03420.03270.03920.03520.03990.04000.03940.04092.25% @100.06170.04490.05270.05310.05330.06180.05400.06440.06270.06570.06762.89% @200.09220.06470.08250.08400.08680.09700.07920.09170.09500.09600.10083.92% @500.14460.09400.14130.14510.15070.16530.12770.16290.14410.15010.17093.39% MAP @50.02180.01590.01790.01820.01780.02110.01900.01820.02110.02110.02319.48% @100.02460.01790.02060.02070.02050.02410.02140.02060.02410.02430.02617.41% @200.02670.01930.02260.02280.02290.02650.02320.02240.02600.02630.02867.92% @500.02830.02020.02450.02480.02490.02830.02470.02390.02810.02840.02892.12% ModelsFREEDOMSMOREGRCNLATTICEBM3PGLSLMRecLGMRecMGCNDRAGONGTCImproved(%) NDCG @50.05080.05310.04390.04770.05040.05080.04930.05150.05180.04990.05473.01% @100.06180.06500.05450.05870.06140.06220.06050.06360.06390.06290.06703.08% @200.07470.07760.06580.07060.07300.07440.07180.07640.07660.07550.08013.22% @500.09100.09500.08130.08640.08960.09270.08710.09420.09340.09260.09823.37% Recall @50.07600.07780.06660.07080.07470.07530.07440.07700.07780.07450.07962.31% @100.11010.11260.09920.10490.10850.11050.10900.11470.11500.11440.11853.04% @200.16040.16340.14360.15180.15420.15830.15320.16430.16480.16400.16932.73% @500.24230.25150.22080.23070.23740.24980.22970.25050.24850.24960.25722.27% MAP @50.04210.04320.03600.03960.04210.04240.04060.04290.04290.04140.04422.31% @100.04660.04810.04030.04410.04660.04700.04510.04780.04780.04660.04912.08% @200.05010.05160.04330.04740.04970.05030.04820.05130.05120.05000.05455.62% @500.05270.05440.04580.04990.05230.05320.05060.05410.05390.05280.05633.49% Table 2: Dataset statistics. Datasets# interactionsparsity useritem numavg interactionnumavg interaction Sports296,33799.95%35,5988.324518,35716.1430 Baby160,79299.88%19,4458.26917,05022.8074 Cell194,43999.93%27,8796.974410,42918.6441 푇=700 for Cell. The regularization weight is set to 0.01. In the interaction-guided denoising, the noise starts at훽 푠 =1푒 −4 and increases to훽 푡 =0.02. Experiments show that variations in the noise levels only affect the model convergence speed and have no obvious impact on model performance. The weights푤 1 and푤 2 are searched in0.1,0.2,0.3,0.4,0.5,0.6,0.7,0.8,0.9,1. For noise steps, we search in 100, 200, 300, 400, 500, 600, 700, 800, 900, 1000. We report recommendation performance using Normalized Discounted Cumulative Gain (NDCG@K), Mean Average Precision (MAP@K), and F1@K for top-퐾items with퐾 ∈ 5,10,20,50. We run experi- ments with 10 different random seeds and report the final results. All experiments are run on either a single NVIDIA V100, DGX A100, or NVIDIA RTX A5000 GPU. 4.2 Overall Performance (RQ1) Table. 1 shows the main results. GTC consistently and significantly outperforms all baselines across all datasets and metrics. On the Sports dataset, GTC achieves 28.30% improvement in NDCG@5 over the strongest baseline. The performance gains on the Baby dataset are also consistent, reaching up to 9.48% in MAP@5. These results provide an affirmative answer to RQ1. Moreover, we observe the following patterns reflecting the limitations of prior methods. Inconsistent Multi-Modal Features Hinder Simple Fusion. SMORE and GRCN generate features for each modality in isolation and com- bine them via simple fusion, leading to the weakest performance. LGMRec and PGL explore multimodal information by capturing global and local structural patterns, but they still lack mechanisms to remove irregular noise. FREEDOM and LATTICE attempt to har- ness textual and visual features in modeling user-item interactions, but they merely use learnable weights to differentiate the modality importance and aggregate them without considering their mutual entanglement. This suggests that overlooking the inconsistencies and potential noise between modalities limits the quality of MMR. User-Aware Conditional Generative Total Correlation Learning for Multi-Modal RecommendationConference’17, July 2017, Washington, DC, USA Table 3: Ablation experiments on Sports (up), Baby dataset (middle), and Cell (bottom). Module RecallNDCG @5@10@20@50@5@10@20@50 GTC Base0.09950.13670.18410.26100.07090.08310.09530.1109 GTC Base w/ DN0.12290.16630.21800.30200.08940.10370.11700.1340 GTC Base w/ TC0.12060.16410.21640.29940.08840.10270.11620.1330 GTC w/o TC0.10990.15060.19820.27940.07850.09180.10410.1206 GTC0.1307 0.1749 0.2275 0.31080.0961 0.1106 0.1242 0.1411 Module RecallNDCG @5@10@20@50@5@10@20@50 GTC Base0.03910.06190.09760.16670.02630.03370.04290.0569 GTC Base w/ DN0.04050.06680.10040.16990.02780.03540.04330.0582 GTC Base w/ TC0.04020.06600.10010.16870.02680.03520.04310.0579 GTC w/o TC0.04010.06520.09960.16860.02720.03530.04300.0576 GTC 0.0409 0.0676 0.1008 0.17090.0283 0.0356 0.0442 0.0585 Module RecallNDCG @5@10@20@50@5@10@20@50 GTC Base0.07720.11260.16090.24410.05200.06350.07580.0924 GTC Base w/ DN0.07670.11410.16440.24590.05160.06380.07650.0929 GTC Base w/ TC0.07790.11600.16690.25500.05250.06480.07780.0954 GTC w/o TC 0.07770.11470.16400.25090.05270.06480.07730.0947 GTC0.0796 0.1185 0.1693 0.25720.0547 0.0670 0.0801 0.0982 Refinement is Beneficial. Methods like BM3, SLMRec, MGCN, and DRAGON, which incorporate modules to denoise or align multi-modal features, demonstrate stronger performance relative to fusion-based methods. Specifically, DRAGON’s empirical evalua- tion shows that indiscriminately fusing multi-modal features can reduce overall performance [34]. MGCN argues that conventional fusion techniques, such as concatenation or mean-pooling, fail to capture the ever-changing importance of features, and therefore implements an attention network to adjust the importance of each modality dynamically. BM3 and SLMRec adopt contrastive learn- ing to ensure effective cross-modal alignment. Their performances confirm the importance of addressing modality inconsistency. Pairwise Alignment is Still a Bottleneck. However, these top- performing baselines rely on aligning modalities through pairwise contrastive objectives. To generate refined embeddings for feature fusion, SLMRec employs DRL-based techniques to build hyper- graphs that capture both shared and unique modality features by optimizing contrastive loss. Likewise, BM3 and MGCN perform inter-modality alignment by focusing on pairwise correlations through contrastive learning. GTC, however, leads over even the strongest baseline, providing evidence that this pairwise approach is a performance bottleneck. It supports our claim that capturing the higher-order correlations unlocks a new level of performance. 4.3 Ablation Study (RQ2) To dissect the contribution of GTC’s key components, we evaluate several variants on the three datasets, shown in Table. 3. Namely, GTC Baseis minimal baseline using only LightGCNs with concate- nation.GTC Base w/ DNis the baseline equipped with the denoising process.GTC Base w/ TCuses total correlation optimization (i.e., GTC variant removing denoisingL GEN ).GTC w/o TCreplaces the to- tal correlation objectiveL CON with a standard pairwise contrastive loss, while retaining the diffusion module. The results clearly show the necessity of both components. Interaction-guided denoising (GTC Base w/ DN) yields a substantial performance improvement (e.g., an 2.34% relative increase in Recall@5 in Sports) compared toGTC Base, indicating that denoising content features with user- specific signals is critical. Similarly, both total correlation (GTC Base w/ TC) and pairwise contrastive objectives (GTC w/o TC) provide 5102050 Top-k Values 0.00 0.02 0.04 0.06 0.08 0.10 0.12 0.14 NDCG Score Ablation NDCG Across Top-k Values inter only w/o vis w/o txt w/o TC GTC 5102050 Top-k Values 0.00 0.05 0.10 0.15 0.20 0.25 0.30 Recall Score Ablation Recall Across Top-k Values inter only w/o vis w/o txt w/o TC GTC 5102050 Top-k Values 0.020 0.025 0.030 0.035 0.040 0.045 0.050 0.055 0.060 NDCG Score Ablation NDCG Across Top-k Values inter only w/o vis w/o txt w/o TC GTC 5102050 Top-k Values 0.04 0.06 0.08 0.10 0.12 0.14 0.16 Recall Score Ablation Recall Across Top-k Values inter only w/o vis w/o txt w/o TC GTC 5102050 Top-k Values 0.05 0.06 0.07 0.08 0.09 NDCG Score Ablation NDCG Across Top-k Values inter only w/o vis w/o txt w/o TC GTC 5102050 Top-k Values 0.075 0.100 0.125 0.150 0.175 0.200 0.225 0.250 Recall Score Ablation Recall Across Top-k Values inter only w/o vis w/o txt w/o TC GTC Figure 4: Impact of content features in Sports (up), Baby (middle), and Cell (bottom) Datasets. consistent gains overGTC Base, highlighting the benefit of enforc- ing cross-modal consistency. However,GTC w/o TCdegradesGTC Base w/ TC’s performance, supporting that higher-order depen- dencies are crucial for alignment. The full GTC model, integrating both components, achieves the best performance. 4.4 Impact of Content Features (RQ3) Fig. 4 shows the performance when modalities are selectively re- moved. Theinter-onlyis a non-multimodal baseline that uses only user-item interaction data. Thew/o visualandw/o textual settings replace the total correlation approach with pairwise corre- lations implemented through standard contrastive learning as only two modalities are involved, confirming the value of side informa- tion with consistently improved performances over theinter-only baseline. Thew/o TCuses all three modalities but with pairwise alignment, which outperforms the two-modality versions, showing that more modalities provide richer information. Still, GTC con- sistently outperforms all these variants on both Sports and Baby datasets. This shows GTC’s ability to not just use, but holistically Conference’17, July 2017, Washington, DC, USAJing Du et al. 050100150200 Epoch 0.575 0.600 0.625 0.650 0.675 0.700 0.725 0.750 Modal Balance Score Sports Dataset 05101520253035 Epoch 0.52 0.54 0.56 0.58 0.60 0.62 0.64 0.66 0.68 Baby Dataset 020406080100 Epoch 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Cell Dataset Figure 5: Modality balance trend during training GTC. integrate and refine information from all available modalities, un- locking performance unattainable by pairwise methods. 4.5 Cross-modal Consistency (RQ4) Lastly, to assess whether GTC produces balanced and consistent cross-modal representations, we introduce the modality balance score. For each user in the test set, we compute the cosine similarities sim( ̄ S, ̄ E 푚 )between the final fused representation ̄ Sand each of the three modality-specific representations ̄ E 푚 for all푚 ∈ I,V,T. Specifically, this balance score is defined as 1−(max(sim( ̄ S, ̄ E 푚 ))− min(sim( ̄ S, ̄ E 푚 ))), where a higher score (closer to 1) indicates less disparity and thus a more balanced integration. It is obvious from Fig. 5 that, the modality balance score obtained by GTC steadily increases during training for both datasets as epochs increases. This upward trend provides evidence that the GTC framework successfully learns to pull the representations from different modalities into a coherent space. Notably, the curve for the Baby dataset is less smooth, suggesting that user preferences in this domain are more heterogeneous, making the alignment task more challenging. Nonetheless, GTC is still able to handle effectively. 4.6 User Preference Consistency (RQ5) To empirically validate the consistency achieved between learned modality features and user preferences, we visualize the dot product trends between user embeddings and the modality embeddings (i.e., ̄ S, ̄ V, ̄ T) within our GTC, the runner-up DRAGON, and MGCN (pair- wise contrastive loss to enforce consistency). Crucially, we clarify that a higher absolute dot product value does not inherently correlate with superior recommendation performance. As shwon in Fig. 6, even without explicit constraints in DRAGON, dot products among three modalities gradually increase as the training process progresses, indicating that better performances implicitly require higher consistency between different modali- ties. DRAGON’s user-interaction curve consistently surpasses its user-image and user-text curves, suggesting a strong reliance on interaction signals with visual and textual modalities contributing to a more limited extent. Similarly, while MGCN utilizes pairwise contrastive loss to ensure steady upward trends of all modalities, it still suffers from an imbalance problem in modality contribu- tions, especially late in training. GTC shows a balanced upward trend across user-interaction, user-visual, and user-text curves, signi- fying that its interaction-guided denoising mechanism effectively enhances the alignment of diverse modalities with user preference signals. This enhanced alignment leads to significant improvements in ranking results, confirming GTC’s superior ability to leverage multimodal information for recommendations. 4.7 Hyperparameter Evaluation (RQ6) GTC involves 3 important hyperparameters: noise step푇, diffusion weight푤 1 , and total correlation weight푤 2 . To investigate the best configuration, we conduct the hyperparameter search on three datasets, shown in Fig. 7. In the user-aware content feature generation, noise step푇controls the granularity of the noise injection and subsequent removal. A larger푇implies a more gradual denoising, while a smaller푇en- forces a coarse, abrupt reversal of noise. On the Sports dataset, the best performance is obtained at푇=300. Differently, on Baby and Cell, GTC performs best with larger푇, achieving the highest scores at푇=600 and푇=700, respectively, indicating that these datasets benefit from finer-grained discrimination than Sports. Given the different scales and sparsity levels of user–item interactions, this pattern suggests that when interactions are relatively scarce or highly sparse (as in Baby and Cell), a more fine-grained denoising schedule enables the model to better exploit cross-modal consistency, leading to improved performance. In all three datasets, performance rises as푇 increases, and subsequently declines. This trend confirms that mod- erate noise injection is beneficial for aligning modality-specific signals, whereas overly strong noise harms meaningful content information. We then perform a grid search over0.1,0.2, . . .,1.0to tune the weight parameter휔 1 , which controls the relative contribution of interaction-guided denoising in the overall objective. Larger values of휔 1 increase the emphasis on noise removal when integrating multimodal embeddings. Specifically, different domains domains differ in their noise levels and signal structure, they require differ- ent degrees of denoising regularization to strike a balance between retaining useful information and suppressing spurious patterns. For the Cell dataset, the best performance is obtained with a relatively small휔 1 , while Sports and Baby perform better with stronger de- noising. Specifically, Fig. 7 (line 2) shows that GTC achieves its peak performance at휔 1 =0.6 on Sports,휔 1 =0.7 on Baby, and 휔 1 =0.3 on Cell. It indicates that different domains exhibit varying degrees of cross-modal alignment, and thus requiring different levels of emphasis on denoising process. We further tune the weight휔 2 , which controls the strength of the total correlation objective. Line 3 of Fig. 7 line 3shows that changing 휔 2 noticeably affects the evaluation metrics, indicating the pres- ence of higher-order inter-modal correlations in all three datasets. Although the response curves differ, each dataset benefits from the total correlation term when휔 2 reaches an appropriate range. On Sports, performance exhibits volatile but generally improving trends as increases from 0.1 to 0.6, with the highest Recall@20 achieved at휔 2 =0.6, after which performance declines as휔 2 grows further. A similar pattern is observed on Cell, where both Recall@20 and NDCG@20 peak at휔 2 =0.5 and then decreases. Differently, on Baby, the evaluation metrics are maximized at휔 2 =0.1 and then decrease with fluctuations as휔 2 increases. These results suggest that a moderate emphasis on total correlation helps extract high-order cross-modal signals, whereas an overlarge weight may interfere with other learning objectives, leading to suboptimal performance. 5 Conclusion We introduced GTC to overcome two critical limitations in MMR. To address the universal feature relevance assumption, GTC employs User-Aware Conditional Generative Total Correlation Learning for Multi-Modal RecommendationConference’17, July 2017, Washington, DC, USA 020406080100120140 Epoch 0.25 0.50 0.75 1.00 1.25 1.50 1.75 2.00 Dot Product Score Dot Product Trends of GTC in Sports User-Interaction User-Image User-Text 0102030405060 Epoch 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 Dot Product Score Dot Product Trends of GTC in Sports User-Interaction User-Image User-Text 0255075100125150175 Epoch 0.06 0.08 0.10 0.12 0.14 0.16 0.18 Dot Product Score Dot Product Trends of GTC in Sports User-Interaction User-Image User-Text 0510152025 Epoch 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 Dot Product Score Dot Product Trends of GTC in Baby User-Interaction User-Image User-Text 01020304050 Epoch 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Dot Product Score Dot Product Trends of GTC in Baby User-Interaction User-Image User-Text 05101520253035 Epoch 0.06 0.08 0.10 0.12 0.14 0.16 0.18 0.20 Dot Product Score Dot Product Trends of GTC in Baby User-Interaction User-Image User-Text 010203040506070 Epoch 0.2 0.4 0.6 0.8 1.0 1.2 Dot Product Score Dot Product Trends of GTC in Cell User-Interaction User-Image User-Text 010203040506070 Epoch 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Dot Product Score Dot Product Trends of GTC in Cell User-Interaction User-Image User-Text 020406080100 Epoch 0.08 0.10 0.12 0.14 0.16 Dot Product Score Dot Product Trends of GTC in Cell User-Interaction User-Image User-Text Figure 6: User preference consistency in the Sports dataset (up) and Baby dataset (down). user-conditional filtering via an interaction-guided diffusion model, which denoises content representations to align with individual user preferences. To achieve a more complete alignment, GTC max- imizes a contrastive lower bound of total correlation across all modalities, capturing the higher-order dependencies that previous studies ignore. This way, GTC produces item representations that are both personalized and coherently aligned. We hope the prin- ciples of user-conditional generation and holistic alignment will open avenues for future MMR research. Conference’17, July 2017, Washington, DC, USAJing Du et al. 1004007001000 noisesteps 0.210 0.212 0.214 0.216 0.218 0.220 0.222 Recall@20 Sports Dataset 1004007001000 noisesteps 0.0985 0.0990 0.0995 0.1000 0.1005 0.1010 Recall@20 Baby Dataset 1004007001000 noisesteps 0.16450 0.16475 0.16500 0.16525 0.16550 0.16575 0.16600 0.16625 Recall@20 Cell Dataset 0.122 0.124 0.126 0.128 0.130 NDCG@20 0.0432 0.0434 0.0436 0.0438 0.0440 0.0442 0.0444 0.0446 NDCG@20 0.0768 0.0770 0.0772 0.0774 0.0776 0.0778 NDCG@20 Recall@20 (Left)NDCG@20 (Right) 0.10.20.30.40.50.60.70.80.91.0 w 1 0.221 0.222 0.223 0.224 0.225 0.226 0.227 0.228 Recall@20 Sport Dataset 0.10.20.30.40.50.60.70.80.91.0 w 1 0.096 0.097 0.098 0.099 0.100 0.101 Recall@20 Baby Dataset 0.10.20.30.40.50.60.70.80.91.0 w 1 0.16500 0.16525 0.16550 0.16575 0.16600 0.16625 0.16650 0.16675 0.16700 Recall@20 Cell Dataset 0.120 0.121 0.122 0.123 0.124 NDCG@20 0.04175 0.04200 0.04225 0.04250 0.04275 0.04300 0.04325 0.04350 0.04375 NDCG@20 0.0770 0.0772 0.0774 0.0776 0.0778 0.0780 0.0782 0.0784 NDCG@20 Recall@20 (L)NDCG@20 (R) 0.10.20.30.40.50.60.70.80.91.0 w 2 0.222 0.223 0.224 0.225 0.226 0.227 0.228 Recall@20 Sport Dataset 0.10.20.30.40.50.60.70.80.91.0 w 2 0.0985 0.0990 0.0995 0.1000 0.1005 0.1010 0.1015 Recall@20 Baby Dataset 0.10.20.30.40.50.60.70.80.91.0 w 2 0.16475 0.16500 0.16525 0.16550 0.16575 0.16600 0.16625 0.16650 Recall@20 Cell Dataset 0.1205 0.1210 0.1215 0.1220 0.1225 0.1230 0.1235 0.1240 NDCG@20 0.0432 0.0434 0.0436 0.0438 0.0440 0.0442 0.0444 0.0446 NDCG@20 0.0750 0.0755 0.0760 0.0765 0.0770 0.0775 0.0780 NDCG@20 Recall@20 (L)NDCG@20 (R) Figure 7: Parameter evaluation in Sports (left), Baby (middle), and Cell (right) datasets. User-Aware Conditional Generative Total Correlation Learning for Multi-Modal RecommendationConference’17, July 2017, Washington, DC, USA References [1]Guojia An, Jie Zou, Jiwei Wei, Chaoning Zhang, Fuming Sun, and Yang Yang. 2025. Beyond whole dialogue modeling: Contextual disentanglement for conversational recommendation. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval. 31–41. [2]Ruichu Cai, Zhifan Jiang, Kaitao Zheng, Zijian Li, Weilin Chen, Xuexin Chen, Yifan Shen, Guangyi Chen, Zhifeng Hao, and Kun Zhang. 2025. Learning disen- tangled representation for multi-modal time-series sensing signals. In Proceedings of the ACM on Web Conference 2025. 3247–3266. [3]Jiangxia Cao, Xixun Lin, Xin Cong, Jing Ya, Tingwen Liu, and Bin Wang. 2022. Disencdr: Learning disentangled representations for cross-domain recommenda- tion. In Proceedings of the 45th International ACM SIGIR conference on research and development in information retrieval. 267–277. [4]Xianshuai Cao, Yuliang Shi, Jihu Wang, Han Yu, Xinjun Wang, and Zhongmin Yan. 2022. Cross-modal knowledge graph contrastive learning for machine learning method recommendation. In Proceedings of the 30th ACM international conference on multimedia. 3694–3702. [5]Jing Du, Zesheng Ye, Bin Guo, Zhiwen Yu, and Lina Yao. 2023. Distributional domain-invariant preference matching for cross-domain recommendation. In 2023 IEEE International Conference on Data Mining (ICDM). IEEE, 81–90. [6]Zhiqiang Guo, Jianjun Li, Guohui Li, Chaoyang Wang, Si Shi, and Bin Ruan. 2024. LGMRec: local and global graph learning for multimodal recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence. 8454–8462. [7]Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition. 770–778. [8] Xiangnan He, Kuan Deng, Xiang Wang, Yan Li, Yongdong Zhang, and Meng Wang. 2020. Lightgcn: Simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval. 639–648. [9]Jonathan Ho, Ajay Jain, and Pieter Abbeel. 2020. Denoising diffusion probabilistic models. Advances in neural information processing systems 33 (2020), 6840–6851. [10]Taeri Kim, Yeon-Chang Lee, Kijung Shin, and Sang-Wook Kim. 2022. MARIO: modality-aware attention and modality-preserving decoders for multimedia recommendation. In Proceedings of the 31st ACM international conference on information & knowledge management. 993–1002. [11]Xixun Lin, Rui Liu, Yanan Cao, Lixin Zou, Qian Li, Yongxuan Wu, Yang Liu, Dawei Yin, and Guandong Xu. 2025. Contrastive Modality-Disentangled Learning for Multimodal Recommendation. ACM Transactions on Information Systems (2025). [12]Qidong Liu, Jiaxi Hu, Yutian Xiao, Xiangyu Zhao, Jingtong Gao, Wanyu Wang, Qing Li, and Jiliang Tang. 2024. Multimodal recommender systems: A survey. Comput. Surveys 57, 2 (2024), 1–17. [13] Zhuang Liu, Yunpu Ma, Matthias Schubert, Yuanxin Ouyang, and Zhang Xiong. 2022. Multi-modal contrastive pre-training for recommendation. In Proceedings of the 2022 International Conference on Multimedia Retrieval. 99–108. [14]Jianxin Ma, Chang Zhou, Peng Cui, Hongxia Yang, and Wenwu Zhu. 2019. Learn- ing disentangled representations for recommendation. Advances in neural infor- mation processing systems 32 (2019). [15]Rongqing Kenneth Ong and Andy WH Khong. 2024. Spectrum-based Modality Representation Fusion Graph Convolutional Network for Multimodal Recom- mendation. arXiv preprint arXiv:2412.14978 (2024). [16] Aaron van den Oord, Yazhe Li, and Oriol Vinyals. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748 (2018). [17] Nils Reimers and Iryna Gurevych. 2019. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084 (2019). [18]Steffen Rendle, Christoph Freudenthaler, Zeno Gantner, and Lars Schmidt-Thieme. 2012. BPR: Bayesian personalized ranking from implicit feedback. arXiv preprint arXiv:1205.2618 (2012). [19]Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015. U-net: Convolu- tional networks for biomedical image segmentation. In Medical image computing and computer-assisted intervention–MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part I 18. Springer, 234–241. [20] Adriel Saporta, Aahlad Manas Puli, Mark Goldstein, and Rajesh Ranganath. 2025. Contrasting with Symile: Simple Model-Agnostic Representation Learning for Unlimited Modalities. Advances in Neural Information Processing Systems 37 (2025), 56919–56957. [21]Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang, Lifang Yang, Xianglin Huang, and Tat-Seng Chua. 2022. Self-supervised learning for multimedia recommenda- tion. IEEE Transactions on Multimedia 25 (2022), 5107–5116. [22]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017). [23]Qifan Wang, Yinwei Wei, Jianhua Yin, Jianlong Wu, Xuemeng Song, and Liqiang Nie. 2021. Dualgnn: Dual graph neural network for multimedia recommendation. IEEE Transactions on Multimedia 25 (2021), 1074–1084. [24]Xin Wang, Hong Chen, Yuwei Zhou, Jianxin Ma, and Wenwu Zhu. 2022. Dis- entangled representation learning for recommendation. IEEE Transactions on Pattern Analysis and Machine Intelligence 45, 1 (2022), 408–424. [25] Satosi Watanabe. 1960. Information theoretical analysis of multivariate correla- tion. IBM Journal of research and development 4, 1 (1960), 66–82. [26] Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, and Tat-Seng Chua. 2020. Graph-refined convolutional network for multimedia recommendation with implicit feedback. In Proceedings of the 28th ACM international conference on multimedia. 3541–3549. [27]Yinwei Wei, Xiang Wang, Liqiang Nie, Xiangnan He, Richang Hong, and Tat-Seng Chua. 2019. MMGCN: Multi-modal graph convolution network for personalized recommendation of micro-video. In Proceedings of the 27th ACM international conference on multimedia. 1437–1445. [28]Xiaolong Xu, Hongsheng Dong, Lianyong Qi, Xuyun Zhang, Haolong Xiang, Xiaoyu Xia, Yanwei Xu, and Wanchun Dou. 2024. Cmclrec: Cross-modal con- trastive learning for user cold-start sequential recommendation. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1589–1598. [29]Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Runsheng Xu, Yue Zhao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. 2023. Diffusion models: A comprehensive survey of methods and applications. Comput. Surveys 56, 4 (2023), 1–39. [30]Penghang Yu, Zhiyi Tan, Guanming Lu, and Bing-Kun Bao. 2023. Multi-view graph convolutional network for multimedia recommendation. In Proceedings of the 31st ACM international conference on multimedia. 6576–6585. [31]Penghang Yu, Zhiyi Tan, Guanming Lu, and Bing-Kun Bao. 2025. Mind individ- ual information! principal graph learning for multimedia recommendation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 13096–13105. [32]Jinghao Zhang, Yanqiao Zhu, Qiang Liu, Shu Wu, Shuhui Wang, and Liang Wang. 2021. Mining latent structures for multimedia recommendation. In Proceedings of the 29th ACM international conference on multimedia. 3872–3880. [33] Yin Zhang, Ziwei Zhu, Yun He, and James Caverlee. 2020. Content-collaborative disentanglement representation learning for enhanced recommendation. In Pro- ceedings of the 14th ACM Conference on Recommender Systems. 43–52. [34]Hongyu Zhou, Xin Zhou, Lingzi Zhang, and Zhiqi Shen. 2023. Enhancing dyadic relations with homogeneous graphs for multimodal recommendation. In ECAI 2023. IOS Press, 3123–3130. [35]Xin Zhou and Zhiqi Shen. 2023. A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation. In Proceedings of the 31st ACM International Conference on Multimedia. 935–943. [36] Xin Zhou, Hongyu Zhou, Yong Liu, Zhiwei Zeng, Chunyan Miao, Pengwei Wang, Yuan You, and Feijun Jiang. 2023. Bootstrap Latent Representations for Multi-Modal Recommendation. In Proceedings of the ACM Web Conference 2023. 845–854.