Paper deep dive
Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM Recommenders
Dojun Hwang, Seunghan Lee, Cheonyoung Park, Sara Yu, SeongKu Kang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/24/2026, 5:20:34 AM
Summary
The paper introduces CAIRO, a user context-aware item profiling framework designed to improve LLM-based reranking in recommender systems. CAIRO addresses challenges related to vast, unstructured item metadata by structuring it into objective features and subjective traits. It employs a lightweight profiler to select relevant information for specific user-item pairs, ensuring concise, context-specific profiles that enhance ranking performance.
Entities (9)
Relation Signals (9)
SeongKu Kang → affiliatedwith → Korea University
confidence 95% · SeongKu Kang Affiliation: Korea University
Dojun Hwang → affiliatedwith → Korea University
confidence 95% · Dojun Hwang Affiliation: Korea University
Seunghan Lee → affiliatedwith → Korea University
confidence 95% · Seunghan Lee Affiliation: Korea University
Cheonyoung Park → affiliatedwith → KT Corporation
confidence 95% · Cheonyoung Park Affiliation: KT Corporation
Sara Yu → affiliatedwith → KT Corporation
confidence 95% · Sara Yu Affiliation: KT Corporation
CAIRO → improves → LLM-based reranking
confidence 95% · Experiments show that CAIRO consistently improves LLM-based reranking
CAIRO → uses → lightweight profiler
confidence 92% · employs a lightweight profiler to select the most relevant information for each user-item pair
CAIRO → structures → raw metadata
confidence 90% · CAIRO first structures raw metadata and reviews into objective features and subjective traits
CAIRO → addresses → context-dependent feature salience
confidence 88% · feature salience is highly context-dependent... CAIRO... selects the most relevant information for each user-item pair
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains challenging. Real-world items are described by vast, heterogeneous, and unstructured metadata, where decision-relevant signals are often implicit, noisy, or buried in long descriptions. Moreover, feature salience is highly context-dependent, varying not only across items but also across users. Existing methods often rely on item titles, fixed attributes, or static item summaries, which limit personalized and fine-grained item understanding. To bridge this gap, we propose CAIRO, a user context-aware item profiling framework for LLM-based reranking. CAIRO first structures raw metadata and reviews into objective features and subjective traits, and employs a lightweight profiler to select the most relevant information for each user-item pair with limited serving-time overhead. The resulting profiles are concise and context-specific, providing relevant item-side evidence for the LLM's ranking decision. Experiments show that CAIRO consistently improves LLM-based reranking, highlighting the importance of item profiling that effectively exploits vast item-side information.
Tags
Links
- Source: https://arxiv.org/abs/2608.20801v1
- Canonical: https://arxiv.org/abs/2608.20801v1
Trouble viewing inline? Open PDF directly →
Full Text
80,632 characters extracted from source content.
Expand or collapse full text
Profiling What Matters: Context-Aware Item Profiles from Large-Scale Metadata for LLM RecommendersConference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 07–11, 2026, Rome, ItalyDOI: 10.1145/3799682.3840943ISBN: 979-8-4007-2539-5/2026/11CCS: Information systems Recommender systems Dojun Hwang Affiliation: Korea University , Seoul , Republic of Korea email: dojun2006@korea.ac.kr , Seunghan Lee Affiliation: Korea University , Seoul , Republic of Korea email: seunghanlee@korea.ac.kr , Cheonyoung Park Affiliation: KT Corporation , Seoul , Republic of Korea email: park.cheonyoung@kt.com , Sara Yu Affiliation: KT Corporation , Seoul , Republic of Korea email: sara.yu@kt.com and SeongKu Kang Affiliation: Korea University , Seoul , Republic of Korea Note: Corresponding author. email: seongkukang@korea.ac.kr 2026; © c Abstract. While Large Language Models (LLMs) have significantly advanced reranking in recommendation, effectively leveraging item-side information remains challenging. Real-world items are described by vast, heterogeneous, and unstructured metadata, where decision-relevant signals are often implicit, noisy, or buried in long descriptions. Moreover, feature salience is highly context-dependent, varying not only across items but also across users. Existing methods often rely on item titles, fixed attributes, or static item summaries, which limit personalized and fine-grained item understanding. To bridge this gap, we propose , a user context-aware item profiling framework for LLM-based reranking. structures raw metadata and reviews into objective features and subjective traits, and employs a lightweight profiler to select the most relevant information for each user–item pair with limited serving-time overhead. The resulting profiles are concise and context-specific, providing relevant item-side evidence for the LLM’s ranking decision. Experiments show that improves LLM-based reranking, highlighting the importance of item profiling that effectively exploits vast item-side information. Our implementation is available at: . Keywords: Recommender systems, Large-scale Metadata, Item Profiling †c-license: by 1. Introduction Figure 1. A conceptual comparison of item description exploitation between previous approaches and our approach. Recently, recommender systems have been embracing a paradigm shift driven by Large Language Models (LLMs). Leveraging extensive world knowledge and reasoning capabilities, LLMs have successfully addressed the data sparsity issue of traditional ID-based methods by generating rich semantic representations of users and items (55). Within standard multi-stage ranking pipelines (32), LLMs are widely adopted in the final reranking stage (52; 2; 11; 39). Provided with a user’s historical interactions and a set of candidate items as input, LLMs perform fine-grained relevance estimation to re-rank candidate items, improving recommendation quality beyond traditional methods (16; 8; 54; 30). To provide richer information to the LLM, recent efforts have focused on user profiling, which generates representative and descriptive textual profiles of users. As raw interaction data is inherently noisy and only indirectly reveals user preferences, merely injecting such data into prompts often yields suboptimal results (56). Instead, user profiling consolidates historical signals into structured, high-level representations that more faithfully capture underlying preferences. Specifically, LLMs leverage their reasoning capabilities to infer user preferences from past interactions (6; 46; 21), reviews (19; 12), and demographic features (11; 48). These profiles articulate latent preferences not explicitly observable from raw interactions alone, such as specific likes and dislikes (19), behavioral habits (56), and long- and short-term interests (3). Prior work shows that integrating these user profiles into input prompts facilitates a deeper understanding of users and improves LLM-based reranking (57; 42). Despite progress on the user-side, comparatively little effort has been devoted to item-side profiling. In real-world applications, each item is associated with numerous attributes, from tens to even hundreds, spanning diverse aspects, from functionality (e.g., battery size, WiFi version) to aesthetics (e.g., color, material). As evidenced by prior work on feature-based recommendation (26; 7; 28), a deep understanding of item content and fine-grained user-item feature interactions is critical. However, LLMs struggle to utilize such item-side information due to three key challenges: (C1) Vast unstructured data overwhelm LLM context window: Item-side information is vast in scale and often inconsistently formatted across providers, appearing in diverse and non-standardized forms (e.g., bullet lists, tables, or free-form text). Moreover, many features are spread across long textual descriptions rather than being clearly structured. LLMs operate under strict input context limits, and struggle to process information within long contexts (31). Therefore, simply including all item data not only wastes the limited context budget but can also lead to suboptimal results. (C2) Important features differ across items: The importance of item features is highly item-dependent. For example, think of the refresh rate of monitors. Most monitors share common specifications such as a 60Hz refresh rate, and therefore it is not a notable feature for most items in the category. However, it becomes a distinguishing factor for a specific subset, like gaming monitors at 360Hz. Therefore, prompts that include the same set of attributes for all items can obscure truly discriminative attributes with generic ones, making it harder for LLMs to judge fine-grained relevance. (C3) Even for a single item, important features vary by user: Even among the salient features of a single item, different users may find different features appealing. This user-dependent variation arises from differences in their preferences and usage contexts. For example, consider Apple’s AirPods Max: users with strong brand loyalty to Apple may value its seamless integration within the Apple ecosystem, whereas some may prioritize functional aspects such as high-quality noise canceling. Given that feature importance varies across users, ignoring such variation can lead to important features for a given user being underrepresented, making it harder for LLMs to accurately assess personalized item relevance. Still, existing LLM-based reranking approaches only partially address these challenges. Many methods rely on item titles (41; 3; 14) or a small set of manually selected attributes (13; 11; 6; 35; 25; 42; 56; 40). They not only provide an incomplete understanding of items but also fail to scale to vast metadata, as they merely inject the given features, leaving (C1) unresolved. Another line of work constructs item-specific profile by summarizing item information into concise forms, such as keywords (51), motivational traits (5), and review-based liked/disliked points (19; 9). While scalable, these approaches still rely on a fixed summarization scheme (e.g., extracting ten keywords of each item (51)) applied to a small set of manually selected attributes. This limits their ability to identify item-specific salient features from vast metadata and reviews, leaving (C2) insufficiently addressed. More importantly, as they generate a static description of an item regardless of the user, they fail to capture user-dependent variations in feature salience, leaving (C3) underexplored. As a solution, we propose , a user Context-Aware Item pROfiling framework for LLM-based reranking. transforms large-scale, heterogeneous item metadata into a unified structured representation by extracting domain-specific objective features and multi-faceted subjective traits from metadata and reviews. It then builds a lightweight item profiler that, given the current user’s information, selects the objective and subjective information relevant to that user’s decision-making context. In particular, rather than relying on fixed heuristics, the profiler learns to quantify the importance of features from collaborative signals, allowing it to identify user-relevant information from vast metadata in an efficient manner. In this way, relying on fixed manually selected attributes or static item summaries, while enabling scalable and user-adaptive item profiling for LLM reranking, resolving all the aforementioned challenges. Furthermore, a dedicated refinement process to mitigate potential deficiencies in LLM-inferred information and further improve the quality of item profiling. Our main contributions are as follows: • We identify an underexplored challenge in LLM-based reranking: effectively leveraging vast and heterogeneous item metadata while capturing item- and user-dependent feature salience. • We propose , a novel item profiling framework that transforms raw metadata into objective features and subjective traits, refines them, and selectively constructs user-specific item profiles through lightweight selectors. • Extensive experiments show that our framework consistently improves LLM-based reranking, highlighting the value of scalable and user-adaptive item profiling. 2. Related Work LLM-based recommendation. Equipped with vast world knowledge and reasoning capability, LLMs are widely adopted in recommendation. For example, LLMRank (16) shows that LLMs, even when given only item titles, can effectively rank candidate items by leveraging their parametric knowledge. To exploit LLMs as recommenders, many methods (1; 13; 53; 51) fine-tune LLM for specific downstream tasks. Although fine-tuning can align an LLM’s knowledge and the reasoning process with a target recommendation task, it requires substantial computational resources and risks compromising the model’s generalization ability (55). The prompt augmentation paradigm instead incorporates auxiliary information into the input prompts, thereby preserving the generalization ability of pretrained LLMs. To enrich prompts, prior studies have explored diverse sources of auxiliary information, with each framework adopting different forms of augmentation. For the user-side, prior work has incorporated demographic information (11), user histories (56; 6), and reviews (35; 19). For the item-side, prior studies have explored selected attribute fields, such as category (56), brand and category (40), and genre (42), as well as graph-derived knowledge (44). This paradigm enables pretrained LLMs to generalize across diverse recommendation domains without costly fine-tuning, and our work also follows this direction. User and item profiling in recommendation. To provide rich information in prompts, textual compression of useful information, referred to as profiling, has been widely studied on the user side (35; 56; 6; 3; 42). Early studies (56; 6) summarize user histories into a static profile, while later work explores more adaptive forms of profiling: (35) allows users to update their profiles explicitly, (3) updates profiles to reflect evolving preferences, and (42) studies profile formats tailored to downstream tasks. These efforts help LLMs better understand users and improve recommendation accuracy. By contrast, item-side profiling in LLM-based recommendation has received comparatively less attention. Most previous studies resort to a fixed set of item attributes (6; 56; 42), or construct item profiles using simple keywords summarized from descriptions (51; 49) or reviews (9). Among recent efforts to advance item profiling, the most closely related are EXP3RT (19) and M-LLM3Rec (5). In addition to a few selected attributes, EXP3RT extracts preference information from reviews and aggregates them per item, to capture subjective perspectives. Similarly, M-LLM3Rec uses phrases containing functional purposes of an item as item profiles. Although these methods are state-of-the-art item profiling methods that effectively capture valuable item information, they have limited capability in identifying salient features of each item from vast metadata, and they use a static item profile for all users, making LLMs struggle to judge personalized item relevance. In this work, we propose an item profiling strategy that effectively exploits vast item metadata and selectively includes item information relevant to recommendation contexts. 3. Problem Formulation Notations. Let U and ℐI denote the set of users and items in a recommendation dataset, respectively. Each user u∈u has an interaction history ℋu=r1,…,r|ℋu|H_u=\r_1,…,r_|H_u|\. Each item i∈ℐi is associated with metadata ℳiM_i and a set of user reviews ℛiR_i. In many real-world applications, ℳiM_i consists of a large collection of unprocessed information collected from various sources. It is vast in scale, often containing tens to hundreds of attributes spanning diverse aspects (10; 24). Moreover, item metadata is often inconsistently formatted across providers, appearing in non-standardized forms such as bullet lists, tables, or free-form text. These attributes span multiple types, including numerical (e.g., price, ratings), categorical (e.g., brand, color), and ordinal values (e.g., size levels, quality tiers). Furthermore, many important signals are not explicitly structured but are buried in long textual descriptions (Figure 1). Recommendation with LLM-based reranking. We adopt the standard two-stage recommendation (32; 17), consisting of candidate generation followed by reranking. Let u be a user with history ℋuH_u, and uC_u be the set of candidate items for u retrieved from the full item set using a lightweight model. LLMs rerank items in uC_u based on their relevance to user u, generating a reranked list u∗C_u^*: (1) u∗=LLM([Pi∣i∈ℋu],Pi∣i∈u,Pu,prerank),C_u^*=LLM([P_i i _u],\P_i i _u\,P_u;p_rerank), where prerankp_rerank is a prompt for the reranking task. PuP_u and PiP_i denote textual profiles of user and item, respectively. As aforementioned, while many efforts have focused on PuP_u, constructing PiP_i that fully leverages ℳiM_i remains underexplored. Problem definition. Our goal is to develop an item profiling function ϕφ that constructs Pi|u=ϕ(ℳi,u)P_i|u=φ(M_i,u), where Pi|uP_i|u is a compact yet informative profile of item i specialized for user u. We aim to produce profiles that address the aforementioned challenges: (C1) being structured for limited context window, (C2) capturing item-specific salient features, and (C3) adapting to user context. 4. Proposed framework: Figure 2. Overview of the . We propose , which constructs structured and context-specific item profiles for LLM-based reranking. We first organize vast item-side knowledge into structured representations (section 4.1), and then build a lightweight profiler that selects context-relevant information (section 4.2). Using this profiler, context-specific item profiles online with relatively negligible latency (section 4.3). We further introduce an optional refinement stage to further enhance the recommendation quality (section 4.4). 4.1. Structuring Item Information Raw item metadata ℳii∈ℐ\M_i\_i comprise large-scale, heterogeneous information from diverse sources, and item reviews ℛii∈ℐ\R_i\_i , written by diverse users, vary widely in writing styles and expressions, making both highly unstructured and difficult to directly utilize. As a first step, we organize item knowledge into a structured form, represented as a dictionary, facilitating systematic usage and compatibility with LLM prompting in the later stages. Specifically, we consider item knowledge to be two-fold: objective and subjective. Objective knowledge refers to factual information that is static across users and is primarily derived from item metadata, such as Weight: 3kg. In contrast, subjective knowledge captures aspects that can be perceived differently across users and is mainly inferred from user reviews, such as Brand loyalty: The brand has strong heritage.... To represent these complementary aspects, we construct two types of records for each item, objective features and subjective traits. 4.1.1. Domain-specific Key Exploration To extract objective features from metadata, we first define feature fields that serve as keys in the structured dictionary. These keys should be tailored to the domain, capturing essential aspects in a consistent manner. Here, the challenges are two-fold: (i) feature fields in raw metadata are often inconsistent, with semantically identical attributes expressed under different names (e.g., DLC vs. Additional Content); (i) important attributes are often not explicitly provided but are instead embedded in textual descriptions. For example, in the video game domain, features such as co-op support are often mentioned only in text descriptions. Therefore, we first identify a set of representative keys for the domain, referred to as domain-specific keys, and then organize scattered item information based on them. A straightforward approach is to prompt an LLM with the domain name and directly generate keys. However, such an approach may fail to cover all subcategories within the domain and is prone to hallucination. Instead, we adopt a data-driven strategy that lets the LLM extract keys grounded in the actual metadata. To ensure balanced coverage of diverse subcategories within a domain, we first identify clusters of similar items, generate candidate keys for each cluster, and then merge them into a unified set. We apply k-means clustering to the embeddings of all item metadata.11 1 We linearize the raw metadata of each item and encode it using BGE-M3 (4). For each cluster ℐjI_j, we randomly sample a small set of items ℐ~j⊆ℐj I_j _j and let LLM identify representative keys for each cluster: (2) ℱj=LLM(ℳi∣i∈ℐ~j;pexplore-keys),j=1,2,…,kitems,F_j=LLM(\M_i i∈ I_j\;p_explore-keys), j=1,2,…,k_items, where pexplore-keysp_explore-keys is the prompt for key exploration. This yields kitemsk_items sets of candidate keys, each reflecting cluster-specific characteristics. We set kitems=20k_items=20. We then reassess and consolidate the candidate keys to build a unified set. Using its reasoning capability, the LLM evaluates whether each key is relevant to user decision-making in the domain, while merging semantically similar keys across clusters (e.g., Water-proof vs. Water-resistance), as: (3) ℱ=LLM(ℱj∣j=1,2,…,kitems;preassess-&-merge),F=LLM(\F_j j=1,2,…,k_items\;p_reassess-\&-merge), where preassess-&-mergep_reassess-\&-merge is the corresponding prompt. Note that we do not apply frequency-based filtering, so as to preserve sparse but informative keys and ensure broad coverage within the domain. The resulting set ℱF defines the final domain-specific keys that comprehensively cover item characteristics within the domain. In our experiments, |ℱ||F| ranges from 30 to 50 depending on the dataset. 4.1.2. Objective Item Features Extraction Based on the domain-specific keys ℱF, we organize the metadata of all items under a unified schema. Specifically, for each item i, we prompt the LLM to extract the value of each key in ℱF, producing a structured dictionary of objective item features: (4) i=LLM(ℳi,ℱ,pextract-obj),i∈ℐ,O_i=LLM(M_i,F;p_extract-obj), i , where i=(f:vf)∣f∈ℱO_i=\(f:v_f) f \ denotes the resulting feature dictionary. iO_i includes not only attributes that are already structured in raw metadata (e.g., price), but also diverse attributes embedded in textual descriptions. To prevent hallucination, we design the reasoning process such that the LLM first verifies whether supporting evidence for each key can be found in the metadata, and assigns a value only when such evidence is found; otherwise, it outputs null. 4.1.3. Subjective Item Traits Extraction Subjective traits capture aspects of an item that may be perceived differently across users. We use user reviews as the primary source of such traits, as they reflect subjective experiences, preferences, and reasons for purchase. Inspired by prior work that extracts aspect-level signals from reviews (e.g., keywords (33), liked/disliked points (19)), we first leverage the LLM to convert reviews written in varied styles into a unified schema. Specifically, for each review r, the LLM produces a structured representation ara_r consisting of four aspects: purpose of purchase, usage context, degree of satisfaction, and category preference. As in objective feature extraction, if a particular aspect cannot be identified from the review, the LLM outputs null. Now, item traits can be constructed by summarizing the extracted aspects for item i’s reviews: ar∣r∈ℛi\a_r r _i\. However, we argue that a naïve single summary is insufficient to capture the diverse subjective perspectives, as users may value different facets, such as functionalities, aesthetic design, or brand affinity. We therefore model subjective traits as a multi-faceted set, where each trait captures a distinct perspective commonly reflected in user reviews. To this end, we design the reasoning process so that the LLM infers multiple traits by reasoning over the extracted aspects, while grounding them in the item’s objective features: (5) i=LLM(ar∣r∈Ri,i,pextract-sub),i∈ℐS_i=LLM(\a_r r∈ R_i\,O_i\,;\,p_extract-sub), i The prompt pextract-subp_extract-sub constrains the number of generated traits |i||S_i| to 3–7 to avoid collapsing into a single trait or generating redundant ones. In addition, for each trait, we identify a small set of supporting features from iO_i that directly support the trait. For example, a trait related to ‘Vivid Color’ can be supported by the feature color: red. In practice, each trait is associated with about two objective features on average. The resulting iS_i is a dictionary, i.e., i=tim:(dim,im)m=1|i|S_i=\t_i_m:(d_i_m,O_i_m)\_m=1^|S_i|, where timt_i_m denotes the trait title, dimd_i_m its description, and imO_i_m its supporting features. Figure 2 shows a video game whose subjective traits are multi-faceted: ‘Character-driven Engagement’ reflects experiences anchored in character development, while ‘Long-term Loyalty’ captures the appeal to long-time franchise fans. Remarks on potential deviations. Despite the explicit grounding on reviews and factual item information, the generated subjective traits may deviate from true user intents. This is because, unlike objective features obtained by directly identifying corresponding values, subjective traits are extracted through the LLM’s reasoning. To cope with such potential deviations, introduces an auxiliary refinement process for subjective traits (Section 4.4). Figure 3. Overview of the feature-grounded trait refinement stage (Section 4.4). 4.2. User Context-aware Item Profiler We now obtain structured objective and subjective information for each item. However, providing all this information to LLMs is suboptimal, as the aspects most relevant to decision-making vary across items and users. We aim to profile each item using only the information relevant to each user’s context and decision-making. A naïve solution is to let LLMs perform this selection at serving time, but this incurs high latency due to additional online LLM calls. To address this, we introduce a lightweight item profiler that generates a user-specific item profile based on the current user’s information. It consists of two submodules for selecting objective features and subjective traits, respectively. Prepared offline, the profiler enables relatively efficient online selection without additional LLM calls. In this section, we describe the profiling mechanism used to select objective features and subjective traits, and the final online prompt construction is presented in Section 4.3. 4.2.1. User Profile Construction LLM-based rerankers typically take a textual user profile PuP_u as input (Eq. 1). Accordingly, we use it as the main source of user information. As constructing elaborate user profiles is not our main contribution, we largely follow a simple approach from prior work (33; 12), which summarizes user reviews for profiling. Similar to subjective trait extraction, we aggregate review aspects for each user, and prompt the LLM to summarize the user’s preferences and interests: (6) Pu=LLM(ar∣r∈ℛu,pextract-user-profile),u∈,P_u=LLM(\a_r r _u\;p_extract-user-profile), u , where ℛuR_u is a set of reviews written by user u. For the user profiling prompt pextract-user-profilep_extract-user-profile, we follow the prompt used in (12). 4.2.2. Sub-module: Objective Feature Selector For each user u, this sub-module aims to select a small number of features from iO_i. To make such selection effective, it is essential to consider both the user’s context and its high-level relationships with item features. As multiple features may contribute collaboratively rather than independently, it is infeasible to manually define deterministic rules. We need a mechanism that automatically quantifies which features are most relevant to each user’s decision-making. We draw inspiration from a line of research on quantifying feature importance from collaborative signals (28; 45; 23). A wisdom from this literature is that, during training, a model can estimate the importance of each feature based on its contribution to prediction. A backbone recommender. Following this idea, we use a lightweight recommendation model gRS(⋅)g_RS(·), which predicts the interaction score as y^ui=gRS(ui) y_ui=g_RS(e_ui).22 2 We use a 2-layer MLP, following prior work (28; 47). Here, uie_ui denotes the input representation of the user–item pair (u,i)(u,i). By default, uie_ui is formed by encoding the user and item information and concatenating the resulting embeddings. Specifically, we encode each item feature in iO_i as fh_f and user profile PuP_u as Puh_P_u, using encoders corresponding to their modalities.33 3 Specifically, numerical values are binned, categorical values are embedded, and textual values are encoded with a text embedding model (4) followed by a projection layer. Also, we treat the user and item IDs as categorical features with embeddings uh_u and ih_i, respectively. Concatenating them yields: (7) ui=[f1;f2;…;fn;Pu;u;i]∈ℝ(|ℱ|+3)⋅d.e_ui=[h_f_1;h_f_2;…;h_f_n;h_P_u;h_u;h_i] ^(|F|+3)· d. Feature importance learning. We introduce a controller network gw:ℝ(|ℱ|+3)⋅d→ℝ|ℱ|+3g_w:R^(|F|+3)· d ^|F|+3 which estimates the importance of each feature in uie_ui, as: (8) =softmax(gw(ui))=[wf1,wf2,…,wfn,wPu,wu,wi],w=softmax(g_w(e_ui))=[w_f_1,w_f_2,…,w_f_n,w_P_u,w_u,w_i], where each weight indicates the estimated importance of the corresponding feature. A weighted input representation ~ui e_ui is then constructed by multiplying each importance score with its corresponding embedding: (9) ~ui=[wf1f1;wf2f2;…;wfnfn;wPuPu;wuu;wii]. e_ui=[w_f_1h_f_1;w_f_2h_f_2;…;w_f_nh_f_n;w_P_uh_P_u;w_uh_u;w_ih_i]. With the reweighted features, the model makes prediction as: y^ui=gRS(~ui) y_ui=g_RS( e_ui), and is optimized using the binary cross-entropy loss: (10) ℒBCE=−∑(u,i)[yuilogy^ui+(1−yui)log(1−y^ui)].L_BCE=- _(u,i) [y_ui y_ui+(1-y_ui) (1- y_ui) ]. where yui=1y_ui=1 if (u,i)(u,i) is observed, and 00 otherwise. During training, the recommender and controller are optimized jointly, leading the controller to assign high weights to the objective features most predictive for each user–item pair (28). This enables automatic quantification of feature importance from collaborative patterns. Feature selection. For each (u,i)(u,i), we apply a pooling operation to =softmax(gw(ui))w=softmax(g_w(e_ui)) to select the most important features: (11) ℱu,i=Pool(ℱ,)F_u,i=Pool(F,w) Following (28), we employ k-max pooling, which selects the top-k important features. In our experiments, we set k=4k=4. By grounding feature selection in collaboratively learned importance weights, this process can capture complex combinations of features that jointly influence user decisions. 4.2.3. Sub-module: Subjective Trait Selector For each user u, this sub-module aims to select the subjective trait from iS_i that is most aligned with the user’s preferences. As subjective traits are already expressed as elaborated textual descriptions, we select the most aligned trait based on its semantic similarity to the user profile: (12) su,i=argmaxs∈icos(Pu,s),s_u,i= _s _icos(e_P_u,e_s), where Pue_P_u and se_s denote the embeddings of the user profile and subjective trait from the text embedding model, respectively. The selected trait su,is_u,i serves as the subjective aspect of the user-specific profile of item i. 4.3. Online Prompt Construction Using the proposed mechanisms, we construct the final reranking prompt with item profiles specialized to each user’s context. Input: User profile PuP_u, history HuH_u, candidates CuC_u, objective features ii\O_i\_i, subjective traits ii\S_i\_i Output: Final reranking prompt pufinalp_u^final 1 Batch items for parallel processing Bu←Hu∪CuB_u← H_u∪ C_u 2 // Objective feature selection Construct the input matrix u=[ui]i∈BuE_u=[e_ui]_i∈ B_u (Eq. 7) Compute the importance matrix u=softmax(gw(u))W_u=softmax(g_w(E_u)) (Eq. 8) Apply row-wise pooling on uW_u to obtain ℱu,ii∈Bu\F_u,i\_i∈ B_u (Eq. 11) 3 // Subjective trait selection Construct the trait embedding matrices [s]s∈ii∈Bu\[E_s]_s _i\_i∈ B_u Select su,i←argmaxs∈icos(Pu,s)s_u,i← _s _icos(e_P_u,E_s) for each i∈Bui∈ B_u (Eq. 12) Collect the supporting features ℱsu,ii∈Bu←isu,ii∈Bu\F_s_u,i\_i∈ B_u←\O_i_s_u,i\_i∈ B_u 4 // Prompt construction Assemble Pi|ui∈Bu\P_i|u\_i∈ B_u with Pi|u=[su,i,ℱu,i∪ℱsu,i]P_i|u=[\,s_u,i,\ F_u,i _s_u,i\,] pufinal←Prompt(Pi|ui∈Hu,Pi|ui∈Cu,Pu)p_u^final (\P_i|u\_i∈ H_u,\ \P_i|u\_i∈ C_u,\ P_u) 5 return pufinalp_u^final Algorithm 1 Online Prompt Construction by Construction process. Algorithm 1 details the construction process of the reranking prompt pufinalp_u^final of user u. Objective features are selected by gwg_w, implemented as a linear layer followed by pooling (Lines 2--4), while the subjective trait is selected via cosine similarity with the user profile (Lines 5--6). We additionally include the objective features that directly support the selected subjective trait (Line 7).44 4 In practice, each item profile contains around five objective features in total. The resulting profiles are then assembled into the final reranking prompt, while the remaining instructions follow standard LLM-based reranking. Each item is profiled according to the current user’s context, unlike existing approaches that rely on a single static profile for all users. Figure 2 shows an example of user-specific prompt construction. For the first user, the selected subjective trait emphasizes character-driven engagement, and the feature selection module identifies objective features such as platform compatibility and rating as relevant to the user’s decision making based on collaborative patterns. For the second user, the selected trait instead highlights long-term loyalty, while a different set of objective features, including age rating, is selected. This illustrates that both subjective traits and objective features to each user-item context. Faithfulness of the profiles. CAIRO is equipped with multiple mechanisms to ensure profile faithfulness. Objective features are extracted only when supported by metadata, with unsupported fields assigned null. Subjective traits are inferred from review-derived aspects and grounded in objective item features, and the final profile includes supporting features for the selected trait. As a result, all profile information is tightly grounded in the given item information, making the profiles more traceable and less prone to hallucination than free-form item summaries. Online efficiency of . A key requirement at this stage is low latency. To satisfy this, all item-side information and profiling modules offline, so that the online stage involves only a few lightweight operations. In particular, all item-side embeddings are precomputed and stored offline, and item profiling is executed in matrix form through parallel GPU computation. Through this lightweight design, user-specific item profiles with relatively limited additional overhead in real time. A detailed efficiency analysis is provided in Section 5.2.3. 4.4. Feature-grounded Trait Refinement Although the constructed profiles capture core item aspects, there remains room for improvement, particularly for subjective traits. Unlike objective features, which are directly evidenced by explicit values, subjective traits are inferred through the LLM’s reasoning process, leaving room for further optimization by aligning them with preference signals from user–item interactions. To this end, we introduce a trait refinement process that diagnoses weaknesses of subjective traits from interaction-based prediction errors and revises them accordingly. The revision is grounded in objective features to ensure faithfulness. This refinement is conducted offline as an auxiliary step to further improve profile quality. In alignment with recent studies on prompt optimization (29; 34), our refinement process consists of three stages: error-case collection, defect diagnosis, and trait refinement. Error collection from proxy task. To refine each subjective trait, we first collect cases in which the trait fails to support successful interaction prediction. To make this process efficient, we adopt simplified next-item prediction as a proxy task. For each item i, we first retrieve histories associated with each subjective trait. Specifically, from the full interaction histories, we collect subsequences whose last item is i and group them by the corresponding trait. Let ℋi,sH_i,s denote the set of histories associated with trait s of item i. Given H∈ℋi,sH _i,s, we treat i as the ground-truth next item and use the preceding interactions as the user history. We then ask the LLM to predict the most probable next item from a candidate set that includes item i and Nproxy-candN_proxy-cand randomly sampled negative items. We set Nproxy-cand=4N_proxy-cand=4. The re-ranking prompt is constructed as described in the previous section. As i is profiled with trait s, an error case where the LLM fails to select i provides evidence that the current trait may not sufficiently support preference prediction. Defect diagnosis. From the proxy task, we obtain error cases in which the LLM fails to predict the ground-truth next item. Let ℰi,s=H1,H2,…,H|ℰi,s|E_i,s=\H_1,H_2,…,H_|E_i,s|\ denote the set of error cases from ℋi,sH_i,s. We then instruct the LLM to analyze each error case and identify the deficiency of the current trait s. Formally, we obtain the diagnosis DkD_k for each error case HkH_k as follows: (13) Dk=LLM(Hk;panalyze-error),k=1,2,…,|ℰi,s|.D_k=LLM(H_k;p_analyze-error), k=1,2,…,|E_i,s|. The resulting diagnoses Dkk\D_k\_k are collectively leveraged to capture failure patterns that consistently appear across multiple error cases. Figure 3 illustrates an example of this diagnosis step, where the LLM identifies that the current trait overemphasizes general brand reliability while missing performance-related aspects. Objective feature-grounded trait refinement. Using the collected diagnoses, we refine the subjective trait by addressing recurring deficiencies. To prevent hallucination, we ground the reasoning process in the objective features of each item. Specifically, the LLM revises the trait by jointly considering the original trait, the diagnosis set, and the objective features. Following the common practice of sampling multiple LLM outputs and selecting the most suitable one to improve generation quality (43; 27), we generate NcandN_cand candidate revisions: (14) sq′=LLM(s,Dkk=1|ℰi,s|,i;prefine-trait),q=1,2,…,Ncand.s _q=LLM(s,\D_k\_k=1^|E_i,s|,O_i;p_refine-trait), q=1,2,…,N_cand. We set Ncand=2N_cand=2. Finally, among the original and generated traits, we select the trait most aligned with the users associated with ℋi,sH_i,s. Alignment is measured by the cosine similarity between each trait embedding and the average embedding of the corresponding user profiles, and the trait with the highest similarity is selected as the refined trait. Figure 3 shows that the refined trait better captures performance-related aspects aligned with user preferences. Overall, this refinement process directly reflects prediction failures through diagnosis-based revision, while objective feature grounding preserves faithfulness of profiles. 5. Experiments Table 1. Overall performance comparison (*: p-value < .05). For feature-enhanced LLM-based reranking frameworks (+Feat), metrics colored red denote under-performance over LLMRank, which uses only item titles. Methods Video Games Sports and Outdoors Electronics nDCG@5 HR@5 nDCG@10 HR@10 nDCG@5 HR@5 nDCG@10 HR@10 nDCG@5 HR@5 nDCG@10 HR@10 BPR 0.0698 0.1211 0.1252 0.2961 0.0793 0.1305 0.1303 0.3082 0.0540 0.0960 0.1086 0.2619 SASRec 0.1599 0.2205 0.1952 0.3349 0.1658 0.2203 0.1968 0.3172 0.1399 0.1835 0.1650 0.2868 BERT4Rec 0.1488 0.2123 0.1842 0.3226 0.1247 0.1825 0.1613 0.2967 0.0938 0.1372 0.1215 0.2237 xDeepFM 0.0932 0.1679 0.1621 0.3834 0.0897 0.1637 0.1638 0.3959 0.0879 0.1491 0.1497 0.3444 AdaFS 0.1401 0.2140 0.2030 0.4119 0.1247 0.2109 0.1937 0.4273 0.1029 0.1752 0.1762 0.4054 REACTION 0.1404 0.2293 0.2005 0.4185 0.1271 0.2240 0.2035 0.4441 0.1055 0.1901 0.1807 0.4257 LLMRank 0.1626 0.2411 0.2180 0.4154 0.1652 0.2434 0.2131 0.3934 0.1214 0.1779 0.1607 0.3012 LLMRank (+Feat) 0.1122 0.1728 0.1527 0.3026 0.1114 0.1725 0.1513 0.2978 0.1070 0.1631 0.1470 0.2885 EXP3RT 0.1644 0.2414 0.2172 0.4073 0.1796 0.2630 0.2453 0.4690 0.1604 0.2276 0.2152 0.4001 EXP3RT (+Feat) 0.1303 0.2052 0.1851 0.3775 0.1420 0.2134 0.2017 0.4011 0.1109 0.1713 0.1656 0.3444 M-LLM3Rec 0.1674 0.2578 0.2290 0.4505 0.1734 0.2667 0.2409 0.4777 0.1581 0.2422 0.2181 0.4308 M-LLM3Rec (+Feat) 0.1519 0.2373 0.2119 0.4202 0.1676 0.2592 0.2309 0.4574 0.1472 0.2304 0.2069 0.4173 0.1839* (+9.86%) 0.2761* (+7.10%) 0.2435* (+6.33%) 0.4605* (+2.22%) 0.1852* (+3.12%) 0.2793* (+4.72%) 0.2512* (+2.41%) 0.4865* (+1.84%) 0.1741* (+9.98%) 0.2624* (+9.06%) 0.2405* (+10.47%) 0.4706* (+9.09%) + Refine 0.1895* (+12.84%) 0.2827* (+10.40%) 0.2467* (+7.51%) 0.4642* (+3.15%) 0.1981* (+10.30%) 0.2877* (+7.87%) 0.2643* (+7.75%) 0.4959* (+3.81%) 0.1866* (+17.88%) 0.2693* (+11.93%) 0.2540* (+16.67%) 0.4804* (+11.36%) 5.1. Experimental Setup The key instructions of each prompt are presented in the main text and figures, and the full prompts are provided in the code.55 5 5.1.1. Datasets We conduct experiments on Amazon datasets (15), which, to our knowledge, provide the largest sources of vast and heterogeneous item metadata with user reviews, which allows meaningful analysis on item-side information. We select three domains with distinct characteristics: Video Games, Sports and Outdoors, and Electronics. We closely follow the preprocessing of prior studies (19; 5; 52; 1): considering the significant cost of generating profiles with competitive commercial LLMs, we sample 200K–300K interactions from each dataset and apply 5-core filtering on both users and items. To enable meaningful analysis of item-side information, we filter out items with missing titles or metadata. For item metadata, we use both common and optional features. The common features consist of 11 fields observed across all items and datasets, while optional features vary across items and may appear under different provider-specific field names. Among the common features, the features and description fields are written as unstructured sentences. Detailed dataset statistics and the used metadata fields are reported in Table 2. 5.1.2. Baselines We consider a variety of baseline methods. Traditional ranking models. We compare against three methods: ∙ BPR (36) is an embedding-based model with pairwise ranking. ∙ SASRec (18) and BERT4Rec (37) are transformer-based models with unidirectional/bidirectional attention, respectively. Feature-based models. We compare against competitive methods that learn feature importance from vast metadata and incorporate fine-grained user–item feature interactions into recommendation: ∙ xDeepFM (26) explicitly models high-order feature interactions via a cross network with the cross-product operation. ∙ AdaFS (28) learns interaction-specific feature importance and selects a small set of important features for prediction. ∙ REACTION (47) is a state-of-the-art method that jointly considers feature redundancy and parameter efficiency by formulating feature selection from an information-theoretic perspective. LLM-based reranking models. We compare against state-of-the-art LLM-based reranking methods to examine how different profiling strategies affect reranking capability. All methods use GPT-4o-mini, and the core difference lies in how each item is profiled. ∙ LLMRank (16) is a fundamental LLM-based reranking method that uses only item titles, aiming to leverage the LLM’s parametric knowledge for recommendation. ∙ EXP3RT (19) employs an advanced profiling that extracts liked/disliked points from reviews to construct both user and item profiles, while also incorporating item features as additional signals. ∙ M-LLM3LLM^3Rec (5) employs a motivation-oriented profiling to capture user motivations, summarizing item metadata into keyword-level item profiles and constructing user profiles from reviews. While EXP3RT and M-LLM3LLM^3Rec use item metadata, they remain limited to a small set of preprocessed common features (Table 2), as item profiling is not their primary focus. To give them more knowledge in terms of the item-side information, we consider feature-enhanced variants of the LLM-based methods, denoted as +Feat. Here we additionally provide the not-null optional fields in Table 2 not originally used by each method, after linearizing them. Agentic RAG profiling. Another possible direction for item profiling is to fully exploit LLM reasoning; LLMs iteratively determine what information is needed for the current decision, retrieve the relevant evidence, and refine the profile over multiple rounds. ∙ REAP (58) is an agentic RAG framework that iteratively re-plans sub-tasks and retrieves supporting facts. We provide the available metadata fields for each item as the retrieval space and allow the model to select among them for up to 5 rounds. As the RAG-based approach incurs substantial LLM cost, we report a focused evaluation on a subset of users in Section 5.2.2. Note that we exclude methods that use LLMs only in auxiliary roles (e.g., embedding generation (46; 48; 33)) and fine-tune the LLM itself (53; 51), since our goal is enhancing general-purpose LLMs through profiling. For our proposed method, we consider two variants: with profiles having original traits from Section 4.1, and + Refine with profiles having refined traits from Section 4.4. 5.1.3. Evaluation Setting Following prior studies (11; 13), we use each user’s last interaction for testing, the second-to-last interaction for validation, and the remaining for training. To prevent information leakage, user and item profiles are generated using only the training data. For reranking, following prior studies on LLM-based reranking (16; 20; 22), we give 20 candidate items, consisting of one ground-truth item and 19 negative samples from BPR-MF, whose order is shuffled to avoid positional bias. We report Hit Rate (HR@K) and nDCG (nDCG@K) at K∈5,10K∈\5,10\. 5.1.4. Implementation Details We use gpt-4o-mini for all LLM-based methods, including ours. EXP3RT (19) is also implemented with GPT-4o-mini for both profiling and reranking, without distillation to a smaller model for efficiency, to ensure a fair comparison under the same LLM setting. We set the temperature to 0.2 in reranking stage to make the results more deterministic, following (16), but 0.9 for other stages to promote creativity, following (5). For , we set kitems=20k_items=20, k=4k=4, Nproxy-cand=4N_proxy-cand=4, and Ncand=2N_cand=2. For the text embedding models, we use BGE-M3 (4). For the objective feature selector, the learning rate is chosen from 5e-3, 2e-3, 1e-3, and other hyperparameters are chosen following (28). All baseline hyperparameter ranges follow their original papers or official implementations. All feature-based models use both metadata features and user/item IDs, with the same preprocessing scheme as in our objective feature selector. We conduct the experiment on a server with Intel Xeon Gold 6338 CPUs and four RTX A5000 GPUs. Table 2. Statistics and features of the used datasets. Dataset Video Games Sports and Outdoors Electronics # Interactions 263,782 232,923 223,173 # Users 31,937 30,734 30,403 # Items 9,233 19,549 13,711 Sparsity 99.91% 99.96% 99.95% Common features main_category, title, average_rating, rating_number, price, store, features, description, categories, parent_asin, bought_together Optional features language, rated, genre, batteries, power source, ⋯·s (total 266) color, size, material, sport type, model year, ⋯·s (total 964) item weight, material, color, voltage, country of origin, ⋯·s (total 705) 5.2. Effect of the Proposed Profiling Strategy 5.2.1. Overall Performance Comparison Table 1 compares the baselines. We observe that outperforms across all datasets and metrics. In particular, the performance gains over LLM-based reranking baselines are most pronounced at K=5K=5, evidencing that structured item profiles help LLMs make finer-grained distinctions among top-ranked candidates. Furthermore, outperforms all baselines even without refinement, supporting the effectiveness of the core profiling strategy; the refinement stage further provides consistent additional gains, serving as a complement to the core profiling design. A key observation is that feature-enhanced variants (+Feat) of existing LLM-based methods degrade across all datasets and metrics, often falling below LLMRank, which uses only item titles. This supports that naïvely injecting raw metadata can harm LLM reranking by introducing unstructured and possibly noisy information. In contrast, from richer item information by structuring it and selecting user-relevant evidence. This underscores the importance of what information is provided to the LLM, rather than simply how much information is provided. Table 3. Comparison of RAG-based method (REAP) and ours on Electronics and Sports and Outdoors datasets. Profiling Time stands for the average time taken to make all candidate item profiles of each user, measured by seconds. Electronics Method nDCG@5 HR@5 nDCG@10 HR@10 Profiling Time REAP (5-hop) 0.1844 0.2728 0.2460 0.4657 815.70s REAP (1-hop) 0.1894 0.2729 0.2556 0.4800 346.63s 0.1892 0.2771 0.2673 0.5214 0.32s Sports and Outdoors Method nDCG@5 HR@5 nDCG@10 HR@10 Profiling Time REAP (5-hop) 0.2092 0.3043 0.2743 0.5100 736.94s REAP (1-hop) 0.2109 0.3014 0.2788 0.5143 315.87s 0.1986 0.3000 0.2663 0.5143 0.25s 5.2.2. Comparison with RAG-based Profiling To compare RAG-based profiling, we consider two REAP variants: REAP (5-hop), which iteratively retrieves information over up to 5 rounds, and REAP (1-hop), which collapses retrieval into a single step. These variants require repeated LLM calls to identify important information in each context. Due to the substantial cost of REAP, we randomly sample 700 users from each test set here. Table 3 reports performance and average per-user profiling time. We observe that, while both REAP variants require the prohibitively heavy latency due to multiple LLM callings, a significant reduction in the average profiling time, while still showing competitive performance to them. This latency gain is attributable to our LLM-free, offline-prepared profiler that requires only lightweight matrix operations at serving time. Table 4. Offline and online efficiency comparison. Dataset Method Offline Online # LLM Calls Avg. Token Profiling Time Reranking Time† Video Games EXP3RT 7.55 4584.6 – 2.89s 7.84 4112.4 0.21s 3.03s Sports and Outdoors EXP3RT 7.55 4887.6 – 3.15s 8.18 4954.4 0.22s 3.02s Electronics EXP3RT 6.97 4337.6 – 3.52s 7.42 4100.0 0.21s 3.20s • †: Reranking time is sensitive to server status; for reference only. Practically both methods have comparable reranking times. 5.2.3. Computational Cost Analysis In Table 4, we compare computational cost involved offline and online to EXP3RT, a review-based profiling method like ours. We measure the average number of LLM calling involved in the offline profiling process, the average number of tokens in the final prompts, and the average latencies involved in online profiling and LLM calling, all per a test user. To measure the reranking time, we made asynchronous LLM calls with the maximum concurrency 5. We can observe that a marginally higher number of offline LLM calls, while online serving latency is comparable to EXP3RT without the online profiling. Together, these results show that ’s consistent performance gains over EXP3RT come at a relatively negligible additional cost. 5.3. Study of Table 5. Ablation study on key/trait construction strategies. K- and T- stand for the variants applied to key exploration stage and trait extraction stage, respectively. Dataset Strategy nDCG@5 HR@5 nDCG@10 HR@10 Video Games K-Frequency 0.1732 0.2585 0.2330 0.4461 K-MergeMore 0.1734 0.2617 0.2402 0.4711 T-Single 0.1840 0.2723 0.2366 0.4370 0.1843 0.2778 0.2438 0.4622 Electronics K-Frequency 0.1713 0.2560 0.2359 0.4587 K-MergeMore 0.1735 0.2618 0.2342 0.4529 T-Single 0.1715 0.2562 0.2372 0.4329 0.1757 0.2639 0.2406 0.4678 Table 6. Ablation study on features/trait selection strategies. O- and S- mean the strategies applied to the objective feature selection and the subjective trait selection, respectively. Dataset Strategy nDCG@5 HR@5 nDCG@10 HR@10 Video Games O-Random 0.1799 0.2701 0.2393 0.4503 O-All 0.1782 0.2643 0.2304 0.4290 O-Sim 0.1832 0.2679 0.2394 0.4441 S-Random 0.1797 0.2658 0.2368 0.4448 S-All 0.1744 0.2555 0.2299 0.4302 0.1843 0.2778 0.2438 0.4622 Electronics O-Random 0.1782 0.2518 0.2332 0.4292 O-All 0.1601 0.2279 0.2121 0.3918 O-Sim 0.1778 0.2528 0.2351 0.4326 S-Random 0.1785 0.2541 0.2345 0.4298 S-All 0.1674 0.2393 0.2208 0.4070 0.1797 0.2571 0.2366 0.4359 5.3.1. Ablation Study We provide ablation study results of Video Games and Electronics datasets. Effect of key/trait construction. Table 5 ablates our key exploration and trait extraction designs in Section 4.1. For domain-specific key exploration, we compare ours with K-Frequency, which selects the same number of keys from the most frequent optional features for each domain, and K-MergeMore, which modifies preassess-&-mergep_reassess-\&-merge (Eq. 3) to ignore sparse features and make a smaller key set. Both generally perform worse than , supporting that our key generation strategy allows fine-grained distinctions between diverse types of items and helps LLMs reason about item differences. For trait extraction, T-Single, which generates only one trait per item, similarly degrades performance. It evidences that user opinions on an item are multi-faceted and therefore it is critical to provide diverse trait candidates from which the selector can identify the one most aligned with each user’s context. Effect of feature/trait selection. To observe the effect of our feature/trait selection strategy in Section 4.2, we compare the selection strategy of different selection strategies of objective features and subjective traits: (1) Random randomly samples a subset of objective features or subjective traits; (2) All includes all available features or traits in the profile. For objective feature selection, we also consider Sim, which selects the features by cosine similarity between feature embeddings and the user profile embedding. Table 6 reports the results on Video Games and Electronics datasets. For O-Random, we sample four features, matching our selector. We first observe that the best performance in all cases, supporting that each module identifies decision-relevant information for each user-item context. In contrast, O-All and S-All consistently perform worse, evidencing that indiscriminate feature injection causes distraction of LLM rather than mere redundancy. We also note that the performance gaps remain moderate, likely because the non-ablated component continues to provide well-constructed signals even when the other is ablated. Lastly, although O-Sim outperforms random selection, it remains below our strategy, showing that objective feature importance is better captured through collaborative patterns than semantic similarity alone. "A chart showing transferability results." Figure 4. Recommendation performance of diverse LLM-based recommendation methods with profiling on Electronics dataset applied to different families and sizes of LLMs. 5.3.2. Transferability to Other LLMs To examine the transferability of our generated item profiles to diverse LLM families and scales, we conduct experiments using four LLMs spanning two families: Qwen2.5-32B-Instruct, Qwen2.5-7B-Instruct (50), Llama3.1-8B-Instruct and Llama3.2-3B-Instruct (38). These models range from 3B to 32B parameters, covering both small and large variants. The results of the comparison is reported in Figure 4. outperforms both baselines across all LLM families and the parameter scales. This supports that our generated item profiles are not tailored to a specific LLM architecture or scale, but instead hold useful item knowledge that helps diverse LLMs better understand the recommendation context. Furthermore, even the smallest 3B model, evidencing that well-structured profiles can compensate for limited model capacity. Table 7. Case study on Electronics dataset. Red text can work as signals for functionality, and blue text for brand-loyalty. Item: B0B6WTFTG9 (Google Pixel Buds Pro - Noise Canceling Earbuds …) Case #1: User ID AHTS2DVIJSO6IARPLRV55RGDHMMA User profile: This user consumes electronics primarily to enhance the protection and functionality of their devices. While they value high quality and durability, they consistently criticize comfort issues and poor functionality, …. Item profile: Audio Technology: Active Noise Cancellation with custom 11 m drivers Included Accessories: Earbuds, Eartips, Wireless Charging Case User-specific Trait: Comparative Evaluation; Users are actively comparing these earbuds with competitors, focusing on sound performance, noise cancellation, and overall features. … ▶ Rank: 2 Case #2: User ID AGS4TRMTPU3Y2CMJ52E6DOLX4GPA User profile: This consumer primarily seeks electronics that offer seamless compatibility with their devices, often using them in specific situations such as charging earphones or operating peripherals with Chromebooks. … Our item profile: Device/OS Compatibility: Google Assistant-enabled Android 6.0 Smart and AI Features: Hands-free Google Assistant support Store: Google User-specific Trait: Brand Loyalty; There are consumers who exhibit a strong preference for Google products, appreciating the familiarity and perceived reliability associated with the brand. … ▶ Rank: 1 cf. EXP3RT item profile: [Like] - Sound quality is improved compared to the Pixel Buds A Series. - It includes additional features such as ANC and Transparency modes. - It works well with other Google products. - Best stock sound quality of all the earbuds, with excellent touch controls and responsive features (Google Pixelbuds Pro). … [Dislike] - It has insecure fit as they do not go into the ear canal or have wingtips for stability. - The case feels cheap and bulky. - It is quite overpriced compared to others (Samsung Galaxy Buds 2 Pro). … [Average Rating] 4.7 ▶ Rank: 19 on Case #1, 7 on Case #2 5.3.3. Case Study Table 7 presents item profiles generated by EXP3RT for the same item across two different users. For Case #1, who criticizes poor functionality, features around its performance (e.g., noise cancelling). On the other hand, for Case #2, who seeks brand ecosystem compatibility, it instead presents brand-specific information, where the combination of these features collectively reinforces the ecosystem angle to form a coherent insight. EXP3RT, by contrast, applies a single static profile to both users; although the resulting profile contains signals relevant to each user, these are buried along with context-irrelevant information, making the LLM struggle to identify what matters for each of the user. This evidences that user-adaptive profile construction is essential; the same item information can either guide or mislead the LLM depending on how selectively it is presented. 6. Conclusion In this paper, we propose address underexplored challenges in item profiling from vast item metadata and item/user context-relevant information. structures raw metadata and reviews into objective features and subjective traits, and employs lightweight selectors to construct user-specialized item profiles with reduced serving-time overhead. Extensive experiments show that outperforms existing baselines, supporting that careful organization and user-adaptive selection of item information are as important as providing valuable information. We expect this study to lay the groundwork for item-side profiling in recommendation, extending beyond the prevailing user-side focus. Acknowledgments This work was the result of project supported by KT (Korea Telecom)-Korea University AICT R&D Center. This work was also supported by the IITP-ICT Creative Consilience Program grant funded by the MSIT (IITP-2026-RS-2020-I201819), Basic Science Research Program through the NRF funded by the Ministry of Education (NRF-2021R1A6A1A03045425), the NRF grant funded by the MSIT (RS-2026-25486220), and the IITP grant funded by the MSIT (IITP-2026-RS-2026-25616664, AI Star Fellowship Support Program). GenAI Usage Disclosure GPT-4o-mini, Llama3.1, Llama3.2 and Qwen2.5 were used as a methodological component of the proposed framework, as described in the Experiment section. These models were used only within the reported experimental settings specified in the work. Generative AI tools (ChatGPT and Claude) were used during manuscript preparation solely for minor language editing, specifically for grammar checking, spelling and typo correction. These tools were used to suggest debugging directions in code development. All AI-assisted edits were manually reviewed and verified by the authors before being incorporated into the work. Apart from the uses disclosed above, no other generative AI tools were used for experimentation, data analysis, or content generation. References Bao et al. (2023) K. Bao, J. Zhang, Y. Zhang, W. Wang, F. Feng, and X. He TALLRec: an effective and efficient tuning framework to align large language model with recommendation. In Proceedings of the 17th ACM Conference on Recommender Systems, RecSys 2023, Singapore, Singapore, September 18-22, 2023, p. 1007–1014. Cited by: §2, §5.1.1. Carraro and Bridge (2026) D. Carraro and D. G. Bridge Enhancing recommendation diversity by re-ranking with large language models. Trans. Recomm. Syst. 4 (2), p. 18:1–18:40. Cited by: §1. Chen et al. (2025a) A. Chen, C. Du, J. Chen, J. Xu, Y. Zhang, S. Yuan, Z. Chen, L. Li, and Y. Xiao DEEPER insight into your user: directed persona refinement for dynamic persona modeling. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 24157–24180. Cited by: §1, §1, §2. Chen et al. (2024) J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu BGE m3-embedding: multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. CoRR abs/2402.03216. External Links: 2402.03216 Cited by: §5.1.4, footnote 1, footnote 3. Chen et al. (2025b) L. Chen, Q. Zeng, and H. Chen M-LLM3 3rec: A motivation-aware user-item interaction framework for enhancing recommendation accuracy with llms. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM 2025, Seoul, Republic of Korea, November 10-14, 2025, p. 291–300. Cited by: §1, §2, 3rd item, §5.1.1, §5.1.4. Chen (2023) Z. Chen PALR: personalization aware llms for recommendation. CoRR abs/2305.07622. External Links: 2305.07622 Cited by: §1, §1, §2, §2, §2. Cheng et al. (2016) H. Cheng, L. Koc, J. Harmsen, T. Shaked, T. Chandra, H. Aradhye, G. Anderson, G. Corrado, W. Chai, M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain, X. Liu, and H. Shah Wide & deep learning for recommender systems. In Proceedings of the 1st Workshop on Deep Learning for Recommender Systems, DLRS@RecSys 2016, Boston, MA, USA, September 15, 2016, p. 7–10. Cited by: §1. Dai et al. (2023) S. Dai, N. Shao, H. Zhao, W. Yu, Z. Si, C. Xu, Z. Sun, X. Zhang, and J. Xu Uncovering chatgpt’s capabilities in recommender systems. In Proceedings of the 17th ACM Conference on Recommender Systems, RecSys 2023, Singapore, Singapore, September 18-22, 2023, p. 1126–1132. Cited by: §1. Fang et al. (2025) Y. Fang, W. Wang, Y. Zhang, F. Zhu, Q. Wang, F. Feng, and X. He Reason4Rec: large language models for recommendation with deliberative user preference alignment. CoRR abs/2502.02061. External Links: 2502.02061 Cited by: §1, §2. Gao et al. (2022) C. Gao, S. Li, Y. Zhang, J. Chen, B. Li, W. Lei, P. Jiang, and X. He KuaiRand: an unbiased sequential recommendation dataset with randomly exposed videos. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management, Atlanta, GA, USA, October 17-21, 2022, p. 3953–3957. Cited by: §3. Gao et al. (2025a) J. Gao, B. Chen, X. Zhao, W. Liu, X. Li, Y. Wang, W. Wang, H. Guo, and R. Tang LLM4Rerank: llm-based auto-reranking framework for recommendations. In Proceedings of the ACM on Web Conference 2025, W 2025, Sydney, NSW, Australia, 28 April 2025- 2 May 2025, p. 228–239. Cited by: §1, §1, §1, §2, §5.1.3. Gao et al. (2025b) Z. Gao, J. Zhou, Y. Dai, and T. Joachims LangPTune: optimizing language-based user profiles for recommendation. In Proceedings of the 34th ACM International Conference on Information and Knowledge Management, CIKM 2025, Seoul, Republic of Korea, November 10-14, 2025, p. 707–717. Cited by: §1, §4.2.1, §4.2.1. Geng et al. (2022) S. Geng, S. Liu, Z. Fu, Y. Ge, and Y. Zhang Recommendation as language processing (RLP): A unified pretrain, personalized prompt & predict paradigm (P5). In RecSys ’22: Sixteenth ACM Conference on Recommender Systems, Seattle, WA, USA, September 18 - 23, 2022, p. 299–315. Cited by: §1, §2, §5.1.3. Hou et al. (2026) M. Hou, X. Liu, L. Wu, C. He, H. Liu, Z. Li, X. Li, and S. Wei WeaveRec: an llm-based cross-domain sequential recommendation framework with model merging. In Proceedings of the ACM Web Conference 2026, W 2026, Dubai, United Arab Emirates, originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026, p. 6342–6353. Cited by: §1. Hou et al. (2024a) Y. Hou, J. Li, Z. He, A. Yan, X. Chen, and J. J. McAuley Bridging language and items for retrieval and recommendation. CoRR abs/2403.03952. External Links: 2403.03952 Cited by: §5.1.1. Hou et al. (2024b) Y. Hou, J. Zhang, Z. Lin, H. Lu, R. Xie, J. J. McAuley, and W. X. Zhao Large language models are zero-shot rankers for recommender systems. In Advances in Information Retrieval - 46th European Conference on Information Retrieval, ECIR 2024, Glasgow, UK, March 24-28, 2024, Proceedings, Part I, Lecture Notes in Computer Science, p. 364–381. Cited by: §1, §2, 1st item, §5.1.3, §5.1.4. Huzhang et al. (2023) G. Huzhang, Z. Pang, Y. Gao, Y. Liu, W. Shen, W. Zhou, Q. Lin, Q. Da, A. Zeng, H. Yu, Y. Yu, and Z. Zhou AliExpress learning-to-rank: maximizing online model performance without going online. IEEE Trans. Knowl. Data Eng. 35 (2), p. 1214–1226. Cited by: §3. Kang and McAuley (2018) W. Kang and J. J. McAuley Self-attentive sequential recommendation. In IEEE International Conference on Data Mining, ICDM 2018, Singapore, November 17-20, 2018, p. 197–206. Cited by: 2nd item. Kim et al. (2025) J. Kim, H. Kim, H. Cho, S. Kang, B. Chang, J. Yeo, and D. Lee Review-driven personalized preference reasoning with large language models for recommendation. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025, Padua, Italy, July 13-18, 2025, p. 1697–1706. Cited by: §1, §1, §2, §2, §4.1.3, 2nd item, §5.1.1, §5.1.4. Kweon et al. (2025) W. Kweon, S. Jang, S. Kang, and H. Yu Uncertainty quantification and decomposition for llm-based recommendation. In Proceedings of the ACM on Web Conference 2025, W 2025, Sydney, NSW, Australia, 28 April 2025- 2 May 2025, p. 4889–4901. Cited by: §5.1.3. Lee et al. (2026a) G. Lee, W. Kweon, Z. Yue, Y. Liu, Y. Liu, S. Yoon, D. Wang, and S. Kang SPRINT: scalable and predictive intent refinement for llm-enhanced session-based recommendation. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2026, MelbourneVICAustralia, July 20-24, 2026, p. 902–912. Cited by: §1. Lee et al. (2026b) J. Lee, S. Jang, S. Kang, and H. Yu Filling the gaps: selective knowledge augmentation for LLM recommenders. In Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2026, MelbourneVICAustralia, July 20-24, 2026, p. 891–901. Cited by: §5.1.3. Lee et al. (2023) Y. Lee, Y. Jeong, K. Park, and S. Kang MvFS: multi-view feature selection for recommender system. In Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, CIKM 2023, Birmingham, United Kingdom, October 21-25, 2023, p. 4048–4052. Cited by: §4.2.2. Li et al. (2010) L. Li, W. Chu, J. Langford, and R. E. Schapire A contextual-bandit approach to personalized news article recommendation. In Proceedings of the 19th International Conference on World Wide Web, W 2010, Raleigh, North Carolina, USA, April 26-30, 2010, p. 661–670. Cited by: §3. Li et al. (2026) S. Li, Y. Wang, J. Wang, Y. Li, J. Ghosh, and A. Cocos LLM reasoning for cold-start item recommendation. In Proceedings of the ACM Web Conference 2026, W 2026, Dubai, United Arab Emirates, originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026, p. 8409–8412. Cited by: §1. Lian et al. (2018) J. Lian, X. Zhou, F. Zhang, Z. Chen, X. Xie, and G. Sun XDeepFM: combining explicit and implicit feature interactions for recommender systems. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, KDD 2018, London, UK, August 19-23, 2018, p. 1754–1763. Cited by: §1, 1st item. Lightman et al. (2024) H. Lightman, V. Kosaraju, Y. Burda, H. Edwards, B. Baker, T. Lee, J. Leike, J. Schulman, I. Sutskever, and K. Cobbe Let’s verify step by step. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §4.4. Lin et al. (2022) W. Lin, X. Zhao, Y. Wang, T. Xu, and X. Wu AdaFS: adaptive feature selection in deep recommender system. In KDD ’22: The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Washington, DC, USA, August 14 - 18, 2022, p. 3309–3317. Cited by: §1, §4.2.2, §4.2.2, §4.2.2, 2nd item, §5.1.4, footnote 2. Liu et al. (2026) H. Liu, Z. Sun, T. Wei, Y. Wang, J. Zhu, and X. Qu Diagnostic-guided dynamic profile optimization for llm-based user simulators in sequential recommendation. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, p. 15306–15314. Cited by: §4.4. Liu et al. (2023) J. Liu, C. Liu, R. Lv, K. Zhou, and Y. Zhang Is chatgpt a good recommender? A preliminary study. CoRR abs/2304.10149. External Links: 2304.10149 Cited by: §1. Liu et al. (2024) N. F. Liu, K. Lin, J. Hewitt, A. Paranjape, M. Bevilacqua, F. Petroni, and P. Liang Lost in the middle: how language models use long contexts. Trans. Assoc. Comput. Linguistics 12, p. 157–173. Cited by: §1. Liu et al. (2022) W. Liu, J. Qin, R. Tang, and B. Chen Neural re-ranking for multi-stage recommender systems. In RecSys ’22: Sixteenth ACM Conference on Recommender Systems, Seattle, WA, USA, September 18 - 23, 2022, p. 698–699. Cited by: §1, §3. Nie and Sun (2026) Z. Nie and P. Sun HADSF: aspect aware semantic control for explainable recommendation. In Proceedings of the Nineteenth ACM International Conference on Web Search and Data Mining, WSDM 2026, Boise, ID, USA, February 22-26, 2026, p. 509–519. Cited by: §4.1.3, §4.2.1, §5.1.2. Pryzant et al. (2023) R. Pryzant, D. Iter, J. Li, Y. T. Lee, C. Zhu, and M. Zeng Automatic prompt optimization with "gradient descent" and beam search. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, p. 7957–7968. Cited by: §4.4. Ramos et al. (2024) J. Ramos, H. A. Rahmani, X. Wang, X. Fu, and A. Lipani Transparent and scrutable recommendations using natural language user profiles. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 13971–13984. Cited by: §1, §2, §2. Rendle et al. (2009) S. Rendle, C. Freudenthaler, Z. Gantner, and L. Schmidt-Thieme BPR: bayesian personalized ranking from implicit feedback. In UAI 2009, Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, Montreal, QC, Canada, June 18-21, 2009, p. 452–461. Cited by: 1st item. Sun et al. (2019) F. Sun, J. Liu, J. Wu, C. Pei, X. Lin, W. Ou, and P. Jiang BERT4Rec: sequential recommendation with bidirectional encoder representations from transformer. In Proceedings of the 28th ACM International Conference on Information and Knowledge Management, CIKM 2019, Beijing, China, November 3-7, 2019, p. 1441–1450. Cited by: 2nd item. Team (2024) L. Team The llama 3 herd of models. Vol. abs/2407.21783. External Links: 2407.21783 Cited by: §5.3.2. Tian et al. (2026) R. Tian, X. Xu, B. Jin, S. Kang, and J. Han LLM-based compact reranking with document features for scientific retrieval. In Proceedings of the 32st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, KDD 2026, Jeju, Korea, August 9-13, 2026, p. 12114–12125. Cited by: §1. Tsai et al. (2024) A. Tsai, A. Kraft, L. Jin, C. Cai, A. Hosseini, T. Xu, Z. Zhang, L. Hong, E. H. Chi, and X. Yi Leveraging LLM reasoning enhances personalized recommender systems. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, Findings of ACL, p. 13176–13188. Cited by: §1, §2. Wang et al. (2025a) B. Wang, F. Liu, J. Chen, X. Lou, C. Zhang, J. Wang, Y. Sun, Y. Feng, C. Chen, and C. Wang MSL: not all tokens are what you need for tuning LLM as a recommender. In Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR 2025, Padua, Italy, July 13-18, 2025, p. 1912–1922. Cited by: §1. Wang et al. (2025b) L. Wang, D. Zhang, F. Yang, P. Zhao, J. Liu, Y. Zhan, H. Sun, Q. Lin, W. Deng, D. Zhang, F. Sun, and Q. Zhang LettinGo: explore user profile generation for recommendation system. In Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining, V.2, KDD 2025, Toronto ON, Canada, August 3-7, 2025, L. Antonie, J. Pei, X. Yu, F. Chierichetti, H. W. Lauw, Y. Sun, and S. Parthasarathy (Eds.), p. 2985–2995. Cited by: §1, §1, §2, §2, §2. Wang et al. (2024) P. Wang, L. Li, Z. Shao, R. Xu, D. Dai, Y. Li, D. Chen, Y. Wu, and Z. Sui Math-shepherd: verify and reinforce llms step-by-step without human annotations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, p. 9426–9439. Cited by: §4.4. Wang et al. (2025c) S. Wang, W. Fan, Y. Feng, S. Lin, X. Ma, S. Wang, and D. Yin Knowledge graph retrieval-augmented generation for llm-based recommendation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, p. 27152–27168. Cited by: §2. Wang et al. (2022) Y. Wang, X. Zhao, T. Xu, and X. Wu AutoField: automating feature selection in deep recommender systems. In W ’22: The ACM Web Conference 2022, Virtual Event, Lyon, France, April 25 - 29, 2022, p. 1977–1986. Cited by: §4.2.2. Wei et al. (2024) W. Wei, X. Ren, J. Tang, Q. Wang, L. Su, S. Cheng, J. Wang, D. Yin, and C. Huang LLMRec: large language models with graph augmentation for recommendation. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining, WSDM 2024, Merida, Mexico, March 4-8, 2024, p. 806–815. Cited by: §1, §5.1.2. Wu et al. (2026) S. Wu, Z. Du, Q. Jia, and Z. Dong REACTION: parameter-efficient learning for recommendation. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, p. 15977–15985. Cited by: 3rd item, footnote 2. Xi et al. (2024) Y. Xi, W. Liu, J. Lin, X. Cai, H. Zhu, J. Zhu, B. Chen, R. Tang, W. Zhang, and Y. Yu Towards open-world recommendation with knowledge augmentation from large language models. In Proceedings of the 18th ACM Conference on Recommender Systems, RecSys 2024, Bari, Italy, October 14-18, 2024, p. 12–22. Cited by: §1, §5.1.2. Xie et al. (2026) Z. Xie, B. Peng, Z. He, Z. Chen, A. Han, I. Ye, B. Coleman, N. Sachdeva, F. Pereira, J. J. McAuley, W. Kang, D. Z. Cheng, B. Wang, and R. Brown AgenticTagger: structured item representation for recommendation with LLM agents. CoRR abs/2602.05945. External Links: 2602.05945 Cited by: §2. Yang et al. (2024) A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu Qwen2.5 technical report. CoRR abs/2412.15115. External Links: 2412.15115 Cited by: §5.3.2. Yu et al. (2026) Q. Yu, K. Fu, Z. Lv, S. Zhang, X. Wu, C. Lin, F. Wei, B. Zheng, and F. Wu ThinkRec: thinking-based recommendation via LLM. In Proceedings of the ACM Web Conference 2026, W 2026, Dubai, United Arab Emirates, originally scheduled for April 13-17, 2026, rescheduled for June 29 - July 3, 2026, p. 5698–5709. Cited by: §1, §2, §2, §5.1.2. Yue et al. (2023) Z. Yue, S. Rabhi, G. de Souza Pereira Moreira, D. Wang, and E. Oldridge LlamaRec: two-stage recommendation using large language models for ranking. CoRR abs/2311.02089. External Links: 2311.02089 Cited by: §1, §5.1.1. Zhang et al. (2025) Y. Zhang, F. Feng, J. Zhang, K. Bao, Q. Wang, and X. He CoLLM: integrating collaborative embeddings into large language models for recommendation. IEEE Trans. Knowl. Data Eng. 37 (5), p. 2329–2340. Cited by: §2, §5.1.2. Zhang et al. (2021) Y. Zhang, H. Ding, Z. Shui, Y. Ma, J. Zou, A. Deoras, and H. Wang Language models as recommender systems: evaluations and limitations. In I (Still) Can’t Believe It’s Not Better Workshop at NeurIPS, Cited by: §1. Zhao et al. (2024) Z. Zhao, W. Fan, J. Li, Y. Liu, X. Mei, Y. Wang, Z. Wen, F. Wang, X. Zhao, J. Tang, and Q. Li Recommender systems in the era of large language models (llms). IEEE Trans. Knowl. Data Eng. 36 (11), p. 6889–6907. Cited by: §1, §2. Zheng et al. (2024) Z. Zheng, W. Chao, Z. Qiu, H. Zhu, and H. Xiong Harnessing large language models for text-rich sequential recommendation. In Proceedings of the ACM on Web Conference 2024, W 2024, Singapore, May 13-17, 2024, T. Chua, C. Ngo, R. Kumar, H. W. Lauw, and R. K. Lee (Eds.), p. 3207–3216. Cited by: §1, §1, §2, §2, §2. Zhou et al. (2024) J. Zhou, Y. Dai, and T. Joachims Language-based user profiles for recommendation. CoRR abs/2402.15623. External Links: 2402.15623 Cited by: §1. Zhu et al. (2026) Y. Zhu, H. Zhou, W. Hong, T. Liu, and N. Wang REAP: enhancing RAG with recursive evaluation and adaptive planning for multi-hop question answering. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, p. 35230–35238. Cited by: 1st item.