Paper deep dive
Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy
Liwei Deng, Jing Jiang, Zhiwei Li, Yang Wang, Guodong Long
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/17/2026, 4:45:09 AM
Summary
This paper introduces the Content Depth Score (CDS), a seven-level metric based on cognitive psychology theories, to quantify the depth of short-video content. The authors present SCOPE-Bench, a benchmark with CDS annotations for 150K videos, to evaluate 13 representative Recommender Systems (RSs). The study reveals that current RSs, optimized for engagement, consistently recommend shallow content with depth scores comparable to random selection, highlighting a decoupling between engagement and cognitive value.
Entities (9)
Relation Signals (8)
SCOPE-Bench → containsannotationsfor → Content Depth Score
confidence 98% · SCOPE-Bench provides CDS annotations for 150K videos
Content Depth Score → isbasedon → Cognitive Psychology
confidence 95% · CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning.
Recommender Systems → optimizesfor → User Engagement
confidence 95% · short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement
Content Depth Score → measures → Cognitive Processes
confidence 94% · CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes
Recommender Systems → recommends → Shallow Content
confidence 92% · These systems inherently favor shallow-content videos that are effective at attracting immediate attention.
Dual-Process Theory → anchors → Content Depth Score
confidence 90% · The resulting rubric ranges from immediate affective responses to increasingly complex cognitive processes... grounded in Dual Process Theory
Bloom's Taxonomy → anchors → Content Depth Score
confidence 90% · Bloom’s Taxonomy organizes cognitive processes into six levels of increasing complexity... used to define the remaining levels.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users' cognitive engagement and mental well-being, raising concerns about the long-term societal impact of the short-video platform. To tackle this challenge, this paper introduces a new metric, the \textbf{Content Depth Score (CDS)}, to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present \textbf{SCOPE-Bench}, the first benchmark for content-depth evaluation in short-video recommendation. Built upon a large-scale open-source short-video dataset, SCOPE-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive-content perspective. Leveraging SCOPE-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow-content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2608.13990v1
- Canonical: https://arxiv.org/abs/2608.13990v1
Trouble viewing inline? Open PDF directly →
Full Text
48,223 characters extracted from source content.
Expand or collapse full text
Content Depth Matters in Short-Video Recommendation: Rethinking the Attention Economy Liwei Deng Jing Jiang Zhiwei Li Thanks: Corresponding author: Zhiwei Li (zhw.li@outlook.com). Yang Wang Guodong Long Abstract Driven by the attention economy, short-video Recommender Systems (RSs) are primarily optimized to maximize user engagement by promoting videos that capture attention within seconds. These systems inherently favor shallow-content videos that are effective at attracting immediate attention. However, growing evidence suggests that prolonged exposure to such content may negatively affect users’ cognitive engagement and mental well-being, raising concerns about the long-term societal impact of the short-video platform. To tackle this challenge, this paper introduces a new metric, the Content Depth Score (CDS), to quantify the content depth of short videos. CDS measures the extent to which a video is expected to stimulate higher-order cognitive processes, using a seven-level scale grounded in established theories of cognitive psychology and learning. As an initial step toward this vision, we present SCOPE-Bench, the first benchmark for content-depth evaluation in short-video recommendation. Built upon a large-scale open-source short-video dataset, SCOPE-Bench provides CDS annotations for 150K videos, enabling systematic evaluation of RSs from a cognitive-content perspective. Leveraging SCOPE-Bench, we evaluate 13 representative RSs and reveal a consistent preference for shallow-content videos. Moreover, we find that these algorithms recommending cognitively deep content are only marginally better than random selection, highlighting a previously overlooked limitation of existing recommendation objectives. Our code and datasets are available at https://liweidengdavid.github.io/SCOPE-Bench/. 1 Introduction In the attention economy, short-video feeds on platforms such as YouTube (Ørmen and Gregersen 2023), TikTok (TikTok Team 2021), and Kuaishou (Kuaishou Technology 2025) have become a central channel for everyday content consumption. Adult users now spend more than one hour per day on these platforms (Kemp 2025). Their Recommender Systems (RSs) (Zou and Sun 2025; Li et al. 2026) optimize engagement signals, such as watch time and clicks, to turn user attention into revenue (García-Rapp 2017), and therefore favor videos capable of rapidly capturing users’ attention. Over time, continuous exposure to such content may weaken users’ sustained attention (Mahakud and Thapliyal 2026) and affect their long-term well-being (Tang et al. 2026). Figure 1: Evaluating 13 Recommender Systems (RSs) on both Engagement metric (X axis) and our Content-Depth metric (Y axis). The regions are split into three levels corresponding to content-depth metric: Low (blue), Medium (green), and High (red). The dashed line marks the content depth of a random selection baseline. Existing RSs perform competitively on the engagement metric while their content-depth metric is low and near to the random selection. Optimizing for engagement rewards videos that retain user attention, which raises a more fundamental question: Do RSs favor attention-grabbing content over videos with greater depth? Content depth reflects how deeply a video develops its information, providing users with richer opportunities for understanding, reasoning, and reflection (Chi and Wylie 2014). We examine this question by comparing the engagement of 13 representative RSs with the content depth of the videos they recommend. As Figure 1 shows, the content-depth performance of every RS remains close to that of random recommendation. Therefore, stronger engagement does not translate into greater recommended content depth. The situation has already drawn awareness and responses beyond research. For example, Australia (Australian Government 2025) and the UK (UK Government 2026) have imposed under-16 restrictions on short-video platforms, and platforms themselves cap adolescents’ daily usage (Keenan 2023). This convergence of regulators signals that the harm to young users is now widely acknowledged. Yet these interventions act on access rather than content. One important reason is the absence of any measure for the content itself. Inspired by the need to promote healthier short video and build a sustainable short-video ecosystem, we propose a new metric to make content depth measurable, enabling it to serve as an explicit objective in RS evaluation and optimization. This paper proposes the Content Depth Score (CDS), a novel metric to measure how deeply a short video presents and delivers information to humans. CDS scores a video on a seven-level rubric grounded in theories of cognition and learning (Wason and Evans 1974; Anderson and Krathwohl 2001; Biggs and Collis 2014), ranging from low-level emotional stimulation to higher-order cognitive processes11 1 Specifically, in this paper, cognitive processes refer to the mental operations through which individuals acquire, process, store, and use information (Sternberg, Sternberg, and Mio 2006). CDS measures it as the opportunities a video’s content provides for these operations.. To validate CDS at both the item and recommendation-list levels, we construct SCOPE-Bench, a Short-video COntent dePth Evaluation Benchmark, by extending the existing open-source ShortVideo dataset (Shang et al. 2025). Specifically, we annotate its 150K videos with CDS labels, covering about 1M user-item interactions from 10K users. Our main contributions are summarized as follows: • To the best of our knowledge, we are the first to formulate and systematically investigate the lack of attention to content depth in engagement-optimized RSs. • We propose CDS, the first quantitative metric for video content depth, based on a seven-level rubric grounded in theories of cognition and learning. We further extended CDS to a list-level metric for recommended lists. • We establish a comprehensive evaluation framework for video content depth and construct SCOPE-Bench, a benchmark of 150K videos with human-aligned CDS annotations and about 1M interactions from 10K users. • Experiments on SCOPE-Bench confirm that the CDS evaluation protocol closely aligns with human judgments. Using CDS to evaluate 13 representative RSs, we find that their recommended lists achieve content-depth performance close to that of random recommendation, revealing that engagement and content depth are decoupled. 2 Related Works 2.1 Content Quality Assessment Content quality assessment can be broadly categorized into metric-based assessment for explicitly quantifiable properties and judgment-based assessment for open-ended or interpretive properties that are difficult to capture with conventional metrics. The former mainly covers video quality assessment, including perceptual fidelity (Wang et al. 2004), temporal consistency (Wang, Lu, and Bovik 2004), generative realism (Heusel et al. 2017), and cross-modal alignment (Radford et al. 2021), as well as text quality assessment, covering linguistic (Napoles, Sakaguchi, and Tetreault 2017) and semantic properties (Barzilay and Lapata 2008), trustworthiness (Lin, Hilton, and Evans 2022), and safety-related dimensions (Gehman et al. 2020). For open-ended and interpretive properties, recent studies have increasingly adopted LLM-as-a-Judge as a flexible judgment-based evaluation paradigm. Existing approaches range from scalar scoring (Liu et al. 2023) and pairwise comparison (Zheng et al. 2023) to rubric-based evaluation (Kim et al. 2024; Wang et al. 2024) and fine-grained assessment protocols (Ye et al. 2023). While effective, these methods provide only a limited conceptualization of content depth. In the paper, we systematically define content depth and propose the first quantitative metric, termed the CDS, to measure the content depth. 2.2 Evaluation of Recommender Systems Existing evaluation of RSs can be broadly categorized into system-centric evaluation and user-centric evaluation. System-centric evaluation covers accuracy-oriented evaluation and beyond-accuracy evaluation (Zangerle and Bauer 2022). The former assesses whether relevant items are accurately retrieved and ranked using metrics such as Precision (Herlocker et al. 2004), Recall (Allen et al. 1955), and NDCG (Järvelin and Kekäläinen 2002), whereas the latter considers complementary recommendation qualities, including diversity (Ziegler et al. 2005), novelty (Vargas and Castells 2011), and catalog coverage (Ge, Delgado-Battenfeld, and Jannach 2010). In contrast, user-centric evaluation examines users’ subjective perceptions and experiences, including choice difficulty (Knijnenburg et al. 2012), usefulness (Pu, Chen, and Hu 2011), and satisfaction (Hijikata, Kai, and Nishida 2012). Although prior studies have attempted to evaluate RSs beyond conventional engagement metrics, little research has quantitatively examined the content depth of recommended videos. To address this gap, we introduce the List-wise CDS (LCDS), the first list-level metric for quantifying the content depth of recommended lists. 3 Content Depth Score Metric CDS Level Rubric Label Operational Criterion Dual Process Theory Bloom’s Taxonomy SOLO Taxonomy Level 0 Affect Mainly evokes affect, humor, spectacle, or atmosphere. System 1 – – Level 1 Point Presents an isolated opinion, or label without explaining why or how. System 2 Remember Prestructural Level 2 Concept Defines or illustrates a single idea with simple background or explanation. System 2 Understand Unistructural Level 3 Procedure Shows how a method can be used in concrete cases, typically through multiple steps or illustrative examples. System 2 Apply Multistructural Level 4 Mechanism Explains mechanisms, variables, constraints, conditions, causal links, or system relationships. System 2 Analyze Relational Level 5 Judgment Weighs evidence or competing explanations, including limitations, uncertainty, or counterexamples. System 2 Evaluate Relational / Extended Abstract Level 6 Model Builds a generalizable model, framework, principle, or decision rule that can transfer across contexts. System 2 Create Extended Abstract Table 1: Seven-level scoring rubric for CDS and its approximate theoretical anchors. Each level is defined by an operational criterion for consistent annotation. A dash indicates that no direct theoretical correspondence is assigned. The rubric captures a progression from affective responses to increasingly complex reasoning and transferable model construction. 3.1 Content Depth To measure content depth, we begin by defining it precisely. Definition 1 (Content Depth). Content depth of a video is the degree to which it develops a topic from isolated information into structured understanding by explaining concepts, demonstrating procedures, analyzing mechanisms, forming evaluative judgments and generalizable insights. Based on Definition 1, we propose CDS, the first metric to quantify the content depth of short videos. A video receives a higher CDS when it presents richer semantic information, clearer explanations, and higher-order reasoning structures. Content depth is cognitively meaningful because deeper content offers viewers richer opportunities for higher-order cognitive processes (Anderson and Krathwohl 2001), which support cognitive development and intellectual growth. 3.2 Seven-level Scoring Rubric Theoretical Foundation. We operationalize CDS as a seven-level rubric grounded in Dual Process Theory (Wason and Evans 1974), the Revised Bloom’s Taxonomy22 2 For simplicity, we refer to it as “Bloom’s Taxonomy” hereafter.(Anderson and Krathwohl 2001), and the SOLO Taxonomy (Biggs and Collis 2014). Specifically, Dual Process Theory distinguishes between immediate affective and deliberate processing. Bloom’s Taxonomy organizes cognitive processes into six levels of increasing complexity, whereas the SOLO Taxonomy classifies the degree of understanding demonstrated by learners into five levels. Based on these complementary perspectives, we use System 1 in Dual Process Theory to describe the lowest level and the six cognitive levels of Bloom’s Taxonomy to define the remaining levels. We further use the SOLO Taxonomy as complementary guidance for distinguishing the complexity of understanding across levels. The resulting rubric ranges from immediate affective responses to increasingly complex cognitive processes. Rubric Structure. Table 1 presents the full rubric and its approximate theoretical anchors, grouping the seven levels into three groups. Low-CDS group (Level 0) represents content that mainly elicits affective or entertaining responses, with limited explicit knowledge value. Medium-CDS group (Levels 1–3) captures progressively more structured knowledge transmission, moving from isolated information to conceptual explanation and practical application. High-CDS group (Levels 4–6) reflects higher-order cognitive processes, ranging from analytical reasoning to evaluative judgment and the development of transferable models or frameworks. Detailed indicators and examples, together with the theoretical foundations and construction of the rubric, are provided in Appendices A and B, respectively. 4 Benchmark for CDS Evaluation Figure 2: Overview of the scalable CDS annotation workflow, which applies the rubric grounded in theories of cognition and learning. A gold-standard subset is used to select the LLM evaluator best aligned with human judgments for large-scale labeling. We introduce SCOPE-Bench, the first benchmark that supports content-depth assessment from two different perspectives. At the item level (Figure 2), SCOPE-Bench assesses the content depth of individual short videos from the user perspective. At the recommendation-list level (Figure 3), it evaluates the content depth of recommendation lists from the platform perspective, enabling platforms to assess RSs beyond conventional engagement-oriented performance. Dataset #Users #Items #Inter. Caption Image ASR Sparsity Full 10,000 153,561 1,007,746 100.00% 88.56% 88.56% 99.93% Sampled 6,654 31,496 128,105 100.00% 99.26% 99.26% 99.94% Note: The modality coverage is not complete because captions, visual features, and Automatic Speech Recognition (ASR) transcripts in the publicly released data are not uniformly available for all videos. Table 2: Statistics and modality coverage of the dataset. 4.1 Video Dataset We build SCOPE-Bench upon ShortVideo33 3 https://github.com/tsinghua-fib-lab/ShortVideo˙dataset (Shang et al. 2025), a publicly available short-video dataset containing user-item interactions and multimodal information. ShortVideo records approximately 1M chronologically ordered interactions generated by 10K real-world users over a one-week period. The dataset provides two versions, Full and Sampled, which cover the same data collection period. The Sampled version contains a subset of the users included in Full, although the sampling strategy used to construct this subset is not documented by the (Shang et al. 2025). We perform necessary preprocessing to improve data consistency and retain the content signals required for CDS assessment, as detailed in Appendix C. The statistics of the resulting Full and Sampled versions are summarized in Table 2. 4.2 Evaluation Protocol Our protocol adopts a content-based setting, where the video content is the primary object of CDS assessment. Instead of directly processing raw visual and audio streams, we use three textual signals as input, which capture complementary aspects of the video content: the caption c summarizes the main theme, the category label k provides topic-level context, and the ASR transcript a captures the spoken content. Given a video x=(c,k,a)x=(c,k,a) and the system prompt P, an LLM evaluator ℳM produces a structured output y=(s,ℓ,r,e,q)y=(s, ,r,e,q), comprising the CDS s, its level name ℓ , a reason r, supporting evidence e, and a confidence q. The score s is ordinal, taking values in 0,1,…,6,∅\0,1,…,6, \, where ∅ marks insufficient information for a reliable assessment. Because this setting relies on text, videos whose depth resides mainly in visual or auditory content, or that lack a usable transcript, receive ∅ . 4.3 Annotation Gold-Standard Subset Construction. To establish reliable reference labels for CDS assessment, we construct a gold-standard subset (gold set) through human annotation. Specifically, we conducted a human evaluation with five human evaluators from computer science and psychology. Following the Delphi method (Hong et al. 2019), the annotation process consists of three steps. In Step ①, the evaluators independently annotated the sampled videos according to scoring rubric based on video content. In Step ②, the human evaluators conducted a second round of annotation for videos showing substantial disagreement across evaluators. In Step ③, we aggregated the annotations using the median to reduce the influence of outlier judgments and improve the reliability of the resulting gold-standard labels. The gold set serves two purposes: validating interpretability and practical usability of our rubric and providing a human reference for selecting the LLM evaluator used in automatic annotation. Details about human evaluators, the annotation procedure, and the size of the gold set are provided in Appendix E. Automatic Scalable Annotation. To enable scalable annotation while maintaining alignment with human judgments, we first evaluate a set of candidate LLMs on the gold set. We compare their CDS predictions with human annotations using multiple agreement and error metrics, and select the model with the strongest overall alignment as the default evaluator. The evaluation results are reported in Section 5.2. We then apply the selected evaluator to the short-video dataset to generate CDS annotations. By augmenting the original dataset with annotations, we construct SCOPE-Bench dataset, which supports two complementary evaluation settings at both item and list levels for recommendation. Figure 3: Overview of the content-depth-aware RS evaluation framework for jointly assessing engagement and content depth. 4.4 List-Level Content-Depth Evaluation As illustrated in Figure 3, the list-level content-depth evaluation workflow evaluates the content depth of top-K recommendation lists. Specifically, we introduce List-wise Content Depth Score (LCDS), a complementary metric that aggregates the CDS of the top-K lists. Let si∈0,…,6s_i∈\0,…,6\44 4 We assign a score of 0 to items with si=∅s_i= . Such cases are primarily caused by insufficient information from ASR, so the resulting values should be interpreted as lower-bound estimation. denote the CDS of the video at rank i. LCDS@K is defined as: LCDS,β,@K _ α,β, w@K =(∑i=1Kwi[g(si)]β∑i=1Kwi)1/β, = ( _i=1^Kw_i [g_ α(s_i) ]^β _i=1^Kw_i )^1/β, (1) whereg(s) \;g_ α(s) =0,s=0,∑ℓ=1sαℓ∑r=16αr,s∈1,…,6. = cases0,&s=0,\\ _ =1^s _ _r=1^6 _r,&s∈\1,…,6\. cases (2) In Eq. (2), αℓ>0 _ >0 denotes the marginal value of moving from Level ℓ−1 -1 to Level ℓ , and g(s)∈[0,1]g_ α(s)∈[0,1] maps an ordinal CDS label to a normalized gain. w model position-dependent exposure and satisfy both wi≥0w_i≥ 0 and ∑i=1Kwi>0 _i=1^Kw_i>0. Let w¯i=wi/∑j=1Kwj w_i=w_i/ _j=1^Kw_j denote the normalized rank i exposure weight, and ℐ+:=i∈1,…,K∣wi>0I_+:= \i∈\1,…,K\ w_i>0 \. The aggregation behavior of LCDS can be characterized by: LCDSβ@K _β@K =∏i∈ℐ+g(si)w¯i,β→0+,∑i=1Kw¯ig(si),β=1,maxi∈ℐ+g(si),β→∞. = cases _i _+g_ α(s_i) w_i,&β→ 0^+,\\[5.69054pt] _i=1^K w_ig_ α(s_i),&β=1,\\[5.69054pt] _i _+g_ α(s_i),&β→∞. cases (3) Hence, larger values of β produce a more peak-oriented evaluation, whereas smaller values make LCDS more sensitive to low-CDS positions and therefore favor lists that sustain CDS. Following the grouping of the original CDS rubric, we further define three corresponding interpretive levels for LCDS. The thresholds are defined directly on the transformed scale as Low LCDS: [0,τL)[0, _L), Medium LCDS: [τL,τH)[ _L, _H), and High LCDS:[τH,1][ _H,1], where τL=g(1) _L=g_ α(1), τH=g(4) _H=g_ α(4). For the default setting, we assume equal marginal values, i.e., = α= 1, and use arithmetic aggregation with β=1β=1, yielding g(si)=si/6g(s_i)=s_i/6. Uniform rank weights (wi=1w_i=1) give Top-K Average LCDS (A-LCDS@K) definition as follows: A-LCDS@K=1K∑i=1Kg(si),A -LCDS@K= 1K _i=1^Kg(s_i), (4) which measures the average content depth of the recommendation list. Moreover, to account for greater exposure at higher ranks, we set wi=1/log2(i+1)w_i=1/ _2(i+1) followed by NDCG (Järvelin and Kekäläinen 2002), and define Top-K Exposure-weighted LCDS (E-LCDS@K) as follows: E-LCDS@K=∑i=1Kg(si)log2(i+1)∑j=1K1log2(j+1).E -LCDS@K= _i=1^K g(s_i) _2(i+1) _j=1^K 1 _2(j+1). (5) Both metrics lie in [0,1][0,1], with larger values indicating greater content depth of top-K recommendation lists. 5 Experiment (a) Agreement under score-difference tolerance. (b) Bootstrap distribution of Ordinal Krippendorff’s α. Figure 4: Human annotation reliability analysis, showing high agreement in their CDS annotations. 5.1 Human Annotation Reliability We assess the reliability of the independent human annotations collected during the construction of the gold set. Specifically, we analyze the Step ① annotations before conflict resolution, thereby evaluating whether different human evaluators can consistently apply the proposed CDS rubric without consensus-based adjustment, following prior evaluation practices (Wang et al. 2024; Han et al. 2025). As shown in Figure 4, the human evaluators exhibit a high level of agreement in their CDS annotations, which remains strong even after accounting for chance agreement. Overall, these results indicate that the proposed scoring rubric can be reliably applied by different human evaluators. Model Spearman↑ Kendall↑ Pearson↑ Exact↑ MAE↓ GLM-5.1 0.6658 0.6410 0.7663 71.67% 0.3337 GPT-5.5 0.6414 0.6108 0.7312 67.35% 0.4418 Gemini3.1-Pro 0.7270 0.7020 0.7683 76.71% 0.3097 Kimi-K2.6 0.7003 0.6714 0.7610 71.43% 0.3505 MiMo-V2.5-Pro 0.6356 0.6027 0.6944 67.35% 0.4538 Qwen3.7-Max 0.7177 0.6960 0.7934 78.03% 0.2605 Table 3: Comparison of LLM evaluators against human CDS judgments. ↑ and ↓ indicate that higher and lower values are better, respectively, and bold denotes the best result. 5.2 LLMs-Human Agreement To identify the most suitable LLMs for our evaluation protocol, we compare six leading models from the leaderboard55 5 https://artificialanalysis.ai/leaderboards/models: three open-weight models, Kimi-K2.6 (Moonshot AI 2026), MiMo-V2.5-Pro (Xiaomi MiMo Team 2026), and GLM-5.1 (Z.ai 2026); and three proprietary models, Gemini 3.1-Pro (Google DeepMind 2026), GPT-5.5 (OpenAI 2026), and Qwen3.7-Max (Qwen Team 2026). Following prior work (Liu et al. 2023; Kim et al. 2024; Ye et al. 2023), we evaluate all models on the same gold set and compare their scores sis_i with human annotations using Spearman’s ρ, Kendall’s τ, Pearson correlation, exact-match accuracy, and MAE. As shown in Table 3, Qwen3.7-Max achieves the strongest human alignment in three of the five metrics and is therefore adopted as the default evaluator. Most LLM models also correlate well with the human scores, suggesting that our protocol enables LLMs to reproduce judgments broadly shared by human evaluators. Additional agreement results are reported in Appendix F. 5.3 Analysis of CDS Figure 5: Distribution of CDS across all videos. Most videos fall into the low-CDS group, with few in the high-CDS group. CDS Distribution. Figure 5 shows that the majority of videos fall into the low-CDS or NaN group66 6 NaN cases mainly result from missing raw videos or ASR transcripts that are too short or noisy for reliable CDS assessment.. Overall, the distribution suggests that a large ratio of videos in short-video platforms imposes limited content depth. This observation is consistent with the attention-economy nature of platforms, where entertaining and attention-grabbing content is prevalent. Rank Low Medium High 1 Dance Health Finance 2 Comedy Law Military 3 Beauty Science History 4 Music History Law 5 Short Dramas Finance Science 6 Casual Videos Real Estate Information 7 Anime Digital Products Health Table 4: Top seven categories in each CDS group, ranked by proportions. Low-CDS categories differ from the others, while medium- and high-CDS categories largely overlap. Category Distribution. Table 4 presents the top categories across three CDS group. We observe that medium- and high-CDS videos are more frequently associated with knowledge-intensive categories, such as Finance, History, and Law. In contrast, low-CDS videos are more commonly associated with entertainment-oriented categories, such as Dance, Comedy, and Beauty. These category-level patterns are consistent with our design intuition of CDS: videos involving domain knowledge tend to receive higher CDS, whereas videos primarily designed for entertainment tend to receive lower CDS. See Appendix G for distribution details. Figure 6: Lexicon-based enrichment across CDS levels. Each cell reports the relative enrichment of signal words at CDS levels, with positive and negative values indicating values above and below the overall average, respectively. The smallest enrichment value in each row is shown in underlined. CDS Word Frequency. Figure 6 shows the enrichment patterns of different CDS levels based on our predefined theory-oriented lexicon77 7 We built the lexicon by extracting the top 500 words in gold subset and the top 200 words at each CDS level, grouping them into theoretical categories, and expanding each category.. Entertainment reaction and plot/dialogue show a decreasing trend as the CDS level rises, whereas most other categories exhibit the opposite trend. These findings are consistent with our expectations: low-CDS videos are more likely to focus on entertainment-oriented reactions and plot-level descriptions, whereas high-CDS videos are more likely to contain signals related to explanation, evidence evaluation, generalization, transfer, and conceptualized expression. We also provide details representative words for each CDS level in Appendix H. Variable Association Explained variance Caption length Spearman ρ=0.0828∗ρ=0.0828^* R2=0.57%R^2=0.57\% Log-ASR lengtha Spearman ρ=0.2182∗ρ=0.2182^* R2=6.09%R^2=6.09\% Category Cramér V=0.2469∗V=0.2469^* η2=25.34%η^2=25.34\% a ASR length |a||a| exhibits a long-tailed distribution. Therefore, we use log-transformed form log(1+|a|) (1+|a|) to estimate the relationship. • ∗ denotes statistical significance at p<0.001p<0.001. Table 5: Associations between valid CDS (s≠∅s≠ ) and surface-level input attributes. CDS is more strongly associated with ASR length and category than with caption length. Surface-Level Correlates of CDS. Table 5 summarizes the relationships between valid CDS (s≠∅s≠ ) and three surface-level input attributes: Caption length, ASR length, and category. The results show that caption length has only a weak association with CDS, whereas ASR length and category exhibit stronger associations. This pattern is consistent with high-CDS content requiring sufficient textual space to express explanations, procedures, and reasoning structures. Category differences may similarly arise from their inherent content orientation, with entertainment-oriented and knowledge-oriented categories tending to receive lower and higher CDS scores, respectively, as shown in Table 4. 5.4 Evaluation Protocol Robustness In Section 5.3, we observe that CDS exhibit certain correlations with video category and ASR length. These observations naturally raise two robustness concerns: Whether the evaluation protocol has category prior bias and verbosity bias. To examine these issues, we conduct paired counterfactual robustness tests (Zheng et al. 2023), where only one input field is perturbed at a time. For each video i, we compare the original CDS score sis_i and ASR length |ai||a_i| with their counterfactual values si′s_i and |ai′||a_i |. We define the absolute score change and relative ASR length change as Δsi=|si′−si| s_i=|s_i -s_i| and δ|ai|=(|ai′|−|ai|)/|ai|δ|a_i|=(|a_i |-|a_i|)/|a_i|, respectively. Perturbation N P(Δs=1 s=1) P(Δs=2 s=2) P(Δs>2 s>2) Low → Medium 50 0.00% 0.00% 0.00% Low → High 50 0.00% 0.00% 0.00% Medium → Low 60 6.67% 3.33% 0.00% High → Low 46 13.04% 13.04% 0.00% Table 6: Category counterfactual perturbation results measured by CDS shift magnitude. Source→TargetSource indicates perturbing videos from the Source CDS group to Target-associated categories. Category perturbations leave low-CDS videos unchanged and cause only limited one- or two-level shifts for medium- and high-CDS videos. Robustness to Category Prior Bias. We examine whether the evaluation protocol relies on category-level shortcuts by overemphasizing category labels while underutilizing ASR transcripts, which provide more direct evidence of a video’s semantic content. As shown in Table 6, changing low-CDS videos to either medium- or high-CDS-associated categories does not change their CDS. For the reverse direction, mapping high- or medium-CDS videos to low-CDS-associated categories may introduce slight downward shifts, with more pronounced changes observed for high-CDS videos. This is consistent with our conservative scoring principle: when the available evidence does not fully support a higher CDS level, the evaluator tends to assign a lower score. Perturbation Group N δ|a|δ|a| P(Δs=1 s=1) P(Δs=2 s=2) P(Δs>2 s>2) ASR length ↑ Low 50 +52.05% 2.00% 0.00% 0.00% Medium 60 +51.39% 5.00% 0.00% 0.00% ASR length ↓ Medium 60 -24.49% 8.33% 0.00% 0.00% High 46 -27.35% 23.91% 15.22% 0.00% Table 7: ASR counterfactual perturbation results measured by CDS shifts. Δsi s_i and δ|ai|δ|a_i| denote the absolute CDS change and relative ASR length change, respectively. CDS remains largely stable under lengthening and changes modestly under shortening, with no shifts exceeding two levels. Robustness to Verbosity Bias. We examine whether the evaluation protocol favors longer ASR transcripts without additional meaningful information. Table 7 shows that meaningless length expansion has minimal impact, whereas transcript shortening affects high-CDS videos more than medium-CDS videos, with no CDS change exceeding two levels. This sensitivity is consistent with our conservative scoring principle, as transcript compression may weaken the reasoning evidence required for high CDS levels, including argument logic and evidence connections. Overall, these results suggest that our evaluation protocol primarily responds to meaningful informational and reasoning structures in the videos, rather than superficial cues such as category labels or ASR transcript length. Details and stability experiment are provided in Appendix I. 6 SCOPE-Bench Leaderboard Method ShortVideoSampled ShortVideoFull R@20↑ A@20↑ E@20↑ R@20↑ A@20↑ E@20↑ ID-based Recommendation BPR 3.30 6.90 6.85 2.22 7.06 7.10 NCF 3.01 6.49 6.42 2.09 7.46 7.53 LightGCN 3.54 7.36 7.33 2.37 6.89 7.08 Multimodal Recommendation VBPR 2.73 6.22 5.70 1.80 7.44 7.95 GRCN 2.83 7.92 7.85 1.64 6.72 6.73 LATTICE 2.70 7.13 7.04 2.25 6.97 7.10 BM3 2.88 6.72 6.66 2.41 6.70 6.80 FREEDOM 3.30 7.13 7.30 2.51 6.81 6.84 MGCN 3.82 7.38 7.58 2.65 6.89 6.99 LGMRec 3.40 6.71 6.63 2.42 7.24 7.52 DiffMM 3.24 7.24 7.35 2.39 7.10 7.18 REARM 3.76 7.71 7.76 2.42 6.81 6.93 FITMM 3.25 8.12 8.35 2.41 7.59 7.88 Random 0.08 6.90 6.90 0.02 6.59 6.59 Table 8: Engagement and content-depth performance of baselines on two datasets. All metrics are scaled by ×100× 100, where R denotes Recall, A/E denote A-LCDS/E-LCDS, and the best results are bolded. Baselines are competitive on engagement but remain in the low-LCDS range and close to random recommendation on content-depth metric. We evaluate 13 baselines on SCOPE-Bench, including three ID-based methods, BPR (Rendle et al. 2009), NCF (He et al. 2017), and LightGCN (He et al. 2020), and ten multimodal methods, VBPR (He and McAuley 2016), GRCN (Wei et al. 2020), LATTICE (Zhang et al. 2021), BM3 (Zhou et al. 2023), FREEDOM (Zhou and Shen 2023), MGCN (Yu et al. 2023), LGMRec (Guo et al. 2024), DiffMM (Jiang et al. 2024), REARM (Ma et al. 2025), and FITMM (Yang et al. 2025). The Random baseline reports the expected performance of uniformly sampling items from each candidate set. We use an 8:1:1 training-validation-test split (Shang et al. 2025). As shown in Table 8, existing methods achieve competitive performance on engagement metrics, whereas their A-LCDS and E-LCDS scores consistently remain within the Low-LCDS range, i.e, [0,1/6)[0,1/6). Moreover, most methods perform close to the Random baseline. These results indicate that stronger engagement performance does not necessarily translate into the recommendation of content with greater depth. Appendix C presents the experimental setup, results under an alternative treatment of si=∅s_i= , training trajectories of engagement and content-depth metrics, and CDS-aware optimization. These results further demonstrate that engagement and content depth are currently decoupled. 7 Conclusion In this paper, we propose a new metric, termed CDS, to measure the content depth of short videos. CDS provides a principled basis for evaluating and optimizing content depth in existing RSs. To comprehensively evaluate this dimension, we construct SCOPE-Bench, the first benchmark that supports both item- and list-level content-depth evaluation. Empirical results demonstrate the interpretability, practical utility, and robustness of the proposed evaluation framework. Our experiments further reveal that existing RSs tend to favor low-depth videos, and remain limited in recommending videos with high CDS. Hereby, CDS and SCOPE-Bench establish a new evaluation axis for developing short-video RSs that jointly consider content depth and user engagement. References Allen et al. (1955) Allen, K.; Berry, M. M.; Luehrs Jr, F. U.; and Perry, J. W. 1955. Machine literature searching VIII. Operational criteria for designing information retrieval systems. American documentation (pre-1986), 6(2): 93. Anderson and Krathwohl (2001) Anderson, L. W.; and Krathwohl, D. R. 2001. A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives: complete edition. Addison Wesley Longman, Inc. Australian Government (2025) Australian Government. 2025. Social media minimum age. https://w.infrastructure.gov.au/media-communications/internet/online-safety/social-media-minimum-age. Barzilay and Lapata (2008) Barzilay, R.; and Lapata, M. 2008. Modeling local coherence: An entity-based approach. Computational Linguistics, 34(1): 1–34. Biggs and Collis (2014) Biggs, J. B.; and Collis, K. F. 2014. Evaluating the quality of learning: The SOLO taxonomy (Structure of the Observed Learning Outcome). Academic press. Chi and Wylie (2014) Chi, M. T.; and Wylie, R. 2014. The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational psychologist, 49(4): 219–243. García-Rapp (2017) García-Rapp, F. 2017. Popularity markers on YouTube’s attention economy: the case of Bubzbeauty. Celebrity studies, 8(2): 228–245. Ge, Delgado-Battenfeld, and Jannach (2010) Ge, M.; Delgado-Battenfeld, C.; and Jannach, D. 2010. Beyond accuracy: Evaluating recommender systems by coverage and serendipity. In Proceedings of the Fourth ACM Conference on Recommender Systems, 257–260. Gehman et al. (2020) Gehman, S.; Gururangan, S.; Sap, M.; Choi, Y.; and Smith, N. A. 2020. Realtoxicityprompts: Evaluating neural toxic degeneration in language models. In Findings of the association for computational linguistics: EMNLP 2020, 3356–3369. Google DeepMind (2026) Google DeepMind. 2026. Gemini 3.1 Pro. https://deepmind.google/models/model-cards/gemini-3-1-pro/. Guo et al. (2024) Guo, Z.; Li, J.; Li, G.; Wang, C.; Shi, S.; and Ruan, B. 2024. Lgmrec: Local and global graph learning for multimodal recommendation. In AAAI, 8454–8462. Han et al. (2025) Han, H.; Li, S.; Chen, J.; Yuan, Y.; Wu, Y.; Deng, Y.; Leong, C. T.; Du, H.; Fu, J.; Li, Y.; et al. 2025. Video-bench: Human-aligned video generation benchmark. In Proceedings of the Computer Vision and Pattern Recognition Conference, 18858–18868. He and McAuley (2016) He, R.; and McAuley, J. 2016. VBPR: visual bayesian personalized ranking from implicit feedback. In AAAI, volume 30. He et al. (2020) He, X.; Deng, K.; Wang, X.; Li, Y.; Zhang, Y.; and Wang, M. 2020. Lightgcn: Simplifying and powering graph convolution network for recommendation. In Proceedings of the 43rd International ACM SIGIR conference on research and development in Information Retrieval, 639–648. He et al. (2017) He, X.; Liao, L.; Zhang, H.; Nie, L.; Hu, X.; and Chua, T.-S. 2017. Neural collaborative filtering. In Proceedings of the 26th international conference on world wide web, 173–182. Herlocker et al. (2004) Herlocker, J. L.; Konstan, J. A.; Terveen, L. G.; and Riedl, J. T. 2004. Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems, 22(1): 5–53. Heusel et al. (2017) Heusel, M.; Ramsauer, H.; Unterthiner, T.; Nessler, B.; and Hochreiter, S. 2017. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30. Hijikata, Kai, and Nishida (2012) Hijikata, Y.; Kai, Y.; and Nishida, S. 2012. The relation between user intervention and user satisfaction for information recommendation. In Proceedings of the 27th Annual ACM Symposium on Applied Computing, 2002–2007. Hong et al. (2019) Hong, Q. N.; Pluye, P.; Fàbregues, S.; Bartlett, G.; Boardman, F.; Cargo, M.; Dagenais, P.; Gagnon, M.-P.; Griffiths, F.; Nicolau, B.; et al. 2019. Improving the content validity of the mixed methods appraisal tool: a modified e-Delphi study. Journal of clinical epidemiology, 111: 49–59. Järvelin and Kekäläinen (2002) Järvelin, K.; and Kekäläinen, J. 2002. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems (TOIS), 20(4): 422–446. Jiang et al. (2024) Jiang, Y.; Xia, L.; Wei, W.; Luo, D.; Lin, K.; and Huang, C. 2024. DiffMM: Multi-Modal Diffusion Model for Recommendation. arXiv preprint arXiv:2406.11781. Keenan (2023) Keenan, C. 2023. New Features for Teens and Families on TikTok. https://newsroom.tiktok.com/new-features-for-teens-and-families-on-tiktok-au?lang=en-AU. Kemp (2025) Kemp, S. 2025. Digital 2025 July Global Statshot Report. https://datareportal.com/reports/digital-2025-july-global-statshot. Kim et al. (2024) Kim, S.; Shin, J.; Jang, J.; Longpre, S.; Lee, H.; Yun, S.; Shin, R.; Kim, S.; Thorne, J.; Seo, M.; et al. 2024. Prometheus: Inducing fine-grained evaluation capability in language models. In International Conference on Learning Representations, volume 2024, 29927–29962. Knijnenburg et al. (2012) Knijnenburg, B. P.; Willemsen, M. C.; Gantner, Z.; Soncu, H.; and Newell, C. 2012. Explaining the user experience of recommender systems. User Modeling and User-Adapted Interaction, 22(4): 441–504. Kuaishou Technology (2025) Kuaishou Technology. 2025. Kuaishou Technology Announces Third Quarter 2025 Unaudited Financial Results. https://ir.kuaishou.com/news-releases/news-release-details/kuaishou-technology-announces-third-quarter-2025-unaudited/. Li et al. (2026) Li, Z.; Long, G.; Jiang, J.; Zhang, C.; and Yang, Q. 2026. Federated Vision-Language-Recommendation with Personalized Fusion. In AAAI, volume 40, 23337–23345. Lin, Hilton, and Evans (2022) Lin, S.; Hilton, J.; and Evans, O. 2022. Truthfulqa: Measuring how models mimic human falsehoods. In Proceedings of the 60th annual meeting of the association for computational linguistics (volume 1: long papers), 3214–3252. Liu et al. (2023) Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; and Zhu, C. 2023. G-eval: NLG evaluation using gpt-4 with better human alignment. In Proceedings of the 2023 conference on empirical methods in natural language processing, 2511–2522. Ma et al. (2025) Ma, S.; Zeng, Y.; Wu, S.; and Xu, G. 2025. Refining Contrastive Learning and Homography Relations for Multi-Modal Recommendation. In Proceedings of the 33rd ACM International Conference on Multimedia, 6316–6324. Mahakud and Thapliyal (2026) Mahakud, G. C.; and Thapliyal, A. 2026. Assessing the Impact of Different Social Media Video Content Formats on Sustained Attention and Working Memory. Annals of Neurosciences, 09727531261424994. Moonshot AI (2026) Moonshot AI. 2026. Meet Kimi K2.6: Advancing Open-Source Coding. https://forum.moonshot.ai/t/meet-kimi-k2-6-advancing-open-source-coding/369. Napoles, Sakaguchi, and Tetreault (2017) Napoles, C.; Sakaguchi, K.; and Tetreault, J. 2017. JFLEG: A fluency corpus and benchmark for grammatical error correction. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, 229–234. OpenAI (2026) OpenAI. 2026. GPT-5.5 System Card. https://openai.com/index/gpt-5-5-system-card/. Ørmen and Gregersen (2023) Ørmen, J.; and Gregersen, A. 2023. Towards the engagement economy: interconnected processes of commodification on YouTube. Media, Culture & Society, 45(2): 225–245. Pu, Chen, and Hu (2011) Pu, P.; Chen, L.; and Hu, R. 2011. A user-centric evaluation framework for recommender systems. In Proceedings of the Fifth ACM Conference on Recommender Systems, 157–164. Qwen Team (2026) Qwen Team. 2026. Qwen3.7: The Agent Frontier. https://qwen.ai/blog?id=qwen3.7. Radford et al. (2021) Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. 2021. Learning transferable visual models from natural language supervision. In International conference on machine learning, 8748–8763. PmLR. Rendle et al. (2009) Rendle, S.; Freudenthaler, C.; Gantner, Z.; and Schmidt-Thieme, L. 2009. BPR: Bayesian Personalized Ranking from Implicit Feedback. In Proceedings of the Twenty-Fifth Conference on Uncertainty in Artificial Intelligence, 452–461. AUAI Press. Shang et al. (2025) Shang, Y.; Gao, C.; Li, N.; and Li, Y. 2025. A Large-scale Dataset with Behavior, Attributes, and Content of Mobile Short-video Platform. In Companion Proceedings of the ACM on Web Conference 2025, 793–796. New York, NY, USA: Association for Computing Machinery. Sternberg, Sternberg, and Mio (2006) Sternberg, R. J.; Sternberg, K.; and Mio, J. 2006. Cognitive psychology. Thomson/Wadsworth Belmont, CA. Tang et al. (2026) Tang, D.; Zhang, X.; Gou, P.; Feng, J.; Hu, R.; and Sum, K.-w. R. 2026. Association between short-form video use and mental health: systematic review and Meta-analysis. Journal of Medical Internet Research, 28: e82503. TikTok Team (2021) TikTok Team. 2021. Thanks a Billion! https://newsroom.tiktok.com/1-billion-people-on-tiktok?lang=en. UK Government (2026) UK Government. 2026. Fact sheet: New rules to protect children online. https://w.gov.uk/government/publications/fact-sheet-new-rules-to-protect-children-online. Vargas and Castells (2011) Vargas, S.; and Castells, P. 2011. Rank and relevance in novelty and diversity metrics for recommender systems. In Proceedings of the Fifth ACM Conference on Recommender Systems, 109–116. Wang et al. (2024) Wang, Y.; Yu, Z.; Yao, W.; Zeng, Z.; Yang, L.; Wang, C.; Chen, H.; Jiang, C.; Xie, R.; Wang, J.; et al. 2024. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization. In International Conference on Learning Representations, volume 2024, 43573–43593. Wang et al. (2004) Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image quality assessment: from error visibility to structural similarity. TIP, 13(4): 600–612. Wang, Lu, and Bovik (2004) Wang, Z.; Lu, L.; and Bovik, A. C. 2004. Video Quality Assessment Based on Structural Distortion Measurement. Signal Processing: Image Communication, 19(2): 121–132. Wason and Evans (1974) Wason, P. C.; and Evans, J. S. B. 1974. Dual processes in reasoning? Cognition, 3(2): 141–154. Wei et al. (2020) Wei, Y.; Wang, X.; Nie, L.; He, X.; and Chua, T.-S. 2020. Graph-refined convolutional network for multimedia recommendation with implicit feedback. In Proceedings of the 28th ACM international conference on multimedia, 3541–3549. Xiaomi MiMo Team (2026) Xiaomi MiMo Team. 2026. Xiaomi MiMo-V2.5-Pro. https://mimo.xiaomi.com/mimo-v2-5-pro/. Yang et al. (2025) Yang, W.; Zhong, R.; Chen, Y.; Li, S.; Ping, H.; Lu, C.; and Jiang, P. 2025. FITMM: Adaptive Frequency-Aware Multimodal Recommendation via Information-Theoretic Representation Learning. In Proceedings of the 33rd ACM International Conference on Multimedia, 6193–6202. Ye et al. (2023) Ye, S.; Kim, D.; Kim, S.; Hwang, H.; Kim, S.; Jo, Y.; Thorne, J.; Kim, J.; and Seo, M. 2023. Flask: Fine-grained language model evaluation based on alignment skill sets. arXiv preprint arXiv:2307.10928. Yu et al. (2023) Yu, P.; Tan, Z.; Lu, G.; and Bao, B.-K. 2023. Multi-view graph convolutional network for multimedia recommendation. In Proceedings of the 31st ACM international conference on multimedia, 6576–6585. Z.ai (2026) Z.ai. 2026. GLM-5.1: Towards Long-Horizon Tasks. https://z.ai/blog/glm-5.1. Zangerle and Bauer (2022) Zangerle, E.; and Bauer, C. 2022. Evaluating recommender systems: survey and framework. ACM computing surveys, 55(8): 1–38. Zhang et al. (2021) Zhang, J.; Zhu, Y.; Liu, Q.; Wu, S.; Wang, S.; and Wang, L. 2021. Mining latent structures for multimedia recommendation. In Proceedings of the 29th ACM international conference on multimedia, 3872–3880. Zheng et al. (2023) Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in neural information processing systems, 36: 46595–46623. Zhou and Shen (2023) Zhou, X.; and Shen, Z. 2023. A tale of two graphs: Freezing and denoising graph structures for multimodal recommendation. In Proceedings of the 31st ACM international conference on multimedia, 935–943. Zhou et al. (2023) Zhou, X.; Zhou, H.; Liu, Y.; Zeng, Z.; Miao, C.; Wang, P.; You, Y.; and Jiang, F. 2023. Bootstrap latent representations for multi-modal recommendation. In Proceedings of the ACM web conference 2023, 845–854. Ziegler et al. (2005) Ziegler, C.-N.; McNee, S. M.; Konstan, J. A.; and Lausen, G. 2005. Improving recommendation lists through topic diversification. In Proceedings of the 14th International Conference on World Wide Web, 22–32. Zou and Sun (2025) Zou, K.; and Sun, A. 2025. A Survey of Real-World Recommender Systems: Challenges, Constraints, and Industrial Perspectives. arXiv preprint arXiv:2509.06002.