Paper deep dive
Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads
Kunpeng Zhang, Poppy Zhang, Shawndra Hill, Amel Awadelkarim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/20/2026, 7:50:44 PM
Summary
This paper introduces MLLM-VAU, a framework for analyzing the 'hooking period' (first three seconds) of video ads using multimodal large language models. It employs two frame sampling strategies (uniform random and key frame selection) to extract visual features, processes audio attributes, and uses BERTopic for topic modeling of MLLM-generated descriptions. The framework correlates these multimodal features with performance metrics like conversion per investment (CPI) and engagement rates, demonstrating improved predictive power and interpretability compared to traditional black-box methods.
Entities (9)
Relation Signals (7)
MLLM-VAU → uses → Multimodal Large Language Models
confidence 98% · This study presents a framework using transformer-based multimodal large language models (MLLMs) to analyze the hooking period of video ads.
MLLM-VAU → uses → BERTopic
confidence 96% · which are distilled into coherent topics using BERTopic for high-level abstraction.
MLLM-VAU → employs → Key Frame Selection
confidence 95% · It tests two frame sampling strategies, uniform random sampling and key frame selection
MLLM-VAU → employs → Uniform Random Sampling
confidence 95% · It tests two frame sampling strategies, uniform random sampling and key frame selection
Hooking Period → influences → Conversion per Investment
confidence 94% · revealing correlations between hooking period features and key performance metrics like conversion per investment.
Llama Multimodal Model → isusedby → MLLM-VAU
confidence 92% · This employs Llama Multimodal Model 1 , which leverages the advanced capabilities of transformer-based architectures
Key Frame Selection → uses → Structural Similarity Index Measure
confidence 90% · Calculating the difference... using a suitable metric, such as the Structural Similarity Index Measure (SSIM).
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Video-based ads are a vital medium for brands to engage consumers, with social media platforms leveraging user data to optimize ad delivery and boost engagement. A crucial but under-explored aspect is the 'hooking period', the first three seconds that capture viewer attention and influence engagement metrics. Analyzing this brief window is challenging due to the multimodal nature of video content, which blends visual, auditory, and textual elements. Traditional methods often miss the nuanced interplay of these components, requiring advanced frameworks for thorough evaluation. This study presents a framework using transformer-based multimodal large language models (MLLMs) to analyze the hooking period of video ads. It tests two frame sampling strategies, uniform random sampling and key frame selection, to ensure balanced and representative acoustic feature extraction, capturing the full range of design elements. The hooking video is processed by state-of-the-art MLLMs to generate descriptive analyses of the ad's initial impact, which are distilled into coherent topics using BERTopic for high-level abstraction. The framework also integrates features such as audio attributes and aggregated ad targeting information, enriching the feature set for further analysis. Empirical validation on large-scale real-world data from social media platforms demonstrates the efficacy of our framework, revealing correlations between hooking period features and key performance metrics like conversion per investment. The results highlight the practical applicability and predictive power of the approach, offering valuable insights for optimizing video ad strategies. This study advances video ad analysis by providing a scalable methodology for understanding and enhancing the initial moments of video advertisements.
Tags
Links
- Source: https://arxiv.org/abs/2602.22299v1
- Canonical: https://arxiv.org/abs/2602.22299v1
Trouble viewing inline? Open PDF directly →
Full Text
53,678 characters extracted from source content.
Expand or collapse full text
Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video Ads Kunpeng Zhang kpzhang@umd.edu University of Marland, College Park College Park, Maryland, USA Poppy Zhang poppyzhang@meta.com Meta Platforms, Inc. New York, USA Shawndra Hill shawndrahill@meta.com Meta Platforms, Inc. New York, USA Amel Awadelkarim ameloa@meta.com Meta Platforms, Inc. California, USA Abstract Video-based advertisements have become an important medium for brands to engage consumers, with social media platforms lever- aging extensive user data to optimize ad delivery and enhance engagement. An under-explored aspect of video ad effectiveness is the initial “hooking period" — the first three seconds that capture viewer attention and influence subsequent engagement metrics. Analyzing the factors that drive performance during this brief time frame is challenging due to the multimodal nature of video content, which integrates visual, auditory, and textual elements. Traditional analysis methods often fall short in capturing the nuanced inter- play of these components, necessitating advanced frameworks for comprehensive evaluation. This study introduces a framework that employs transformer- based multimodal large language models (MLLMs) to dissect the hooking period of video advertisements. It tests two different frame sampling strategies — uniform random sampling and key frame selection — to ensure a balanced and representative acoustic fea- ture extraction, capturing the full spectrum of design elements. The hooking video is processed by state-of-the-art MLLMs to generate descriptive analyses of the ad’s initial impact, which are then dis- tilled into coherent topics using BERTopic for high-level abstraction. Additionally, the framework integrates additional features such as audio attributes and aggregated ad targeting information, enriching the feature set for subsequent analysis. Empirical validation on large-scale real-world data from social media platforms demonstrates the efficacy of our framework, re- vealing correlations between hooking period features and the key performance metrics like conversion per investment. Our results highlight the practical applicability and predictive power of the proposed approach, offering valuable insights for optimizing video ad strategies. This study advances the state-of-the-art in video ad analysis by providing a comprehensive and scalable methodology Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Conference, USA © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-X-X/2018/06 https://doi.org/X.X for understanding and enhancing the initial moments of video ad- vertisements. Keywords Video ads, Multimodal LLM, Hooking period, Conversion, Feature extraction 1 Introduction In the rapidly evolving digital landscape, video-based advertise- ments have emerged as a pivotal medium for brands to engage consumers. Social media platforms have capitalized on this trend, leveraging vast amounts of user data to optimize ad delivery and enhance user engagement [31]. Understanding the elements that contribute to the effectiveness of video ads is paramount for both advertisers and platform providers. Specifically, the initial moments of an advertisement—the so-called "hooking period" comprising the first three seconds—are important in capturing viewer atten- tion and influencing subsequent engagement metrics. Despite the significance of this brief yet impactful timeframe, comprehensively analyzing the factors that drive ad performance during the hooking period remains a formidable challenge. Video advertisements are a cornerstone of digital marketing strategies, offering dynamic and immersive experiences that static ads cannot match [19]. The ability to convey compelling narratives, evoke emotional responses, and showcase products or services within a limited timeframe makes video ads exceptionally effective [18]. For social media platforms, which facilitate a considerable amount of ad impressions daily, optimizing ad performance directly correlates with user satisfaction and platform growth [24]. The hooking period is particularly influential, as it determines whether a viewer continues watching the ad or scrolls past it. A well-crafted hooking period can significantly enhance metrics such as impres- sions, conversion per investment (CPI, investment referring to ad- vertisement budget), and user engagement rates, thereby delivering substantial returns on advertising investments [29]. However, the complexity of video content, which encompasses visual elements, audio cues, and temporal dynamics, poses signif- icant challenges for analysis. Traditional methods often rely on manual annotation or simplistic feature extraction techniques that fail to capture the nuanced interplay of multimodal elements within the hooking period [5]. Consequently, there is a pressing need for advanced analytical frameworks that can automatically extract and arXiv:2602.22299v1 [cs.M] 25 Feb 2026 Conference, 2026, USAZhang et al. interpret the important features of video ads, thereby informing the design and optimization of more effective advertising strategies. The task of dissecting and understanding the hooking period of video ads involves several intricate challenges. First, the multimodal nature of video content—integrating visual, auditory, and textual information—necessitates sophisticated models capable of process- ing and interpreting diverse data types simultaneously. Traditional machine learning approaches may struggle with this complexity, of- ten requiring separate processing pipelines for different modalities, which can lead to fragmented and less coherent feature represen- tations [3]. Second, existing methodologies in video ad analysis predominantly focus on either broad content classification or super- ficial feature extraction [15], inadequately addressing the specific dynamics of the initial three seconds. These methods often overlook the subtle yet impactful design elements that characterize success- ful hooking periods, such as emotional appeal, visual aesthetics, interactivity, and challenges posed to the viewer [8]. Moreover, the temporal aspect of video content, where the sequence and timing of elements play an important role in viewer retention, is frequently under-explored [30]. Another significant challenge lies in linking the extracted features of the hooking period to concrete perfor- mance metrics. While various studies have attempted to correlate ad attributes with outcomes like impressions and engagement rates, establishing a robust and predictive relationship remains elusive. This gap is exacerbated by the sheer volume of ad data, which demands scalable and efficient analytical frameworks capable of handling large-scale datasets without compromising on the granu- larity of insights. Addressing these challenges necessitates a multifaceted approach that leverages the latest advancements in machine learning, partic- ularly in the realm of multimodal large language models (MLLMs) [17]. Our proposed framework harnesses the power of transformer- based architectures [28] to comprehensively analyze the hooking period of video ads, extracting and interpreting key features that drive performance metrics. The process begins with the conversion of each video ad into a sequence of frames, employing two distinct sampling strategies: uniform random sampling and key frame se- lection [26]. Uniform random sampling ensures a representative distribution of frames across the entire hooking period, while key frame selection targets frames that encapsulate significant visual or narrative shifts. This dual approach facilitates a more nuanced understanding of the temporal dynamics within the hooking pe- riod. Subsequently, the extracted frames are fed into state-of-the-art MLLM, which are adept at processing and interpreting multimodal data. These models generate text-based reasoning descriptions that encapsulate the design methodologies employed in the hooking pe- riod. For instance, the model may identify elements related to emo- tional appeal, visual aesthetics, interactivity, or viewer challenges. Such descriptive analyses provide a structured representation of the ad’s initial impact, facilitating deeper insights into the factors that influence viewer engagement. To further distill the rich textual descriptions generated by the MLLMs, we employ BERTopic [14], a topic modeling technique that leverages transformer-based em- beddings to identify coherent topics within large text corpora. By summarizing the methodology descriptions into a few salient topics, we obtain a high-level overview of the prevalent design strategies within the hooking period. This topic-level abstraction enables the integration of qualitative insights with quantitative performance metrics. In addition to the multimodal features extracted from the hooking period, our framework incorporates auxiliary features such as audio attributes and aggregated-level ad information (e.g., tar- geting, ad placement). By synthesizing these diverse data sources, we construct a comprehensive feature set that encapsulates both the intrinsic qualities of the ad content and the contextual factors influencing its performance. Finally, we employ predictive model- ing techniques to establish the relationships between the extracted features and key performance metrics, including impressions, con- version per investment, and engagement rates. By training models on this enriched feature set, we aim to achieve improved predic- tive performance, thereby enabling more accurate forecasting of ad success based on the characteristics of the hooking period. In summary, this paper presents a comprehensive and innovative approach to understanding the hooking period of video ads through the lens of multimodal large language models. By addressing the existing research gaps and introducing novel analytical techniques, we contribute valuable insights and methodologies that advance the state-of-the-art in video ad analysis and optimization. Specifically, we contribute to the literature as follows. (1)Innovative Multimodal Analysis Framework: We introduce a framework that leverages multimodal large language mod- els to extract and interpret key features from the hooking period of video ads. This approach effectively integrates vi- sual, auditory, and textual data, providing a comprehensive understanding of the elements that drive ad performance. (2)Frame Sampling Strategies for Acoustic Feature Extraction: By testing both uniform random sampling and key frame selection, our method ensures a balanced and representative extraction of frames, capturing the full spectrum of design elements within the hooking period. This enhances the ro- bustness and granularity of the feature extraction process. (3) Integration of Auxiliary Features: Our framework seamlessly incorporates additional features such as audio attributes and aggregated ad information, enriching the feature set and en- hancing the predictive capabilities of the model. This holistic approach ensures that both content-specific and contextual factors are accounted for in the analysis. (4)Empirical Validation on Real-World Data: Utilizing five cate- gories of real-world data from a social platform, we demon- strate the efficacy of our interpretable framework through extensive experiments. The results reveal insightful findings regarding the relationship between hooking period features and ad performance metrics, showcasing the practical appli- cability and predictive power of our approach. 2 Related Work We review the relevant research from the following two perspec- tives, including the use of multi-modal large language models for video understanding and ads performance prediction. Video understanding Recent advances in video understand- ing have significantly benefited from deep learning techniques, particularly leveraging architectures such as convolutional neural networks (CNNs) and Transformers. Early methods include extend- ing 2D CNNs into 3D variants, such as C3D [27] and I3D [6], to Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video AdsConference, 2026, USA effectively capture temporal information within videos, thereby achieving significant improvements in tasks like video classifica- tion and event detection. More recently, Transformer-based mod- els, including Video Vision Transformer (ViViT) [1] and TimeS- former [4], have further advanced the state-of-the-art by explicitly modeling spatio-temporal relationships using self-attention mech- anisms. These approaches excel at modeling long-range depen- dencies across video frames, enhancing their capability to capture complex semantic contexts. However, despite these advancements, accurately interpreting content effectiveness in video advertisements remains challenging, largely due to multimodal dynamics involving visual, audio, and textual cues [12,22]. Recent studies have explored multimodal fu- sion methods that integrate different modalities to better predict ad performance metrics. For example, multimodal Transformer-based models jointly encoding visual, audio, and textual features have demonstrated promising results in unerstanding viewer reactions and improving content personalization [11,20]. To sum up, opti- mizing such models for video ad performance prediction remains an active research area, particularly in identifying effective meth- ods for integrating multimodal data and interpreting the complex interactions among multiple modalities. Ads performance prediction Ad performance prediction using images or videos has gained considerable attention in recent years due to the rich content embedded in these mediums. Traditional methods of predicting ad performance mainly relied on text-based features and user interactions, but the inclusion of images or videos allows for a deeper analysis of the content itself. Researchers have leveraged advanced deep learning techniques, particularly Convo- lutional Neural Networks (CNNs) and Recurrent Neural Networks (RNNs), to successfully extract features from visual, audio, and tex- tual data that are correlated with key performance metrics such as click-through rates (CTR), conversion rates, and overall user engagement [7]. They include color, objects, and facial expressions in images, motion and scene changes in videos, as well as text-based features. Moreover, recent advancements in multimodal learning have enabled the integration of visual, audio, and textual data for even more accurate performance forecasting [13]. Such a cross- modal approach captures the complex relationships between visual appeal and user responses, considering factors like emotional tone, product presentation, and visual aesthetics [16]. As digital advertis- ing continues to evlove, these image- and video-based prediction models are expected to play a key role in optimizing ad strategies, ensuring that marketers can more effectively target their audiences and maximize return on investment (ROI) [21]. 3 Methodology In this section, we present a detailed description of our proposed method, MLLM-VAU (Multimodal LLM-based Video Ad Under- standing). The overall framework of MLLM-VAU is illustrated in Figure 1. We first introduce an overview of the preliminaries in Section 3.1. Next, we detail the four core components of MML-VAU: the video processor (Sec. 3.2), prompt-based vision insights extrac- tor (Sec. 3.3), audio attributes extractor (Sec. 3.4), and the predictive analyzer (Sec. 3.5). 3.1 Preliminary LetV= 푣 1 ,푣 2 ,· ,푣 푛 denote a set of videos in a specific industry vertical (e.g., Ecommerce). Each video푣 푖 =F 푖 ,A 푖 ,T 푖 has three components, whereF 푖 represents a sequence of푚image frames ex- tracted from the hooking period of the video,F 푖 =I 푖 1 ,I 푖 2 ,· ,I 푖 푚 , A 푖 is the audio component in the hooking period, and potentially T 푖 represents the text if the hooking period has human speech rather just background music.T 푖 can also be some textual de- scriptions of the video such as the title or the short summary. The objective of MLLM-VAU is to understand which design feature from three components have an impact on some predefined metrics such as performance (e.g., click-through rate, impression, conversion per investment) or engagement (e.g., likes, shares). In recent practice, researchers and practitioners mainly design various methods to extract features from a video. Each individual feature is often a well-trained model based on some manually labeled datasets. In addition, there exist many approaches that convert a video into an embedding upon which a predictive model is then built. One of the major drawbacks is lack of interpretability due to its nature of ‘black-box’, which does not provide any guidance for advertisers and thus has limited practical value. 3.2 Video Processor To harness the capabilities of multimodal large language models (MLLMs) for analyzing the hooking period of video advertisements, a sophisticated video processing framework is essential. This frame- work is designed to preprocess raw video data by extracting se- quences of image frames, audio attributes, and generating corre- sponding textual descriptions. These multimodal inputs provide a comprehensive foundation for subsequent analysis, ensuring that the intricate interplay between visual, auditory, and textual ele- ments within the hooking period is effectively captured and inter- preted. The initial phase of the video processing framework involves the extraction of image frames and audio attributes from each video advertisement. Image frames are sampled to represent the temporal dynamics of the first three seconds—the hooking period—of the ad- vertisement. Concurrently, the audio stream is processed to extract relevant attributes such as volume levels, frequency components, and temporal variations, which are important for understanding the auditory appeal of the ads. Additionally, automatic speech recogni- tion (ASR) systems are employed to transcribe any spoken content if it exists, thereby generating textual descriptions that complement the visual and auditory data. This multimodal preprocessing en- sures that the MLLMs receive a rich and diverse set of features, enabling a more nuanced analysis of the ad’s initial impact. Two Frame Sampling Strategies A pivotal component of the video processing framework is the selection of frames that best represent the hooking period. We have developed and implemented two distinct frame sampling strategies: random sampling and key frame selection. Random sampling involves selecting frames at a uniform interval throughout the hooking period without any bias towards specific content changes or significant visual events. It can be formalized as uniformly and independently selecting a subset of 푚frames from the total of퐾frames, where퐾is determined by the frames per second (fps). For example, the hooking period has 90 Conference, 2026, USAZhang et al. Figure 1: Overview of our proposed framework: multimodal LLM-based video ad understanding (MLLM-VAU) frames if fps=30.F 푖 =. This ensures a broad and unbiased repre- sentation of the entire hooking period, capturing a diverse array of visual and auditory elements. Additionally, it is straightforward to implement and computationally efficient, making it suitable for processing large volumes of video data rapidly. However, it may miss key moments or significant transitions within the hooking period that are important for capturing the ad’s impact. This ap- proach does not account for the semantic or emotional significance of specific frames, potentially diluting the relevance of the extracted features. On the other hand, key frame selection focuses on identifying and extracting frames that encapsulate significant visual or narrative shifts within the hooking period. The process involves three key steps: •Calculating the difference퐷 푖 between consecutive frames (퐷 푖 =1− 푆퐼푀(퐼 푖 ,퐼 푖+1 ),퐼 푖 is the푖 푡ℎ 푓푟푎푚푒) using a suit- able metric, such as the Structural Similarity Index Measure (SSIM). •Determine a threshold휏to identify significant changes:휏= 훼 ∗ 푚푎푥(퐷 푖 ,퐷 2 ,· ,퐷 퐾 ), where훼is a scaling factor (e.g., 0.5). •Select frames where the difference exceeds the threshold: F= 퐼 푖 |퐷 푖 > 휏. To ensure temporal diversity, impose a minimum intervalΔ푡between consecutive key frames: F=퐼 푖 |퐷 푖 > 휏 and 푖− 푗 ≥Δ푡 ∀ selected frame, 푗< 푖. This strategy leverages algorithms designed to detect changes in scene composition, motion intensity, and other salient features that indicate pivotal moments in the advertisement. As we can see that this strategy requires additional processing overhead, as it involves running algorithms to detect significant changes or events within the video. It may also introduce bias by prioritizing certain types of changes or events, potentially overlooking other relevant aspects of the hooking period. To sum up, the frame selection between random sampling and key frame selection presents a trade- off between computational efficiency and the depth of contextual representation. 3.3 Prompt-based Vision Insights Extractor To extract nuanced insights from each video advertisement, we employ prompt-based multimodal large language models (MLLMs). Specifically, we designed an extractor: vision design methodology extractor. Vision Design Methodology Extractor This employs Llama Multimodal Model 1 , which leverages the advanced capabilities of transformer-based architectures to interpret and analyze the com- plex interplay among a sequence of selected image frames within the hooking period of video ads. By designing tailored prompts, we guide the model to assess the primary engagement strategies employed in each advertisement, thereby facilitating a deeper un- derstanding of the factors that drive ad performance. The effectiveness of MLLMs in extracting meaningful insights heavily relies on the design of the input prompts. Our prompts are meticulously crafted to include both the title and a detailed textual description of the video ad, providing comprehensive context for the model. The prompt is specifically tailored to elicit the model’s assessment of the primary engagement strategy used as the hook in the first three seconds of the ad. The prompt we used is below. After examining the video and text advertisement titled "ad title text" with the body texts "ad body text", determine the primary method used by the advertiser to engage the audience. Base your selection on the actual content of the advertisement without making assumptions or interpretations. Respond using the JSON format. **JSON Response Format:** "methodology": "Methodology chosen by the advertiser", "rationale": "Provide a concise explanation based on specific elements observed in the advertisement 1 https://w.llama.com/ Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video AdsConference, 2026, USA that supports why this option best represents the primary engagement strategy used." This prompt structure ensures that the model receives sufficient contextual information to make an informed assessment. By explic- itly asking for the identification of engagement strategies and a corresponding rationale, we obtain both categorical and explana- tory outputs that enrich the feature set used for predictive modeling. Upon processing the crafted prompt, MLLM generates a structured output that includes both the assessment of the primary engage- ment strategy and a textual rationale explaining the reasoning behind the assessment. This dual-output mechanism provides a clear and interpretable understanding of the strategies used in the hooking period, allowing for more granular analysis and feature extraction. Employing MLLM for feature extraction offers several distinct ad- vantages over traditional methods. First, the adaptability of prompt- based interactions allows for enhanced contextual understanding. Because LLMs, by design, are trained on vast corpora of data encom- passing diverse contexts and can interpret image frames through a lens that incorporates not only visual features but also contex- tual and semantic information. Second, LLMs offer unparalleled flexibility through prompt-based interactions, allowing for the ex- traction of a wide array of features without the need for retrain- ing or extensive reconfiguration. By designing specific prompts, researchers can tailor the feature extraction process to focus on particular aspects of the image that are hypothesized to influence advertisement performance. Third, traditional ML approaches often necessitate significant manual effort in feature engineering, involv- ing the selection, transformation, and combination of various visual attributes to create meaningful features. This process is not only time-consuming but also prone to human bias and may miss subtle yet impactful features. LLM-based feature extraction automates much of this process by leveraging the model’s inherent ability to understand and synthesize information based on the prompts provided. To further distill the rich textual rationales generated by Llama, we employ BERTopic, a topic modeling technique that leverages transformer-based embeddings to identify coherent topics within large text corpora. This step transforms the qualitative rationales into a structured set of latent topics, facilitating their integration into subsequent predictive models. The latent topics derived from BERTopic are integrated with other multimodal features such as audio attributes and aggregated ad information. This comprehen- sive feature set encapsulates both the qualitative insights from the engagement strategies and the quantitative aspects of the video con- tent. Overall, employing prompt-based MLLMs in conjunction with BERTopic offers several advantages: (i) Contextual Understanding: The ability to process and interpret multimodal inputs allows for a more nuanced analysis of video content compared to traditional feature extraction methods. (i) Detailed Insights: The generation of both categorical assessments and rationales, followed by topic modeling, provides a layered understanding of the underlying en- gagement strategies. (i) Scalability: Automated prompt-based anal- ysis and topic modeling enable the processing of large-scale video datasets efficiently. (iv) Enhanced Interpretability: Latent topics summarize complex rationales into coherent themes, facilitating easier interpretation and integration into predictive models. 3.4 Audio Attributes Extractor Beyond visual stimuli, acoustic elements also play a pivotal role in shaping the viewer’s perception and emotional response. This study emphasizes the extraction and analysis of specific acous- tic features during this hooking period to enhance the predictive accuracy of advertisement performance models. The acoustic fea- tures we focus on include decibels (dB), jitter, tempo, degree of dynamic pitch (DDP), pitch (maximum, minimum, mean), power, peak, and shimmer. These features collectively provide a compre- hensive understanding of the audio dynamics that contribute to an advertisement’s effectiveness. Details about each acoustic feature is below. • Decibels (dB): Decibels measure the loudness of the audio signal. In the context of advertisements, variations in volume can influence attention and emotional intensity. A sudden increase in dB may be used to highlight key moments or transitions, thereby enhancing the memorability of the ad. •Jitter: Jitter refers to the frequency variation or instability in the audio signal. High jitter levels can indicate a more dynamic or erratic audio pattern, which may be used to convey excitement or urgency. Conversely, low jitter can create a sense of calmness and stability, aligning with the intended message of the advertisement. •Tempo: Tempo denotes the speed or pace of the audio track. A faster tempo can energize viewers and create a sense of urgency, while a slower tempo may evoke relaxation and contemplation. Understanding the tempo during the hooking period helps in assessing how the audio pace aligns with the visual content to engage viewers effectively. •Degree of Dynamic Pitch (DDP): DDP captures the variability and changes in pitch over time. This feature helps identi- fying tonal shifts that can signal different emotional states or transitions within the advertisement. By analyzing DDP, we can better understand how pitch variations contribute to maintaining viewer interest during the initial moments. •Pitch (Maximum, Minimum, Mean): Pitch analysis provides insights into the fundamental frequency of the audio. The maximum and minimum pitch values, along with the mean pitch, help in characterizing the overall tonal range and emo- tional tone of the advertisement. For instance, higher pitches may be associated with excitement or positivity, while lower pitches might convey seriousness or authority. •Power: Power measures the energy of the audio signal, re- flecting its overall strength and presence. High power levels can make an advertisement more impactful and memorable, whereas lower power may be used to create subtlety or in- timacy. Analyzing power during the hooking period helps in understanding how audio intensity influences viewer en- gagement. •Peak: Peak detection identifies the highest amplitude points in the audio signal. Peaks often correspond to key moments or emphases in the advertisement, such as a dramatic sound Conference, 2026, USAZhang et al. effect or a vocal highlight. Recognizing these peaks is essen- tial for assessing how audio cues are used to draw attention and reinforce the advertisement’s message. •Shimmer: Shimmer quantifies the amplitude variation or instability in the audio signal. This feature is indicative of subtle fluctuations in volume that can add expressiveness and nuance to the audio track. Shimmer analysis helps in evalu- ating how these subtle variations contribute to the overall emotional and psychological impact of the advertisement. 3.5 Predictor To understand ad performance, we integrate the rich features ex- tracted from the hooking period of video ads (e.g., those discussed above) with ad characteristics. The hooking period features include visual design methodology regarding engaging audiences derived from MLLM , and acoustic characteristics, while the aggregated user data are gender, age bucket, advertiser size, zip code, and others. We then link these features to the ad performance metric, conversion per investment (CPI), using Gradient Boosting Decision Tree (GBDT)[10,25]. By training on historical ad performance data, we are able to identify and quantify the correlations between spe- cific features and performance metrics. This not only enhances the predictive accuracy but also provides valuable insights into which aspects of the hooking period – be it certain visual compositions or acoustic dynamics — are most influential in driving successful ad engagements. Ultimately, this approach enables advertisers to optimize their content strategically, leveraging data-driven insights to enhance the effectiveness and impact of their campaigns. 4 Experiments and Results In this section, we first introduce the data and experimental setup (see Sec. 4.1). Then we show the performance comparison results of our method with two baselines: one strongly designed one and one intuitively simple one in Sec 4.2. Major findings are detailed in Sec. 4.3, including the topics summarizing the hook responses into different categories, overall important features that affect the CPI, and the partial dependency plots. Finally, we showcase sev- eral examples to reinforce the practical values of our study in the Appendix. 4.1 Data and Experimental Settings In this study, we focus on advertisers whose ads with one video have a minimum spend of certain amount for the first day and for the 56 days since their launch. The descriptive statistics of our final dataset are described in Table 1. All videos are in the format of MP4. The median video length is 29 seconds. Implementation Details 2 We use the Llama MLLM as the vi- sion design methodology extractor. For the GBDT model, we use the grid search to tune the hyperparameters. The final optimal setting is: #of trees=740, tree depth=12, learning rate=0.0764, minimum number of data points in a node that is required for the node to be split further=50, sample size = 0.86 (the proportion of data that is exposed to the fitting routine), and the others are default values. 2 Sample data and all code can be provided upon request. Table 1: Descriptive statistics of datasets. Note that “CPG" refers to “Consumer Packaged Goods". # of videos[min,max,mean]% of videos Category (Binned)CPIw/ audios Ecommerce100k - 150k[0, 52.60, 0.028]94.72 Healthcare 25k-50k[0, 33.20, 0.045]87.11 CPG75k-100k[0, 37.25, 0.092]91.87 Automobile10k-25k[0, 56.25, 0.219]87.43 Entertainment 10k-25k[0, 13.31, 0.048]95.43 The acoustic features are extracted using the python package li- brosa. 3 Note that all experiments can be done within a few hours (e.g., about 20 hours for the Ecommerce) in total, including feature extraction, training, and testing. Benchmark and Metrics To validate our method, we evaluate its performance and compare it with a strong, carefully designed baseline, widely adopted in video analytics - ViViT 4 and X-CLIP 5 . ViViT is a transformer-based neural network model designed for video analysis that extends Vision Transformers (ViT [9]) by incor- porating spatiotemporal attention to effectively capture both spatial and temporal information in video sequences. X-CLIP is a straight- forward yet powerful architecture designed to adapt pretrained language-image models for direct video recognition. It compre- hends video clips by employing a cross-frame attention mechanism, which explicitly facilitates information exchange across multiple frames. For the efficiency purpose, we uniformly sample 8 frames from the hooking period. We use ViVit/X-CLIP to obtain the embed- ding of a video, which is fed into GBDT model for CPI prediction along with the same set of variables related to ads targeting. We also benchmark against a simple baseline, which we refer to as a “junk" predictor. It converts each video hook into a vector by aggregating the pixel values of each frame, upon which the CPI is predicted along with other targeting variables. We use standard evaluation metrics: R-squared value (푅 2 ) and mean square error (MSE) to demon- strate the model fit and the magnitude of the difference between truth and prediction, respectively. 4.2 Performance Comparison The comparison results are shown in Table 2. From the table, we have the following observations. First, our method outperforms both the strong and the weak baselines for Ecommerce, CPG, and Automobile on both푅 2 and MSE. Second, ViVit achieves the best performance for videos in the vertical of Entertainment. One possi- ble reason is that each entertainment video has a lot more than the fixed number of key frames that are sampled and processed by the Llama MLLM model. ViVit uses all frames in the hooking period, which prevents from losing much information. However, it is worth noting that ViViT is the black-box model and cannot provide any actionable insights to advertisers. Third, the “Junk” benchmark does perform the best for the Healthcare category in terms of푅 2 , but comparable to our method regarding MSE. This could be due to the 3 https://librosa.org/doc/latest/index.html 4 https://huggingface.co/docs/transformers/en/model_doc/vivit 5 https://huggingface.co/docs/transformers/en/model_doc/xclip Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video AdsConference, 2026, USA Table 2: Performance comparison results. Note: The best performance are highlighted in bold. Vertical R-squared (푅 2 )MSE OursViViT [2]X-CLIP [23]“Junk"OursViViTX-CLIP“Junk" Ecommerce0.210.080.200.080.180.231.130.23 Healthcare 0.330.280.270.420.170.220.180.15 CPG0.500.340.070.270.210.210.390.21 Automobile 0.660.570.460.111.952.302.093.51 Entertainment0.250.69-1.180.070.12 0.020.360.15 fact that salient features of videos in this category are primarily demo/products (see Table 3). This indicates that raw pixels of image frames in the hooking period for this vertical might be enough. Finally, X-CLP does not perform well, due to two possible reasons. (i) Many CLIP-based approaches, including X-CLIP, often process videos by extracting frame-level features and then averaging them. This can miss important temporal dynamics, such as motion or ac- tion sequences, which are crucial for understanding video content beyond static scenes. (i) CLIP and similar models are pretrained on large-scale image-text pairs, not video data. This means they may lack the ability to model temporal relationships and actions that are unique to videos. 4.3 Further Findings In this section, we present top 3 visual design methodologies (top- ics), top 3 acoustic features correlated to CPI, and their partial dependencies to CPI for all five verticals. Table 3: Top visual and acoustic features that are correlated with CPI by vertical VerticalVisualAcoustic Ecommerce Topic 6: Interactive contentdB Topic 7: Interactionpower Topic 10: Connectionmaximum pitch Healthcare Topic 10: Demo / productpower Topic 9: Connectionpeak Topic 4: Endorsement / celebrityshimmer CPG Topic 14: Visual appealspower Topic 16: Humorpeak Topic 12: Visual aestheticsddp Automobile Topic 12: Visual appealsmaximum pitch Topic 15: Realismtempo Topic 2: Storytellingpower Entertainment Topic 14: EntertainmentdB Topic 8: Endorsement / celebrityaverage pitch Topic 16: Humormaximum pitch Feature Importance Table 3 presents the top topics and acoustic features of video hooking periods that correlate with CPI. Specif- ically, “Interactive content" emerges as the most effective design methodology for engaging audiences in Ecommerce video ads, fol- lowed by “Interaction" and “Connection." However, effectiveness varies notably across different verticals. For instance, “Demo/Product" is the leading methodology in Healthcare, whereas “Visual appeals" excels in the Consumer Packaged Goods sector. The topics were derived using BERTopic, based on the design methodology and accompanying reasoning illustrations provided by the vision de- sign methodology extractor. Overall, we identified 17 key meth- ods/topics frequently employed by advertisers to engage their au- dience 6 (see the Ecommerce topics in Fig. 2 for details). Figure 2: Topics of hooking period design methods for videos in Ecommerce. Only top 10 words are shown for each topic Partial Dependence To further examine the relationship be- tween a specific feature and the predicted outcome of our model while controlling for the effects of other features, we present partial 6 The number of topics (e.g., 17) is chosen based on the relatively optimal perplexity score among a few choices. Conference, 2026, USAZhang et al. Figure 3: PDP for top 3 visual topics (left) and top 3 acoustic features (right) for Ecommerce dependence plots (PDPs) [10]. An upward trend in the PDP line indicates a positive relationship, meaning the prediction increases as the feature value increases. Conversely, a downward trend in- dicates a negative relationship. Figure 3 displays the PDPs for the top three visual and acoustic features, respectively (similar plots for other verticals are provided in the Appendix). The visual feature analysis suggests that incorporating more interactive content during the hooking period of Ecommerce videos can effectively increase CPI, exerting a stronger influence compared to interaction and connection design elements. In contrast, the acoustic features from topic 3 demonstrate non-linear relationships. Specifically, an optimal range for decibel levels (dB) and maximum pitch can lead to higher CPI, whereas the power feature shows a threshold effect. 5 Conclusion In conclusion, this paper presents a framework for systematically analyzing the hooking period of video advertisements through advanced multimodal large language models. Our framework ef- fectively leverages transformer-based architectures to integrate visual, acoustic, and textual features, enabling a comprehensive understanding of the nuanced interplay among these multimodal elements. The empirical evaluation conducted on large-scale, real- world data demonstrates that our approach can significantly en- hance the prediction of advertisement performance metrics, notably the Conversion Per Investment (CPI). The findings from our analysis provide actionable insights that can guide advertisers in developing more effective and targeted video ad strategies tailored to specific industry verticals. By identi- fying important features that influence viewer engagement during the initial moments of ads, our method empowers advertisers to optimize content strategically and maximize viewer attention and subsequent performance outcomes. Future research could further expand this framework by incorporating additional data modalities such as viewer emotional responses, gaze tracking, or user interac- tion metrics, and exploring more sophisticated multimodal fusion techniques to enhance predictive accuracy and generalizability. Despite the promising results, our study is not without limi- tations that present opportunities for future research. First, the analysis is confined to the initial three seconds of a video ad, which may not capture the full dynamics of viewer engagement across the entire video. Second, the reliance on pretrained multimodal large language models introduces potential biases and sensitivity to prompt design. Furthermore, the dataset, though extensive, is platform-specific and thus may not generalize to other environ- ments. Addressing these limitations in future work would further strengthen the robustness and applicability of the proposed frame- work. Although our system is designed to extract actionable insights from video-based ads using Multimodal Large Language Models (MLLMs), deployment to live users has been blocked due to real- world challenges. Specifically, regulatory constraints around user privacy and ad targeting have prevented us from launching the system at scale. We believe that demonstrating the benefit of this approach can motivate both scholars and practitioners to further explore LLM applications for multimodal data. We provide detailed documentation of our deployment attempts and the key barriers encountered. Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video AdsConference, 2026, USA References [1]Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. ViViT: A Video Vision Transformer. In IEEE/CVF Interna- tional Conference on Computer Vision. 6816–6826. [2]Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021.ViViT: A Video Vision Transformer. arXiv:2103.15691 [cs.CV] [3]Tadas Baltrusaitis, Chaitanya Ahuja, and Louis-Philippe Morency. 2019. Multi- modal Machine Learning: A Survey and Taxonomy. IEEE Trans. Pattern Anal. Mach. Intell. 41, 2 (2019), 423–443. [4]Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is Space-Time Atten- tion All You Need for Video Understanding?. In Proceedings of the 38th Interna- tional Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 139). 813–824. [5]Qifeng Cai, Hao Liang, Hejun Dong, Meiyi Qiang, Ruichuan An, Zhaoyang Han, Zhengzhou Zhu, Bin Cui, and Wentao Zhang. 2025. LoVR: A Benchmark for Long Video Retrieval in Multimodal Contexts. arXiv:2505.13928 [cs.CV] [6]João Carreira and Andrew Zisserman. 2017. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 4724–4733. [7]Junxuan Chen, Baigui Sun, Hao Li, Hongtao Lu, and Xian-Sheng Hua. 2016. Deep CTR Prediction in Display Advertising. In Proceedings of the 24th ACM International Conference on Multimedia. 811–820. [8]Kesha Coker, Richard Flight, and Dominic Baima. 2021. Video storytelling ads vs argumentative ads: how hooking viewers enhances consumer engagement. Journal of Research in Interactive Marketing ahead-of-print (07 2021). [9] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. arXiv:2010.11929 [cs.CV] [10] Jerome H. Friedman. 2001. Greedy function approximation: A gradient boosting machine. The Annals of Statistics 29, 5 (2001), 1189 – 1232. [11] Valentin Gabeur, Chen Sun, Karteek Alahari, and Cordelia Schmid. 2020. Multi- modal Transformer for Video Retrieval. In ECCV 2020. 214–229. [12] Ruohan Gao, Tae-Hyun Oh, Kristen Grauman, and Lorenzo Torresani. 2020. Listen to Look: Action Recognition by Previewing Audio. In IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10454–10464. [13] Zhabiz Gharibshah and Xingquan Zhu. 2021. User Response Prediction in Online Advertising. ACM Comput. Surv. 54, 3 (2021). [14] Maarten Grootendorst. 2022. BERTopic: Neural topic modeling with a class-based TF-IDF procedure. arXiv preprint arXiv:2203.05794 (2022). [15] Daya Guo and Zhaoyang Zeng. 2021. Multi-modal Representation Learning for Video Advertisement Content Structuring. In Proceedings of the 29th ACM International Conference on Multimedia. 4770–4774. [16]Jun Ikeda, Hiroyuki Seshime, Xueting Wang, and Toshihiko Yamasaki. 2021. Predicting Online Video Advertising Effects with Multimodal Deep Learning. In 2020 25th International Conference on Pattern Recognition (ICPR). 2995–3002. [17]Yizhang Jin, Jian Li, Yexin Liu, Tianjun Gu, Kai Wu, Zhengkai Jiang, Muyang He, Bo Zhao, Xin Tan, Zhenye Gan, Yabiao Wang, Chengjie Wang, and Lizhuang Ma. 2024. Efficient Multimodal Large Language Models: A Survey. arXiv:2405.10739 [cs.CV] [18]Jin-Ae Kang, Sookyeong Hong, and Glenn Hubbard. 2020. The role of storytelling in advertising: Consumer emotion, narrative engagement level, and word-of- mouth intention. Journal of Consumer Behaviour 19 (01 2020), 47–56. [19]S. Shunmuga Krishnan and Ramesh K. Sitaraman. 2013. Understanding the effec- tiveness of video ads: a measurement study. In Proceedings of the 2013 Conference on Internet Measurement Conference. Association for Computing Machinery, New York, NY, USA, 149–162. [20]Sangho Lee, Youngjae Yu, Gunhee Kim, Thomas Breuel, Jan Kautz, and Yale Song. 2021. PARAMETER EFFICIENT MULTIMODAL TRANSFORMERS FOR VIDEO REPRESENTATION LEARNING. In ICLR. [21] So-Hyun Lee, Sang-Hyeak Yoon, and Hee-Woong Kim. 2021. Prediction of Online Video Advertising Inventory Based on TV Programs: A Deep Learning Approach. In IEEE Access, Vol. 9. 22516–22527. [22]Jie Lei, Linjie Li, Luowei Zhou, Zhe Gan, Tamara L. Berg, Mohit Bansal, and Jingjing Liu. 2021. Less is More: CLIPBERT for Video-and-Language Learning via Sparse Sampling . In IEEE/CVF Conference on Computer Vision and Pattern Recognition. 7327–7337. [23]Bolin Ni, Houwen Peng, Minghao Chen, Songyang Zhang, Gaofeng Meng, Jian- long Fu, Shiming Xiang, and Haibin Ling. 2022. Expanding Language-Image Pretrained Models for General Video Recognition. In ECCV 2022. 1–18. [24]Hitesh Sagtani, Madan Gopal Jhawar, Rishabh Mehrotra, and Olivier Jeunen. 2024. Ad-load Balancing via Off-policy Learning in a Content Marketplace. In Proceedings of the 17th ACM International Conference on Web Search and Data Mining. New York, NY, USA, 586–595. [25]Si Si, Huan Zhang, S. Sathiya Keerthi, Dhruv Mahajan, Inderjit S. Dhillon, and Cho-Jui Hsieh. 2017. Gradient boosted decision trees for high dimensional sparse output. In Proceedings of the 34th International Conference on Machine Learning. 3182–3190. [26]Xi Tang, Jihao Qiu, Lingxi Xie, Yunjie Tian, Jianbin Jiao, and Qixiang Ye. 2025.Adaptive Keyframe Sampling for Long Video Understanding. arXiv:2502.21271 [cs.CV] [27]Du Tran, Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri. 2015. Learning Spatiotemporal Features with 3D Convolutional Networks. In Pro- ceedings of the 2015 IEEE International Conference on Computer Vision. 4489–4497. [28]Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2023. Attention Is All You Need. arXiv:1706.03762 [cs.CL] [29]Nikhita Vedula, Wei Sun, Hyunhwan Lee, Harsh Gupta, Mitsunori Ogihara, Joseph Johnson, Gang Ren, and Srinivasan Parthasarathy. 2017. Multimodal Content Analysis for Effective Advertisements on YouTube. In 2017 IEEE International Conference on Data Mining (ICDM). 1123–1128. [30] Jinhui Ye, Zihan Wang, Haosen Sun, Keshigeyan Chandrasegaran, Zane Durante, Cristobal Eyzaguirre, Yonatan Bisk, Juan Carlos Niebles, Ehsan Adeli, Li Fei-Fei, Jiajun Wu, and Manling Li. 2025. Re-thinking Temporal Search for Long-Form Video Understanding. arXiv:2504.02259 [cs.CV] [31] Wei Zhang, Dai Li, Chen Liang, Fang Zhou, Zhongke Zhang, Xuewei Wang, Ru Li, Yi Zhou, Yaning Huang, Dong Liang, Kai Wang, Zhangyuan Wang, Zhengxing Chen, Fenggang Wu, Minghai Chen, Huayu Li, Yunnan Wu, Zhan Shu, Mindi Yuan, and Sri Reddy. 2024. Scaling User Modeling: Large-scale Online User Representations for Ads Personalization in Meta. In Companion Proceedings of the ACM Web Conference. 47–55. Conference, 2026, USAZhang et al. Appendix: Showcase We showcase some of the top-performing video ads from our ex- tensive dataset (see Fig. 4).These examples are public ads that can be seen via the link provided, and we are showing here mockups to respect the creative rights and these serve to provide face validity for the effectiveness of our method, particularly emphasizing the robustness of our feature extraction process facilitated by MLLM. By analyzing these standout ads, we can better understand the elements that contribute to their success and how our approach successfully identifies and leverages these key features. Example 1 highlights a human action centered around demonstrating a cough- ing symptom with the produce included in the initial hooking period of the advertisement. This relatable activity effectively captures the audience’s attention right from the start, establishing an emotional connection that engages viewers and encourages them to continue watching. Example 2 focuses on the use of celebrity endorsement or influencer collaboration, where the featured individual is seen posing in a manner that resonates with the target audience. The presence of a well-known personality not only enhances the ad’s credibility but also leverages the influencer’s existing fan base to broaden the advertisement’s reach and impact. This strategic po- sitioning helps in building trust and persuading viewers through association with a familiar and respected figure. Example 3 revolves around a promotional strategy that incorporates a direct response. This approach is designed to prompt immediate viewer interaction, e.g., making a purchase, or visiting the retailer’s website. By clearly communicating the desired action, the ad effectively drives conver- sions and achieves specific marketing objectives, demonstrating the practical application of persuasive techniques within the pro- motional content. Example 4 is the ad from P&G, showing some humor (including a message in the first 3 seconds: “hard launch ♥New Scent of the Year, vanilla Suede") with a visual appealing background (e.g., green plants and wall art decoration). Example 5 features a video introducing a conceptual luxury car through a narrative-driven presentation, telling a story that highlights the car’s design philosophy and innovative features. These examples collectively illustrate the diverse strategies em- ployed in successful video advertisements. They also demonstrate how our method accurately identifies and extracts critical features that contribute to high performance. By leveraging multimodal LLMs for feature extraction, our approach not only validates the effectiveness of these advertising techniques but also provides a scalable framework for analyzing and optimizing future video ad campaigns. Decoding the Hook: A Multimodal LLM Framework for Analyzing the Hooking Period of Video AdsConference, 2026, USA Figure 4: Top row: Example 1 (Healthcare, @benylinsa): https://w.instagram.com/benylinsa/reel/C8667kyKkNS, Example 2 (Entertainment, @mayadeluxeuk):https://w.instagram.com/mayadeluxeuk/reel/C7WiUL2tMAv, and Example 3 (Ecommerce, @momglamlife): https://w.facebook.com/reel/644847075355183; Bottom row: Example 4 (Consumerm @proctergamble):https://w.instagram.com/reel/DLCbpbjI55Y and Example 5 (Automobile, @douradoluxurycars):https://w.instagram.com/reel/DEz6LLdtDGw