Paper deep dive
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti, Andrej Radonjic, Thaddäus Wiedemer, Christoph Schuhmann, Romain Beaumont, Wieland Brendel, Bernhard Schölkopf, A. Sophia Koepke, Jenia Jitsev, Matthias Bethge
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
Tags
Links
- Source: https://arxiv.org/abs/2608.24845v1
- Canonical: https://arxiv.org/abs/2608.24845v1
Trouble viewing inline? Open PDF directly →
Full Text
107,855 characters extracted from source content.
Expand or collapse full text
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Andreas Hochlehnert 1,∗ Marianna Nezhurina 2,3,∗ Mehdi Cherti 2,3 Andrej Radonjic 4 Thaddäus Wiedemer 5,1 Christoph Schuhmann 2 Romain Beaumont 2 Wieland Brendel 5 Bernhard Schölkopf 5 A. Sophia Koepke 1,6,⋄ Jenia Jitsev 2,3,⋄ Matthias Bethge 1,⋄ ∗ Shared first authors ⋄ Shared last authors 1 University of Tübingen, Tübingen AI Center 2 LAION 3 JSC, FZJ 4 Wynd Labs 5 MPI for Intelligent Systems, ELLIS Institute Tübingen 6 Technical University of Munich, MCML https://projects.laion.ai/bvd/ Abstract We present LAION-BVD 1 , a large-scale open video dataset for multimodal learn- ing, which contains 1.3B platform-specific video URLs collected from Common- Crawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale. 1 Introduction The current paradigm for training capable, open multimodal embedding and foundation models relies on accessible, large-scale datasets. Many publicly available datasets exist for text [105,38,11,99, 98,142] and image-text data [128,113,119,15,31,111,12,51,135,37,145,74,112]. While not quite reaching the scale of proprietary datasets, they are sufficiently large to train competitive open models [112, 23, 108, 52]. As a prominent example, LAION-5B [112] and its 2B English-language subset are among the largest public image-text datasets (with Re-LAION-5B [66] being their recent 1 LAION - Big Video Dataset Preprint. arXiv:2608.24845v1 [cs.CV] 25 Aug 2026 MSR-VTT DideMo VideoCC3M WebVid10M Panda-70M HD-VILA-100M InternVid LAION-BVD (Total) LAION-BVD (Captioned) 10 2 10 3 10 4 10 5 10 6 10 7 Total Duration (hours) 40 87 17.5k 52k 167k 371.5k 760.3k 10M 55k Video Datasets Average Retrieval Average Classification 0 10 20 30 40 50 60 70 80 Accuracy (%) ViCLIP DataComp-1B InternVid-10M BVD-V-10M BVD-V-50M 200300400 Model Size 42 43 44 45 46 47 CLAP Laion-Audio BVD-A-10M Figure 1: LAION-BVD is an open web video dataset at unprecedented scale. It facilitates strong downstream performance for multimodal pretraining. Left: dataset scale in comparison with existing video datasets. Right: average downstream performance of ViCLIP [141] (video-text, Table 6) and scaling-trends of CLAP [146] (audio-text, Table 7) on standard benchmarks. These results demonstrate that LAION-BVD serves as valuable training data across both video and audio modalities. update) and were used to train popular models like OpenCLIP [23], KOSMOS-1 [52], and Stable Diffusion [108]. In contrast, open video-language resources remain comparatively small. The largest dataset to date, InternVid [141], contains 7 million videos, which is relatively modest compared to modern text and image corpora. Video data can offer a particularly rich training signal, containing not only large-scale visual content but also aligned audio that can support audio-language learning. The primary bottleneck for large-scale video-text datasets is not the availability of videos per se – millions are accessible on the web – but the significant compute, memory, and engineering effort required to process them at scale. Unlike images, videos are often hosted on large commercial platforms that can restrict or gate access, making large-scale collection challenging in practice. In this work, we address this gap with LAION-BVD, a large-scale open web video dataset. Starting from 1.3B platform-specific video URLs collected from CommonCrawl, we download 10M video hours and process a subset of them into scene-level clips with aligned audio segments as well as individual frames. We then generate modality-specific captions: clip-level video captions for video- language models, audio captions for audio-language models, and frame captions for image-text analysis. In order to validate the quality of the video data for multimodal pre-training we use these captions to train ViCLIP [141], CLAP [146] and CLIP [23] models at different model and data scales. We publish the video URLs for LAION-BVD along with captions for a subset of it on HuggingFace 2 . Furthermore, research institutions can download the raw data after accepting our terms of use. Overall, our contribution can be summarized as follows: •We release LAION-BVD, a large-scale open web video dataset built from 1.3B extracted video URLs and 80M videos with a total duration of 10 million hours. • We introduce a scalable multimodal annotation pipeline that segments videos into scenes and produces clip-level video captions, audio captions, and frame-caption pairs. • Using this pipeline, we release 55M clips (BVD-V-55M) and 300M scene-changing frames (BVD-I-300M) from a randomly sampled subset of the data. • We show that LAION-BVD is a useful open resource for multimodal pre-training. Our ViCLIP and CLAP evaluations support video-language and audio-language learning at scale, while frame-based CLIP experiments provide complementary evidence through strong retrieval performance. 2 https://huggingface.co/collections/laion/bvd-big-video-dataset 2 2 Related Work Multimodal datasets. Training multimodal foundation models requires large-scale datasets contain- ing data from at least two modalities. Video-text datasets usually provide captions for clips ranging from three seconds to a few minutes [141, 123]. Many early datasets are specific to domains like movies [107,117], cooking [169,29], instruction following [110,91], or action recognition [118,13,115,57,62,116,44,139,120]. More relevant for foundation model training are more general datasets that are not tied to one specific domain. Amongst those, smaller datasets can be manually annotated [152,48,151,81], but most recent and larger datasets use automatic speech recognition (ASR), subtitles, alt-text, image captions [165,155,8,93], and/or utilize LLMs [141,138,18,20,56,148,40,69,94,68,136,124]. Some newer resources also broaden supervision beyond purely visual captions by incorporating audio or other multimodal signals [68,81,40,19]. An overview of video-text datasets is provided in Table 1. Table 1: LAION-BVD compared to public open-ended video-text datasets. The LAION-BVD row contextualizes the full released corpus scale; our current ViCLIP and CLAP experiments use a recaptioned subset of 55M clips from roughly 2.4M videos from LAION-BVD. DatasetSourceCaptionsVideosClips Duration in h Median Resolution MSR-VTT [152]YouTubeManual7.2 k10 k40240p DideMo [48]FlickrManual10.5 k27 k87– LongVale [40]ACAV100MGenerated8.4 k105 k550– OpenVid-1M [94]Panda-70M & OthersGenerated–1 M2.1 k– VidGen [124]HD-VILA-100MGenerated3.8 M1 M3 k720p LVD-2M [148]YouTube + Stock footageGenerated2 M2 M11.2 k720p MiraData [56]YouTube + Stock footageGenerated330 k330 k16 k720p VideoCC3M [93]YouTubeTransfer6.3 M10.3 M17.5 k– WebVid10M [8] * Stock footageAlt-text10.7 M10.7 M52 k360p OpenHumanVid [69]Films, Series, DocumentariesGenerated134 k52.3 M70.6 k720p VAST-27M [19]HD-VILA-100MGenerated–27 M77 k720p HowTo100M [91]YouTubeASR1.2 M136 M134.5 k240p Panda-70M [20]HD-VILA-100MGenerated3.8 M70.7 M167 k720p Koala-36M [134]HD-VILA-100MGenerated–36 M172 k720p HD-VG-130M [136]YouTubeGenerated–130 M181 k720p ACAV100M [68]YouTubeGenerated–100 M272 k720p HD-VILA-100M [155]YouTubeASR3.3 M103 M371.5 k720p YT-Temporal-180M [165]YouTubeASR6 M180 M– InternVid [141]YouTubeGenerated7.1 M234 M760.3 k720p LAION-BVDYouTube, Vimeo, DailymotionGenerated80 M1.8 B † 10 M720p † Estimated from the 10M-hour duration using average clip lengths from our captioning pipeline. We release the 80M videos unsplit, along with the analyzed subsets. * this dataset is now defunct Audio-text datasets are comparatively smaller, but they enable text supervision for speech, music, and ambient sounds. Table 2 summarizes representative existing audio-language corpora: Audio- Caps [58] and Clotho [34] offer human-written captions for short general-audio clips. At larger scale, WavCaps [90] and AudioSetCaps [6] broaden coverage through weak or automatically generated captions, and audio-language pretraining efforts such as LAION-CLAP further aggregate large-scale audio-text pairs from open sources [146]. In contrast to these dedicated audio corpora, LAION-BVD is sourced from web videos, retaining the paired visual stream alongside the audio. Image-text pairs are easy to collect at scale, since many images on the web are captioned or come with descriptive alt-text. As a result, open image-text datasets have grown rapidly over the past decade [128,113,119,15,31,111,12,51,135,37,145,74], from just 330k English image-text pairs in MS-COCO [76] to over 2B in LAION-5B [112]’s English-language subset (see Table 3 for more details). Proprietary datasets have been scaled even further to at least 100B image-text pairs [103,55,21,101,100,32,140]. An overview of non-proprietary image-text datasets is provided in Table 3. Interleaved image-text data [2,171,67,52,47,89,73,36,4] can be scaled even further and provides more surrounding context, albeit often with weaker local alignment between modalities. In this space, OmniCorpus [73] experimented with increasing data diversity by including keyframes and video transcriptions. However, both large paired image-text corpora and interleaved web documents ultimately draw from closely related web-image distributions, and image-only scaling already shows diminishing returns on traditional benchmarks [140]. In our setting, captioned video frames are therefore most relevant as an additional source of image-text supervision with a distinct visual distribution. 3 Table 2: LAION-BVD compared to representative publicly accessible audio-text datasets. The captioned LAION-BVD row corresponds to the 55M recaptioned clips from roughly 2.4M videos used in our audio-language experiments, while the total row reports the scale of the fully released (captioned and uncaptioned) corpus. DatasetClipsVideosDuration (in h) Clotho [34]7k-29.1-58.1 AudioCaps [58]46k-142.5 LAION-Audio-630K [146]634k-4.3k AudioSet [39]2.1M2.1M5.8k WavCaps [90]403k-7.6k AudioVerse [154]1.4M-10.1k SoundVECaps [163]1.7M-18.4k Auto-ACD [122]1.5M-4.2k AudioSetCaps [6]6.1M-17k VAST-27M [19]27M-- LAION-BVD55M2.4M56k LAION-BVD (total, uncaptioned)-80M10M Table 3: LAION-BVD compared to publicly accessible image-text datasets. As the images are sampled from videos rather than web-crawled image pages, they complement existing datasets. DatasetSourceCaption Source English Img-Txt Pairs MS-COCO [76]FlickrManual330K Visual Genome [63]MS-COCO + YFCC100MManual5.4M YFCC100M [128]FlickrAlt-text + Title99M C3M [113]Custom web crawlAlt-text3.3M WIT [119]WikipediaAlt-text + Caption5.5M C12M [15]Custom web crawlAlt-text12M RedCaps [31]RedditSubreddit name + Title12M ALT200M [51]Custom web crawlAlt-text203M COYO-300M [12]CommonCrawlGenerated300M LAION-400M [111]CommonCrawl [26]Alt-text413M COYO-700M [12]CommonCrawlAlt-text747 M DataComp-1B [37]CommonCrawlAlt-text1.4B LAION-5B [112]CommonCrawlAlt-text2.3B BVD-I-300MYouTube, Vimeo, DailymotionGenerated300M Multimodal foundation models. Paired image-text, video-text, and audio-text data is used to train several families of foundation models [41]. Within this landscape, LAION-BVD is most directly relevant to video-language and audio-language representation learning, while its frame-caption pairs provide supervision for image-text training. Embedding models like CLIP [103,54,23], ImageBind [42], and other variants [9,150,72] learn shared cross-modal embedding spaces and are common components in multimodal systems. Video- Clip [149], VideoMAE [129], and ViCLIP [141] are important video counterparts, while LAION- CLAP [146] provides an audio-text setup for contrastive audio-language learning [146]. Multimodal LLMs like Flamingo [2], BLIP [27,70,71], LLaVA [78,79,82,80,167], recent GPT variants [170,16,97,17], DeepSeek-VL [84,147] and many others [67,21,100,5,33,102,77, 85,159,137] can process images as an input modality, and output text. Even without an explicit temporal dimension during training, some image-text trained multimodal LLMs exhibit good video understanding [59]. Other multimodal LLMs like Llama model variants [166,75], some GPT variants [132,7,87,121], and others [167,81,86,157,168] are specifically trained with video data. Large multimodal models like recent Gemini [127] models, CoDi [125,126], Next-GPT [144], and VideoPoet [61] handle images and videos as input and output. Image, video, and audio captioning. Image and video captioning have a long history in deep learning, both as benchmark tasks for perceptual understanding and as tools for summarization and 4 abstraction [131,160,45,113,46,114]. Audio captioning benchmarks such as AudioCaps [58] and Clotho [34] extend this to non-visual modalities, supporting the audio-text retrieval task and benchmark introduced by [95,60], while open audio-language models such as Audio Flamingo 3 build on this paired data for audio understanding [43]. Automatic captioning pipelines for multimodal data curation at scale commonly follow a shared approach [145,74,141,156,20,40]: Images (including the center frame for videos) are captioned using a strong pretrained multimodal LLM. Videos are first split into short clips and might be annotated by an existing video-language model like LLaVA-NeXT-Video [167] or frame-by-frame by a more lightweight model like Tag2Text [53]. Audio captions from models like Qwen-Audio / Qwen- Omni [25,153] or transcriptions from models like Whisper-Large [104] can also be incorporated. Final video captions are generated either by synthesizing the partial captions with pretrained language models such as T5 [105], Vicuna [24], Gemini [127], and Claude [3], or by directly captioning videos with multimodal models [158]. 3 Dataset 3.1 Data curation Figure 2: LAION-BVD data curation pipeline. We source data from CommonCrawl [26] and filter for platform-specific video URLs, resulting in 1.3B high-quality video candidates. Out of these, we successfully downloaded 10M video hours using a distributed infrastructure from which we create the BVD-* subsets. In this section, we outline the data collection process for LAION-BVD and provide key dataset statistics. Specifically, Section 3.1 details the data curation procedure, Section 3.2 describes the captioning pipeline, and Section 3.3 presents the dataset statistics. Sources. Large-scale multimodal datasets like LAION-5B [112]/Re-LAION-5B [66] and DataComp- 1B [37] rely on CommonCrawl [26] as a source of raw data. Following this approach, we use Com- monCrawl WAT (Web Archive Transformation) files, which provide essential metadata about archived web pages, including HTTP headers and hyperlinks. We extract platform-specific video links using yt-dlp[162] extractors. To efficiently process this data at scale, we employ thecc2dataset[10] tool in conjunction with an Apache Spark [164] cluster. This setup enables the rapid extraction of video URLs and their associated metadata. We use all CommonCrawl dumps available as of March 2024, resulting in a corpus of 4.7B candidate URLs. To ensure the quality and accessibility of the videos, we filter for links from major supported platforms – YouTube [161], Vimeo [130], and Dailymotion [28] – yielding a final set of 1.3B video URLs of which we download a subset of size ∼80M, which amounts to 10M video hours (see Figure 2). We did not apply additional safety filters at the data collection stage, as these platforms are moderated for harmful and unsafe content. Video download. We download videos using a distributed setup of 2,000 virtual servers coordinated via a cluster built onCelery[14] and powered byyt-dlp[162]. To ensure robust access to video content, we employ a residential proxy network throughout the download process. Overall, we attempted to download 130M videos and achieved a link success rate of approximately60 %, resulting in 80M successfully retrieved videos with a total duration of 10M hours. Clip extraction for video and audio training. We randomly sample 2.4M videos from the LAION- BVD pool and extract 55M clips using a shared pre-processing pipeline for both video-language and audio-language experiments. The pre-processing pipeline is as follows. First, we discard videos shorter than 10 seconds or longer than 30 minutes. We then apply content-based scene detection using 5 051015202530+ Duration (min) 0 10 Percentage (%) Duration Mean: 7.7 min Median: 3.7 min 2006200820102012201420162018202020222024 Year 0 10 Upload Year Figure 3: Dataset statistics for LAION-BVD. Top row: Videos are sourced from YouTube, Vimeo, and Dailymotion, and cover a diverse set of topics. While English remains the most prominent language, nearly half of the videos feature other languages. Bottom row: The mean duration of videos in LAION-BVD is 7.7 minutes and videos cover almost two decades. PySceneDetectwith a threshold of 30 and split each video into scene-level segments. To further improve clip quality, we remove segments that are effectively static by estimating frame-to-frame motion on a low-resolution sample and discarding clips with no meaningful visual change. These video clips are then captioned and used for the ViCLIP experiments, while the associated audio segments are captioned and used for the CLAP experiments (see Figure 4 for example clips and caption statistics). Frame filtering for image training. From the full LAION-BVD pool, we randomly sample videos and extract keyframes usingffmpeg. We then filter out black frames and retain scene-change frames usingffmpeg’s scene detection with a threshold of0.1. We continue this process until reaching approximately 300M filtered frames, resulting in the BVD-I-300M dataset. 3.2 Captioning pipeline Following the download and pre-processing described above we annotate the resulting clips and frames with modality specific captioners. Video captioning. We caption videos using Qwen3-VL-2B-Instruct served through vLLM. For each clip, we sample up to 32 frames uniformly across the video and provide them to the model together with the prompt: “Describe the video in 20 words or less.”. The resulting output is used as a video caption for the clip (see Figure 4 for caption statistics). Audio captioning. We caption audio clips with Audio Flamingo 3 [43]. For each audio segment, we prompt the model with: “Describe the audio sounds in 10 words or less.” This is motivated by the caption length in the AudioCaps [58] and Clotho [34] audio-text datasets, which average around 8 and 11 words respectively (see Figure 4 for caption statistics). Frame recaptioning for image-text analysis. To preserve comparability with existing web-image analyses, we additionally recaption sampled keyframes with DeepSeek-VL2-tiny [147]. These captions are used for the image-text analysis. A detailed account on the caption pipeline and model selection can be found in Section A.3.2. 3.3 Dataset statistics LAION-BVD comprises 10 million video hours, yielding a large corpus of audio-visual content. Figure 1 and Tables 1 to 3 contextualize our dataset within existing video- and audio-language as well as image-text pair datasets. As is the case for other video-language datasets, most videos (94 %) are sourced from YouTube. Much smaller fractions (4 % and 2 %) come from Vimeo and Dailymotion. 6 <2 2-5 5-10 10-2020-3030-60 >60 Time (s) 0 20 40 Percentage (%) BVD-V-* Clip Duration <5 5-10 10-1515-2020-2525-30 >30 Words BVD-V-* Caption Length <5 5-10 10-1515-2020-2525-30 >30 Words BVD-A-* Caption Length Figure 4: Overview of the 55M extracted clips from the LAION-BVD corpus. Top: Sampled clips from the 55M extracted clips with corresponding video and audio captions. Bottom: Distribution of clip and caption lengths for both video (BVD-V- ∗ ) and audio (BVD-A- ∗ ) modalities. LAION-BVD is diverse in its topic coverage. Vlogs (20 %) and music (17 %) lead only narrowly, reflecting a broad and diverse content distribution. Furthermore, our dataset is noticeably multi- lingual. While the majority (almost60 %) of video content is in English, other languages are present in significant proportions. Like other video datasets, most videos in LAION-BVD are below 5 minutes long, with an average duration of 7.7 minutes. However, the distribution is rather long-tailed, and almost6 %of videos are 30 minutes or longer, supporting long-horizon video tasks. LAION-BVD includes relatively recent uploads, with the latest entries from early 2024. For our modality-specific analyses, we create two subsets from LAION-BVD as described in Sections 3.1 and 3.2: Captioned video/audio subset. Our video-text (ViCLIP) and audio-text (CLAP) data validation experiments in Section 4 operate on 55M clips constructed from 2.4M videos. Statistics of these clips, including clip lengths, caption characteristics, and representative examples, are summarized in Figure 4. Captioned frame subset. The frame-captions are created for 300M scene-changing keyframes from LAION-BVD. As an additional image-centric analysis, we quantify distributional differences between web-image data and our video-derived frames by computing a Fréchet Inception Distance (FID) over CLIP-ViT-B/32 embeddings using 100k randomly sampled images from each source. We observe a substantial shift between the distributions of ReLAION and LAION-BVD frames (FID = 33.92), while two independent 100k samples from ReLAION yield a near-zero FID (0.16). This confirms that LAION-BVD provides a complementary visual distribution to existing web-scale image datasets, supporting its value as an additional pre-training source. 7 Table 4: L-14 scaling with LAION-BVD data on classification and retrieval tasks. LAION- BVD-trained ViCLIP outperforms both InternVid-trained ViCLIP and the image-only CLIP baseline trained on DataComp-1B where video embeddings are obtained by averaging frame embeddings. Performance consistently improves with increased training data and dataset scale, with BVD-V-50M achieving the best overall average. (* denotes results reported in [141]; all other results are evaluated using our pipeline.) DatasetSamplesK400UCF-101HMDB51MSR-VTTMSVDAvg seentop-1top-1top-1T2V2T2V2T DataComp-1B (image only)-61.278.148.137.326.948.363.353.2 InternVid-10M*10M56.7--36.437.145.269.8- InternVid-50M*50M57.2--39.740.746.572.2- InternVid-10M-FLT*10M64.8--42.441.349.175.1- InternVid-10M-FLT10M61.878.655.138.739.850.974.158.0 BVD-V-10M10M62.779.560.442.742.953.281.061.3 BVD-V-10M50M63.179.460.341.343.153.283.061.4 BVD-V-50M10M63.378.960.443.542.653.682.161.5 BVD-V-50M50M63.379.761.442.843.253.783.962.0 Table 5: Scaling behavior of ViCLIP on LAION-BVD across classification and retrieval tasks. Results demonstrate consistent performance gains with increased model size and samples seen. ModelDatasetSamplesK400UCF-101HMDB51MSR-VTTMSVDAvg seentop-1top-1top-1T2V2T2V2T B-32BVD-V-55M10M49.970.244.634.238.144.274.851.4 B-32BVD-V-55M50M51.670.445.435.639.046.377.752.7 B-16BVD-V-55M10M54.671.851.736.040.048.679.055.1 B-16BVD-V-55M50M55.975.352.437.740.849.481.356.8 L-14BVD-V-55M10M63.378.560.643.644.353.783.461.9 L-14BVD-V-55M50M63.380.560.442.743.053.684.362.0 4 Experiments validating LAION-BVD We evaluate LAION-BVD on three modalities: video, audio, and static frames. For video and audio, we study whether 55M synthetically annotated clips sourced from 2.4M randomly sampled videos from the LAION-BVD corpus support video-text and audio-text learning with ViCLIP and CLAP respectively. For static frame analysis, we assess LAION-BVD as a source of image-text data by captioning 300M individual frames and training CLIP. Hyperparameters and experimental details, are provided in Section A. 4.1 Video-text data validation with ViCLIP For training ViCLIP, we use the 55M-clip subset described in Section 3.1. Following [141], we further construct two smaller video subsets of 10M and 50M samples, sampled uniformly from the 55M-clip subset. We denote them as BVD-V-10M and BVD-V-50M. We evaluate video performance on three zero-shot action recognition benchmarks (Kinetics-400 [57], UCF-101 [118], and HMDB51 [64]) and two video-text retrieval datasets (MSR-VTT and MSVD [152]), reporting a total of three classification metrics (top-1 accuracy) and four retrieval metrics (video-to-text retrieval recall@1 and text-to-video retrieval recall@1). To provide a concise summary of these results, we report an overall average (denoted Avg in Table 4) defined as the arithmetic mean of two components: the average classification accuracy across the three classification metrics and the average of the four retrieval metrics across both retrieval datasets. ViCLIP trained on LAION-BVD outperforms InternVid-trained models. Table 4 illustrates performance as a function of increasing LAION-BVD data scale and number of training samples. Our LAION-BVD subsets achieve better performance than InternVid-10M and InternVid-50M. In particular, at 50M samples seen, BVD-V-10M and BVD-V-50M outperform the more heavily filtered subset InternVid-10M-FLT by 3.3 and 4.0 percentage points respectively, in terms of the overall average metric. Furthermore, we observe clear scaling trends with both increasing dataset size and compute, where compute is measured by the number of samples seen; this is particularly notable given the minimal filtering applied to LAION-BVD. This trend is further supported by Table 5, where we observe a strictly monotonic increase in average performance with increasing model size and compute. Together, these scaling behaviors and competitive results provide strong evidence that LAION-BVD is a high-quality dataset for video pre-training. 8 Table 6: Improved ViCLIP L-14 results with LAION-BVD data using checkpoint merging. Applying WiSE-FT checkpoint merging [143] improves performance of our models. For InternVid- 10M-FLT, it narrows the gap to the results reported by [141] and improves the overall average by 2.2p (* denotes results reported in [141]; all other results are evaluated using our pipeline.) DatasetSamplesWiSE-FTK400UCF-101HMDB51MSR-VTTMSVDAvg seentop-1top-1top-1T2V2T2V2T InternVid-10M-FLT*10M✓64.8--42.441.349.175.1- InternVid-10M-FLT10M✗61.878.655.138.739.850.974.158.0 InternVid-10M-FLT10M✓65.080.857.343.541.751.474.560.2 BVD-V-10M10M✗62.779.560.442.742.953.281.061.3 BVD-V-10M10M✓64.180.560.044.545.354.580.862.3 BVD-V-50M50M✗63.379.761.442.843.253.783.962.0 BVD-V-50M50M✓64.379.961.044.643.754.584.662.6 Improved ViCLIP L-14 results with WiSE-FT checkpoint merging. As shown in Table 4, our initial evaluation of InternVid-10M-FLT yielded lower performance than that reported by [141]. We subsequently found a discussion in the official InternVid GitHub repository in which one of the authors indicated that the reported results were obtained using WiSE-FT checkpoint merging [143]. Specifically, the InternVid-10M-FLT checkpoint was averaged with the original CLIP checkpoint at test time. Following this, we evaluate our models and InternVid-10M-FLT with WiSE-FT and report the results in Table 6. For InternVid-10M-FLT, we use an interpolation coefficient ofα = 0.5, correspond- ing to the equal-weight averaging described by the authors. For BVD-V-50M, we selectαfrom 0.5, 0.6, 0.7, 0.8, 0.9which leads to the highest average metric based on the validation sets, and report the results of the merged model on the test sets. Applying WiSE-FT increases the overall average for InternVid-10M-FLT from 58.0 to 60.2 and substantially narrows the gap to the reported results in[141], suggesting that checkpoint merging accounts for at least part of the discrepancy. WiSE-FT also improves the overall average of our BVD-V-10M model from 61.3 to 62.3 and our BVD-V-50M model from 62.0 to 62.6. Takeaway Minimally curated LAION-BVD samples outperform InternVid-trained ViCLIP (FLT) by2.1 percentage points on an aggregate metric encompassing zero-shot classification and retrieval benchmarks. This indicates that LAION-BVD is a valuable source for video-text pre-training. 4.2 Audio-text data validation with CLAP For CLAP training, we utilize the audio data from the 55M clips introduced in Section 3.1. We further construct two randomly sampled audio subsets of sizes 1.7M and 10M, referred to as BVD- A-1.7M and BVD-A-10M. We evaluate audio performance on one classification benchmark (Ur- banSound8K [109]) and text-audio retrieval [95,60] on AudioCaps [58] and Clotho [34], reporting audio-to-text and text-to-audio recall@5. As a summary metric, we report the average across classification and retrieval. BVD-A-10M outperforms the larger LAION-Audio dataset in single-dataset training. Table 7 isolates the effect of the training data source on four model scales, comparing LAION-Audio (LA), Table 7: Average zero-shot audio performance for pure training data sources at 30M samples seen with CLAP models. Each cell reports the best average score per model scale, aggregated over UrbanSound8K [109] classification and AudioCaps/Clotho retrieval [58,34,95,60]. Large-scale LAION-BVD audio alone matches or exceeds LAION-Audio (LA) across model scales, although the strongest results still come from adding curated audio data from AudioSet (LA+AS). Despite minimal curation, BVD-A-10M exhibits consistent scaling behavior with increasing model size. ModelParamsSamplesLALA+ASBVD-A-1.7MBVD-A-10M HTSAT-tiny-Roberta-base158M30M42.156.442.442.9 HTSAT-base-Roberta-base200M30M43.255.545.744.8 HTSAT-tiny-Roberta-large388M30M45.055.845.645.9 HTSAT-base-Roberta-large431M30M44.856.644.546.8 9 Table 8: Average zero-shot audio performance with CLAP models for mixed training datasets at∼28M samples seen. Each cell reports the best average score per model scale, aggregated over UrbanSound8K [109] classification and AudioCaps/Clotho retrieval [58,34]. Across all model scales, replacing LAION-Audio (LA) [146] with BVD-A-1.7M yields competitive performance, while incorporating AudioSet (AS) consistently achieves the strongest results. ModelParamsSamplesAC+CL+BVD-A-1.7MAC+CL+LAAC+CL+LA+AS HTSAT-tiny-Roberta-base158M28.2M60.662.266.4 HTSAT-base-Roberta-base200M27.1M62.064.164.1 HTSAT-tiny-Roberta-large388M27.1M61.363.966.6 HTSAT-base-Roberta-large431M28.2M62.363.765.7 Table 9: Scaling behavior of pure BVD-A-10M CLAP training. We compare best-checkpoint results at roughly 30M and 110M samples seen across four model scales. More seen audio-text samples help primarily at larger model sizes: the 431M model reaches the strongest overall average (48.7) and best retrieval scores at 110M samples seen. AudioCaps R@5Clotho R@5 ScaleModelParamsSamplesUS8KT2A2T2A2TAvg ∼ 30MHTSAT-tiny-Roberta-base158M30.0M59.245.352.525.132.242.9 HTSAT-base-Roberta-base200M30.0M59.046.956.526.335.544.8 HTSAT-tiny-Roberta-large388M30.0M60.548.558.326.835.445.9 HTSAT-base-Roberta-large431M30.0M64.848.659.226.934.546.8 ∼ 110MHTSAT-tiny-Roberta-base158M110M50.943.552.222.530.940.0 HTSAT-base-Roberta-base200M110M60.541.950.522.830.741.3 HTSAT-tiny-Roberta-large388M110M68.846.756.926.333.446.4 HTSAT-base-Roberta-large431M110M70.650.059.327.236.448.7 LAION-Audio + AudioSet (LA+AS), as well as our BVD-A-1.7M and BVD-A-10M subsets. Here, we can see that both BVD-A-1.7M and BVD-A-10M match or exceed LAION-Audio across different model scales for matched samples seen. Only LAION-Audio + AudioSet (which contains data specifically curated for audio tasks) yields better results. Furthermore, despite minimal curation, both LAION-BVD audio subsets exhibit consistent scaling behavior with increasing model size. BVD-A-1.7M remains competitive with LAION-Audio in mixed-dataset training. Table 8 compares CLAP training on four model scales across three training sets: AudioCaps + Clotho + BVD-A-1.7M (AC+CL+BVD-A-1.7M); AudioCaps + Clotho + LAION-Audio (AC+CL+LA); and AudioCaps + Clotho + LAION-Audio + AudioSet (AC+CL+LA+AS), which matches the original LAION-CLAP [146] training data. We observe that despite only minimal curation of BVD-A-1.7M, AC+CL+BVD-A-1.7M remains close to AC+CL+LA while AC+CL+LA+AS, containing the highly curated AudioSet, is strongest overall. BVD-A-10M exhibits consistent scaling behavior across model sizes. Table 9 shows the scaling behavior of BVD-A-10M on two data scales (30M and 110M samples seen) and four model sizes. Increasing model size consistently improves average benchmark performance. Increasing the number of samples helps only for larger models (388M+). This is an expected bottleneck effect, where smaller model scales cannot benefit from increasing samples scale, with a further factor being the diminishing gains from the high number of repetitions (110M samples scale) over only 10M unique samples in BVD-A-10M. The scaling effect is also present in mixed dataset training (Table 8). However, it is weaker, as the highly curated small scale AudioCaps/Clotho training datasets dominate the model performance. Takeaway LAION-BVD provides a viable source of audio data for multimodal pre-training, supported by consistent performance across single- and mixed-dataset settings and favorable scaling behavior. 4.3 Image-text validation with CLIP We train CLIP models on extracted frame–caption pairs from the BVD-I-300M subset of LAION- BVD and compare them to models trained on DataComp-1B [37] and Re-LAION-2B [66] using 10 Table 10: Image-text training with CLIP across datasets and scales. LAION-BVD achieves the strongest COCO retrieval results, while web-image datasets remain stronger on ImageNet-style classification. AccCOCO Retrieval R@5Acc ModelDatasetSamplesImNet-1kT2I2TImNet-RImNet-SketchImNet-V2 ViT-B-32Re-LAION30.7M0.170.200.320.220.100.15 64M0.260.300.450.320.170.22 128M0.350.380.530.400.240.28 300M0.440.460.640.520.320.37 ViT-B-32DataComp30.7M0.200.210.320.250.120.17 64M0.310.290.430.360.210.26 128M0.400.370.540.450.290.33 300M0.500.450.630.570.390.42 ViT-B-32LAION-BVD30.7M0.080.270.400.130.040.08 64M0.130.400.560.180.060.11 128M0.190.480.670.260.110.16 300M0.240.560.740.340.160.21 ViT-B-16Re-LAION128M0.410.450.620.480.290.34 300M0.510.530.710.590.380.44 ViT-B-16DataComp30.7M0.220.260.380.280.150.21 64M0.380.360.520.420.260.32 128M0.470.430.600.520.340.40 300M0.580.520.700.640.440.49 ViT-B-16LAION-BVD30.7M0.100.350.510.160.050.10 64M0.180.500.670.240.090.15 128M0.220.560.740.300.140.19 300M0.280.630.800.400.190.23 the same training recipe and data scales, evaluating on zero-shot ImageNet-1k classification and MS-COCO retrieval (Table 10). BVD-I-300M achieves strong retrieval performance. Overall, LAION-BVD demonstrates strong performance on text-to-image and image-to-text retrieval, while lagging behind on ImageNet-1k zero-shot classification (see Appendix B for further analysis). We observe comparable performances for CLIP models trained on LAION-BVD and models trained on DataComp-1B recaptioned with our captioning pipeline (e.g. 0.23 ImageNet-1k accuracy compared to 0.19 with LAION-BVD when training on 128M frames as can be seen in Tables 10 and 25). In addition to ImageNet-1k, we also evaluate our models on ImageNet-R [49], ImageNet-Sketch [133], and ImageNet-V2 [106], which further support the generalization capabilities of models trained on LAION-BVD. BVD-I-300M shows strong scaling trends for retrieval tasks. The plots in Figure 5 show strong scaling trends for LAION-BVD across tasks. While all datasets yield improved performance when using larger sets of image-text pairs for training, LAION-BVD exhibits weaker scaling for ImageNet-1k zero-shot classification compared to Re-LAION and DataComp. This again confirms that its synthetic captions may be less effective for fine-grained, label-centric classification. However, we emphasize that strong retrieval performance is indicative of high-quality embeddings, making LAION-BVD particularly well-suited for representation learning and downstream applications such as similarity search or kNN-based indexing. In this context, the lower ImageNet-1k performance is less concerning and likely reflects a weaker alignment with the specific class distribution and biases of ImageNet-1k: biases that are explicitly or implicitly encoded in datasets like DataComp (e.g., via clustering towards ImageNet classes) and Re-LAION/OpenCLIP (e.g., via CLIP-based filtering grounded in WordNet hierarchies). In contrast, LAION-BVD demonstrates strong scaling behavior on MS-COCO text-to-image retrieval, outperforming other datasets with a steeper slope and 11 10 9 10 10 Compute C [GFLOPs] 4 × 10 1 6 × 10 1 ImageNet1k 0-shot [Error Rate] DataComp 1B: 37.49 * x 0.19 Re-LAION 2B EN: 19.79 * x 0.16 LAION-BVD (Ours): 3.93 * x 0.07 (a) ImageNet-1k zero-shot classification error rate 10 9 10 10 Compute C [GFLOPs] 4 × 10 1 6 × 10 1 MS-COCO Image Retrieval R@5 [Error Rate] DataComp 1B: 14.92 * x 0.15 Re-LAION 2B EN: 19.32 * x 0.16 LAION-BVD (Ours): 35.92 * x 0.20 (b) MS-COCO image retrieval R@5 error rate Figure 5: Image-text scaling trends for LAION-BVD, Re-LAION, and DataComp-1B. LAION- BVD scales especially well on retrieval, while its ImageNet zero-shot classification performance remains weaker than web alt-text baselines. lower error rates when using more data, further underscoring its effectiveness for retrieval-focused applications. Takeaway Frame-caption pairs from LAION-BVD provide strong supervision for image-text retrieval and scale favorably with data size, but yield weaker performance on classification benchmarks due to differences in caption style and class coverage. This shows that videos from LAION-BVD can be used as a new source of image-text pairs. 5 Limitations Caption quality and style. All captions in LAION-BVD are automatically generated and intention- ally short using small captioning models. This enables scalable processing across large volumes of video data, but it constrains the richness of the captions. It may also introduce systematic errors or biases from the captioning model. Model scope. Our evaluation focuses on contrastive video-text training using ViCLIP and does not cover generative paradigms such as video-language models or diffusion-based approaches. While we believe that the validation experiments with ViCLIP, CLAP, and CLIP provide a useful signal of dataset quality, future work should further evaluate the data on generative tasks. Audio-visual integration. We evaluate video-text, audio-text, and image-text settings separately, and do not train joint audio-visual models. This leaves open how well LAION-BVD supports learning unified representations across modalities, particularly for tasks that require tight synchronization between visual and auditory signals. Data biases. LAION-BVD is constructed from video platforms with minimal filtering. While the underlying platforms apply moderation, the dataset may still contain biases, stereotypes, or uneven representation across languages, regions, and topics. Models trained on this dataset can inherit such biases. 6 Discussion and Conclusion The LAION-BVD dataset is designed for multimodal pre-training at scale. It features 1.3B video platform links, 10 million video hours, 55M clip-level video and audio captions, and 300M frame captions for image-text training. Our analysis shows that all modalities (video, audio, individual frames) provide a strong training signal. ViCLIP models trained on LAION-BVD outperform InternVid-based models on matched model and data scale. For audio, CLAP achieves competitive results despite minimal curation, and CLIP models trained on extracted video frames perform strongly on retrieval benchmarks. Across 12 modalities, we observe consistent scaling trends with increasing data and compute, indicating that LAION-BVD is a reliable source for multimodal pre-training. LAION-BVD broadens access to large-scale open web video and supports more reproducible multimodal foundation model research. By releasing the dataset and metadata to research institutions, and by documenting a scalable curation pipeline, we aim to lower the barrier to studying video, audio, and image data jointly at this scale. Acknowledgments This work was supported in part by the German Federal Ministry of Research, Technology and Space (BMFTR) under grants FKZ 01IS24060, FKZ 01IS18039A, FKZ 16IS24079A (XEI), 01IS24085C (OPENHAFM), 01IS22094B (WestAI – AI Service Center West), and 16HPC117K (MINERVA), as well as by the German Research Foundation (DFG) through SFB 1233 (project number 276693517) and utilized compute resources at the Tübingen Machine Learning Cloud, DFG FKZ INST 37/1057-1 FUGG. AH acknowledges support from the International Max Planck Research School for Intelligent Systems (IMPRS-IS). MB additionally acknowledges support from Coefficient Giving funded by the Good Science Ventures Foundation (Open Philanthropy). MB, MN, MC, and J further acknowledge funding from the European Union Horizon Programme under grant no. 101214398 (ELLIOT), co-funding from the Digital Europe Programme under grant no. 101195233 (openEuroLLM) and grant no. 101198470 (LLMs4EU), and support from the EuroHPC Joint Undertaking under grant no. 101182737 (MINERVA). WB acknowledges financial support via an Emmy Noether Grant funded by the German Research Foundation (DFG) under grant no. BR 6382/1-1 and WB is a member of the Machine Learning Cluster of Excellence, EXC number 2064/1 - Project number 390727645. We gratefully acknowledge the Gauss Centre for Supercomputing e.V. (GCS) for providing computing resources through the John von Neumann Institute for Computing (NIC) on the JUWELS Booster and JUPITER supercomputers at the Jülich Supercomputing Centre (JSC). We further acknowledge the EuroHPC Joint Undertaking for computing time and storage on the LEONARDO supercomputer hosted by CINECA (Italy) and the LEONARDO consortium through the EuroHPC Extreme Access grant EHPC-EXT-2023E02-068. We also acknowledge storage resources on JUST operated by JSC and supported by the Helmholtz Data Federation (HDF), computing time granted by JARA and JSC on the JURECA supercomputer, access to the JEDI prototype system through the JUREAP (JUPITER Early Access Program), and computing time provided through the Gauss AI Competition (reformo) on JUPITER via GCS and BMFTR. LAION further acknowledges a public storage grant from Hugging Face, enabling broad community access to open-source research outputs through Hugging Face repositories. We also thank the teams operating the supercomputing facilities for their support, especially Björn Hagemeier for assistance with data storage and handling. References [1] M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, N. Bach, A. Bahree, A. Bakhtiari, J. Bao, H. Behl, et al. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219, 2024. [2]J.-B. Alayrac, J. Donahue, P. Luc, A. Miech, I. Barr, Y. Hasson, K. Lenc, A. Mensch, K. Milli- can, M. Reynolds, et al. Flamingo: a visual language model for few-shot learning. Advances in neural information processing systems, 35:23716–23736, 2022. [3] Anthropic.Claude3-haiku.https://w.anthropic.com/news/claude-3-family, 2024. Accessed: 2025-05-14. [4]A. Awadalla, L. Xue, O. Lo, M. Shu, H. Lee, E. Guha, S. Shen, M. Awadalla, S. Savarese, C. Xiong, et al. Mint-1t: Scaling open-source multimodal data by 10x: A multimodal dataset with one trillion tokens. Advances in Neural Information Processing Systems, 37:36805–36828, 2024. [5]J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou. Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023. URL https://arxiv.org/abs/2308.12966. 13 [6]J. Bai, H. Liu, M. Wang, D. Shi, W. Wang, M. D. Plumbley, W.-S. Gan, and J. Chen. Audioset- caps: An enriched audio-caption dataset using automated generation pipeline with large audio and language models. IEEE Transactions on Audio, Speech and Language Processing, 2025. [7]S. Bai, Y. Cai, R. Chen, K. Chen, X. Chen, Z. Cheng, L. Deng, W. Ding, C. Gao, C. Ge, W. Ge, Z. Guo, Q. Huang, J. Huang, F. Huang, B. Hui, S. Jiang, Z. Li, M. Li, M. Li, K. Li, Z. Lin, J. Lin, X. Liu, J. Liu, C. Liu, Y. Liu, D. Liu, S. Liu, D. Lu, R. Luo, C. Lv, R. Men, L. Meng, X. Ren, X. Ren, S. Song, Y. Sun, J. Tang, J. Tu, J. Wan, P. Wang, P. Wang, Q. Wang, Y. Wang, T. Xie, Y. Xu, H. Xu, J. Xu, Z. Yang, M. Yang, J. Yang, A. Yang, B. Yu, F. Zhang, H. Zhang, X. Zhang, B. Zheng, H. Zhong, J. Zhou, F. Zhou, J. Zhou, Y. Zhu, and K. Zhu. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. [8]M. Bain, A. Nagrani, G. Varol, and A. Zisserman. Frozen in time: A joint video and image encoder for end-to-end retrieval. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1728–1738, 10 2021. [9]H. Bao, W. Wang, L. Dong, Q. Liu, O. K. Mohammed, K. Aggarwal, S. Som, S. Piao, and F. Wei. Vlmo: Unified vision-language pre-training with mixture-of-modality-experts. Advances in Neural Information Processing Systems, 35:32897–32912, 2022. [10] R. Beaumont. c2dataset. https://github.com/rom1504/c2dataset, 2022. [11] S. Biderman, K. Bicheno, and L. Gao. Datasheet for the pile. arXiv preprint arXiv:2201.07311, 2022. [12]M. Byeon, B. Park, H. Kim, S. Lee, W. Baek, and S. Kim. Coyo-700m: Image-text pair dataset. https://github.com/kakaobrain/coyo-dataset, 2022. [13]F. Caba Heilbron, V. Escorcia, B. Ghanem, and J. Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceedings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. [14]Celery. Celery: Distributed task queue.https://docs.celeryq.dev/. Accessed: 2025- 05-14. [15] S. Changpinyo, P. Sharma, N. Ding, and R. Soricut. Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3558–3568, 2021. [16] J. Chen, D. Zhu, X. Shen, X. Li, Z. Liu, P. Zhang, R. Krishnamoorthi, V. Chandra, Y. Xiong, and M. Elhoseiny. Minigpt-v2: large language model as a unified interface for vision-language multi-task learning. arXiv preprint arXiv:2310.09478, 2023. [17] L. Chen, J. Li, X. Dong, P. Zhang, C. He, J. Wang, F. Zhao, and D. Lin. Sharegpt4v: Improving large multi-modal models with better captions. In European Conference on Computer Vision, pages 370–387. Springer, 2024. [18] S. Chen, H. Li, Q. Wang, Z. Zhao, M. Sun, X. Zhu, and J. Liu.Vast: A vision- audio-subtitle-text omni-modality foundation model and dataset. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural In- formation Processing Systems, volume 36, pages 72842–72866. Curran Associates, Inc., 2023.URLhttps://proceedings.neurips.c/paper_files/paper/2023/file/ e6b2b48b5ed90d07c305932729927781-Paper-Conference.pdf. [19] S. Chen, H. Li, Q. Wang, Z. Zhao, M. Sun, X. Zhu, and J. Liu. Vast: A vision-audio-subtitle- text omni-modality foundation model and dataset. Advances in Neural Information Processing Systems, 36:72842–72866, 2023. [20]T.-S. Chen, A. Siarohin, W. Menapace, E. Deyneka, H.-w. Chao, B. E. Jeon, Y. Fang, H.-Y. Lee, J. Ren, M.-H. Yang, and S. Tulyakov. Panda-70m: Captioning 70m videos with multiple cross-modality teachers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 13320–13331, 6 2024. 14 [21]X. Chen, X. Wang, S. Changpinyo, A. Piergiovanni, P. Padlewski, D. Salz, S. Goodman, A. Grycner, B. Mustafa, L. Beyer, et al. Pali: A jointly-scaled multilingual language-image model. arXiv preprint arXiv:2209.06794, 2022. [22]Z. Chen, J. Wu, W. Wang, W. Su, G. Chen, S. Xing, M. Zhong, Q. Zhang, X. Zhu, L. Lu, B. Li, P. Luo, T. Lu, Y. Qiao, and J. Dai. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024. URL https://arxiv.org/abs/2312.14238. [23]M. Cherti, R. Beaumont, R. Wightman, M. Wortsman, G. Ilharco, C. Gordon, C. Schuhmann, L. Schmidt, and J. Jitsev. Reproducible scaling laws for contrastive language-image learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2818–2829, 2023. [24]W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, I. Stoica, and E. P. Xing. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, 3 2023. URL https://lmsys.org/blog/2023-03-30-vicuna/. [25]Y. Chu, J. Xu, X. Zhou, Q. Yang, S. Zhang, Z. Yan, C. Zhou, and J. Zhou. Qwen-audio: Advancing universal audio understanding via unified large-scale audio-language models. arXiv preprint arXiv:2311.07919, 2023. [26] Common Crawl. Common crawl. https://commoncrawl.org/. Accessed: 2025-05-14. [27]W. Dai, J. Li, D. Li, A. M. H. Tiong, J. Zhao, W. Wang, B. Li, P. Fung, and S. Hoi. Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023. URLhttps: //arxiv.org/abs/2305.06500. [28] Dailymotion. Dailymotion. https://w.dailymotion.com/. Accessed: 2025-05-14. [29] D. Damen, H. Doughty, G. M. Farinella, S. Fidler, A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T. Perrett, W. Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European conference on computer vision (ECCV), pages 720–736, 2018. [30] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition, pages 248–255. Ieee, 2009. [31] K. Desai, G. Kaul, Z. Aysola, and J. Johnson. Redcaps: Web-curated image-text data created by the people, for the people. arXiv preprint arXiv:2111.11431, 2021. [32] H. Dong, Z. Kang, W. Yin, X. Liang, C. Feng, and J. Ran. Scalable vision language model training via high quality data curation. arXiv preprint arXiv:2501.05952, 2025. [33] D. Driess, F. Xia, M. S. Sajjadi, C. Lynch, A. Chowdhery, A. Wahid, J. Tompson, Q. Vuong, T. Yu, W. Huang, et al. Palm-e: An embodied multimodal language model, 2023. [34] K. Drossos, S. Lipping, and T. Virtanen. Clotho: An audio captioning dataset. In ICASSP 2020- 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 736–740. IEEE, 2020. [35] H. Duan, J. Yang, Y. Qiao, X. Fang, L. Chen, Y. Liu, X. Dong, Y. Zang, P. Zhang, J. Wang, et al. Vlmevalkit: An open-source toolkit for evaluating large multi-modality models. In Proceedings of the 32nd ACM International Conference on Multimedia, pages 11198–11201, 2024. [36]M. Futeral, A. Zebaze, P. O. Suarez, J. Abadji, R. Lacroix, C. Schmid, R. Bawden, and B. Sagot. moscar: A large-scale multilingual and multimodal document-level corpus. arXiv preprint arXiv:2406.08707, 2024. [37]S. Y. Gadre, G. Ilharco, A. Fang, J. Hayase, G. Smyrnis, T. Nguyen, R. Marten, M. Wortsman, D. Ghosh, J. Zhang, et al. Datacomp: In search of the next generation of multimodal datasets. Advances in Neural Information Processing Systems, 36:27092–27112, 2023. 15 [38]L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, et al. The Pile: An 800GB dataset of diverse text for language modeling. arXiv preprint arXiv:2101.00027, 2020. [39] J. F. Gemmeke, D. P. Ellis, D. Freedman, A. Jansen, W. Lawrence, R. C. Moore, M. Plakal, and M. Ritter. Audio set: An ontology and human-labeled dataset for audio events. In 2017 IEEE international conference on acoustics, speech and signal processing (ICASSP), pages 776–780. IEEE, 2017. [40]T. Geng, J. Zhang, Q. Wang, T. Wang, J. Duan, and F. Zheng. Longvale: Vision-audio- language-event benchmark towards time-aware omni-modal perception of long videos. arXiv preprint arXiv:2411.19772, 2024. [41] A. Ghosh, A. Acharya, S. Saha, V. Jain, and A. Chadha. Exploring the frontier of vision- language models: A survey of current methodologies and future directions. arXiv preprint arXiv:2404.07214, 2024. [42]R. Girdhar, A. El-Nouby, Z. Liu, M. Singh, K. V. Alwala, A. Joulin, and I. Misra. Imagebind: One embedding space to bind them all. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 15180–15190, 2023. [43] A. Goel, S. Ghosh, J. Kim, S. Kumar, Z. Kong, S.-g. Lee, C.-H. H. Yang, R. Duraiswami, D. Manocha, R. Valle, et al. Audio flamingo 3: Advancing audio intelligence with fully open large audio language models. arXiv preprint arXiv:2507.08128, 2025. [44] R. Goyal, S. Ebrahimi Kahou, V. Michalski, J. Materzynska, S. Westphal, H. Kim, V. Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, et al. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision, pages 5842–5850, 2017. [45] J. Gu, J. Bradbury, C. Xiong, V. O. Li, and R. Socher. Non-autoregressive neural machine translation. arXiv preprint arXiv:1711.02281, 2017. [46] L. Guo, J. Liu, X. Zhu, X. He, J. Jiang, and H. Lu. Non-autoregressive image captioning with counterfactuals-critical multi-agent learning. arXiv preprint arXiv:2005.04690, 2020. [47]C. He, Z. Jin, C. Xu, J. Qiu, B. Wang, W. Li, H. Yan, J. Wang, and D. Lin. Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models. arXiv preprint arXiv:2308.10755, 2023. [48]L. A. Hendricks, O. Wang, E. Shechtman, J. Sivic, T. Darrell, and B. Russell. Localizing moments in video with natural language, 2017. URLhttps://arxiv.org/abs/1708. 01641. [49]D. Hendrycks, S. Basart, N. Mu, S. Kadavath, F. Wang, E. Dorundo, R. Desai, T. Zhu, S. Parajuli, M. Guo, D. Song, J. Steinhardt, and J. Gilmer. The many faces of robustness: A critical analysis of out-of-distribution generalization. ICCV, 2021. [50]J. Hessel, A. Holtzman, M. Forbes, R. L. Bras, and Y. Choi. Clipscore: A reference-free evaluation metric for image captioning. arXiv preprint arXiv:2104.08718, 2021. [51]X. Hu, Z. Gan, J. Wang, Z. Yang, Z. Liu, Y. Lu, and L. Wang. Scaling up vision-language pre-training for image captioning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 17980–17989, 2022. [52] S. Huang, L. Dong, W. Wang, Y. Hao, S. Singhal, S. Ma, T. Lv, L. Cui, O. K. Mohammed, B. Patra, et al. Language is not all you need: Aligning perception with language models. Advances in Neural Information Processing Systems, 36:72096–72109, 2023. [53] X. Huang, Y. Zhang, J. Ma, W. Tian, R. Feng, Y. Zhang, Y. Li, Y. Guo, and L. Zhang. Tag2text: Guiding vision-language model via image tagging. arXiv preprint arXiv:2303.05657, 2023. 16 [54]G. Ilharco, M. Wortsman, R. Wightman, C. Gordon, N. Carlini, R. Taori, A. Dave, V. Shankar, H. Namkoong, J. Miller, H. Hajishirzi, A. Farhadi, and L. Schmidt. Openclip, July 2021. URL https://doi.org/10.5281/zenodo.5143773. If you use this software, please cite it as below. [55] C. Jia, Y. Yang, Y. Xia, Y.-T. Chen, Z. Parekh, H. Pham, Q. Le, Y.-H. Sung, Z. Li, and T. Duerig. Scaling up visual and vision-language representation learning with noisy text supervision. In International conference on machine learning, pages 4904–4916. PMLR, 2021. [56]X. Ju, Y. Gao, Z. Zhang, Z. Yuan, X. Wang, A. Zeng, Y. Xiong, Q. Xu, and Y. Shan. Miradata: A large-scale video dataset with long durations and structured captions. In A. Globerson, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. Tomczak, and C. Zhang, editors, Advances in Neural Information Processing Systems, volume 37, pages 48955–48970. Curran Associates, Inc., 2024.URLhttps://proceedings.neurips.c/paper_files/paper/2024/ file/57f6683e550eb067936c9e9f0bcb8e31-Paper-Datasets_and_Benchmarks_ Track.pdf. [57]W. Kay, J. Carreira, K. Simonyan, B. Zhang, C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green, T. Back, P. Natsev, et al. The kinetics human action video dataset. arXiv preprint arXiv:1705.06950, 2017. [58] C. D. Kim, B. Kim, H. Lee, and G. Kim. Audiocaps: Generating captions for audios in the wild. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 119–132, 2019. [59]W. Kim, C. Choi, W. Lee, and W. Rhee. An image grid can be worth a video: Zero-shot video question answering using a vlm. IEEE Access, 2024. [60] A. S. Koepke, A.-M. Oncescu, J. F. Henriques, Z. Akata, and S. Albanie. Audio retrieval with natural language queries: A benchmark study. IEEE Transactions on Multimedia, 25: 2675–2685, 2022. [61] D. Kondratyuk, L. Yu, X. Gu, J. Lezama, J. Huang, G. Schindler, R. Hornung, V. Birodkar, J. Yan, M.-C. Chiu, et al. Videopoet: A large language model for zero-shot video generation. arXiv preprint arXiv:2312.14125, 2023. [62]R. Krishna, K. Hata, F. Ren, L. Fei-Fei, and J. Carlos Niebles. Dense-captioning events in videos. In Proceedings of the IEEE international conference on computer vision, pages 706–715, 2017. [63]R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma, et al. Visual genome: Connecting language and vision using crowdsourced dense image annotations. International journal of computer vision, 123:32–73, 2017. [64] H. Kuehne, H. Jhuang, E. Garrote, T. Poggio, and T. Serre. Hmdb: a large video database for human motion recognition. In 2011 International conference on computer vision, pages 2556–2563. IEEE, 2011. [65] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica. Efficient memory management for large language model serving with pagedattention. In Proceedings of the ACM SIGOPS 29th Symposium on Operating Systems Principles, 2023. [66]LAION. Releasing re-laion 5b: transparent iteration on laion-5b with additional safety fixes. https://laion.ai/blog/relaion-5b/, 2024. Accessed: 30 aug, 2024. [67]H. Laurençon, L. Saulnier, L. Tronchon, S. Bekman, A. Singh, A. Lozhkov, T. Wang, S. Karam- cheti, A. Rush, D. Kiela, et al. Obelics: An open web-scale filtered dataset of interleaved image-text documents. Advances in Neural Information Processing Systems, 36:71683–71702, 2023. [68] S. Lee, J. Chung, Y. Yu, G. Kim, T. Breuel, G. Chechik, and Y. Song. Acav100m: Automatic curation of large-scale datasets for audio-visual video representation learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10274–10284, 2021. 17 [69]H. Li, M. Xu, Y. Zhan, S. Mu, J. Li, K. Cheng, Y. Chen, T. Chen, M. Ye, J. Wang, et al. Open- humanvid: A large-scale high-quality dataset for enhancing human-centric video generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 7752–7762, 2025. [70]J. Li, D. Li, C. Xiong, and S. Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning, pages 12888–12900. PMLR, 2022. [71]J. Li, D. Li, S. Savarese, and S. Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pages 19730–19742. PMLR, 2023. [72]L. H. Li, P. Zhang, H. Zhang, J. Yang, C. Li, Y. Zhong, L. Wang, L. Yuan, L. Zhang, J.- N. Hwang, K.-W. Chang, and J. Gao. Grounded language-image pre-training, 2022. URL https://arxiv.org/abs/2112.03857. [73]Q. Li, Z. Chen, W. Wang, W. Wang, S. Ye, Z. Jin, G. Chen, Y. He, Z. Gao, E. Cui, et al. Omnicorpus: A unified multimodal corpus of 10 billion-level images interleaved with text. arXiv preprint arXiv:2406.08418, 2024. [74]X. Li, H. Tu, M. Hui, Z. Wang, B. Zhao, J. Xiao, S. Ren, J. Mei, Q. Liu, H. Zheng, Y. Zhou, and C. Xie. What if we recaption billions of web images with llama-3?, 2024. URLhttps: //arxiv.org/abs/2406.08478. [75]Y. Li, C. Wang, and J. Jia. Llama-vid: An image is worth 2 tokens in large language models. In European Conference on Computer Vision, pages 323–340. Springer, 2024. [76] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick. Microsoft coco: Common objects in context. In Computer vision–ECCV 2014: 13th European conference, zurich, Switzerland, September 6-12, 2014, proceedings, part v 13, pages 740–755. Springer, 2014. [77] Z. Lin, C. Liu, R. Zhang, P. Gao, L. Qiu, H. Xiao, H. Qiu, C. Lin, W. Shao, K. Chen, et al. Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models. arXiv preprint arXiv:2311.07575, 2023. [78] H. Liu, C. Li, Q. Wu, and Y. J. Lee. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023. [79] H. Liu, C. Li, Y. Li, and Y. J. Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26296–26306, 2024. [80]H. Liu, C. Li, Y. Li, B. Li, Y. Zhang, S. Shen, and Y. J. Lee. Llavanext: Improved reasoning, ocr, and world knowledge, 2024. [81]J. Liu, S. Chen, X. He, L. Guo, X. Zhu, W. Wang, and J. Tang. Valor: Vision-audio-language omni-perception pretraining model and dataset. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024. [82]S. Liu, H. Cheng, H. Liu, H. Zhang, F. Li, T. Ren, X. Zou, J. Yang, H. Su, J. Zhu, et al. Llava-plus: Learning to use tools for creating multimodal agents. In European Conference on Computer Vision, pages 126–142. Springer, 2024. [83] I. Loshchilov and F. Hutter.Decoupled weight decay regularization.arXiv preprint arXiv:1711.05101, 2017. [84]H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yang, et al. Deepseek-vl: towards real-world vision-language understanding. arXiv preprint arXiv:2403.05525, 2024. [85]G. Luo, Y. Zhou, T. Ren, S. Chen, X. Sun, and R. Ji. Cheap and quick: Efficient vision- language instruction tuning for large language models. Advances in Neural Information Processing Systems, 36:29615–29627, 2023. 18 [86]C. Lyu, M. Wu, L. Wang, X. Huang, B. Liu, Z. Du, S. Shi, and Z. Tu. Macaw-llm: Multi- modal language modeling with image, audio, video, and text integration. arXiv preprint arXiv:2306.09093, 2023. [87]M. Maaz, H. Rasheed, S. Khan, and F. S. Khan. Video-chatgpt: Towards detailed video understanding via large vision and language models. arXiv preprint arXiv:2306.05424, 2023. [88]A. Marafioti, O. Zohar, M. Farré, M. Noyan, E. Bakouch, P. Cuenca, C. Zakka, L. B. Allal, A. Lozhkov, N. Tazi, V. Srivastav, J. Lochner, H. Larcher, M. Morlon, L. Tunstall, L. von Werra, and T. Wolf. Smolvlm: Redefining small and efficient multimodal models. arXiv preprint arXiv:2504.05299, 2025. [89]B. McKinzie, Z. Gan, J.-P. Fauconnier, S. Dodge, B. Zhang, P. Dufter, D. Shah, X. Du, F. Peng, A. Belyi, et al. Mm1: methods, analysis and insights from multimodal llm pre-training. In European Conference on Computer Vision, pages 304–323. Springer, 2024. [90]X. Mei, C. Meng, H. Liu, Q. Kong, T. Ko, C. Zhao, M. D. Plumbley, Y. Zou, and W. Wang. Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 32:3339–3354, 2024. [91] A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic. Howto100m: Learning a text-video embedding by watching hundred million narrated video clips. In Proceedings of the IEEE/CVF international conference on computer vision, pages 2630–2640, 2019. [92]P. Moritz, R. Nishihara, S. Wang, A. Tumanov, R. Liaw, E. Liang, M. Elibol, Z. Yang, W. Paul, M. I. Jordan, et al. Ray: a distributed framework for emerging ai applications. In Proceedings of the 13th USENIX conference on Operating Systems Design and Implementation, pages 561–577, 2018. [93] A. Nagrani, P. H. Seo, B. Seybold, A. Hauth, S. Manen, C. Sun, and C. Schmid. Learning audio-video modalities from image captions. In European Conference on Computer Vision, pages 407–426. Springer, 2022. [94]K. Nan, R. Xie, P. Zhou, T. Fan, Z. Yang, Z. Chen, X. Li, J. Yang, and Y. Tai. Openvid-1m: A large-scale high-quality dataset for text-to-video generation. In International Conference on Learning Representations, volume 2025, pages 1045–1064, 2025. [95]A.-M. Oncescu, A. Koepke, J. F. Henriques, Z. Akata, and S. Albanie. Audio retrieval with natural language queries. In INTERSPEECH, 2021. [96]OpenAI.CLIP ViT-L/14-336 model card.https://huggingface.co/openai/ clip-vit-large-patch14-336, 2021. Accessed: 2025-05-14. [97] OpenAI. GPT-4o System Card.https://cdn.openai.com/gpt-4o-system-card.pdf, Apr. 2024. Accessed: 2024-04-12. [98]G. Penedo, Q. Malartic, D. Hesslow, R. Cojocaru, A. Cappelli, H. Alobeidli, B. Pannier, E. Almazrouei, and J. Launay. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only. arXiv preprint arXiv:2306.01116, 2023. URL https://arxiv.org/abs/2306.01116. [99]G. Penedo, H. Kydlí ˇ cek, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The fineweb datasets: Decanting the web for the finest text data at scale. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024. URL https://openreview.net/forum?id=n6SCkn2QaG. [100] Z. Peng, W. Wang, L. Dong, Y. Hao, S. Huang, S. Ma, and F. Wei. Kosmos-2: Grounding multimodal large language models to the world. arXiv preprint arXiv:2306.14824, 2023. [101]H. Pham, Z. Dai, G. Ghiasi, K. Kawaguchi, H. Liu, A. W. Yu, J. Yu, Y.-T. Chen, M.-T. Luong, Y. Wu, et al. Combined scaling for zero-shot transfer learning. Neurocomputing, 555:126658, 2023. 19 [102]A. Piergiovanni, I. Noble, D. Kim, M. S. Ryoo, V. Gomes, and A. Angelova. Mirasol3b: A multimodal autoregressive model for time-aligned and contextual modalities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26804–26814, 2024. [103] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learning, pages 8748–8763. PmLR, 2021. [104] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492–28518. PMLR, 2023. [105]C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research, 21(140):1–67, 2020. [106]B. Recht, R. Roelofs, L. Schmidt, and V. Shankar. Do imagenet classifiers generalize to imagenet? In International conference on machine learning, pages 5389–5400. PMLR, 2019. [107]A. Rohrbach, A. Torabi, M. Rohrbach, N. Tandon, C. Pal, H. Larochelle, A. Courville, and B. Schiele. Movie description. International Journal of Computer Vision, 123:94–120, 2017. [108]R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 10684–10695, 2022. [109]J. Salamon, C. Jacoby, and J. P. Bello. A dataset and taxonomy for urban sound research. In 22nd ACM International Conference on Multimedia (ACM-M’14), pages 1041–1044, Orlando, FL, USA, Nov. 2014. [110]R. Sanabria, O. Caglayan, S. Palaskar, D. Elliott, L. Barrault, L. Specia, and F. Metze. How2: a large-scale dataset for multimodal language understanding. arXiv preprint arXiv:1811.00347, 2018. [111]C. Schuhmann, R. Vencu, R. Beaumont, R. Kaczmarczyk, C. Mullis, A. Katta, T. Coombes, J. Jitsev, and A. Komatsuzaki. Laion-400m: Open dataset of clip-filtered 400 million image-text pairs. arXiv preprint arXiv:2111.02114, 2021. [112] C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al. Laion-5b: An open large-scale dataset for training next generation image-text models. Advances in neural information processing systems, 35: 25278–25294, 2022. [113] P. Sharma, N. Ding, S. Goodman, and R. Soricut. Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2556–2565, 2018. [114]O. Sidorov, R. Hu, M. Rohrbach, and A. Singh. Textcaps: a dataset for image captioning with reading comprehension. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16, pages 742–758. Springer, 2020. [115] G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. Gupta. Hollywood in homes: Crowdsourcing data collection for activity understanding. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part I 14, pages 510–526. Springer, 2016. [116]G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari. Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626, 2018. 20 [117]M. Soldan, A. Pardo, J. L. Alcázar, F. Caba, C. Zhao, S. Giancola, and B. Ghanem. Mad: A scalable dataset for language grounding in videos from movie audio descriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 5026–5035, 2022. [118]K. Soomro, A. R. Zamir, and M. Shah. Ucf101: A dataset of 101 human actions classes from videos in the wild. arXiv preprint arXiv:1212.0402, 2012. [119]K. Srinivasan, K. Raman, J. Chen, M. Bendersky, and M. Najork. Wit: Wikipedia-based image text dataset for multimodal multilingual machine learning. In Proceedings of the 44th international ACM SIGIR conference on research and development in information retrieval, pages 2443–2449, 2021. [120] J. C. Stroud, Z. Lu, C. Sun, J. Deng, R. Sukthankar, C. Schmid, and D. A. Ross. Learning video representations from textual web supervision. arXiv preprint arXiv:2007.14937, 2020. [121]Y. Su, T. Lan, H. Li, J. Xu, Y. Wang, and D. Cai. Pandagpt: One model to instruction-follow them all. arXiv preprint arXiv:2305.16355, 2023. [122] L. Sun, X. Xu, M. Wu, and W. Xie. A large-scale dataset for audio-language representation learning. arXiv preprint arXiv:2309.11500, 6, 2023. [123]R. Sun, Y. Zhang, T. Shah, J. Sun, S. Zhang, W. Li, H. Duan, B. Wei, and R. Ranjan. From sora what we can see: A survey of text-to-video generation. arXiv preprint arXiv:2405.10674, 2024. [124]Z. Tan, X. Yang, L. Qin, and H. Li. Vidgen-1m: A large-scale dataset for text-to-video generation. arXiv preprint arXiv:2408.02629, 2024. [125]Z. Tang, Z. Yang, C. Zhu, M. Zeng, and M. Bansal. Any-to-any generation via composable diffusion. Advances in Neural Information Processing Systems, 36:16083–16099, 2023. [126]Z. Tang, Z. Yang, M. Khademi, Y. Liu, C. Zhu, and M. Bansal. Codi-2: In-context interleaved and interactive any-to-any generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 27425–27434, 2024. [127] G. Team, P. Georgiev, V. I. Lei, R. Burnell, L. Bai, A. Gulati, G. Tanzer, D. Vincent, Z. Pan, S. Wang, et al. Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context. arXiv preprint arXiv:2403.05530, 2024. [128]B. Thomee, D. A. Shamma, G. Friedland, B. Elizalde, K. Ni, D. Poland, D. Borth, and L.-J. Li. Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016. [129]Z. Tong, Y. Song, J. Wang, and L. Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35:10078–10093, 2022. [130] Vimeo. Vimeo. https://vimeo.com/. Accessed: 2025-05-14. [131]O. Vinyals, A. Toshev, S. Bengio, and D. Erhan. Show and tell: Lessons learned from the 2015 mscoco image captioning challenge. IEEE transactions on pattern analysis and machine intelligence, 39(4):652–663, 2016. [132]C. Wang, Y. Zhu, Y. Xu, J. Yang, Z. Yan, Y. Wang, Y. Wang, and L. Wang. Internvideo-next: Towards general video foundation models without video-text supervision. arXiv preprint arXiv:2512.01342, 2025. [133]H. Wang, S. Ge, Z. Lipton, and E. P. Xing. Learning robust global representations by penalizing local predictive power. Advances in neural information processing systems, 32, 2019. [134]Q. Wang, Y. Shi, J. Ou, R. Chen, K. Lin, J. Wang, B. Jiang, H. Yang, M. Zheng, X. Tao, et al. Koala-36m: A large-scale video dataset improving consistency between fine-grained conditions and video content. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 8428–8437, 2025. 21 [135]W. Wang, M. Shi, Q. Li, W. Wang, Z. Huang, L. Xing, Z. Chen, H. Li, X. Zhu, Z. Cao, et al. The all-seeing project: Towards panoptic visual recognition and understanding of the open world. arXiv preprint arXiv:2308.01907, 2023. [136]W. Wang, H. Yang, Z. Tuo, H. He, J. Zhu, J. Fu, and J. Liu. Videofactory: Swap attention in spatiotemporal diffusions for text-to-video generation. arXiv preprint arXiv:2305.10874, 2023. [137]W. Wang, Q. Lv, W. Yu, W. Hong, J. Qi, Y. Wang, J. Ji, Z. Yang, L. Zhao, S. XiXuan, et al. Cogvlm: Visual expert for pretrained language models. Advances in Neural Information Processing Systems, 37:121475–121499, 2024. [138]W. Wang, H. Yang, Z. Tuo, H. He, J. Zhu, J. Fu, and J. Liu. Swap attention in spatiotemporal diffusions for text-to-video generation, 2024. URLhttps://arxiv.org/abs/2305.10874. [139]X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang. Vatex: A large-scale, high- quality multilingual dataset for video-and-language research. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4581–4591, 2019. [140]X. Wang, I. Alabdulmohsin, D. Salz, Z. Li, K. Rong, and X. Zhai. Scaling pre-training to one hundred billion data for vision language models. arXiv preprint arXiv:2502.07617, 2025. [141]Y. Wang, Y. He, Y. Li, K. Li, J. Yu, X. Ma, X. Li, G. Chen, X. Chen, Y. Wang, et al. Internvid: A large-scale video-text dataset for multimodal understanding and generation. arXiv preprint arXiv:2307.06942, 2023. [142] M. Weber, D. Y. Fu, Q. Anthony, Y. Oren, S. Adams, A. Alexandrov, X. Lyu, H. Nguyen, X. Yao, V. Adams, B. Athiwaratkun, R. Chalamala, K. Chen, M. Ryabinin, T. Dao, P. Liang, C. Ré, I. Rish, and C. Zhang. Redpajama: an open dataset for training large language models. NeurIPS Datasets and Benchmarks Track, 2024. [143]M. Wortsman, G. Ilharco, J. W. Kim, M. Li, S. Kornblith, R. Roelofs, R. G. Lopes, H. Hajishirzi, A. Farhadi, H. Namkoong, et al. Robust fine-tuning of zero-shot models. In 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7949–7961. IEEE, 2022. [144]S. Wu, H. Fei, L. Qu, W. Ji, and T.-S. Chua. Next-gpt: Any-to-any multimodal llm. In Forty-first International Conference on Machine Learning, 2024. [145]W. Wu, K. Zheng, S. Ma, F. Lu, Y. Guo, Y. Zhang, W. Chen, Q. Guo, Y. Shen, and Z. Zheng-Jun. Lotlip: Improving language-image pre-training for long text understanding. arXiv, 2024. [146]Y. Wu, K. Chen, T. Zhang, Y. Hui, T. Berg-Kirkpatrick, and S. Dubnov. Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation. In IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP, 2023. [147]Z. Wu, X. Chen, Z. Pan, X. Liu, W. Liu, D. Dai, H. Gao, Y. Ma, C. Wu, B. Wang, Z. Xie, Y. Wu, K. Hu, J. Wang, Y. Sun, Y. Li, Y. Piao, K. Guan, A. Liu, X. Xie, Y. You, K. Dong, X. Yu, H. Zhang, L. Zhao, Y. Wang, and C. Ruan. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding, 2024. URLhttps://arxiv.org/abs/ 2412.10302. [148]T. Xiong, Y. Wang, D. Zhou, Z. Lin, J. Feng, and X. Liu. Lvd-2m: A long-take video dataset with temporally dense captions, 2024. URL https://arxiv.org/abs/2410.10816. [149] H. Xu, G. Ghosh, P.-Y. Huang, D. Okhonko, A. Aghajanyan, F. Metze, L. Zettlemoyer, and C. Feichtenhofer. Videoclip: Contrastive pre-training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084, 2021. [150]H. Xu, S. Xie, X. E. Tan, P.-Y. Huang, R. Howes, V. Sharma, S.-W. Li, G. Ghosh, L. Zettle- moyer, and C. Feichtenhofer. Demystifying clip data. arXiv preprint arXiv:2309.16671, 2023. 22 [151]H. Xu, Q. Ye, X. Wu, M. Yan, Y. Miao, J. Ye, G. Xu, A. Hu, Y. Shi, G. Xu, C. Li, Q. Qian, M. Que, J. Zhang, X. Zeng, and F. Huang. Youku-mplug: A 10 million large-scale chinese video-language dataset for pre-training and benchmarks, 2023. URLhttps://arxiv.org/ abs/2306.04362. [152]J. Xu, T. Mei, T. Yao, and Y. Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 6 2016. [153] J. Xu, Z. Guo, H. Hu, Y. Chu, X. Wang, J. He, Y. Wang, X. Shi, T. He, X. Zhu, Y. Lv, Y. Wang, D. Guo, H. Wang, L. Ma, P. Zhang, X. Zhang, H. Hao, Z. Guo, B. Yang, B. Zhang, Z. Ma, X. Wei, S. Bai, K. Chen, X. Liu, P. Wang, M. Yang, D. Liu, X. Ren, B. Zheng, R. Men, F. Zhou, B. Yu, J. Yang, L. Yu, J. Zhou, and J. Lin. Qwen3-omni technical report. arXiv preprint arXiv:2509.17765, 2025. [154]J. Xu, C. Thomé, D. Horak, W. Xie, and A. Zisserman. Scaling audio-text retrieval with multimodal large language models. arXiv preprint arXiv:2602.18010, 2026. [155]H. Xue, T. Hang, Y. Zeng, Y. Sun, B. Liu, H. Yang, J. Fu, and B. Guo. Advancing high- resolution video-language representation with large-scale video transcriptions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5036–5045, 6 2022. [156] Z. Xue, J. An, X. Yang, and K. Grauman. Progress-aware video frame captioning. arXiv preprint arXiv:2412.02071, 2024. [157]S. Yan, T. Zhu, Z. Wang, Y. Cao, M. Zhang, S. Ghosh, Y. Wu, and J. Yu. Videococa: Video-text modeling with zero-shot transfer from contrastive captioners. arXiv preprint arXiv:2212.04979, 2022. [158]T. Yang, S. Ren, A. Yuille, and F. Wang. Vimix-14m: A curated multi-source video-text dataset with long-form, high-quality captions and crawl-free access. arXiv preprint arXiv:2511.18382, 2025. [159] H. You, H. Zhang, Z. Gan, X. Du, B. Zhang, Z. Wang, L. Cao, S.-F. Chang, and Y. Yang. Ferret: Refer and ground anything anywhere at any granularity. arXiv preprint arXiv:2310.07704, 2023. [160] Q. You, H. Jin, Z. Wang, C. Fang, and J. Luo. Image captioning with semantic attention. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4651–4659, 2016. [161] YouTube. Youtube. https://w.youtube.com/. Accessed: 2025-05-14. [162] yt dlp. yt-dlp. https://github.com/yt-dlp/yt-dlp, 2021. [163] Y. Yuan, D. Jia, X. Zhuang, Y. Chen, Z. Chen, Y. Wang, Y. Wang, X. Liu, X. Kang, M. D. Plumbley, et al. Sound-vecaps: Improving audio generation with visually enhanced cap- tions. In ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5. IEEE, 2025. [164] M. Zaharia, R. S. Xin, P. Wendell, T. Das, M. Armbrust, A. Dave, X. Meng, J. Rosen, S. Venkataraman, M. J. Franklin, A. Ghodsi, J. Gonzalez, S. Shenker, and I. Stoica. Apache spark: a unified engine for big data processing. Commun. ACM, 59(11):56–65, Oct. 2016. ISSN 0001-0782. doi: 10.1145/2934664. URL https://doi.org/10.1145/2934664. [165]R. Zellers, X. Lu, J. Hessel, Y. Yu, J. S. Park, J. Cao, A. Farhadi, and Y. Choi. Merlot: Multimodal neural script knowledge models.In M. Ranzato, A. Beygelz- imer, Y. Dauphin, P. Liang, and J. W. Vaughan, editors, Advances in Neural Infor- mation Processing Systems, volume 34, pages 23634–23651. Curran Associates, Inc., 2021.URLhttps://proceedings.neurips.c/paper_files/paper/2021/file/ c6d4eb15f1e84a36eff58eca3627c82e-Paper.pdf. 23 [166]H. Zhang, X. Li, and L. Bing. Video-llama: An instruction-tuned audio-visual language model for video understanding. arXiv preprint arXiv:2306.02858, 2023. [167]Y. Zhang, B. Li, h. Liu, Y. j. Lee, L. Gui, D. Fu, J. Feng, Z. Liu, and C. Li. Llava-next: A strong zero-shot video understanding model, 4 2024. URLhttps://llava-vl.github. io/blog/2024-04-30-llava-next-video/. [168]Z. Zhao, L. Guo, T. Yue, S. Chen, S. Shao, X. Zhu, Z. Yuan, and J. Liu. Chatbridge: Bridging modalities with large language model as a language catalyst. arXiv preprint arXiv:2305.16103, 2023. [169]L. Zhou, C. Xu, and J. Corso. Towards automatic learning of procedures from web instructional videos. In Proceedings of the AAAI conference on artificial intelligence, volume 32, 2018. [170]D. Zhu, J. Chen, X. Shen, X. Li, and M. Elhoseiny. Minigpt-4: Enhancing vision-language understanding with advanced large language models. arXiv preprint arXiv:2304.10592, 2023. [171]W. Zhu, J. Hessel, A. Awadalla, S. Y. Gadre, J. Dodge, A. Fang, Y. Yu, L. Schmidt, W. Y. Wang, and Y. Choi. Multimodal c4: An open, billion-scale corpus of images interleaved with text. Advances in Neural Information Processing Systems, 36:8958–8974, 2023. 24 A Additional Experimental Details A.1 Video-Language (ViCLIP) In this section, we present additional details about our video-language experiments. A.1.1 Experimental setup ViCLIP[141] models consist of two components, a video encoder and a text encoder. The video and text encoders are initialized from a pre-trained CLIP model’s vision and text encoder respectively, with a modification of the vision encoder to process a set of frames instead of a single image, while the text encoder remains unchanged. The main architectural change compared to standard CLIP is that the vision encoder receives patches from all the frames, on which we apply both spatial and temporal positional embeddings. For attention layers, any patch token can attend to any patch token across the frames. Training is performed using standard InfoNCE loss[96] computed using video and text embeddings. We train our ViCLIP models by initializing the video frame encoder and text encoder from CLIP pre-trained checkpoints on DataComp-1B [37]. We uniformly sample 8 frames from each video following [141]. For every training run in Tables 4 and 5, we performed learning rate hyperparameter sweeps covering1e-6,4e-6,1e-5,4e-5,1e-4,4e-4. We also experimented with two global batch sizes of8kand32k. For optimization, we used AdamW[83] withβ 1 = 0.9andβ 2 = 0.98, a weight decay of0.2, gradient clipping with a norm of1. We used the standard cosine learning rate schedule with a warmup of 100 in all experiments. We trained our models on 100 NVIDIA GH200 GPUs. To construct a validation set, we randomly sampled a maximum of 10K examples from the training sets of each downstream dataset (K400, UCF-101, HMDB51, MSR-VTT, MSVD) and evaluated all our trained models on the validation sets. We then used the average metric (defined in Section 4.1) on the validation sets for hyper-parameter selection. In Table 11, we show the best hyperparameters used to train our ViCLIP models and report the final test metrics of each task and the overall average in Tables 4 and 5. Table 11: Hyperparameters used for training ViCLIP models. ModelDatasetSamples seenBatch SizeLRWDFLOPsGPU-hours B-32BVD-V-55M10M32k1e-40.27.4e1712 B-32BVD-V-55M50M32k1e-40.23.7e1865 B-16BVD-V-55M10M32k1e-40.22.7e1825 B-16BVD-V-55M50M32k4e-50.21.4e19125 L-14BVD-V-10M10M32k1e-40.21.3e1990 L-14BVD-V-10M50M32k4e-050.26.3e19455 L-14BVD-V-50M10M32k1e-40.21.3e1990 L-14BVD-V-55M10M32k1e-40.21.3e1990 L-14BVD-V-50M50M32k4e-050.26.3e19455 L-14BVD-V-55M50M32k4e-050.26.3e19455 A.1.2 Additional results Statistical uncertainty analysis. To assess whether observed differences exceed run-to-run variation, Table 12 reports mean, standard deviation, and 95% confidence intervals across runs for BVD-V- 10M and BVD-V-50M at fixed training compute (50M samples seen). While several metrics show only minor differences within the uncertainty range, we observe statistically clear improvements on HMDB51, MSR-VTT V2T, MSVD T2V, and the overall average metric when training on the larger dataset. Overlap and decontamination analysis.. We perform an overlap analysis between BVD-V-55M and the K400, MSVD, and MSR-VTT test sets using unique YouTube video IDs which is shown in Table 13. We then construct decontaminated versions of the downstream test sets by removing all examples whose YouTube video IDs are present in BVD-V-55M. The resulting test-set sizes are shown in Table 14. For MSR-VTT and MSVD, the 885 and 522 entries in Table 13 denote the number of unique YouTube video IDs, whereas the corresponding test sets contain 1,000 and 670 examples, 25 Table 12: Comparison between BVD-V-10M and BVD-V-50M at fixed training compute (50M samples seen). We report mean and standard deviation across five runs with different seeds together with the performance difference and corresponding 95% confidence interval. Larger datasets primarily improve retrieval performance and HMDB51 classification, while gains on K400 and UCF-101 remain within the estimated uncertainty range. BenchmarkMetricBVD-V-10MBVD-V-50MDifference95% CIConclusion K400top-163.78 ±0.15 63.84 ±0.17 +0.06[-0.20, +0.32]not clear UCF-101top-180.28 ±0.37 80.42 ±0.25 +0.14[-0.35, +0.63]not clear HMDB51top-155.50 ±0.23 57.48 ±0.22 +1.98[+1.54, +2.42]clear improvement MSR-VTTT2V42.58 ±0.39 42.62 ±0.20 +0.04[-0.49, +0.57]not clear V2T42.32 ±0.26 43.28 ±0.54 +0.96[+0.06, +1.86]clear improvement MSVDT2V53.30 ±0.07 53.52 ±0.11 +0.22[+0.06, +0.38]clear improvement V2T83.32 ±0.46 83.76 ±0.79 +0.44[-0.78, +1.66]not clear OverallAvg60.94 ±0.11 61.52 ±0.18 +0.58[+0.40, +0.76]clear improvement Table 13: Overlap of BVD-V-55M with evaluation datasets. We report the number and percentage of unique YouTube video IDs from K400, MSR-VTT, and MSVD that overlap with BVD-V-55M. DatasetOverlapUnique VideosOverlap (%) K40010339,8050.26 MSR-VTT498855.5 MSVD335226.3 respectively. This discrepancy arises because multiple test examples may correspond to clips from the same source video. We re-evaluate our models on the decontaminated MSR-VTT and MSVD test sets and compare the results with those obtained on the original test sets in Tables 15 and 16 respectively. Overall, performance on the decontaminated test sets remains comparable to that on the original test sets, indicating that the observed overlap with BVD-V-55M does not materially affect the reported results. Human audit of video captions. To provide a direct assessment of video-caption faithfulness, we conducted a human audit of 134 captioned videos. We classified each caption as: (i) accurate, when it faithfully described the video; (i) minor error, when it was broadly correct but contained a localized mistake, such as an incorrect action, object, or color; or (i) major error, when it substantially misrepresented the video. The results are as follows: • Accurate: 106/134 (79.1%) • Minor error: 25/134 (18.7%) • Major error: 3/134 (2.2%) • No major error: 131/134 (97.8%) The observed errors involved incorrect actions (17 cases), objects (7), or scene descriptions (4). We emphasize that a primary goal of this study is to validate the quality of the dataset itself. The audit indicates that most video captions are faithful, while the downstream improvements over InternVid (Table 5) suggest that they provide useful training supervision. A.2 Audio-Language (CLAP) A.2.1 Experimental setup We validate the audio component of LAION-BVD using CLAP [146], a contrastive audio-language model trained on paired audio-caption data. Audio waveforms are resampled to48 kHzand converted into log Mel spectrogram representations before being processed by the audio encoder. Text captions are tokenized and encoded using the CLAP text encoder. The audio and text embeddings are projected into a shared embedding space and trained using a symmetric contrastive loss over matched and mismatched audio-text pairs. We train models on BVD- A-1.7M and BVD-A-10M to study scaling behavior across both dataset and model size. Table 17 reports the computational training cost of different CLAP backbones at a fixed number of training samples seen. 26 Table 14: Evaluation dataset sizes after decontamination. We report the original and decontami- nated test-set sizes after removing examples whose YouTube video IDs are present in BVD-V-55M, together with the number of removed examples (∆). DatasetOriginalDecontaminated∆ K40038,68638,583103 MSVD67063139 MSR-VTT1,00094357 Table 15: Decontamination has minimal impact on MSR-VTT retrieval performance. We compare text-to-video (T2V) and video-to-text (V2T) retrieval performance on the original and decontaminated MSR-VTT test sets across pretraining datasets, model sizes, and training compute. ModelPretraining DatasetSamples SeenT2V2T OriginalDecont.OriginalDecont. ViCLIP ViT-L/14BVD-V-10M10M42.742.342.942.9 ViCLIP ViT-L/14BVD-V-10M50M41.341.043.142.9 ViCLIP ViT-L/14BVD-V-50M10M43.543.342.642.7 ViCLIP ViT-L/14BVD-V-50M50M42.842.543.243.4 ViCLIP ViT-B/32BVD-V-55M10M34.234.538.138.5 ViCLIP ViT-B/32BVD-V-55M50M35.635.739.039.4 ViCLIP ViT-B/16BVD-V-55M10M36.036.140.040.2 ViCLIP ViT-B/16BVD-V-55M50M37.737.840.841.1 ViCLIP ViT-L/14BVD-V-55M10M43.643.644.344.0 ViCLIP ViT-L/14BVD-V-55M50M42.241.843.843.7 Table 16: Decontamination has minimal impact on MSVD retrieval performance. We compare text-to-video (T2V) and video-to-text (V2T) retrieval performance on the original and decontaminated MSVD test sets across pretraining datasets, model sizes, and training compute. ModelPretraining DatasetSamples SeenT2V2T OriginalDecont.OriginalDecont. ViCLIP ViT-L/14BVD-V-10M10M53.253.181.080.8 ViCLIP ViT-L/14BVD-V-10M50M53.253.183.083.0 ViCLIP ViT-L/14BVD-V-50M10M53.653.582.181.7 ViCLIP ViT-L/14BVD-V-50M50M53.753.583.983.8 ViCLIP ViT-B/32BVD-V-55M10M44.244.074.874.7 ViCLIP ViT-B/32BVD-V-55M50M46.346.377.777.9 ViCLIP ViT-B/16BVD-V-55M10M48.648.679.079.0 ViCLIP ViT-B/16BVD-V-55M50M49.449.381.380.9 ViCLIP ViT-L/14BVD-V-55M10M53.753.683.483.0 ViCLIP ViT-L/14BVD-V-55M50M53.653.584.084.0 27 Evaluation follows standard CLAP protocols using zero-shot classification on UrbanSound8K and retrieval benchmarks on AudioCaps and Clotho. We report text-to-audio and audio-to-text Recall@5 metrics. Table 17: Computational training cost across different CLAP backbones at a fixed number of training samples seen. ModelParamsSamplesFLOPs per sampleTotal FLOPs HTSAT-tiny-Roberta-base158M28.2M1.1e119.4e18 HTSAT-base-Roberta-base200M27.1M1.3e111.1e19 HTSAT-tiny-Roberta-large388M27.1M3.5e112.9e19 HTSAT-base-Roberta-large431M28.2M3.7e113.1e19 A.2.2 Additional results In this section, we present additional results to complement our audio-language validation. Table 18 compares our CLAP pipeline to the original LAION-CLAP model trained on the same dataset. Our implementation consistently achieves higher performance, showing the correctness of our training pipeline. Table 18: Audio-language comparison at matched 158M model scale using different training datasets. Our open CLAP pipeline reproduces and exceeds the published LAION-CLAP [146] baseline at comparable or lower training compute. Replacing LAION-Audio with BVD-A-1.7M remains competitive overall and is particularly strong on UrbanSound8K and AudioCaps retrieval. TrainsetD (K)S (M)C (FLOPs)AvgUS8KAC aR@5AC tR@5Cl aR@5Cl tR@5 LAION-CLAP (AC+CL+LA+AS)2621118 3.9× 10 19 61.473.270.579.539.044.6 AC+CL+LA59027.1 9.1× 10 18 62.278.669.074.741.347.5 AC+CL+LA+AS262127.0 9.0× 10 18 66.479.675.585.941.149.7 AC+CL+BVD-A-1.7M176728.2 9.4× 10 18 60.676.270.380.432.143.8 A.2.3 Human audit of audio captions We conducted a small human audit of the audio captions using 134 randomly sampled captioned clips. The results are: • Accurate: 106/134 (79.1%) • Minor error: 20/134 (14.9%) • Major error: 8/134 (6.0%) • No major error: 126/134 (94.0%) The main hallucination types involved incorrect sound events (10 cases), speech content (5), music descriptions (4), and speaker attributes (1). These results show that most audio captions are faithful, while also revealing that audio captioning remains more prone to substantial errors than video captioning. A.2.4 Contamination analysis of Audio Flamingo 3 with downstream test sets We cross-reference the released Audio Flamingo 3 (AF3) training manifests with the AudioCaps and Clotho test sets. For Clotho, we find no test-set overlap; the 367 matching clips all belong to the Clotho training split. For AudioCaps, 356 of 975 test video IDs occur in the AF3 AudioSkills subset. However, these examples are paired with AudioSet-derived captions rather than the AudioCaps reference captions used for evaluation. Thus, AF3 was exposed to a subset of AudioCaps test audio, but not to its evaluation captions. To test whether this overlap affects retrieval performance, we re-evaluate all 140 trained runs (2,118 checkpoints) on an AF3-decontaminated AudioCaps subset. Of the 975 test clips, 619 are absent 28 Table 19: Removing AF3-seen AudioCaps clips does not disproportionately affect BVD-trained models. We report the mean change in audio-to-text (aR@5) and text-to-audio (tR@5) retrieval when evaluating on 607 AF3-clean AudioCaps clips relative to a fixed, randomly sampled control set of the same size. Negative values indicate lower performance on the AF3-clean subset. We additionally indicate whether each configuration was trained on BVD and whether its retrieval training data include AudioCaps or AudioSet-derived data. Training DataaR@5tR@5BVDAC/AS-derived BVD-A-1.7M−0.39 −4.50YesNo BVD-A-10M−0.81 −2.52YesNo AC+CL+BVD-A-1.7M −2.40 −3.44YesYes LA+0.97 −1.48NoNo AC+CL+LA−2.98 −3.52NoYes AC+CL+LA+AS −3.19 −2.91NoYes LA+AS−3.49 −3.85NoYes BVD-trained (pooled) −1.20 −3.49– Non-BVD (pooled) −2.56 −3.09– Table 20: BVD exhibits a broader caption-length distribution than the AudioCaps and Clotho test sets. We report the percentage of captions falling into different caption-length intervals, measured in words. Dataset< 55–1010–1515–2020–2525–30 > 30 BVD-A-1.7M2.0%34.7%25.4%16.0%11.8%6.0%4.0% AudioCaps (Test)10.6%48.3%24.3%11.0%4.4%1.3%0.2% Clotho (Test)0.0%44.8%44.6%10.6%0.0%0.0%0.0% from AF3 training and 607 are available in our evaluation shards. Since reducing the retrieval gallery from 963 to 607 clips itself increases R@5 by approximately 8.3 percentage points, we compare the 607 clean clips with a fixed randomly sampled control set of the same size. The resulting decontamination effects are shown in Table 19. BVD-trained models are not more affected by removing AF3-seen examples; pooled BVD models show a smaller reduction than non-BVD models. Instead, the effect is larger for retrieval models trained directly on AudioCaps or AudioSet-derived data, independent of whether BVD is included. This suggests that AF3 overlap does not explain the retrieval gains observed for BVD. Caption-style analysis.. We further compare caption statistics across BVD, AudioCaps, and Clotho (Table 20). AudioCaps and Clotho have similar mean caption lengths (approximately 9 and 11 words), while BVD has a broader length distribution. If AF3 recaptioning transferred benchmark-specific caption style to BVD, we would also expect improved performance on Clotho, whose training captions were available to AF3 as metadata. Instead, BVD underperforms LA on Clotho. The differing retrieval behavior is therefore more consistent with differences in benchmark construction and caption style than with contamination. A.3 Image-Text (Frame-based CLIP) A.3.1 Experimental setup For training CLIP models we used the OpenCLIP [54] codebase. We trained models on different datasets with the same setup and hyperparameters outlined in Table 21. For synthetic frame-caption generation, thevLLM[65] framework was used, alongside a Ray [92] cluster to distribute the workload across nodes. We used 256 NVIDIA-A100 GPUs to efficiently produce captions for 300M frames. Since frames extracted from LAION-BVD have a higher resolution than DataComp images (720px vs256px), we could not achieve the same throughput (8.3 img/s). The number of GPU hours for CLIP training and captioning is summarized in Tab. 22. The hyperparameters used for captioning can be found in Tab. 23 29 Table 21: Hyperparameters for CLIP ViT-B/32 and ViT-B/16. ModelSamples SeenLR SchedulerWarmup StepsLRGPUsBatch Size ViT-B/32 12.8Mcosine40000.00116 GPUs A1002048 30.7Mcosine30000.00116 GPUs A1004096 64.0Mcosine40000.00216 GPUs A1008192 128.0Mcosine40000.00264 GPUs A1008192 307.2Mcosine80000.00264 GPUs A10016384 ViT-B/16 12.8Mcosine40000.00116 GPUs A1002048 30.7Mcosine40000.00216 GPUs A1004096 64.0Mcosine40000.00216 GPUs A1008192 128.0Mcosine60000.00264 GPUs A1008192 307.2Mcosine40000.00264 GPUs A10016384 ExperimentA100 Hours CLIP Training2187 Captioning12200 Table 22: Total GPU hours consumed A.3.2 Captioning We select a VLM that balances quality and efficiency, and then apply it to generate frame-level captions. We validate the effectiveness of our frame captions through downstream performance on standard benchmarks and explore the impact of caption length on the performance. Selecting a captioning model. We considered models for image captioning that balance quality and throughput, selecting the seven top-performing models from the OpenVLM leaderboard [35] as our starting point. Due to practical constraints on overall compute and GPU memory, we limited our selection to VLMs with fewer than 4B parameters. The resulting candidate pool is provided in Table 24. We use CLIPScore [50] as a proxy to measure the quality of generated captions. Specifically, we calculate CLIPScore using OpenAI CLIP L-14-336 [96] on a representative subset of generated captions for 100k images from DataComp-1B [37]. We select DeepSeek-VL2-tiny [147] for captioning our dataset, as it yielded the highest CLIPScore and second-best throughput. Validation. To assess caption quality, we recaption subsets of DataComp-1B [37] using DeepSeek- VL2-tiny and train CLIP-ViT-B-32. For validation, we consider the zero-shot image classification performance on ImageNet-1k [30], ImageNet-R [49], and ImageNet-Sketch [133] alongside image retrieval performance on COCO Captions [76]. The results are summarized in Table 25. While our automatically generated captions lead to a drop in classification accuracy (36 %for ImageNet,11 % and16 %for ImageNet-R and ImageNet-Sketch respectively at the largest data scale), image retrieval performance increases by27 %compared to the original captions. Overall, we find the caption quality acceptable considering the reduced annotation cost and increased dataset scale. Caption length. We explore different input prompts for captioning and consequently the impact of different resulting caption lengths on downstream model performance. Specifically, we use the following prompt as our default choice:Provideaverycoarsebriefsinglelineofcaption fortheimage.We compare this to using the following prompt, which resulted in caption lengths limited to around 7 words on average (short):Provideaverycoarsebriefsinglelineof captionforthemainobjectintheimage.Don’tworryaboutthedetails,justavery high-leveldescription.Thedescriptionshouldbeasshortaspossible-1-3words. The impact on model performance of these short captions is also included in Table 25. While we observe a small improvement in classification accuracy on 30.7M training samples, we ultimately decide against artificially constraining the caption length. A.3.3 Frame filtering Table 26 shows the number of frames extracted and the corresponding percentage of videos. 30 HyperparameterValue Temperature0.5 Max Model Length4096 Max Sequences16 Max Tokens20 Table 23: Frame-Level Captioning Hyperparameters Table 24: Candidate captioner models. We choose DeepSeek-VL2-tiny, which achieves the highest CLIPScore and thesecond-highest throughput. Throughput ModelLanguage ModelVision ModelParamsCLIPScorein img/s InternVL2.5-1B [22]Qwen-2.5-0.5BInternViT-300M-v2.51 B0.4310.41 InternVL2.5-2B-MPO [22]InternLM2.5-1.8BInternViT-300M-v2.52 B0.276.41 InternVL2.5-2B [22]InternLM2.5-1.8BInternViT-300M-v2.52 B0.216.21 SmolVLM-Instruct [88]SmolLM2-1.7BSigLIP-400M2.3 B0.502.71 DeepSeek-VL2-tiny [147]DeepSeekMoE-3BSigLIP-400M3.4 B0.628.20 Qwen2.5-VL-3B-Instruct [5]Qwen2.5-3BQwenViT3.75 B0.552.46 Phi-3.5-vision-instruct [1]Phi-3.5-mini-instructOpenAI CLIP L-14-3364.15 B0.593.74 B Why Strong Retrieval Does Not Translate to ImageNet-1k Accuracy Here we want to expand our analysis of the surprisingly strong retrieval performance (and suboptimal ImageNet-1k classification accuracy) on the filtered LAION-BVD captioned frames as shown in Table 10. Table 27 shows that models trained on LAION-BVD subsets also achieve strong Flickr30k image-to-text retrieval performance (up to 0.87 R@5). Additionally, we analyze both the caption length distribution and the coverage of ImageNet-1k class names in two 300M-sample subsets: (i) synthetic captions generated for LAION-BVD and (i) original (alt-text) captions from DataComp. Caption length distribution. Figure 6 shows the frequency distribution of the number of tokens per caption. Captions in the LAION-BVD subset exhibit a noticeably narrower distribution, with longer captions dominating. In contrast, captions in the DataComp subset display a broader distribution, with a substantial fraction of short captions. This indicates that synthetic LAION-BVD captions resemble verbose, COCO-style descriptions, while real DataComp captions reflect the more diverse and often concise nature of web text. ImageNet-1k class-name coverage. We further count the occurrences of ImageNet-1k class names (based on the corresponding synsets) in the two subsets. We identify approximately 21M class-name mentions in the LAION-BVD captions, compared to 141M in the DataComp captions. Figure 7 and Figure 8 show the top-50 most frequent ImageNet-1k class names for the respective datasets. This substantial difference in class-name coverage provides a plausible explanation for the weaker ImageNet-1k zero-shot classification performance of models trained on LAION-BVD subsets (see Table 27). Real captions from DataComp appear to be more strongly aligned with the semantic struc- ture of ImageNet-1k, containing a far greater density and diversity of class-associated terminology. Taken together, these findings suggest that caption length or fluency alone is insufficient for improving downstream transfer. Instead, explicit alignment with class-centric semantics, by, for instance, increasing the inclusion of ImageNet-1k class names or concept-oriented vocabulary during synthetic caption generation may be a more effective strategy for boosting zero-shot classification performance on ImageNet-like benchmarks. Disclaimer for use of LLMs Large Language Models (LLMs) were used for language polishing and proofreading to improve clarity and readability, as well as for caption generation (see Section 3.2). In addition, LLM-based coding agents were used to support coding tasks. All ideas, experimental design, and scientific content originated from the authors. 31 Table 25: Validation of caption quality. We report the performance of CLIP-ViT-B-32 trained on subsets of DataComp-1B [37] with their original captions, our automatically generated captions (Recap), and length-constrained captions with only 7 words on average (Short). Zero-Shot Classification Acc@1Img-Retrieval Recall@5 ImageNet-1kImageNet-RImageNet-SketchCOCO Captions SamplesOriginalRecapShortOriginalRecapOriginalRecapOriginalRecapShort 12.8 M0.100.090.100.120.150.040.060.100.210.19 30 M0.200.150.170.210.240.100.130.180.310.28 128 M0.400.23–0.410.360.250.210.330.42– Table 26: Number of extracted frames per video using our filtering pipeline and corresponding percentage of videos. Extracted framesPercentage of videos 112.93% 27.96% 36.31% 45.36% 54.74% 64.21% 73.79% 83.44% 9+51.27% Table 27: Image-text retrieval performance (R@5) on Flickr30k and MS-COCO for models trained on different 300M-30M subsets of LAION-BVD, DataComp, and Re-LAION. Scaling laws for LAION-BVD persist on different models. ModelDatasetSamplesFlickr30k I→T R@5MS-COCO I→T R@5 ViT-B-16-text-plusLAION-BVD300M0.870.63 ViT-B-16-text-plusLAION-BVD128M0.820.56 ViT-B-32LAION-BVD300M0.820.56 ViT-B-16-text-plusRe-LAION300M0.800.53 ViT-B-16-text-plusLAION-BVD64M0.790.50 ViT-B-16-text-plusDataComp300M0.750.52 ViT-B-32LAION-BVD128M0.750.48 ViT-B-32Re-LAION300M0.720.46 ViT-B-16-text-plusRe-LAION128M0.710.45 ViT-B-16-text-plusLAION-BVD30M0.670.40 ViT-B-32LAION-BVD64M0.670.40 ViT-B-32DataComp300M0.660.45 ViT-B-16-text-plusDataComp128M0.650.43 ViT-B-32Re-LAION128M0.610.38 ViT-B-32DataComp128M0.550.37 ViT-B-16-text-plusDataComp64M0.550.36 ViT-B-32LAION-BVD30M0.540.31 ViT-B-32Re-LAION64M0.510.30 ViT-B-32DataComp64M0.440.29 ViT-B-16-text-plusDataComp30M0.380.26 ViT-B-32Re-LAION30M0.370.22 ViT-B-32DataComp30M0.320.21 32 1020304050 Tokens 0 50000 100000 150000 200000 250000 300000 350000 Distribution of Text Lengths for BVD-I-300M 01020304050607080 Tokens 0 10000 20000 30000 40000 50000 60000 Distribution of Text Lengths for DataComp Figure 6: Caption length distribution for BVD-I-300M synthetically captioned subset (left) and the 300M DataComp subset (right). LAION-BVD captions are longer with a narrower distribution, while DataComp captions show a broader distribution dominated by short captions. T-shirt necklace pillow chain desk website laptop computer patio restaurant sunglasses backpack valley wallet sweatshirt jeep candle microphone refrigerator music speaker ram (adult male sheep) vase pizza storage chest vacuum cleaner taxicab bathtub tray tiger metal nail notebook computer barn ski cougar bucket banana lion palace tractor perfume microwave oven washing machine quilt parallel bars balloon barrel buckle projector leopard dishwasher gown 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1e7 Top 50 IN1k class names in DataComp (Total Count: 141,495,067) Figure 7: Top-50 most frequent ImageNet-1k class names appearing in the 300M DataComp captions. Real captions from DataComp strongly align with IN1K semantic categories, exhibiting 141M total class-name occurrences. microphone desk sunglasses television website T-shirt laptop computer apron necklace backpack music speaker restaurant cornet lipstick metal nail rifle parallel bars tractor violin cliff ski sweatshirt bikini Cardigan Welsh Corgi dough notebook computer altar chain leopard stove balance beam tray storage chest scoreboard saxophone pole lion projector refrigerator candle gown bucket pizza pig vacuum cleaner oscilloscope Boxer accordion vase balloon 0 1 2 3 4 1e6 Top 50 IN1k class names in BVD-I-300M (Total Count: 21,003,027) Figure 8: Top-50 most frequent ImageNet-1k class names appearing in the 300M LAION-BVD synthetic captions. LAION-BVD captions contain 21M total class-name occurrences, substantially fewer than DataComp, consistent with the observed differences in IN1K zero-shot performance. 33