【文章标题】:Laion Big Video Dataset 【文章标题】:Laion大型视频数据集

【文章正文】: Overview 【文章正文】: 概述

We present LAION-BVD (LAION — Big Video Dataset), a large-scale open video dataset for multimodal learning, containing 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale. 我们推出了LAION-BVD(LAION——大型视频数据集),这是一个用于多模态学习的大规模开放视频数据集,包含从CommonCrawl收集的13亿个特定平台视频URL。从中,我们下载了8000万个视频,总时长达1000万小时。该数据集专为跨视频、音频和图像模态的多模态预训练而设计。利用内容感知的场景检测技术,我们提取了视频片段,并为其合成了视频和音频描述。在这些数据上训练的模型在标准的视频-文本和音频-文本基准测试中取得了具有竞争力的表现,并且随着训练规模或模型规模的增加,性能持续提升。此外,我们通过提取场景切换帧,探索将视频帧作为图像-文本数据的替代来源。这些帧展现出与标准网络图像语料库截然不同的视觉分布,在该数据集上训练的模型在图像-文本检索任务中表现优异。我们向研究社区开源了LAION-BVD。它以前所未有的规模显著扩大了多模态视频的开放获取范围。

By the Numbers 数据概览

The largest openly accessible video corpus for multimodal learning research 多模态学习研究中最大规模的公开可访问视频语料库

Platform-specific URLs collected from Common Crawl 从Common Crawl收集的特定平台URL

Successfully downloaded and processed videos 成功下载并处理的视频数量

Combined video content across all downloads 所有下载内容的视频总时长

Clips with generated video captions 生成视频描述的片段数量

Video frames for image-text pre-training 用于图像-文本预训练的视频帧

Benchmarks 基准测试

Models trained on LAION-BVD achieve competitive performance across video, audio, and image-text benchmarks 在LAION-BVD上训练的模型在视频、音频和图像-文本基准测试中均取得了具有竞争力的表现

ViCLIP models trained on LAION-BVD match or exceed InternVid-trained models by up to 2.1% on standard video-text benchmarks, with consistent improvements as training scale grows from 10M to 50M clips. 在LAION-BVD上训练的ViCLIP模型在标准视频-文本基准测试中,表现与InternVid训练的模型持平或高出多达2.1%,且随着训练规模从1000万片段增长到5000万片段,性能持续提升。

CLAP models trained on LAION-BVD achieve competitive performance against other large-scale uncurated audio datasets, leveraging rich in-the-wild soundscapes extracted directly from video. 在LAION-BVD上训练的CLAP模型利用直接从视频中提取的丰富真实环境声景,在与其他大规模未经筛选音频数据集的对比中取得了具有竞争力的表现。

Frame-based CLIP models achieve strong image-text retrieval performance on standard benchmarks. Video frames exhibit a visual distribution distinct from typical web corpora, complementing existing image pre-training sources. 基于帧的CLIP模型在标准基准测试中实现了强大的图像-文本检索性能。视频帧展现出与典型网络语料库截然不同的视觉分布,是对现有图像预训练数据源的有效补充。

Responsible Use 负责任的使用

LAION-BVD is released to support open and reproducible multimodal research at scale. Large-scale video datasets and the models trained on them are increasingly concentrated within a small number of predominantly proprietary technology companies, limiting independent scientific investigation and reproducibility. By providing an open resource for academic research, we aim to broaden access to multimodal training data and enable more transparent evaluation of large-scale video, audio, and image models. LAION-BVD的发布旨在支持大规模、开放且可复现的多模态研究。大规模视频数据集及其训练模型正日益集中在少数主要为专有技术的科技公司手中,这限制了独立的科学研究和可复现性。通过为学术研究提供开放资源,我们旨在拓宽多模态训练数据的获取渠道,并实现对大规模视频、音频和图像模型更透明的评估。

LAION-BVD is released exclusively for research purposes and not for commercial use. The dataset is intended to support scientific research, reproducibility, safety analysis, and the study of multimodal foundation models and related systems. We encourage users to respect the rights and copyright of content creators and to use the dataset responsibly and in accordance with applicable laws and platform terms. LAION-BVD仅用于研究目的发布,严禁用于商业用途。该数据集旨在支持科学研究、可复现性验证、安全分析以及多模态基础模型和相关系统的研究。我们鼓励用户尊重内容创作者的权利和版权,负责任地使用本数据集,并遵守适用的法律法规和平台条款。

Like other large-scale web datasets, LAION-BVD may contain biases, stereotypes, and uneven representation across languages, regions, and topics. Models trained on this data may inherit such biases. Researchers using the dataset should be aware of these limitations and, where relevant, evaluate and report them alongside model capabilities. 与其他大规模网络数据集一样,LAION-BVD可能包含不同语言、地区和主题之间的偏见、刻板印象以及代表性不均的问题。在此数据上训练的模型可能会继承这些偏见。使用该数据集的研究人员应意识到这些局限性,并在相关情况下,在评估模型能力的同时对这些局限性进行评估和报告。

Reference 参考文献

@misc{laionbvd2026, title={LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training}, author={Andreas Hochlehnert and Marianna Nezhurina and Mehdi Cherti and Andrej Radonjic and Thaddäus Wiedemer and Christoph Schuhmann and Romain Beaumont and Wieland Brendel and Bernhard Schölkopf and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge}, year={2026}, eprint={2608.24845}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2608.24845}, } @misc{laionbvd2026, title={LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training}, author={Andreas Hochlehnert and Marianna Nezhurina and Mehdi Cherti and Andrej Radonjic and Thaddäus Wiedemer and Christoph Schuhmann and Romain Beaumont and Wieland Brendel and Bernhard Schölkopf and A. Sophia Koepke and Jenia Jitsev and Matthias Bethge}, year={2026}, eprint={2608.24845}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2608.24845}, }