COMPANY

LAION

companytopic-noteopen-sourcedataset

Overview

LAION (Large-scale Artificial Intelligence Open Network) is a German non-profit that ships large open datasets — LAION-5B (5B image-text pairs) is the reference pretraining corpus for open text-to-image work. In August 2026 LAION released BVD, an open-video corpus (80M videos, 10M hours, 55M clips, 300M stills) with auto-generated video + audio captions, distributed under a research-only license via LAION’s projects portal with code on GitHub — the open counterpart to the proprietary video corpora frontier video-generation labs have been assembling privately.

Timeline

  • 2026-08-30-AI-DigestLAION released BVD — 80M videos, 10M hours of footage, 55M individual clips, and 300M associated stills, all with auto-generated video + audio captions, distributed under a research-only license via LAION’s projects portal with code on GitHub (The Decoder). Positioned as the open counterpart to the proprietary corpora frontier video-generation labs have been assembling privately. Narrow read the digest carries: the dataset is real, the numbers are LAION’s own, and the delivery mechanism (portal + GitHub) matches LAION’s prior LAION-5B distribution pattern; auto-generated captions carry the usual quality caveat, but for pre-training scale that has historically been fine. Structural read: do NOT frame this as “video models about to catch up to closed labs” — the gap Sora / Runway / DeepMind Genie-style systems have opened is on compute and post-training, not just data. Correct frame: the open-corpus floor for video just moved up by an order of magnitude, which does most of its work on academic reproducibility (PAWBench-style evaluations, distribution-alignment papers, world-model scaling laws) and on the second-tier vendor tier that could not previously afford proprietary video-training deals. Pair with today’s PAWBench paper — the community now has both an open pre-training corpus and an open distribution-alignment benchmark landing in the same 48 hours. Log against MOC - Open Source Models and MOC - AI Infrastructure.

Key Developments

  1. BVD Open-Video Corpus — 80M Videos / 10M Hours / 55M Clips / 300M Stills (August 30, 2026): LAION drops the largest open video dataset to date — 80M videos, 10M hours of footage, 55M individual clips, and 300M associated stills with auto-generated video + audio captions, research-only license, distributed via LAION’s projects portal with code on GitHub. Load-bearing framing to carry: numbers are LAION’s own; delivery mechanism matches the LAION-5B distribution pattern; auto-generated captions carry the usual quality caveat but pre-training scale has historically been fine on that count. Structural framing: the open-corpus floor for video just moved up by an order of magnitude, but the closed-vs-open gap remains on compute and post-training, not just data — Sora / Runway / DeepMind Genie-style systems have opened the gap on compute and post-training axes that a corpus release cannot close. Correct read: BVD does its work on academic reproducibility (PAWBench-style evaluations, distribution-alignment papers, world-model scaling laws) and on the second-tier vendor tier that could not previously afford proprietary video-training deals. Pair with today’s PAWBench paper (arXiv:2608.27345) — open pre-training corpus and open distribution-alignment benchmark landing in the same 48 hours. 30 / 60 / 90-day watch: whether an academic group reproduces a video-generator baseline trained purely on BVD; whether the auto-generated captions get pressure-tested for quality drop at scale; whether the second-tier vendor tier (below Sora / Runway / Genie) surfaces BVD-trained checkpoints inside 90 days.

See also: Hugging Face, DeepMind, MOC - Open Source Models, MOC - AI Infrastructure.