Source-linked AI summary
Sekai: A Video Dataset towards World Exploration
Zhen Li, Chuanhao Li, Xiaofeng Mao, Shaoheng Lin, Ming Li, Shitian Zhao, Zhaopan Xu, Xinyue Li, Yukang Feng, Jianwen Sun, Zizhen Li, Fanrui Zhang, Jiaxin Ai, Zhixiang Wang, Yuwei Wu, Tong He, Jiangmiao Pang, Yu Qiao, Yunde Jia, Kaipeng Zhang
TL;DR
Existing video-generation datasets lack the geographic breadth, duration, dynamic scenes, and annotations needed for world exploration. Sekai introduces a large egocentric worldwide dataset and curation pipeline, and experiments report improved generation quality and interaction following. The dataset is intended to support video generation and world exploration research.
Problem
Existing video-generation datasets are not well-suited to world exploration because they have limited locations, short duration, static scenes, and insufficient exploration and world annotations.
Method
Sekai combines worldwide walking and drone-view videos with annotations and a pipeline for collecting, preprocessing, filtering, annotating, and sampling data.
Results
Experiments report consistent gains in video-generation quality and substantially improved interaction following, including reduced generated-to-target trajectory error.
Takeaways & Limitations
Sekai provides a large, diverse, richly annotated resource for training video-generation models for world-exploration scenarios.
Abstract
from arXiv · showhide
Video generation techniques have made remarkable progress, promising to be the foundation of interactive world exploration. However, existing video generation datasets are not well-suited for world exploration training as they suffer from some limitations: limited locations, short duration, static scenes, and a lack of annotations about exploration and the world. In this paper, we introduce Sekai (meaning "world" in Japanese), a high-quality first-person view worldwide video dataset with rich annotations for world exploration. It consists of over 5,000 hours of walking or drone view (FPV and UVA) videos from over 100 countries and regions across 750 cities. We develop an efficient and effective toolbox to collect, pre-process and annotate videos with location, scene, weather, crowd density, captions, and camera trajectories. Comprehensive analyses and experiments demonstrate the dataset's scale, diversity, annotation quality, and effectiveness for training video generation models. We believe Sekai will benefit the area of video generation and world exploration, and motivate valuable applications. The project page is https://lixsp11.github.io/sekai-project/.
1 Introduction
Sekai addresses the lack of long, diverse, richly annotated data for world-exploration video generation by introducing a worldwide egocentric dataset and curation pipeline. Experiments validate its scale, annotation quality, and usefulness across video-generation tasks.
- Existing video-generation datasets are limited by locations, duration, scene dynamics, and exploration- or world-related annotations.
- Sekai provides over 5,000 hours of egocentric walking and drone-view videos across 101 countries and more than 750 cities.Sekai-Real contributes YouTube videos, while Sekai-Game contributes videos from a realistic video game.
- The dataset includes diverse weather, times, dynamic scenes, long walking videos, audio, and annotations for location, scene, crowd density, captions, and camera trajectories.Walking videos are at least 60 seconds long and recorded at 720p and 30 FPS.
- A curation pipeline collects, filters, and annotates videos from YouTube and video games for practical model training.The pipeline includes video preprocessing and annotation of location, scene type, weather, crowd density, captions, and camera trajectories.
- Experiments show consistent gains in video-generation quality and substantially improved interaction following using Sekai’s data and camera trajectories.Trajectory-aware training significantly reduces the error between generated and target camera trajectories.
2 Related Work
Prior work spans general video, 3D, and 4D generation, while existing video datasets are organized around specific or open scenarios. Many existing datasets remain limited in scale, duration, or scenario coverage.
- Recent research advances text-to-video, image-to-video, 3D, and 4D generation as foundations for world-generation models.
- Existing video-generation datasets include specific-scenario and open-scenario collections with differing coverage and intended uses.
- Typical specific-scenario datasets contain less than 800 total hours and have limited individual video duration.
3 Dataset Curation
Sekai’s curation process combines collection, preprocessing, annotation, and sampling to produce training-ready real and game-derived videos. It uses quality, diversity, and trajectory-aware procedures to select and enrich the dataset.
- The curation pipeline has four stages: video collection, preprocessing, annotation, and sampling.
- Video Collection: The collection stage gathers YouTube walking and drone videos plus realistic game videos with accessible ground-truth annotations.The game source provides location, weather, and camera-trajectory ground truth and is inexpensive to scale.
- Pre-processing: Preprocessing trims source videos, detects shot boundaries, extracts one-minute clips, standardizes encoding, and removes unsuitable footage.Filtering addresses brightness, technical quality, subtitles, and implausible camera trajectories.
- Video Annotation: The annotation process labels geographic location, scene, weather, time of day, crowd density, captions, and per-frame camera trajectories.YouTube metadata and video content support annotations, while multiple trajectory methods are evaluated for real videos.
- Video Annotation: Sekai-Game annotations are captured through an Unreal Engine toolchain that records camera poses and aligns them with video frames.Captured poses are calibrated for delay compensation and interpolated for synchronization.
- Sampling: Sekai-Real-HQ samples 400 hours of clips according to aesthetic quality, semantic quality, location diversity, and other diversity criteria.Location diversity explicitly considers the distribution of clips across cities.
4 Dataset Statistics
Sekai-Real spans 101 countries and multiple environmental dimensions, while Sekai-Real-HQ provides a more balanced, higher-quality subset for training.
- 101 countries and regions are represented in Sekai-Real, with the top eight countries contributing about 60% of total video duration.
- Sekai-Real covers four weather types, four scene types, four time-of-day categories, and five crowd-density levels.
- Sekai-Real-HQ has a higher mean video quality and lower variance than Sekai-Real, addressing its long-tail quality distribution.
- Sekai-Real-HQ has a more balanced location distribution, which helps mitigate potential bias during model training.
- Sekai-Real provides richer textual supervision than OpenVid-1M through a higher average caption token count.
5 Experiments
The experiments assess annotation quality and test Sekai for text-to-video, image-to-video, and camera-trajectory-guided generation. The dataset yields reliable annotations and improves both generation quality and interactive control.
- 5.1 Evaluation of Annotation Quality: Location annotation issues occurred in fewer than 5% of 500 randomly sampled videos, indicating high overall location quality.
- 5.1 Evaluation of Annotation Quality: MegaSAM produces smoother camera trajectories, whereas VGGT provides faster inference but lower annotation quality.
- 5.1 Evaluation of Annotation Quality: Overall agreement between Qwen2.5-VL and human category annotations exceeds 90%, with most weather discrepancies involving cloudy, foggy, and rainy labels.
- 5.2 Quantitative Results: Fine-tuning on Sekai-Real-HQ consistently improves text-to-video and image-to-video quality across most metrics and raises the Overall Score.
- 5.2 Quantitative Results: Dynamic Degree decreases after early peaks as training shifts from exaggerated motion toward sharper details, cleaner frames, and more moderate motion.
- 5.3 Quantitative Results: Fine-tuning on drone-view Sekai-Real reduces TransErr by △11.13 and RotErr by △7.31 while improving other interactive-generation metrics.
- 5.3 Quantitative Results: Fine-tuning on Sekai-Game reduces camera-control error rates by more than 30% on average.
6 Conclusion
The paper introduces Sekai as a large-scale, annotated video dataset for video-generation-based world exploration. Its analyses and experiments support its scale, diversity, annotation quality, and usefulness for model training.
- Sekai contains over 5,000 hours of walking or drone-view videos from 101 countries and more than 750 cities.
- The dataset pipeline processes, filters, annotates, and samples videos with location, scene, weather, crowd density, captions, and camera trajectories.
- Comprehensive analyses and experiments validate Sekai’s scale, diversity, annotation quality, and effectiveness for training world-exploration video-generation models.