Source-linked AI summary
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
Matthew Wallingford, Anand Bhattad, Aditya Kusupati, Vivek Ramanujan, Matt Deitke, Sham Kakade, Aniruddha Kembhavi, Roozbeh Mottaghi, Wei-Chiu Ma, Ali Farhadi
TL;DR
Real-world 3D learning lacks large-scale data with diverse corresponding views, and standard videos restrict camera access through fixed viewpoints. The paper introduces 360-1M and a correspondence-processing pipeline, then trains ODIN to generate geometrically consistent novel views and reconstruct scenes. ODIN outperforms existing methods on novel-view synthesis and 3D reconstruction benchmarks, while dynamic elements remain a limitation.
Problem
Real-world 3D generative modeling lacks large-scale data with diverse corresponding views, while standard videos use fixed viewpoints that hinder scalable multi-view construction.
Method
The paper collects one million 360° videos, efficiently extracts corresponding multi-view frames, and trains a diffusion-based novel-view synthesis model named ODIN.
Results
ODIN improves existing methods on standard novel-view synthesis and 3D reconstruction benchmarks without fine-tuning to target data.
Takeaways & Limitations
360-1M enables ODIN to generate 3D-consistent novel views of real-world scenes with free camera movement beyond previous methods.
Takeaways & Limitations
Motion masking filters dynamic scene elements rather than modeling them, leaving generalized 4D modeling largely unexplored.
Abstract
from arXiv · showhide
Three-dimensional (3D) understanding of objects and scenes play a key role in humans' ability to interact with the world and has been an active area of research in computer vision, graphics, and robotics. Large scale synthetic and object-centric 3D datasets have shown to be effective in training models that have 3D understanding of objects. However, applying a similar approach to real-world objects and scenes is difficult due to a lack of large-scale data. Videos are a potential source for real-world 3D data, but finding diverse yet corresponding views of the same content has shown to be difficult at scale. Furthermore, standard videos come with fixed viewpoints, determined at the time of capture. This restricts the ability to access scenes from a variety of more diverse and potentially useful perspectives. We argue that large scale 360 videos can address these limitations to provide: scalable corresponding frames from diverse views. In this paper, we introduce 360-1M, a 360 video dataset, and a process for efficiently finding corresponding frames from diverse viewpoints at scale. We train our diffusion-based model, Odin, on 360-1M. Empowered by the largest real-world, multi-view dataset to date, Odin is able to freely generate novel views of real-world scenes. Unlike previous methods, Odin can move the camera through the environment, enabling the model to infer the geometry and layout of the scene. Additionally, we show improved performance on standard novel view synthesis and 3D reconstruction benchmarks.
1 Introduction
Real-world 3D generative modeling remains difficult because large-scale scene data and diverse corresponding views are scarce. The paper addresses this with 360-1M and ODIN, enabling novel-view synthesis and geometry reconstruction from a single image.
- Real-world 3D generative models remain limited by the lack of large-scale datasets for diverse objects and scenes.
- Standard videos contain 3D information, but fixed capture viewpoints and sparse correspondences make scalable multi-view construction difficult.
- 360-1M collects one million 360° YouTube videos and provides an efficient process for transforming them into multi-view training data.
- ODIN is trained on 360-1M to synthesize novel views of real-world scenes and reconstruct their geometry from a single image.
- ODIN improves performance on standard novel-view synthesis and 3D reconstruction benchmarks without fine-tuning on target data.
2 Related Work
Prior multi-view and generative methods provide useful 3D data or novel-view synthesis, but are often constrained by scale, diversity, controlled settings, or fixed camera coverage. The paper uses large-scale 360-degree YouTube videos to broaden real-world multi-view data.
- Novel View Synthesis: NeRF-based methods typically rely on densely sampled images and known camera poses, whereas this approach targets widely varying real-world camera views.
- Novel View Synthesis: Generative diffusion methods have expanded novel-view synthesis from objects to scenes, including camera-conditioned approaches for single-image generation.
- Multi-View Datasets: Existing multi-view datasets span real-world sequences, synthetic assets, and task-specific collections but remain constrained by their environments, objects, or capture procedures.
- Multi-View Datasets: The paper identifies scale and real-world data as jointly missing from current multi-view datasets.
- Multi-View Datasets: 360-1M uses large-scale 360-degree YouTube videos to provide more diverse data, substantial camera changes, and broader real-world coverage than controlled datasets.
3 Multi-View Data from 360◦Video
The paper converts 360° videos into scalable multi-view data by searching for overlapping views, propagating correspondences over a graph, and calibrating relative poses to metric scale. This preserves long-range viewpoint variation while limiting computational cost.
- 360° video enables controllable alignment of views across frames, addressing the fixed-viewpoint problem in standard video.
- 3.1 Scalable Correspondence Search: The correspondence search samples one frame per second, compares frames within a 20-frame window, generates four panoramic views, and filters pairs using Dust3R confidence.
- 3.1 Scalable Correspondence Search: The method refines projection pitch and yaw to maximize overlap and discards pairs with relative translation below 0.25 m.
- Current multi-view collection is limited by the cost of finding high-quality correspondences and relative poses, especially across long video sequences.
- 3.2 Correspondence Propagation: Graph-based correspondence propagation maintains a small search window while recovering long-range correspondences through connected frames.
- 3.3 Resolving Scale Ambiguity: Because Dust3R produces dimensionless poses, the method fuses its point maps with depth estimates to recover a scale factor and convert translations into metric poses.
4 Dataset Collection and Statistics
The authors build 360-1M from over one million 360° YouTube videos and derive a large collection of frame correspondences with relative camera poses.
- Dataset Collection: The scalable correspondence-search pipeline enables 360-1M to provide diverse multi-view data beyond the constraints of controlled datasets.The dataset is introduced as the largest 360° video dataset to date and is used to support large-scale multi-view generation.
- Dataset Collection: The dataset is assembled by filtering YouTube metadata for equirectangular 360° videos, yielding 1,076,592 videos for download.The metadata includes duration, view count, format, and subject category.
- Dataset Collection: Duplicate videos are removed with a thumbnail-based deduplication model, although this does not guarantee unique content across all frames.Running deduplication over every frame is described as computationally infeasible.
- Dataset Statistics: 360-1M contains 80,567,325 unique frames from 1,076,592 videos and provides 363,417,730 frame correspondences with relative camera poses.Videos average 74.83 unique frames, and the collection spans 15 subject categories.
5 Method
ODIN is a diffusion-based model that generates viewpoint-conditioned images from a single scene image, using translation and rotation conditioning, trajectory sampling, and motion masking for dynamic videos.
- 5.1 Viewpoint Conditioned Diffusion: The model generates a sequence of images from a single input image and target viewpoints using a latent diffusion architecture with an encoder, denoising U-net, and decoder.The target viewpoint is represented by relative rotation and translation between views.
- 5.1 Viewpoint Conditioned Diffusion: ODIN conditions novel-view generation on both camera rotation R and translation t, enabling freer camera movement than methods restricted to rotating around a scene center.The model targets image sequences along viewpoint trajectories and uses long-range correspondences from the training data.
- 5.2 Motion Masking: Motion masking filters dynamic scene elements from the training loss so the model can focus on static content in in-the-wild videos.A dense soft mask is predicted by an additional U-net output channel and applied through elementwise multiplication.
- 5.2 Motion Masking: The motion-masking objective requires an auxiliary loss because directly optimizing the masked loss can produce a degenerate all-zero mask.The auxiliary term incentivizes the mask to remain non-zero.
- 5.1 Viewpoint Conditioned Diffusion: ODIN samples images along smooth camera trajectories to support multi-view generation despite the flexibility of its image-to-image training setup.This trajectory-based sampling approach is not restricted to simple rotations.
6 Experiments
ODIN improves novel view synthesis on DTU and Mip-NeRF 360 and supports 3D reconstruction from generated viewpoint trajectories, without fine-tuning on target data.
- 6.2 Novel View Synthesis: ODIN improves performance on DTU and Mip-NeRF 360 without fine-tuning on the target task.The benchmarks report standard novel view synthesis metrics, with LPIPS emphasized for Mip-NeRF 360.
- 6.2 Novel View Synthesis: On Mip-NeRF 360, ODIN shows significant gains on real-world scenes and handles views that differ substantially from the input better than competing methods.Zero1-to-3 struggles to generate full real scenes, while ZeroNVS produces more plausible but still weaker views for complex scenes.
- 6.3 3D Reconstruction: ODIN reconstructs 3D scenes by generating images along viewpoint trajectories and applying Dust3R, with results reported on Google Scanned Objects.The comparison includes volumetric IoU for Google Scanned Objects and Chamfer Distance for the 360-1M held-out set.
- 6.3 3D Reconstruction: On Google Scanned Objects, ODIN performs comparably to Zero1-to-3 and outperforms the other compared reconstruction methods.Zero1-to-3 is a particularly relevant comparison because it was designed for synthetic objects.
- 6.3 3D Reconstruction: For held-out 360-1M scene reconstruction, the comparison is limited to ZeroNVS because the other methods cannot generate scenes.This evaluation uses pseudo-ground truth derived from Dust3R applied to ground-truth video views.
7 Limitations and Broader Impact
The authors identify dynamic scene content as a key modeling limitation and note potential benefits and risks from generated 3D assets.
- Limitations: ODIN filters dynamic elements with a motion mask but does not yet model those elements directly.The authors identify generalized 4D models, which vary camera view across both time and space, as largely unexplored.
- Broader Impact: The work could support 3D asset creation for AR, VR, and robotic navigation, but could also generate fake images or inappropriate scenes.These are presented as potential positive and negative societal impacts.
8 Conclusion
The paper introduces scalable real-world multi-view data through 360-1M and trains ODIN to generate 3D-consistent novel views with free camera movement.
- Conclusion: 360-1M provides large-scale, diverse, long-range correspondences for training ODIN on real-world multi-view data.The authors describe it as the largest multi-view dataset to date.
- Conclusion: ODIN generates 3D-consistent novel views of real-world scenes with free camera movement and outperforms existing methods on novel view synthesis and 3D reconstruction without target-data fine-tuning.The authors identify modeling dynamics for 4D scene generation as an important next step.
A Dataset Statistics
The dataset statistics section presents distributions of video duration, subject categories, and language in 360-1M.
- Dataset Statistics: Video duration in 360-1M follows the distribution shown in Figure 5.
- Dataset Statistics: Video categories in 360-1M follow the distribution shown in Figure 6.
- Dataset Statistics: Video languages in 360-1M follow the distribution shown in Figure 7.
C 3D Reconstruction Evaluation
This section reports 3D reconstruction results on 360-1M and compares them with Zero 1-to-3, alongside the model’s training and release details.
- 3D reconstruction on 360-1M is compared with Zero 1-to-3.
- The released data requires applications to obtain video links and metadata.
- ODIN was trained for 2 weeks on 16 A40 GPUs for 100 epochs.
G Ablations
The ablations examine sampling frame rates and the motion-masking coefficient, finding that higher sampling rates improve performance while 1 FPS balances performance and compute cost.
- Frame-rate sampling: Higher sampling FPS improves LPIPS, PSNR, and SSIM, with 1 FPS selected for scaling to one million videos.The choice is described as a balance between performance and compute cost.
- MVImageNet provides correspondence examples but was previously the largest multiview dataset.
- Motion masking: Motion masking is ablated across different λ values using novel view synthesis metrics.