Source-linked AI summary

A Cookbook of 3D Vision: Data, Learning Paradigms, and Application

Hongyang Du, Zongxia Li, Dawei Liu, Runhao Li, Haoyuan Song, Qingyu Zhang, Yubo Wang, Jingcheng Ni, Shihang Gui, Congchao Dong, Tao Hu

arXiv:2606.04291v1cs.CV

TL;DR

3D vision lacks a unified, data-centric framework connecting its diverse representations, datasets, and modeling paradigms. This paper develops such a taxonomy and synthesizes how their trade-offs and relationships shape reconstruction, generation, and video modeling.

  • Problem

    Existing 3D vision reviews are representation-centric or task-specific rather than unifying data structures, benchmark datasets, and modeling paradigms.

  • Method

    The paper presents a data-centric taxonomy analyzing 3D representations, acquisition pipelines, datasets, benchmarks, supervision regimes, learning paradigms, and applications.

  • Results

    The synthesis clarifies how representation trade-offs, benchmark design, and supervision regimes relate to efficiency, fidelity, scalability, and downstream 3D tasks.

  • Takeaways & Limitations

    A consolidated view links evolving representations, datasets, and learning paradigms while highlighting multimodal geometric grounding and emerging 4D world-modeling directions.

  • Takeaways & Limitations

    Current benchmarks lack large-scale multimodal coverage combining heterogeneous representations, temporal consistency, and open-world generalization within unified protocols.

Abstract

from arXiv · show

3D vision has rapidly evolved, driven by increasingly diverse data representations, learning paradigms, and modeling strategies. Yet the field remains fragmented across representations and benchmarks, making it difficult to develop unified perspectives on efficiency, fidelity, and scalability. This work provides a data-centric taxonomy of 3D vision that connects geometric representations, datasets, learning frameworks, and applications within a single conceptual map. We begin by analysing the principal structural representations of 3D data--point clouds, meshes, voxels, and 3D Gaussians--along with their acquisition pipelines. We then examine how dataset design, benchmark construction, and supervision regimes shape recent advances, spanning 2D-supervised 3D learning, implicit neural representations, and 4D world modeling. Through this integrative lens, we clarify the relationships among representations, learning paradigms, and downstream tasks in reconstruction, generation, and video modeling, offering a consolidated view of emerging trends toward balancing efficiency and fidelity and toward multimodal geometric grounding.

1 Introduction

3D vision is increasingly practical and widespread but remains complex and fragmented across representations, learning pipelines, datasets, and tasks. This section presents a unified, data-centric framework connecting these elements and situating emerging paradigms around efficiency, fidelity, and accessibility.

  • Motivation: 3D vision spans diverse representations, structural assumptions, learning pipelines, computational trade-offs, and downstream tasks, making it fundamentally more complex than 2D vision.Covered formats include point clouds, meshes, voxel grids, RGB-D images, multi-view images, CAD models, neural implicit fields, and 3D Gaussians.
  • Motivation: Existing reviews are typically architecture-centric, representation-centric, or task-specific rather than unifying data structures, benchmark datasets, and modeling paradigms.The paper identifies the absence of a unified, data-centric view connecting these components in one framework.
  • Contributions: The cookbook maps how major 3D data formats are represented, stored, and processed across computers and machine learning systems.The unified map covers point clouds, meshes, voxel grids, RGB-D images, CAD models, implicit fields, and 3D Gaussians.
  • Contributions: It explains how datasets and benchmarks shape 3D learning through data structures, supervision formats, evaluation, and scalability constraints, while contextualizing 2D-supervised learning, implicit fields, and 4D world modeling.These trends are framed within broader concerns about efficiency, fidelity, and accessibility.

2 Scope of the Paper

This survey covers 3D vision through three connected axes: data representations, datasets and benchmarks, and modeling paradigms. It distinguishes itself by linking these dimensions across representations, supervision strategies, and cross-task scalability rather than focusing on isolated architectures, paradigms, or applications.

  • Data Representations: The survey analyzes major 3D data forms, including point clouds, meshes, voxel grids, images, CAD/B-Rep models, neural implicit fields, and 3D Gaussian.It examines their efficiency–fidelity trade-offs.
  • Datasets and Benchmarks: It explores datasets and benchmarks across modalities and tasks, emphasizing how benchmark design enables progress while constraining model development.
  • Modeling Paradigms: The review summarizes classical geometry-based pipelines and neural approaches spanning 2D-supervised 3D learning, implicit neural fields, and 4D video/world modeling.
  • Positioning: Unlike prior reviews, it connects datasets and representations while considering supervision strategies and scalability across tasks.Architecture-centric, topic-centric, and task-oriented reviews typically focus on network families, individual paradigms, or applications in isolation.

3 A Taxonomy of 3D Representations

3D vision uses diverse representations—including RGB-D, point clouds, voxels, meshes, CAD models, and 3D Gaussians—whose structures, acquisition pipelines, and processing efficiencies differ. These representations support tasks ranging from reconstruction and recognition to rendering, simulation, segmentation, and neural rendering.

  • RGB-D: RGB-D combines per-pixel color and depth into a structured 2.5D format that supports 3D recovery and efficient 2D CNN processing at O(H × W).RGB-D is commonly captured with Kinect, Intel RealSense, or Structure Sensor devices and is widely used for indoor scene understanding, pose estimation, and SLAM.
  • Point Clouds: Point clouds represent 3D space as discrete points acquired by LiDAR, RGB-D, or photogrammetry, with computational cost ranging from O(N) to O(N 2) depending on the architecture.PointNet and PointMamba operate in O(N), whereas PointTransformer scales as O(N 2); acquisition may be direct or reconstructed from image collections using SfM and MVS.
  • Voxel Grids: Voxel grids discretize 3D space into uniform N × N × N cells storing occupancy, color, density, or semantic attributes, making them compatible with 3D CNNs.Voxel data typically come from voxelized meshes, CAD surfaces, dense point clouds, or fused multi-view depth observations rather than direct sensing.
  • Meshes: Meshes explicitly encode vertices, edges, and faces, preserving shape and topology for rendering, CAD design, and physical simulation but challenging standard grid-based learning.Pipelines often convert meshes to point clouds or voxels, while direct mesh networks such as MeshCNN remain more specialized.
  • 3D Gaussians: 3D Gaussians provide compact continuous primitives for neural rendering, encoding position, covariance, opacity, and spherical-harmonics coefficients for geometry and view-dependent color.They are typically initialized from SfM point clouds and optimized by gradient descent to minimize rendering loss.

4 3D Learning Paradigms and Applications

3D learning has shifted from explicit geometry supervision toward differentiable rendering, image-aligned representations, generative priors, and structured 3D latents. These paradigms support reconstruction, generation, consistent video, dynamic world modeling, and spatially grounded embodied control.

  • Core 3D Learning and Rendering Paradigms: Differentiable rendering replaced costly direct 3D losses by enabling image-space supervision, while 3D Gaussian Splatting reduced rendering from seconds to milliseconds.Early methods explicitly optimized Chamfer distance, Earth Mover’s Distance, or volumetric TSDF errors in 3D space; optimized α-blending rasterization enabled large feed-forward foundation models.
  • Core 3D Learning and Rendering Paradigms: Image-aligned representations preserve dense per-pixel structure while keeping learning in the 2D domain, supporting direct prediction of geometry, depth, and camera information.DUSt3R, VGGT, RayZer, π3, and Depth Anything 3 instantiate confidence-weighted, multi-task, self-supervised, permutation-equivariant, and unified depth-plus-ray formulations.
  • Core 3D Learning and Rendering Paradigms: When explicit 3D data is scarce, methods distill priors from large-scale 2D models or use structured latents, with TRELLIS decoding into radiance fields, Gaussians, or meshes.DreamFusion and Magic3D optimize neural fields through Score Distillation Sampling, while native 3D geometric foundation models learn structured representations.
  • Core 3D Learning and Rendering Paradigms: Reconstruction and generation increasingly reinforce each other: generative priors hallucinate missing geometry in sparse-view reconstruction, while reconstructed scaffolds constrain physically consistent generation.This coupling connects previously separate domains through methods including RCFM and diffusion-based approaches.
  • Downstream Applications: Applications extend from end-to-end reconstruction and feed-forward asset generation to geometry-regulated video, temporally persistent world models, and 3D-grounded vision-language-action systems.These systems respectively recover geometry from imagery, map multi-view diffusion outputs into 3D assets, regulate temporal consistency, model dynamic states, and ground robotic control in shared 3D representations.

5 Dataset and Benchmark

The section presents a four-axis taxonomy of 3D datasets and benchmarks spanning modality, spatial granularity, task formulation, and temporal dimension. It also highlights rapid, burst-like benchmark growth, increasingly model-aware construction, and persistent gaps in multimodal, temporally consistent, open-world coverage.

  • Dataset taxonomy: Datasets are categorized by data modality, spatial granularity, task formulation, and temporal dimension.The modalities include RGB-D, point clouds, meshes, multi-view images, implicit fields, and Gaussians; tasks include segmentation, correspondence, reconstruction, and generation.
  • Dataset taxonomy: Recent benchmarks encode assumptions of modern 3D pipelines, including image-aligned reconstruction and 3DGS-native learning.Benchmark design therefore extends beyond data collection to reflect representation- and pipeline-specific learning assumptions.
  • Benchmark growth: Dataset releases have surged over the past decade, with post-2020 growth occurring in bursts linked to new sensing pipelines and model families.The section identifies high-fidelity real capture in curated settings as one of three recent scaling directions, though the supplied passage truncates the examples.
  • Remaining gaps: Current benchmarks lack large-scale multimodal coverage supporting heterogeneous representations, temporal consistency, and open-world generalization.The section contrasts ScanNet++ and DL3DV-10K for geometry and view diversity with WildRGB-D for real-world object capture and PointOdyssey as a synthetic dataset.

6 Conclusion

This conclusion presents a data-centric framework unifying 3D representations, datasets, and learning paradigms while identifying persistent scalability, comparison, and generalization challenges. It highlights future directions in unified evaluation, cross-modal geometric grounding, and scalable representations balancing efficiency with fidelity.

  • Contributions: The paper unifies 3D representations, datasets, and learning paradigms into a coherent data-centric framework.It clarifies how efficiency, fidelity, and scalability shape representation design.
  • Challenges: Fragmented datasets hinder fair comparison, voxel- and mesh-based methods face scalability problems, and generalization beyond curated domains remains limited.The conclusion also identifies 4D spatiotemporal reasoning, physics-aware modeling, and world-consistent video generation as emerging integration challenges.
  • Future directions: Future work should develop unified benchmarks spanning objects, scenes, and dynamics, alongside cross-modal and 2D-supervised learning that preserves geometric grounding.These strategies exploit large-scale image data while retaining geometric structure.
  • Future directions: Scalable real-time representations, from Gaussian splats to parametric CAD, should balance efficiency with fidelity.The conclusion presents these representations as a promising direction.
Loading 2606.04291v1…