Source-linked AI summary

OmniObject3D: Large-Vocabulary 3D Object Dataset for Realistic Perception, Reconstruction and Generation

Tong Wu, Jiarui Zhang, Xiao Fu, Yuxin Wang, Jiawei Ren, Liang Pan, Wayne Wu, Lei Yang, Jiaqi Wang, Chen Qian, Dahua Lin, Ziwei Liu

arXiv:2301.07525v2cs.CV

TL;DR

OmniObject3D targets the lack of large-scale real-scanned 3D object data for realistic 3D vision. It builds a 6,000-object, 190-category dataset with multimodal annotations and evaluates it across four tracks, while exposing semantic-distribution bias and varying generation difficulty.

  • Problem

    The lack of large-scale real-scanned 3D databases leaves many 3D methods reliant on synthetic data despite synthetic–real appearance and distribution gaps.

  • Method

    The paper constructs OmniObject3D from professionally scanned objects with textured meshes, point clouds, rendered multiview images, and annotated real videos, then defines four evaluation tracks.

  • Results

    The four benchmarks reveal observations, challenges, and opportunities across robust perception, novel-view synthesis, neural surface reconstruction, and 3D object generation.

  • Takeaways & Limitations

    OmniObject3D provides a database for examining realistic 3D vision, including robustness, generalizable modeling, reconstruction, and large-vocabulary object generation.

  • Takeaways & Limitations

    Large-vocabulary realistic generation remains challenging, with semantic distribution bias and varying exploration difficulties across object groups.

Abstract

from arXiv · show

Recent advances in modeling 3D objects mostly rely on synthetic datasets due to the lack of large-scale realscanned 3D databases. To facilitate the development of 3D perception, reconstruction, and generation in the real world, we propose OmniObject3D, a large vocabulary 3D object dataset with massive high-quality real-scanned 3D objects. OmniObject3D has several appealing properties: 1) Large Vocabulary: It comprises 6,000 scanned objects in 190 daily categories, sharing common classes with popular 2D datasets (e.g., ImageNet and LVIS), benefiting the pursuit of generalizable 3D representations. 2) Rich Annotations: Each 3D object is captured with both 2D and 3D sensors, providing textured meshes, point clouds, multiview rendered images, and multiple real-captured videos. 3) Realistic Scans: The professional scanners support highquality object scans with precise shapes and realistic appearances. With the vast exploration space offered by OmniObject3D, we carefully set up four evaluation tracks: a) robust 3D perception, b) novel-view synthesis, c) neural surface reconstruction, and d) 3D object generation. Extensive studies are performed on these four benchmarks, revealing new observations, challenges, and opportunities for future research in realistic 3D vision.

1. Introduction

OmniObject3D addresses the shortage of large-scale, high-quality real-world 3D object data with a broad, richly annotated collection of realistic scans. It also establishes four evaluation tracks spanning perception, synthesis, reconstruction, and generation.

  • Synthetic-data reliance and the synthetic–real appearance and distribution gaps motivate a large-scale, high-quality real-world 3D object dataset.
  • OmniObject3D contains 6,000 textured meshes across 190 daily categories, with shared classes from ImageNet, LVIS, and ShapeNet.The dataset is described as the largest real-world collection with accurate 3D meshes.
  • Each object includes textured meshes, sampled point clouds, posed Blender-rendered multiview images, and real video frames with masks and COLMAP poses.
  • Professional scanners provide high-fidelity shapes, geometric details, and realistic high-frequency textures.
  • The benchmark suite covers robust 3D perception, novel-view synthesis, neural surface reconstruction, and 3D object generation.The authors use these tracks to expose observations, challenges, and future opportunities in realistic 3D vision.

2. Related Works

Existing 3D object resources trade off scale, realism, completeness, or semantic breadth, leaving a need for a broad real-world dataset. Related work spans synthetic models, multiview reconstruction, point clouds, and textured-mesh generation.

  • Synthetic datasets offer scale but retain an inevitable gap from real objects, while realistic multiview datasets are small or lack category annotations.
  • Figure 2 depicts OmniObject3D’s 190-category long-tailed semantic distribution and its overlap with popular 2D and 3D datasets.
  • CO3D provides 19,000 object-centric videos, but only 20% have accurate COLMAP point clouds and it lacks meshes and textures.
  • 3D generation research includes voxel, point-cloud, octree, implicit, and textured-mesh formulations, but complex textured-surface generation remains challenging.The paper evaluates GET3D on OmniObject3D to reveal challenges and future opportunities.

3. The OmniObject3D Dataset

OmniObject3D is constructed through category-driven collection, professional scanning, and multimodal video annotation. Its 6,000 models span 190 long-tailed categories and support configurable generated data.

  • The category list is initialized from popular 2D and 3D datasets and dynamically expanded, producing 190 widely distributed categories.The design targets diverse texture, geometry, and semantic information.
  • Professional 3D scanners collect high-resolution textured meshes from varied objects in each category.
  • The dataset also provides a generation pipeline with user-defined camera distributions, lighting, and point-sampling methods.
  • Videos capture each object through a full 360° range around a calibration board, with blurry frames filtered using recognized calibration corners.
  • COLMAP estimates camera poses, calibration-board scales recover absolute SfM scale, and U2Net plus FBA produce foreground masks.
  • The collection contains 6,000 models across 190 categories, averaging around 30 objects per category, with a long-tailed distribution.

4. Experiments

Experiments use OmniObject3D to evaluate robustness, novel-view synthesis, surface reconstruction, and 3D generation across realistic objects and challenging settings. Results expose architecture-specific robustness patterns, generalization benefits, reconstruction artifacts, and generation trade-offs.

  • Robust 3D Perception: OmniObject3D evaluates point-cloud classifiers separately on OOD styles and OOD corruptions using clean real scans and OmniObject3D-C.Models trained on ModelNet are tested on OmniObject3D for style robustness and on corrupted OmniObject3D-C for corruption robustness.
  • Robust 3D Perception: Clean accuracy has little correlation with OOD-style robustness, while advanced point grouping improves robustness to both OOD styles and corruptions.CurveNet and GDANet are identified as robust to both challenges, whereas combined OOD style and corruption remains difficult.
  • Novel View Synthesis: Plenoxels achieve the best average PSNR, SSIM, and LPIPS in single-scene novel-view synthesis but are less stable and produce artifacts on concave geometry.The comparison uses dense captured images and reports standard deviation across training samples.
  • Novel View Synthesis: In cross-scene novel-view synthesis, models trained across categories can match or outperform category-specific training, indicating useful priors for unseen scenes.MVSNeRFAll* is comparable to MVSNeRFCat., while IBRNetAll* and pixelNeRFAll* outperform their corresponding category-specific models on visual metrics.
  • Neural Surface Reconstruction: Sparse-view reconstruction shows artifacts across methods; SparseNeuS performs best on average, while NeuS, MonoSDF, pixelNeRF, and MVSNeRF exhibit different geometry strengths and failures.NeuS can recover coherent global shapes for thin structures, MonoSDF can reduce ambiguity with geometry cues, and generalized NeRF surfaces are relatively low quality.
  • 3D Object Generation: Generation experiments find realistic textures and coherent shapes, while data splits reveal semantic distribution bias, varying difficulty, and a quality–diversity trade-off.Furniture is lowest quality, toys achieve the best quality, and Rand-100 is the most difficult case.

5. Conclusion and Outlook

OmniObject3D is a large-vocabulary dataset of 6,000 real-scanned objects across 190 categories, with rich multimodal annotations and four evaluation tracks for realistic 3D vision.

  • OmniObject3D includes 6,000 real-scanned objects from 190 categories, with textured meshes, point clouds, rendered multiview images, and real-captured video frames.
  • The authors establish four evaluation tracks to reveal observations, challenges, and opportunities in realistic 3D vision.
  • The project will regulate data usage to avoid potential negative social impacts.

A. Additional Information of OmniObject3D

The supplementary information details OmniObject3D’s category coverage, object data, and scan-quality comparison with COLMAP reconstruction.

  • Most categories contain between 10 and 40 objects, with the full class distribution listed in Figure S1.
  • The dataset includes commonly manipulated objects and provides each object’s textured 3D mesh together with several surrounding videos.
  • The supplementary material compares textured scan meshes with COLMAP sparse reconstructions to demonstrate scan completeness and quality.

B. Related Works

Related work spans realistic 3D datasets, robust point-cloud perception, neural radiance fields, novel-view synthesis, and 3D object generation.

  • Robust Point Cloud Perception: Prior robustness benchmarks separately study synthetic corruptions or sim-to-real differences, whereas OmniObject3D evaluates OOD styles and corruptions together.
  • Neural Radiance Field: Neural radiance fields represent scenes with neural networks that map sampled points along camera rays to predicted color and density.
  • Neural Radiance Field: OmniObject3D supports generalizable novel-view synthesis and surface reconstruction through realistic photos, meshes, broad vocabulary, and diverse shapes and appearances.
  • Table R1 organizes perception results by architecture, reporting OOD-style accuracy on OmniObject3D and mean Corruption Error on OmniObject3D-C.
  • Neural Radiance Field: Qualitative NVS comparisons examine single-scene methods across rendered scenes and different data types.
  • 3D Object Generation: Existing 3D generation methods include coarse geometry, implicit surfaces, 3D-aware image synthesis, and textured mesh generation from templates.

C.1. Robust 3D Perception

The supplementary benchmark evaluates point-cloud robustness under seven OOD corruption types and compares NVS performance across data-acquisition settings.

  • The robustness study applies seven OOD corruptions: Scale, Jitter, Drop Global/Local, Add Global/Local, and Rotate.
  • Mean Corruption Error (mCE) averages the errors across corruption types for the evaluation.
  • For three single-scene NVS methods, Blender performs best, SfM-wo-bg is slightly worse, and SfM-w-bg achieves the lowest PSNR.

C.2.1 Single-Scene NVS

Single-scene NVS comparisons show that Plenoxels capture high-frequency textures particularly well, whereas NeRF and mip-NeRF are more robust to dark textures and concave geometry.

  • Plenoxels are especially effective at modeling high-frequency textures such as those on the coconut.
  • NeRF and mip-NeRF are more robust than Plenoxels for dark textures and concave geometry.Plenoxels can suffer from inaccurate geometry in these cases.
  • The dataset enables comprehensive evaluation of these differing method behaviors.

Comparisons of NVS on rendered images and iPhone

The experiments compare NVS and reconstruction methods across rendered and real-captured imagery, sparse views, and varying scene or category settings. Results expose challenges from backgrounds, viewpoint changes, depth guidance, and category-dependent geometry priors.

  • Rendered images and iPhone: Blender-rendered inputs achieve the best visual quality and highest PSNR, while SfM-with-background inputs suffer from significant background error.SfM-without-background foreground quality is only slightly better than SfM-with-background.
  • Rendered images and iPhone: Real-captured videos introduce additional challenges for NeRF-like novel-view synthesis methods.The comparison uses SfM settings derived from iPhone videos and COLMAP camera parameters.
  • Cross-scene NVS: Cross-scene NVS evaluates 3 unseen scenes per category using 3 source views and 10 widely distributed test frames.
  • Cross-scene NVS: Unaligned coordinate systems can impair category-specific spatial priors learned when xyz is provided to the network.The experiments manually introduce nonalignment and suggest naturally unaligned coordinates may cause a more severe impairment.
  • Sparse-view surface reconstruction: 8-view NeuS reaches 7.96 Chamfer distance, remaining worse than the 100-view setting at 6.09.NeuS accuracy improves significantly from 2 to 8 views, while MonoSDF gains slow from 5 to 8 views.
  • Sparse-view surface reconstruction: MonoSDF’s sparse-view performance may be limited by inaccurate depth guidance, while NeuS can preserve coherent shapes and geometry details in some thin-structure cases.
Loading 2301.07525v2…