Source-linked AI summary
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis
Cheng-You Lu, Yi-Shan Hung, Wei-Ling Chi, Hao-Ping Wang, Charlie Li-Ting Tsai, Yu-Cheng Chang, Yu-Lun Liu, Thomas Do, Chin-Teng Lin
TL;DR
Distractor-free radiance fields lack a large-scale, challenging benchmark with clean and cluttered images per scene. DF3DV-1K fills this gap with a diverse dataset, systematic benchmark, and diffusion-based enhancer, identifying robust methods and improving rendering quality by 0.96 dB PSNR and 0.057 LPIPS.
Problem
Distractor-free radiance fields lack a large-scale, challenging benchmark providing clean and cluttered images for each scene.
Method
The paper constructs DF3DV-1K and DF3DV-41, benchmarks ten methods across scenarios, and fine-tunes a diffusion-based enhancer on the dataset.
Results
The benchmark identifies comparatively robust methods and challenging scenarios, while DI2FIX achieves a 0.96 dB PSNR gain and 0.057 LPIPS reduction.
Takeaways & Limitations
DF3DV-1K enables large-scale comparisons and progress beyond scene-specific vision for distractor-free radiance fields.
Takeaways & Limitations
The enhancer may fail when input views are severely corrupted or exhibit confirmation bias.
Abstract
from arXiv · showhide
Advances in radiance fields have enabled photorealistic novel view synthesis. In several domains, large-scale real-world datasets have been developed to support comprehensive benchmarking and to facilitate progress beyond scene-specific reconstruction. However, for distractor-free radiance fields, a large-scale dataset with clean and cluttered images per scene remains lacking, limiting the development. To address this gap, we introduce DF3DV-1K, a large-scale real-world dataset comprising 1,048 scenes, each providing clean and cluttered image sets for benchmarking. In total, the dataset contains 89,924 images captured using consumer cameras to mimic casual capture, spanning 128 distractor types and 161 scene themes across indoor and outdoor environments. A curated subset of 41 scenes, DF3DV-41, is systematically designed to evaluate the robustness of distractor-free radiance field methods under challenging scenarios. Using DF3DV-1K, we benchmark nine recent distractor-free radiance field methods and 3D Gaussian Splatting, identifying the most robust methods and the most challenging scenarios. Beyond benchmarking, we demonstrate an application of DF3DV-1K by fine-tuning a diffusion-based 2D enhancer to improve radiance field methods, achieving average improvements of 0.96 dB PSNR and 0.057 LPIPS on the held-out set (e.g., DF3DV-41) and the On-the-go dataset. We hope DF3DV-1K facilitates the development of distractor-free vision and promotes progress beyond scene-specific approaches. The dataset and leaderboard are available at https://johnnylu305.github.io/df3dv1k_web/.
1 Introduction
Distractor-free radiance field research lacks a large-scale real-world benchmark with clean and cluttered images, limiting reliable evaluation and progress toward generalizable solutions. DF3DV-1K addresses this gap with a diverse dataset, a systematically designed challenging subset, and diffusion-based enhancement results.
- Motivation: Distractor-free radiance field research lacks a large-scale real-world dataset with clean and cluttered images per scene, hindering identification of method strengths and limitations.This also hinders progress from scene-specific methods toward generalizable solutions.
- DF3DV-1K: 1,048 indoor and outdoor scenes, 128 distractor types, 161 scene themes, and 89,924 images comprise DF3DV-1K, with clean and cluttered images for each scene.The dataset mimics real-world capture conditions for distractor-free novel view synthesis research focused on casually captured images.
- DF3DV-41: 41 systematically designed capture scenes form DF3DV-41, covering 17 distractor and scene scenarios for robustness evaluation under challenging conditions.The analysis identifies semantically similar and fluid distractors and nighttime scenes as challenging cases.
- Application: 0.96 dB PSNR and 0.057 LPIPS improvements are achieved by DI2FIX after fine-tuning on DF3DV-1K* and evaluating on DF3DV-41 and the On-the-go dataset.DI2FIX is a distractor-free diffusion-based enhancement model.
2 Related Work
Related work spans generalizable, dynamic, and in-the-wild radiance fields, while distractor-free methods have largely remained scene-specific. Existing benchmarks provide paired clean and cluttered images but are limited or increasingly saturated, motivating a larger real-world dataset.
- Radiance field research: Radiance fields support photorealistic novel view synthesis, with research extending from scene-specific models to generalizable, dynamic, and in-the-wild settings.The survey excludes multi-modal, task-specific, and domain-specific approaches because it focuses on general scenarios.
- Distractor-free radiance fields: Generalizable distractor-free methods emerged only recently, while most distractor-free radiance field methods remain scene-specific and require per-scene optimization.The first generalizable distractor-free 3DGS was fine-tuned from a pretrained generalizable model using synthetic data and a small real-world dataset.
- Existing benchmarks: RobustNeRF introduced a five-scene indoor tabletop dataset spanning four themes and four distractor types, with clean and cluttered image sets created through manual object repositioning.The benchmark was widely adopted but became less challenging for recent state-of-the-art methods, limiting performance differentiation.
- Existing benchmarks: Existing outdoor evaluation datasets offered more challenging scenes but have also become saturated, prompting methods to emphasize subtle visual differences in background regions.The passage also identifies RealX3D and D-RE10K-iPhone as concurrent works providing paired clean and cluttered data.
3 Data Acquisition and Curation
DF3DV-1K is constructed through customized scene planning and manual curation to provide clean and cluttered views for distractor-free radiance-field benchmarking. Its 1,048 scenes and 89,924 consumer-camera images span diverse indoor/outdoor environments, distractor conditions, and capture settings.
- Scene Design: Scene planning specifies each scene’s theme, potential distractors, viewpoint coverage, orientation, and approximate clean and cluttered image counts.Viewpoint coverage ranges from 180° to 360°.
- Data Review and Pose Processing: Operators remove low-quality images and verify that clean image sets contain no distractors, while allowing lower-quality distractors that reflect casual capture.Camera poses are jointly estimated from clean and cluttered images using COLMAP.
- Scale: 1,048 scenes—726 indoor and 322 outdoor—with clean and cluttered image sets contain 89,924 images, averaging approximately 50 images per scene.The clean-image ratio skews lower for efficient benchmarking, whereas cluttered-image ratios are higher to cover diverse distractors.
- Semantic Diversity: DF3DV-41 systematically designs distractor types and scene themes, including semantically similar distractors that challenge feature-based methods.The curated subset is presented as part of the dataset’s diversity-oriented design.
- Capture Diversity: Over nine months, the dataset was captured with 12 consumer cameras, from iPhone to Samsung smartphones, across nine image resolutions.The collection uses consumer cameras and spans indoor and outdoor scenes.
4 Benchmarks & Experiments
The benchmarks show that DF3DV-1K and DF3DV-41 are challenging, diverse testbeds where AsymGS and RobustSplat are the most robust evaluated methods. The experiments also introduce DI2FIX, a plug-and-play enhancer trained from filtered DF3DV-1K reconstructions to remove distractor artifacts without modifying radiance-field models.
- Benchmark Methods and Robustness: AsymGS and RobustSplat are the most robust methods on DF3DV-1K and DF3DV-41, followed by OCSplats and DeGauss.The benchmark evaluates these methods alongside SLS, DeSplat, WildGaussians, T-3DGS, T-3DGS-TMR, and 3DGS.
- Benchmark Difficulty: DF3DV-1K and DF3DV-41 are larger, more diverse, and more challenging than prior benchmarks, discouraging per-scene tuning.Compared with RobustNeRF, they produce a wider performance challenge, whereas RobustNeRF shows modest improvements over 3DGS and many methods exceed 29 dB PSNR.
- Observed Limitations: Semantic similarity and semi-transparent, spatially scattered fluid distractors make distractor identification and masking difficult for existing methods.Methods relying on semantic features struggle when distractors resemble static objects, while online masking can remove parts of fluid phenomena unintentionally.
- DI2FIX: DI2FIX is a plug-and-play 2D enhancer fine-tuned from DIFIX using 316,890 candidate pairs reconstructed from cluttered DF3DV-1K* scenes and rendered at clean-image views.Pairs are retained when LPIPS≤γ, with γ = 0.5 in the final setting; DF3DV-41 is excluded from training.
- Training Data Scale: Larger and more diverse training sets improve DI2FIX’s performance and help it identify distractors while avoiding edits to static regions.Experiments vary training sets of 250, 500, 750, and 1,007 scenes, while qualitative results show progressively improved artifact removal and preservation of static content.
- Training Data Filtering: A moderate LPIPS filtering threshold provides a better balance between training-pair quality and diversity than overly strict or loose thresholds.Strict thresholds exclude challenging pairs, whereas loose thresholds add noisy samples that can cause undesired scene modifications.
5 Conclusion
The paper introduces DF3DV-1K as a large-scale real-world dataset and benchmark for distractor-free novel view synthesis, including a challenging subset for scenario-wise evaluation. It also identifies limitations under severely corrupted inputs or confirmation bias and suggests multi-reference frameworks as future work.
- DF3DV-1K contains 1,048 indoor and outdoor scenes with clean and cluttered images, spanning 161 scene themes and 128 distractor types.
- DF3DV-41 is a systematically designed challenging subset enabling scenario-wise evaluation within a comprehensive benchmark across 10 representative methods.
- Inputs that are severely corrupted or exhibit confirmation bias may cause failure, motivating multi-reference frameworks as a future direction.
S.1 Author Statement
The supplementary material documents DF3DV-1K’s benchmark setup and provides a standardized, user-friendly dataset organization for radiance-field research. It describes dataset access, licensing, scene storage, reconstruction outputs, annotations, and camera-parameter formats.
- Dataset Organization and File Structure: The dataset and leaderboard are publicly available, data collection followed applicable local laws and regulations, and the dataset uses the CC BY-NC 4.0 license.The dataset and leaderboard are hosted at the project website.
- Dataset Organization and File Structure: Scene files include clean and cluttered images, COLMAP sparse reconstructions, undistorted images and camera parameters, and instant-ngp JSON files for all, cluttered, or clean images.Cluttered images use the “clutter_xxx.JPG” naming pattern, while clean images use “extra_xxx.JPG”.
- Dataset Organization and File Structure: The approximately 2 TB dataset provides per-scene JSON annotations and a layout designed for straightforward integration with existing radiance-field training and evaluation pipelines.The organization follows commonly adopted structures in distractor-free radiance-field research, with minor renaming for clarity.
S.2 DF3DV-1K and DF3DV-41 Benchmarks
The DF3DV-1K and DF3DV-41 benchmarks evaluate nine distractor-free radiance field methods and 3DGS across challenging scenarios, identifying robust methods, difficult conditions, and the tradeoff of removing distractors. Scenario-wise results highlight AsymGS, RobustSplat, and OCSplats as consistently robust, while semantically similar distractors, fluid distractors, and nighttime scenes remain especially challenging.
- Benchmark Setup: Nine recent open-source distractor-free radiance field methods and 3DGS are evaluated on the DF3DV-1K and DF3DV-41 benchmarks.The evaluation includes AsymGS, RobustSplat, OCSplats, DeGauss, SLS, DeSplat, WildGaussians, T-3DGS, and T-3DGS-TMR.
- Experimental Details: The evaluation standardizes data preprocessing and training configuration across official method implementations for fair comparison.The compared methods otherwise use differences such as on-the-fly undistortion, scene-dependent parameters, or pre-undistortion downsampling.
- Scenario-wise Performance: AsymGS, RobustSplat, and OCSplats perform robustly across all DF3DV-41 scenarios, achieving consistent improvements over 3DGS.Scenario-wise improvements are reported for PSNR, SSIM, and LPIPS relative to 3DGS.
- Scenario-wise Performance: Semantically similar distractors, fluid distractors, and nighttime scenes are the most challenging DF3DV-41 scenarios.They show low 3DGS performance and relatively limited improvements from distractor-free radiance field methods compared with other scenarios.
- View-dependent Static Objects Tradeoff: The benefits of removing distractors outweigh occasional misclassification of view-dependent static objects.DF3DV-1K contains 388 scenes with strong view-dependent static objects, including 18 selected DF3DV-41 scenes evaluated with human-annotated masks.
S.3 Beyond Scene-Specific Methods
This section presents DI2FIX, a plug-and-play 2D enhancement module adapted for distractor-free radiance-field renderings without clean reference images. DI2FIX consistently improves PSNR and LPIPS across methods, generalizes to unseen methods, and has limitations under severe degradation and strong distractors.
- Experimental Details: DI2FIX is a plug-and-play 2D enhancement module designed to improve distractor-free radiance-field rendering quality.It is built upon the diffusion-based DIFIX image-to-image model and is adapted to settings where clean reference images are unavailable.
- Per-method Improvement: DI2FIX consistently improves all enhanced methods in PSNR and LPIPS, while SSIM changes remain relatively limited.The limited SSIM changes are attributed to DIFIX not using an SSIM loss and to related perceptual trade-offs reported in prior work.
- Per-method Improvement: Domain-specific fine-tuning is important: original DIFIX yields marginal improvements, and DIFIX+RobustNeRF produces slightly better but inconsistent gains.These comparisons highlight the value of large-scale, domain-specific datasets such as DF3DV-1K.
- Our-of-distribution Test: Worst-case out-of-distribution changes remain bounded within 0.1 PSNR, a 0.005 decrease in SSIM, and a 0.009 increase in LPIPS.The leave-one-method-out protocol indicates stable generalization to radiance-field methods not observed during training.
- Limitations of DI2FIX: DI2FIX may fail under extreme degradation or confirmation bias, especially when corrupted inputs lack reliable evidence or multiple views contain a strong distractor.In such cases, diffusion priors can reduce reconstruction fidelity or reinforce the distractor.
S.4 DF3DV-1K and DF3DV-41
DF3DV-1K introduces a large-scale dataset with paired clean and cluttered images for distractor-free novel view synthesis, including the systematically designed DF3DV-41 benchmark. Its evaluations show realistic, challenging conditions that distinguish method robustness, expose difficult distractor scenarios, and support improvements from distractor-aware enhancement.
- Dataset construction: DF3DV-1K contains 1,048 indoor and outdoor scenes, 128 distractor types, 161 scene themes, and 89,924 images with clean and cluttered views per scene.The data are captured to reflect practical usage scenarios.
- Benchmark findings: Ranking trends on DF3DV-1K broadly align with the chronological progression of recent research, while the benchmark also identifies weaknesses in recent distractor-free radiance field methods.This alignment supports DF3DV-1K as a realistic and challenging benchmark that reflects progress in the field.
- Per-scenario evaluation: AsymGS, RobustSplat, and OCSplats show consistent PSNR robustness across scenarios, while semantically similar distractors are identified as the most challenging PSNR scenario.The table reports PSNR for 3DGS and performance differences relative to 3DGS for other methods.
- View-dependent static objects: Distractor-free radiance field methods outperform 3DGS on view-dependent static objects, indicating that removing distractors can improve rendering despite possible effects on view-dependent appearance.AsymGS also renders a console reflected on a screen, demonstrating separation of view-dependent static objects from distractors.
- Benchmark design: DF3DV-41 systematically covers diverse challenging distractor and scene scenarios, enabling clear qualitative comparison and robustness evaluation across radiance field methods.Its scenarios include color-similar, fluid, frontal-occlusion, highly reflective, large-scale, semantically similar, semi-transparent, nighttime, and other distractors.
- Enhancer evaluation: DI2FIX, a DIFIX enhancer fine-tuned on DF3DV-1K*, shows promising mitigation of distractor artifacts and improved visual quality across radiance field methods, although it cannot always restore severely degraded regions.The reported enhancer comparison measures mean performance change on On-the-go and DF3DV-41 relative to Vanilla.