Source-linked AI summary

NTIRE 2026 3D Restoration and Reconstruction in Real-world Adverse Conditions: RealX3D Challenge Results

Shuhong Liu, Chenyu Bao, Ziteng Cui, Xuangeng Chu, Bin Ren, Lin Gu, Xiang Chen, Mingrui Li, Long Ma, Marcos V. Conde, Radu Timofte, Yun Liu, Ryo Umagami, Tomohiro Hashimoto, Zijian Hu, Yuan Gan, Tianhan Xu, Yusuke Kurose, Tatsuya Harada, Junwei Yuan, Gengjia Chang, Xining Ge, Mache You, Qida Cao, Zeliang Li, Xinyuan Hu, Hongde Gu, Changyue Shi, Jiajun Ding, Zhou Yu, Jun Yu, Seungsang Oh, Fei Wang, Donggun Kim, Zhiliang Wu, Seho Ahn, Xinye Zheng, Kun Li, Yanyan Wei, Weisi Lin, Dizhe Zhang, Yuchao Chen, Meixi Song, Hanqing Wang, Haoran Feng, Lu Qi, Jiaao Shan, Yang Gu, Jiacheng Liu, Shiyu Liu, Kui Jiang, Junjun Jiang, Runyu Zhu, Sixun Dong, Qingxia Ye, Zhiqiang Zhang, Zhihua Xu, Zhiwei Wang, Phan The Son, Zhimiao Shi, Zixuan Guo, Xueming Fu, Lixia Han, Changhe Liu, Zhenyu Zhao, Manabu Tsukada, Zheng Zhang, Zihan Zhai, Tingting Li, Ziyang Zheng, Yuhao Liu, Dingju Wang, Jeongbin You, Younghyuk Kim, Il-Youp Kwak, Mingzhe Lyu, Junbo Yang, Wenhan Yang, Hongsen Zhang, Jinqiang Cui, Hong Zhang, Haojie Guo, Hantang Li, Qiang Zhu, Bowen He, Xiandong Meng, Debin Zhao, Xiaopeng Fan, Wei Zhou, Linzhe Jiang, Linfeng Li, Louzhe Xu, Qi Xu, Hang Song, Chenkun Guo, Weizhi Nie, Yufei Li, Xingan Zhan, Zhanqi Shi, Dufeng Zhang, Boyuan Tian, Jingshuo Zeng, Gang He, Yubao Fu, Weijie Wang, Cunchuan Huang

arXiv:2604.04135v2cs.CV

TL;DR

The paper examines how to reconstruct 3D scenes from real-world low-light and smoke-degraded multiview images, where existing methods are primarily validated on clean captures. It reviews the NTIRE 2026 3DRR Challenge, its benchmark and evaluation protocol, and the submitted restoration-and-reconstruction pipelines. The challenge organizes evaluation around low-light enhancement and smoke restoration using RealX3D degraded views and clean-capture novel-view targets.

  • Problem

    Existing NeRF and 3DGS methods are predominantly developed and benchmarked on clean captures, while faithful reconstruction under real-world adverse conditions remains an open challenge.

  • Method

    The paper reviews a challenge using RealX3D multiview data, two degradation tracks, and submitted pipelines that combine image restoration, geometric priors, and 3D reconstruction.

  • Results

    The challenge evaluates reconstructed novel views against clean captures using average PSNR for ranking and SSIM as a tiebreaker.

  • Takeaways & Limitations

    The challenge provides a dedicated benchmark and evaluation setting for developing reconstruction methods under low-light and smoke degradation.

Abstract

from arXiv · show

This paper presents a comprehensive review of the NTIRE 2026 3D Restoration and Reconstruction (3DRR) Challenge, detailing the proposed methods and results. The challenge seeks to identify robust reconstruction pipelines that are robust under real-world adverse conditions, specifically extreme low-light and smoke-degraded environments, as captured by our RealX3D benchmark. A total of 279 participants registered for the competition, of whom 33 teams submitted valid results. We thoroughly evaluate the submitted approaches against state-of-the-art baselines, revealing significant progress in 3D reconstruction under adverse conditions. Our analysis highlights shared design principles among top-performing methods and provides insights into effective strategies for handling 3D scene degradation.

1. Introduction

The paper frames 3D reconstruction under real-world adverse conditions as an open problem because existing NeRF and 3DGS methods are mainly developed for clean captures. The NTIRE 2026 challenge addresses this gap through dedicated evaluation of low-light and smoke degradation.

  • Real-world adverse conditions severely corrupt input images, while accurate 3D reconstruction requires high-fidelity multi-view inputs for geometry and appearance.
  • Existing NeRF and 3DGS methods are predominantly developed and benchmarked under controlled, clean-capture settings.
  • The first NTIRE 2026 3DRR Challenge was launched to highlight adverse-condition reconstruction, provide a comprehensive benchmark, and drive research progress.
  • The challenge features two tracks targeting low-light and smoke degradations.

2. Tracks and Competition

The competition evaluates 3D low-light enhancement and smoke restoration on RealX3D scenes with degraded inputs, ground-truth poses, and clean-capture novel-view targets. Ranking is based primarily on average PSNR, with SSIM as a tiebreaker.

  • The challenge dataset is a RealX3D subset containing diverse real-world multiview pairs of degraded and clean captures.
  • The two tracks cover 3D low-light enhancement and 3D smoke restoration, with seven unique scenes per track.
  • Each scene provides degraded views and ground-truth camera poses, while participants submit reconstructed scenes and rendered novel-view synthesis results.
  • Final ranking uses average PSNR across scenes, with SSIM serving only as a tiebreaker.

3. Track 1: 3D Lowlight Enhancement

Track 1 methods improve low-light reconstruction by combining image enhancement with geometry-aware or multi-branch 3DGS optimization. Their designs use complementary pseudo-targets, fusion strategies, and staged training to stabilize appearance and structure.

  • FuME-GS: Fusion-guided enhancement generates candidates from multiple low-light models, fuses them regionally, and uses depth and enhanced appearance for 3DGS initialization.
  • CISP-GS: 3DGS4LL optimizes three branches using analytical, calibrated-ISP, and frequency-split pseudo-ground-truth supervision.
  • CISP-GS: The 3DGS4LL backbone trains from scratch with early densification, progressive spherical-harmonic growth, and frozen networks used only for pseudo-ground-truth generation.

3.3. TCIDNet-IBGS

The section presents low-light reconstruction pipelines that combine illumination-aware restoration, geometry-driven initialization, appearance disentanglement, and point-cloud refinement. These methods differ in how they protect 3D structure from noisy enhancement artifacts.

  • TCIDNet-IBGS: TCIDNet-IBGS combines CIDNet-based HVI restoration with geometry-driven rendering, using coordinated intensity and chromatic streams plus adaptive tone correction.
  • IDEAL: IDEAL uses pseudo-RGB and pseudo-depth inputs, while an Appearance Decoder absorbs view-dependent illumination noise so the base 3DGS preserves physical geometry.
  • IDEAL: IDEAL discards the Appearance MLP during novel-view synthesis after using it during training to model view-dependent artifacts.
  • NAKA-GS: NAKA-GS combines Naka-guided chroma correction with VGGT dense point clouds and progressive point pruning before 3DGS initialization.

3.6. GREP-GS: Global-Residual Decomposition with Hybrid Edge Prior for Low-Light 3DGS

GREP-GS uses enhanced images and relative depth priors to optimize 3DGS in a canonical display space while modeling view-specific enhancement residuals.

  • Pseudo-enhanced targets and relative depth priors are extracted before optimization in the canonical display space.
  • Global and per-view adapters address enhancement bias that varies across input views.
  • Retinexformer and MoGe2 generate enhanced images and relative depth maps from low-light inputs.
  • The method separates geometry-consistent canonical appearance from view-dependent residual corrections caused by single-image enhancement.
  • 3DGS jointly optimizes geometry, opacity, anisotropic scale, and spherical-harmonic coefficients from randomly initialized Gaussians.The renderer predicts RGB, opacity, and expected depth for each camera pose.

3.8. Space-GS

Space-GS combines structure-adaptive low-light enhancement with 3DGS to preserve geometry and multi-view consistency through dual-stream feature interaction.

  • Space-GS combines GeoLight with 3DGS to preserve geometric structures and multi-view consistency.
  • SAHT maps RGB inputs to HVI space, separating illumination from hue and saturation.
  • A dual-stream backbone processes chrominance and luminance features and exchanges information through DMMA.The architecture uses separate chrominance and luminance streams.
  • The model is trained on 256×256 patches with flipping and random Gamma perturbations.Optimization uses Adam with cosine-annealed learning rates and RGB, HVI, and Fourier-domain losses.

3.9. ELoG-GS: Extreme Low-light Optimized Gaussian Splatting

ELoG-GS follows a restoration-then-reconstruction design, using zero-shot restoration and depth-based initialization before hybrid reconstruction and post-enhancement.

  • ELoG-GS explicitly separates restoration, hybrid dual-branch reconstruction, and post-enhancement into three stages.
  • Zero-shot Retinexformer recovers latent scene information, while VGGT depth estimation and voxelized fusion initialize geometry without relying on SfM.
  • AdaTone-GS generates scene-adaptive pseudo ground truths using a soft-exponential tone curve controlled by the 90th luminance percentile.The gain is set as g = −ln(0.2)/p90 to adapt to scene exposure.
  • Each Gaussian in AdaTone-GS includes depth, structure-prior, noise, and illumination features, totaling 11 channels.
  • A cascaded denoiser and learnable tonemapper reconstruct low-light observations, while depth reprojection enforces cross-view geometric consistency.
  • AdaTone-GS training combines pseudo-GT reconstruction, low-light reconstruction, depth, structure, cross-view consistency, and denoiser total-variation losses.

3.11. SLL-GS: Staged Low-Light 3DGS

SLL-GS uses staged optimization to separate luminance recovery from chroma correction, combining sparse geometric initialization, depth priors, and multi-view consistency.

  • The pipeline restores appearance in Stage 3 with an illumination head and a YCbCr chroma residual.
  • SLL-GS decouples luminance recovery from chroma correction through a three-stage low-light 3DGS pipeline.
  • Stage 1 initializes Gaussians from fixed-pose COLMAP sparse points and random points, using RGB, exposure, depth, and reprojection supervision.
  • Stage 2 refines geometry without densification while reusing COLMAP sparse points as weak geometric guidance.
  • Training runs for 15,000 iterations in Stage 1 and 5,000 iterations in each of Stages 2 and 3.
  • GammaGS instead applies uniform gamma correction before 3DGS reconstruction and learns a per-channel affine transform for final color calibration.The affine calibration is learned by least squares and applied to final renders.

3.13. DarkIR-GS

DarkIR-GS jointly optimizes low-light illumination restoration and 3D Gaussian Splatting reconstruction through differentiable rasterization. A related two-stage low-light pipeline enhances views before geometric filtering and Luminance-GS training.

  • Differentiable rasterization back-propagates RGB reconstruction gradients into both the 3DGS parameters and enhancement module.
  • DarkIR-GS restores illumination with a lightweight encoder-decoder before processing enhanced images with 3DGS.Its total loss combines L1 RGB reconstruction and pixel-wise enhancement losses.
  • The two-stage 3DLLR pipeline applies RetinexFormer enhancement, COLMAP geometric consistency filtering, and a modified Luminance-GS training objective.The modified objective combines log Charbonnier and SSIM losses.
  • 3DLLR training uses 10,000 iterations with Adam and a loss combining low-light reconstruction, enhancement, spatial-consistency, and histogram-prior terms.Gaussian densification runs every 100 iterations from steps 500 to 8000.

3.15. LLE-GS: A Unified Reconstruction Pipeline Combining Lowlight Restoration with 3DGS

LLE-GS combines transformer-based low-light restoration with a view-adaptive 3D reconstruction backend. Its hierarchical representation uses visible anchors and local conditions to support adaptive rendering.

  • LLE-GS uses an illumination-guided transformer with self-attention to estimate illumination and repair degradations in low-light multi-view inputs.
  • Its backend constructs a hierarchical, region-aware 3D representation from SfM points and dynamically decodes neural Gaussians from visible anchors.Decoding is conditioned on local features, viewing distance, and relative direction.
  • The front end is trained with rotation, flipping, and crop augmentation using Adam and cosine annealing to minimize MAE.
  • Back-end anchors and MLPs are optimized end-to-end with L1, SSIM, and volume regularization losses.An error-driven strategy grows anchors from accumulated gradients and prunes low-opacity anchors.

3.16. RunAI-Harmony3D

RunAI-Harmony3D combines global photometric harmonization with complementary geometry-stable and low-light-recovery 3D branches. It uses 2D photometric and pose-conditioned depth priors within a four-stage pipeline.

  • RunAI-Harmony3D combines I2-NeRF as a geometry-stable anchor with LITA-GS for low-light recovery.
  • HVI-CIDNet, DarkIR, and DA3 provide 2D photometric and pose-conditioned depth priors.
  • The pipeline comprises scene diagnostics, scene-global photometric harmonization, weak-prior creation, and dual-branch optimization.
  • Development scenes tune routing weights and harmony strength, while official training views are used for blind-scene runs.
  • For MilkCookie, original low-light RGB images remain optimization inputs, while a harmonized copy is generated only for the DA3 depth prior.This restricts LITA-GS to the darkest flat regions.

4. Track 2: 3D Smoke Restoration

Smoke-restoration methods combine image enhancement, physical degradation modeling, generative refinement, and 3DGS reconstruction. The approaches differ in whether they restore views before reconstruction, model view-dependent smoke directly, or use hybrid residual and physics-based representations.

  • GenSmoke-GS: GenSmoke-GS sequentially applies preliminary restoration, dark-channel dehazing, illumination correction, and MLLM-guided refinement before 3DGS reconstruction.The MLLM prompt preserves geometry, layout, and scene structure while improving visibility and local details.
  • Smoke-GS: Smoke-GS models view-dependent smoke appearance by encoding pixel ray directions and predicting medium RGB, scattering, and attenuation parameters.The predicted medium terms are fused with the 3DGS rendering.
  • Dehaze-then-reconstruct: The dehaze-then-reconstruct pipeline generates pseudo-clean images with Nano Banana Pro, normalizes brightness across views, and reconstructs scenes with physics-informed 3DGS.Its composite loss includes RGB, SSIM, dark-channel, pseudo-depth, and gradient terms.
  • MSDG: MSDG applies mild pre-dehazing, 3D-UIR reconstruction, scene-specific dehazing, Difix3D+ repair, mean fusion, and Gaussian smoothing.
  • DiT-IBGS: DiT-IBGS uses a parameter-efficient FLUX.1-dev latent-diffusion framework with frozen VAE and text encoders and LoRA-based transformer fine-tuning.
  • DePhy-GS: DePhy-GS combines conservative dehazing, deterministic MLP refinement, dual-layer 3DGS residual modeling, and a physical head predicting transmission and airlight.The final clean rendering combines base and residual Gaussian layers with a learnable residual gain.
Loading 2604.04135v2…