Source-linked AI summary
Neural Reflectance Fields for Appearance Acquisition
Sai Bi, Zexiang Xu, Pratul Srinivasan, Ben Mildenhall, Kalyan Sunkavalli, Miloš Hašan, Yannick Hold-Geoffroy, David Kriegman, Ravi Ramamoorthi
TL;DR
Existing reconstruction methods struggle to reproduce complex real-scene appearance from images. The paper introduces Neural Reflectance Fields with differentiable physically based ray marching, learns them from collocated cellphone flash images, and reports photo-realistic view synthesis and relighting that significantly outperform previous mesh- and volume-based methods while supporting standard rendering engines.
Problem
Existing reconstruction methods can produce inaccurate reconstructions and artifacts, while mesh-based methods struggle with thin structures and sharp specularities in complex scenes.
Method
A fully connected neural network maps 3D points to volume density, normals, and reflectance properties for differentiable physically based ray marching.
Results
The method produces high-quality view synthesis and relighting with challenging specularities, shadows, occlusions, and fine textures, significantly better than previous mesh- and volume-based methods.
Takeaways & Limitations
The learned representation can be rendered in standard graphics engines and composed with traditional 3D models for scene modeling applications.
Takeaways & Limitations
Too many details can cause slight blurring, and inconsistent adaptive samples may introduce minor relighting flicker in videos.
Abstract
from arXiv · showhide
We present Neural Reflectance Fields, a novel deep scene representation that encodes volume density, normal and reflectance properties at any 3D point in a scene using a fully-connected neural network. We combine this representation with a physically-based differentiable ray marching framework that can render images from a neural reflectance field under any viewpoint and light. We demonstrate that neural reflectance fields can be estimated from images captured with a simple collocated camera-light setup, and accurately model the appearance of real-world scenes with complex geometry and reflectance. Once estimated, they can be used to render photo-realistic images under novel viewpoint and (non-collocated) lighting conditions and accurately reproduce challenging effects like specularities, shadows and occlusions. This allows us to perform high-quality view synthesis and relighting that is significantly better than previous methods. We also demonstrate that we can compose the estimated neural reflectance field of a real scene with traditional scene models and render them using standard Monte Carlo rendering engines. Our work thus enables a complete pipeline from high-quality and practical appearance acquisition to 3D scene composition and rendering.
1 INTRODUCTION
Neural Reflectance Fields jointly represent scene geometry and reflectance, enabling differentiable rendering from practical flash-image captures under novel views and lighting.
- Explicit reconstruction methods often produce inaccurate geometry and rendered images with significant artifacts.
- Neural Reflectance Fields represent both scene geometry and reflectance, unlike prior neural representations focused on color or radiance.
- A differentiable physically based ray-marching framework renders the representation under arbitrary viewpoints and lighting.
- The method reconstructs neural reflectance fields from unstructured cellphone images captured with a built-in flash and collocated camera-light setup.
- Results model complex geometry and reflectance, including intricate geometry, highly specular materials, furry objects, and human portraits, significantly better than stated prior methods.
- The pipeline supports view synthesis, relighting, and composition with traditional 3D models in modern rendering engines.
2 RELATED WORK
The paper addresses limitations of explicit meshes and view-dependent neural representations by combining implicit geometry-reflectance modeling with reflectance-aware differentiable ray marching.
- Neural scene representations have modeled geometry with volumes, point clouds, and implicit functions, while neural networks have also been applied to reflectance.
- Many image-based methods avoid explicit geometry, but view-dependent representations can support only limited viewing ranges or require special fusion techniques.
- Discrete reflectance volumes support relighting but are constrained by fixed resolution, whereas the proposed framework regresses decomposed shading components.
- Traditional reflectance acquisition may require sophisticated devices, while cellphone cameras with built-in flashes provide practical collocated reflectance samples.
- The neural reflectance field bypasses explicit mesh reconstruction and targets thin structures and sharp specularities that mesh-based methods struggle to recover.
- Reflectance-aware ray marching incorporates light transmittance, allowing the method to recover challenging hard shadows and support relighting.
- Unlike approaches that feed viewing or lighting directions into networks, this method uses classical reflectance models and keeps the field direction-independent.
3 REFLECTANCE-AWARE RAY MARCHING
The paper extends differentiable volume ray marching with explicit reflectance and light transport, enabling view synthesis and relighting with realistic shading and shadows. The framework integrates neural scene properties with physically based rendering and supports flexible reflectance models.
- Rendering equation: The framework models radiance by integrating scattered light along camera rays through a volume.Volume density determines extinction and camera-ray transmittance, while scattered light accounts for illumination and reflectance.
- Rendering equation: The rendering equation includes complete one-bounce camera-volume-light paths and explicitly models light transmittance to render realistic shadows.Unlike methods modeling only view transmittance or opacity, the framework accounts for attenuation between each shading point and the light.
- Rendering equation: Explicit reflectance parameters separate light transmittance, reflectance, and incident illumination for reflectance-aware rendering.This supports both view synthesis and relighting with realistic shading and shadowing effects.
- Ray marching: Ray marching estimates the continuous rendering integral by sampling shading points along camera rays and numerically evaluating camera and light transmittance.Light transmittance normally requires an additional ray from the light to each shading point.
- Ray marching: A collocated camera-light setup makes training efficient, while an adaptive transmittance volume enables inference under arbitrary point lights.The collocated setup makes camera and light rays identical during training; inference uses precomputation to approximate light transmittance efficiently.
- Reflectance models: The framework accepts any differentiable reflectance model and demonstrates analytic BRDF and hair/fur reflectance models.The analytic BRDF represents opaque surfaces with diffuse albedo and specular roughness, while hair/fur models extend the formulation to furry objects.
4 NEURAL REFLECTANCE FIELDS
Neural Reflectance Fields represent scene geometry and reflectance continuously and combine with differentiable ray marching for appearance acquisition, view synthesis, and relighting. Adaptive sampling and precomputed light transmittance support efficient rendering under novel viewpoints and lighting, including realistic shadows.
- 4.1 Network: Neural Reflectance Fields use an MLP to represent volume density, normals, and reflectance properties at every 3D scene point.The representation outputs a (4 + m)-dimensional vector containing density, normals, and reflectance parameters.
- 4.1 Network: Frequency-based positional encodings of 3D coordinates are input to the MLP, which predicts view- and light-independent scene properties.The highest frequency level is W = 10; shading and lighting are computed in the ray-marching framework.
- 4.2 Learning neural reflectance fields from flash images: The differentiable representation is trained by minimizing the error between rendered and captured images, using collocated flash light and view to equate camera and light transmittance.This avoids marching an additional light ray at every shading point during training.
- 4.2 Learning neural reflectance fields from flash images: Coarse-to-fine adaptive sampling uses visibility-derived contribution weights to concentrate camera-ray samples around scene structures before final rendering.A coarse network produces the sampling distribution, and a fine network uses the combined samples to compute final radiance.
- 4.3 Efficient rendering under novel light and view: An adaptive transmittance volume precomputes light transmittance, which is interpolated during ray marching to enable efficient rendering with realistic shadows under novel lights and viewpoints.The volume is built by marching rays from a virtual image plane at the point light source, while inference samples both light and camera rays coarsely to finely.
- 4.3 Efficient rendering under novel light and view: Despite training only on collocated images without shadows, the method synthesizes novel non-collocated views and lighting with realistic shadows, specularities, and other appearance effects.The learned volume density meaningfully expresses scene geometry even though the training images contain no shadows.
5 IMPLEMENTATION
The implementation uses practical cellphone-based acquisition, a microfacet reflectance model, coarse-to-fine sampling, and L2-based training with transmittance regularization. Training takes about two days on four RTX 2080Ti GPUs, while inference renders a 512 × 512 image in about 30 seconds.
- Data acquisition: A handheld cellphone with flash can capture the collocated input data; one portrait experiment used 150 selected video frames.Other results used a robotic arm holding a cellphone to facilitate acquisition.
- Reflectance model: The practical reflectance model combines diffuse Lambertian and GGX specular terms, with diffuse albedo and specular roughness among its parameters.The neural reflectance field MLP outputs an 8-D vector for this model.
- Training parameters and loss function: Training randomly samples 50 × 50 pixel rays per batch and uses Adam with an initial learning rate of 0.0001.The method uses 64 coarse and 128 fine samples for adaptive light and camera-ray sampling.
- Training parameters and loss function: The loss combines L2 supervision of coarse and fine radiance with transmittance regularization that encourages opaque-object ray transmittance toward 0 or 1.The regularization strength is β = 0.0001.
- Run time: Training each reflectance-field network takes about 2 days on 4 NVIDIA RTX 2080Ti GPUs, while inference renders a 512 × 512 image in about 30 seconds.Inference uses the adaptive transmittance volume.
6 RESULTS
The results evaluate Neural Reflectance Fields for view synthesis, relighting, diverse real scenes, and integration with Monte Carlo rendering. Across these settings, the method reproduces detailed appearance effects and outperforms prior mesh- and volume-based approaches, while retaining practical limitations.
- Comparisons with previous methods: The method outperforms mesh-based reconstruction by avoiding distorted or missing geometry and recovering realistic geometric details, high specularities, and hard shadows.The comparison attributes mesh failures to challenging regions with little texture, high specularity, or thin structures.
- Comparisons with previous methods: The method outperforms discrete volume rendering by using a continuous representation that recovers high-frequency appearance beyond fixed-resolution volume limits.The neural reflectance field is also compact, with weights consuming only 5 MB of memory.
- Additional results on diverse real scenes: Additional results reproduce detailed geometry, complex textures, specularities, and hard shadows across complex objects, furry objects, and human portraits.The representation also works with the classical fur reflectance model and practical cellphone flash capture for facial appearance.
- Synthetic results: On a synthetic scene, the method accurately reproduces high-frequency textures, specularities, and hard shadows under non-collocated lighting and viewpoints.The rendered images are reported to be very close to ground truth.
- Comparisons with previous methods: Neural Reflectance Fields produce realistic renderings with high-frequency textures, specular highlights, and complex shadowing on complex real scenes.Figure 4 compares captured images with renderings under novel collocated and non-collocated camera-light conditions.
- Integrating with Monte-Carlo renderers: Neural Reflectance Fields can be rendered with standard Monte Carlo engines and composed with traditional 3D models.The authors demonstrate this by converting a captured field into discrete 512 × 512 × 512 volumes for Mitsuba rendering.
- Limitations: The method may produce slightly blurry results for scenes with too many details, occasional dark floaters, and minor relighting flicker in videos.The paper attributes flicker to inconsistent adaptive samples across frames and suggests increased network capacity, masking, or sampling as possible remedies.
7 CONCLUSION
The paper presents a practical neural reflectance-field pipeline that supports high-quality view synthesis, relighting, and scene composition from simple mobile-phone capture. It reproduces challenging appearance effects and integrates real scenes with traditional models in standard rendering engines.
- 7 CONCLUSION: A neural reflectance field encodes volume-rendering properties for real-scene geometry and reflectance acquisition.The representation is learned through a deep training process from cellphone flash images captured with collocated camera and light.
- 7 CONCLUSION: The method renders photo-realistic images under arbitrary camera and non-collocated light positions for high-quality relighting and view synthesis.The reported results include novel lighting and viewpoint renderings compared with synthetic ground truth.
- 7 CONCLUSION: It reproduces challenging effects including specularities, shadows, occlusions, and fine textures more effectively than previous mesh-based and volume-based methods.Synthetic evaluations report images close to ground truth while reproducing high-frequency textures, specularities, and hard shadows.
- 7 CONCLUSION: Estimated neural reflectance fields can be composed with traditional scene models and rendered using standard graphics rendering engines.The paper demonstrates composition of a real scene with a synthetic model under complex illumination.