Source-linked AI summary
Learning Physics-guided Face Relighting under Directional Light
Thomas Nestmeyer, Jean-François Lalonde, Iain Matthews, Andreas M. Lehrmann
TL;DR
Realistic face relighting must transfer captured faces into new lighting while handling complex reflectance and directional shadows. The paper uses end-to-end intrinsic decomposition, diffuse physics-based rendering, and residual refinement, and reports realistic results on a varied light-stage dataset. Its scope is limited by failures in extremely dark input regions and by assumptions challenged by uncontrolled multi-source illumination.
Problem
Realistic AR and telepresence require faces captured under one illumination to match a different environment, but simplified models omit complex reflectance, cast shadows, and specularities.
Method
The model jointly de-lights and relights faces using intrinsic components, diffuse directional-light rendering, and a neural residual conditioned on albedo, normals, and the diffuse render.
Results
The method produces realistic directional-light relighting with non-diffuse effects including specularities and hard-cast shadows, including challenging poses and illumination conditions.
Takeaways & Limitations
Structured physics-guided decomposition improves interpretability and enables direct manipulation or extraction of semantically meaningful intermediate layers for downstream tasks.
Takeaways & Limitations
Performance worsens when the input is so dark that the camera mostly returns noise, and uncontrolled portraits violate the single-directional-light assumption.
Abstract
from arXiv · showhide
Relighting is an essential step in realistically transferring objects from a captured image into another environment. For example, authentic telepresence in Augmented Reality requires faces to be displayed and relit consistent with the observer's scene lighting. We investigate end-to-end deep learning architectures that both de-light and relight an image of a human face. Our model decomposes the input image into intrinsic components according to a diffuse physics-based image formation model. We enable non-diffuse effects including cast shadows and specular highlights by predicting a residual correction to the diffuse render. To train and evaluate our model, we collected a portrait database of 21 subjects with various expressions and poses. Each sample is captured in a controlled light stage setup with 32 individual light sources. Our method creates precise and believable relighting results and generalizes to complex illumination conditions and challenging poses, including when the subject is not looking straight at the camera.
1. Introduction
Face relighting must recover intrinsic properties and reproduce target illumination despite complex reflectance, hard shadows, and limited training data. The proposed end-to-end model combines physics-based rendering with neural residual refinement to relight faces under directional lighting.
- Relighting requires de-lighting reflectance, geometry, and lighting before rendering under a desired target illumination.
- Human-face relighting is difficult because faces exhibit subsurface scattering, view-dependent reflectance, spatial variation, and perceptually salient rendering errors.
- Simplified diffuse and smooth-lighting assumptions cannot represent effects such as specularities and hard shadows from point-like sources.
- The model combines diffuse physics-based rendering of intrinsic components with residual refinement for shadows and other non-diffuse effects.The refinement is conditioned on albedo, normals, and the diffuse rendering.
- The approach is trained end-to-end to de-light and relight faces, using directional lighting that can generalize to complex outputs as sums of point lights.
- The study introduces a portrait dataset with varied lighting conditions and poses and reports realistic relighting of complex non-diffuse faces from a single input image.
2. Related work
Prior work estimates intrinsic properties, lighting, or relit appearance using inverse-rendering, neural, and image-translation methods. The paper distinguishes its single-image face relighting setting by targeting directional-light shadows and non-diffuse effects without requiring special test-time capture.
- Earlier intrinsic-decomposition methods recover shape, illumination, reflectance, or shading using priors, filtering, or deep networks.
- Relighting studies include multi-image or selected-direction methods, while image-to-image translation with multi-illumination data targets full-scene relighting.
- Methods using spherical harmonics or similar smooth-lighting models do not handle sharp cast shadows, whereas this work targets directional-light face relighting.
- Prior face-relighting approaches include recognition normalization, geometry or texture reconstruction, and neural decomposition into normals, albedo, and spherical-harmonic lighting.
- A related directional-light method requires spherical-gradient image pairs at test time, while portrait lighting transfer fails when adding or removing non-diffuse effects.
3. Architecture
The architecture combines explicit diffuse rendering with a learned residual to model non-diffuse facial effects. It is trained end-to-end to infer intrinsic components and generate relit images under directional lighting.
- 3. Architecture: The model decomposes relighting into physics-based diffuse rendering followed by physics-guided residual refinement for non-diffuse effects.The residual can represent specularities, cast shadows, and effects outside the diffuse rendering model.
- 3.1. Image formation process: The image-formation model assumes directional lighting and diffuse materials, while the unconstrained residual captures effects outside the diffuse model.The residual is conditioned on albedo, normals, and diffuse rendering and is trained against photometric-stereo guidance.
- 3.1. Image formation process: The hybrid design assigns most image intensity to an explicit physics-based model while leaving specular highlights and other unconstrained effects to the residual.This is intended to improve physical consistency and ease CNN learning for non-diffuse appearance.
- 3.2. Physics-guided relighting: The differentiable recognition and generative components are stacked for end-to-end learning from an input image to a relit result.Losses supervise albedo, normals, shading, diffuse rendering, visibility, and residual predictions.
- 3.2. Physics-guided relighting: Stage 1 predicts albedo and normals, computes target-light shading, and produces a diffuse render; Stage 2 predicts a residual and visibility map from the inferred modalities.The generator uses a U-Net with grouped convolutions for independent intrinsic predictions, with normals re-normalized to unit vectors.
4. Data
The dataset uses a calibrated multi-view light stage to capture diverse facial expressions and lighting conditions. It yields millions of relighting pairs from 21 subjects, with photometric-stereo decompositions and augmentation for training.
- 4. Data: The capture system uses 6 synchronized cameras and 32 white LEDs, recording calibrated HDR imagery at 2048×1080 and 60 fps.Each LED is flashed for one frame while subjects hold a static expression through the 32-frame cycle.
- 4. Data: 2,961,408 relighting pairs were collected from 482 sequences involving 21 subjects, using 32 input and output lights within each sequence and camera.The split assigns 81% of subjects to training and 9.5% each to validation and testing.
- 4. Data: Photometric stereo separates each input frame into albedo, shading with normals, and a non-diffuse residual for training supervision.The residual is computed as R = I − A ⊙ S.
- 4.2. Augmentation: Training augmentation flips images, perturbs light calibration with Gaussian noise, and rescales images and lighting to improve coverage of the relighting space.Linear intensity scaling was tested but did not provide substantial benefits over training without scaling.
- 4.2. Augmentation: Table 1 evaluates five training losses against five validation metrics for both pix2pix and the structured approach.The table compares all pairwise loss–metric combinations and marks the best model for each evaluation metric.
5. Experiments
The experiments compare the proposed physics-guided relighting architecture with photometric-stereo, SfSNet, and pix2pix baselines under qualitative and quantitative settings. Results show advantages from modeling non-diffuse effects while retaining physics-based intrinsic guidance.
- Evaluation metric: DSSIM-trained models generalize better on the validation set for most error metrics, so the final test models use DSSIM training.LPIPS is the exception: models trained with LPIPS perform better when evaluated with LPIPS.
- Qualitative evaluation: Retrained SfSNet is more accurate than its pretrained variant, but its diffuse assumption produces flatter results and misses specularities.The pretrained model also shows bias toward an albedo resembling skin color in its training data.
- Qualitative evaluation: Pix2pix generates promising results but often produces physically implausible artifacts, including missing shadows.Its domain-agnostic architecture lacks the structured image-formation guidance used by the proposed model.
- Qualitative evaluation: The proposed architecture typically produces the most realistic relit faces, including specular highlights and cast shadows that competing methods often miss.It also handles a hand occluding the face, where strong cast shadows must be introduced or removed.
- Quantitative evaluation: Table 2 compares test-set performance with known and unknown source illumination, using DSSIM-trained models across the reported baselines.The evaluation includes SfSNet, pix2pix, and a photometric-stereo diffuse reconstruction reference.
- Qualitative evaluation: Compared with mass-transport relighting and spherical-harmonics approaches, the proposed technique captures cast shadows and specular highlights under portrait relighting conditions.The cited comparison reports that mass transport fails to generate these effects, while spherical harmonics handles smooth lighting exclusively.
6. Extensions
The model generalizes beyond explicit source illumination to environment-map targets and portraits captured outside the light-stage domain, including uncontrolled office lighting. These extensions demonstrate broader operating conditions while exposing limitations from unknown illumination and imaging pipelines.
- The model generalizes to unknown input illumination, environment-map relighting, and portraits captured in the wild.These scenarios are presented as extensions beyond the controlled capture setting.
- 6.1. Relighting with unknown source illumination: Without explicit source illumination, the model incurs a small performance drop but can match or outperform pix2pix given source-light access.This comparison is reported for the model variant trained without explicit input-light information.
- 6.2. Relighting with environment maps: Environment maps are approximated by assigning one directional light to each 64×32 pixel and mixing predictions by color and intensity.The procedure exploits additive light contributions rather than using importance sampling.
- 6.3. Relighting in the wild: In-the-wild portraits are relit toward three target lights despite uncontrolled office illumination and imaging-pipeline discrepancies.The experiment uses manually selected source lighting, color correction, and hand-masked backgrounds.
- 6.3. Relighting in the wild: In-the-wild results cannot be quantitatively or qualitatively compared with ground truth because corresponding relit images do not exist.The uncontrolled source lighting and imaging pipeline also violate the single-directional-light assumption.
7. Conclusion
The paper presents an end-to-end, structured face-relighting method that combines intrinsic decomposition, diffuse rendering, and neural residual refinement. It reports realistic non-diffuse effects and broader interpretability, while noting degraded performance in extremely dark inputs.
- The method accurately reproduces specularities and hard-cast shadows through structured intrinsic decomposition and re-rendering.Its architecture combines an explicit generative rendering process with a non-diffuse neural refinement layer.
- On the directional-light dataset, the integrated physics-based renderer and neural refinement layer proved superior to the baselines.The authors also report qualitatively closer shadows and albedo maps than all baselines.
- The structured representation improves interpretability and permits direct manipulation or extraction of semantically meaningful intermediate layers.The claimed benefits extend beyond raw performance to downstream face-centric applications.
- Performance worsens when extremely low illumination causes the camera to return mostly noise.The paper suggests explicitly identifying such pixels and applying dedicated context-conditioned infilling as future work.