Source-linked AI summary

Neural Rerendering in the Wild

Moustafa Meshry, Dan B Goldman, Sameh Khamis, Hugues Hoppe, Rohit Pandey, Noah Snavely, Ricardo Martin-Brualla

arXiv:1904.04290v1cs.CVcs.GR

TL;DR

Total scene capture requires realistic views of a scene across varying appearances and transient states from publicly available photos. The paper uses explicit 3D proxy geometry, staged appearance learning, and semantic conditioning to rerender such images. It reports realistic results across five challenging datasets, with its outputs preferred 69.9% of the time in a user study, while infrequent appearances such as night remain difficult to model.

  • Problem

    Total scene capture seeks realistic arbitrary viewpoints across diverse illumination and transient states, but uncontrolled internet photos make appearance diversity difficult to capture and existing reconstructions lack image realism.

  • Method

    The method reconstructs a dense point-cloud proxy, factorizes viewpoint, appearance, and semantics, and trains a neural rerenderer with staged appearance encoding and semantic-mask conditioning.

  • Results

    69.9% of user-study judgments preferred images generated by the authors’ system over the comparison outputs, with preferences on all but one image set.

  • Takeaways & Limitations

    The system provides a first step toward realistic viewpoint and appearance manipulation from unstructured internet photos, including semantic control over transient objects.

  • Takeaways & Limitations

    The approach does not reliably model infrequent appearances such as nighttime scenes, and focuses primarily on static scene parts.

Abstract

from arXiv · show

We explore total scene capture -- recording, modeling, and rerendering a scene under varying appearance such as season and time of day. Starting from internet photos of a tourist landmark, we apply traditional 3D reconstruction to register the photos and approximate the scene as a point cloud. For each photo, we render the scene points into a deep framebuffer, and train a neural network to learn the mapping of these initial renderings to the actual photos. This rerendering network also takes as input a latent appearance vector and a semantic mask indicating the location of transient objects like pedestrians. The model is evaluated on several datasets of publicly available images spanning a broad range of illumination conditions. We create short videos demonstrating realistic manipulation of the image viewpoint, appearance, and semantic labeling. We also compare results with prior work on scene reconstruction from internet photos.

1. Introduction

The paper targets total scene capture from uncontrolled internet photos, combining explicit 3D scaffolding with neural rerendering to generate realistic views across appearance conditions and editable transient-object semantics.

  • Motivation: Total scene capture records a scene across diverse lighting and transient states while enabling arbitrary viewpoints under those conditions.The target includes crowded, rainy, snowy, sunrise, spotlit, and nighttime appearances.
  • Motivation: Community photos provide abundant samples of scene appearances over many years, despite the challenge of uncontrolled variation.Alternative sources trade away viewpoint diversity or appearance diversity.
  • Approach: The approach factorizes images into viewpoint, appearance conditions, and transient objects, using a dense but noisy point cloud as explicit 3D scaffolding.This separates scene structure from appearance and transient-object handling.
  • Approach: A deferred-shading framebuffer containing albedo, depth, and other attributes is pixelwise aligned with each photo to train multimodal image translation.The network learns to rerender an approximate initial scene rendering as a realistic image under different appearances.
  • Training: Staged appearance training pretrains an appearance encoder, trains rerendering with fixed embeddings, and then jointly fine-tunes both networks.The authors report that this strategy better captures scene appearance while supporting simpler networks on large datasets.
  • Transient objects: Semantic-mask conditioning helps ignore transient objects, discard poorly reconstructed thin features, and render scenes without people.The mask specifies the desired semantic labeling of the output image.
  • Contributions: The contributions include total scene capture from in-the-wild collections, factorized rerendering, proxy-loss appearance learning, and evaluations on five large datasets.The paper also reports view and appearance interpolation and comparisons with previous methods.

2. Related work

Prior work reconstructs geometry, synthesizes views, models appearance, and translates images, but commonly assumes limited appearance variation or controlled data. This paper combines these directions for multimodal neural rerendering from uncontrolled internet photos.

  • Scene reconstruction: Traditional reconstruction uses structure-from-motion followed by dense reconstruction, but usually assumes one appearance or recovers an average appearance.The paper uses dense MVS point clouds as proxy geometry for neural rerendering.
  • Image-based rendering: Image-based rendering creates new viewpoints by warping pixels with proxy geometry, but generally assumes static scene appearance.That assumption makes it poorly suited to scenes whose appearance varies across images.
  • Neural scene rendering: Neural scene rendering learns latent scene representations for novel views, but prior work cited here is limited to simple synthetic geometry.
  • Appearance modeling: Scene-appearance research addresses changes across time, weather, and years, with internet-photo appearance often having relatively low dimensionality apart from transient outliers.
  • Appearance modeling: Unlike methods requiring direct lighting or geometry supervision, this paper learns an implicit appearance representation directly from the input image distribution.
  • Deep image synthesis: Image-to-image translation has expanded from paired domain translation to unpaired and multimodal settings that preserve content across domains.
  • Neural rerendering: The paper casts in-the-wild neural rerendering as multimodal synthesis, combining latent appearance control with editable output semantics for a given viewpoint.

3. Total scene capture

Total scene capture models a scene across viewpoints, appearances, and transient-object states using an explicit 3D proxy and neural rerendering. The framework combines multimodal appearance modeling, staged training, and semantic conditioning to generate realistic views from internet photos.

  • Problem: Total scene capture seeks a generative model that encodes 3D structure, renders varied lighting and weather, and controls transient objects.The stated target includes arbitrary viewpoints, appearances, and optional pedestrians or cars.
  • 3.1. Neural rerendering framework: The pipeline reconstructs a dense colored point cloud from unstructured internet photos using Structure-from-Motion and Multi-View Stereo.The point cloud serves as an explicit proxy for the scene.
  • 3.2. Appearance modeling: The rerenderer maps pixel-aligned deep-buffer renderings to realistic images while using a latent appearance vector to represent output variations unavailable from static geometry.The deep buffer contains scene attributes such as albedo and depth, and the appearance encoder observes both the image and buffer.
  • 3.2. Appearance modeling: Staged appearance training pretrains the appearance encoder, trains the rerenderer with fixed embeddings, and then jointly fine-tunes both networks.This regime removes several auxiliary consistency and latent-reconstruction losses, retaining direct reconstruction and GAN losses.
  • 3.2. Appearance modeling: The baseline multimodal model struggles with infrequent appearances such as night scenes, motivating staged training to capture more complex appearance variation.The paper attributes this limitation to insufficient expressiveness in the jointly trained appearance encoder.
  • 3.3. Semantic conditioning: Semantic conditioning supplies transient-object labels and helps the network avoid encoding their locations in appearance or viewpoint while recovering poorly reconstructed static features.The labels also help place objects such as lampposts where they are detected and provide semantic context for appearance estimation.

4. Evaluation

The evaluation tests reconstruction quality, appearance modeling, interpolation, transfer, semantic conditioning, and comparison with prior 3D reconstruction. Staged appearance training generally improves results, while artifacts remain in difficult transitions and small-data settings.

  • Evaluation setup: The system is evaluated on five datasets reconstructed from public images, using aligned point-cloud renderings and 100-image validation sets per dataset.Each dataset receives a separately trained model.
  • Quantitative evaluation: Table 1 compares I2I, semantic conditioning, baseline appearance modeling, and staged appearance training using perceptual loss, L1 loss, and PSNR.Perceptual and L1 losses are lower-is-better; PSNR is higher-is-better.
  • Quantitative evaluation: Staged appearance training performs better than the baseline on all but the smallest Sacre Coeur dataset, where it overfits and fails to generalize.The baseline is less prone to overfitting but models appearance less effectively.
  • Appearance interpolation: Appearance interpolation produces more complex changes with staged training, but day-to-night transitions can contain unrealistic twilight artifacts.The baseline produces relatively linear interpolations and struggles with complex sunset and night scenes.
  • Appearance transfer: Appearance transfer renders multiple viewpoints under appearances from other photos, although fine details can flicker during smooth camera or latent-space interpolation.The demonstrated transfer includes sunny highlights and spotlit night illumination on statues.
  • Semantic consistency: Predicted semantic masks produce similar building results and remove people, but pedestrians appearing in the mask may be rendered as black, ghostly figures.Semantic conditioning also supports handling transient objects and poorly reconstructed thin features.

5. Discussion

The system is presented as a first attempt at total scene capture using internet photos, but its reliance on segmentation introduces sensitivity to segmentation errors.

  • Segmentation: The rerenderer relies heavily on segmentation masks to synthesize regions absent from the proxy geometry, including ground and sky.
  • Segmentation: Segmentation errors can produce artifacts such as incorrect sky regions and a “ghost pole” in San Marco.
  • Segmentation: Jointly training the neural rerenderer and segmentation network is proposed as a way to reduce such artifacts.
  • The paper presents the system as a first attempt to solve total scene capture from unstructured internet photos.

A. Supplementary Results

Supplementary results show realistic appearance variation, qualitative comparison against prior work, and quantitative evaluation using estimated segmentation masks.

  • Appearance variation: The staged-training method produces realistic renderings across five scenes or viewpoints and four appearances in the San Marco results.
  • Qualitative comparison: The Colosseum evaluation compares the technique with Shan et al. using 20 held-out output-image sets in a user study.
  • Quantitative evaluation: Quantitative metrics are recomputed using estimated validation-set segmentation masks rather than only ground-truth masks.

B. Implementation Details

The supplementary materials describe different network choices for staged and baseline training and include aligned-dataset frames illustrating occlusion handling.

  • Network configurations: The staged-training and baseline modes use different network architectures, with each model receiving its best results from its selected network.
  • Aligned dataset: Aligned-dataset frames show neural rerendering avoiding artifacts caused by interior structures visible through walls in point-cloud renderings.

B.1. Neural rerender network architecture

The rerendering network uses a symmetric encoder-decoder with skip connections for 256 × 256 conditional image translation.

  • The rerendering network is a symmetric encoder-decoder with skip connections operating at 256 × 256 resolution.It uses six downsampling and upsampling blocks, with feature-map concatenation between corresponding encoder and decoder stages.

B.2. Appearance encoder architecture

The appearance encoder uses pixel-wise normalization and an 8-dimensional latent appearance vector integrated into the rendering network.

  • B.2. Appearance encoder architecture: Pixel-norm layers follow each downsampling block in the appearance encoder.The normalization stabilizes training without mixing information between pixels as instance normalization or batch normalization do.
  • B.2. Appearance encoder architecture: The appearance representation is an 8-dimensional latent vector za ∈R8.The vector is injected at the bottleneck between the encoder and decoder in the rendering network.
  • B.2. Appearance encoder architecture: The latent appearance vector is tiled to match the rendering network dimensions.This allows the appearance code to be incorporated at the encoder-decoder bottleneck.

B.3. Baseline architecture

The baseline uses a faithful TensorFlow implementation of the prior encoder-decoder and appearance encoder, adapted for a single-domain supervised setting.

  • B.3. Baseline architecture: The baseline implements the encoder-decoder network and appearance encoder from [25].The implementation follows the released PyTorch code from [25] as a guideline.
  • B.3. Baseline architecture: The implementation uses TensorFlow rather than the released PyTorch code.The PyTorch release serves as an implementation reference, not the stated framework.
  • B.3. Baseline architecture: The training pipeline is adapted to the single-domain supervised setup described in Section 3.2.

B.4. Aligned datasets

The aligned datasets support visualizing rerendered views, comparisons with Shan et al., and the learned appearance latent space. The staged-training embedding forms meaningful clusters after pretraining and reaches quality comparable to BicycleGAN after finetuning.

  • B.4. Aligned datasets: Figure 11 presents sample frames from the aligned datasets generated as described in Section 3.1.
  • B.4. Aligned datasets: Figure 15 visualizes latent appearance spaces after pretraining, after finetuning, and for the BicycleGAN baseline.The pretrained embedding forms meaningful clusters, while the finetuned embedding has higher quality and is comparable to BicycleGAN’s.
  • B.4. Aligned datasets: Figure 12 shows original-image appearances in the first row and rerendered viewpoints under those appearances, with rendered point-cloud images as inputs in the first column.
  • B.4. Aligned datasets: Figures 13 and 14 compare Shan et al. [39] in the first and third columns with the authors’ results in the second and fourth columns.
Loading 1904.04290v1…