Source-linked AI summary

One Noise to Rule Them All: Learning a Unified Model of Spatially-Varying Noise Patterns

Arman Maesumi, Dylan Hu, Krishi Saripalli, Vladimir G. Kim, Matthew Fisher, Sören Pirk, Daniel Ritchie

arXiv:2404.16292v1cs.GRcs.CVcs.LG

TL;DR

Procedural material design relies on distinct noise types, but designers may need patterns between those types and inverse design can require searching over unknown discrete choices. This paper trains a controllable diffusion model to blend noise, including spatial variation absent from training data, and reports higher-fidelity material reconstructions without pre-specifying noise types.

  • Problem

    Designers may need noise characteristics between available types, while unknown noise-node types can make inverse material search combinatorial and limit gradient-based optimization.

  • Method

    A denoising diffusion model with spatially varying conditioning learns a continuous, controllable space of noise patterns from data without spatially varying observations.

  • Results

    The model supports spatially varying blends despite lacking spatially varying training data, and its use as a material-graph noise node can yield higher-fidelity reconstructions without a pre-specified noise type.

  • Takeaways & Limitations

    The learned model provides a single controllable noise generator for blends between noise types and a proof-of-concept use in inverse procedural material design.

  • Takeaways & Limitations

    Deterministic pattern generators are an unnatural fit for diffusion generation, so including them would require changes to the model or training procedure.

Abstract

from arXiv · show

Procedural noise is a fundamental component of computer graphics pipelines, offering a flexible way to generate textures that exhibit "natural" random variation. Many different types of noise exist, each produced by a separate algorithm. In this paper, we present a single generative model which can learn to generate multiple types of noise as well as blend between them. In addition, it is capable of producing spatially-varying noise blends despite not having access to such data for training. These features are enabled by training a denoising diffusion model using a novel combination of data augmentation and network conditioning techniques. Like procedural noise generators, the model's behavior is controllable via interpretable parameters and a source of randomness. We use our model to produce a variety of visually compelling noise textures. We also present an application of our model to improving inverse procedural material design; using our model in place of fixed-type noise nodes in a procedural material graph results in higher-fidelity material reconstructions without needing to know the type of noise in advance.

1 INTRODUCTION

The paper replaces discrete noise-type choices with a controllable model that blends noise patterns, including spatially varying blends learned without spatially varying training data. It applies the model to inverse material design and reports higher-fidelity reconstructions without pre-specifying noise types.

  • 1 INTRODUCTION: Discrete noise choices constrain design and inverse material search, while alpha-blending can leave overlapping features and transitions that do not sensibly interpolate.Unknown node types can make the inverse-design search space combinatorial and limit gradient-based optimization.
  • 1 INTRODUCTION: A denoising diffusion model with spatially varying conditioning learns continuous noise blends from data without spatially varying observations.The model is controlled through interpretable parameters or a random seed, and can generate images beyond training size and seamlessly tileable images.
  • 1 INTRODUCTION: The evaluation includes noise blends, spatially varying patterns driven by image masks, and a painting interface for authoring spatial variation.Videos of the interface are included in the supplemental material.
  • 1 INTRODUCTION: Using the model as a material-graph noise node can produce higher-fidelity reconstructions without knowing that node’s noise type in advance.The introduction presents this as an application to inverse procedural material design.
  • 1 INTRODUCTION: A CutMix-based training scheme supports spatially varying noise generation without spatially varying training data.The introduction identifies this scheme as a contribution alongside the generative model and material-design application.

2 RELATED WORK

Related work spans procedural noise, neural and non-parametric texture synthesis, and inverse procedural material design. The paper distinguishes its unified stochastic noise model from prior texture and pattern approaches.

  • 2 RELATED WORK: Procedural noise functions algorithmically synthesize patterns that mimic natural randomness, with established examples including Perlin, Wavelet, and Gabor noise.Such patterns support applications including materials, terrain, and motion.
  • 2 RELATED WORK: Neural parametric texture synthesis includes exemplar-driven methods and a GAN that learns a continuous texture space, but that GAN entangles style with spatial randomness and lacks interpretable noise controls.The authors report that their initial GAN experiments struggled with the many modes represented by distinct noise types.
  • 2 RELATED WORK: Non-parametric texture methods include morphable texture spaces and patch-based in-filling, rather than learned parametric generation.The cited examples define interpolatable textures with a similarity-induced simplicial complex and fill regions using a regularized screened Poisson equation.
  • 2 RELATED WORK: Inverse material-design research predicts graph parameters, optimizes continuous parameters, or synthesizes graphs, while learned proxies can omit stochastic noise and require separate models for separate patterns.The unified noise DDPM instead captures multiple non-deterministic noise functions in one continuous space.

3 METHOD

The method learns a conditional generative model from noise samples with globally uniform properties, then uses spatially varying conditioning to synthesize smooth blends at inference. CutMix training helps the model respond to localized conditioning even though the training data lacks spatially varying samples.

  • 3 METHOD: The model learns noise textures conditioned on noise type and parameters, using a DDPM designed to support spatially varying noise despite training only on globally uniform samples.The method treats noise type and parameters as conditioning inputs and uses CutMix to regulate the model's response to granular signals.
  • 3.1 Spatially-varying conditioning: An MLP maps class and parameter vectors into a feature grid that conditions U-Net GroupNorm through SPADE; interpolating grid features enables smooth spatial blends at inference.The training grid is spatially uniform, while inference can use artificially constructed grids to interpolate between noise types and parameters.
  • 3.1 Spatially-varying conditioning: Spherical regularization penalizes class-embedding norm deviations and improves texture blending between classes, but this post-publication enhancement is absent from the primary results.The target norm is defined from the expected squared norm of a d-dimensional Gaussian vector.
  • 3.2 Enhancing localized conditioning: The U-Net's wide receptive field makes localized conditioning changes affect much of the canvas, so CutMix trains it to respond to spatially localized signals.CutMix combines noise textures with their corresponding conditioning grids, enriching training examples while teaching localized responses.
  • 3.2 Enhancing localized conditioning: CutMix is applied to half of training samples, mixing a base texture with one to four randomly cropped patches from unique noise types.Only the crop mask is rotated; the texture itself is not.

4 IMPLEMENTATION DETAILS

Implementation uses a large dataset of parameterized noise images to train a conditioned U-Net diffusion model. The paper also compares visual blending quality and reports inference speed and a tileable-noise modification.

  • 4 IMPLEMENTATION DETAILS: The dataset contains about 1.2 million images from 18 noise functions, with 16,384 parameter sets and four seeds per type at 512 × 512 pixels.The dataset contains no spatially varying samples, as illustrated by the sampled noise types and parameters.
  • 4 IMPLEMENTATION DETAILS: The conditioned U-Net has about 5.1 million parameters and is trained with a noise-prediction objective using AdamW for about 300,000 steps.Training uses batches of 128, 8 NVIDIA RTX 3090 GPUs, and images downsampled to 256 × 256 pixels.
  • 4 IMPLEMENTATION DETAILS: Our method blends smoothly and synthesizes novel details, while PSGAN and Image Melding exhibit artifacts and repeated details in the qualitative comparison.Image Melding is given the first and last image quarters and fills the interior; PSGAN also produces anisotropic features unlike the data distribution.
  • 4 IMPLEMENTATION DETAILS: At 256 × 256 resolution, inference reaches 80 diffusion steps per second on one NVIDIA RTX 3090 GPU; circular padding makes noise maps tileable.Inference speed scales quadratically with resolution.

5 RESULTS AND EVALUATION

The model generates spatially varying and interpolated noise, and performs better than PSGAN on the reported FID comparison. It also supports inverse material design, tileable textures, and synthesis at larger canvas sizes.

  • Spatially and temporally varying noise: Blending maps produce spatially varying noise, and interpolating conditioning maps or diffusion noise yields continuously varying textures.Figure 7 shows examples from blended feature grids; Figure 11 interpolates conditioning parameters horizontally and noise seeds vertically.
  • Quantitative evaluation: FID mean 20.9 and median 13.1 for our method, versus PSGAN’s 99.2 and 87.5, respectively.Lower is better; the evaluation samples 20,000 images per noise type.
  • Inverse material graph design: In material graph optimization, replacing one noise node with the model provides a prior that helps recover more accurate target-photo reconstructions.Figure 8 compares results using MATch’s feature-based similarity and LPIPS scores; Figure 9 shows edits to optimized graphs.
  • Tileability and texture size agnosticism: The model produces tileable noise at arbitrary sizes, supporting large canvases without excessive repeated visual patterns.The paper shows a 2048 × 2048 Damascus pattern generated in one diffusion process.
  • CutMix augmentation: Removing CutMix severely weakens local conditioning responses; using one or four patches yields similar qualitative results, with four sometimes interpolating more smoothly.The authors note that the comparison is qualitative because meaningful ground-truth data is unavailable.

6 CONCLUSION

The model learns controllable blends across procedural noise types and can generate spatially varying blends without such training examples. The conclusion notes limits in blend quality and mode coverage, and frames inverse material design as an early proof of concept.

  • 6 CONCLUSION: The model learns a continuous space of controllable noise patterns, including spatially varying blends without spatially varying training data, and demonstrates a proof-of-concept use in inverse material design.The paper also reports visually compelling textures and higher-fidelity material reconstruction using the model as a noise generator.
  • 6 CONCLUSION: Noise pairs with sufficiently dissimilar geometric features can produce blurry or visually awkward intermediate blend regions.The figure reports artifacts in transition regions for some noise-type blends.
  • 6 CONCLUSION: Without CutMix, the U-Net fails on non-uniform noise maps, including a class blend where galvanic noise disappears; uniform outputs remain similar.The uniform, no-blending panel serves as a control, and its outputs are similar across conditions.
  • 6 CONCLUSION: Some low-density modes of noise distributions are poorly captured, so distributions with such modes may need more careful parameter-space sampling.One suggested approach is to sample some parameter regions more frequently.
  • 6 CONCLUSION: Deterministic pattern generators are an unnatural fit for denoising diffusion, so including them in the learned blend space would require model or training changes.
  • 6 CONCLUSION: A proposed direction for interactive spatially varying textures is to place noise droplets that spread and interact, potentially through a diffusion equation.
  • 6 CONCLUSION: The inverse material-design application remains an early proof of concept; one-shot diffusion could make replacing every noise node tractable.With differentiable graph operations, the paper suggests continuous optimization of an over-complete graph with edge weights and a sparsity prior.

A.1 Noise dataset details

The appendix describes the sampled noise types and parameter handling, along with offset-noise and random-output details.

  • A.1 Noise dataset details: The dataset table enumerates sampled noise types, parameters, and parameter ranges, excluding parameters that only act as color correction.Cells 4, cells 1, and voronoi are identified as Worley-noise variants.
  • A.1 Noise dataset details: The conditioning vector treats identically named parameters as separate entries, except scale, and independently normalizes entries to [0, 1].
  • A.1 Noise dataset details: Without offset noise, the network sometimes fails to synthesize images with extremely dark or bright intensity distributions.The passage describes representative examples of this failure.
  • A.1 Noise dataset details: The appendix shows output samples with all listed noise parameters sampled randomly.Examples appear in Figures 19 to 23.

A.2 PSGAN Baseline

The PSGAN baseline is modified with conditioning and training changes to support the comparison in Figure 6.

  • A.2 PSGAN Baseline: For Figure 6, the PSGAN architecture is equipped with SPADE conditioning blocks identical to those used in the proposed method.
  • A.2 PSGAN Baseline: The baseline uses a slightly larger discriminator to compensate for added generator parameters and retains most other implementation and training components.
  • A.2 PSGAN Baseline: The PSGAN baseline uses WGAN loss, and its generator has 30 million parameters.

A.3 Inverse Material Design Details

Inverse material optimization exposes class, noise-parameter, and diffusion-noise inputs, with regularization and a scheduled temperature to guide class selection.

  • A.3 Inverse Material Design Details: The optimizer exposes a soft-class vector, the parameter vector, and the diffusion Gaussian noise image; the soft-class vector forms a convex combination of class embeddings.
  • A.3 Inverse Material Design Details: L1 regularization with weight 0.1 is applied to both the soft-class vector and the parameter vector.
  • A.3 Inverse Material Design Details: The softmax temperature starts at 0.25, is multiplied by 0.97 each optimization step, and is clamped at 0.01 to encourage selection of a single class.At the minimum temperature, the optimizer is warm-restarted and noise is added to the exposed parameters to avoid local minima.

B TRAINING WITH ADDITIONAL NOISE TYPES

The model is trained on Phasor and Gabor noise using the same sampling method as Section 4, with examples of parameter and class interpolation. Its FID is notably higher for these noise types, reflecting difficulty capturing sparse parameter configurations.

  • B TRAINING WITH ADDITIONAL NOISE TYPES: Phasor and Gabor training data use the same sampling method as Section 4, and examples show parameter and class interpolation for both types.The released implementations were used to accrue the training data; sampled parameters and ranges are listed in Table 3.
  • B TRAINING WITH ADDITIONAL NOISE TYPES: FID was 93.6 for Phasor and 129.3 for Gabor noise, notably higher than the scores for other noise types.Some parameter configurations are not well captured because narrow regions of the parameter space contribute fewer training samples.
  • B TRAINING WITH ADDITIONAL NOISE TYPES: More frequent sampling of low-density parameter regions may be needed when the desired noise distribution contains such modes.The passage attributes weaker performance in these regions to diffusion models representing low-density data regions less accurately.

C MODEL ARCHITECTURES

The paper compares three U-Net architectures and uses the lightweight Model-XS for its main-text figures and inverse material design. Additional figures show interpolations and random samples across noise types.

  • C MODEL ARCHITECTURES: Model-XS has about 5.1 million parameters, with two downsampling blocks, a bottleneck, two upsampling blocks, and SPADE conditioning by 128-dimensional noise embeddings.Each block contains two ResNet sub-blocks, with one additional ResNet sub-block at the network end.
  • C MODEL ARCHITECTURES: Figures 17–18 show parameter and class interpolations for Phasor and Gabor noise.The examples include transitions within a noise type and between different types.
  • C MODEL ARCHITECTURES: The primary model is significantly faster than the larger architectures, but compromises slightly on FID.Table 4 summarizes the architectures, parameter counts, design choices, FID scores, and inference performance.
  • C MODEL ARCHITECTURES: Model-S adds linear attention layers and another downsampling/upsampling set to Model-XS, while Model-M further increases channel dimensions.The architectures were trained for the same number of optimization steps.
  • C MODEL ARCHITECTURES: Model-XS was chosen for inverse material design because attention layers add memory cost; larger models may suit applications without that lightweight constraint..
  • C MODEL ARCHITECTURES: Figures 19–24 show random 256 × 256 samples from the model: cells 4 and 1, Voronoi, microscope view, bnw spots1, liquid, several grunge types, messy fibers 3, Perlin, Gaussian, and clouds 1–3.
Loading 2404.16292v1…