Source-linked AI summary

Albedo Estimation via Latent Bridge Matching

Carme Corbi, David Serrano-Lozano, Javier Vazquez-Corral, Maria Vanrell

arXiv:2609.09884v1cs.CVcs.AI

TL;DR

IID methods struggle with physical consistency, inference cost, and generalization. This paper applies LBM to albedo estimation with source-anchored transport, reconstruction loss, and shading conditioning, reporting competitive performance across five datasets with substantially lower computational cost.

  • Problem

    Generative IID methods remain limited by insufficient physical consistency, high inference cost, residual noise, and weak generalization across datasets.

  • Method

    The paper uses LBM for source-to-albedo latent transport, adds pixel reconstruction loss for physical consistency, and incorporates shading conditioning.

  • Results

    LBM achieves up to 75% lower inference time, improves accuracy by 3 to 6 PSNR units, and reduces reconstruction error to 50% compared with other generative approaches.

  • Takeaways & Limitations

    Albedo and shading provide mutually informative conditioning signals, while the single-step model is competitive with substantially more expensive multi-step diffusion pipelines across five datasets.

  • Takeaways & Limitations

    Synthetic training datasets introduce domain gaps, supervision ambiguity, corrupted scenes, and color-material biases that constrain real-world generalization.

Abstract

from arXiv · show

Recent advances in Intrinsic Image Decomposition (IID) have increasingly relied on generative models. However, progress remains limited by three key challenges: (a) insufficient physical consistency, (b) high computational cost at inference time, and (c) limited generalization capabilities. In this work, we show that latent bridge matching (LBM) effectively addresses these limitations for albedo estimation. We introduce a novel LBM-based architecture that enforces physical consistency through a pixel reconstruction loss, benefits from the inherent efficiency of LBM low-cost inference, and improves generalization across diverse datasets by incorporating a shading conditioning. In this extended version, we additionally show that conditioning the shading estimator itself on the predicted albedo further improves reconstruction fidelity, and we benchmark our best model against stateof-the-art IID methods across five real and synthetic datasets.

Introduction

Intrinsic Image Decomposition separates an image into albedo and shading, but remains ill-posed and difficult to solve physically, efficiently, and robustly. The proposed LBM approach uses pixel reconstruction, conditioning, and low-cost inference to address these challenges.

  • IID decomposes a single RGB image into albedo, representing surface color, and shading, representing illumination and geometry effects.
  • The basic Lambertian model is frequently violated, so real-world images often require residual effects beyond albedo and shading.
  • IID is severely ill-posed because infinitely many albedo-shading pairs or albedo-shading-residual triplets can explain the same image.
  • Generative IID methods face weak physical consistency, residual noise, high inference cost, and limited generalization linked to synthetic training data.
  • LBM uses pixel-wise reconstruction loss and conditioning inputs to incorporate image-based physical constraints into IID.
  • LBM reduces inference time by up to 75%, improves accuracy by 3 to 6 PSNR units, and reduces reconstruction error to 50%.

Related works

IID research spans classical priors, learning-based models, and generative approaches, while benchmark datasets trade off realism, scale, and annotation quality. Different training data and evaluation metrics make fair comparison difficult.

  • Classical IID methods use Retinex assumptions, hand-crafted optimization priors, or learning-based encoder-decoder models.
  • Optimization-based methods require no labeled training data but rely on assumptions often violated by complex, non-Lambertian scenes.
  • Generative IID methods include GAN-based layer separation and diffusion models that leverage pretrained generative priors.
  • Synthetic datasets provide dense annotations at scale, whereas controlled real-world and sparse-supervision datasets offer narrower or less complete coverage.
  • Different dataset combinations and dataset-specific metrics prevent a common evaluation protocol and make fair comparison difficult.

Background

Diffusion and Flow Matching generate outputs through trajectories originating from Gaussian noise, limiting pixel anchoring and efficient physical constraints. LBM instead transports source images directly to albedo targets, enabling fast inference and reconstruction-based physical grounding.

  • Diffusion reverses a gradual noising process, while Flow Matching learns a continuous velocity field for transport with fewer sampling iterations.
  • Both diffusion and Flow Matching begin image-to-image trajectories from a fixed Gaussian prior, so outputs are not anchored to observed pixels and may retain noise variability.
  • LBM builds a stochastic interpolant between natural-image and albedo distributions, learning transport between arbitrary source and target distributions.
  • LBM predicts the target from the source in as little as one step and operates in the latent space of a pretrained autoencoder.
  • Because LBM begins at the source image, its decoded prediction remains pixel-anchored and supports an explicit reconstruction loss enforcing the image formation model.
  • The paper treats diffusion, Flow Matching, and LBM as a hierarchy and reuses the same UNet backbone and training pipeline across variants.

Method

The method formulates albedo estimation as latent image-to-image transport with Latent Bridge Matching, pixel-level reconstruction supervision, and optional intrinsic-component conditioning. It uses a four-step latent bridge process with a frozen VAE and Stable Diffusion XL UNet to combine efficiency with physically grounded decomposition.

  • Albedo estimation: Albedo estimation is posed as translation from an observed RGB image to illumination-independent ground-truth albedo in a compressed VAE latent space.The source and target are encoded as z0 = E(x0) and z1 = E(x1).
  • Latent Bridge Matching formulation: LBM constructs a stochastic latent bridge between source and target representations, then predicts the velocity toward the clean target latent with a UNet.The predicted target is recovered in closed form, and the network is trained with a bridge-matching regression objective.
  • Latent Bridge Matching formulation: Sampling from a small set of equally spaced timesteps aligns training and inference and caps the latent transport at four steps.This timestep design is identified as the primary source of LBM's efficiency over diffusion training.
  • Reconstruction Loss: The latent objective is augmented with pixel-space supervision and an image-formation reconstruction term to enforce consistency after decoding.The reconstruction term compares the input image with an image re-synthesized from the estimated albedo, while Lpixel uses LPIPS.
  • Conditioning Albedo Estimation: Conditional LBM concatenates an encoded auxiliary cue with the interpolated latent, and the method explores surface normals and shading as conditioning signals.Shading provides an illumination cue that helps recover residual reflectance and reduces incorporation of soft shadows and inter-reflections into albedo.

Experiments and Results

Experiments evaluate LBM-based albedo estimation across five benchmarks, conditioning strategies, and comparisons with prior methods. The results show strong efficiency, reconstruction consistency, and competitive quality, while shading-conditioned and bidirectionally conditioned variants improve albedo estimates.

  • Baseline evaluation: LBM-AID achieves strong albedo-estimation performance while reducing inference time relative to Stable Diffusion-based generation.The study evaluates MIT, IIW, and Hypersim, with additional comparisons on ARAP and InteriorVerse.
  • Conditioning ablation: Surface-normal conditioning provides no consistent benefit, whereas shading conditioning is the strongest tested strategy for albedo estimation.The comparison includes foundation-model normals, LBM-based normal estimation, and LBM-based shading estimation.
  • Bidirectional conditioning: Albedo-conditioned shading improves the auxiliary shading estimate and produces the lowest reconstruction error on MIT, ARAP, and Hypersim.The resulting errors are 0.0104 on MIT, 0.0139 on ARAP, and 0.0250 on Hypersim.
  • Physical consistency: The reconstruction loss and shading conditioning both improve physical consistency with the image-formation model.The reconstruction loss is especially aligned with pixel-level fidelity in albedo estimation, while shading conditioning supplies complementary information.
  • Comparison with prior methods: Across state-of-the-art comparisons, the method suppresses shadows and reflections and preserves surface texture, but remains weaker on some benchmarks and materials.It performs best on LMSE on MIT, trails IntrinsicDiffusion on ARAP, and falls short of reinforcement-learning-guided variants on some datasets; transparent and metallic regions remain difficult.
  • Comparison with prior methods: The method is competitive with expensive multi-step diffusion pipelines despite training exclusively on synthetic data, with remaining gaps attributed to the lack of real-world supervision.The comparison includes methods fine-tuned with reinforcement learning on real-world feedback.

Limitations and Future Work

The paper identifies constraints from model capacity, synthetic-data domain gaps, and fragmented evaluation, while pointing to improved supervision and broader conditioning as future directions.

  • Model capacity and efficiency: The 2.5B-parameter backbone requires substantial memory, restricts training to 256 × 256 patches, and degrades beyond roughly 2K inference resolution.LoRA fine-tuning did not converge satisfactorily; Diffusion Transformer backbones are suggested for scalability.
  • Qualitative failure cases: Qualitative InteriorVerse comparisons show hallucinated reflections, near-zero reflectance collapse, and color shifts among compared methods and the proposed method.The highlighted failures involve reflective, transparent, and color-sensitive regions.
  • Dependence on synthetic training data: Synthetic training data constrain real-world generalization through ambiguous Hypersim supervision, corrupted scenes, and InteriorVerse color-material bias.The bias likely explains observed color shifts, while ambiguous near-zero albedo labels affect dark opaque and non-diffuse materials.
  • Evaluation inconsistencies across the literature: IID comparisons remain difficult because methods use heterogeneous dataset-specific metrics and few report all five benchmarks.The paper motivates standardized, multi-dataset evaluation protocols.
  • Future directions: Future work includes reducing reliance on dense annotations through perceptual-feedback reinforcement learning and extending conditioning to specular reflections and interreflections.ReasonX is presented as a model-agnostic extension for the LBM pipeline.
  • Qualitative failure cases: Transparent regions remain a main failure case on Hypersim, while featureless estimates occur in RGB↔X low-light regions.The Hypersim comparison also reports preserved surface color and structural detail for the proposed method.

Conclusions

The paper presents LBM as a physically coherent, conditioning-compatible framework for albedo estimation and evaluates a single-step model across five datasets. It remains competitive with more expensive diffusion pipelines, while real-world feedback and synthetic-data dependence remain open issues.

  • Conclusions: The proposed LBM model enforces the image formation model and uses mutually informative albedo-shading conditioning, whereas surface normals do not improve conditioning.The model is benchmarked against state-of-the-art IID methods on five datasets.
  • Conclusions: The single-step model is competitive with substantially more expensive multi-step diffusion pipelines but trails methods fine-tuned on real-world feedback.This conclusion is supported by the five-dataset benchmark summary.
  • Conclusions: IID progress appears bottlenecked by inconsistent cross-dataset performance, fragmented metrics, and dependence on large-scale synthetic training data.The paper identifies improved real-world generalization as a key future direction.
Loading 2609.09884v1…