Source-linked AI summary

On the Transfer of Inductive Bias from Simulation to the Real World: a New Disentanglement Dataset

Muhammad Waleed Gondal, Manuel Wüthrich, Đorđe Miladinović, Francesco Locatello, Martin Breidt, Valentin Volchkov, Joel Akpo, Olivier Bachem, Bernhard Schölkopf, Stefan Bauer

arXiv:1906.03292v3stat.MLcs.LG

TL;DR

Disentanglement methods have largely been developed and evaluated on synthetic toy data because real-world data are costly and difficult to control. This paper introduces matched real and simulated datasets built around controlled physical objects, then evaluates transfer between them. Learned representations transfer poorly, whereas selecting models and hyperparameters in simulation transfers useful information to real-world data.

  • Problem

    Performance of state-of-the-art disentanglement learning on real-world data is unknown because real-world data are costly, confounded, and difficult to control.

  • Method

    The paper builds a controlled real-world 3D dataset with seven variation factors and two matched simulated datasets with different realism levels, then evaluates disentanglement transfer.

  • Results

    Learned representations transfer poorly from simulated to real data, while simulation-based model selection outperforms random selection 72% of the time for realistic renderings and 78% for simpler synthetic images.

  • Takeaways & Limitations

    Model and hyperparameter selection can transfer information from simulation to the real world even when direct representation transfer performs poorly.

  • Takeaways & Limitations

    The framework assumes generative factors are independent, although some factors may require a hierarchical structure to satisfy that assumption.

Abstract

from arXiv · show

Learning meaningful and compact representations with disentangled semantic aspects is considered to be of key importance in representation learning. Since real-world data is notoriously costly to collect, many recent state-of-the-art disentanglement models have heavily relied on synthetic toy data-sets. In this paper, we propose a novel data-set which consists of over one million images of physical 3D objects with seven factors of variation, such as object color, shape, size and position. In order to be able to control all the factors of variation precisely, we built an experimental platform where the objects are being moved by a robotic arm. In addition, we provide two more datasets which consist of simulations of the experimental setup. These datasets provide for the first time the possibility to systematically investigate how well different disentanglement methods perform on real data in comparison to simulation, and how simulated data can be leveraged to build better representations of the real world. We provide a first experimental study of these questions and our results indicate that learned models transfer poorly, but that model and hyperparameter selection is an effective means of transferring information to the real world.

1 Introduction

Disentanglement research has advanced largely on synthetic toy datasets because they are controllable and inexpensive, while performance on real-world data remains unknown. The paper introduces controlled real and simulated datasets to study transfer and inductive bias.

  • Disentangled representations aim to encode low-dimensional generative factors such as shape, size, and color in compact latent variables.
  • Synthetic datasets are widely used because they are cheap to generate and allow precise control of independent generative factors.
  • Real-world recordings contain camera and surface-property imperfections, leaving state-of-the-art disentanglement performance on real data unknown.
  • The recording platform targets transfer from rendered images to physical recordings, simulation realism, supervision, image quality, confounding, and causal-mechanism disentanglement.
  • The paper introduces a controlled real-world 3D dataset with 7 variation factors and two simulated datasets with different realism levels.
  • The datasets support systematic investigation of how model and hyperparameter choices transfer between simulations and real-world data.

2 Background and Related Work

Disentanglement methods seek latent variables corresponding to independent generative factors, but existing approaches and evaluations rely heavily on synthetic datasets. The background reviews the causal-factor formulation, VAE-based methods, and the need for controlled real-world data.

  • The framework models observations X as generated by K unobserved, independently changeable causes G, with learned features Z intended to capture individual factors.
  • Disentanglement commonly means that each learned feature in Z captures one factor of variation in G.
  • State-of-the-art approaches commonly use VAEs, whose neural generative model and proxy posterior are optimized through the variational lower bound.
  • Because the basic VAE objective does not enforce latent structure beyond similarity to an isotropic Gaussian prior, later methods add supervision or regularization.
  • Real-world data are costly and often lack ground truth, motivating reliance on synthetic toy datasets whose real-world validity requires investigation.

3 Bridging the Gap Between Simulation and the Real World: A Novel Dataset

The paper introduces a controlled platform and matched real and simulated datasets for studying disentanglement and simulation-to-reality transfer. The setup varies independent object, camera, background, and robotic-motion factors while keeping the three datasets aligned.

  • A controlled dataset is needed because existing real-world recordings do not support quantitative investigation of inductive biases, sample complexity, and simulation–reality interplay.
  • The platform uses three cameras, a robotic manipulator, a rotating table, shielding, and internal lighting to control physical 3D-object recordings.
  • Factors of Variation: The recording factors include object color, background color, two object sizes, camera height, and two independent robotic-arm rotations.
  • Factors of Variation: The dataset varies colors, shapes, sizes, background color, and camera height while holding the other factors constant in the corresponding comparisons.
  • Two simulated datasets use the same factors as the real dataset, with one designed for realism and the other deliberately simplified.
  • The synthetic data uses CAD-based scene components, tuned materials and camera poses, and computer-graphics rendering to approximate the physical setup.

4 First Experimental Evaluations of (unsupervised) Disentanglement Methods on Real-World Data

Experiments compare six disentanglement methods across real, realistic-synthetic, and toy-synthetic data. Direct representation transfer to real data performs poorly, whereas simulation-based model selection transfers useful information.

  • Six established methods are evaluated across the three datasets using 64x64 images and the same scores as a prior large-scale study.
  • Reconstruction Across Datasets: Real data has the lowest reconstruction score, followed by realistic simulation and then toy simulation, while method rankings remain similar across datasets.
  • Direct Transfer of Representations: MIG performance is higher when models are trained and evaluated on the same dataset, whereas direct transfer from simulation to real data works poorly.The authors caution that high variance makes conclusive statements difficult.
  • Transfer of Hyperparameters: The study tests whether model and hyperparameter choices can transfer as an inductive bias instead of transferring learned representations directly.
  • Transfer of Hyperparameters: 72% of the time, selection from realistic renderings beats random selection, while toy-to-real selection beats random selection 78% of the time.
  • Transfer of Hyperparameters: Cross-dataset rank correlations show that model performance, including hyperparameters, is highly correlated across datasets, with similar patterns across most disentanglement metrics.

5 Conclusions

The work establishes real-world evaluation of disentanglement methods and provides matched rendered datasets for studying transfer, inductive biases, and related representation-learning problems. Its dataset design also supports future extensions to more complex objects, textures, and dependent factors.

  • The dataset complements prior efforts by enabling real-world studies of inductive biases, sample complexity, transfer learning, and label use.
  • Matched rendered and real recordings support investigation of how disentangled representations transfer and how transferability depends on simulation realism.
  • The experimental setup can also support 3D reconstruction, scene rendering, and learning compositional visual concepts.
  • Planned extensions include more complicated shapes and textures, harder conditions, and dependence among factors.

A Platform

The platform is documented through views of the recording setup and the mechanical apparatus used to collect the real-world dataset.

  • The recording setup is illustrated with rendered images of the experimental environment.
  • Together, the figures document both the setup view and the physical recording platform.
  • The mechanical platform is shown as the apparatus for recording the real-world dataset.

A.1 Difference between Realistic Simulations and Real-World Images

The realistic simulated and real-world images closely overlap because the recording procedure and CAD-based construction reproduce the physical setup. Differences are confined mainly to fine visual details and resolution.

  • The realistic simulation closely matches the real-world image because the procedure and CAD files reproduce the robotic arm and objects.The remaining differences include small details, lighting shadings, and a floor crack visible only in the real image.
  • Both simulated and real images have a resolution of 512 × 512.

B A Discussion on Disentangled Representations and their Transfer

Disentanglement is imperfect and varies across factors, with models favoring visually prominent environmental changes. Transfers from simple simulation fail, whereas realistic simulation preserves several environmental factors but not object properties.

  • Models disentangle camera height, background color, object size, and robotic-arm motions better than some object shapes and colors.Pyramid, cone, olive, and brown factors are particularly difficult, possibly because they produce less pixel variation.
  • Reconstructions become blurrier as the data move from simple simulation to realistic simulation and then to real-world images.The paper links this degradation to overregularization and higher KL-divergence weighting in VAE models.
  • Models trained on simple simulation completely fail to transfer representations to real-world data.
  • Realistic-simulation transfer usually reconstructs background color, manipulator pose, and camera height, but differs substantially in object properties.In complex environments, models focus more on environmental factors than on smaller object changes.

C Details of the Experimental Protocol

The experimental protocol evaluates disentanglement using ground-truth factors, predictive and interventional metrics, and latent traversals. Models share a fixed architecture and latent size, while training hyperparameters are fixed within the considered methods.

  • Disentanglement metrics use known generative factors and assess mutual information, predictive feature importance, or interventional robustness.
  • Latent traversals provide a visual validation method when the generative factors are not known.
  • All models use the same convolutional encoder-decoder architecture with a fixed latent size of 10.
  • Training hyperparameters are kept fixed for each considered method.

D Detailed Experimental Results

The experiments compare six disentanglement methods across synthetic and real datasets using multiple metrics and train-test configurations. Metric correlations are generally positive, and hyperparameter rankings transfer across datasets for several metrics.

  • Six methods are compared across five training-and-evaluation configurations spanning realistic synthetic, toy synthetic, and real data.The methods are β-VAE, FactorVAE, β-TCVAE, DIP-VAE-I, DIP-VAE-II, and AnnealedVAE.
  • All metrics except Modularity are at least mildly correlated across datasets.
  • Disentanglement scores on real and simulated data show positive correlation for at least three metrics.
  • Good hyperparameters transfer well between datasets according to at least three metrics.
Loading 1906.03292v3…