Source-linked AI summary

Beyond Uniform Local Isometry and Topology: FactoMap for Disentangled Representations

Sohini Gupta, Bahareh Tolooshams

arXiv:2608.24762v1cs.LGcs.AI

TL;DR

Disentanglement methods often assume Euclidean factor coordinates even when factors wrap, collapse, or have position-dependent geometry, motivating a richer factor-space description. The paper formalizes these structures and introduces FactoMap, which learns prototypes on a matching factor-space lattice. Experiments show that matched structure preserves factor continuity and enables disentanglement of underlying factors, while hue and scale exhibit anisotropy that fixed rescaling cannot remove.

  • Problem

    Unlabelled observations do not determine disentangled factors without assumptions, and statistical independence does not ensure geometric separability when factor effects vary with position.

  • Method

    FactoMap combines factor domains, generator-induced identifications, and position-dependent scales in a lattice-indexed topographic prototype representation.

  • Results

    Matching factor-space structure preserves factor continuity and enables disentanglement, including periodic and non-uniform factor organization.

  • Takeaways & Limitations

    Disentangled representations can require lattice geometry that matches factor periodicity, identifications, and local scales rather than fixed Euclidean coordinates.

Abstract

from arXiv · show

Many disentanglement methods represent generative factors using Euclidean product coordinates, although the underlying factor spaces may wrap, collapse, or have position-dependent geometry. We introduce factor-space structure, combining factor domains, generator-induced identifications, and position-dependent scales to distinguish topologically equivalent spaces with different factor geometries. We show that statistically independent factors need not be geometrically separable: hue and scale produce effects that grow at different rates, yielding anisotropy that no fixed rescaling removes. We propose the Factor-Space Topographic Map (FactoMap), which learns interpretable prototypes indexed by a factor-space lattice. Topographic learning transfers the lattice's periodicity, collapses, and non-uniform extent to the representation. Experiments show that matching this structure preserves factor continuity and enables disentanglement of the underlying factors.

1 Introduction

The paper argues that disentangled representations require assumptions about factor geometry beyond statistical independence, topology, and uniform local isometry. It introduces factor-space structure and FactoMap to represent generator-induced identifications and position-dependent scales.

  • Motivation: Disentangled representations aim to recover generative factors as independently interpretable coordinates, but unlabelled observations cannot identify them without additional assumptions.Such coordinates are motivated by interpretation, generalization, fairness, and control.
  • Motivation: Existing methods impose statistical or geometric inductive biases, including local isometry that treats equal factor changes as equal observation changes up to fixed units.This assumption supports identifiability under suitable conditions and motivates flat or distance-preserving representations.
  • Limitations of existing assumptions: Independent hue and scale can remain geometrically coupled because their observation effects grow at different rates, producing position-dependent anisotropy that fixed rescaling cannot remove.For rendered objects, hue affects area while scale primarily affects boundary, so their relative effects change as the object grows.
  • Limitations of existing assumptions: Local geometry does not reveal whether a factor terminates or closes, and topology alone cannot distinguish homeomorphic spaces with different factor organizations, such as disks and cones.A circular factor may retain constant extent, shrink, or collapse at one point.
  • Contribution: FactoMap combines factor domains, generator-induced identifications, and factor-wise scales in a lattice whose topographic prototypes align representation coordinates with underlying factors.Its controlled evaluation compares matched and mismatched prototype domains in one- and two-factor settings.

2 Factor-Space Topographic Map (FactoMap)

The proposed factor-space formalism records domains, generator-induced identifications, and local factor scales, while FactoMap realizes that structure through a geometry-aware prototype lattice and topographic learning.

  • Factor-space structure: A generator-induced factor-space structure comprises factor domains, identifications defined by equal generator outputs, and local scales given by generator derivative norms.The realized factor space uses the quotient topology induced by those identifications.
  • Factor-space structure: Topology determines the organization of factor-space sites, while the scale vector records factor-wise local data-space effects; FactoMap models diagonal scales but not renderer-induced cross-terms.The cross-term limitation is analyzed separately for the HSV renderer.
  • Beyond uniform Euclidean geometry: Uniform Euclidean local isometry requires G(z) = C^2I up to global scale, but unequal, position-dependent diagonal scales or nonzero cross-terms violate this condition.FactoMap addresses position-dependent diagonal scales by varying lattice geometry according to the scale functions.
  • Beyond uniform Euclidean geometry: For hue and scale, the generator is defined on a circular hue factor and an interval scale factor, with rendering exponent γ = 1/2 for a fixed-width transition.The model specifies g(h, s)_xy = m_xy(s)ρ(h), with s constrained to a positive interval.
  • Beyond uniform Euclidean geometry: The induced metric is G = diag(a^2v^2s^2, b^2q^2s^(2γ)), and its anisotropy ratio changes with scale, so no fixed diagonal re-weighting makes it uniformly isotropic.The ratio satisfies κ(smax)/κ(smin) = (smax/smin)^(1−γ) over a non-degenerate scale interval.
  • FactoMap: A matched lattice can vary the cyclic extent along scale, unlike a uniform cylinder, because independent factors need not be geometrically separable.This varying extent represents the dependence of hue’s observation-space effect on scale.
  • FactoMap: FactoMap uses a factor-space lattice, data-space prototype field, and topographic objective; its analytic distance encodes wrapping, fibre identifications, and relative local extent.Each lattice site indexes a learnable prototype, and the best-matching site supplies a factor-space-indexed discrete coordinate.
  • FactoMap: Topographic cooperation encourages nearby factor-space sites to represent nearby observations, making the learned prototype field realize the supplied topology and geometry.The lattice and prototype map jointly define the learned FactoMap.

3 Results

Experiments show that FactoMap preserves factor continuity when lattice topology and local scale match the generative structure, whereas mismatched lattices mix hue and scale. The matched cone represents periodic hue and position-dependent geometry, enabling factor-aligned prototypes and improved disentanglement.

  • FactoShapes contains four controlled factors, with hue and scale evaluated as the varying factors in the reported experiments.
  • A line lattice learns scale as an ordered sequence with smoothly changing object size, while an open hue lattice cuts adjacent endpoint hues apart.The ring lattice reduces the endpoint separation and produces a continuous hue cycle.
  • Measured hue and scale speeds follow different power laws, with c_h(s) ∝ s and c_s(s) ∝ √s, producing position-dependent anisotropy.Across s ∈ [0.75, 1.5], anisotropy changes by approximately 1.404, close to the analytical prediction of 1.41.
  • The rectangular lattice mixes hue and scale because it neither wraps hue nor varies the cyclic extent with scale.
  • The matched cone lattice wraps hue and varies its fibre circumference with scale, yielding factor-aligned prototypes and substantially improving disentanglement over the mismatched grid.Because s_min > 0, the experimental lattice is a conical frustum rather than containing a collapsed fibre.
  • Table 1 compares disentanglement for the matched cone and mismatched grid.

A Factor-space and FactoMap Details

The paper defines notation for factor configurations, observations, factor-space structures, finite lattices, prototype fields, encoding, neighbourhood kernels, objectives, and prototype counts.

  • The generator g maps factor configurations z ∈ Z to observations x, while F_g denotes the induced factor-space structure and F_g = Z/≈_g its realized space.
  • A supplied structure F is realized by a finite lattice Λ_F with analytic distance d_F; each lattice site indexes a prototype w_k in observation space.
  • The encoder q_W returns the best-matching lattice site, and H^t_σ(j, k) denotes the topographic neighbourhood kernel at iteration t.
  • L_t(W) and bL_t(W) denote population and minibatch objectives, while p is the number of factors and K := |Λ_F| is the number of prototypes.

A.1 Remarks on the factor-space structure

The realized factor space inherits quotient topology from generator-induced identifications, while scale functions remain defined on the original domain and may fail to descend through collapsed fibres. FactoMap uses diagonal scales and separable lattice distances, and the finite lattice realization is supplied rather than uniquely determined.

  • The realized factor space F_g = Z/≈_g carries quotient topology induced by the canonical projection π: Z → F_g.
  • Scale functions are defined on Z and need not descend to the quotient, especially when the generator identifies an entire coordinate fibre but remaining scales vary along it.
  • The complete pullback metric is distinguished from the present FactoMap construction, which uses diagonal scales and analytic separable lattice distances.The construction does not represent nonzero off-diagonal terms.
  • Neither the scale functions nor quotient topology uniquely determines a global lattice distance, so (Λ_F, d_F) is a supplied finite realization and part of the model.

A.2 Derivation of the hue-scale metric

The hue-scale metric has factor-dependent diagonal scales, so their relative geometry changes with scale and cannot be made uniformly Euclidean by fixed rescaling. The HSV renderer additionally introduces generally nonzero cross-terms, while the model’s construction represents only diagonal scales.

  • Metric derivation: The hue and scale directions have diagonal scales ch(h,s)=av s and cs(h,s)=bq s^γ, producing scale-dependent anisotropy.Their ratio varies as (smax/smin)^(1−γ), so the relative effect of the two factors changes across the scale interval.
  • Metric derivation: No fixed diagonal re-weighting makes the metric uniformly Euclidean throughout a non-degenerate scale interval.The hue-hue entry remains proportional to s^2, preserving positional dependence after constant coordinate rescaling.
  • Metric derivation: A factor-wise reparameterization can equalize diagonal entries point-wise, but their common value remains position-dependent rather than uniformly isometric.The transformed scales retain distinct dependencies on the reparameterized hue and scale coordinates.
  • Renderer dependence: The scale exponent γ depends on the renderer and is measured near the fixed-width prediction, but the non-uniform-geometry conclusion does not depend on its precise value.The measured exponent is γ≈0.56, interpreted as a property of the rendering pipeline.
  • Collapse: Extending the model to zero scale identifies the entire hue fibre there while the hue scale simultaneously vanishes.Thus the global identification and local scale collapse describe the same degeneracy.

A.3 Analytic lattice distances

FactoMap takes a finite factor-space lattice and its distance as inputs, using distance definitions to encode open, wrapped, collapsed, and non-uniform structures. Neighborhood scaling must be specified jointly with lattice-distance normalization.

  • Inputs: FactoMap receives a finite lattice ΛF and distance dF; the present work does not estimate topology, identifications, or scales.The lattice structure is supplied rather than selected or learned by the method.
  • Product lattices: Weighted product distances assign positive lattice units λi to factor directions, while wrapped axes represent rings, cylinders, and tori.Open axes instead produce lines and grids.
  • Structured distances: Collapsed or non-uniform structures require analytic distances that specify both collapse implementation and how one factor’s extent depends on another.This extends beyond the regular product-lattice distance.
  • Neighborhoods: Because neighborhoods depend on dF/σt, globally rescaling lattice distance is equivalent to rescaling the neighborhood schedule.Distance normalization and the σt schedule therefore need to be specified together.
  • Optimization: For each minibatch, the method computes assignments, holds them fixed during the gradient step, and updates prototypes without differentiating through the arg min.Adam updates prototypes and BMUs are recomputed after every update.

B.1 Datasets

FactoShapes is a 64×64×3 synthetic dataset with hue, scale, and position factors arranged as S1×I3, and the experiments vary hue and scale continuously. Its HSV renderer introduces a nonzero metric cross-term, while the supplied lattice matches only part of the geometry.

  • Dataset construction: FactoShapes uses factors v=(object_hue, scale, posx, posy) with factor space S1×I3: hue is periodic, while scale and positions are intervals.The position bound is chosen so the projected object remains fully inside the frame at the largest scale.
  • Sampling: Continuous mode draws each factor independently from a continuous range, produces no repeated images, and sets integer class labels to −1.Grid and random-binned modes instead retain discrete factor classes.
  • Experimental datasets: The hue-and-scale experiments generate 100,000 continuous-mode samples with both varying factors drawn independently and uniformly.Position factors remain fixed in these experiments.
  • Renderer geometry: The HSV colour path is non-constant in RGB norm and piecewise differentiable across six linear hue sectors, making the renderer-induced cross-term generally nonzero.Within the first sector, the path is ρ(h)=(1,6h,0).
  • Geometric match: The supplied lattice matches factor topology and scale’s power-law dependence, but not the HSV cross-term or hue-dependent scale prefactor.The resulting evidence concerns a partial geometric match rather than general handling of non-diagonal metrics.
  • Dataset visualization: Figure 3 presents the uniform sampling distribution and visual samples while object_hue and scale vary.

B.2 Training and model parameters

Training uses Adam to learn grid and cone prototype lattices, with a scheduled learning rate and neighborhood radius. The cone is chosen to encode periodic hue and scale-dependent geometric growth.

  • Training schedule: Training runs 60,000 updates with batch size 128, Adam, initial learning rate η0=0.1, and exponential learning-rate decay.The neighborhood radius decays exponentially to σT=1.0 during the first 65% of training and then remains constant.
  • Model parameters: The cone FactoMap uses a 40×20 lattice with K=800 prototypes, whose wrapped rows form rings with radii from 18.4 to 37.4.The ring circumference is 2πkr with k=0.243.
  • Lattice geometry: A cone avoids the seam of an open rectangular grid and accommodates a non-product generator metric whose hue and scale coefficients grow as s^1.05 and s^0.56.The resulting full-hue arc length relative to radial spacing grows with scale.

C Additional Visualizations

Figure 4 visualizes learned prototypes on grid and cone lattices, showing how cone geometry supports alignment of hue and scale factors.

  • On the cone lattice, hue aligns with the circular direction while scale aligns with the axial direction.The grid lattice does not systematically align either factor with an axis.
  • The cone is cut open to visualize the alignment of prototypes across its lattice.
Loading 2608.24762v1…