Source-linked AI summary
The Universal Normal Embedding
Chen Tasker, Roy Betser, Eyal Gofer, Meir Yossef Levi, Guy Gilboa
TL;DR
Generative models and vision encoders lack a specified shared latent geometry and operational connection; this paper proposes the approximately Gaussian UNE and finds aligned semantic structure supporting controllable edits.
Problem
Existing work suggests encoder and generator latents are compatible but does not specify their shared geometry or an operational mechanism for using it.
Method
The paper formalizes UNE as an approximately Gaussian shared latent space, relates model latents through linear projections, and edits DDIM-inverted noise along probe-derived orthogonalized directions.
Results
DDIM-inverted noise and encoder embeddings support strong, aligned attribute prediction, while noise-space directions enable faithful controllable edits without architectural changes.
Takeaways & Limitations
The findings provide preliminary empirical support for a shared Gaussian-like latent geometry linking encoding and generation.
Takeaways & Limitations
The UNE hypothesis assumes an invertible mapping from data to a Gaussian latent space with linearly separable semantic properties, while models recover only noisy projections of it.
Abstract
from arXiv · showhide
Generative models and vision encoders have largely advanced on separate tracks, optimized for different goals and grounded in different mathematical principles. Yet, they share a fundamental property: latent space Gaussianity. Generative models map Gaussian noise to images, while encoders map images to semantic embeddings whose coordinates empirically behave as Gaussian. We hypothesize that both are views of a shared latent source, the Universal Normal Embedding (UNE): an approximately Gaussian latent space from which encoder embeddings and DDIM-inverted noise arise as noisy linear projections. To test our hypothesis, we introduce NoiseZoo, a dataset of per-image latents comprising DDIM-inverted diffusion noise and matching encoder representations (CLIP, DINO). On CelebA, linear probes in both spaces yield strong, aligned attribute predictions, indicating that generative noise encodes meaningful semantics along linear directions. These directions further enable faithful, controllable edits (e.g., smile, gender, age) without architectural changes, where simple orthogonalization mitigates spurious entanglements. Taken together, our results provide empirical support for the UNE hypothesis and reveal a shared Gaussian-like latent geometry that concretely links encoding and generation. Code and data are available https://rbetser.github.io/UNE/
1. Introduction
The paper proposes the Universal Normal Embedding (UNE), an approximately Gaussian latent space linking generative noise and encoder representations as noisy linear projections. It motivates UNE through shared latent geometry and develops empirical tests of Gaussianity, semantic separability, alignment, and controllability.
- UNE hypothesis: UNE posits a shared, approximately Gaussian latent space from which generative-model latents and vision-encoder representations arise as noisy linear projections.The hypothesis directly links generative noise with representations from encoders such as CLIP and DINO.
- Semantic geometry: UNE treats semantic variation as linear directions, enabling linear probes for classification and controllable edits through latent perturbations.The conceptual formulation describes class separation by hyperplanes and continuous-attribute editing along single latent directions.
- Motivation: The hypothesis is motivated by Gaussian priors in generative models, approximately Gaussian encoder coordinates, encoder identifiability up to linear transformations, and shared geometry across latent spaces.Prior work also shows that generative and representation spaces can be linearly stitched across models, architectures, and modalities.
- Empirical program: The empirical investigation analyzes per-image latents from multiple diffusion models and vision encoders against predicted properties including Gaussianity, semantic separability, cross-model alignment, and linear controllability.The study also examines multi-view intersections of latent spaces and introduces a unified per-image dataset for these analyses.
- Contributions: The stated contributions formalize UNE, relate it to real latents, explore a multi-view estimator for a shared k-dimensional intersection subspace, and study semantic structure in generative noise.The contribution list identifies DDIM-inverted noise as a target for semantic analysis, though the supplied passage is truncated before reporting the full findings.
2. Related Work
Prior work documents alignment among latent spaces and proposes shared representations, including theoretical identifiability and cross-encoder alignment results. This paper distinguishes its approach by modeling the shared space as approximately Gaussian and explicitly exploiting that geometry.
- Latent alignment and shared geometry: Latent spaces of VAEs, GANs, normalizing flows, and diffusion models often exhibit alignment, with linear mappings translating between independently trained spaces.These mappings can operate across models with different dimensionalities.
- Latent alignment and shared geometry: Theoretical frameworks propose shared latent descriptions, while identifiability results establish recovery of generative factors up to invertible transforms and tighter linear cross-encoder alignment.The cited progression runs from component-wise invertible transforms to linear identifiability.
- Gaussianity of representation spaces: Existing theories assume a shared space without specifying its geometry, and empirical alignment studies provide no operational mechanism for using it.This paper instead proposes an approximately Gaussian shared space supporting linear classification, semantic manipulation, and shared-space constructions.
3. Universal Normal Embedding (UNE)
The UNE hypothesis posits an information-preserving Gaussian latent space where semantic properties are linearly separable, while model representations are noisy linear projections of it. This framework explains Gaussian-like latents, aligned linear semantics, controllable edits, and shared structure across models.
- Universal Normal Embedding (UNE): The UNE is hypothesized as an N(0, I_D) latent space with an invertible mapping to data and linearly separable semantic properties.The latent dimension D is unknown, and the mapping is intended to preserve information.
- Induced Normal Embeddings (INE): Each model’s latent code is hypothesized to be an approximately Gaussian, noisy linear projection of the same underlying UNE.Model-specific dimensionality, objectives, architectures, modalities, transforms, and noise cause representations to expose different parts of the shared structure.
- Mapping Between Models: A shared low-dimensional proxy can extract latent structure consistently expressed across models, with deviations attributed to noise or unused dimensions.This construction uses linear mappings from each model’s INE and is presented as one initial alternative among many.
- Empirical Gaussianity: Across eight models, more than 90% of latent dimensions satisfy Gaussianity tests in most models, while nuisance dimensions coexist with Gaussian-like directions.The models include three Stable Diffusion generators and five encoders, evaluated using Anderson-Darling and D’Agostino-Pearson tests.
- Linear Semantics and Editing: Linear probes expose semantic attributes in both encoder representations and DDIM-inverted diffusion noise, supporting linearly accessible semantics beyond ideal UNEs.The framework motivates linear classifiers, regressors, and latent edits for approximately Gaussian attributes such as age, height, and smile intensity.
- Linear Semantics and Editing: Orthogonalizing semantic directions mitigates unintended attribute changes caused by spurious alignment between directions.The strategy is demonstrated for linear editing in DDIM-inverted latent space.
4. Experiments
Experiments evaluate NoiseZoo through cross-space classification, controllable editing, and recovery of shared latent structure. Across these tests, generative noise latents show meaningful attribute information, align with encoder spaces, and support shared low-dimensional representations and disentangled edits.
- Experimental setup: NoiseZoo evaluates linear classification within and across latent spaces, controllable probe-direction editing, and recovery of a shared k-dimensional core.The dataset contains per-image latents designed to support all three experimental axes.
- Experimental setup: CelebA experiments use approximately 19k validation images and latents from five vision encoders plus approximately 16k-dimensional DDIM-inverted diffusion noise.The validation set is split into 15k training and 4k test samples.
- Gaussianity: Generative models approach the theoretical 95% acceptance rate in random-projection normality tests, while reference spaces perform substantially worse.Gaussianity is assessed with Anderson-Darling, D’Agostino-Pearson, and Shapiro-Wilk tests over 5,000 projections from 250 sampled points per model.
- Classification and transfer: DDIM-inverted noise latents achieve attribute separability only slightly below leading encoders, with highly correlated per-attribute performance.Transferred-latent evaluations report low MSE, high cosine similarity, and accuracy drops below 0.3%.
- Controllable editing: Probe-derived directions enable smooth, local, controllable edits across six CelebA attributes without prompts or fine-tuning, while orthogonalization suppresses spurious attribute changes.Edit strength varies with α, and Figure 5 shows orthogonalized directions isolating target attributes such as smiles, gender, age, or facial features.
- Shared latent spaces: Shared spaces built from four or six balanced encoder and generative sources retain strong attribute classification, concentrate information in shared directions, and exhibit consistently high neighborhood-structure correlations.Retrieval analysis on 10k images compares similarity vectors using Spearman rank correlations across shared spaces.
5. Conclusions
The paper introduces the Universal Normal Embedding hypothesis, which proposes that generative and representation models approximate a shared Gaussian latent geometry with semantic factors represented by linear directions. It reports that DDIM-inverted noise codes and representation embeddings encode comparable, linearly decodable semantic structure.
- 5. Conclusions: The Universal Normal Embedding hypothesis posits a shared Gaussian latent geometry between generative and representation models.Under this hypothesis, semantic factors correspond to linear directions.
- 5. Conclusions: A direct consequence of UNE is that generators contain linearly separable semantics, similarly to encoders.
- 5. Conclusions: Empirically, DDIM-inverted noise codes and representation embeddings encode comparable semantic structure with linearly decodable attributes.
Overview
The supplementary material adds implementation and experimental details for reproducibility, plus further analyses, qualitative examples, and experiments on an additional dataset.
- Section A: The supplement provides implementation and experimental details to support full reproducibility.These materials are presented in Section A.
- Section B: It also extends the linear editing analysis with qualitative examples and experiments on an additional dataset.These additions are presented in Section B.
A.1. NoiseZoo construction details
NoiseZoo contains per-image diffusion and encoder latents for all 19,867 CelebA validation images, using specified Stable Diffusion, CLIP, OpenCLIP, and DINOv3 variants. The dataset is split into 15,893 training and 3,974 test samples, with distinct preprocessing and extraction procedures for diffusion and encoder representations.
- Dataset composition: NoiseZoo covers all 19,867 unfiltered images in the CelebA validation split and includes latents from three Stable Diffusion, four CLIP/OpenCLIP, and one DINOv3 variant.The dataset uses SD 1.5, SD 2.1, LCM, CLIP ViT-B/16 and ViT-L/14, OpenCLIP ViT-B/16 and ViT-L/14, and DINOv3 ViT-L/16.
- Dataset split: The dataset was randomly divided into 15,893 training samples and 3,974 test samples.
- Stable Diffusion latents: Diffusion latents were extracted by DDIM inversion after center-cropping and resizing images to 512×512, using an empty prompt, guidance scale 3.5, seed 42, and model-specific step counts.SD 1.5 and SD 2.1 used 50 DDIM steps, while LCM used 150 steps with a DDIMScheduler.
- Encoder latents: Encoder latents were obtained from original CelebA images using each model’s default preprocessing, without additional normalization, with dimensions of 512 or 768.DINOv3 images were center-cropped and resized to 224x224; CLIP and OpenCLIP dimensions were 512 for ViT-B/16 and 768 for ViT-L/14, while DINOv3 was 768.
A.2. Experimental details · B. Editing Examples · B.1. Comparison of editing in different models
The experiments used regularized linear probes and ridge mappings to analyze and translate latent representations. Editing comparisons showed that diffusion latents preserve input images better than CLIP embeddings while achieving target attributes.
- A.2. Experimental details: Each feature set used PCA, standard scaling, and separate logistic regression classifiers for 40 attributes.PCA retained 500 components for generative models and 310 for encoders; logistic regression used saga, L2 regularization, 25 iterations, and 30 parallel jobs.
- A.2. Experimental details: Cross-space transfer learned a linear ridge mapping from paired training representations and evaluated target-space classifiers on translated test representations.The mapping used paired samples from the training set, with evaluation reported in Table 2.
- A.2. Experimental details: Ridge regularization was scaled by source-feature energy to maintain consistent penalties across latent representations.The effective penalty was defined as αeff = α ||Xsource||2_F / d, with α = 1.0 in the reported results.
- A.2. Experimental details: Shared-latent experiments used five specified combinations of diffusion models, encoders, and CLIP variants.The combinations included SD 1.5, SD 2.1, LCM, CLIP B/16 or L/14, OpenCLIP B/16, and DINOv3.
- B. Editing Examples: Figure 7 compared linear editing in SD 1.5, LCM, and CLIP ViT-L/14 latent spaces.CLIP embeddings were inverted with the UnCLIP variant of Stable Diffusion.
- B.1. Comparison of editing in different models: Diffusion latents faithfully modified original images, whereas CLIP edits achieved target attributes but failed to preserve the input.The comparison covered SD 1.5, LCM, and CLIP ViT-L/14 latent spaces.
B.2. Quantitative analysis of editing
Quantitative evaluation on SD 1.5 latents shows that increasing edit intensity strengthens alignment with the target attribute while original-image similarity peaks at zero intensity.
- Quantitative editing results: As attribute intensity increases, similarity between edited images and the CLIP text embedding of the attribute name increases.Intensity is normalized by distance to the classifier’s decision plane; x = 0 denotes editing to that plane, not no editing.
- Quantitative editing results: Similarity between edited and original images peaks at zero intensity.The evaluation uses SD 1.5 latents and measures both image preservation and target-attribute alignment.
B.3. Effect of model scale, conditioning, and pixel space
Generative latent spaces support substantially stronger linear attribute classification than pixel space, while broad training appears more important than model scale or CelebA fine-tuning for linear separability.
- Representation and model comparisons: Pixel-space representations yield substantially lower linear attribute-classification performance than generative latent spaces.Figure 9 compares pixel space with SD 1.5, SD CelebA, CelebA Diff, and SDXL representations.
- Representation and model comparisons: The smaller CelebA-only diffusion model shows a clear degradation in linear separability.The comparison includes CelebA Diff, a smaller unconditional model trained only on CelebA.
- Representation and model comparisons: Increasing model scale from SD 1.5/2.1 to SDXL produces only marginal improvements in linear separability.The passage contrasts SDXL with SD 1.5/2.1 and characterizes the gains as marginal.
- Representation and model comparisons: Fine-tuning on CelebA reduces linear separability even when evaluated on the same dataset, underscoring the importance of broad, diverse training.SD CelebA is the CelebA-finetuned comparison model.
B.4. Evaluation on additional datasets
On AFHQ’s diverse animal faces, semantic categories remain linearly separable across species, and classifier-direction latent shifts produce realistic edits that preserve original image structure.
- AFHQ dataset: AFHQ spans Cat, Dog, and Wild animal-face categories, with finer Wild labels assigned using CLIP scores and dominant class labels.The evaluation tests generalization beyond CelebA on visually diverse animal faces.
- Classification: Semantic categories remain structured and linearly separable across species under pairwise classification with CLIP-prompt-defined sub-categories.Figure 9 reports AFHQ binary-classification AUC values using SD 1.5 latents.
- Latent editing: Shifting AFHQ latents along learned classifier directions produces realistic animal-face edits while preserving the original image structure.The procedure follows Section 4.3 and indicates that learned directions capture shared high-level semantic information despite visual diversity.