Source-linked AI summary
What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
Hao Chen
TL;DR
FID can miss distributional differences and cannot indicate whether dispersion is too low or too high. ZID combines complementary statistics with permutation calibration and signed readouts, providing broad detection, severity ranking, and dispersion diagnosis across controlled departures and generative-model guidance sweeps.
Problem
FID’s reliance on feature means and covariances can make distinct distributions indistinguishable and cannot encode dispersion-change direction.
Method
ZID combines six complementary graph- and kernel-based arms with outer permutation calibration, reporting departure magnitude, equality-test evidence, and signed dispersion direction.
Results
ZID provides broad detection coverage, consistently ranks increasing severity, and identifies under- versus over-dispersion when direction is identifiable.
Takeaways & Limitations
The evidence supports pairing calibrated distributional-difference testing with diagnostic readouts instead of relying on an unsigned moment summary alone.
Takeaways & Limitations
In reference-adaptive matched-moment stress tests, ZID’s permutation tail is descriptive rather than a calibrated population two-sample p-value.
Abstract
from arXiv · showhide
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce \textbf{ZID} (\emph{Z-resolved Integrated Diagnostic}), which combines six standardized location- and dispersion-sensitive arms from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths). Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis. In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion.
1 Introduction
FID and KID can miss important distributional departures, reverse rankings, and provide limited directional or calibrated-testing information. ZID addresses these gaps with broad-coverage detection, severity ranking, permutation-calibrated equality testing, and signed dispersion diagnosis.
- Matched-moment analysis: FID’s first-two-moment restriction lets visual noise score 24.7 versus 58.6 for held-out real images, reversing the nominal ranking.The noise-initialized images were optimized only to match the reference Inception mean and covariance.
- Matched-moment analysis: 19.5 versus 0.003: FID distinguishes real from held-out real features more than from a moment-restored bimodal batch, while KID estimates also invert.The corresponding KID estimates are 9.5×10−5 and −3.2×10−4, respectively.
- Directional diagnosis: FID and KID cannot distinguish under-dispersion from over-dispersion because their scalar discrepancies are unchanged when the two samples are exchanged.ZID adds a separate signed dispersion readout for directionally identifiable departures.
- Calibrated detection: At d = 100 and n = 100, KID power stays below .25 across 5%–25% dispersion shrinkages, while FID reaches only .75 power for 5% shrinkage at n = 1000.These results show sensitivity to sampling variation can limit scalar discrepancy tests.
- ZID contributions: ZID combines graph- and kernel-based statistics with permutation calibration to provide a departure index, equality-test p-value, and signed under- or over-dispersion label.Across controlled evaluations, it offers broad detection coverage and consistently ranks increasing severity, including cases where FID is flat or reversed.
2 Background and related metrics
Existing generative-evaluation metrics compare feature distributions through moments or kernels, while other measures target diversity, coverage, fidelity, or dispersion. Related work also addresses FID’s finite-sample bias and multivariate dispersion testing.
- Distributional metrics: FID compares fitted feature-space Gaussians using means and covariances, whereas KID is MMD with a degree-3 polynomial kernel.Gaussian-kernel MMD and CMMD are kernel two-sample discrepancies distinguished by their kernels or feature representations.
- Distributional metrics: FID’s finite-sample bias motivated the extrapolation-based FID∞ correction.
- Diversity, coverage, and dispersion: Vendi Score summarizes sample diversity, Precision/Recall separates fidelity from coverage, and Density/Coverage refines neighborhood-based estimates.
- Diversity, coverage, and dispersion: PERMDISP tests homogeneity of multivariate dispersion using distances to group centroids.
3 The ZID diagnostic: score, test, and signed readout
ZID separates generative evaluation into an omnibus departure score, a permutation-calibrated equality test, and a signed dispersion readout. Its six rank- and kernel-based arms detect departures that scalar moment-based discrepancies can miss while distinguishing under- from over-dispersion.
- Signed readout: The signed ZD coordinate contrasts within-sample quantities, reverses sign when samples are exchanged, and supports an under- versus over-dispersion diagnosis.ZW captures aligned displacement, whereas ZD captures opposing displacement; scale-dominated alternatives can signal primarily through signed ZD.
- Motivation: FID can be exactly zero for distinct distributions sharing mean and covariance, establishing a two-moment blind spot that motivates ZID’s broader diagnostics.The proposition states that for every finite-second-moment distribution with nonzero covariance, a distinct distribution can share its mean and covariance, implying FID(P, Q) = 0.
- ZID score: ZID combines six standardized arms from rank-based RISE and Gaussian kernels at two bandwidths into an omnibus departure index.The construction retains RISE, GPK-med, and GPK-small, with six total arms.
- ZID score: The flat-Simes score is monotone in standardized arm magnitudes, unbounded, and driven by strong signals in one or several arms.It is a reference-scale ranking quantity rather than a finite-sample arm p-value; larger values indicate more extreme departures.
- Permutation test: The outer plus-one Monte Carlo permutation procedure recomputes the full construction under B relabelings and provides a finite-sample-valid p-value without independence or Gaussian assumptions.At nominal level α, the test rejects when bpZID ≤ α.
4 Feature-space evaluation under controlled departures
Across controlled feature-space departures, ZID combines broad detection power with reliable severity ranking and signed dispersion diagnosis, outperforming FID and KID especially on moment-restored and dispersion-collapse sweeps. Its calibration is near nominal, and its component and cross-dataset analyses show complementary arms and persistent response patterns.
- Controlled departures: ZID is the only compared method with power at least .70 across all eight departure columns, leading nonlinear dependence (.97), skewness (.79), kurtosis (.83), and matched-moment multimodality (.71).KID leads location (.89), PERMDISP dispersion (.87), FID linear dependence (.87), and Density off-manifold support (.93).
- Calibration: Null rejection rates are .048 at α = .05 and .010 at α = .01 for same-pool real-feature samples, supporting calibration at the tested representation and sample size.The experiment used 500 repetitions with 999 outer permutations and reported a Wilson 95% CI of [.032, .070] for α = .05.
- Directional diagnosis: The formal DZID readout is negative for displayed contractions and positive for displayed expansions, distinguishing under-dispersion from over-dispersion.The ungated net signed dispersion score changes sign across the full-rank reference-scale sweep.
- Severity ranking: ZID maintains positive rank association on every severity sweep, while FID reverses dispersion-collapse ordering (ρ = −.41) and is essentially flat (|ρ| ≤.05) on four moment-restored departures.KID remains weak on dispersion collapse and the four moment-restored sweeps (ρ ≤.20).
- Severity ranking: On dispersion collapse and four moment-restored departures, ZID orders 83–100% of pairs correctly, versus FID at or below chance (27–55%) and KID at 48–64%.All three methods are near-perfect on location, linear dependence, and off-manifold support.
- ZID construction: Ablations show complementary arms: omitting RISE causes the largest power loss on six departures, GPK-med lowers location power from .830 to .237, and GPK-small lowers nonlinear-dependence power from .980 to .860.GPK-small’s incremental nonlinear-dependence contribution grows from +.013 at multiplier 1.25 to +.343 at .50.
5 Evaluation of pretrained generative models
Across five pretrained generator evaluations, ZID detects departures while ranking their magnitude and diagnosing dispersion direction. It reveals under-dispersion in truncation and high-guidance-collapse settings, including cases where conventional distances provide no direction or are misleading.
- BigGAN-deep-128: KID rises from .0075 to .0173 and ZID from 328 to 1150 as BigGAN-deep-128 truncation tightens, while ZID rejects all comparisons and reports under-dispersion.FID is lowest at ψ = .8 and rises under stronger truncation; the ZID dispersion SD increases from 42.2 to 176.7.
- CIFAR DDPM: FID decreases from 173.6 to 130.8 toward the real–real value 121.0 as CIFAR-10 DDIM steps increase, while ZID rejects every comparison and its direction changes at later steps.ZID reports under-dispersion at five steps and a member-sign conflict at 15 and 25 steps.
- Guidance sweeps: On DiT-XL/2, ZID’s W-component magnitude rises from CFG 2 through 8, while the aggregate readout changes from over-dispersion at CFG 1 to under-dispersion at CFG 2, 4, and 8.The result holds across all six ImageNet classes, with positive DZID at low guidance and negative DZID as guidance rises.
- Guidance sweeps: On SiT-XL/2, the signed readout indicates over-dispersion at CFG = 1 and under-dispersion at CFG ∈{4, 8, 16}, with five classes under-dispersed at CFG = 2.At CFG = 16, FID exceeds its CFG = 1 value on five classes but remains lower on one, illustrating differing behavior across metrics.
- StyleGAN2-ADA: FID, KID, and Gaussian-kernel MMD increase monotonically with StyleGAN2-ADA truncation, while ZID separates every setting and diagnoses increasingly strong under-dispersion as ψ falls.Precision and Density increase as Recall declines, emphasizing different aspects of the precision–recall tradeoff.
- Cross-generator findings: Across five generator families, ZID combines scalar ordering, calibrated detection, and directional diagnosis, distinguishing departures on both sides of guidance sweeps and identifying truncation-induced under-dispersion.Conventional distances can respond to departures but do not provide dispersion direction.
6 Gaming FID in pixel space
Pixel-level moment matching can make visually unrecognizable images achieve lower FID than real–real baselines, exposing FID’s blind spot. ZID still flags these optimized batches as extreme, but reference-adaptive optimization makes the resulting permutation tail descriptive rather than calibrated.
- Pixel-space moment matching: 512 images optimized in pixel space to match ImageNet Inception mean and covariance remained visually unrecognizable while targeting the moments that determine FID.The optimization began from noise at 64×64 resolution, with the pretrained feature extractor fixed and only pixels updated.
- Pixel-space moment matching: 24.7 FID for the gamed ImageNet images fell below the 58.6 real–real baseline, despite their visual unrecognizability.On the same feature banks, real–holdout and real–gamed KID estimates were 8.6×10−6 and −5.2×10−5, respectively.
- ZID diagnosis: ZID’s gamed-pair component magnitudes were SW=17.9 and SD=9.8, with SD exceeding all 499 relabeled values.The RISE dispersion arm was not reference-active, while two active GPK arms had opposite signs, making DZID undefined and the dispersion diagnosis member-sign conflict.
- Inference limitation: The 512-image batch was reference-adaptive rather than an i.i.d. draw, so its plus-one permutation tail of .002 was descriptive, not a calibrated population two-sample p-value.None of the 499 random outer relabelings produced an equally or more extreme score for this fixed pair.
- Replication: 31.1 FID for the CIFAR-10 gamed images fell below the 76.0 real–real baseline, while the ZID score exceeded all 499 randomly relabeled scores.The replication used the same N=512 optimization setup.
7 Conclusion · A Foundations and controlled validation
FID’s mean-and-covariance compression can hide distinct distributions and cannot reveal whether diversity changes are upward or downward. ZID instead separates severity ranking, calibrated equality testing, and diagnostic readouts, with controlled evidence showing broad detection and progression tracking.
- 7 Conclusion: FID’s mean-and-covariance summary can make distinct feature distributions indistinguishable and cannot encode diversity-change direction.
- 7 Conclusion: ZID separates severity ordering, distributional-equality testing, and component-level diagnostic readouts instead of assigning all roles to one scalar.Its six-arm score orders settings within a fixed protocol, while an outer-permutation p-value tests equality.
- A Foundations and controlled validation: Across controlled departures, real-feature replications, and pretrained-model evaluations, ZID provides broad detection and tracks progression as severity increases.
- A Foundations and controlled validation: ZID distinguishes low-guidance over-dispersion from high-guidance collapse.
- 7 Conclusion: Retained member signs reveal when dispersion direction differs across scales.
- 7 Conclusion: The evidence supports pairing calibrated distributional-difference evidence with diagnostic readouts rather than relying on an unsigned scalar.
A.1 Matched-moment feature controls
Matched-moment controls test whether FID’s feature-space inversion persists across embeddings and datasets, while additional controls restore reference moments or use adapted optimized batches. These experiments clarify the finite-sample baselines and limit interpretation of permutation tails for reference-adaptive constructions.
- Matched-moment feature controls: Table 9 tests whether the feature-space FID inversion is specific to one dataset or embedding using 2,500 samples projected onto 128 pooled principal components.The Gaussian construction is drawn from the Gaussian fitted to the reference sample.
- Matched-moment feature controls: The large-sample population-matched alternative follows the real–real finite-sample empirical FID baseline rather than yielding numerically zero, with n = 25k observations per sample.The six-arm ZID comparison uses n = 2k, 20 repetitions, and 99 outer permutations.
- Matched-moment feature controls: Table 10 compares raw Inception-2048 features with m = n = 2,500 after recoloring a matched-bimodal sample to the reference mean and covariance.KID is an unbiased degree-3 polynomial-kernel MMD estimate and may be slightly negative.
- Matched-moment feature controls: In the pixel-space stress test, KID is 8.6×10−6 for real versus held-out real and −5.2×10−5 for real versus the optimized batch.The comparison uses a separate pair of 512 × 2048 feature banks.
- Matched-moment feature controls: The CIFAR-10 replication optimizes a noise-initialized batch to match one real sample’s feature mean and covariance, so its permutation tail is descriptive rather than a calibrated population two-sample p-value.A disjoint real sample supplies the real–real baseline.
A.2 Proof of Proposition 1 · A.3 Controlled-departure overview · A.4 Outer-permutation calibration
The proof shows that matching only the first two moments cannot guarantee distributional equality. Controlled departures span location, dispersion, dependence, higher moments, multimodality, and low-variance noise, while outer-permutation calibration maintains near-nominal null rejection rates.
- A.2 Proof of Proposition 1: A distribution Q can differ from P while sharing its mean and covariance, so the moment-based discrepancy is zero despite unequal distributions.The construction covers both non-Gaussian P and Gaussian P, including Gaussian laws supported on a proper subspace.
- A.3 Controlled-departure overview: Eight controlled transformations test leading-direction location shifts, isotropic contraction or expansion, dependence, higher moments, multimodality, and low-variance-coordinate noise.Dependence and higher-order transformations are followed by affine moment restoration where required.
- A.3 Controlled-departure overview: The evaluations use standard Inception-v3 penultimate features, DINO ViT-S/16 CLS embeddings, and DINOv2 for the DDPM representation analysis.The pixel-space stress test and controlled feature-space studies use separate standard Inception implementations, so absolute FID and KID values are interpreted on their respective implementations.
- A.3 Controlled-departure overview: Moment-restored resampling varies departure through the replacement fraction and perturbation scale ε relative to coordinatewise standard deviations.The procedure begins with an independent real sample, replaces a specified fraction with perturbed resamples, and restores empirical mean and covariance.
- A.3 Controlled-departure overview: Comparator protocols standardize projections, nearest neighbors, cosine-similarity Gram matrices, centroid distances, and unbiased Gaussian-kernel MMD across methods.Permutation-calibrated rows reuse pooled samples and the statistic-specific representation.
- A.3 Controlled-departure overview: C2ST uses a stratified half split and a fixed one-hidden-layer MLP, with calibration relabeling held-out labels while keeping predictions fixed.Held-out accuracy is evaluated on the second half after training on the first.
- A.3 Controlled-departure overview: ECS sensitivity depends on frequency parameter T: smaller values emphasize location and lower-order differences, whereas larger values emphasize higher moments and tails.An independent preliminary study selected T = .28, and results remain conditional on the PCA coordinates because coordinatewise ECS is not invariant to rescaling or rotation.
- A.4 Outer-permutation calibration: Table 11’s six-arm equality test has rejection rates near nominal levels across two datasets and two embeddings, supporting outer-permutation calibration.The study uses 500 repetitions per setting, m = n = 200, and 99 permutations; the main-text check uses 999 permutations for finer p-value resolution.
A.5 GPK, GET, and RISE constructions and sensitivity to arm-level tail probabilities … B.2 BigGAN truncation control
The supporting analyses distinguish the component constructions and calibration roles, find broadly similar arm-tail behavior, and show strong high-dimensional dispersion sensitivity for GPK. Pretrained-model evaluations use fixed protocols, while BigGAN truncation produces rising ZID and under-dispersion as diversity tightens.
- A.5 GPK, GET, and RISE constructions and sensitivity to arm-level tail probabilities: GPK uses dense Gaussian-kernel weights, GET uses edge-disjoint minimum spanning trees, and RISE uses a rank-weighted directed nearest-neighbor graph.Standalone tests use fixed graph specifications and pooled median-distance GPK bandwidths.
- A.5 GPK, GET, and RISE constructions and sensitivity to arm-level tail probabilities: The final six-arm ZID p-value is an outer permutation tail, whereas individual GPK coordinates use normal-reference tails and standalone GPK tests use label-permutation calibration.These family-specific calibrations apply only to the standalone rows.
- A.5 GPK, GET, and RISE constructions and sensitivity to arm-level tail probabilities: .964 and .958 are the pooled probabilities that alternative scores exceed paired null scores under Gaussian-reference and permutation arm tails, respectively.Across 2,400 paired null comparisons, rejection rates are .050 and .037; Gaussian-reference tails have somewhat higher finite-sample detection power.
- A.6 High-dimensional dispersion sensitivity and coordinate separation: .98 is GPK power for 1% shrinkage at d = 2048 and n = 100; at d = 10, power is .493 for 7% shrinkage and .84 for 10% shrinkage.The comparison uses the same Gaussian reference–alternative pairs and 99 relabelings, with power estimated over 300 repetitions.
- B.1 Sampling and calibration for pretrained generative models: Pretrained-model sweeps hold model weights, representations, pipelines, and reference sets fixed, changing only the listed sampling or truncation parameter.ZID equality-test and D-component p-values use the same six-arm construction and 499 outer relabelings.
- B.2 BigGAN truncation control: 328 to 1150 is the ZID-score increase as BigGAN truncation tightens, while aggregated SD rises from 42.2 to 176.7 and the diagnosis remains under-dispersion.For m = n = 450 over 10 ImageNet classes, real–real FID is 70.3 and generator–real FID values are 89–102.
B.3 Detailed CFG sweeps: DiT and SiT
Detailed guidance sweeps evaluate DiT-XL/2 and SiT-XL/2 against class-matched real ImageNet samples across six classes. ZID provides calibrated detection and directional diagnosis, including a clear shift from over-dispersion to under-dispersion at high guidance.
- DiT-XL/2: Table 14 reports per-class DiT-XL/2 classifier-free-guidance results for six ImageNet classes, with m = n = 500 and all bpZID = .002 using 499 permutations.Parentheses report bpD, while “over” and “under” abbreviate dispersion diagnoses.
- SiT-XL/2: Figure 9 summarizes the corresponding SiT-XL/2 guidance transition over CFG ∈ {1, 2, 4, 8, 16} using the same six classes and n = 500.SiT-XL/2 is described as an interpolant transformer with a continuous flow objective.
- SiT-XL/2: FID, KID, and Frobenius covariance discrepancy are minimized near CFG = 2 and increase on both sides, but their scalar values do not identify over- versus under-dispersion.The sweep therefore requires directional readouts beyond scalar discrepancy magnitudes.
- SiT-XL/2: On class 10, ZID’s D-component magnitude SD rises from 42.3 at CFG = 1 to 176.6 at CFG = 16 while its direction changes from over-dispersion to under-dispersion.This final aggregate readout supplies the dispersion direction across the guidance sweep.
C Additional aggregation, ranking, and bandwidth controls · D Detection across severity, dimension, and sample size · D.1 Detection across increasing severity
The controls support ZID’s aggregation, ranking, and bandwidth choices, while severity sweeps show detection power increasing across controlled departures. ZID ranks strongly against external comparators and remains stable near its selected GPK-small bandwidth.
- C Additional aggregation, ranking, and bandwidth controls: Flat Simes has the highest mean power and strongest ranking summaries across full severity ladders, with only small descriptive differences from max and Cauchy.Cauchy has a one-point higher minimum power estimate, while sum-of-squares and Fisher are weaker in the reported comparison.
- C Additional aggregation, ranking, and bandwidth controls: ZID ranks first across all four expanded ranking summaries against FID, KID, multiscale MMD, MIND, and ECS.The summaries are mean and minimum Spearman correlations and mean and minimum pairwise ordering accuracies across eight departures.
- C Additional aggregation, ranking, and bandwidth controls: .969 versus .968 is ZID’s mean pairwise ordering accuracy compared with the best tied standalone test, while ZID’s minimum accuracy is .90 versus .83.ZID also has higher mean and minimum Spearman correlations; full-ladder monotonicity differs from mild-versus-final ordering.
- C Additional aggregation, ranking, and bandwidth controls: .70 to .71 is the observed range of minimum power as the GPK-small multiplier varies from .10 to .25.Mean null rejection ranges from .035 to .048 while samples and relabelings remain fixed.
- C Additional aggregation, ranking, and bandwidth controls: .15 and .20 bandwidth ratios differ from .175 by at most .05 in power on every departure, with paired decision agreement at least .95 under alternatives and .98 under the null.Only the GPK-small multiplier changes; the other arms, samples, and relabelings are shared within replicates.
- D.1 Detection across increasing severity: Five-level severity ladders extend the Fig. 3 signal using multiples {0, .5, .75, 1, 1.25} for seven departures and replacement fractions λ ∈{0, .25, .5, .75, 1} for multimodality.Each level uses 300 repetitions and 499 outer permutations per cell.
- D.1 Detection across increasing severity: Detection rejection rates are evaluated at the 0.05 significance level across increasing-severity ladders, with Wilson 95% intervals and reference lines at .05 and .80.For multimodality, severity is the fraction of initially real observations replaced by moment-matched bimodal counterparts; other departures use signal multiples.
D.2 Dimension–sample-size grid
The dimension–sample-size grid evaluates detection across three feature dimensions and three sample sizes using eight controlled departures. Results are compared within cells under shared signals, while null-rejection estimates summarize worst-panel calibration across methods.
- D.2 Dimension–sample-size grid: The grid spans d ∈{128, 512, 2048} and n ∈{50, 200, 600} across eight controlled-departure constructions.
- D.2 Dimension–sample-size grid: Comparisons are within cells because one panel-selected signal is shared across methods to limit widespread ceiling effects.
- D.2 Dimension–sample-size grid: Figure 11 reports power over nine (d, n) cells and eight departures, using 100 repetitions at the 0.05 significance level.The experiment uses CIFAR-10 Inception features projected to d pooled principal components; permutation-calibrated rows use 99 relabelings.
- D.2 Dimension–sample-size grid: Table 18 summarizes each method’s maximum panel-specific null rejection estimate for every (d, n) cell.Values range from .00 to .10, with rejection defined by p ≤.05.