Source-linked AI summary
DefaultShift: Auditing Semantic Default Shift in Accelerated Text-to-Image Models
Xuanhua Yin, Chuanzhi Xu, Shunqi Mao, Wei Guo, Weidong Cai
TL;DR
Accelerated text-to-image replacements can alter prompt-unspecified semantic distributions despite plausible individual outputs, and existing evaluations do not directly test their preservation. The paper introduces DefaultShift, a paired audit of repeated labeled samples and probability-mass movement, plus DefaultShift-Select for offline calibration. Across 14 pairs, adjusted color discrepancies span 0.054–0.303; selection reduces human-measured shift by 10.3%–35.1% and recovers balanced downstream performance.
Problem
Existing quality, preference, and diversity evaluations do not directly test whether an accelerated replacement preserves its reference model’s distributions over unspecified attributes.
Method
DefaultShift repeatedly samples reference–replacement pairs, labels outputs with closed semantic vocabularies, measures magnitude and direction of probability-mass movement, and separates ranking from confirmatory inference.
Results
Across 14 reference–replacement pairs, adjusted color discrepancies range from 0.054 to 0.303, while DefaultShift-Select reduces human-measured shift by 10.3%–35.1% across Turbo, DMD2, and FLUX without material quality loss.
Takeaways & Limitations
DefaultShift makes semantic preservation under acceleration measurable, and DefaultShift-Select provides an offline way to reduce measured shift while improving balanced downstream performance.
Takeaways & Limitations
The audit score is not an unbiased estimator of population TV, so q = 32 is used for uniform-cost screening; evidence is strongest for color and balanced downstream evaluation.
Abstract
from arXiv · showhide
Few-step text-to-image models increasingly replace slower generators, yet acceleration can silently change distributions over unspecified attributes even when individual outputs remain plausible and aligned. We call these distributions semantic defaults and their change under replacement semantic default shift. Existing quality, preference, and diversity evaluations do not test whether a replacement preserves its reference model's semantic defaults. We introduce DefaultShift, a paired audit that labels repeated samples with closed semantic vocabularies, measures probability-mass movement, and separates interpretable ranking from confirmatory cross-fit inference. Across 14 reference and replacement pairs, adjusted color discrepancies range from 0.054 to 0.303 with recipe-specific directions. A 1,000-image human audit reproduces the ordering. We further introduce DefaultShift-Select, an offline calibration method that reduces human-measured shift by 10.3 percent to 35.1 percent across Turbo, DMD2, and FLUX without material quality loss. Under balanced evaluation, selected data recover 4.3 accuracy points and 7.5 worst-group points over uncalibrated replacement data. DefaultShift makes semantic preservation under acceleration measurable and actionable.
1 Introduction
Accelerated text-to-image replacements can preserve individual plausibility while shifting distributions over prompt-unspecified attributes. DefaultShift audits this reference–replacement change and shows recipe-specific shifts that standard evaluations may miss.
- Accelerated replacements can favor different unspecified attributes even when individual images remain realistic and prompt-aligned.These prompt-conditioned distributions are called semantic defaults, and their change under replacement is semantic default shift.
- Current quality, alignment, preference, diversity, and coverage evaluations do not directly test preservation of a reference model’s semantic distributions.Existing attribute-level studies characterize defaults within individual models rather than tracking probability-mass movement between a declared reference–replacement pair.
- DefaultShift repeatedly samples both models, assigns labels from a closed semantic vocabulary, and compares attribute distributions.The audit measures probability-mass movement while accounting for finite-sample effects, with separate scores for interpretable ranking and confirmatory inference.
- Across 14 reference–replacement pairs, adjusted color discrepancies range from 0.054 to 0.303 and shift directions depend on the acceleration recipe.DMD2 and Turbo reduce gray mass and increase warm-color mass, whereas FLUX moves in the opposite direction.
- DefaultShift-Select selects a quality-filtered subset of replacement outputs to better match the reference distribution, reducing human-measured shift by 10.3%–35.1% without material quality loss.In balanced synthetic-data classification, selected data recover 4.3 accuracy points and 7.5 worst-group points over uncalibrated replacement data.
2 Related Work
Prior work emphasizes fidelity, coverage, alignment, preference, or within-model attribute diversity, while this paper frames replacement auditing as set-level semantic distribution matching. It therefore compares declared reference–replacement pairs rather than treating fast models as interchangeable.
- Fast text-to-image methods combine solver, distillation, adversarial, distribution-matching, and flow-compression recipes optimized for fidelity and latency.These recipes need not preserve the semantic behavior of a slower reference model.
- DefaultShift audits each declared deployment replacement instead of treating all accelerated models as equivalent.The valid comparison is reference-relative; for example, NitroSD-Realism is distilled from DMD2 rather than directly from SDXL.
- Aggregate fidelity, coverage, alignment, and preference metrics complement but do not directly compare semantic defaults across a reference–replacement pair.GRADE and DIMCIM expose attribute diversity within one model, whereas DefaultShift tracks distributional change between models.
- Preference ranking evaluates candidates one image at a time, while DefaultShift-Select performs set-level distribution matching.The selection procedure targets the joint semantic composition of the released set rather than only individual image scores.
3 Methodology
DefaultShift audits whether an accelerated replacement preserves a declared reference model’s conditional semantic-attribute distributions. It uses matched sampling, closed-vocabulary labeling, bias-reduced distance measures, directional analysis, and reference-guided offline subset selection.
- Attribute Measurement: Repeated matched samples from reference and replacement models are labeled with closed semantic vocabularies to estimate conditional attribute distributions.The protocol supports evaluator swaps and treats unknown labels separately from valid-label histogram normalization.
- Distribution Comparison: Cell-level distribution discrepancies are aggregated to rank replacement recipes and localize probability-mass movement by category.Signed category-wise movement identifies gains and losses under replacement.
- Bias-Reduced Distribution Distance: Raw total variation is corrected with a matched reference noise floor because finite samples create positive empirical distance even under equal distributions.The adjusted score subtracts the estimated reference noise floor and retains negative values rather than clipping them.
- Inference: TVadj supports interpretable recipe ranking, while q = 64 signed cross-fit sensitivity is required for confirmatory effect-size claims.Cross-fit learns category signs on one half and evaluates signed contrasts on the other before swapping halves and averaging.
- DefaultShift-Select: DefaultShift-Select estimates reference target histograms and selects a quality-filtered replacement subset whose normalized attributes match those targets.Greedy updates approximate the discrete selection objective, while held-out splits keep reference labels used for evaluation unavailable to selection.
4 Experiments and Results
The experiments show that semantic default shift varies continuously by acceleration recipe, is primarily color-led, and remains detectable across robustness and human-validation tests. Reference-guided selection reduces this shift without material quality loss and improves balanced downstream performance.
- Benchmark and measurement: The benchmark audits 14 reference–replacement pairs across multiple acceleration families using 48 objects, four prompt templates, and paired seeds.The unified analysis contains 192 audit cells, with q = 32 used for screening and q = 64 for confirmatory evaluation.
- Benchmark and measurement: Color is the primary attribute, with Qwen2.5-VL-3B-Instruct as the main evaluator and BLIP and LLaVA used as evaluator swaps.Background, lighting, and viewpoint are also measured with adjusted and sensitivity divergences.
- Cross-recipe default shift: 0.303 to 0.054: unified q = 64 color TVadj spans NitroFusion-4 to LCM, while Turbo and DMD2 reach 0.254 and 0.235.The q = 32 and q = 64 rankings have Spearman correlation 0.996, and the discrepancy is color-led rather than uniform across attributes.
- Cross-recipe default shift: DMD2 and Turbo lose gray and gain warm mass, whereas FLUX moves oppositely despite all three belonging to the aggressive group.Category-wise direction distinguishes recipe fingerprints that a scalar magnitude alone cannot separate.
- Validation and scope: Human annotations reproduce the replacement ordering, with human and VLM rankings correlating at 0.943.Grouped human labels estimate color TV of 0.18 for Turbo and 0.10 for LCM, with paired difference 0.08.
- Calibration and downstream effects: DefaultShift-Select lowers human-measured shift by 10.3%, 11.7%, and 35.1% for Turbo, DMD2, and FLUX at a fourfold candidate budget.Noninferiority tests find no material quality loss; selection costs 0.652 of reference inference, preserving a 1.53-fold speedup.
- Calibration and downstream effects: Selected data recover 4.3 accuracy points and 7.5 worst-group points relative to uncalibrated replacement data under balanced evaluation.The gain on a natural-frequency test set is 1.1 points with an interval spanning −0.8 to 3.0.
- Validation and scope: The evidence is strongest for color and balanced downstream evaluation; background effects are smaller, viewpoint is inconclusive, and multilingual coverage is lower.Mechanism probes distinguish recipes but do not establish complete causal identification.
5 Conclusion
DefaultShift frames accelerated text-to-image deployment as a semantic preservation problem, comparing attribute distributions between reference and replacement models. Across 14 pairs, it identifies recipe-specific color shifts, while DefaultShift-Select reduces human-measured shift and improves balanced downstream performance.
- DefaultShift compares prompt-conditioned attribute distributions between declared reference and replacement models, accounting for finite-sample inflation and reporting shift magnitude and direction.
- Across 14 pairs, the confirmatory audit reveals color-led, recipe-specific shifts that quality, preference, and global diversity measures do not reliably characterize.
- Human validation, simulations, and stress tests support the audit's ordering across replacement pairs.
- DefaultShift-Select reduces human-measured shift across Turbo, DMD2, and FLUX while preserving quality and improving balanced downstream performance.
- The staged design connects low-cost screening, confirmatory inference, and a bounded offline response for joint decisions about speed, quality, and semantic preservation.
Supplementary Material
The supplement documents the benchmark's models, sampling design, semantic evaluation, estimators, and clustered uncertainty procedures. It standardizes evaluation across 48 objects, four templates, and repeated reference–replacement samples.
- The unified confirmatory benchmark uses 48 objects, four frozen prompt templates, and q = 64 queries per prompt cell.These choices produce 192 object-template cells per replacement pair.
- The supplement preserves the main paper's terminology and numerical conventions while providing complete configurations, benchmark results, diagnostics, and downstream evaluations.
- Each object-template cell is evaluated with repeated samples from reference and replacement models, with paired seeds when the interface permits.
- Color labels come from a frozen vocabulary, while unclear, unparsable, and safety-blocked outputs receive the unknown label ⊥.Unknown outputs remain in coverage statistics but are excluded from histogram normalization.
- The primary score adjusts empirical total variation using a matched reference split-half floor, while signed cross-fit estimates category signs and evaluates held-out contrasts.
- Confidence intervals cluster repeated measurements by object, with bootstrap repetition counts specified separately for the main audit, selection, and human comparisons.
C Complete Cross-Recipe Benchmark
The cross-recipe benchmark finds that acceleration-related semantic shifts vary in magnitude and direction across recipes and attributes. Screening remains useful for ranking, while q = 64 supports confirmatory effect claims.
- The screen and endpoint rankings have Spearman correlation 0.996, and TVadj correlates 0.994 with signed cross-fit TV.
- The SDXL aggressive-minus-mild contrast is 0.207 with interval [0.171, 0.243], while the FLUX family range is 0.021 with interval [-0.008, 0.051].
- The lighting aggressive-minus-mild contrast is 0.051 with interval [0.030, 0.073], whereas viewpoint is 0.016 with interval [-0.001, 0.033].These results support a color-led and background-supported pattern while leaving the general viewpoint effect inconclusive.
- NitroFusion-Realism is interpreted against SDXL as end-to-end deployment-chain drift because its direct teacher is DMD2.
- At q = 32, Turbo's absolute error relative to q = 96 is 0.044, falling to 0.017 at q = 64.The Turbo-minus-LCM separation is 0.250 at q = 32 and 0.291 at q = 64.
- The q = 32 screen remains useful for ranking, while q = 64 is the frozen minimum for confirmatory effects.
- Signed cross-fit is used for confirmatory inference because its null behavior is calibrated, whereas TVadj preserves ordering after removing most raw-TV inflation.
D.3 Unknown Labels and Label Noise
Unknown-label handling and label-noise analyses test how robust semantic-shift rankings are to evaluator coverage and errors. The results support coverage-adjusted background analysis but show that source-specific replacement errors can materially degrade ranking reliability.
- Treating unknown as an explicit background category raises background ranking correlation from 0.60 at q = 32 to 0.83 at q = 64.Coverage-adjusted q = 64 sensitivity reaches 0.97.
- Under symmetric label noise, ranking correlation is 0.996, 0.978, and 0.941 at noise rates of 0%, 10%, and 20%.
- A source-specific extra replacement error of 5 percentage points reduces correlation to 0.926 and induces mean spurious shift 0.031.
- At 10 percentage points of source-specific error, correlation falls to 0.824 and spurious shift rises to 0.064.
- Matched-resolution estimates remain 0.270 [0.252, 0.287] for Turbo at 512 pixels and 0.250 [0.221, 0.278] for DMD2 at 1024 pixels.
- A mixed model retains an aggressive-recipe coefficient of 0.118 with interval [0.079, 0.157] after accounting for teacher entropy, CFG, resolution, and sampling steps.
E Prompt and Metric Stress Tests
Stress tests show that DefaultShift’s recipe ordering is robust across metrics, prompt suites, evaluator changes, and human annotations, while standard quality and diversity measures do not recover semantic category movement. Specifying an attribute sharply reduces measured drift, supporting a prompt-conditional interpretation.
- E Prompt and Metric Stress Tests: 0.87 minimum cross-suite recipe-ranking correlation supports stable ordering across prompt suites.The interval is [0.71, 0.96].
- E Prompt and Metric Stress Tests: 0.131 aggressive coefficient after controlling explicit-attribute rate indicates residual prompt-suite separation.Its interval is [0.096, 0.166].
- E.1 Specified-Attribute Negative Control: All five specified-attribute intervals exclude zero, while the contraction supports prompt-conditional interpretation rather than universal evaluator distance.
- E Prompt and Metric Stress Tests: 0.987, 0.974, 0.934, and 0.994 rank correlations link TVadj with Jensen–Shannon, Hellinger, categorical MMD, and signed cross-fit TV.
- E Prompt and Metric Stress Tests: 0.048, 0.005, 0.000, 0.008, and 0.049 variance fractions from CLIP, aesthetic score, PickScore, LPIPS, and DINO spread do not recover category identity or direction.
- E Prompt and Metric Stress Tests: Explicitly naming color sharply reduces drift, and aggressive-minus-mild separation persists across five prompt suites.
- F.3 Six-Pair Human Expansion: 0.943 human–VLM recipe-ranking correlation reproduces the ordering in the 1,000-image audit, with a pooled aggressive-minus-mild contrast of 0.155.The interval for the contrast is [0.079, 0.231].
G Semantic Decomposition
Semantic decomposition shows that DefaultShift overlaps with diversity degradation but adds category-level information. Entropy-preserved subsets can still exhibit substantial category movement, while teacher-first hybrid inference reduces measured shift at the cost of changing the sampling trajectory and requiring reference computation.
- G Semantic Decomposition: 0.829 rank correlation links absolute normalized entropy change with TVadj, showing substantial overlap with diversity degradation.
- G Semantic Decomposition: 67 of 1,152 cells have entropy change no greater than 0.05 while raw TV is at least 0.30.
- G Semantic Decomposition: 0.093 Turbo mean TVadj falls to 0.026 when entropy is preserved and 0.013 when both entropy and Vendi diversity are preserved.The panel contains 100 objects, four templates, four attributes, and q = 32.
- G Semantic Decomposition: 0.054 mode-flip statistic remains in the entropy-preserved subset, supporting semantic localization rather than independence from diversity degradation.Its interval is [0.036, 0.074].
- G Semantic Decomposition: 0.201 to 0.141 DMD2 shift and 0.297 to 0.160 Turbo shift reductions occur under teacher-first hybrid inference.The probes use separate q = 32 panels; hybrid inference changes the sampling trajectory and requires reference-model computation.
H Recipe-Specific Mechanism Probes
Mechanism probes distinguish acceleration recipes through guidance-related and entropy-related associations, but the evidence does not identify a complete causal mechanism. An explicit under-guidance intervention lowers DMD2 shift while also reducing CLIP quality.
- H Recipe-Specific Mechanism Probes: 0.923 cosine similarity links DMD2’s shift direction to increasing teacher CFG, with an effective CFG near 15.
- H Recipe-Specific Mechanism Probes: 0.670 and 0.575 cosine similarities align LCM and Lightning with the teacher-CFG direction, versus 0.269 for Turbo.
- H Recipe-Specific Mechanism Probes: 0.596 Spearman correlation links Turbo’s object-level shift with reference color entropy, while peak-mode probability rises by 0.194.
- H Recipe-Specific Mechanism Probes: Mechanism probes remain associative unless an intervention changes the proposed factor, and these probes do not establish complete causal identification.
- H Recipe-Specific Mechanism Probes: 0.442 to 0.384 DMD2 TV reduction under guidance weight 1.0 to 0.5 accompanies CLIP decline from 0.250 to 0.225.At weight 0.3, CLIP falls to 0.133; the intervention does not provide a quality-preserving repair.
I DefaultShift-Select
DefaultShift-Select performs reference-guided offline set selection: it uses shared candidate pools and quality floors to match replacement subsets to reference semantic histograms. Across models, it reduces semantic shift, retains speedups and quality, and improves balanced downstream classification, while evidence remains bounded to evaluated settings.
- I DefaultShift-Select: 32 images selected from 128 candidates per object under the fourfold setting, with shared pools, held-out objects, and a common quality floor.
- I DefaultShift-Select: The selected subset reproduces one frozen FLUX object’s empirical reference histogram, while aggregate evidence uses the held-out 24-object evaluation.
- I DefaultShift-Select: 0.775, 0.774, and 0.865 of candidate-oracle reduction is reached by DefaultShift-Select for Turbo, DMD2, and FLUX.
- I DefaultShift-Select: 1.4, 2.1, and 1.9 percentage points separate integer optimization from the simpler selection rule for Turbo, DMD2, and FLUX.
- I DefaultShift-Select: 1.00 recipe-ordering Spearman correlation holds across primary VLM, second VLM, and human evaluation.Second-VLM intervals are [4.8, 19.6], [6.1, 21.4], and [28.4, 47.9].
- I DefaultShift-Select: 2.98 and 1.53 speedups remain for twofold and fourfold selection, whose normalized costs are 0.336 and 0.652.VLM evaluation contributes 0.31 of fourfold selection cost.
- I DefaultShift-Select: 4.3 balanced-accuracy points and 7.5 worst-color points improve over random replacement data under balanced evaluation.The corresponding intervals are [2.2, 6.4] and [3.4, 11.6].
- I DefaultShift-Select: −2.0 points is the selected-versus-reference balanced-accuracy difference, with interval [-4.1, 0.1].The evidence supports utility under balanced evaluation, not improvement across all deployment settings.