Source-linked AI summary
FASH-iCNN: Making Editorial Fashion Identity Inspectable Through Multimodal CNN Probing
Morayo Danielle Adeyemi, Ryan A. Rossi, Franck Dernoncourt
TL;DR
Fashion AI can encode houses, editors, and historical moments without making that cultural influence visible. FASH-iCNN probes garment and supplementary inputs across a large Vogue runway corpus, recovering house, era, and color-tradition signals while identifying the visual channels carrying them. Its results show that editorial culture is recoverable and can be made inspectable, though evaluation remains bounded to held-out Vogue runway data and related scope constraints.
Problem
Fashion AI systems encode specific houses, editors, and historical moments without disclosing the editorial traditions shaping their guidance.
Method
FASH-iCNN trains a multimodal system on annotated Vogue runway imagery to recover house identity, temporal era, and color tradition while testing garment, face, and metadata inputs.
Results
Garment crops recover house identity, decade, and specific year, while texture and luminance carry more house-identity signal than color or shape.
Takeaways & Limitations
Making editorial reference frames visible is a viable design principle for multimodal fashion systems, grounding outputs in specific, nameable traditions.
Takeaways & Limitations
Evaluation uses held-out Vogue runway data, leaving non-editorial, nonluxury, and non-Western fashion contexts untested.
Abstract
from arXiv · showhide
Fashion AI systems routinely encode the aesthetic logic of specific houses, editors, and historical moments without disclosing it. We present FASH-iCNN, a multimodal system trained on 87,547 Vogue runway images across 15 fashion houses spanning 1991-2024 that makes this cultural logic inspectable. Given a photograph of a garment, the system recovers which house produced it, which era it belongs to, and which color tradition it reflects. A clothing-only model identifies the fashion house at 78.2% top-1 across 14 houses, the decade at 88.6% top-1, and the specific year at 58.3% top-1 across 34 years with a mean error of just 2.2 years. Probing which visual channels carry this signal reveals a sharp dissociation: removing color costs only 10.6pp of house identity accuracy, while removing texture costs 37.6pp, establishing texture and luminance as the primary carriers of editorial identity. FASH-iCNN treats editorial culture as the signal rather than background noise, identifying which houses, eras, and color traditions shaped each output so that users can see not just what the system predicts but which houses, editors, and historical moments are encoded in that prediction.
1 Introduction
FASH-iCNN makes the cultural authorship embedded in fashion imagery inspectable by recovering houses, eras, and color traditions from garment photographs. It tests which multimodal inputs add signal and frames editorial culture as a recoverable component of fashion AI outputs.
- FASH-iCNN treats cultural authorship as information users can inspect, question, and recognize in fashion guidance.
- The system tests whether multimodal inputs contribute meaningful signal beyond what the garment crop alone encodes.
- FASH-iCNN exposes the houses, eras, and color traditions shaping each prediction instead of leaving editorial lineage invisible.The system accepts a garment photograph and can optionally incorporate face, designer, season, and year signals.
- Garment appearance encodes editorial culture as a structured, recoverable signal across house identity, temporal era, and color regime.
2 Related Work
Prior fashion AI commonly uses behavioral signals and treats editorial metadata as filtering information rather than as a traceable source of aesthetic taste. FASH-iCNN instead grounds prediction in named editorial precedents while probing when supplementary inputs add information beyond the primary visual stream.
- Computational fashion systems and taste-based recommendation: Prior fashion systems address compatibility, attributes, retrieval, and recommendation, but commonly learn from purchase history, ratings, and clicks.
- Computational fashion systems and taste-based recommendation: FASH-iCNN uses designer, collection, season, and year metadata as primary signals encoding aesthetic taste rather than merely as filtering tags.
- Multimodal fusion with supplementary inputs: Multimodal fusion raises the design question of when supplementary inputs contribute substantively versus duplicating information already present in the primary stream.
- Hierarchical and perceptually grounded color prediction: FASH-iCNN operationalizes color prediction through a Berlin–Kay to CSS to LAB hierarchy spanning named color families, finer labels, and perceptual regression.
3 Dataset and Corpus
The corpus treats Vogue runway imagery as culturally structured data produced through coordinated editorial decisions within Western luxury fashion. FASH-iCNN builds annotated garment, face, color, skin-tone, designer, season, and year data while restricting color experiments to a chromatic subset.
- Vogue runway images encode coordinated aesthetic decisions from creative direction, casting, styling, and editorial selection.
- The corpus represents one Western luxury fashion tradition rather than fashion universally.
- 87,547 Vogue runway images span 15 fashion houses from 1991–2024, with 84,596 remaining after quality filtering.
- 65,541 garment crops are produced using clothing-region extraction, while annotations include dominant CIELAB colors, Berlin–Kay terms, CSS colors, skin tone, designer, season, and year.
- 24,500 chromatic images are used for color prediction after removing black- and gray-dominant images from the with-face-crop set.The corpus is 68.3% low-saturation, and white remains a chromatic class because it functions as a deliberate stylistic choice in editorial fashion.
4 Multimodal Color Prediction System
FASH-iCNN combines clothing and optional face inputs to predict editorial color structure, while probing which visual information supports color and designer identity. The experiments show that dominant color is recoverable, but richer palette structure remains difficult and designer identity depends strongly on texture and luminance.
- System architecture: Independent EfficientNet-B0 streams process clothing and optional face crops before feature concatenation and classification.The model uses 224×224 RGB inputs and a two-layer prediction head.
- Per-house color prediction: 93.4% BK9 top-1 is achieved for Calvin Klein Collection, compared with 91.0% for Chanel and 82.3% for Alexander McQueen.These are per-house constrained models trained and evaluated within each house’s chromatic subset.
- Visual abstraction analysis: 78.2% top-1 identifies the fashion house from full-color garment appearance, while removing color costs only 10.6pp and removing texture causes a 37.6pp drop.Edge maps and silhouettes perform nearly identically at 30.7% and 30.0%, indicating limited added value from filled shape beyond contour.
- Temporal identity: 88.6% top-1 decade accuracy and 58.3% top-1 year accuracy show that clothing crops encode temporal identity across 1991–2024.Year prediction falls within two years 73.2% of the time, with a mean absolute error of 2.2 years across 34 years.
- Multimodal compensation: Face input changes color prediction by −0.6pp with full-color garments but lifts it by +20.8pp on silhouettes and +20.5pp on edge maps.The face modality contributes more when garment information is sparse.
- Single-color vs. multi-slot prediction: A dominant-color swatch reduces CSS top-1 accuracy by only 0.5pp, whereas later palette slots degrade sharply and reach approximately 17 ΔE00 median error by slot 4.Multi-label prediction reaches precision@1 of 0.858 but loses ordering information, so multi-color palette prediction remains open on this corpus.
5 Discussion
FASH-iCNN presents editorial fashion information at multiple levels, from house and decade to color tradition and perceptual coordinate. Its scope remains constrained by the Vogue-centered corpus, untested cross-house color generalization, and limited palette prediction.
- Interaction Implications: A layered output connects predicted house and decade with color tradition, its Berlin–Kay family, CSS named hue, and CIELAB coordinate.The system serves broad provenance questions, color-lineage exploration, and precise styling or design decisions.
- Interaction Implications: Single-dominant-color prediction reflects the learnability boundary rather than fabricating richer palette outputs.Palette-level prediction beyond the dominant slot remains an open problem.
- Portability and Future Work: Retraining on non-Western archives or regional dress collections would preserve the technical structure while producing culturally distinct models.Corpus diversification beyond Vogue’s Western, luxury-centric coverage is identified as an immediate extension.
- Limitations: Cross-house color generalization is untested because per-designer constrained color models are trained and evaluated within one house.This boundary limits conclusions about color prediction across fashion houses.
- Limitations: Evaluation uses held-out Vogue runway data, leaving non-editorial, nonluxury, and non-Western fashion contexts untested.The corpus represents one editorial tradition rather than fashion universally.
6 Conclusion
FASH-iCNN treats garment appearance as a culturally structured signal and makes the editorial reference frame behind fashion predictions inspectable. Its outputs ground predictions in specific, nameable traditions rather than leaving that cultural structure invisible.
- Conclusion: FASH-iCNN demonstrates that garment appearance supports independent recovery of house identity, temporal era, and color regime.The conclusion frames cultural structure as a recoverable signal for multimodal fashion systems.
- Conclusion: The paper positions cultural transparency as a viable design principle for multimodal fashion systems.This conclusion follows from making editorial structure visible in system outputs.
- Conclusion: The system makes its cultural reference frame inspectable rather than invisible.Outputs are grounded in a specific, nameable editorial tradition.