Source-linked AI summary
Sketch2Inspire: Structure-Sensitive Evaluation for Product Retrieval
Ge Kong
TL;DR
Early-stage product-design retrieval needs references matching both semantic intent and rough structure, while existing resources rarely separate category relevance from within-category structural fit. Sketch2Inspire addresses this gap with aligned text and sketch-proxy queries, complementary relevance protocols, and a lightweight CLIP-family reference system. Late fusion ranks highest across broad, automatic structure-sensitive, and human-graded evaluations, but the gain depends on how relevance is defined.
Problem
Existing product-image resources and generic retrieval benchmarks rarely distinguish broad category retrieval from within-category structural fit for early-stage product-design search.
Method
Sketch2Inspire builds a curated ABO-based resource with aligned text, sketch-proxy, and fused queries, evaluating transparent CLIP-family retrieval modes under complementary relevance protocols.
Results
Late fusion obtains the highest nDCG under broad, automatic structure-sensitive, and human-graded relevance, with the gain depending on the relevance definition.
Takeaways & Limitations
Sketch2Inspire provides a controlled diagnostic resource for measuring modality contribution and evaluating structure-sensitive product retrieval.
Takeaways & Limitations
The edge-based sketch proxies are more visually coupled to product photographs than freehand ideation sketches, limiting claims about real sketch abstraction and ambiguity.
Abstract
from arXiv · showhide
Early-stage product design retrieval often requires more than category recognition: designers may need reference examples that match both a short semantic intent and a rough structural cue. Existing product-image resources and generic image--text retrieval benchmarks rarely separate category retrieval from within-category structural fit. We present Sketch2Inspire, built from a curated subset of Amazon Berkeley Objects with aligned text queries, edge-based sketch-proxy queries, and fused text--sketch queries. The resource separates broad category-level retrieval from structure-sensitive within-category retrieval and includes a human-graded reference protocol for calibration. We evaluate a lightweight reference system based on pretrained CLIP-family encoders, comparing text-only retrieval, sketch-only retrieval, weighted late fusion, and text-first reranking without updating model weights. Under broad relevance, late fusion obtains the highest score (nDCG = 0.9962). Under automatic structure-sensitive relevance, late fusion again obtains the highest score (nDCG = 0.7015), exceeding text-only retrieval (nDCG = 0.5912). In the human-graded results, late fusion obtains the highest nDCG@10 (0.9133), while text-only retrieval ranks second (0.9030). These results show that the retrieval gain from multimodal input depends on how relevance is defined. Sketch2Inspire therefore provides a diagnostic resource for evaluating modality contribution and supports the development of structure-aware product-retrieval protocols with independent human annotation.
1 Introduction
Sketch2Inspire targets early-stage product-design retrieval, where useful references must match both semantic category and structural cues. It introduces a traceable multimodal resource and evaluation protocol, with late fusion leading across relevance definitions.
- Motivation: Early-stage design retrieval requires references matching product category alongside shapes, silhouettes, proportions, and visual analogies.
- Motivation: Existing product datasets and retrieval protocols rarely align sketch, text, and product queries or distinguish within-category structural fit.
- Resource and protocol: Sketch2Inspire constructs DesignABO with broad category-level, structure-sensitive, and human-graded evaluation views.
- Reference system: The reference system compares text-only, sketch-only, weighted late fusion, and text-first reranking using pretrained CLIP-family encoders without model training changes.
- Findings: Late fusion achieves the highest nDCG across broad, automatic structure-sensitive, and human-graded relevance, although gains depend on the relevance definition.
2 Related Work
Prior work provides strong product, fashion, vision-language, and sketch-retrieval foundations, but these resources do not directly isolate design-oriented structural retrieval. Sketch2Inspire instead uses transparent score-level fusion to test whether sketch cues alter rankings beyond text matching.
- Product and design retrieval: ABO and fashion-oriented benchmarks provide product images, metadata, and retrieval capabilities but are not designed for early-stage product-design inspiration search.
- Vision-language and composed retrieval: Vision-language and composed-retrieval research supplies shared representations and multimodal query strategies across several architectures and sketch-driven settings.
- Reference comparison: The experiments use common gallery embeddings, query sets, and relevance protocols so ranking changes can be attributed to text, sketch, or fused evidence.
- Scope: Stronger retrieval architectures remain important comparisons for absolute performance, while this study isolates whether evaluation protocols reveal sketch-driven ranking changes.
3 Evaluation Resource and Reference Retrieval
Sketch2Inspire builds a traceable product-retrieval resource with broad and structure-sensitive evaluation settings, then probes modality contribution using fixed pretrained representations and multiple ranking modes.
- Evaluation Resource: DesignABO uses curated product subsets with fixed galleries, category labels, query records, and relevance lists for traceable evaluation.The Expanded Setting contains 9,000 images from 300 categories, while the Structure-Sensitive Setting contains 3,900 images from 260 categories.
- Evaluation Resource: 3900 images from 260 categories form the Structure-Sensitive Setting, which retains within-category variation in silhouette, aspect ratio, part layout, occupancy, and contour complexity.Fan and humidifier are excluded because their curated examples were less informative for silhouette-driven disambiguation.
- Query Construction: Edge-based sketch proxies are generated from anchor product images, aligned with text queries, and stripped of the anchor from candidate rankings to prevent direct self-match leakage.The proxy construction uses image preprocessing and an edge operator, making the structural cue reproducible but narrower than real hand-drawn sketches.
- Reference Retrieval: The reference system compares text-only, sketch-only, weighted late fusion, and text-first reranking using pretrained CLIP-family image and text encoders without updating model weights.Shared gallery embeddings, query sets, and relevance protocols make ranking changes attributable to modality availability and relevance definition.
- Relevance Protocols: Broad relevance treats non-anchor items from the same category as relevant, whereas structure-sensitive relevance selects focused within-category matches using shape-weighted descriptors.The automatic structure-sensitive rule is a controlled diagnostic rather than independent ground-truth annotation because it includes heuristic descriptors and a CLIP-family image-space component.
4 Experiments
Sketch2Inspire evaluates product retrieval under broad category relevance, structure-sensitive relevance, and a separate human-graded protocol. Across these settings, late fusion performs best, while the value of sketch cues becomes most visible when relevance rewards within-category structure.
- Settings and Metrics: The experiments compare aligned automatic and human-graded protocols, with the human-graded reference check reported separately from automatic structure-sensitive evaluation.The automatic settings report Recall@5, Recall@10, mAP, and nDCG; the human protocol reports nDCG@10, mAP@10, and Recall@5.
- Broad Category-Level Retrieval: 0.9962 nDCG: late fusion achieves the highest score under broad category-level relevance.Text-only retrieval remains a strong baseline, with Recall@10 = 0.4983 and mAP = 0.9839.
- Broad Category-Level Retrieval: Under broad same-category scoring, high mAP and nDCG mainly reflect strong category clustering rather than solved inspiration retrieval.This establishes a coarse-relevance baseline that does not test the structural aspect of design search.
- Structure-Sensitive Retrieval: Under fine-grained structural relevance, late fusion reaches Recall@10 = 0.9000, mAP = 0.5290, and nDCG = 0.7015, the highest nDCG and mAP.Text-first reranking is also competitive, testing whether sketch information helps after text constrains the semantic candidate pool.
- Structure-Sensitive Retrieval: The category-versus-structure contrast shows that sketch-only retrieval improves over text-only on Recall@5 and nDCG only under fine-grained structural relevance.Adding sketch information does not materially improve over text-only retrieval under category-level relevance.
- Structure-Sensitive Retrieval: α = 0.40: late-fusion performance peaks in the middle range, indicating that sketch evidence complements rather than replaces text under structural relevance.The sweep is a controlled probe of useful structural evidence, not a learned weighting strategy.
- Human-Graded Evidence: 0.9133 nDCG@10: late fusion ranks highest in the human-graded protocol, while text-only retrieval ranks second at 0.9030.The manual grades provide a calibration separate from automatic structure-sensitive relevance.
- Human-Graded Evidence: The human-graded results show that conclusions about the same retrieval system depend on whether relevance is broad, automatically structure-sensitive, or manually graded.The protocol supports separating structure-aware evaluation from category-level retrieval while indicating that stronger human validation is needed for robust fusion claims.
5 Discussion
The experiments establish that sketch cues matter most when relevance measures within-category structural fit rather than category membership alone. Sketch2Inspire therefore frames multimodal retrieval as a diagnostic evaluation problem whose conclusions depend on the relevance protocol.
- Discussion: Late fusion achieves the highest nDCG under broad category-level, automatic structure-sensitive, and human-graded relevance, but the gain depends on the relevance definition.The results support comparing modalities across complementary evaluation protocols rather than treating one score as definitive.
- Discussion: Under fine-grained structural relevance, sketch cues improve ranking and late fusion achieves the highest automatic nDCG and mAP.This setting exposes ranking differences that category-level relevance can obscure.
- Discussion: Under category-level relevance, adding sketch information does not materially improve over text-only retrieval.Broad same-category scoring can make text-only category matching appear sufficient even when structural fit is not measured.
- Discussion: The resource measures modality contribution under controlled relevance definitions and is intentionally diagnostic rather than definitive.Its scope is to identify when multimodal input improves ranking beyond category-level text matching and when sketch signals provide design-relevant structure.
6 Limitations
The study’s conclusions are bounded by proxy sketches, coupled automatic relevance heuristics, and a transparent zero-shot reference method. These constraints limit claims about real designer sketches, independent ground truth, and absolute retrieval competitiveness.
- Sketch realism: Edge-based sketches generated from product photos may overestimate alignment with gallery images compared with freehand ideation sketches.The proxy design improves control and reproducibility but limits claims about roughness, ambiguity, and abstraction in real sketches.
- Automatic relevance coupling: The automatic structure-sensitive protocol is not independent ground truth because it couples CLIP similarity to the reference retriever’s model family.Its shape descriptors are also hand-designed proxies, while human grading is limited to a shared candidate pool.
- Method scope: The zero-shot retrieval method supports modality-attribution claims within this setting, not absolute competitiveness against stronger retrieval architectures.Future comparisons should add stronger composed-retrieval models, diffusion-assisted matching, and larger-scale multimodal backbones under the same protocols.
7 Conclusion
Sketch2Inspire tests when sketch input contributes beyond text-only category matching in product-inspiration retrieval. Late fusion leads across the reported relevance settings, but the relevance-dependent results remain a controlled foundation rather than a universal retrieval claim.
- Conclusion: Late fusion achieves the highest nDCG under broad category-level relevance and the highest automatic scores when relevance rewards within-category structure.The human-graded protocol also ranks late fusion first on nDCG@10, with text-only retrieval second.
- Conclusion: The fusion gain is visible under ranking-sensitive relevance but should be interpreted as relevance-dependent rather than universal retrieval improvement.This conclusion follows from the differing roles of category-level and structure-sensitive evaluation.
- Conclusion: Sketch2Inspire contributes a controlled diagnostic setting for measuring modality contribution under explicit relevance definitions.The current galleries and edge-based sketch proxies motivate larger resources with real designer sketches, independent annotation, broader baselines, and user-centred validation.