Source-linked AI summary

SOCO: Benchmarking Semantic Object Correspondence in Vision Foundation Models

Olaf Dünkel, Basavaraj Sunagad, Haoran Wang, David T. Hoffmann, Christian Theobalt, Adam Kortylewski

arXiv:2605.31597v3cs.CV

TL;DR

Structured object understanding is difficult to evaluate because existing semantic correspondence benchmarks use inconsistent protocols, limited part-level supervision, and insufficient distinctions among correspondence abilities. SOCO addresses this gap with a taxonomy-driven benchmark spanning 100 categories, over 1M correspondence pairs, and keypoint language descriptions. Experiments show weaknesses in object-relative geometry, cross-category transfer, and visual-reference matching, while SOC better predicts dense downstream performance than ImageNet classification.

  • Problem

    Existing semantic correspondence evaluations provide limited evidence about structured object understanding because they conflate correspondence abilities, omit cross-category transfer, and use ambiguous or limited keypoint annotations.

  • Method

    SOCO introduces a taxonomy-driven Semantic Object Correspondence formulation and benchmark with hierarchical keypoint annotations, cross-category pairs, and language descriptions across 100 categories.

  • Results

    Experiments show strong concept matching but weaker object-level geometry and cross-category transfer, LVLMs outperforming visual-reference matching on text-prompted localization, and SOC correlating more strongly than ImageNet classification with dense downstream tasks.

  • Takeaways & Limitations

    SOCO provides a unified benchmark for diagnosing fine-grained visual and multimodal representation quality through separately measurable correspondence failure modes.

  • Takeaways & Limitations

    Existing semantic correspondence annotations are constrained by geometric definitions, ambiguity under variability or symmetry, viewpoint-sensitive 2D projections, and omission of cross-category relationships.

Abstract

from arXiv · show

Measuring structured object understanding in vision foundation models remains challenging due to inconsistent evaluation protocols and limited part-level supervision. Semantic correspondence (SC) evaluates this capability by testing whether object parts can be matched across instances and categories under large variations in appearance, viewpoint, and geometry. To enable a systematic SC evaluation, we introduce SOCO, a new benchmark for Semantic Object Correspondence that introduces a taxonomy of correspondence types and provides consistent, functionally meaningful keypoint annotations across 100 categories and over 1M correspondence pairs. In addition, SOCO includes keypoint language descriptions, enabling the evaluation of large vision-language models (LVLMs) and their fine-grained part-level understanding. Comprehensive experiments reveal that (i) vision foundation backbones encode strong semantic structure but transfer correspondences poorly across related categories and only partially capture object-part position, (ii) LVLMs are stronger at text-prompted part localization than at visual-reference cross-image matching, exposing a gap between language-grounded localization and fine-grained visual correspondence, and (iii) correspondence performance predicts performance on dense downstream tasks, including segmentation, tracking, 3D pose estimation, and 3D detection, more strongly than ImageNet classification. Together, these findings position SOCO as a benchmark for structured, part-level representation quality in vision and multimodal foundation models.

1 Introduction

SOCO addresses limitations in semantic correspondence evaluation by separating concept recognition, object-relative identity, and cross-category transfer, then benchmarking these abilities with large-scale taxonomy-driven annotations. Experiments show distinct weaknesses in vision and multimodal models, while correspondence quality tracks dense downstream performance more strongly than ImageNet classification.

  • Motivation: Existing benchmarks do not adequately measure structured object understanding or distinguish local concept recognition from repeated-part identity and cross-category transfer.They largely report a single within-category score and omit semantically related categories.
  • Task formulation: SOC decomposes semantic correspondence into concept correspondence, semantic object correspondence, and cross-category correspondence.This taxonomy separates what part is matched, which instance of that part is intended, and whether the relation transfers across related categories.
  • Model findings: Vision foundation models recognize local concepts but show repeated-part confusion and limited cross-category abstraction.These weaknesses appear as CC→SOC and SOC→Cross-SOC performance drops.
  • Model findings: LVLMs perform better on text-prompted part localization than on visual-reference cross-image matching.This exposes a gap between language-grounded localization and fine-grained visual correspondence.
  • Dataset: SOCO provides taxonomy-driven keypoint annotations, language descriptions, and cross-category correspondence pairs across 100 categories.The benchmark combines visual correspondence evaluation with language-grounded analysis of multimodal models.

2 Related work

Prior semantic correspondence benchmarks established keypoint matching across object instances but remain limited in scale, category diversity, and coverage of fine-grained multimodal understanding. Foundation-model research extends zero-shot correspondence analysis, while LVLM evaluation has largely emphasized retrieval, captioning, and broad visual reasoning rather than fine-grained spatial matching.

  • Semantic correspondence benchmarks: Early semantic correspondence datasets were limited in scale and category diversity, while SPair-71k became a standard benchmark with 71k image pairs across 18 categories.Its class selection favors quadruped animals and vehicles.
  • Semantic correspondence benchmarks: Existing benchmark coverage omits important category types, including man-made objects in animal-focused datasets.This limits evaluation of general object-level understanding across diverse keypoint types.
  • Foundation models: Self-supervised and multimodal foundation-model features support zero-shot semantic correspondence but do not encode 3D part composition particularly well.This motivates using correspondence as a probe of representation quality.
  • Vision-language models: Vision-language model evaluation has focused mainly on retrieval and captioning rather than fine-grained spatial understanding.Modern LVLMs extend multimodal reasoning, but their evaluation remains dominated by broader tasks.

3 A Taxonomy for Semantic Correspondence

The paper defines Semantic Object Correspondence as a taxonomy that separates local concept semantics from object-relative position and supports correspondence within and across related categories.

  • Limitations of Current SC Keypoint Annotations: Existing semantic-correspondence benchmarks lack systematic hierarchical annotations and omit cross-category transfer, limiting structured evaluation.Their keypoints can be geometrically defined, ambiguous under variability or symmetry, viewpoint-sensitive, and internally inconsistent.
  • A Taxonomy of Semantic Object Correspondence: SOC separates local semantics from spatial configuration to test both concept matching and position-aware semantic object keypoint matching.This distinction addresses the difference between recognizing a part and identifying its object-relative instance.
  • A Taxonomy of Semantic Object Correspondence: A semantic concept is a uniquely identifiable location within an object part that is typically shared across instances of a category.Concepts capture local semantics and immediate functional context, such as a car door handle.
  • A Taxonomy of Semantic Object Correspondence: Semantic object keypoints inherit concepts but add positional attributes such as left, right, bottom, or rear, making correspondence unique and more demanding.Matching requires capturing both semantic identity and geometric placement in the object context.
  • A Taxonomy of Semantic Object Correspondence: A taxonomy spanning categories, super-categories, and shared concepts enables correspondence evaluation within categories and across related object classes.Cross-SOC can match shared concepts such as wheels across passenger cars, school buses, and tractors.
  • A Taxonomy of Semantic Object Correspondence: The formulation provides the conceptual foundation for SOCO's coherent and extensible semantic-correspondence annotation framework.It is designed to support systematic annotation and evaluation of semantic object correspondence.

4 The SOCO dataset

SOCO is a taxonomy-driven dataset for structured semantic correspondence, combining standardized keypoints, hierarchical cross-category pairs, and language descriptions across 100 categories.

  • Dataset Creation: SOCO provides a standardized semantically grounded keypoint schema, hierarchical cross-category image-keypoint pairs, broader category coverage, and language descriptions.The dataset is designed to address key limitations of prior correspondence benchmarks.
  • Dataset Creation: SOCO uses ImageNet images with pose metadata, single salient objects, and sufficiently large object size, supplemented by ImageNet3D and Animal3D annotations.Man-made objects rely on ImageNet3D 2D and 3D annotations, while animal categories use Animal3D keypoints.
  • Dataset Statistics: SOCO contains 100 categories across Transportation, Hand-held Objects, Furniture, and Animals, with 31, 20, 9, and 40 classes respectively.These four super-categories define the dataset's broad category distribution.
  • Dataset Creation: Each annotated keypoint receives a language description combining category, concept, and geometric attributes, such as the center of a bus's front-left wheel.Descriptions are generated from category, concept, keypoint position within the part, and part position within the object.
  • Dataset Statistics: SOCO annotates 40 images per category, yielding 4000 images with diverse viewpoints, shapes, and instance-level variation.Within-category SOC pairs require at least three shared semantic keypoints.
  • Dataset Statistics: Model performance drops across CC, SOC, and Cross-SOC as geometric awareness and semantic abstraction demands increase, while rankings change for SOC-geo.The results indicate that models encode object geometry and semantics to different extents.
  • Dataset Statistics: Around 73k SOC pairs contain approximately 560k keypoint correspondences, while Cross-SOC generation produces around 1.3M cross-category correspondence pairs.CC uses the same within-category pairs, and the pairing regimes increase in difficulty from concepts to keypoints and categories.

5 Experiments

Experiments evaluate vision and multimodal foundation models on SOCO’s correspondence subtasks using zero-shot feature matching and PCK, then relate SOC to downstream-task performance. Results show a persistent gap between semantic concept matching and geometric part awareness, weaker cross-category transfer, LVLM difficulty with visual-reference matching, and stronger SOC associations with dense tasks than ImageNet kNN.

  • Evaluation setup: SOCO evaluates concept correspondence (CC), semantic object correspondence (SOC), and Cross-SOC, using zero-shot nearest-feature matching and PCK at α = 0.1.The evaluation also isolates geometric awareness with SOC-geo, which tests repeated keypoint positions within the same category.
  • Vision foundation models: All evaluated models show substantial CC→SOC drops, indicating that strong semantic representations do not guarantee geometric part awareness.Even strong models such as DINOv2 struggle to disambiguate repeated parts, with further degradation in Cross-SOC.
  • Vision foundation models: The CC→SOC gap is largest for Furniture and Transportation, where repeated symmetric parts make object-relative identity difficult to distinguish.For DINOv2 on Furniture, SOC is 45.5 versus CC 77.5.
  • Vision foundation models: Dense self-supervised objectives produce stronger semantic correspondence representations than global alignment objectives, while DINO-family models perform particularly well on CC.Reconstruction-based models generally perform poorly, although scaling reconstruction objectives improves dense correspondence for PIXIO.
  • LVLMs: LVLMs localize text-described parts more accurately than they transfer visual references across images.Adding keypoint descriptions improves performance over purely visual queries, and description-only queries outperform visual-reference queries.
  • LVLMs: 81.0% for DINOv2 adapted to the 4-choice setting exceeds the best evaluated LVLM at 54.0%.The comparison indicates that current LVLMs remain limited on fine-grained cross-image correspondence despite stronger text-guided localization.

6 Conclusion

The paper introduces SOC and SOCO to separate semantic and geometric correspondence abilities in a large benchmark with hierarchical, cross-category, and language-grounded annotations. Evaluations show distinct weaknesses across vision and multimodal models, while SOC aligns more strongly with dense vision tasks than ImageNet kNN.

  • Conclusion: SOC separates geometric matching from semantic object-level understanding by explicitly modeling object parts in relation to overall object structure.This formulation makes correspondence capabilities more clearly distinguishable than a single undifferentiated score.
  • Conclusion: SOCO provides hierarchical part annotations, cross-category correspondences, and language descriptions for evaluating vision and multimodal foundation models.The benchmark is designed to address limitations of existing correspondence datasets.
  • Conclusion: SOCO exposes separate failure modes for repeated-part disambiguation, cross-category transfer, and visual-reference versus language-grounded localization.These gaps correspond respectively to CC→SOC, SOC→Cross-SOC, and Vis. versus Desc. comparisons.
  • Conclusion: SOC performance correlates more strongly with dense vision tasks than ImageNet kNN, supporting SOC as a zero-shot diagnostic of representation quality.The conclusion frames SOCO as a testbed for structured, part-level visual and multimodal understanding.

Supplementary Material

The supplementary material extends SOCO analysis across categories, evaluation protocols, thresholds, viewpoints, and additional models. These results show category- and threshold-dependent performance, viewpoint-sensitive geometric matching, and differences between zero-shot and supervised transfer across datasets.

  • Scope: Supplementary experiments evaluate SOCO subsets, PCK levels, additional models, and per-category results.The material includes analyses of supercategories, complementary protocols, geometric awareness, thresholds, viewpoints, and further models.
  • Category results: Furniture and Transportation show large CC→SOC drops because they contain locally similar or repeated parts such as chair legs and vehicle wheels.DINOv3 outperforms DINOv2 on Furniture SOC, indicating better geometric-position encoding in that comparison.
  • Evaluation protocols: Alternative protocols include window softargmax, a supervised shared linear probe, and fixed patch counts.Window softargmax consistently improves results modestly, while the linear probe improves performance substantially across evaluated settings.
  • Threshold analysis: SOC performance substantially drops at smaller PCK thresholds, with a smaller decline for SD than for other models.The supplementary results report threshold sensitivity under pair-averaged and per-keypoint evaluations.
  • Additional models: Supervised models trained on SPair-71k show lower performance on SOCO than on the larger SPair-71k test set, indicating that SOCO contains different categories.Weak supervision improves semantic correspondence on SOCO, while larger model variants typically outperform base models.

A.7 Category-Specific Results

Category-specific supplementary results show that SOC–CC gaps and threshold sensitivity vary by category. Viewpoint analysis further separates stable concept matching from geometry-sensitive object correspondence.

  • Category-specific results: SOC–CC gaps depend substantially on the category, and lowering the PCK threshold affects categories differently.The table reports DINOv2-B results across multiple PCK thresholds.
  • Annotations: The supplementary figures illustrate both category diversity and keypoints that are unique or shared across semantic concepts.These annotations distinguish semantic concepts from more specific instance-level keypoint realizations.

B.1 Limitations of Existing SC Annotations

Existing semantic-correspondence annotations suffer from weak semantic grounding, ambiguity, and inconsistent definitions, while model rankings vary across datasets and category compositions.

  • Existing SC keypoints can lack semantic grounding, creating ambiguous definitions for categories with substantial intra-class variability.
  • Symmetric objects produce nonunique keypoints when definitions rely on 2D projections.
  • Some existing benchmarks use inconsistent keypoint definitions across instances, including for trains.
  • Adding man-made objects changes rankings for the whole SOCO dataset.
  • DINOv2 remains the best-performing model across MISC210K, SPair-71k, and AP-10K, but other model rankings vary.

C.2 Other Downstream Tasks

The study compares SOCO correspondence with ImageNet kNN classification and several downstream tasks, finding stronger alignment between SOCO and many downstream capabilities.

  • SOCO performance correlates with downstream segmentation, detection, pose, multi-view correspondence, and tracking performance.
  • 0.943 Pearson r links SOCO performance with multi-view correspondence, compared with 0.266 for ImageNet kNN.
  • 0.907 Pearson r links SOCO performance with tracking, compared with 0.286 for ImageNet kNN.
  • 0.892 Pearson r links SOCO performance with 3D detection, compared with 0.393 for ImageNet kNN.
  • SOCO and ImageNet kNN are both negatively correlated with surface-normal and depth metrics in the reported analysis.

D.3 Details on Other Downstream Tasks

The downstream evaluation suite probes frozen visual backbones across correspondence, classification, dense prediction, geometry, tracking, and 3D detection using zero-shot matching or lightweight probes.

  • The unified suite spans monocular geometry, multi-view correspondence, semantic segmentation, tracking, classification, and 3D detection.
  • Correspondence: Zero-shot correspondence uses nearest-neighbor feature matching, with PCK@0.1 reported for semantic and multiview correspondence tasks.
  • Classification: ImageNet classification uses k-nearest neighbors, selecting k for best validation accuracy and using CLS tokens or averaged dense tokens.
  • Backbones: The evaluated models include self-supervised, vision-language, generative, and 3D-aware backbones, with architecture, supervision, and pretraining data recorded.
  • Probed tasks: Frozen backbone features support lightweight probes for semantic segmentation, depth, surface normals, 3D pose, and 3D detection.
  • LVLM evaluation: Qwen2.5-VL-7B-Instruct performs best with arrow markers at 30.0 mean and red markers at 30.0 mean in the visual-prompt study.

F Limitations

SOCO’s evaluation is bounded by sparse keypoints, curated image sources, template-based language prompts, zero-shot matching, and a defined cross-category taxonomy.

  • Sparse keypoints diagnose structured part-level understanding but do not support dense pixel-wise correspondence evaluation.
  • ImageNet3D and Animal3D sourcing biases SOCO toward salient curated views and limits out-of-distribution evaluation.
  • Template-based keypoint descriptions constrain the prompted LVLM setting, and richer language could improve performance further.
  • Nearest-neighbor matching is intentionally simple and provides a lower bound on performance achievable with supervised adaptation.
  • Cross-category correspondences remain limited to SOCO’s concept hierarchy, excluding broader functional analogies across distant categories.
Loading 2605.31597v3…