Source-linked AI summary

End-to-end weakly-supervised semantic alignment

Ignacio Rocco, Relja Arandjelović, Josef Sivic

arXiv:1712.06861v2cs.CVcs.LG

TL;DR

Dense semantic alignment is difficult under intra-class variation, viewpoint changes, and clutter, while prior methods lack end-to-end weakly supervised training. This paper introduces a differentiable soft-inlier alignment network trained from matching image pairs and achieves state-of-the-art performance on multiple benchmarks.

  • Problem

    Semantic alignment remains challenging under intra-class variation, viewpoint changes, and clutter, while existing methods lack jointly trained representations and alignment models or require ground-truth correspondences.

  • Method

    The paper develops an end-to-end convolutional alignment network trained from matching image pairs with a differentiable RANSAC-inspired soft-inlier scoring module.

  • Results

    75.8% overall PCK on PF-PASCAL sets a new state of the art, improving 3.6% over the best competitor, while the method also leads on multiple benchmarks.

  • Takeaways & Limitations

    The results show that weak supervision from matching image pairs can support state-of-the-art semantic alignment, including against methods using geometric training supervision.

  • Takeaways & Limitations

    Existing semantic alignment methods either train image representations and geometric models separately or require ground-truth correspondences that are difficult to obtain at scale.

Abstract

from arXiv · show

We tackle the task of semantic alignment where the goal is to compute dense semantic correspondence aligning two images depicting objects of the same category. This is a challenging task due to large intra-class variation, changes in viewpoint and background clutter. We present the following three principal contributions. First, we develop a convolutional neural network architecture for semantic alignment that is trainable in an end-to-end manner from weak image-level supervision in the form of matching image pairs. The outcome is that parameters are learnt from rich appearance variation present in different but semantically related images without the need for tedious manual annotation of correspondences at training time. Second, the main component of this architecture is a differentiable soft inlier scoring module, inspired by the RANSAC inlier scoring procedure, that computes the quality of the alignment based on only geometrically consistent correspondences thereby reducing the effect of background clutter. Third, we demonstrate that the proposed approach achieves state-of-the-art performance on multiple standard benchmarks for semantic alignment.

1. Introduction

The paper studies category-level semantic alignment: establishing dense correspondence between different objects of the same category despite intra-class variation, viewpoint changes, and background clutter. It addresses limitations of prior methods with an end-to-end CNN trained from matching image pairs and a differentiable soft inlier scoring module.

  • Problem: Semantic alignment establishes dense correspondence between different objects belonging to the same category.The task extends correspondence beyond identical objects or scenes to category-level matching.
  • Problem: The task is challenging because of large intra-class variation, viewpoint changes, and background clutter.
  • Limitations: Prior methods combine CNN-based image representations with geometric deformation models but do not train representation and alignment jointly end-to-end.
  • Limitations: Trainable geometric alignment models require ground-truth correspondences, which are difficult to obtain for diverse real images at large scale.
  • Contribution: The proposed CNN architecture is end-to-end trainable from weak supervision consisting of matching image pairs without ground-truth correspondences.Its differentiable soft inlier scoring module is inspired by RANSAC inlier scoring.

2. Related work

Prior semantic alignment methods progressed from hand-engineered descriptors and geometric models toward trainable components, but existing approaches remained limited by non-end-to-end processing or strong synthetic supervision. The proposed architecture trains both components end-to-end from weakly supervised real image pairs, avoiding manual correspondences and enabling richer appearance variation.

  • Hand-engineered methods: Early methods combined hand-engineered descriptors such as SIFT or HOG with hand-engineered alignment models, sometimes using fixed pretrained CNN descriptors.These methods did not train the descriptor or geometric model directly for semantic alignment.
  • Trainable descriptors: Trainable image descriptors were also combined with hand-engineered alignment models, preventing end-to-end trainability.This line of work improved descriptor learning but retained a non-trainable alignment component.
  • Trainable alignment: Recent methods jointly trained CNN descriptors and geometric alignment methods, but one required non-trainable fusion while another relied on strong synthetic correspondence supervision.Synthetic warping limited appearance variation in the training data.
  • Contributions: The proposed network makes both the descriptor and alignment model trainable end-to-end from weakly supervised real images with rich appearance variation.The training avoids manual ground-truth correspondences.
  • Contributions: The approach significantly improves alignment results and achieves state-of-the-art performance on several semantic-alignment datasets.This is presented as a contribution of the proposed architecture.

3. Weakly-supervised semantic alignment

The method trains semantic alignment end-to-end from matching image pairs without ground-truth geometric transformations. It uses a differentiable soft-inlier count to reward geometrically consistent feature matches and provide the learning signal.

  • Weak supervision: The model learns alignments from image-pair matching supervision, without requiring manual geometric transformation annotations.The network is driven to estimate good alignments because high soft-inlier counts are rewarded for matching image pairs.
  • Alignment network: A Siamese CNN performs feature extraction, pairwise feature matching, and geometric transformation estimation in three differentiable stages.Shared-weight fully convolutional branches produce L2-normalized dense local features, whose similarities are passed to a transformation regression CNN.
  • Feature matching: Normalized correlation computes all pairwise local-feature similarities while down-weighting ambiguous matches with multiple highly rated correspondences.This normalization follows the classical second-nearest-neighbour test of Lowe.
  • Soft-inlier count: The soft-inlier count extends RANSAC by summing match scores over all possible matches that satisfy an estimated geometric transformation’s inlier threshold.It replaces nondifferentiable sparse inlier counting with masked scores over the discretized four-dimensional match space.
  • Training objective: The soft-inlier count is differentiable with respect to transformation parameters and match scores, enabling back-propagation through geometric alignment and feature extraction.The objective maximizes c for matching image pairs, equivalently minimizing L = −c.

4. Evaluation and results

The evaluation uses a single weakly supervised semantic alignment network across PF-PASCAL, Caltech-101, and TSS, with comparisons against state-of-the-art methods. The method achieves leading results on PF-PASCAL, Caltech-101’s main metrics, and two TSS subsets, while qualitative results show robustness to viewpoint changes, clutter, and intra-class variation.

  • Implementation details: The network uses a ResNet-101 feature extractor cropped after conv4-23 and predicts an 18-degree-of-freedom thin-plate spline transformation.The model is trained end-to-end, including affine batch-normalization parameters, while keeping running averages fixed.
  • Benchmarks: Evaluation covers three standard benchmarks: PF-PASCAL, Caltech-101, and TSS.PF-PASCAL contains 1351 image pairs across 20 categories; Caltech-101 contains 1515 pairs across 101 categories; TSS contains 400 pairs across three subsets.
  • Assessing generalization: A single model trained on 700 PF-PASCAL training pairs without keypoint annotations is evaluated unchanged on PF-PASCAL, Caltech-101, and TSS.The weakly supervised objective uses only that each image pair should match, testing generalization across different categories and image types.
  • PF-PASCAL: 75.8% overall PCK on PF-PASCAL establishes the state of the art, improving 3.6% over the best competitor despite using weak supervision instead of bounding-box annotations.Weakly supervised fine-tuning also provides a 3.9% boost over synthetically trained ResNet-101+CNNGeo.
  • Caltech-101: The method surpasses state-of-the-art results on Caltech-101 for label-transfer accuracy and intersection-over-union, with a 2% gain over synthetically trained ResNet-101+CNNGeo.It does not attain state-of-the-art localization-error performance, whose method ordering conflicts with the other two metrics.
  • TSS and qualitative results: The method sets the state of the art on two TSS subsets, FG3DCar and JODS, although weak supervision yields only a modest gain over ResNet-101+CNNGeo.Qualitative results show alignment across prominent viewpoint changes, significant clutter, and large intra-class variations.

5. Conclusions

The paper presents a semantic image alignment network and training procedure inspired by RANSAC robust inlier scoring, requiring only matching image pairs for supervision. It achieves state-of-the-art results on multiple standard benchmarks, while handling multiple objects and non-matching pairs remains open.

  • Conclusions: The proposed network architecture and training procedure are inspired by robust inlier scoring used in RANSAC.The approach targets semantic image alignment.
  • Conclusions: Training requires supervision only in the form of matching image pairs.This avoids requiring geometric supervision at training time.
  • Conclusions: The method sets the new state-of-the-art on multiple standard semantic alignment benchmarks, even beating methods requiring geometric training supervision.The comparison concerns semantic alignment benchmark performance.
  • Conclusions: Handling multiple objects and non-matching image pairs remains an open challenge.These limitations are identified as unresolved directions.
  • Conclusions: The results open the possibility of learning powerful correspondence networks from large-scale datasets such as ImageNet.The conclusion points to scaling correspondence learning with large datasets.
Loading 1712.06861v2…