Source-linked AI summary
Learning SO(3) Equivariant Representations with Spherical CNNs
Carlos Esteves, Christine Allen-Blanchette, Ameesh Makadia, Kostas Daniilidis
TL;DR
The paper studies 3D data for alignment, retrieval, and classification, where existing volumetric and point-cloud representations provide translation and scale invariance. It introduces Spherical CNNs using spherical convolutions and spectral-domain operations, achieving near-state-of-the-art performance with lower network capacity and small input sizes.
Problem
3D data analysis for alignment, retrieval, and classification requires representations beyond the translation and scale invariance provided by volumetric and point-cloud approaches.
Method
Spherical CNNs evaluate spherical convolutions in the spectral domain and use spectral-domain pooling and filter parameterization.
Results
Near state of the art performance is achieved on SHREC’17 and ModelNet40 with much lower network capacity.
Takeaways & Limitations
The model can handle arbitrary input orientations with relatively few parameters and small input sizes across classification, retrieval, and alignment tasks.
Takeaways & Limitations
The spherical representation is suitable for star-shaped objects, but the paper does not check whether this condition holds in practice, even when the representation may be ambiguous or non-invertible.
Abstract
from arXiv · showhide
We address the problem of 3D rotation equivariance in convolutional neural networks. 3D rotations have been a challenging nuisance in 3D classification tasks requiring higher capacity and extended data augmentation in order to tackle it. We model 3D data with multi-valued spherical functions and we propose a novel spherical convolutional network that implements exact convolutions on the sphere by realizing them in the spherical harmonic domain. Resulting filters have local symmetry and are localized by enforcing smooth spectra. We apply a novel pooling on the spectral domain and our operations are independent of the underlying spherical resolution throughout the network. We show that networks with much lower capacity and without requiring data augmentation can exhibit performance comparable to the state of the art in standard retrieval and classification benchmarks.
1 Introduction
The paper targets the persistent difficulty of handling arbitrary 3D rotations by introducing an SO(3)-equivariant spherical CNN. It combines exact spherical convolutions, spectral pooling, and resolution-independent parameterization to achieve competitive performance with lower capacity.
- Motivation: Equivariant networks retain information about group actions across layers, linking feature transformations directly to spatial transformations of the input.
- Motivation: 3D rotations remain challenging for current alignment, retrieval, and classification approaches, especially under arbitrary orientations.Conventional methods suffer significant classification drops when arbitrary rotations are introduced.
- Approach: The proposed network models 3D data as R^n-valued spherical functions and uses exact spherical convolutions with zonal filters computed in the spherical harmonic domain.The paper distinguishes spherical-output convolution from correlation with outputs on SO(3).
- Approach: Spectral pooling preserves equivariance, while weighted averaging pooling uses weights proportional to spherical cell area.The rectifying nonlinearity is the only operation that requires returning to the spatial domain.
- Evaluation: Experiments on SHREC’17 and ModelNet40 cover retrieval, classification, and alignment, achieving near-state-of-the-art performance with much lower network capacity.
- Approach: Smooth spectral parameterization localizes filters while keeping the number of weights independent of spatial resolution.Weights are learned at a few anchor frequencies and interpolated between them.
2 Related work
Related work establishes two broad routes to equivariance and contrasts spherical CNNs with graph, volumetric, and SO(3)-based approaches. The paper emphasizes smaller filters, faster spherical convolutions, smooth spectral localization, and alternative pooling schemes.
- Group equivariance: Equivariant CNNs obtain group equivariance either by constraining filter structure or by using an equivariant filter orbit.
- Graph methods: Graph convolutional methods learn filters on irregular structured graphs, whereas this work explicitly constructs equivariant and invariant representations for spherical 3D data under rotations.
- Spherical methods: Compared with a concurrent spherical-correlation approach mapping inputs to SO(3), spherical convolutions use one fewer filter-map dimension and are potentially one order of magnitude faster.
- Spherical methods: The paper adds smooth spectral parameterization for better spherical receptive-field localization and uses either spectral low-pass or spatial weighted-average pooling.
- Volumetric methods: Volumetric CNNs adapt 2D architectures with 3D filters but require substantial computation at basic voxel resolutions and higher capacity.
- Volumetric methods: Earlier volumetric models include fully volumetric networks and subvolume-classification strategies addressing end-to-end training difficulties.
3 Preliminaries
The preliminaries define group-equivariant representations and convolutions, then specialize them to spherical signals under SO(3). They motivate spectral-domain computation as a practical response to sphere-discretization constraints.
- Equivariance: Equivariance preserves information about group actions by relating transformations of feature maps directly to transformations of inputs.
- Group Convolution: Group convolution computes inner products with transformed copies of a filter and is equivariant to transformations in the group.
- Spherical Discretization: No sphere discretization simultaneously provides well-distributed compact cells and transitivity, complicating cascaded spherical convolutions.
- Spherical Harmonics: Spectral-domain evaluation applies the spherical convolution theorem: expand signal and filter into spherical harmonics, multiply coefficients pointwise, then invert the expansion.
- Spherical Convolution: Spherical convolution produces spherical outputs by marginalizing rotation about the filter’s north pole, yielding zonal filters rather than SO(3) responses.
- Implementation: The method evaluates spherical Fourier transforms on an equiangular grid with sample weights, using matrix operations and sums that are differentiable in automatic-differentiation frameworks.
4 Method
The method builds spherical CNNs from spectral-domain convolutions, filter parameterizations, and pooling operations designed to preserve rotation equivariance while producing invariant descriptors.
- Architecture: The architecture applies spherical convolutional layers, optional pooling, nonlinearities, and weighted global average pooling to obtain an invariant descriptor.The network block consists of convolution, optional pooling, and nonlinearity; WGAP is applied at the last layer.
- Spherical convolution: Spherical convolution is computed exactly in the spherical harmonic domain, where convolution becomes pointwise multiplication of spherical Fourier transforms.Only order m = 0 filter coefficients are used, implying that learnable filters are zonal and constant along latitudes.
- Filter parameterization: Filters may be parameterized by all m = 0 spectral coefficients or by sparse anchor points whose missing degrees are linearly interpolated.For 32 × 32 inputs, the full-spectrum example uses 16 learned parameters, while localized filters use 4 anchor points.
- Filter parameterization: Spectral smoothness is used to obtain spatially localized filters because smooth spectra correspond to spatial decay.Locality is not guaranteed for the full-spectrum parameterization, although it may be learned.
- Pooling: Spectral pooling removes coefficients with degree larger or equal than b/2, acting as a low-pass operation that preserves equivariance.Compared with weighted average pooling, it is faster and reduces equivariance error, but also reduces classification accuracy; the preferred method depends on the application.
- Global pooling: Weighted global average pooling uses weights proportional to spherical cell areas, with each cell weight given by the sine of its latitude.This compensates for unequal areas in equiangular sampling and yields rotation-invariant descriptors.
5 Experiments
The experiments focus on 3D shape tasks using spherical representations derived from meshes or voxel grids, while documenting representation assumptions and training details.
- Experimental scope: The model targets shape classification, retrieval, and alignment in arbitrary orientations because these tasks benefit from inherent SO(3) equivariance.The representation is also applicable to other data that can be mapped to the sphere, such as panoramas.
- Spherical representation: The spherical conversion must itself be equivariant to rotations; otherwise, the learned representation will not be equivariant.This makes preprocessing a direct scope condition for the claimed equivariance.
- Spherical representation: Mesh and voxel inputs are converted into spherical functions by casting n×n equiangular rays from the bounding-sphere center and recording farthest intersection distances.Mesh inputs may include a second channel containing sin α, where α is the angle between the ray and the surface normal.
- Representation limitations: The representation is suitable for star-shaped objects whose bounding-sphere center is an interior point from which the whole boundary is visible.The authors do not check these conditions in practice, even when the representation is ambiguous or non-invertible.
- Training: Training uses ADAM for 48 epochs with an initial learning rate of 10^-3, divided by 5 at epochs 32 and 40.
Training:
ModelNet40 evaluation compares azimuthal and arbitrary rotation settings, showing that Spherical CNNs are more robust to unseen orientations while using fewer parameters and faster training.
- Training:: The evaluation considers azimuthal-to-azimuthal, arbitrary-to-arbitrary, and azimuthal-to-arbitrary rotation settings.
- Training:: All competing methods suffer a sharp performance drop when arbitrary rotations are present, including when those rotations appear during training.
- Training:: Spherical CNNs are more robust to arbitrary rotations, although performance drops in the azimuthal-to-arbitrary setting because of sampling effects.
- Training:: Spherical CNNs use one order of magnitude fewer parameters and train faster than competing methods.
- Training:: Equivariance to SO(3) is not needed when inputs contain only azimuthal rotations, so that setting does not exercise the model’s full potential.
5.3 3D object retrieval
On SHREC’17 retrieval with random SO(3) perturbations, the model matches state-of-the-art performance while using fewer parameters, smaller inputs, and no pre-training.
- 5.3 3D object retrieval: Retrieval uses ShapeNet Core55 with random SO(3) perturbations and combines classification training with an in-batch triplet loss.
- 5.3 3D object retrieval: The invariant descriptor is compared with cosine distance using class-specific thresholds and same-class predictions.
- 5.3 3D object retrieval: The model matches state-of-the-art retrieval performance with significantly fewer parameters, smaller input size, and no pre-training.
- 5.3 3D object retrieval: The SHREC’17 evaluation reports precision, recall, and micro- and macro-mean average precision, with their sum used for ranking.
5.4 Shape alignment
The model aligns differently oriented shapes through spherical correlation of learned feature maps, with intermediate layers providing the best alignment performance.
- 5.4 Shape alignment: Given same-category shapes in arbitrary orientations, corresponding feature maps are correlated and summed to estimate the aligning SO(3) rotation.
- 5.4 Shape alignment: Alignment is evaluated across network layers and against spherical correlation applied directly to the unlearned input representation.
- 5.4 Shape alignment: Learned features outperform the handcrafted spherical representation, with best performance from intermediate layers.
- 5.4 Shape alignment: The evaluation uses nonsymmetric ModelNet10 categories so each test shape has a unique ground-truth rotation and measurable angular error.
5.5 Equivariance error analysis
The analysis measures equivariance errors beyond the exact bandlimited convolution and pooling guarantees, identifying nonlinearities, sampling, and pooling as important sources.
- 5.5 Equivariance error analysis: For bandlimited inputs, spherical convolutions and spectral pooling preserve SO(3) equivariance, but other pipeline components can introduce errors.
- 5.5 Equivariance error analysis: The experiment compares each test entry with a randomly rotated version by measuring average relative feature-map error.
- 5.5 Equivariance error analysis: Pointwise nonlinearities introduce equivariance errors by generating frequencies beyond the bandlimit, while the mesh-to-sphere map is only approximately equivariant.
- 5.5 Equivariance error analysis: Larger input dimensions mitigate mesh-to-sphere mapping error, and bandlimited inputs produce smaller equivariance error.
- 5.5 Equivariance error analysis: Spectral pooling is exactly equivariant, whereas max-pooling introduces higher frequencies and has larger error than weighted average pooling.
- 5.5 Equivariance error analysis: Conventional planar CNNs likewise exhibit some translational equivariance error from max-pooling and discretization.
5.6 Ablation study
The ablation study finds that pooling, filter localization, and network size materially affect performance, while bandwidth and receptive-field choices trade off equivariance and accuracy.
- WAP, WGAP, and localized filters significantly improve performance in the ablation study.
- Larger networks further improve performance, indicating sensitivity to network size.
- Increasing bandwidth, such as through max-pooling, increases equivariance error and may reduce accuracy.
- Global operations in early layers, such as non-local filters, escape the receptive field and reduce accuracy.
6 Conclusion
Spherical CNNs use spherical convolutions to achieve SO(3) equivariance for 3D classification, retrieval, and alignment. The model handles arbitrary input orientations with relatively few parameters and small input sizes.
- Spherical CNNs leverage spherical convolutions to achieve equivariance to SO(3) perturbations.
- The network is applied to 3D object classification, retrieval, and alignment.
- The model naturally handles arbitrary input orientations while requiring relatively few parameters and small input sizes.