Source-linked AI summary

RAFT-DVC: Resolution-Aware Machine Learning-Based Digital Volume Correlation

Zixiang Tong, Lehu Bu, Jin Yang

arXiv:2609.01876v1cs.CVcond-mat.mtrl-sci

TL;DR

The paper addresses limited understanding of how machine-learning DVC resolution affects displacement accuracy and operating range. It develops and systematically evaluates RAFT-DVC solvers with matched feature-grid designs, finding resolution-dependent accuracy scaling and complementary regimes governed by displacement reach and volumetric-texture compatibility. The results support selecting a trained solver according to the measurement regime while identifying scope limits in the current evidence.

  • Problem

    How internal feature-grid resolution, volumetric texture, displacement reach, spatial resolution, computational cost, and out-of-distribution performance interact in correspondence-based machine-learning DVC remains insufficiently characterized.

  • Method

    RAFT-DVC is a family of three RAFT-based DVC solvers with encoder downsampling factors s = 2, 4, and 8, evaluated using matched synthetic designs and additional experimental texture-transfer tests.

  • Results

    RAFT-DVC shows approximately EPEraw ≈ 0.017s voxel scaling, complementary operating regimes, sampler-consistency gains, and competitive performance relative to tuned classical DVC under coarse-texture, large-displacement conditions.

  • Takeaways & Limitations

    Solver choice is a measurement-design decision that should jointly match displacement reach and volumetric texture rather than assume a universal ranking.

  • Takeaways & Limitations

    Controlled scaling and most operating-range and benchmark experiments use one synthetic particle-rendering pipeline, while the confocal experiment lacks independently known ground-truth displacement.

Abstract

from arXiv · show

Digital volume correlation (DVC) provides three-dimensional full-field displacement measurements from volumetric images, but how the internal resolution of a machine-learning-based DVC model affects accuracy and operating range remains poorly understood. Here, we present RAFT-DVC, a resolution-aware family of recurrent all-pairs field transforms (RAFT)-based DVC solvers with encoder downsampling factors s = 2, 4, and 8. Using a matched design, we find that the three solvers localize displacement to approximately 0.017 feature-grid voxel, giving an empirical raw-volume error scaling of approximately 0.017s voxel. The solvers exhibit complementary operating regimes governed jointly by displacement reach and volumetric-texture compatibility. Synthetic benchmarks show that RAFT-DVC achieves errors of the same order as tuned classical DVC under fine-texture, small-to-moderate-displacement conditions and becomes competitive or advantageous under coarse-texture, large-displacement conditions. Frequency-swept tests quantify deformation spatial resolution, while tiled inference enables dense estimation on large volumes. Evaluation on confocal volumetric images acquired during indentation illustrates the importance of matching solver operating regime to deformation magnitude and image texture. Tests on micro-CT images of elastomeric foam, despite training only on particle-labeled synthetic data, provide evidence of cross-texture transfer. We also identify coordinate-order inconsistencies in three-dimensional RAFT correlation sampling and introduce a non-cubic impulse test to verify sampler geometry independently of network training. Correcting the sampler improves native-input accuracy and generalization to unseen volume dimensions. Together, these results establish RAFT-DVC as a fast, resolution-aware framework for dense DVC with characterized accuracy and operating regimes.

1 Introduction

DVC methods must balance accuracy, spatial resolution, computational cost, volumetric texture, and displacement magnitude. RAFT-DVC addresses these gaps with resolution-aware solvers and systematic operating-regime characterization.

  • Existing DVC methods: Conventional DVC methods differ in local, global, and hybrid formulations, with local methods estimating displacement independently within subvolumes.Global methods impose volume-wide representations or regularization, while hybrid methods combine local correlation with global kinematic consistency.
  • Current challenges: User-selected parameters must be adapted to imaging conditions, volumetric texture, displacement magnitude, and desired spatial resolution.Subvolumes that are too small may lack sufficient image information, whereas oversized subvolumes can smooth local deformation gradients.
  • Current challenges: Volumetric texture varies across imaging modalities in feature size, density, contrast, and noise, preventing a single optimal DVC configuration.The measurement setup must be matched to the available image texture, deformation regime, and research objective.
  • Current challenges: Conventional three-dimensional DVC can require minutes to hours as volume size and measurement resolution increase.Local methods repeat correlation or optimization at measurement points, while global methods repeatedly solve coupled volume-wide problems.
  • Machine-learning DVC limitations: Machine-learning DVC methods accelerated analysis, but their performance outside represented textures and deformation distributions remained incompletely characterized.Prior evaluations also left zero-strain precision, spatial resolution, and operating range insufficiently established.
  • Contributions of this work: RAFT-DVC uses three encoder downsampling factors, s = 2, 4, and 8, to characterize feature-grid resolution across texture and deformation regimes.The framework treats solver selection as a measurement-design choice rather than seeking one universally optimal network.

2 Methodology

RAFT-DVC estimates dense three-dimensional displacement by recurrently refining correspondence between encoded reference and deformed volumes. Three matched encoder arms vary downsampling factor while preserving comparable feature-grid problems, enabling controlled study of raw-volume resolution.

  • RAFT-DVC Architecture: RAFT-DVC encodes both volumes, builds an all-pairs six-dimensional correlation volume, and recurrently accumulates three-component displacement increments.A context encoder initializes the recurrent update operator, which samples correlation features around the current estimate.
  • Encoder Arms and Matched Controlled Design: Matched arms map 32^3, 64^3, and 128^3 raw inputs to the same 16^3 feature grid while scaling particle radius and other problem dimensions with s.Particle radii are 2, 4, and 8 raw voxels, yielding Rfeature = 1 feature voxel for every arm.
  • Encoder Arms and Matched Controlled Design: The s2, s4, and s8 encoder variants share convolutional stages, channel dimensions, and parameter counts but differ in stride-2 downsampling placement.The corresponding feature-grid spacings are 2, 4, and 8 raw voxels.
  • Encoder Arms and Matched Controlled Design: The matched design preserves particle geometry, density, displacement range, and correlation-grid dimensions in feature coordinates while varying raw-volume spacing per feature voxel.This is a controlled comparison of encoder downsampling and raw-coordinate displacement accuracy, not a fixed-input encoder ablation.
  • Synthetic Particle-Volume Generation: Synthetic training volumes use rendered spherical particles with edge taper, Gaussian point-spread convolution, intensity and radius variability, and prescribed polynomial displacement fields.Particle radius and intensity are independently perturbed to introduce intra-volume variability.

Displacement Accuracy

Displacement accuracy is measured with endpoint error in both raw-volume and feature-volume coordinates. The two forms represent the same Euclidean displacement error on different spatial grids, while training and evaluation use different loss norms.

  • Displacement Accuracy: Endpoint error is the mean Euclidean distance between the estimated displacement field and prescribed ground truth over the evaluated region.EPEraw is expressed in raw-volume voxels.
  • Displacement Accuracy: EPEfeature expresses the same displacement error in feature-volume voxels, allowing localization comparisons across encoder downsampling factors.Raw-volume and feature-volume values differ by the spatial-grid scaling.
  • Displacement Accuracy: The training objective uses an ℓ1 sequence loss, whereas evaluation uses the Euclidean ℓ2 endpoint error.This choice follows common practice in optical-flow and learning-based DVC benchmarking.

Reference-free Image-matching Residual

The paper introduces CSSD as a reference-free measure of how well a measured displacement field aligns reference and deformed volumes, while emphasizing that image-matching quality is not displacement accuracy.

  • CSSD evaluates local image-matching quality by measuring residual intensity mismatch after applying a measured displacement field.It complements ground-truth-based EPE for experimental data, where the true displacement is unavailable.
  • The reference and deformed volumes are Gaussian low-pass filtered before computing the image-matching residual.The filtering reduces contributions from high-spatial-frequency acquisition noise.
  • The measured displacement field samples the filtered deformed volume back at corresponding reference coordinates using tricubic B-spline interpolation.This interpolation matches the bicubic convention used by the classical comparison solvers.
  • A local affine intensity correction accounts for brightness and contrast differences, including depth-dependent attenuation in confocal image stacks.Local contrast and brightness coefficients are obtained by minimizing squared intensity mismatch within cubic windows.
  • CSSD can be artificially reduced by displacement fields that overfit image noise, so it is interpreted alongside agreement with classical DVC methods and the null-displacement residual.

Deformation Spatial Resolution

The paper characterizes deformation spatial resolution through attenuation of recovered sinusoidal displacement amplitudes and addresses computational and sampler constraints affecting RAFT-DVC deployment.

  • Deformation Spatial Resolution: DVC attenuates displacement variations at sufficiently short spatial wavelengths, reducing recovered amplitude as deformation wavelength decreases.
  • Deformation Spatial Resolution: The critical wavelength λq is the wavelength at which recovered amplitude reaches fraction q of prescribed amplitude, such as λ85 for 85% recovery.
  • Deformation Spatial Resolution: Frequency-swept sinusoidal tests estimate retained amplitude as upred/uGT along the sweep and fit attenuation with a decreasing sigmoid.The recovered profile is averaged transversely before local extrema are evaluated.
  • Deformation Spatial Resolution: The sigmoid uses a long-wavelength plateau, transition width, and transition center, while its zero lower asymptote represents complete attenuation at vanishing wavelength.The fit is performed in position coordinates and mapped to prescribed wavelength to obtain λq.
  • Sampler Geometry: A non-cubic impulse test detects coordinate-order inconsistencies by revealing whether the sampler returns an impulse at its prescribed location.The corrected sampler was applied to all reported RAFT-DVC results after retraining.
  • Computational Constraints: Three-dimensional all-pairs correlation requires O(n6) entries, making direct inference memory grow rapidly with feature-grid size.Reducing encoder downsampling increases feature-grid dimensions for a fixed raw volume, while overlapping tiled inference makes peak memory depend mainly on tile dimensions.

3 Results

Under a matched feature-grid design, RAFT-DVC arms achieve similar feature-space localization accuracy, while raw-volume error increases with encoder downsampling factor.

  • Matched Feature-Grid Accuracy: 0.0179, 0.0153, and 0.0166 feature voxels are the mean EPEfeature values for s2, s4, and s8, respectively.These values cluster around the common scale EPEfeature ≈ 0.017 feature voxel.
  • Matched Feature-Grid Accuracy: 0.036, 0.061, and 0.133 raw voxels are the corresponding mean EPEraw values for s2, s4, and s8, respectively.Because EPEraw = sEPEfeature, the empirical scaling is EPEraw ≈ cs.
  • Matched Feature-Grid Accuracy: The three arms localize displacement to a similar fraction of a feature-grid voxel, but the same relative error becomes larger in raw-volume coordinates as s increases.Evaluation before final trilinear upsampling shows the same trend.
  • Out-of-Distribution Evaluation: Additional tests vary displacement magnitude and particle size relative to the feature grid to characterize operating reach and robustness outside nominal training conditions.

Displacement Operating Range

RAFT-DVC arms maintain an error floor through their training displacement intervals before accuracy collapses, while particle-size mismatch and implementation choices materially affect performance.

  • Displacement Operating Range: Each solver shows an approximately constant error floor over its training displacement range, followed by rapid accuracy loss as displacement increases.
  • Displacement Operating Range: 4.3 and 5.7 voxels are the accurate-range and collapse thresholds for s2; corresponding thresholds are 9.4 and 12.7 for s4, and 18.9 and 22.2 for s8.Each arm’s accurate range extends beyond its upper training bound.
  • Displacement Operating Range: 16, 32, and 64 raw voxels are the nominal geometric reaches for s2, s4, and s8, exceeding measured operating thresholds.Thus, correlation-lookup extent alone does not determine usable displacement range.
  • Robustness to Particle-Size Variations: Departing from the trained particle size increases EPEfeature approximately two- to five-fold, with undersized particles causing the largest degradation.At Rfeature = 0.5, the penalty increases from approximately 2.4× for s2 to 5.5× for s8.
  • Sampler Correction: Correcting sampler coordinates reduces native s2 EPEraw from 0.089 to 0.036 voxel, an approximately 2.5× reduction, and improves unseen-size generalization.The coordinate-inconsistent model has errors 34–51× larger at dimensions different from training.
  • Tiled Inference: With 50% overlap, tile-boundary EPE is approximately 10–17% larger than interior EPE; increasing overlap to 75% reduces the difference to approximately 4%.Overlapping inference reduces but does not eliminate boundary effects.
  • Displacement Noise: RAFT-DVC displacement noise floors are approximately 0.018–0.047 voxels, while sampler-corrected VolRAFT gives approximately 0.15–0.36 voxel.
  • Strain Noise: 0.60–1.9 × 10^-4 RMS strain is reported for classical DVC, compared with approximately 2.4–2.8 × 10^-4 for s2 and s4 and 9.2 × 10^-4 for s8.The larger s8 strain noise is attributed to differentiation amplifying coarse feature-grid errors.

Benchmark scenarios and comparison protocol

The benchmark spans matched, cross-regime, spatial-resolution, and large-volume deployment tests, comparing three resolution-aware RAFT-DVC solvers with VolRAFT and classical DVC methods under shared evaluation conditions.

  • Scenario design: Seven scenarios vary particle size, displacement magnitude, deformation wavelength, and volume size to assess accuracy, generalization, spatial resolution, and scalability.S1–S5 cover particle-size and displacement regimes; S6 probes spatial resolution; S7 evaluates tiled inference on 512^3 volumes.
  • Scenario design: S1–S3 are matched in-distribution cases pairing raw particle radii 2, 4, and 8 voxels with corresponding displacement ranges for s2, s4, and s8.The matched pairs are (2,), (4,), and (8,) raw-volume voxels.
  • Scenario design: S4–S5 are cross-regime tests combining fine particles with large displacement or large particles with small displacement outside any single arm’s joint training distribution.S4 uses (2,) and S5 uses (8,) raw-volume particle-radius and displacement combinations.
  • Scenario design: S6 prescribes a frequency-swept sinusoidal displacement whose local wavelength decreases across the volume, with bands matched to the s2, s4, and s8 particle regimes.The tested wavelength ranges are 64→8, 96→16, and 128→32 voxels for the fine-, intermediate-, and coarse-particle bands.
  • Comparison protocol: All methods are compared on the same 4096-location grid in the central 128^3 region, while S7 assesses tiled inference, memory, and cost on a 512^3 volume.The benchmark includes RAFT-DVC, two VolRAFT sampler variants, local DVC, finite-element global DVC, and ALDVC.

Displacement Accuracy Within and Across Training Regimes

RAFT-DVC accuracy depends on the joint particle-size and displacement regime: classical DVC leads for fine and intermediate matched cases, whereas s8 is advantageous for coarse particles and large displacement, with performance degrading out of distribution.

  • Matched regimes: 0.045, 0.058, and 0.128 voxels are the S1–S3 mean EPEraw values for matched s2, s4, and s8 solvers.The raw-volume error increases with feature-grid spacing across the matched arms.
  • Matched regimes: 0.021 voxel versus 0.045 voxel in S1 and 0.037 voxel versus 0.058 voxel in S2 shows ALDVC outperforming the matched RAFT-DVC arms for fine and intermediate particles.These comparisons use ALDVC EPEraw against RAFT-DVC s2 in S1 and s4 in S2.
  • Matched regimes: 0.128 voxel versus 0.248 voxel in S3 shows RAFT-DVC s8 achieving an approximately 1.9× lower endpoint error than ALDVC for coarse particles and large displacement.The ordering reverses relative to S1 and S2.
  • Cross-regime generalization: 1.15 voxels is the best RAFT-DVC EPEraw in S4, compared with 0.68, 0.81, and 0.84 voxels for ALDVC, global DVC, and local DVC.In S5, the best classical result is 0.058 voxel versus 0.089 voxel for s8; both are cross-regime out-of-distribution tests.
  • Visual comparisons: Figures 9–11 present median-EPE representative displacement fields and signed error maps for the three in-distribution scenarios, while Figure 12 probes spatial resolution.The displayed comparisons include ground truth, matched RAFT-DVC, corrected VolRAFT, and classical DVC methods.

Deformation Spatial Resolution

Deformation spatial resolution worsens with encoder downsampling but scales sublinearly, so solver selection must consider wavelength resolution alongside displacement reach and texture compatibility.

  • Resolution thresholds: λ85 is 30, 57, and 94 raw-volume voxels for s2, s4, and s8, while λ50 is 19, 36, and 59 voxels.The long-wavelength amplitude plateaus are 0.87, 0.89, and 0.95 for the three bands.
  • Scaling: b = 0.79 ± 0.07, 0.83 ± 0.06, 0.83 ± 0.06, and 0.78 ± 0.05 for q = 50, 75, 80, and 85% indicates sublinear dependence on encoder downsampling.The similar exponents across thresholds distinguish this scaling from strict proportionality.
  • Solver selection: A 50-voxel deformation wavelength exceeds the 85% threshold for s2 and the 50% threshold for s4, but falls below the 50% threshold for s8.Thus spatial variation can constrain solver choice even when displacement magnitude is within operating range.
  • Large-volume inference: Dense tiled inference produces 110.6 million output values in 18.4 s at 0.5 GB peak memory, but 256^3 tiles reduce runtime to 10.1 s while raising memory to 8.3 GB and EPEraw to 0.184 voxel.The outputs are interpolated to the raw grid from displacement refined on the feature grid.
  • Large-volume inference: Runtime comparisons are implementation-level benchmarks rather than hardware-independent algorithmic speedups because hardware, software, output density, and nodal spacing differ.The paper explicitly cautions against interpreting these measurements as normalized algorithmic complexity comparisons.

4 Discussion

RAFT-DVC solvers have complementary operating regimes: accuracy depends jointly on displacement reach and volumetric-texture compatibility, not encoder resolution or displacement magnitude alone. Their performance relative to classical DVC is likewise regime-dependent, while dense output, sampler correctness, and input dimensions impose additional deployment considerations.

  • Feature-grid resolution: 0.017 feature voxel is the matched-solvers’ approximate displacement error in feature-grid coordinates, while raw-volume error scales as approximately 0.017s voxel.The scaling coefficient is empirical and its universality remains unresolved.
  • Joint operating regimes: Solver selection requires both recoverable displacement reach and volumetric texture compatible with the learned representation.The operating-regime experiments show that feature-grid spacing cannot be selected from displacement magnitude alone.
  • Joint operating regimes: In S4, mismatched fine texture and large displacement produced RAFT-DVC EPEraw ≈1.15 voxels, while tuned classical methods remained below one voxel.Neither narrowly trained solver family matched both dimensions of the S4 condition.
  • Joint operating regimes: In S5, the s8 solver achieved EPEraw = 0.089 voxel versus 0.742 voxel for s2 despite the smaller displacement magnitude being within s2’s training range.The result emphasizes that texture compatibility can outweigh displacement magnitude when selecting a solver.
  • Relationship to classical DVC: RAFT-DVC matched tuned classical DVC in error order for S1–S2 and reduced EPE by approximately 1.9× relative to ALDVC in coarse-texture, large-displacement S3.Classical methods retained a modest accuracy advantage in S1–S2, so RAFT-DVC is not a universal replacement.
  • Sampling and deployment: Dense voxelwise output does not imply voxel-scale independent spatial resolution, although matched RAFT-DVC solvers retained deformation spatial resolution comparable to or better than ALDVC.Learning-based inference also reduced per-measurement parameter adjustment and enabled dense inference on large volumes.
  • Sampling and deployment: Coordinate-order errors in volumetric correlation sampling can remain hidden on fixed-size cubic inputs, while tiled accuracy changes with training-unseen dimensions and overlapping boundaries.Sampler geometry should be tested independently, and input dimensions and tiling configuration belong to the characterized operating regime.

5 Conclusion

RAFT-DVC provides resolution-aware solvers with matched feature-grid accuracy but complementary operating regimes. Sampler validation and released resources support reproducible, fast, dense DVC while preserving solver-selection requirements.

  • 0.017 feature voxel localization yields empirical raw-volume error scaling EPEraw ≈ 0.017s voxel across the three matched solver arms.The arms use encoder downsampling factors s = 2, 4, and 8.
  • Finer feature grids reduce raw-volume error, whereas coarser-resolution solvers cover larger displacement ranges and reduce fixed-volume correlation memory.Operating regimes depend jointly on displacement reach and volumetric-texture compatibility.
  • Sampler coordinate-order corrections improve native-input accuracy and generalization to volume dimensions absent from training.A non-cubic impulse test verifies three-dimensional correlation-sampling geometry independently of network training.
  • Learning relocates DVC parameter selection from continuous per-measurement tuning to choosing a trained solver with a characterized operating regime.The authors release three trained models, a correlation-sampling diagnostic, and the synthetic-data generation pipeline.
  • The conclusion identifies reproducibility, speed, density, and quantitative characterization as goals of the released learning-based DVC framework.

Declarations

The declarations report funding, disclose author relationships affecting baseline evaluation and dataset provenance, and state that no financial conflicts or ethics approvals apply.

  • Funding is provided through NSF grants and University of Texas at Austin academic and fellowship programs.
  • Classical-DVC baselines were evaluated with an author-maintained MATLAB implementation, and baseline parameters and headless drivers are released for independent assessment.
  • The confocal indentation images come from the publicly available DVC Challenge 2.0 dataset, whose co-authorship by study authors is disclosed.
  • The authors declare no financial or other competing interests, and refer readers to Appendix A for data and code availability.
  • The study uses synthetic volumes and a previously published public dataset, with no human participants or animals.

A.1 Hardware and software.

The experiments use matched RAFT-DVC arms, controlled synthetic training, explicit displacement conventions, and scenario-specific classical-DVC settings. Hardware, software, solver parameters, and evaluation grids are reported for reproducibility.

  • Hardware and software: All three RAFT-DVC arms share the framework, recurrent architecture, channel dimensions, and optimization procedure while using s = 2, 4, and 8 encoder downsampling.The feature encoder has 128 channels; the recurrent update uses a 96-channel hidden state and 64-channel context representation.
  • Hardware and software: Training uses AdamW, gradient clipping, RAFT sequence loss, a one-cycle learning-rate schedule, 300 epochs, effective batch size 8, and base seed 42.The peak learning rate is 2 × 10^-4, with warm-up fraction 0.2 and cosine annealing.
  • Hardware and software: Synthetic volumes contain tapered spherical particles with specified point-spread, particle-sampling, saturation, Poisson-noise, and Gaussian read-noise models.The point-spread-function width is σ = 0.8 voxel and Poisson gain is 500.
  • Hardware and software: The matched design uses particle radii 2/4/8 voxels, densities 4.8/0.6/0.075 particles per 10^3 raw voxels, and volume sizes 32^3/64^3/128^3 voxels.These choices produce common feature-grid particle radius and displacement range [1,2] voxels.
  • Hardware and software: Prescribed displacement uses a componentwise maximum convention, while reported displacement errors use Euclidean vector magnitudes.The maximum vector magnitude can exceed Uraw by as much as √3 and averages approximately 1.2Uraw.
  • Hardware and software: EPEraw is measured in raw-volume voxels and EPEfeature = EPEraw/s in feature-grid voxels for matched-arm comparisons.Out-of-distribution robustness is reported as EPEfeature while particle size varies at each arm’s training density.
  • Classical DVC methods parameters: Classical-DVC comparisons use scenario-specific nodal spacing, subset sizes scaled to particle radius, and distinct local, ALDVC, and FE-global formulations.Spacing is 8 voxels for S1–S6, 16 voxels for S7, and 12 voxels for the sparse-texture indentation volume; FE elements equal nodal spacing.
  • Classical DVC methods parameters: ALDVC and FE-global DVC use explicitly reported initialization, iteration, tolerance, regularization, and solver parameters, while calculations run in MATLAB with 12 CPU workers.
Loading 2609.01876v1…