Source-linked AI summary

Proximity3D: Shape from Capacitive Proximity on Sensing Manifold

Hao Chen, Chenming Wu, Chun Ping Lam, Xiangjia Chen, Guoxin Fang, Charlie C. L. Wang, Yeung Yam, Juncong Lin, Chengkai Dai

arXiv:2608.30344v2cs.CVcs.CGcs.GRcs.RO

TL;DR

Near-field robotic shape sensing is difficult because capacitive proximity signals are sparse, indirect, and collected on curved sensing surfaces. Proximity3D reconstructs 3D object geometry from multi-view capacitive fields and demonstrates reconstruction across simulated and physical experiments.

  • Problem

    Existing proximity sensing mainly supports collision heuristics or distance estimates, while sparse, indirect fields provide limited evidence for richer object-shape reconstruction.

  • Method

    Proximity3D uses manifold-aware local aggregation, global multi-view fusion, and a pretrained shape decoder to recover complete 3D geometry from capacitive fields.

  • Results

    Physical and simulated experiments demonstrate 3D shape reconstruction from multi-view capacitive proximity fields, with material changes causing at most 8.1% CD and 6.6% EMD changes.

  • Takeaways & Limitations

    The work demonstrates pre-contact geometric awareness from capacitive signals captured on a curved sensing manifold for embodied robotic sensing.

  • Takeaways & Limitations

    The method currently applies only to conductive objects, and the contribution of surrogate-model approximation to reconstruction error remains unquantified.

Abstract

from arXiv · show

Most shape reconstruction methods assume measurements defined over planar sensing domains, such as RGB images or depth maps. In this paper, we use a curved capacitive textile as a shape sensor, treating its surface as a non-planar sensing manifold. Each scan is represented as a capacitive proximity field on this manifold, induced by the interaction between the curved electrode layout and nearby object geometry. We introduce a multi-view feedforward reconstruction model that aggregates these fields across known sensor views and recovers the observed object shape. Simulated and physical experiments demonstrate robust reconstruction from capacitive proximity signals acquired on curved sensing surfaces, pointing toward a new route to robotic near-field geometric awareness via embodied sensing.

1 Introduction

The introduction frames capacitive proximity sensing as an embodied solution to occluded robotic near-field perception, while identifying indirect, sparse signals as a reconstruction challenge. Proximity3D addresses this challenge through manifold-aware attention, multi-view feature fusion, and a surrogate forward model for scalable data synthesis.

  • Motivation: Occlusion in the narrow pre-contact gap motivates using the robot body itself as a sensing domain for embodied near-field perception.Nearby conductive objects perturb the sensor’s electric field, modulating capacitance and generating measurable responses.
  • Challenge: Capacitive fields are indirect and sparse because channel values encode complex coupling between local electrode layouts and nearby object geometry.The introduction states that interpreting these signals requires explicitly integrating the sensor’s local geometry into the representation.
  • Proximity3D: Proximity3D processes sensing-manifold measurements with local electrode frames and non-uniform layouts, using Manifold Sensing Attention to aggregate neighboring channel responses.Per-view sensing features are fused into shape tokens and transformed by a global module into a 3D latent representation.
  • Data synthesis: A surrogate forward model synthesizes realistic proximity fields for unseen meshes and different sensor poses, extending training data beyond physical collection.The model adapts local geometry-derived priors into real sensor responses.
  • Contributions: The framework reconstructs shape from multi-view capacitive proximity fields on a curved sensing manifold and introduces MSA using local tangent frames to capture local sensing context.The passage presents this as the first method to effectively address this challenge, to the authors’ knowledge.

2 Related Work

Prior work spans dense visual reconstruction, contact-based tactile sensing, conformable electronic skins, and geometric learning on non-Euclidean domains. Proximity3D builds on fabrication, capacitive forward modeling, surface learning, and multi-view latent representations while targeting inversion of sparse capacitance readings into 3D geometry.

  • Visual and learned reconstruction: Traditional reconstruction uses dense visual or depth data, while learned representations include DeepSDF, Occupancy, NeRF, and 3D Gaussian Splatting.These approaches are contrasted with the sparse capacitive measurements addressed by Proximity3D.
  • Tactile reconstruction: Optical tactile and visuo-tactile methods integrate localized surface deformations into global shapes using implicit fields, diffusion priors, or neural tracking.GelSight and DIGIT are examples of elastomer-based optical tactile sensors.
  • Electronic skins and proximity sensing: Electronic skins distribute sensing across robot bodies and conformable substrates, while capacitive and multimodal arrays extend sensing beyond physical contact.Reported applications include whole-body collision avoidance, bio-inspired spatial tracking, flexible platforms, and wireless multimodal e-skins.
  • Electronic skins and proximity sensing: Existing proximity systems mainly support threshold-based alarms, gesture recognition, or coarse localization rather than inverting sparse capacitance readings into 3D geometry.This identifies the task that the paper positions as highly ill-posed and insufficiently addressed by prior systems.
  • Methodological foundations: Digital fabrication, capacitive forward models, surface-learning architectures, and multi-view latent-query methods provide foundations for reconstructing geometry on prescribed non-Euclidean sensing domains.Computational textiles enable calibrated surface parameterizations, while geometric learning, transformers, and cross-attention handle surface data and variable-sized observations.

3 Method

Proximity3D reconstructs a canonical-frame object mesh from multi-view capacitive proximity fields, sensor poses, and known manifold connectivity. It combines manifold-aware local feature aggregation, cross-view shape-token fusion, and a pretrained shape prior for sparse-to-complete decoding.

  • Overview: The pipeline maps posed multi-view capacitive proximity fields and known sensor connectivity to an object mesh in the canonical object frame.Each view provides a normalized response at every channel site together with rotation and translation derived from hardware kinematics.
  • Manifold Sensing Attention: Manifold Sensing Attention aggregates neighboring electrode readouts using local tangent-frame geometry and connectivity to encode direction-dependent sensing responses.The same object can produce different channel responses under different orientations relative to local electrode frames.
  • Multi-view Fusion via Shape Tokens: Per-view sensing features are combined with Fourier-encoded poses, while 8 learnable shape tokens exchange cross-view context through self-attention.The model also uses 4 register tokens to provide non-semantic transformer capacity and reduce spurious accumulation in shape tokens.
  • Shape Decoding with Prior: A pretrained TRELLIS-2 decoder supplies a learned shape prior for reconstructing complete geometry from sparse fields containing typically 100–300 electrode responses per view.The decoder is conditioned through a structured 8^3 latent grid with 32-dimensional features per occupied cell.

4 Surrogate Forward Model for Capacitive Proximity Fields

The section introduces a surrogate generative model for capacitive proximity fields, using hemispherical raycasting to derive a sensing geometric prior and a conditional Diffusion Transformer to predict physical responses. Manifold Sensing Attention injects local geometry into denoising, while FEM pre-training followed by physical-data fine-tuning adapts the model to real hardware.

  • The method formulates data generation as a generative problem to mitigate the time cost of physical data acquisition at scale.
  • Each channel response is approximated by hemispherical raycasting in its local tangent frame, using first-hit distances clipped at d_max = 30 mm.For channel site x_j, the method samples K random directions and computes the first-hit distance along each direction.
  • A cosine-weighted hemispherical fraction defines the sensing geometric prior, with local-normal alignment accounting for directional dependency across the sensing manifold.
  • A conditional Diffusion Transformer learns the surrogate distribution p_θ(C | g), training through standard noise prediction on corrupted capacitive fields.The diffusion timestep is uniformly sampled, and Gaussian noise ε ∼ N(0, I) corrupts the capacitive field.
  • The denoising network uses Manifold Sensing Attention to inject local geometry embeddings into self-attention and ground generation in sensing-manifold geometry.
  • The surrogate is pre-trained on FEM-generated synthetic data and then fine-tuned exclusively on a small physical dataset to adapt its prior to real hardware.

5 Experiments

The experiments evaluate Proximity3D using simulated and physical proximity data on a hemispherical 100-channel sensing manifold. They assess reconstruction quality, view-count and attention-module effects, and surrogate forward-model accuracy against held-out hardware measurements.

  • Experimental setup: The physical setup uses a hemispherical sensing manifold with 100 electrode channels and a TM-500 robotic arm controlling sensor pose.
  • Datasets: Training combines 500 FEM-simulated meshes with 25 physical objects sampled at 512 poses, then generates realistic proximity data for 1,624 meshes.The surrogate forward model is trained on simulations and fine-tuned with physical measurements to align real-world proximity responses.
  • Metrics: Evaluation uses Chamfer Distance (CD), Earth Mover’s Distance (EMD), and Average Surface Error to measure local and global reconstruction fidelity.EMD and Average Surface Error complement CD because CD can be biased toward local point density.
  • Evaluation protocol: Reconstruction is evaluated on 200 unseen test shapes using 512 valid views per shape, with a standard random subset of 64 views.Valid views have non-zero sensor responses, and additional analyses test Manifold Sensing Attention (MSA) against vanilla attention and GAT baselines.
  • Ablation and view count: MSA outperforms both baselines, while reconstruction improves through V=64 views and then plateaus.The study therefore uses V=64 by default to balance reconstruction fidelity and computational efficiency; visualizations appear in Fig. 4 and Fig. 5.
  • Surrogate validation: The surrogate forward model is validated on held-out physical object-pose pairs using channel-wise MAE, RMSE, Pearson, and Spearman correlation.These metrics compare generated and measured capacitive proximity fields, quantifying absolute response errors and preservation of channel-wise response patterns.

6 Application: Dexterous Hand with Palm-Mounted Sensing Manifold

The application integrates a woven capacitive sensing manifold into a dexterous hand, using hand motion to collect multi-view proximity signals for object-shape reconstruction and grasp planning. With 64 accepted pre-contact views, the palm-mounted sensor provides sufficient near-field observations for shape recovery, while grasp planning is evaluated on reconstructed and geometric reference representations.

  • Application setup: A woven capacitive sensing manifold mounted on an RH56F1 dexterous hand captures multi-view proximity signals during hand motion for object-shape reconstruction and grasp planning.The platform uses the hand’s motion to acquire views around an object.
  • Application setup: Valid views are selected from randomly sampled hand poses using a prespecified proximity-response band and the proximity response as a pre-contact safety signal.Poses are adjusted closer when the maximum channel response is below the lower threshold; the supplied passage truncates the remaining condition.
  • Reconstruction: 64 accepted pre-contact scans are used by default to reconstruct the object mesh in the object frame.The view count follows view-count calibration, and representative real-world reconstructions are shown for this setting.
  • Reconstruction: The palm-mounted sensing manifold provides sufficient near-field observations for object-shape recovery to facilitate grasp planning.This conclusion is drawn from three representative real-world reconstruction cases.
  • Grasp planning evaluation: Grasp planning is evaluated on reconstructed geometry and the ground-truth convex hull, OBB, and AABB using GraspGen and a standard force-closure check.The evaluation includes 100 objects for which planning on the ground-truth mesh succeeds.

7 Discussion

Proximity3D demonstrates multi-view 3D shape recovery from capacitive proximity fields on curved sensing manifolds, while its current applicability is limited to conductive objects. The discussion identifies decoder bias, weak sensing, and surrogate approximation as important limitations requiring further evaluation.

  • Material applicability: Proximity3D currently applies only to conductive objects because its sensing mechanism relies on electric-field coupling.Material mainly changes overall coupling strength, with weaker influence on the normalized spatial patterns used by the network.
  • Material applicability: A lower-conductivity conductive-fiber composite hand is reconstructed without material-specific retraining, suggesting robustness across the tested conductive objects.Broader material characterization is needed to determine the method’s practical conductivity boundary.
  • Evidence dependence: The pretrained decoder converts sparse multi-view evidence into complete meshes but may bias reconstruction when sensed evidence is weak.Quality improves with additional views, while replacing MSA with vanilla self-attention degrades performance despite an unchanged decoder.
  • Error sources: Surrogate-model approximation may contribute to reconstruction failures, but the current evaluation does not quantify its share of observed reconstruction error.Future work will estimate this contribution using paired physical and surrogate measurements under matched sensing configurations.
  • Overall contribution: Physical experiments validate Proximity3D as a framework for recovering 3D object geometry from multi-view capacitive proximity fields across a curved sensing manifold.The approach demonstrates pre-contact geometric awareness and has potential to enable broader robotic applications.

Supplementary Material … A.3 Sensor Fabrication

The supplementary material describes a woven capacitive sensing manifold, its multiplexed spatial readout, the local proximity-sensing mechanism, and fabrication through computer-controlled 3D freeform weaving. It emphasizes that measurements are spatially indexed and reflect nonlinear, distributed field perturbations rather than direct depth.

  • A.1 Sensor Architecture and Readout: The fabric interlaces stainless-steel warp sensing electrodes with silicone-insulated copper weft driven electrodes, whose coating preserves capacitive coupling while preventing electrical contact.Adjacent conductive threads are grouped into logical channels, so each thread crossing is structural rather than an individual sensing site.
  • A.1 Sensor Architecture and Readout: Time-multiplexed electronics excite selected weft channels and integrate charge transferred to selected warp channels to estimate capacitance.Scanning all channel pairs produces a spatially indexed capacitance vector while preserving each measurement’s physical location on the curved surface.
  • A.2 Proximity Sensing Principle: Each site has a baseline mutual capacitance C0,j comprising intended warp–weft coupling and fixed parasitic contributions from the textile and wiring.A nearby conductive object adds a capacitive path that redistributes the fringing electric field and changes charge received by the sensing electrode.
  • A.2 Proximity Sensing Principle: Object–electrode coupling increases with effective coupled area and decreases with separation.This provides intuition for how nearby geometry influences the capacitive response.
  • A.2 Proximity Sensing Principle: The coupling relation is only a local approximation because curvature, fringing fields, grounding, routing parasitics, and neighboring channels produce nonlinear, spatially distributed responses.Accordingly, a single channel is not treated as a direct depth measurement; baseline correction and calibration are applied before interpreting the site response.
  • A.3 Sensor Fabrication: The textile layout derives from a geodesic stitch-map formulation and is fabricated using a computer-controlled 3D freeform weaving system.The target surface is compiled into machine instructions controlling warp selection, local warp-thread release, and weft insertion to form curvature.

B Implementation Details

The implementation reconstructs objects from 64 posed proximity fields using manifold aggregation, multi-view Transformer fusion, and cell-query decoding. Training uses weighted reconstruction losses and AdamW optimization on eight RTX 4090 GPUs.

  • Network architecture: The reconstruction network uses V=64 posed proximity fields, two MSA layers, a six-layer Transformer, and a four-layer cross-attention cell-query decoder.Its feature width is d=384, and the specified ψ modules are implemented as two-layer MLPs.
  • Loss configuration: The training loss weights are λ_o=1.0, λ_f=1.0, and λ_D=0.1.These weights apply to the reconstruction loss defined in the main paper.
  • Surrogate forward model: The surrogate forward model samples K=512 random directions from the local upper hemisphere S2+ at each channel for geometric conditioning.The sampled directions construct the geometric conditioning signal.
  • Optimization and training: Optimization uses AdamW with a 1×10−4 base learning rate, 0.05 weight decay, gradient clipping at 1.0, 2,000-step warmup, and cosine decay.Gradient computations use PyTorch, and training runs on a single node with 8× NVIDIA RTX 4090 GPUs.

C Architecture Generalization

The study evaluates architecture generalization by reconstructing a hand, cube, and ring across three distinct sensing-manifold geometries. The models are retrained for each manifold while preserving the same architectures and training settings.

  • Architecture generalization: The same hand, cube, and ring are reconstructed using sensing manifolds shaped as a hemisphere, robotic-arm elbow, and dexterous-hand palm.Each manifold uses its own geometry and sensor layout.
  • Architecture generalization: Both surrogate and reconstruction models are retrained separately for each manifold using its geometry and sensor layout, while retaining identical architectures and training settings.This protocol directly tests generalization across sensing-manifold architectures.
  • Architecture generalization: Reconstructions from the palm-mounted sensor are presented for the hand, cube, and ring, while hemisphere and robotic-arm-elbow results are shown in Fig. S2.The figures visualize reconstruction of the same three objects across the tested sensor geometries.
Loading 2608.30344v2…