Source-linked AI summary
A soft thumb-sized vision-based sensor with accurate all-round force perception
Huanbo Sun, Katherine J. Kuchenbecker, Georg Martius
TL;DR
Existing vision-based haptic sensors do not yet fully meet practical requirements for robust robot contact perception and rich force information. Insight combines a soft-stiff sensor design, internal vision with photometric and structured lighting, and learned force estimation. It reports all-surface directional force maps with 0.4 mm localization accuracy, 0.03 N force magnitude accuracy, and 5° force direction accuracy.
Problem
Robot contact perception requires information about where and how contact occurs, but existing approaches face occlusion, small deformation signals, limited force information, and practical design constraints.
Method
Insight uses a soft elastomer over-molded on a stiff skeleton, a single internal camera with photometric and structured lighting, and a neural network that directly estimates distributed contact forces from images.
Results
0.4 mm localization accuracy, 0.03 N force magnitude accuracy, and 5° force direction accuracy were reported, with discrimination of up to five simultaneous contact regions.
Takeaways & Limitations
Insight provides directional force maps across its thumb-shaped 3D surface and its design concepts can be adapted to varied robot parts.
Takeaways & Limitations
The data-driven approach requires a precise test bed to collect reference data, although the test bed can support different geometries with a geometric model.
Abstract
from arXiv · showhide
Vision-based haptic sensors have emerged as a promising approach to robotic touch due to affordable high-resolution cameras and successful computer-vision techniques. However, their physical design and the information they provide do not yet meet the requirements of real applications. We present a robust, soft, low-cost, vision-based, thumb-sized 3D haptic sensor named Insight: it continually provides a directional force-distribution map over its entire conical sensing surface. Constructed around an internal monocular camera, the sensor has only a single layer of elastomer over-molded on a stiff frame to guarantee sensitivity, robustness, and soft contact. Furthermore, Insight is the first system to combine photometric stereo and structured light using a collimator to detect the 3D deformation of its easily replaceable flexible outer shell. The force information is inferred by a deep neural network that maps images to the spatial distribution of 3D contact force (normal and shear). Insight has an overall spatial resolution of 0.4 mm, force magnitude accuracy around 0.03 N, and force direction accuracy around 5 degrees over a range of 0.03--2 N for numerous distinct contacts with varying contact area. The presented hardware and software design concepts can be transferred to a wide variety of robot parts.
1 Introduction
Robots need robust perception of where and how they contact objects, but distant cameras are poorly suited to occluded, small-scale deformations and existing haptic sensors often lack practical designs or rich force information. Insight addresses these gaps with a compact, soft, vision-based sensor that provides all-around force sensing.
- Cameras and computer vision are poorly suited to robot contact perception because of occlusion and the small scale of contact deformations.
- Existing haptic sensors use diverse transduction methods, while vision-based designs commonly remain fragile, bulky, insensitive, inaccurate, or expensive.
- Insight provides a force map across a 3D surface, enabling detailed directional information about simultaneous contacts rather than only single-contact localization or force magnitude.
- Insight is designed as a soft, thumb-sized, durable, compact, sensitive, accurate, affordable sensor with all-around force-sensing capabilities.Its reported cost is less than 100 USD.
- 0.4 mm localization accuracy, 0.03 N force accuracy, and 5° directional estimation error are reported for hemispherical indentation up to 2.0 N.
2 Results
Insight combines a soft-stiff mechanical structure, single-camera photometric and structured-light imaging, and data-driven force inference to provide distributed directional sensing across a 3D surface. The prototype achieves accurate localization and force estimation, while revealing performance boundaries near the stiff frame, camera blind regions, and higher forces.
- Mechanics: Insight uses a flexible elastomer over-molded on an aluminum skeleton to combine compliant contact, sensitivity, shape stability, and robustness under high forces.The elastomer is around 70 kPa, while the aluminum skeleton is around 70 GPa in Young’s modulus.
- Imaging: A single internal camera combines photometric stereo and structured light through a collimator to detect deformation across the full 3D conical shell.Shading captures orientation changes, while structured-light color changes detect surface displacement relative to the camera.
- Force inference: The adapted ResNet maps camera images directly to distributed 3D force maps, avoiding a stiffness-based calibration chain that is difficult for inhomogeneous, nonlinear surfaces.The force map represents normal and shear components at points across the sensing surface.
- Accuracy: direct contact estimation: 0.4 mm localization precision, approximately 0.03 N force precision, and approximately 5° force-direction precision were achieved for single contacts up to 2.0 N.These values were measured on test contact points absent from training data.
- Accuracy: direct contact estimation: Accuracy remained stable across the sensing surface overall, but errors increased near the stiff frame and in two camera-hidden base regions.Force accuracy also worsened above 1.6 N, plausibly because of sparse high-force training data and limited deformation near the frame.
- Force-map performance: 0.08 N median total-force error was reported for force-map estimation, with force underestimation increasing at higher forces.A 12 mm indenter produced 1 mm median position accuracy and 8° median force-direction error; the tactile fovea achieved 0.3 mm position and 0.026 N force errors.
- Multiple simultaneous contacts: Up to five simultaneous contact points were consistently discriminated, with each contact area estimated in a visually accurate manner.Two-finger pinching and twisting were used to demonstrate qualitative force-map behavior during complex contact.
3 Discussion
Insight provides directional force maps across a thumb-shaped surface with fine localization and force accuracy, while supporting multiple contacts and low-force detailed-shape sensing. Its data-driven approach models nonlinear effects automatically, but dynamic motion and inertial deformation constrain some applications.
- Results: 0.4 mm localization, 0.03 N force-magnitude, and 5° force-direction accuracy characterize Insight’s directional force-map output.The sensor also discriminated up to five simultaneous contact regions in evaluation.
- Discussion: End-to-end learning models deformation and optical effects automatically instead of relying on calibrated linear-elastic force computation.The trade-off is a requirement for a precise test bed to collect reference data.
- Limitations: High angular velocities and accelerations may produce tactile artifacts from inertial deformation of the inhomogeneous soft sensing surface.The authors indicate that data collected during dynamic trajectories could potentially mitigate these effects.
- Generalization: The hardware, machine-learning architecture, training, and inference processes can be adapted to robot parts with different shapes and precision requirements.Design parameters such as camera field of view and light-source arrangement can be adjusted for other applications.
II III
Figure 5 evaluates Insight’s self-posture recognition and tactile-fovea sensitivity to detailed shapes. It combines posture inference from image differences with sharpness and edge-discrimination tests.
- A: posture recognition: Posture experiments estimate roll and yaw by mapping current-reference image differences to posture coordinates.The figure includes the setup, inference procedure, and pixel-wise RMS image differences across sensor rotations.
- A: posture recognition: The figure summarizes yaw accuracy and roll accuracy under different yaw angles using statistical plots and a fitted yaw curve.The caption identifies the yaw plot on the left and roll plot on the right.
- B: shape detection: Tactile-fovea tests probe sharpness with a 150° V-shaped wedge and edge count with a nine-sided polygon using a 6 mm indenter.Raw images and their red, green, and blue channels are shown alongside the test configurations.
4 Methods
The methods combine a cone-shaped, camera-based soft sensor with controlled force measurements, structured illumination, and neural-network inference. Experiments compare sensing behavior, operating characteristics, and robustness with existing vision-based sensors.
- Optical design: Eight tri-color LEDs and a 3D-printable collimator provide the structured illumination used by the imaging system.The LEDs are arranged circumferentially with programmed red, green, and blue outputs.
- Mechanical design: A single EcoFlex 00-30 elastomer layer is over-molded around an aluminum skeleton designed for sensitivity, robustness, and low fatigue.Finite-element analysis supports positioning the skeleton closer to the elastomer’s inner surface to improve force sensitivity.
- Data collection: A five-degree-of-freedom test bed applies controlled normal and shear indentations while force-torque measurements and camera images are recorded.The probe moves in 0.2 mm normal steps and applies shear with a 2:1 normal/shear movement ratio.
- Machine learning: An adapted ResNet maps images to contact position, amplitude, and force-distribution outputs using location-split training, validation, and test data.Single-contact data include 187,358 samples from 3,800 randomly selected initial contact locations.
- Evaluation and comparison: Insight’s reported robustness exceeds 400,000 interactions without noticeable damage or performance change, while its current force-map inference runs at 10 fps.The comparison discusses fragility and wear effects in alternative vision-based sensors.
Supplementary Materials for A soft thumb-sized vision-based sensor with
The supplementary materials provide additional analyses of Insight’s imaging, mechanics, data interpretation, machine learning, functionality, and ablation studies.
- Supplementary sections: Supplementary Section A.1 covers imaging-system design, while A.2 covers mechanical tests for the sensor design.These sections begin on pages 26 and 30, respectively.
- Supplementary sections: Supplementary Sections A.3 and A.4 address data interpretation and machine-learning details.They begin on pages 32 and 40, respectively.
- Supplementary sections: Supplementary Sections A.5 and A.6 provide functionality illustrations and ablation studies.They begin on pages 41 and 44, respectively.
A.1 Imaging system design
Insight uses a wide-angle fisheye camera and programmable, collimated LED illumination to image its three-dimensional sensing surface. The camera’s field of view varies by operating mode and is not equal horizontally and vertically.
- Camera geometry and field of view: 40 mm diameter and 70 mm height: The cone-shaped geometry supports distributed 3D haptic sensing within the camera’s field of view.The shape is similar to a human thumb and can be adapted by changing the sensor’s size and shape.
- Camera geometry and field of view: 160°: A fisheye lens provides the wide-angle field of view needed to view much of Insight’s inner surface.Mode 1 was selected mainly to maintain maximal field of view.
- Camera geometry and field of view: The rectangular imaging sensor gives unequal horizontal and vertical viewing capabilities.The difference is shown for the fisheye camera’s operating modes and summarized in Table S1.
- Lighting system: Eight LED units with programmable red, green, and blue channels generate the structured-light illumination.The LED ring emits light into the half 3D space around the camera.
- Lighting system: Collimator diameter D controls light-cone size, while tilt angle α slants the cone radially outward.These parameters are adjusted to shape the projected illumination pattern.
A I II
The sensing surface is tuned through light-attenuation analysis and elastomer composition tests so projected patterns remain observable across the soft shell. The selected illumination uses collimated cones and compensates for unequal camera sensitivity to color.
- Lighting analysis: Approximately quadratic attenuation with distance and approximately linear attenuation with reduced brightness characterize the received light intensity.The projected cone widens with distance, while the camera sees a smaller portion of reflected beams; FWHM is used to calculate cone size.
- Lighting analysis: 1 : 1 : 2: The camera’s sensitivity ratio is measured as red, green, and blue, respectively.The sources are arranged as R, G, B, R, G, R, B, G in response to this analysis.
- Surface composition: The material tests mix aluminum powder, aluminum flakes, and pigments into the elastomer to tune the sensing surface’s reflective properties.The goal is to avoid regions that appear too dark or too bright.
- Surface composition: Aluminum powder gives adequate but relatively dark images, whereas aluminum flakes readily create saturated points through strong specularity.Black pigment absorbs almost all light, making the camera unable to see much of the surface.
A.2 Mechanical tests for the sensor design
The mechanical design combines a soft EcoFlex shell with a stiff skeleton, using over-molding and relative positioning tests to balance durability, robustness, and sensitivity. The design also considers weight, yield strength, and the arrangement of soft sensing areas.
- Soft shell: EcoFlex 00-30: This elastomer is selected for its Young’s modulus, density, curing time, and compatibility with degassing.The material is chosen from the Smooth-On EcoFlex series for weight, durability, and elongation properties.
- Skeleton design: The beam skeleton is designed around weight, robustness, and yield strength while spacing soft sensing areas across the surface.The arrangement also mimics human fingertips by separating soft areas with rigid-frame beams.
- Over-molding: 0.8 mm: This is the minimum over-molding thickness found to provide a robust elastomer–skeleton connection without defects.The test checks whether the elastomer covers a cylinder without defects.
- Relative positioning: An internal skeleton offset near the elastomer surface increases sensitivity by causing greater displacement.This relationship is evaluated with a finite element model based on the tested material properties.
A.3 Data interpretation
Insight’s data interpretation combines force-distribution approximation, learned force-map inference, and evaluations of morphology, accuracy, force direction, and posture sensitivity.
- Surface morphology: Ridges improve localization and force quantification while accelerating machine-learning training and making surface motion easier to track.The comparison is between smooth and ridged internal surfaces.
- Force map evaluation: 0.4 mm and 0.03 N are the direct-estimation median errors for contact localization and force, respectively.Force-map inference is less accurate, with median errors of 0.6 mm and 0.08 N.
- Force map evaluation: 5° is Insight’s average directional estimation error when estimating normal and shear force components.The force map evaluation reports good correspondence between true and predicted force directions.
- Posture sensitivity: 2.5° yaw and 11.6° roll are the mean posture-estimation accuracies obtained from gravity-induced deformations.Roll error is highest when the roll axis aligns with gravity, while gravitational deformation does not significantly affect contact perception.
A.4 Machine Learning Details
Insight uses customized ResNet-18 models to convert camera inputs into direct contact estimates, 3D force maps, and posture predictions.
- Inference formats: Raw images are transformed into contact location, directional force, force-map, and posture outputs through machine-learning inference.The system supports both direct single-contact inference and a three-dimensional force map over the conical surface.
- Direct single-contact inference: Direct inference predicts three contact coordinates and three local force components from RGB differences and a skeleton image.The six output channels represent x, y, z position and shear and normal force components.
- Force map inference: Force-map inference replaces the fully connected output with convolutional operations producing a 64 × 64 map with three force channels.The architecture uses two ResNet blocks followed by up-scaling, convolution, normalization, activation, and force-map projection.
- Network design: 1.5 million parameters support spatial-map prediction while preserving a 52 × 39 intermediate resolution after two ResNet blocks.The architecture is designed around local correspondence between input regions and output pixels.
- Posture inference: ResNet-18 posture inference uses three-channel difference images to predict yaw and roll angles.The same standard architecture is used with task-specific inputs and outputs.
A.5 Functionality illustration
Functionality tests show Insight can detect multiple contacts, recognize shapes and posture, and localize sliding contacts while estimating force direction.
- Multiple contacts: Up to four nearby contacts are demonstrated, although the sensor can theoretically discriminate 16 contacts simultaneously.The theoretical limit corresponds to the number of hollow areas formed by the skeleton.
- Posture recognition: Gravity-induced deformation permits posture recognition, while the reported gravitational effect remains small relative to typical external contacts.The posture experiment rotates the sensor around roll at a yaw angle of 90°.
- Shape detection: 150° V-shaped wedge sharpness and about 9 or 10 polygon edges can be visually discriminated.These tests supplement the shape-detection evaluation shown elsewhere in the paper.
- Dynamic evaluation: Sliding contacts are accurately localized and their changing force direction is discriminated during repeated forward and backward motion.The evaluation includes sliding over 4 mm followed by pauses and repeated cycles.
A.6 Ablation Studies
Ablation studies examine dataset size and network inputs, showing that structured-light color and photometric stereo contribute to force and localization accuracy.
- Dataset size: Dataset-size ablations evaluate direct force prediction, force-map prediction, and posture prediction using progressively larger training subsets.The study is motivated by the data demands of machine-learning-driven sensors.
- Network input: Input ablations compare raw images, reference images, skeleton images, and their combined representation.The chosen input is [Img - Ref, Skel].
- Network input: Removing color from structured light lowers force accuracy, while retaining photometric stereo achieves 0.6 mm average localization accuracy.The 0.6 mm localization result is reported as one order of magnitude better than GelTip’s 5 mm.