Source-linked AI summary

Shape-independent Hardness Estimation Using Deep Learning and a GelSight Tactile Sensor

Wenzhen Yuan, Chenzhuo Zhu, Andrew Owens, Mandayam A. Srinivasan, Edward H. Adelson

arXiv:1704.03955v1cs.RO

TL;DR

Robotic hardness estimation is limited by tactile sensors that provide insufficient information and often requires controlled contact or object geometry. This paper uses GelSight video with convolutional and recurrent neural networks to estimate hardness across shapes and loading conditions, achieving useful estimates for many smooth, simple natural objects while struggling with complicated or ridged shapes.

  • Problem

    Robotic techniques for estimating object hardness are limited, while existing approaches may require strict control of sensor movement and object geometry.

  • Method

    A deep neural network combines convolutional and recurrent processing to regress hardness directly from high-resolution GelSight video sequences under varied contact conditions and object geometries.

  • Results

    The model estimates hardness for silicone samples with similar shapes regardless of loading conditions and roughly measures hardness for many natural objects with smooth surfaces and simple geometries.

  • Takeaways & Limitations

    The method can support recognition of objects with special hardness or selection of fruits by preferable ripeness level within its demonstrated scope.

  • Takeaways & Limitations

    The sensor cannot well differentiate hardness above 70 in Shore 00, and performance decreases for complicated or ridged shapes because the training data are incomplete.

Abstract

from arXiv · show

Hardness is among the most important attributes of an object that humans learn about through touch. However, approaches for robots to estimate hardness are limited, due to the lack of information provided by current tactile sensors. In this work, we address these limitations by introducing a novel method for hardness estimation, based on the GelSight tactile sensor, and the method does not require accurate control of contact conditions or the shape of objects. A GelSight has a soft contact interface, and provides high resolution tactile images of contact geometry, as well as contact force and slip conditions. In this paper, we try to use the sensor to measure hardness of objects with multiple shapes, under a loosely controlled contact condition. The contact is made manually or by a robot hand, while the force and trajectory are unknown and uneven. We analyze the data using a deep constitutional (and recurrent) neural network. Experiments show that the neural net model can estimate the hardness of objects with different shapes and hardness ranging from 8 to 87 in Shore 00 scale.

I. INTRODUCTION

The paper addresses limited robotic hardness estimation by using GelSight’s deformation-sensitive tactile images under loosely controlled contact. A deep convolutional and recurrent model estimates hardness across multiple object shapes.

  • Hardness describes a solid’s resistance to compressive force and is commonly measured through indentation relative to applied force.
  • Robotic hardness estimation is limited because force-based methods require strict control of object geometry and robot movement.
  • GelSight combines a soft elastomer and embedded camera to capture high-resolution contact geometry and approximate force, resembling cutaneous fingertip sensing.
  • The method trains a deep neural network to regress hardness directly from raw GelSight video, using convolutional features and recurrent modeling of gel deformation over time.
  • 8 to 87 in Shore 00 scale: experiments tested silicone objects across this hardness range using manual or robot-mediated pressing.
  • The model estimates basic and natural-object hardness across varied shapes, including samples with different hemisphere or cylinder radii.

II. RELATED WORK

Related work contrasts conventional tactile sensing with optical approaches. GelSight is presented as an optical sensor designed to recover high-resolution contact shape.

  • Piezoresistive, piezoelectric, and capacitive tactile sensors commonly measure local force but can have limited spatial resolution or sensing area.
  • GelSight uses a clear elastomeric slab, reflective membrane, camera, and illumination to obtain a highly accurate deformation height map through photometric stereo.

B. Hardness Measurement

Prior hardness-measurement methods generally rely on controlled force, deformation, or specialized indentation sensing. The paper instead targets direct hardness estimation with more varied geometries and contact conditions.

  • Direct tactile hardness estimation has received less attention than inferring object properties from sequences of contact-force signals.
  • Conventional approaches apply controlled force and measure deformation, requiring tightly controlled sensor movement and object geometry.
  • The present approach directly estimates hardness rather than distinguishing samples and uses more diverse geometries and hardness values.
  • Specialized sensors can estimate hardness by combining force measurements with indentation depth measurements.
  • The authors’ prior GelSight model estimated hemispherical silicone hardness under loosely controlled contact but remained limited to hemispherical samples.

C. Learning for tactile sensing

Learning-based tactile methods infer properties from temporal sensor data. This paper uses GelSight’s visual deformation and marker-motion cues with a recurrent architecture to estimate hardness.

  • Learning-based tactile systems have used recurrent networks to infer qualitative haptic adjectives from time-varying sensor measurements.
  • Hard and soft objects produce visibly different GelSight deformations: hard objects deform little, while soft objects flatten under smaller force.
  • The network represents GelSight frames with VGG16 CNN features and feeds them into an LSTM to model temporal information.
  • GelSight videos encode contact-surface normals through color intensity and contact force through motion of embedded black markers.

A. Neural network design

The model maps GelSight image sequences to hardness estimates by combining convolutional image features with recurrent temporal processing. It produces per-timestep predictions and averages the final three frames for an object-level estimate.

  • GelSight images are represented using penultimate-layer fc7 features from a VGG convolutional network.
  • An LSTM updates its hidden state from the current image features and applies an affine transformation to regress hardness at each timestep.The transformation uses parameters W and b.
  • The prediction y_t is the hardness estimate for the current timestep.
  • The object-level hardness estimate averages predictions from the final 3 frames.
  • Training uses a Huber loss penalizing differences between predicted and ground-truth hardness values on a per-frame basis.Per-frame regression is intended to add robustness when pressing motions differ from the training set.

B. Choosing input sequences

The input sequence is selected from the loading period using intensity-based markers so predictions are less dependent on pressing speed and maximum contact force. Initial-frame subtraction removes preexisting elastomer deformation, while training includes varied contact conditions.

  • A 5-frame loading-period sequence is selected between an intensity-threshold start and the last peak-intensity-change frame.Three intermediate frames are evenly selected according to intensity change.
  • The sequence-selection procedure is designed to make the method invariant to pressing speed and maximum contact force.
  • Subtracting the first frame accounts for preexisting deformations in the GelSight elastomer.
  • The training set contains about 7000 independent human-pressing videos, including basic shapes, complicated shapes, and bad contact conditions.Irregular data are used to help prevent model overfitting.

IV. EXPERIMENTAL SETUP

The experiments use a marked Fingertip GelSight with a soft elastomer and silicone samples spanning multiple geometries and hardness levels. Hardness is calibrated with durometer-tested standardized samples, while the sensor’s own hardness limits its sensitive range.

  • The Fingertip GelSight uses black surface markers to track displacement and captures tactile images with an embedded camera.
  • The sensor records 960 × 720-pixel images over an 18.4mm × 13.8mm area at 30Hz, with a 17 Shore 00 elastomer.
  • Silicone samples include hemispheres, cylinders, and arbitrary shapes made from three materials with varied mixing ratios.
  • Samples span 16 hardness levels from 8 Shore 00 to 45 Shore A, equivalent to 87 Shore 00.
  • The GelSight elastomer’s 17 Shore 00 hardness improves discrimination of very soft objects but limits differentiation above 70 Shore 00.

V. EXPERIMENTAL PROCEDURE

Experiments collect tactile videos during human pressing and robot-gripper squeezing under deliberately varied contact conditions. The dataset covers basic and more complex geometries, with training and testing groups arranged to evaluate generalization.

  • Contact procedures: Human testers press GelSight vertically, while a Weiss WSG 50 gripper squeezes samples at 5–7 mm/s until a randomly selected 5–9N force threshold.Each pressing period averages 20–30 frames.
  • Contact procedures: Both experimental settings produce highly varied contact conditions, and human contact is uncontrolled with small shear force and torque.This setup aims to model natural tactile interaction.
  • Dataset groups: The dataset includes basic shapes, undesired contact conditions, natural objects with simple geometry, complicated chocolate-mold shapes, tomatoes, and candies.
  • Dataset groups: Training uses selected samples from Groups 1, 2, and 4, while testing includes Groups 1, 3, 4, and 5, with both human and robot data in the test set.
  • Evaluation: Figure 6 reports prediction results for basic shapes including hemispheres, cylinders, flat surfaces, edges, and corners.
  • Evaluation: Table I presents network predictions on basic shapes.

VI. EXPERIMENTAL RESULTS

The experiments test whether the model generalizes to new hardness values, unseen shapes, and changed experimental conditions across three test groups.

  • The first experiment tests generalization to new hardness values using test objects with training-set shapes but different hardness ratings.
  • The second experiment evaluates cylindrical or spherical samples whose shapes are absent from the training set.
  • The third experiment uses basic shapes seen during training but changes the experimental conditions.

B. Arbitrary Shapes

On arbitrary shapes, performance declines relative to familiar shapes, with errors concentrated around sharp or multiple ridged surface curvatures; natural-object estimates are more reliable for smooth, simple geometries.

  • B. Arbitrary Shapes: The model often overestimates hardness for unfamiliar shapes with sharp curvatures or multiple ridges.
  • B. Arbitrary Shapes: For shapes included in training, the neural network can estimate the target hardness well.
  • B. Arbitrary Shapes: Basic-shape training generalizes to more complex shapes, but performance decreases and broader shape variation requires a much larger training set.
  • C. Estimation of Natural Objects: For natural objects, estimates align with human hardness rankings for tomatoes and can distinguish ripeness levels despite close hardness values.
  • C. Estimation of Natural Objects: Natural objects with simple, smooth geometries are estimated well, whereas textured or ridged surfaces tend to appear harder because training data are incomplete.

VII. CONCLUSION

The paper uses GelSight image sequences with convolutional and recurrent neural networks to estimate hardness across object shapes. The model performs well under varied loading for familiar silicone shapes and offers rough hardness estimates for smooth natural objects, but struggles with complicated or ridged surfaces.

  • The proposed system estimates object hardness across multiple shapes using a GelSight tactile sensor.
  • GelSight records high-resolution deformation image sequences containing information about shape change and contact force.
  • A convolutional neural network and recurrent neural network extract hardness information from GelSight video sequences.
  • The network predicts hardness well for silicone samples with similar dataset shapes regardless of loading conditions.
  • For complicated shapes or ridged surfaces, hardness estimation is poor because the training dataset is incomplete.
  • For many smooth, simply shaped natural objects, the model roughly measures hardness and can support recognizing special hardness or preferable fruit ripeness.
Loading 1704.03955v1…