Source-linked AI summary
DIGIT: A Novel Design for a Low-Cost Compact High-Resolution Tactile Sensor with Application to In-Hand Manipulation
Mike Lambeta, Po-Wei Chou, Stephen Tian, Brian Yang, Benjamin Maloon, Victoria Rose Most, Dave Stroud, Raymond Santos, Ahmad Byagowi, Gregg Kammerer, Dinesh Jayaraman, Roberto Calandra
TL;DR
Robots still struggle with general-purpose in-hand manipulation partly because precise contact-force sensing is difficult. This paper introduces DIGIT, a compact, inexpensive, high-resolution tactile sensor with improved manufacturing and reliability, and demonstrates tactile model-based control for in-hand marble manipulation. The design and manufacturing process are open-sourced to support access to the platform.
Problem
General-purpose in-hand manipulation remains difficult, partly because robots struggle to precisely sense contact forces needed to control environmental interactions.
Method
The paper designs DIGIT and trains model-based controllers from raw tactile inputs to manipulate glass marbles with a multi-finger robotic hand.
Results
About 25% of trials resulted in the marble dropping before it was fully manipulated to the goal.
Takeaways & Limitations
DIGIT provides a compact, high-resolution tactile platform demonstrated on challenging in-hand marble manipulation, with its design and manufacturing process released openly.
Abstract
from arXiv · showhide
Despite decades of research, general purpose in-hand manipulation remains one of the unsolved challenges of robotics. One of the contributing factors that limit current robotic manipulation systems is the difficulty of precisely sensing contact forces -- sensing and reasoning about contact forces are crucial to accurately control interactions with the environment. As a step towards enabling better robotic manipulation, we introduce DIGIT, an inexpensive, compact, and high-resolution tactile sensor geared towards in-hand manipulation. DIGIT improves upon past vision-based tactile sensors by miniaturizing the form factor to be mountable on multi-fingered hands, and by providing several design improvements that result in an easier, more repeatable manufacturing process, and enhanced reliability. We demonstrate the capabilities of the DIGIT sensor by training deep neural network model-based controllers to manipulate glass marbles in-hand with a multi-finger robotic hand. To provide the robotic community access to reliable and low-cost tactile sensors, we open-source the DIGIT design at https://digit.ml/.
I. INTRODUCTION
DIGIT addresses the lack of tactile sensors that simultaneously provide high resolution, sensitivity, reliability, usability, compactness, and low cost. The paper presents its design and manufacturing process, demonstrates tactile-MPC manipulation with multiple sensors, and releases the platform openly.
- DIGIT targets the combined requirements of high resolution, sensitivity, reliability, ease of use, compactness, and low cost.
- The sensor uses a smaller form factor, streamlined manufacturing, enhanced mechanical reliability, modular components, and a plug-and-play software interface.
- The paper presents DIGIT’s design and manufacturing process, analyzes the resulting sensor, and demonstrates manipulation from raw tactile inputs.
- New dynamics-model-learning and task-specification approaches reduce the computational cost of scaling tactile-MPC to multiple touch sensors.
- The authors release DIGIT’s design to provide an affordable, robust, and easy-to-use tactile sensing platform.
II. RELATED WORK
Related work establishes vision-based tactile sensing as a high-resolution, sensitive, and relatively inexpensive approach, while identifying integration challenges for robotic manipulation. DIGIT builds on this class of sensors and supports tactile learning-based control.
- Vision-based tactile sensors infer contact forces from camera images of deformable elastomers and often offer high spatial resolution, sensitivity, and lower manufacturing cost.
- Existing vision-based sensors include TacTip, FingerVision, GelSight, and other designs, with GelSight using reflective elastomers and surface markers.
- DIGIT improves on existing GelSight sensors through compactness, more durable elastomer gel, and design changes supporting repeatable large-scale production.
- A key bottleneck is extracting and integrating meaningful features from high-resolution tactile sensors into control algorithms.
- Prior in-hand manipulation approaches used model-free reinforcement learning, tracking cameras, or low-cost tactile feedback for selected dexterous tasks.
III. DIGIT: A LOW COST, COMPACT, HIGH-RESOLUTION TACTILE SENSOR
DIGIT retains the high-resolution sensing advantages of vision-based tactile sensors while addressing their bulk, wear, and manufacturing variability. Its design combines compactness, interchangeable components, robust materials, and low-cost repeatable production.
- Previous vision-based tactile sensors are limited by bulky form factors, surface wear, and complex manual manufacturing with high sensor-to-sensor variability.
- DIGIT fits arrays of end effectors or multi-fingered robot arms while retaining the advantages of vision-based tactile sensing.
- Its gel is designed to be more robust and interchangeable, producing a more rugged sensor.
- Automated, tool-less assembly and commercial off-the-shelf components support rapid, repeatable, large-scale manufacture.
- 15 USD is DIGIT’s estimated manufacturing cost per sensor when produced in a batch of 1000.
A. Mechanical Design
DIGIT uses a compact, modular mechanical design with interchangeable components and elastomers. Its custom electronics fit in an area only slightly larger than a human fingertip.
- Dimensions and enclosure: DIGIT measures 20 mm × 27 mm × 18 mm, weighs approximately 20 g, and uses a three-piece plastic enclosure.The enclosure supports 3D-printing for prototyping and injection molding for production.
- Modularity: Press-fit connections allow the camera, gel, and acrylic-gel unit to be swapped after breakage or wear.The plastic housing can also be changed to support different focal lengths.
- Interchangeable elastomers: Three interchangeable elastomers—reflective, reflective with markers, and transparent with markers—support application-specific tactile readings.Reflective elastomers support texture measurement, while marked elastomers support optical-flow computation.
- Custom electronics: Custom electronics controlling the camera, illumination, and video capture fit within 7 cm^2 and connect multiple DIGITs through one USB port.The camera captures color video at 60 fps.
C. Elastomer Design
DIGIT’s elastomer combines a layered silicone construction with a controlled manufacturing process intended to improve reliability and production scalability. The design balances sensing resolution against resistance to damage.
- Layered elastomer: The contact image-transfer layer is vulnerable to wear because its deformations provide the camera-based tactile percepts.Repeated use can alter the characteristics of vision-based tactile sensors.
- Manufacturing process: DIGIT’s elastomer uses a white silicone image-transfer layer, a silicone base layer, and an acrylic window assembled with optically clear adhesive.The image-transfer layer is airbrushed into the mold and cured to achieve controlled uniform thickness.
- Reliability evaluation: Abrasion reliability is measured as the percentage increase in transmittance, with higher values indicating greater coating wear.The DIGIT elastomer shows the lowest degradation among the evaluated gels.
- Thickness trade-off: Thicker image-transfer layers deform less and lose spatial resolution, whereas thinner layers are more prone to damage.The design iterates over layer thickness to trade off ruggedness and sensitivity.
D. Mechanical Robustness of the Elastomer
DIGIT’s elastomer was evaluated against two comparison gels using standardized abrasion testing. After five passes, DIGIT’s gel remained nearly unaffected while both alternatives sustained significant damage.
- Abrasion results: After five linear abrasion passes, DIGIT’s gel was nearly unaffected, while the gels from Yuan et al. and GelSight Inc. showed significant damage.The comparison used a 1.7 N calibrated weight and an H-18 Calibrade abrasive tip.
- Abrasion results: Damage to the two comparison gels produced large tears and removal of surface material, rendering them unusable.The figure shows visible coating-layer damage in both comparison gels.
- Manipulation relevance: DIGIT’s compact form factor allows high-resolution camera-based tactile sensors to be mounted on all fingers of an Allegro robotic hand.The paper uses this configuration for precision-grip marble manipulation.
A. Task Setup
The task setup uses an Allegro hand mounted on a Sawyer arm to manipulate a marble from a pincer grip to target configurations. Autonomous randomized trials provide data for learning tactile dynamics and control.
- Task setup: A linear motor raises each marble, and the arm performs a preprogrammed pickup before the hand learns to roll its fingers over the marble.The manipulation requires modeling slipping and rolling over curved, deformable DIGIT surfaces under varying pressure.
- Self-supervised data collection: Data collection used 4800 trials, with 950 trials reserved for validation.
- Learned control pipeline: The encoder detects the marble position from tactile observations, while a forward dynamics model predicts its next position for model predictive control.At each step, an optimizer selects action sequences intended to move the marble toward a specified target.
- Learned control pipeline: The first optimized action is applied after selecting the best action sequence for moving the marble from its current position to the desired target.
- Self-supervised data collection: Each trial applies 20 random angular-displacement commands to four joint servos per finger, producing an 8-D action space over approximately 10 seconds.Videos, eight servo positions, and commanded displacements are recorded from both DIGITs.
C. Tactile Predictive Model
The tactile predictive model represents DIGIT observations through object keypoints, then learns dynamics over that compact state for video prediction.
- C. Tactile Predictive Model: A structural bottleneck autoencoder learns keypoints that capture factors of variation in tactile images.The encoder produces K feature maps, while the decoder reconstructs images from Gaussian blobs centered at the predicted keypoints.
- C. Tactile Predictive Model: Each keypoint contains a 2D maximum-activation location and an intensity scalar.The representation is k = [x, y, i], with x and y denoting location and i denoting average activation magnitude.
- C. Tactile Predictive Model: The learned dynamics model predicts next state s′ = f(s, a) from the current state and action.Training uses sampled (s, a, s′) tuples, augmented with zero-action self-transitions and lighting perturbations.
- C. Tactile Predictive Model: Because the keypoint state is fully observable, the authors use an MLP instead of the more complex VRNN.The model is trained with zero-action tuples and RGB/gamma augmentation to improve robustness to lighting changes.
D. Model-based Control
The controller uses model-predictive control with cross-entropy optimization to plan marble manipulation in tactile keypoint space.
- D. Model-based Control: MPC with the cross-entropy method plans action sequences using the learned dynamics model.The setup has two DIGIT images and 8 degrees of freedom, making the planning search substantially more complex than the earlier one-sensor, 3-DOF setting.
- D. Model-based Control: The goal DIGIT image is mapped into keypoint space as a target for the left or right finger.In the experiments, target marble locations are directly supplied as keypoint locations.
- D. Model-based Control: The planning cost sums Euclidean distances between current and target positions in (x, y, i) coordinates.This encourages reaching the desired x,y location while avoiding marble drops or excessive pressing force.
- D. Model-based Control: The model-based controller is evaluated on the in-hand tactile manipulation task after separate image-quality and gel-robustness tests.The evaluation concerns manipulating a marble between fingers using DIGIT tactile observations.
A. Video Predictive Model
The video-predictive model produces qualitatively good tactile predictions and enables learned control that moves marbles toward goals more reliably than a hand-tuned proportional controller, although drops remain.
- A. Video Predictive Model: 1.4 seconds per MPC step is required with the Struct-NN keypoint dynamics model in the multi-finger marble-manipulation setup.Each step evaluates about 0.3 million forward passes through the dynamics model using 250 particles, horizon 10, and roughly 120 CEM iterations.
- A. Video Predictive Model: Struct-NN produces qualitatively good predictions but slightly higher per-pixel RMSE than CDNA on BAIR pushing and tactile marble videos.The comparison also considers model size, training time, inference time, and one-step MPC time.
- B. Manipulating Marbles: The learned dynamics controller reduces Euclidean distance to the desired marble goal over time, whereas the hand-tuned P controller increases it.The trajectories are evaluated in terms of (x, y) distance to goal across performed actions.
- B. Manipulating Marbles: The marble trajectories demonstrate that MPC can move the marble accurately toward the desired tactile-space goal.Figure 9 marks the goal with a red dot and the current keypoint with a green dot.
- B. Manipulating Marbles: About 25% of trials end with the marble dropping before complete manipulation to the goal.The authors attribute this partly to planning inaccuracies and hand-joint actuation noise, and hypothesize that better low-level control and more data could help.
- B. Manipulating Marbles: The evaluation supports DIGIT’s high-resolution tactile sensing and Struct-NN’s ability to scale tactile MPC to this task.The claim concerns fine-grained multi-finger in-hand manipulation with the demonstrated setup.
VI. CONCLUSION
The paper presents DIGIT as a compact, high-resolution tactile sensor with improvements in reliability, assembly, component availability, and manufacturing cost. It demonstrates marble manipulation from raw tactile inputs and releases the design and manufacturing process openly.
- VI. CONCLUSION: DIGIT provides rich, high-resolution tactile readings in a compact sensor design.The conclusion also reports improvements in reliability, component availability, ease of assembly, and manufacturing cost.
- VI. CONCLUSION: The authors learn to manipulate glass marbles toward desired target positions from raw tactile inputs using model predictive control.The design and manufacturing process are open-sourced at www.digit.ml.
- VI. CONCLUSION: Future work should further miniaturize DIGIT and develop curved, omni-directional sensing fields.These are stated directions for extending the sensor design.