Source-linked AI summary
$\mathcal{N}_0$-Foundation: Towards the Age of Tactile Intelligence
NeoteAI Team, Fudan TEAI Team
TL;DR
Existing manipulation corpora provide limited tactile coverage and fragmented sensor-specific signals, despite contact-rich tasks requiring physical contact information. N0-Foundation addresses this gap with integrated tactile infrastructure, NeoData, NeoForce, and paired real-sim evaluation; experiments indicate that tactile feedback improves contact-rich manipulation and that force-based representations support transfer across sensor designs.
Problem
Existing embodied datasets are dominated by vision, while tactile signals remain fragmented across hardware-specific formats for contact-rich manipulation.
Method
N0-Foundation combines tactile collection hardware, cross-embodiment NeoData, NeoForce force-field representation learning, and NeoReal–NeoSim evaluation.
Results
Experiments indicate that tactile feedback improves contact-rich manipulation, while latent supervision improves temporal tactile representation learning and NeoForce provides transferable force representations.
Takeaways & Limitations
The integrated resources provide a foundation for more capable and transferable tactile-aware manipulation policies.
Takeaways & Limitations
The unified representation currently targets parallel-jaw visuo-tactile fingers rather than dexterous hands or non-camera-based transducers.
Abstract
from arXiv · showhide
We present $\mathcal{N}_0$-Foundation, a paradigm for tactile-enabled embodied manipulation, which integrates tactile sensing hardware, large-scale multimodal data, tactile representation learning, and standardized evaluation. First, we engineer the infrastructure for scalable data collection, including a vision-based tactile sensor, a tactile Universal Manipulation Interface (UMI), and a synchronized visuo-tactile data collection system supporting both robot embodiments and UMI-based demonstrations. Leveraging this infrastructure, we construct NeoData, which contains more than 30000 hours of synchronized visual and tactile demonstrations, spanning six embodiments, 450 tasks, and billions of paired RGB and tactile frames collected through a mixture of real-robot teleoperation and UMI-based demonstrations. To facilitate open research, we further release OpenNeoData, a 5000-hour open-source subset of NeoData. The dataset addresses a central limitation of existing manipulation corpora, critical for deformable-object manipulation, precise assembly, delicate force control, and sustained surface interaction. Capitalizing on the large-scale, heterogeneous tactile measurements, we propose NeoForce, a visuo-tactile representation model that learn transferable tactile representations across different sensor designs. To enable systematic evaluation of tactile embodied models built upon our infrastructure, datasets and tactile representations, we further propose a comprehensive benchmark, which combines the real-world NeoReal suite and the simulated NeoSim suite for standardized evaluation. Experiments across both suites show that policies benefit from the physical contact state rather than from the device-specific appearance of the tactile signal. We release the dataset, the representation, and the benchmark, aiming at supporting future work on tactile-enabled embodied manipulation.
1 Introduction
N0-Foundation addresses vision-dominated manipulation data and fragmented tactile sensing by integrating scalable infrastructure, NeoData, NeoForce, and standardized evaluation. Its resources target contact-rich manipulation across diverse embodiments and tasks.
- Motivation: Existing manipulation datasets are dominated by visual observations, while contact-rich tasks depend on force, friction, slip, and contact transitions.The limitation is especially relevant to deformable-object manipulation, precise assembly, delicate force control, surface interaction, and bimanual coordination.
- Foundation: N0-Foundation unifies tactile infrastructure, multimodal data acquisition, representation learning, and standardized policy evaluation.The four pillars are tactile infrastructure, NeoData, NeoForce, and evaluation through NeoReal and NeoSim.
- NeoData: NeoData contains more than 30,000 hours of synchronized visual–tactile demonstrations across six embodiments and 450+ tasks.The corpus combines real-robot teleoperation with portable UMI demonstrations and includes billions of paired RGB and tactile frames.
- NeoForce: NeoForce uses dense three-axis force fields to provide a transferable, hardware-agnostic tactile representation across sensor designs.The representation is intended as a common physical supervision space for tactile applications across embodied models.
- Evaluation: The NeoReal and NeoSim suites provide standardized real-world and simulated evaluation for tactile-aware policies.The proposed protocol is used to evaluate large-scale tactile learning and unified tactile representation learning.
- Implication: The framework is presented as a scalable path for incorporating tactile perception into embodied learning beyond vision alone.The stated scope is fine-grained, contact-rich manipulation capabilities.
2 Related Work
Related work spans large-scale robot datasets, reusable embodied policies, tactile representation learning, and simulation benchmarks. N0-Foundation builds on these directions while combining tactile sensing, cross-embodiment collection, and paired sim-real evaluation.
- Data Collection: UMI-style interfaces decouple demonstration collection from physical robots, but prior systems generally target few platforms and retain device-specific tactile streams.NeoData combines robot-free collection with multi-robot teleoperation under a common action and observation schema.
- Datasets: Large-scale manipulation datasets establish data scale and embodiment diversity, but many remain dominated by vision and proprioception.RH20T adds synchronized force and other modalities, while other tactile datasets remain smaller or use heterogeneous tactile encodings.
- Tactile Representation: Vision-based tactile observations are difficult to reuse because sensor optics, geometry, and elastomer appearance entangle contact with hardware-specific signals.Prior work addresses this dependence through force-based pretraining and sensor-invariant contact representations.
- Benchmarks: Existing simulation and real-world benchmarks standardize manipulation evaluation, but vision and proprioception-only protocols do not sense or score contact quality.Paired sim-real evaluation motivates the methodology adopted by NeoReal and NeoSim.
- Tactile Benchmarks: Tactile benchmarks remain scarce because simulation must reproduce both contact mechanics and sensor deformation.Existing efforts range from sensor-specific tactile simulation to multi-sensor and visuo-tactile benchmark suites.
- N0-TacUMI: N0-TacUMI combines a handheld gripper, wide-angle wrist camera, infrared tracking, and magnetic aperture sensing in one collection interface.The device records complementary visual, tactile, pose, and gripper measurements.
3 Data Acquisition Hardware
The data acquisition system combines five tactile-equipped robot platforms with a handheld N0-TacUMI interface and camera-based tactile hardware. Its synchronized streams support unified demonstrations and force-field recovery from tactile images.
- 3.1 Overview: NeoData combines five robot platforms and the handheld N0-TacUMI device under aligned tactile, visual, and action conventions.Robot demonstrations use teleoperation, while N0-TacUMI enables scalable collection beyond fixed robot workcells.
- 3.2 Tactile Sensor: The tactile sensor uses a wear-resistant glass panel, deformable elastic layer, embedded RGB camera, and controller board.External load deforms the sensing layer, and texture displacement encodes applied force for camera-based measurement.
- 3.2 Tactile Sensor: Contact pressure and shear produce texture displacement that the camera records as a tactile image.The tactile signal is therefore a visual measurement of sensing-layer deformation.
- Task Coverage: NeoData covers more than 450 tasks organized into categories including deformable-object skills, precision assembly, transport, and surface interaction.Task frequencies are represented by episode counts for constituent tasks within each category.
- 3.2 Tactile Sensor: Recovering the underlying force field from a tactile image requires inverting the composed transduction and imaging operators.The paper learns this inverse mapping and uses its output as the unified tactile representation.
- 3.1 Overview: N0-TacUMI records synchronized wrist images, left and right tactile images, gripper aperture, and relative end-effector motion.Relative pose and gripper-width actions avoid binding demonstrations to a specific robot base frame.
4 NeoData
NeoData is a large cross-embodiment visuo-tactile corpus combining robot teleoperation and handheld UMI demonstrations, with synchronized streams, broad task coverage, and extensive quality control. Its format preserves aligned visual, tactile, and action data while curation and annotation pipelines improve training readiness and structural consistency.
- Scale and scope: NeoData contains over 30,000 hours across 1.4M episodes, 3.3B timesteps, 8B RGB frames, and 10B tactile frames.The corpus spans six embodiments and was collected by nearly 100 human operators.
- Acquisition settings: The corpus combines five robot platforms with the handheld N0-TacUMI device, integrating physically grounded teleoperation with scalable UMI demonstrations.Robot demonstrations use sensors mounted on physical grippers, while N0-TacUMI provides human-operated handheld trajectories.
- Data format: Each episode synchronizes visual, tactile, action, and metadata streams, with robot records including external and wrist imagery, bilateral tactile images, and end-effector and joint commands.The action format preserves commanded operational motion alongside absolute robot state for replay or embodiment-specific analysis.
- Curation: Five complementary curation checks remove incomplete, motion-degraded, atypical or unfinished, invalid, and non-transferable demonstrations before release.The pipeline verifies stream completeness, motion quality, episode duration and completion, video validity and alignment, and inverse-kinematics feasibility within robot limits.
- Annotation: Episodes receive task, subtask, action, and atomic-segment labels through VLM template proposals, human verification, signal-based segmentation, and hierarchical labeling.The hierarchy separates semantic structure from temporal localization.
5 Tactile Representation
This section introduces a force-based tactile representation designed to unify heterogeneous sensors and preserve contact structure across space and time. NeoForce learns temporally structured visuo-tactile features from force fields, improving reconstructed force accuracy while maintaining contact localization.
- Representation design: The representation preserves spatial contact structure, including patch location, shape, extent, and intensity variation, rather than reducing touch to a single summary.These properties distinguish behaviors such as full grasps, edge contacts, stable holds, and incipient slip.
- Representation design: NeoForce represents tactile observations as dense three-axis force fields, encoding shear and pressure as a hardware-independent physical contact signal.The representation uses two tangential components for shear and one normal component for pressure at each sensing-surface location.
- Implementation: Raw tactile images are converted into unified force fields through a learned dense mapping trained on paired visuo-tactile data.The resulting field is the tactile representation supplied to downstream models and the target on which NeoForce builds.
- NeoForce model: NeoForce patchifies synchronized RGB frames and tactile force fields, fuses their tokens with a shared transformer, and aggregates temporal context.A reconstruction head predicts force fields and contact masks, while a latent prediction head models temporally structured tactile features.
- Evaluation: Adding latent prediction lowers reconstruction MAE from 0.070 to 0.066 and RMSE from 0.095 to 0.089, while contact overlap remains nearly unchanged at 0.966 versus 0.968 mIoU.The reported gains are concentrated in force-magnitude reconstruction rather than contact-region localization.
- Evaluation: Across press, twist, grasp, lift, and hold behaviors, force grounding yields consistent representations focused on interaction regions and their force, geometry, and extent.The authors conclude that force-based encoding generalizes across manipulation behaviors as a stable carrier of tactile interaction.
6 Evaluations on Tactile-aware Manipulation Tasks
The NeoReal and NeoSim suites evaluate tactile-aware policies on contact-rich manipulation under standardized real and simulated protocols. Results show strong difficulty in physical contact and bimanual simulation, while contact representations improve NeoReal performance.
- NeoReal: NeoReal evaluates 10 tactile-relevant tasks spanning deformable shaping, force-guided mating, delicate grasping, insertion, surface contact, and bimanual folding.Each task uses standardized initial states, reset protocols, and binary success criteria over physical robots equipped with tactile fingers.
- NeoReal: 26.5% average success rate makes π0.5 the strongest NeoReal policy among π0.5, LingBot-VA, and Fast-WAM.The policies are evaluated over 20 randomized rollouts per task.
- NeoReal: 38.1% average progressive score for π0.5 exceeds its 26.5% success rate, reflecting partial milestone completion even when full tasks fail.On Socket Plugging, π0.5 rises from 60% success to 73.5% progressive score; on Bag Packing, it rises from 20% to 43.0%.
- Tactile integration: 47.5% progressive score with NeoForce conditioning exceeds 38.1% for the vision-only baseline on NeoReal.Every tactile variant outperforms the vision-only policy, although individual tasks fluctuate across raw tactile conditions.
- NeoSim: 45.8% mean success rate makes π0.5 the NeoSim leader, followed by LingBot-VA at 32.1%, while the remaining baselines stay below 11%.NeoSim evaluates 12 tasks over 100 randomized rollouts per task, including four single-arm and eight dual-arm tasks.
- NeoSim: Dual-arm NeoSim tasks are consistently harder than single-arm tasks, and sustained mutual-contact tasks remain nearly unsolved.The best scores reach only 18% on Place Gears and 25% on Cup Handover; Fast-WAM fails all twelve tasks despite ranking among NeoReal’s strongest policies.
7 Conclusion
The paper concludes with an integrated tactile-manipulation resource spanning hardware, data, representation learning, and evaluation. It also identifies scope boundaries for the current representation and directions for broader deployment.
- Contributions: N0-Foundation unifies scalable tactile infrastructure, multimodal data acquisition, hardware-agnostic representation learning, and standardized policy evaluation.Its resources span hardware, data, representation, and evaluation for tactile-enabled embodied manipulation.
- Contributions: NeoData contains more than 30,000 hours of synchronized visual-tactile demonstrations across six embodiments and 450+ tasks.The corpus combines real-robot teleoperation with UMI demonstrations.
- Conclusions: NeoForce provides a compact physical alternative to raw tactile images, while experiments indicate benefits from tactile feedback and latent supervision.The conclusion presents these as supported experimental findings rather than universal guarantees.
- Open directions: The unified representation currently targets parallel-jaw visuo-tactile fingers, limiting direct applicability to dexterous hands and non-camera-based transducers.The paper identifies extension to those platforms, full-corpus policy learning, and tighter NeoSim-to-real connections as open directions.
- NeoData: NeoData stores synchronized RGB, tactile, action, and derived NeoForce streams, with UMI data using relative end-effector commands and gripper width.The synchronized streams are recorded at 640 × 360 for UMI observations, while NeoForce force fields are stored alongside raw tactile images.
- Annotation pipeline: The annotation pipeline combines VLM-generated task templates, signal-based interval boundaries, VLM labeling of pre-segmented intervals, and deterministic bottom-up temporal assembly.A human expert reviews each task template, while temporal endpoints come from signal computation rather than VLM inference.
D.1 Real Force Calibration Data
The force-calibration pipeline combines controlled real indentation with simulated contact geometries to train a tactile-image-to-force-field converter. The resulting representation is pixel-aligned and predicts three-axis contact forces.
- Real calibration: Controlled six-degree-of-freedom indentation pairs tactile images with reference loads measured by a six-axis force-torque sensor.The indenter is held stationary at target depth under a quasistatic protocol to capture stable tactile-force pairs.
- Real calibration: Approximately 4,300 optical-flow and multidimensional-force pairs are split into 80% training, 10% validation, and 10% test partitions.The test partition is held out exclusively for final evaluation.
- Simulation augmentation: Simulation augments physical calibration by generating supervised deformation-to-force pairs across broader contact geometries.Abaqus finite-element solutions provide displacement and distributed force fields used to render loaded tactile images.
- Simulation augmentation: Each simulated sample contains an undeformed image, a deformed image, and a pixel-aligned three-axis force field.The force-field channels encode x shear, y shear, and normal pressure.
- Force conversion model: The converter uses GoogLeNet with an output layer modified to regress per-pixel shear and pressure forces.Training mixes physically measured calibration data with simulation-augmented data.
D.4 NeoForce Experiment Details
NeoForce is trained on synchronized RGB and tactile-force chunks using reconstruction and teacher-supervised latent prediction. The experiment uses broad real-robot and handheld demonstrations with held-out episode evaluation.
- Data and evaluation: NeoForce trains on 20,000 demonstrations spanning 30 tasks and evaluates on 2,500 held-out episodes.The demonstrations come from real robots and N0-TacUMI, providing contact-rich manipulation priors.
- Architecture: Synchronized RGB and tactile-force chunks of length 4 are patchified independently and jointly processed by a shared ViT-B transformer initialized from DINOv2.The model has reconstruction and latent-prediction heads.
- Training objective: The reconstruction head predicts force fields and contact masks, while the latent head matches exponential-moving-average teacher outputs across visual, tactile, cross-modal, and masked-patch tokens.Training uses masked inputs and a learning rate of 2 × 10^-5 for 100,000 steps on 8× NVIDIA A100 GPUs.
E NeoReal Task Suite
NeoReal is a standardized real-world benchmark for contact-rich manipulation on physical robots with tactile fingers. It evaluates ten tasks using controlled initial states and binary, task-specific success criteria.
- E NeoReal Task Suite: NeoReal contains 10 real-world contact-rich manipulation tasks executed on physical robot setups equipped with tactile fingers.Each task is introduced with a representative benchmark image and its evaluated manipulation skills and tactile capabilities.
- E NeoReal Task Suite: Each task uses a standardized initial-state distribution and reset protocol for repeatable evaluation.
- E NeoReal Task Suite: Success is judged by a binary, task-specific criterion such as complete insertion, stable stacking, correct routing, object placement, or final shape quality.
E.1 Data Demonstrations
The NeoReal suite covers diverse contact-rich manipulation skills, from deformable-object handling and surface interaction to force-limited insertion and stacking. Its task-progress table decomposes each task into ordered milestones with cumulative points reaching 100 at completion.
- E.1 Data Demonstrations: Cardboard Box Folding tests contact-rich bending, crease following, and force-controlled pressing without tearing or collapsing the cardboard.
- E.1 Data Demonstrations: Bag Packing combines cluttered grasping, compliant-container placement, and zipper closing that requires sensing pulling resistance.
- E.1 Data Demonstrations: Cable Winding requires maintaining cable tension, following a target path, and avoiding missed posts, tangles, or crossing errors.
- E.1 Data Demonstrations: Board Wiping evaluates sustained surface contact, stable normal force, and target-area coverage without losing contact or displacing the board.
- E.1 Data Demonstrations: Cup Stacking, Socket Plugging, and Board Insertion respectively test low-force alignment of deformable cups, tactile edge-contact detection, and small-clearance assembly.
- E.1 Data Demonstrations: Bottle Standing, Fruit Collection, and Towel Folding assess controlled release, narrow-window gentle grasping, and coordinated bimanual manipulation of deformable objects.
- E.1 Data Demonstrations: Each task’s ordered milestones award cumulative points, with the row summing to 100 at full task completion.
E.2 NeoReal Task Progressive Score
NeoReal supplements binary success with a progressive score that measures how far an execution advances through ordered task milestones. Policies are pretrained on broad NeoData demonstrations and then specialized with task-specific demonstrations before controlled evaluation.
- E.2 NeoReal Task Progressive Score: The progressive score addresses binary success’s loss of information about partial progress before failure.
- E.2 NeoReal Task Progressive Score: Each task is decomposed into ordered sub-goals, and fixed points accumulate at milestones reached before progress stops or a failure condition occurs.
- E.2 NeoReal Task Progressive Score: A completed rollout passes all milestones and scores 100 points, while a partial failure receives the points accumulated through its furthest milestone.
- E.2 NeoReal Task Progressive Score: Policies are pretrained on over 400,000 NeoData episodes and specialized to each NeoReal task using 300 task-specific demonstrations.
- E.2 NeoReal Task Progressive Score: Evaluation provides third-person and wrist visual inputs, parameterizes actions as end-effector poses, and uses different robot arms for single- and dual-arm tasks.
- E.2 NeoReal Task Progressive Score: Each task is evaluated over 20 trials with randomized initial arm poses and object positions to test generalization across starting configurations and layouts.
- E.2 NeoReal Task Progressive Score: Failures include joint limits, workspace violations, self-locking, collisions, repeated actions without progress, and five-minute timeouts.
F NeoSim Task Suite
NeoSim is a standardized simulated benchmark for contact-rich manipulation, spanning single-arm and dual-arm tasks with tactile observations generated by a visuo-tactile simulation platform. It uses randomized configurations, geometric success criteria, scripted demonstrations, and fixed-seed evaluation.
- F NeoSim Task Suite: NeoSim covers 12 contact-rich tasks, including 4 single-arm and 8 dual-arm tasks.The single-arm tasks stress force regulation, insertion precision, and damage-free grasping; dual-arm tasks additionally test coordination and stable contact.
- F NeoSim Task Suite: The benchmark presents simulated RGB scenes alongside particle-gel tactile images and describes each task’s manipulation skill and tactile-feedback role.
- F NeoSim Task Suite: NeoSim uses UniVTAC with Isaac Sim, Isaac Lab, and TacEx, simulating finite-element gel outputs including tactile RGB images, marker motion, and gel depth maps.
- F NeoSim Task Suite: Task configuration files specify assets, randomized initial layouts, randomization ranges, and observation streams, while success uses task-specific geometric thresholds.
- F NeoSim Task Suite: The suite tests force-sensitive pouring, precision plugging, narrow-force-window grasping, force-guided mating, and tactile alignment during stacking, threading, handover, and unstacking.
- F NeoSim Task Suite: Dual-arm tasks include screw insertion, gear placement, bowl and cup unstacking, cup handover, and centered bowl stacking.
- F NeoSim Task Suite: Demonstrations retain only successful scripted-planner episodes, with 100 successful episodes collected per task and synchronized visual, tactile, and joint data stored.
- F NeoSim Task Suite: Each task is evaluated on 100 rollouts from fixed seeds disjoint from demonstration seeds, and performance is reported as mean success rate.