HomeArticle

Transplant the "sense of touch" of human hands to robots: 500 hours of real operation data, the task success rate has doubled

机器之心2026-10-09 07:32
Haptics is now starting to follow the scaling path that vision has already gone through.

Pick up a raw egg, and your fingers need to apply just the right amount of force: not too loose to let it slip, not too tight to crush it. When wringing a wet towel, the force from your fingertips also needs to be constantly adjusted as the water is squeezed out. For humans, these operations are almost instinctive. However, if relying solely on vision, the camera can capture the trajectory of the hand, but cannot detect the contact position of the fingers and the pressing force.

In the past few years, in the academic community, first-person human videos have rapidly expanded to tens of thousands of hours, and more and more work has begun to see the benefits brought by the growth of data scale. Visual-tactile data is far from reaching such a scale: most of the previously published human visual-tactile datasets only range from a few hours to dozens of hours. If multiple small datasets are simply merged, the tactile sensors, collection processes and data formats of each party are different. Although the amount of data has increased, the hardware and processes have also changed accordingly, making it difficult to independently judge how much contribution "more data" has made.

Recently, researchers from 15 institutions including Texas A&M, Google DeepMind, Carnegie Mellon University (CMU), Stanford University, University of California, Berkeley, Yale University, Microsoft, NVIDIA, University of Liverpool, Meta, University of Washington, Northwestern University, Sony, Georgia Institute of Technology (Georgia Tech), and Overfit Lab (data provider) released TouchScale: a 500-hour human visual-tactile dataset constructed under a unified sensor, unified collection and synchronization process.

The team hopes to answer a key question: instead of splicing data from different sources, but under the same set of sensing and collection systems, when high-quality first-person visual-tactile data is truly expanded to hundreds of hours, will it also show the Scaling benefits that come with the growth of data scale like large-scale videos? How much of these Scaling benefits can be further transformed into the operational capabilities of robots?

What does TouchScale bring?

1.500 hours of data collected by a unified wearable device. TouchScale contains 500 hours of synchronized visual-tactile data of human hands, all collected by the same set of wearable systems, which simultaneously record head RGB-D, dual-wrist RGB and full-hand tactile signals of both hands.

2.Human "hand feel" can be transferred to robots. Without human action labels, nor mapping human hand actions to robots, only using human visual-tactile data for mid-training, the average success rate of the robot in 4 real-machine contact tasks increased from 22.5% to 57.5%.

3.Hundreds of hours of human visual-tactile data have begun to show obvious Scaling benefits. When TouchScale is used for mid-training, the overall performance of real-machine operations becomes stronger as the data scale expands; the same upward trend is also observed in independent zero-shot tactile prediction experiments. The scale benefit extends all the way from perception to robot control.

  • Title: TouchScale: 500 Hours of Human Vision and Touch for Visual-Tactile Learning
  • Paper link: https://arxiv.org/abs/2610.10288
  • Project homepage: https://touch-scale.github.io/
  • Dataset address: https://huggingface.co/datasets/2077AIDataFoundation/TouchScale

For 500 hours of tactile data, the difficulty is not just "collecting more"

For TouchScale, the real difficulty is not simply stacking the number of hours to 500, but continuously expanding under the same set of sensors and the same set of synchronization and collection processes. Most of the previously published human visual-tactile datasets are less than 30 hours; if multiple small datasets are directly spliced together, different hardware and collection processes will change together with the data scale, making it difficult to separately study the impact of the scale itself.

Therefore, TouchScale fixes the hardware from the collection end. The head RGB-D camera records the overall operation process and scene depth, and the RGB cameras on both sides of the wrist record the interaction details between the hand and the object at close range; both hands are equipped with full-hand tactile gloves, each glove contains 880 tactile sensor units (taxel), covering five fingers and the palm, used to record the normal pressure during contact.

In order to ensure the quality of large-scale tactile data, TouchScale jointly controls data quality from both hardware and algorithm ends. The collection end uses high-precision, low-noise flexible skin tactile sensors to stably record the pressure changes of fingers and palms during the interaction process; after the collection is completed, through the VLM-assisted automatic quality inspection process, the grasping, releasing and fingers involved in contact are judged only based on the wrist video, and cross-compared with the actual tactile response. Obvious sensor failures will be directly screened out, and uncertain samples such as missing single finger, severe occlusion or insufficient visual evidence will automatically enter the manual review process.

What exactly was collected in the 500 hours?

TouchScale contains about 87,000 interaction recordings, about 2000 task descriptions and more than 1500 objects, covering 9 types of collection environments, more than 800 scene configurations, and completed by about 20 collectors. The tasks include not only daily operations, but also more fine-grained contact and tool use. For example, pouring water from a kettle into a wide-mouth bottle requires continuous adjustment of the container posture; rolling up the notebook charging cable and fixing it with Velcro involves two-hand cooperation and continuous contact; sealing the top seam of the carton with tape includes operations such as grasping tools, positioning and continuous pressing. Different tasks, objects and execution methods also bring different hand-object postures and tactile contact patterns.

Compared with the existing publicly available human visual-tactile data, TouchScale not only expands the scale to 500 hours, but also always maintains a unified sensor and collection process, and provides full-hand tactile collection for both hands, with 880 taxels per hand. This provides conditions for independently studying the impact brought by the data scale itself.

From perception to real machine: What do 500 hours of data bring?

With the same 16 hours, better cross-sensor generalization

The first set of experiments examines zero-shot cross-sensor tactile prediction. The model predicts hand tactile sensation based on first-person and wrist videos, and then directly tests on EgoTactile collected by another set of tactile gloves without fine-tuning using the target data. Due to the different sensing layouts of different gloves, the results are mapped to a unified 12 hand anatomical regions during evaluation.

First, the amount of training data is controlled at the same about 16 hours: the model trained with the TouchScale subset achieves a contact cIoU (contact intersection over union) of 0.181, which is higher than 0.134 of EgoTouch. This comparison excludes "more data", indicating that under the same training amount, TouchScale is also more conducive to cross-sensor generalization.

Tactile supervision can also help vision in turn

The second set of experiments uses tactile sensation as the supervision signal of the visual encoder, and then transfers the learned visual representation to the action recognition task. The researchers used TouchScale, OpenTouch, FEEL and EgoTouch for pre-training respectively, and kept the same initialization and training budget.

On the three benchmarks of MECCANO, Something-Something V2 and Ego-Exo4D, compared with pre-training using OpenTouch, FEEL and EgoTouch, the visual encoder pre-trained with TouchScale achieves the highest accuracy under both linear probing and end-to-end fine-tuning settings.

Finally look at the real machine: Can human "hand feel" be used by robots?

The real-machine experiment uses the xArm6 robotic arm and BrainCo Revo 2 dexterous hand, and sets up four contact-intensive tasks: sorting soft and hard objects, removing bottle caps, transferring test tubes and wiping whiteboards. Each task uses 50 robot demonstrations and is tested 20 times.

The control setup is very simple: both versions start from the same N0-VTLA checkpoint, both use visual and tactile inputs, and both use exactly the same robot demonstrations for post-training. The only difference is that one version goes through TouchScale mid-training before post-training.

TouchScale mid-training does not use human action labels, nor does it map human hand trajectories to robot actions. It follows the future tactile prediction goal of N0-VTLA, allowing the model to predict subsequent tactile changes based on current vision, language and tactile sensation. In other words, what the robot learns from human data is not "how the hand should move", but "how the contact will occur".

After adding TouchScale mid-training, the average success rate of the four tasks increased from 22.5% to 57.5%, and all four tasks improved:

  • Soft and hard object sorting: 10% → 60%
  • Bottle cap removal: 40% → 70%
  • Test tube transfer: 30% → 60%
  • Whiteboard wiping: 10% → 40%

Will the effect continue to rise as more data is added?

What is really exciting is that when the data scale of TouchScale continues to increase, both the model and the robot continue to get better.

In the tactile prediction task, the model and evaluation settings remain unchanged, only the training data is increased. After TouchScale expands from 10% to 100%, the cIoU increases from 0.311 to 0.383. That is to say, the more human visual-tactile interactions the model has seen, the more accurate contact predictions it can still make when facing a set of never-seen tactile sensors.

And this trend does not stop at the perception level. The real-machine demonstration data used by the robot remains unchanged, and only the human data of TouchScale mid-training is increased, the average success rate of the four tasks increases from 22.5% to 57.5%.

Final note: Tactile sensation has also reached the starting point of Scaling

What TouchScale wants to answer is what can be brought by simply expanding the scale of human visual-tactile data on the premise that the sensor and collection process remain consistent. Now it seems that the answer is quite clear: as the amount of data increases, the cross-sensor tactile prediction becomes more accurate overall, and the robot that has undergone TouchScale mid-training also obtains stronger real-machine operation capabilities.

This means that tactile sensation is no longer only an expensive, scattered additional modality in the robot system. It can also be collected and learned on a large scale like first-person videos, and further transformed into the physical interaction experience required by robots. More importantly, what TouchScale demonstrates is not just a larger dataset, but a new data expansion path: when the "hand feel" in human operations is systematically recorded, robots can obtain transferable physical interaction experience from these human experiences without action labels.

Of course, this road has just begun. The current real-machine experiments only cover one robot platform and four contact-intensive tasks, and there are still many problems worthy of further research on the mechanism through which human tactile experience is transformed into better robot control. But TouchScale has given a clear signal: the Scaling path that vision has gone through, tactile sensation has also started to take.

Author Introduction

The co-first authors of this paper are Dayou Li, PhD student at Texas A&M University, Hao Wang, Senior Research Engineer at Google DeepMind, Qian