SoTa

SoTa: Soft Tactile Skins for
Dexterous Manipulation

Jingyun Yang1*Baiyu Shi1*Timothy Yu1*Haitian Liu2Alberta Longhini1Weichen Wang1Rika Antonova3Zhenan Bao1Jeannette Bohg1
1 Stanford University2 Tsinghua University3 University of Cambridge

* Equal contribution. Correspondence to: jingyuny@stanford.edu

Paper teaser: human and robot tactile skins, paired visual and tactile demonstrations, and cup-picking examples with robot-only, tactile, and human co-training policies
Figure 1. Full-hand tactile skins with matched human–robot layouts, recorded visual–tactile observations, and representative manipulation outcomes.

Abstract

A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scaling this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 202 taxels across corresponding finger and palm regions. Our multilayer design with fabric electrodes enables in-house fabrication of thin, soft skins with customizable geometry for under $10 in materials per skin. The sensor retains over 97% of its initial response span after 10,000 loading-unloading cycles with traces retaining continuity through 1,280 tight-fist folding cycles. The shared taxel layout supports human-robot co-training with a common tactile encoder and no learned cross-sensor mapping. Across three contact-rich manipulation tasks, tactile observations improve in-distribution success over vision-only policies. With a fixed robot-demonstration budget, adding human demonstrations more than doubles mean success across eight evaluation conditions, from 22.8% to 45.9%, improving success in all five out-of-distribution conditions. We plan to open-source the fabrication process, hardware, and firmware.

01 / SENSOR DESIGN

Tactile Sensor Design

ROBOTHUMAN
Robot and human tactile skins with corresponding finger and palm regions
Custom geometry preserves the shared region-level layout.

Full-Hand Tactile Sensing

SoTa combines thin, soft construction with 202 taxels spanning the fingers and palm. Its geometry adapts to human and robot hands while preserving corresponding sensing regions, so demonstrations from both can train a shared tactile encoder without a learned cross-sensor mapping. Each skin costs under $10 in materials and can be fabricated in-house.

Taxels per hand
202
Demonstration streams
30 Hz
Wireless operation
~7 hours
Tested load range
20 mN – 10 N
Exploded view of the capacitive skin stack: electrodes, SEBS dielectric, TPU adhesive, and polyimide substrate
Capacitive Sensor Stack

Sensor Construction

Conductive fabric is heat-pressed onto a thin polyimide support using TPU adhesive, then patterned with a UV laser to form electrically isolated electrodes and traces. A sandpaper-molded SEBS dielectric is sandwiched between the top and bottom electrode layers to form the capacitive sensing stack. The layers are aligned, sealed at the edges with Kapton tape, and connected to the readout board through flexible flat cables.

Fabrication takes approximately six hours of laser machine time and two hours of manual assembly.

02 / DEMONSTRATIONS

SoTa in Action

Training Demonstrations

Box reorientation

Human demonstration
Robot demonstration

Cup picking

Human demonstration
Robot demonstration

Plug insertion

Human demonstration
Robot demonstration
How to read the synchronized tactile display

The hand map shows all 202 recorded taxels in their canonical positions. Left-hand geometry is mirrored without changing channel order. Colors show post-processed tactile activation, not force in newtons. Tactile values undergo exponential normalization and exactly one clipped linear rescaling from [0.1, 2.0] to [0, 1]. Every clip uses the same color scale from 0 to 0.50 (values above 0.50 saturate); the line shows the maximum activation across taxels on the same clipped scale. Taxel colors and activation-dependent sizes match the paper figure renderer.

All videos play at 2x.

03 / TRAINING

Training Pipeline

Training architecture: shared observation paths, an auxiliary language objective, and robot-only action supervision
Pre-training: human and robot observations supervise task-phase prediction. The action branch is frozen and receives no action supervision.

Our policy architecture is built on top of FTP-1. To incorporate human demonstrations, we use a shared tactile encoder and a training-only task-phase prediction auxiliary objective.

Representation pretraining. Human and robot demonstrations supervise task-phase prediction from visual and tactile observations. The action branch remains frozen, and no action loss is used.

Policy co-training. Robot demonstrations supervise action prediction through flow matching, while both human and robot demonstrations supervise the auxiliary task-phase objective. Human demonstrations never contribute to the action loss.

We use 75 robot demonstrations for box reorientation, 50 for cup picking, and 75 for plug insertion. Human–robot co-training adds 150 human demonstrations per task while keeping the robot demonstration budget fixed.

04 / RESULTS

Quantitative Results

With the same robot demonstration budget, full human–robot co-training increases mean success from 22.8% to 45.9% across eight evaluation conditions. Success improves in every reported condition, including all five out-of-distribution conditions.

40 trials per method, task, and condition. Partial completions count as failures. One training seed per method and task. ID: seen in robot training. OOD: absent from both robot and human training.

Qualitative Comparisons

Robot-only with TactileFailure

The robot lifts nested cups together.

Human-robot Co-trainingSuccess

The robot separates and retrieves one cup.

The evaluation stack is 21 mm higher than the training stack.

05 / CHARACTERIZATION

Sensor Characterization

Sensor response over three cycles each at loads from 20 millinewtons to 10 newtons

Load Response

Capacitive response over applied loads from 20 mN to 10 N.
Capacitance over 10,000 loading and unloading cycles at 8 newtons

Cyclic Stability

97.31% of the initial response span retained after 10,000 loading cycles at 8 N over 26.5 hours.
Conductor resistance versus folding cycles comparing four conductor materials

Conductor Flexural Durability

All four fabric-electrode traces retained continuity through 1,280 tight-fist folding cycles.

Motion artifacts are not zero: the largest mean regional RMS during no-contact finger flexion was 6.18%. The paper reports the full protocol and limitations.

06 / DISCUSSION

Discussion

Human demonstrations improve generalization to unseen objects.

Human demonstrations cover three objects per task, compared with one in the robot dataset. The co-training gains are consistent with this broader contact experience supporting generalization; the evaluated OOD objects remain absent from both datasets.

Touch improves trained-task performance, but does not ensure generalization.

Tactile input helps all three in-distribution tasks, but alone is not enough for reliable transfer to unseen cups and plugs. Shared geometry creates a useful interface, while the training distribution still matters.

Limitations

The study uses one training seed per method and task, a fixed set of 150 human demonstrations per task, and short-term hardware tests. Larger and more diverse human datasets, alternative representations, long-term wear, and sustained use remain open questions.

Scope of the Ablation Study

Vision-only auxiliary learning removes touch from the auxiliary objective for both human and robot observations. It tests tactile auxiliary supervision as a whole, rather than isolating human touch. The study also does not compare against every tactile layout or processing method.