SoTa: Soft Tactile Skins for
Dexterous Manipulation

Abstract
A growing body of work suggests that tactile sensing gives robot policies contact information that complements vision in dexterous manipulation. However, visuo-tactile robot data remains scarce: dexterous demonstrations require teleoperating robots, which limits dataset scale. Human demonstrations are far cheaper to collect and offer a path to scaling this data, but only if human and robot hands carry tactile sensors with corresponding signals. This requires sensors that conform to different hand geometries, cover the full hand, and share a common layout across embodiments. We present SoTa, a low-cost capacitive tactile skin that provides full-hand coverage on humans and robots while preserving a shared layout of 202 taxels across corresponding finger and palm regions. Our multilayer design with fabric electrodes enables in-house fabrication of thin, soft skins with customizable geometry for under $10 in materials per skin. The sensor retains over 97% of its initial response span after 10,000 loading-unloading cycles with traces retaining continuity through 1,280 tight-fist folding cycles. The shared taxel layout supports human-robot co-training with a common tactile encoder and no learned cross-sensor mapping. Across three contact-rich manipulation tasks, tactile observations improve in-distribution success over vision-only policies. With a fixed robot-demonstration budget, adding human demonstrations more than doubles mean success across eight evaluation conditions, from 22.8% to 45.9%, improving success in all five out-of-distribution conditions. We plan to open-source the fabrication process, hardware, and firmware.
Tactile Sensor Design

Full-Hand Tactile Sensing
SoTa combines thin, soft construction with 202 taxels spanning the fingers and palm. Its geometry adapts to human and robot hands while preserving corresponding sensing regions, so demonstrations from both can train a shared tactile encoder without a learned cross-sensor mapping. Each skin costs under $10 in materials and can be fabricated in-house.
- Taxels per hand
- 202
- Demonstration streams
- 30 Hz
- Wireless operation
- ~7 hours
- Tested load range
- 20 mN – 10 N
Sensor Construction
Conductive fabric is heat-pressed onto a thin polyimide support using TPU adhesive, then patterned with a UV laser to form electrically isolated electrodes and traces. A sandpaper-molded SEBS dielectric is sandwiched between the top and bottom electrode layers to form the capacitive sensing stack. The layers are aligned, sealed at the edges with Kapton tape, and connected to the readout board through flexible flat cables.
Fabrication takes approximately six hours of laser machine time and two hours of manual assembly.
SoTa in Action
Training Demonstrations
Cup picking
Plug insertion
How to read the synchronized tactile display
The hand map shows all 202 recorded taxels in their canonical positions. Left-hand geometry is mirrored without changing channel order. Colors show post-processed tactile activation, not force in newtons. Tactile values undergo exponential normalization and exactly one clipped linear rescaling from [0.1, 2.0] to [0, 1]. Every clip uses the same color scale from 0 to 0.50 (values above 0.50 saturate); the line shows the maximum activation across taxels on the same clipped scale. Taxel colors and activation-dependent sizes match the paper figure renderer.
All videos play at 2x.
Training Pipeline
Our policy architecture is built on top of FTP-1. To incorporate human demonstrations, we use a shared tactile encoder and a training-only task-phase prediction auxiliary objective.
Representation pretraining. Human and robot demonstrations supervise task-phase prediction from visual and tactile observations. The action branch remains frozen, and no action loss is used.
Policy co-training. Robot demonstrations supervise action prediction through flow matching, while both human and robot demonstrations supervise the auxiliary task-phase objective. Human demonstrations never contribute to the action loss.
We use 75 robot demonstrations for box reorientation, 50 for cup picking, and 75 for plug insertion. Human–robot co-training adds 150 human demonstrations per task while keeping the robot demonstration budget fixed.
Quantitative Results
With the same robot demonstration budget, full human–robot co-training increases mean success from 22.8% to 45.9% across eight evaluation conditions. Success improves in every reported condition, including all five out-of-distribution conditions.
40 trials per method, task, and condition. Partial completions count as failures. One training seed per method and task. ID: seen in robot training. OOD: absent from both robot and human training.
Qualitative Comparisons
The robot lifts nested cups together.
The robot separates and retrieves one cup.
The evaluation stack is 21 mm higher than the training stack.
Sensor Characterization

Load Response

Cyclic Stability

Conductor Flexural Durability
Motion artifacts are not zero: the largest mean regional RMS during no-contact finger flexion was 6.18%. The paper reports the full protocol and limitations.
Discussion
Human demonstrations improve generalization to unseen objects.
Human demonstrations cover three objects per task, compared with one in the robot dataset. The co-training gains are consistent with this broader contact experience supporting generalization; the evaluated OOD objects remain absent from both datasets.
Touch improves trained-task performance, but does not ensure generalization.
Tactile input helps all three in-distribution tasks, but alone is not enough for reliable transfer to unseen cups and plugs. Shared geometry creates a useful interface, while the training distribution still matters.
Limitations
The study uses one training seed per method and task, a fixed set of 150 human demonstrations per task, and short-term hardware tests. Larger and more diverse human datasets, alternative representations, long-term wear, and sustained use remain open questions.
Scope of the Ablation Study
Vision-only auxiliary learning removes touch from the auxiliary objective for both human and robot observations. It tests tactile auxiliary supervision as a whole, rather than isolating human touch. The study also does not compare against every tactile layout or processing method.