Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist
RA-L 2026
TL;DR A soft wrist collects safe tactile demos, and masked trajectory encoding retrieves them to insert unseen pegs at 85% success.
Tactile memory, the ability to store and retrieve touch-based experience, is critical for contact-rich tasks such as peg-in-hole insertion under uncertainty. To replicate this capability, we introduce Tactile Memory with Soft Robot (TaMeSo-bot), a system that integrates a soft wrist with tactile retrieval-based control to enable safe and robust manipulation. The soft wrist allows safe contact exploration during data collection, while tactile memory reuses past demonstrations via retrieval for flexible adaptation to unseen scenarios. The core of this system is the Masked Tactile Trajectory Transformer (MAT3), which jointly models spatiotemporal interactions between robot actions, distributed tactile feedback, force-torque measurements, and proprioceptive signals. Through masked-token prediction, MAT3 learns rich spatiotemporal representations by inferring missing sensory information from context, autonomously extracting task-relevant features without explicit subtask segmentation. We validate our approach on peg-in-hole tasks with diverse pegs and conditions in real-robot experiments. Our extensive evaluation demonstrates that MAT3 achieves higher success rates than the baselines over all conditions and shows remarkable capability to adapt to unseen pegs and conditions.
You drop a peg into a hole you cannot see, guided only by what the fingers feel. Small exploratory motions tell you where the rim catches, and you recall how a peg of this shape slid, jammed, and finally seated the last time. Cognitive science calls this tactile memory: the encoding, storage, and retrieval of touch-based experience. It keeps the temporal structure of contact rather than isolated events, and it works together with proprioception rather than in isolation.
TaMeSo-bot mirrors the three stages. A soft wrist lets the robot press into contact without triggering a protective stop, so teleoperated demonstrations capture rich signals from the distributed taxels (encoding). MAT3 compresses each sub-trajectory of tactile, force-torque, pose, and action signals into one vector, and each vector goes into a database with the action that accompanied it (storage). During execution the robot encodes its current contact history as a query and replays the action attached to the nearest stored experience (retrieval).
Why this framing matters. Prior insertion systems name their contact phases in advance and hand the robot a state machine to follow. Tactile memory needs no such labels. Similarity in the learned representation decides which past experience applies right now, so fit, align, and insert emerge from the data instead of being written down. The soft wrist absorbs whatever mismatch retrieval leaves behind, which is what lets a memory built from two peg shapes carry over to five it has never touched.
Softness for safe data collection. TaMeSo-bot pairs a rigid arm with a highly deformable soft wrist and a two-finger gripper carrying a 3 x 3 array of taxels, each measuring a 3D force. The compliance pays off twice: it absorbs contact forces so teleoperated demonstrations can be collected without triggering protective stops, and it keeps the same contact-rich behavior safe at execution time. Those offline demonstrations become the tactile memory, a database the robot queries instead of rolling out a learned policy.
MAT3 encodes what the fingers felt. The keys of that database come from a bidirectional Transformer over sub-trajectories of taxel readings and actions. Taxels and the action are the base tokens, while force-torque, arm pose, and gripper pose are fused in as auxiliary global context; sinusoidal encodings place each taxel on the sensor grid and the action token at its centre. Training is masked-token reconstruction: a random fraction of the input is hidden and the encoder infers it from the surrounding spatial and temporal context, so task-relevant structure emerges without any subtask labels.
Retrieval is the policy. At execution time the current action token is masked, the encoder and average pooling produce a query embedding, and the nearest stored embeddings are found by L2 distance over an HNSW index, fast enough for real-time control. One neighbour is sampled and its action is executed, so the database itself acts as a non-parametric policy whose behavior stays bounded by the demonstration data.
Demonstrations were collected on two peg shapes only, and the policy is then asked to insert five shapes it has never touched. MAT3 matches the masking-free variant on the seen pegs and pulls clearly ahead on the unseen ones: 85% versus 71% over 200 trials. We also compare against Action Chunking Transformer (ACT), a parametric policy trained on the same 64 demonstrations with tactile-proprioceptive inputs: it reaches 63.8% on seen pegs but falls to 53% on unseen shapes, its dominant failure being downward force applied before the peg is aligned, which triggers a protective stop. The tactile Transformer baseline stays below 20% throughout.
| Shape | Tactile Transformer | MAT3 w/o Mask | MAT3 (Ours) | ACT |
|---|---|---|---|---|
| Seen | ||||
| Square | 20% 8/40 |
90% 36/40 |
87.5% 35/40 |
32.5% 13/40 |
| Cyl. (⌀40) | 25% 10/40 |
90% 36/40 |
90% 36/40 |
95% 38/40 |
| Total (Seen) | 22.5% 18/80 |
90% 72/80 |
88.8% 71/80 |
63.8% 51/80 |
| Unseen | ||||
| Cyl. (⌀30) | 30% 12/40 |
82.5% 33/40 |
87.5% 35/40 |
92.5% 37/40 |
| Rectangle | 15% 6/40 |
65% 26/40 |
80% 32/40 |
25% 10/40 |
| Oval | 17.5% 7/40 |
70% 28/40 |
85% 34/40 |
65% 26/40 |
| Hexagon | 12.5% 5/40 |
75% 30/40 |
90% 36/40 |
50% 20/40 |
| Pentagon | 12.5% 5/40 |
62.5% 25/40 |
82.5% 33/40 |
32.5% 13/40 |
| Total (Unseen) | 17.5% 35/200 |
71% 142/200 |
85% 170/200 |
53% 106/200 |
Retrieval is what keeps the policy inside the demonstration distribution. A parametric policy has no such guarantee, and ACT extrapolates beyond it, executing insertion-phase behavior before the alignment is done.
The gap widens once the setup itself changes. Under starting positions outside the demonstration range, added friction between peg and fixture, and a tilted grasp, MAT3 succeeds on 57.5% of trials, while MAT3 w/o Mask drops to 29.2%, ACT to 14.2%, and the tactile Transformer to 7.5%, evidence that masked training buys robustness rather than just accuracy.
| Condition | Tactile Transformer | MAT3 w/o Mask | MAT3 (Ours) | ACT |
|---|---|---|---|---|
| Unseen Starting Positions | 5% 2/40 |
22.5% 9/40 |
50% 20/40 |
10% 4/40 |
| Increased Friction | 5% 2/40 |
17.5% 7/40 |
50% 20/40 |
7.5% 3/40 |
| 5° Tilted Grasp | 12.5% 5/40 |
47.5% 19/40 |
72.5% 29/40 |
25% 10/40 |
| Total | 7.5% 9/120 |
29.2% 35/120 |
57.5% 69/120 |
14.2% 17/120 |
@article{kamijo2026tactile,
title={Tactile Memory with Soft Robot: Robust Object Insertion via Masked Encoding and Soft Wrist},
author={Kamijo, Tatsuya and Nishimura, Mai and Shibasaki, Nodoka and Siburian, Jeremy and Beltran-Hernandez, Cristian C and Hamaya, Masashi},
journal={IEEE Robotics and Automation Letters},
year={2026},
publisher={IEEE}
}