VisTacAlign: Co-Training Dexterous Policies on Tactile Human and Robot Demonstrations

Published in arXiv, 2026

Check out our website: VisTacAlign! Arxiv here

Abstract:

Human demonstrations are a cheap source of data for dexterous manipulation, but co-training a robot policy on them requires closing the human-robot gap in every modality the policy consumes. We present VisTacAlign, a framework for co-training 3D-visual-tactile dexterous policies on human and robot demonstrations. Glove-tracked human hand motion is retargeted to a 17-DoF tactile robot hand with a one-time fingertip correction. The human hand is then erased from both stereo views and replaced by a posed robot-hand mesh painted with pixels from robot recordings, and a real-time stereo foundation model is re-run on the composite, so the human point clouds carry the same stereo errors and visibility as the robot ones. Finally, a capacitive tactile glove is aligned to the robot’s fingertip sensors in its signal space, giving one interpretable per-finger force representation. A diffusion transformer consumes point-cloud, proprioceptive, and per-finger tactile tokens. On three real-world tasks requiring precise force (Lego assembly, plucking strawberries of varying size, and activating and lifting a power drill), adding aligned human demonstrations to existing robot data improves over robot-only policies, and ablations show that both tactile input and visual alignment are necessary.