liangnjupt/VisTouch

A large-scale synchronized vision–touch–audio dataset of robotic sliding contact

What it solves

VisTouch provides a large-scale, synchronized dataset of robotic sliding contact to help AI models learn to recognize materials and understand the relationship between vision, sound, and touch. It addresses the lack of public datasets that strictly align these three modalities for robotic interaction.

How it works

Using a robot arm and a dexterous hand, the system records simultaneous video, audio, and force-tactile trajectories as the hand slides across eight different materials (such as brass, silk, and wood). The data is strictly synchronized using a shared wall clock, ensuring that every frame of video, sound sample, and tactile reading is aligned to the same event.

Who it’s for

Researchers and developers working on multimodal learning, robotic perception, and haptic feedback systems who need high-quality, paired data for training models.

Highlights

  • Strict Synchronization: The first public dataset to provide event-level synchronization across vision, sound, and force-tactile signals.
  • Diverse Materials: Includes 10,498 synchronized clips covering eight distinct materials.
  • Multimodal Benchmarks: Includes baseline models for material recognition, tactile super-resolution, and cross-modal generation (e.g., generating tactile force from sound).
  • Comprehensive Hardware: Data captured using a 2K camera, professional microphone, and a dexterous hand's integrated pressure/force sensors.

Related

  • Project
  • Project
  • Project
  • Project
  • Project