vb000/LookOnceToHear

A novel human-interaction method for real-time speech extraction on headphones.

What it solves

It addresses the challenge of target speech hearing in noisy environments, allowing a user to select a specific speaker they want to hear by simply looking at them for a few seconds.

How it works

The system uses a combination of audio and visual cues. It is trained on synthetic audio mixtures generated on-the-fly using the Scaper toolkit, incorporating clean speech, background noise, and spatial audio properties like head-related transfer functions (HRTFs) and binaural room impulse responses (BRIRs).

Who it’s for

Researchers and developers working on intelligent hearables, audio-spatial perception, and multimodal AI systems for hearing assistance.

Highlights

  • Best paper honorable mention at CHI 2024.
  • Uses synthetic audio generation for training data.
  • Supports binaural audio processing to simulate real-world acoustic environments.

Related

  • Project
  • Project
  • Project
  • Project
  • Project