vb000/LookOnceToHear
A novel human-interaction method for real-time speech extraction on headphones.
What it solves
It addresses the challenge of target speech hearing in noisy environments, allowing a user to select a specific speaker they want to hear by simply looking at them for a few seconds.
How it works
The system uses a combination of audio and visual cues. It is trained on synthetic audio mixtures generated on-the-fly using the Scaper toolkit, incorporating clean speech, background noise, and spatial audio properties like head-related transfer functions (HRTFs) and binaural room impulse responses (BRIRs).
Who it’s for
Researchers and developers working on intelligent hearables, audio-spatial perception, and multimodal AI systems for hearing assistance.
Highlights
- Best paper honorable mention at CHI 2024.
- Uses synthetic audio generation for training data.
- Supports binaural audio processing to simulate real-world acoustic environments.
Related
- Project
- Project
- Project
- Project
- Project