jcvasquezc/DisVoice
feature extraction from speech signals
What it solves
DisVoice provides a standardized way to extract acoustic and linguistic features from speech files to help identify paralinguistic aspects of voice. This is primarily used to recognize emotions or detect communication impairments caused by speech disorders, such as Parkinson's, Huntington's, larynx cancer, or cleft-lip and palate, as well as mood disorders like depression.
How it works
The framework computes a wide variety of speech features from both sustained vowels and continuous speech. It utilizes several extraction strategies:
- Glottal, Phonation, Articulation, and Prosody: Traditional acoustic analysis.
- Phonological: Analysis of speech sounds.
- Representation Learning: Using autoencoders to learn feature representations from the audio data.
It integrates with external tools like Praat and Kaldi for the underlying signal processing and phonological analysis.
Who it’s for
Researchers and clinicians focusing on speech pathology, medical diagnostics via audio, and emotion recognition.
Highlights
- Supports extraction of glottal, phonation, articulation, prosody, and phonological features.
- Uses autoencoders for representation learning.
- Capable of analyzing both sustained vowels and continuous speech.
- Designed for the classification of pathological speech and mood-related patterns.
Related
- Project
- Project
- Project
- Project
- Project