xzf-thu/Mega-ASR
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
What it solves
Mega-ASR addresses the failure of standard Automatic Speech Recognition (ASR) models in "in-the-wild" environments. Most models struggle when faced with real-world acoustic challenges like background noise, echoes, or electronic distortion, which often lead to high Word Error Rates (WER) or completely empty outputs.
How it works
The project scales up real-world acoustic simulation to create a robust foundation model. It was trained on 2.4 million samples covering 7 atomic acoustic conditions and 54 compound scenarios, including:
- Noise and far-field speech
- Obstructions and reverberation
- Recording artifacts and electronic distortion
- Transmission dropouts
To optimize performance, it employs A2S-SFT (Supervised Fine-Tuning) and DG-WGPO based RL (Reinforcement Learning).
Who it’s for
It is designed for developers and researchers needing high-accuracy speech-to-text transcription in challenging, non-studio acoustic environments where traditional SOTA models typically fail.
Highlights
- Robustness: Specifically targets full-scenario robust speech recognition.
- Scale: Trained on a massive dataset of 2.4M samples simulating diverse real-world conditions.
- Performance: Achieves up to nearly 30% gains over leading open and closed source SOTA models in difficult environments.
- Comprehensive Simulation: Covers a wide array of acoustic distortions from transmission dropouts to far-field speech.
Related
- Project
- Project
- Project
- Project