Nari Labs Qwen3-ASR Fast and Qwen3-TTS Fast Lead Coval Voice AI Benchmarks (Sept 2026)
TL;DR – Nari Labs dominates Coval’s voice AI benchmark
Nari Labs’ Qwen3‑ASR Fast and Qwen3‑TTS Fast models occupy the quality‑latency Pareto frontier for Speech‑to‑Text (STT) and Text‑to‑Speech (TTS) respectively, and they also sit on the latency‑cost and quality‑cost frontiers among all publicly listed models.
Speech‑to‑Text performance
Conclusion: Qwen3‑ASR Fast delivers the fastest median time‑to‑final‑segment (TTFS) at 44 ms while keeping word error rate (WER) at 3.6%, ranking #1 for latency and #2 for accuracy on Coval’s benchmark.
- Latency: p50 TTFS = 44 ms, the lowest reported value across the Coval STT leaderboard.
- Accuracy: WER = 3.6%, only 0.1 % higher than AssemblyAI’s Universal 3.5 Pro (WER = 3.5 %).
- Cost: $0.12 / hour for the Fast endpoint, tying for the second‑lowest price among models with published rates. The Standard endpoint is even cheaper at $0.06 / hour.
- Competitive pricing: AssemblyAI’s Universal 3.5 Pro costs 3.75× more; Deepgram Nova 3 costs 2.4× more.
“Our Qwen3‑ASR Fast model is ranked #1 in Time‑to‑Final‑Segment (TTFS), at p50 of 44 ms and WER of 3.6%, placing #2 behind AssemblyAI’s Universal 3.5 Pro at 3.5%.” – Nari Labs blog
Text‑to‑Speech performance
Conclusion: Qwen3‑TTS Fast achieves the second‑fastest median time‑to‑first‑audio (TTFA) at 63 ms and the best WER at 3.8%, outperforming larger commercial services while costing less.
- Latency: p50 TTFA = 63 ms, second only to Fluxions’ vui (49 ms) which uses a 300 M‑parameter model.
- Accuracy: WER = 3.8%, the lowest among all models evaluated.
- Cost: $10 / 1M characters, tied for the cheapest price on Coval’s pricing directory. The Standard endpoint is $5 / 1M characters.
- Competitive pricing: ElevenLabs Eleven v3 Conversational is ~5× more expensive; Cartesia Sonic 3.6 is ~6.5× more expensive.
- Model size: Qwen3‑TTS Fast runs a 1.7 B‑parameter model, offering better quality‑latency than the 300 M‑parameter vui.
“Our Qwen3‑TTS Fast model is ranked #2 in Time‑to‑First‑Audio (TTFA), at p50 of 63 ms and WER of 3.8%, coming in at #1.” – Nari Labs blog
Cost‑efficiency across the board
Conclusion: Nari Labs leads the latency‑cost and quality‑cost Pareto frontiers, meaning you get the fastest, most accurate STT/TTS at the lowest price among publicly documented services.
- STT: $0.12 / hour (Fast) vs. $0.45 / hour for AssemblyAI’s comparable offering.
- TTS: $10 / 1M characters (Fast) vs. $50‑$65 / 1M characters for ElevenLabs and Cartesia.
- Pricing methodology: Comparisons use Coval’s pricing directory, which aggregates known public rates for each endpoint.
Community reaction on Hacker News
Conclusion: The discussion highlights enthusiasm for real‑time TTS, calls for independent verification, and curiosity about technical shortcuts.
- Real‑time viability: A user (mowmiatlas) shared a GitHub project for real‑time TTS, noting the trend toward ubiquitous low‑latency synthesis.
- Competitive landscape: iharnoor warned that the TTS market will become more crowded, emphasizing that voice models differ from LLM APIs in market dynamics.
- Quality concerns: A comment (rahimnathwani) reported a voice‑switching glitch in a 33‑second clip, suggesting the need for robustness testing.
- Verification demand: yoloakki urged independent evaluations (e.g., Datapoint AI) to confirm Nari’s claims.
- Technical curiosity: asaiacai asked about the primary levers for speeding up TTS, hypothesizing model distillation and architecture‑aware optimizations.
- Demo expectations: ipsum2 reminded that public announcements benefit from interactive demos; Nari Labs currently offers a free trial period.
- Open‑source curiosity: meatmanek inquired whether the ASR inference engine is open source; Nari Labs has not publicly disclosed this.
What the benchmark numbers mean for developers
Conclusion: For production voice agents, choosing Nari Labs’ Fast endpoints can reduce perceived latency by tens of milliseconds while cutting operating costs by up to 75 % compared to leading commercial alternatives.
- User experience: Lower TTFA/TTFS directly translates to faster turn‑around in conversational interfaces, reducing the “thinking” gap for end users.
- Scalability: The cheap per‑hour (STT) and per‑character (TTS) rates enable high‑volume deployments without prohibitive cloud spend.
- Integration: Nari Labs provides RESTful APIs with free‑trial access; a $20 credit is offered to early adopters as the service moves from beta to GA.
Takeaway: Nari Labs’ Qwen3‑ASR Fast and Qwen3‑TTS Fast set new industry standards for voice AI latency, accuracy, and cost, positioning them as compelling choices for developers building real‑time speech applications.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch