NVIDIA Nemotron 3 Diarization 100M‑parameter model release
TL;DR
NVIDIA announced the open‑weight Nemotron‑3 Diarization model—a 100 M‑parameter transformer that supports up to eight speakers, achieves a 14.72 % diarization error rate (DER) to rank #1 on VoiceArena’s Diarization‑Bench, and works for both offline and real‑time streaming use cases.
Why speaker diarization matters
Speaker diarization adds the "who spoke when" layer to speech transcripts. Without speaker attribution, a transcript cannot reliably link commitments, objections, or interruptions to the correct participant, limiting search, summarisation, action‑item extraction, and voice‑agent memory.
Model overview and performance
- Model size: 100 M parameters, open‑weight checkpoint on Hugging Face.
- Speaker capacity: Up to eight anonymous speaker channels.
- Leaderboard ranking: #1 on VoiceArena’s Diarization‑Bench (14.72 % DER) out of 12 systems and 17 configurations.
- Latency options: Input‑buffer latencies of 30.4 s (offline style), 1.04 s, 0.64 s, and 0.32 s.
- Throughput: 15,113× real‑time factor (RTFx) at 30.4 s latency (batch‑size 32, BF16, torch.compile) versus 2,619× for the previous four‑speaker Sortformer baseline.
Architecture and inference flow
- Audio preprocessing – 16 kHz mono audio → Mel‑spectrogram with 10 ms frame step, stacked 8× to produce 80 ms frames.
- Transformer encoder – 31‑layer transformer with rotary positional embeddings (RoPE).
- Conv1D up‑sampler – Restores predictions to the original feature resolution.
- Output tensor – Shape
[T, 8]where each column is the probability that a given speaker channel is active at time step T (default 10 ms stride). - Streaming memory –
- Arrival‑Order Speaker Cache (AOSC) retains speaker embeddings across chunks, ordered by first appearance.
- FIFO queue supplies recent frame context.
- Right‑context buffering – Provides future audio to improve transition handling; configurable to trade latency vs. accuracy.
The model follows the Sortformer principle: speakers are ordered by arrival time, giving stable generic labels (speaker_0, speaker_1, …) and eliminating per‑chunk permutation solving.
Training data and impact of licensed sources
Training used public speech corpora plus licensed multi‑speaker conversations from David AI. Adding the David AI data reduced compound DER by 0.77 absolute points (from 11.19 % to 10.42 %) at both offline and ultra‑low‑latency operating points.
Offline vs. streaming diarization
- Offline: Model sees the entire recording, allowing global context for speaker assignment.
- Streaming: Model processes fixed‑size chunks with limited left/right context, relying on AOSC and FIFO to maintain speaker identity across chunks. The same checkpoint serves both modes, simplifying deployment.
Benchmark results and relative improvements
| Dataset (subset) | Baseline DER | Nemotron‑3 DER | Relative reduction |
|---|---|---|---|
| VoiceArena (overall) | 19.3 % | 14.72 % | ~24 % |
| DIHARD III (full) | 19.09 % | 12.73 % | 33 % |
| CALLHOME‑Part2 (1.04 s) | 10.32 % | 9.10 % | 12 % |
| NOTSOFAR1 MHM (1.04 s) | 27.5 % | 9.6 % | 65 % |
Across eight evaluation conditions, the unweighted mean relative DER reduction at 1.04‑second latency is 41 %. Improvements are larger for higher speaker‑count subsets (five‑plus speakers) and consistent across all latency settings.
Accuracy vs. throughput trade‑offs
- 30.4 s buffer: 15,113× RTFx, DER 12.73 % (DIHARD III).
- 1.04 s buffer: 865× RTFx, DER 13.18 % (DIHARD III).
- 0.32 s buffer: Lowest latency, but DER rises modestly; still outperforms the prior Sortformer baseline.
Developers should benchmark end‑to‑end latency, including audio I/O, diarization, ASR, and downstream processing, on target hardware.
Integration with speaker‑attributed ASR
Diarization supplies timestamps per speaker channel; ASR supplies word‑level timestamps. A simple midpoint‑alignment heuristic can map each word to the active speaker:
midpoint = (word['start'] + word['end']) / 2
label = speaker_at(midpoint)
More sophisticated overlap handling may be required for production‑grade transcripts.
Deployment considerations
- Speaker limit: Model supports a maximum of eight concurrent speakers; exceeding this may cause missed speech or mis‑assignments.
- Acoustic robustness: Severe noise, reverberation, far‑field capture, or domain shift can degrade performance.
- Uncertainty handling: Downstream systems should retain confidence scores rather than treating assignments as absolute.
- License: Use is governed by the OpenMDW License v1.1.
Getting started with NVIDIA NeMo Speech
apt-get update && apt-get install -y libsndfile1 ffmpeg
uv pip install Cython packaging
uv pip install 'nemo-toolkit[asr]'
from nemo.collections.asr.models import SortformerEncLabelModel
diarmodel = SortformerEncLabelModel.from_pretrained(
"nvidia/Nemotron-3-Diarization"
)
diarmodel.eval()
# Example streaming parameters for 1.04 s latency
diarmodel.sortformer_modules.spkcache_len = 264
ndiarmodel.sortformer_modules.fifo_len = 40
ndiarmodel.sortformer_modules.chunk_len = 9 # 720 ms
ndiarmodel.sortformer_modules.chunk_right_context = 4
ndiarmodel._check_streaming_parameters()
segments = diarmodel.diarize(["/path/to/conversation.wav"], batch_size=1)
print(segments)
The diarize() method returns speaker‑marked segments (start end speaker_id). Set include_tensor_outputs=True to obtain raw probability tensors.
Real‑time demo and ecosystem support
- Live demo: https://huggingface.co/spaces/nvidia/nemotron-diarization – supports synthetic eight‑speaker conversations, live microphone input, and multilingual streams.
- Argmax Pro SDK 3: Adds on‑device, real‑time speaker attribution for up to eight speakers and a pre‑diarized transcription API.
- Partner deployments: Baseten, DigitalOcean, and Argmax provide production‑grade inference, scaling, and mobile integration pathways.
Resources
- Model card: https://huggingface.co/nvidia/Nemotron-3-Diarization
- VoiceArena Diarization‑Bench leaderboard
- Argmax Pro SDK 3 announcement
- NVIDIA NeMo Speech repository
- Sortformer paper (arXiv:2409.06656) and Streaming Sortformer paper (arXiv:2507.18446)
- Quick start guide for diarization + ASR integration
- Interactive demo space
Implications
Nemotron‑3 Diarization demonstrates that a single, relatively small (100 M) transformer can deliver state‑of‑the‑art speaker‑attribution accuracy across eight speakers while supporting a wide range of latency budgets. This lowers the barrier for real‑time, multi‑speaker transcription in meetings, call‑center analytics, podcasts, and on‑device voice assistants, and it establishes a new baseline for future diarization research and productization.