A transcript can be right and still miss the conversation
NVIDIA’s Nemotron 3 Diarization marks speaker activity over time, including overlapping speech. Transformers 5.18.0 added support for it. The open-weight model has about 100 million parameters and supports up to eight channels.
The channels are generic labels such as speaker_0, not verified identities. Diarization assigns time intervals to voices; speech recognition turns audio into words. Combining them can tie each phrase to a speaker and timestamp. Sources: https://github.com/huggingface/transformers/releases/tag/v5.18.0 and https://huggingface.co/blog/nvidia/nemotron-diarization
One model covers recorded and streaming audio
NVIDIA recommends buffer settings from 0.32 seconds for very low latency to 30.4 seconds for offline-style processing. Shorter buffers can return results sooner, while the company says more context generally improves accuracy and throughput.
Those buffer durations do not include model computation, network transport, speech recognition or application processing. A real deployment must test the full pipeline on its own audio and hardware.
The benchmark result is strong—and specific
In an initial VoiceArena evaluation, NVIDIA reports a 14.72% diarization error rate across 139 English conversations totaling about 22 hours. It ranked first among 12 systems and 17 configurations; the next-ranked system scored 19.3%.
The initial evaluation includes overlapping speech, and NVIDIA says the leaderboard analysis is still being completed. Its batch throughput tests on an RTX PRO 5000 are not end-to-end latency for one live call. Source: https://huggingface.co/blog/nvidia/nemotron-diarization
Source published September 30, 2026. Coverage is based on the maker’s announcement and demonstration.
