Nvidia released an open model that tells 8 speakers apart in real time
Nemotron 3 Diarization is a 100 million parameter model that works out who spoke when in an audio recording. It runs both on live streams and on recordings, and can separate up to eight speakers, including overlapping speech. It ranks first on VoiceArena’s Diarization-Bench with a 14.72% error rate, an average 41% relative reduction against Nvidia’s previous four-speaker model. The weights are on Hugging Face under the OpenMDW license, which permits commercial use. Uses include meeting transcripts, call centers and voice assistants.
Sources
- Nvidia (Hugging Face blogu), “Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization”, (huggingface.co)
- Nvidia (Hugging Face), “nvidia/Nemotron-3-Diarization”, Model card, (huggingface.co)
- Voice Arena, “Diarization Bench: speaker diarization leaderboard, early results”, Leaderboard (voicearena.com)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 30, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/nemotron-diarization
This story in Turkish: Nvidia, gerçek zamanlı olarak 8 konuşmacıyı ayırt eden açık bir model yayınladı
On the same topic
AWS adds Chinese lab Z.ai’s GLM 5.3 model to Amazon Bedrock
AWS added GLM 5.3, a 753-billion-parameter model from Beijing-based Z.ai (formerly Zhipu AI), to Amazon Bedrock for eligible enterprise customers on October 5.
Nvidia unveils a 64 GB DGX Spark that runs models of up to 100 billion parameters locally
Nvidia announced a 64 GB configuration of its DGX Spark desktop AI computer on October 2, sold through partners from October 23 at a starting price of $4,999.
Microsoft releases three new voice models, one transcribing 60 languages in real time
On October 1, Microsoft AI announced MAI-Transcribe-2-Streaming, which transcribes 60 languages in real time, and the 23-language MAI-Voice-2.1 and Flash.