Nvidia released an open model that tells 8 speakers apart in real time
Nemotron 3 Diarization is a 100 million parameter model that works out who spoke when in an audio recording. It runs both on live streams and on recordings, and can separate up to eight speakers, including overlapping speech. It ranks first on VoiceArena’s Diarization-Bench with a 14.72% error rate, an average 41% relative reduction against Nvidia’s previous four-speaker model. The weights are on Hugging Face under the OpenMDW license, which permits commercial use. Uses include meeting transcripts, call centers and voice assistants.
Sources
- Nvidia (Hugging Face blogu), “Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization”, (huggingface.co)
- Nvidia (Hugging Face), “nvidia/Nemotron-3-Diarization”, Model card, (huggingface.co)
- Voice Arena, “Diarization Bench: speaker diarization leaderboard, early results”, Leaderboard (voicearena.com)
About this story
This story was posted on Instagram by @jarrus.tech on Sept. 30, 2026.
Spotted an error in this story? [email protected] · Instagram
Short link: thejarrus.com/en/h9-26
This story in Turkish: Nvidia, gerçek zamanlı olarak 8 konuşmacıyı ayırt eden açık bir model yayınladı
On the same topic
Anthropic releases Claude Haiku 5.5 at 90% lower prices
Claude Haiku 5.5, released Oct. 7, costs 90% less than Haiku 4.5 for prompts up to 100,000 tokens. It leads on benchmarks but uses more tokens per task.
OpenAI is rolling out GPT-6 to everyone in ChatGPT
OpenAI began rolling out GPT-6 to everyone in ChatGPT on October 7. Its new Intelligent UI builds answers from text, visuals and interactive elements.
Stripe bought AI model marketplace OpenRouter for $8 billion. Nvidia was reportedly getting ready to make a last-minute counteroffer
According to The Information, Nvidia told OpenRouter it was prepared to make a counteroffer to Stripe’s deal, but it never submitted a formal bid.