Skip to content
Türkçe

Nvidia · Models

Nvidia released an open model that tells 8 speakers apart in real time

Published: 3 sourcesTürkçe

Nemotron 3 Diarization is a 100 million parameter model that works out who spoke when in an audio recording. It runs both on live streams and on recordings, and can separate up to eight speakers, including overlapping speech. It ranks first on VoiceArena’s Diarization-Bench with a 14.72% error rate, an average 41% relative reduction against Nvidia’s previous four-speaker model. The weights are on Hugging Face under the OpenMDW license, which permits commercial use. Uses include meeting transcripts, call centers and voice assistants.

Sources

  1. Nvidia (Hugging Face blogu), “Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization”, (huggingface.co)
  2. Nvidia (Hugging Face), “nvidia/Nemotron-3-Diarization”, Model card, (huggingface.co)
  3. Voice Arena, “Diarization Bench: speaker diarization leaderboard, early results”, Leaderboard (voicearena.com)

About this story

This story was posted on Instagram by @jarrus.tech on Sept. 30, 2026.

Spotted an error in this story? [email protected] · Instagram

This story in Turkish: Nvidia, gerçek zamanlı olarak 8 konuşmacıyı ayırt eden açık bir model yayınladı

On the same topic