Sep 27, 2026
ManyPress

Advertisement

Artificial Intelligence

Nvidia has launched a 100-million parameter model capable of identifying up to eight speakers in real-time audio.

ManyPress

ManyPress

ManyPress Editorial

2 min readSource:The Decoder
Nvidia releases open-source Nemotron 3 Diarization AI model

Key facts

  • •The Nemotron 3 Diarization model contains approximately 100 million parameters.
  • •It can track up to eight distinct speakers simultaneously.
  • •The model holds the top position on the VoiceArena Diarization-Bench with a 14.72% error rate.
  • •It outperforms the previous Streaming Sortformer model by an average of 41%.
  • •The system supports adjustable audio buffer settings ranging from 0.32 to 30.4 seconds.

Nvidia has released Nemotron 3 Diarization, an AI model designed to identify individual speakers within a conversation. The model, which features 100 million parameters and is freely available, can distinguish between up to eight speakers and detect instances of overlapping speech. It is compatible with both live audio and recorded files.

By the numbers

100 million
number of model parameters
14.7%
Diarization Error Rate on VoiceArena benchmark
41%
average error rate reduction over predecessor

Performance and Benchmarking

In the VoiceArena Diarization Benchmark v1, Nemotron 3 achieved a Diarization Error Rate (DER) of 14.7%, placing it first in the current rankings. This performance represents a 41% improvement over Nvidia's previous model, Streaming Sortformer, across eight test scenarios using a 1.04-second buffer. The benchmark is noted for its strict scoring, which counts overlapping speech and minor misalignments at speaker transitions as errors. The model's accuracy is influenced by the audio buffer size, which can be adjusted between 0.32 and 30.4 seconds, with shorter buffers generally resulting in higher error rates. Additionally, performance can be impacted by heavy background noise, significant reverb, or a high number of participants.

Integration and Capabilities

When paired with a speech recognition system like Parakeet, Nemotron 3 can generate transcripts that include speaker labels, such as "speaker_2." While the model identifies the presence of different speakers, it provides anonymous labels rather than identifying specific individuals.

Advertisement

This article was independently rewritten by ManyPress editorial AI from reporting originally published by The Decoder.

Artificial Intelligence