NVIDIA Just Open-Sourced a 100M-Parameter Model That Knows Who's Talking
Ever read a meeting transcript where every line is correct but you can't tell who said what? NVIDIA just fixed that — and gave the fix away for free. On September 23, NVIDIA released Nemotron 3 Diarization , an open-weight, 100-million-parameter model that answers one of audio AI's most annoying questions: "who spoke when," in live or recorded audio. It handles up to 8 simultaneous speakers with speaker-activity probabilities at 10-millisecond resolution — granular enough to catch rapid-fire crosstalk and interruptions. Streaming latency profiles go as low as about 320 ms, and there's an offline mode for full recordings. The model ranked #1 on VoiceArena's Diarization-Bench leaderboard with a 14.72% diarization error rate. It was trained on roughly 10,000 hours of real conversations plus 82,611 hours of simulated multi-talker mixtures. Under the hood is a neat engineering choice: a single-pass end-to-end design NVIDIA calls the "Arrival-Order Speake...


Comments
Post a Comment