NVIDIA Leverages SDAIA's SADA Dataset to Enhance Nemotron 3.5 ASR Model, Cut Saudi Dialect Speech Recognition Errors by Half
Global computing and artificial intelligence leader NVIDIA announced that it utilized the Saudi Audio Dataset for Arabic (SADA), launched by the Saudi Data and AI Authority (SDAIA) in partnership with the Saudi Broadcasting Authority, to train its "Nemotron 3.5 ASR" multilingual automatic speech recognition model. According to a technical study published on NVIDIA Developer blog, integrating the dataset reduced comprehension errors for Saudi dialects by nearly half.
The model, which supports real-time transcription across about 40 languages and dialects including Arabic, initially misidentified about 55 out of every 100 words in Saudi dialects. After training on 133.7 hours of Najdi and Hijazi audio from SADA, the word error rate dropped to roughly 30 words. Across the entire multi-dialect dataset, the error rate decreased from 58.8% to 35.6%, while character-level errors fell from 31.6% to 12.2%, alongside simultaneous performance gains in Modern Standard Arabic and English. The training was completed in just four and a half hours using two GPUs.
Optimized for real-time interactive systems, the enhanced model features latency starting at 80 milliseconds, enabling applications in intelligent voice assistants, conversational agents, live broadcast subtitling, and the transcription of call centers, media archives, and meetings. NVIDIA has also shared the required workflows and tools to allow global developers to replicate the methodology across other languages.
SDAIA previously released SADA dataset on the Kaggle platform to support international researchers and developers. The dataset spans some 667 hours of transcribed audio, including over 600 hours provided by the Saudi Broadcasting Authority from 57 television programs and series covering more than 10 Saudi dialects, with 20 hours reserved for validation. Comprising over 125,000 categorized clips, it enabled NVIDIA to target specific Najdi and Hijazi speech data.
The SADA initiative supports academic and technical efforts to develop advanced audio models, including speech recognition, text-to-speech synthesis, speaker diarization, and demographic classification, while enriching digital Arabic content as the language of the Holy Quran spoken by millions globally.
NVIDIA's integration of the dataset highlights the pivotal role of high-quality national data in aligning global AI systems with local dialects and improving technologies tailored for Arabic speakers.



