NVIDIA Nemotron 3.5 ASR 如何微调适配沙特阿拉伯语方言
Fine-Tuning NVIDIA Nemotron for Saudi Arabic Dialects, with a Path to Other Languages
NVIDIA 发布技术教程,介绍如何用 NeMo 框架微调 Nemotron 3.5 ASR,使其适配沙特 Najdi 和 Hijazi 方言。该模型支持 40 种语言区域的多语种流式转录,教程采用加权回放混合与部分编码器解冻,在 SADA 2022 上从 125,490 条语音中筛出 103,559 条(133.7 小时)用于训练。
Automatic speech recognition must handle how people actually speak, not only the languages and styles that dominate pretraining data. Regional dialects and local recording conditions are often underrepresented, so a multilingual model that performs well on broad benchmarks may still fall short in deployment.
Saudi Arabic makes that concrete. A model may recognize Modern Standard Arabic or English yet struggle with Najdi and Hijazi speech, or local recording conditions. Fine-tuning only on the target dialect can improve it while weakening other languages.
NVIDIA Nemotron 3.5 ASR supports multilingual streaming transcription across 40 language-locales, including transcription-ready Arabic, but deployment-specific dialects and recording conditions still benefit from fine-tuning.
This post shows how to adapt it with the NVIDIA NeMo framework and the ASR fine-tuning recipe: curate a low-resource corpus, build a weighted replay mix, fine-tune with efficient batching, and evaluate transcription quality on an independent set.
When to use this workflow
This pipeline is useful when you have enough labeled speech to specialize an ASR model, but not enough to train one from scratch: dialect adaptation, domain-specific transcription, deployments that must retain existing languages.
Its techniques solve different problems:
- Minimal curation removes unusable labels and obvious alignment failures without discarding scarce, difficult speech.
- Replay mixing interleaves a small amount of previously learned data to reduce catastrophic forgetting.
- Partial encoder unfreezing limits how many parameters change — cheaper and faster than a full fine-tune, at some cost to accuracy.
- Length bucketing reduces padding and makes training practical for streaming encoders.
- Beam search and a larger attention context can improve offline accuracy without retraining, at the cost of latency and compute.
These are not universal defaults. Replay protects only what its data represents; partial unfreezing needs re-tuning when the mix changes. And this workflow doesn’t generalize into evidence for every Arabic dialect or deployment environment.

Fine-tuning walkthrough
Prerequisites
- NVIDIA NeMo, PyTorch, Python, OmegaConf
- This tutorial uses SADA 2022 and FLEURS
- Working knowledge of Python, model fine-tuning, WER and CER
- Two GPUs were used for the 12,000-step baseline experiment; exact GPU model and memory: NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
1. Curate the target corpus without filtering away the problem
Start by selecting the dialects you intend to deploy. For the initial SADA experiment, that meant Najdi and Hijazi:
SAUDI_DIALECTS = {"najdi", "hijazi"}
seen = set()
for item in manifest:
dialect = str(item.get("speaker_dialect", "")).lower().strip()
seen.add(dialect)
if dialect not in SAUDI_DIALECTS:
continue
keep(item)
missing = SAUDI_DIALECTS - seen
assert not missing, f"never matched: {missing}; observed {sorted(seen)}"
Next, remove references the model cannot learn and clips that are probably misaligned.
SADA marks inaudible speech with غيرواضح; because the model cannot emit that annotation, every occurrence creates an unavoidable error.
In the code below, the markers appear as Unicode escapes so it renders left-to-right; they spell غيرواضح and غير واضح. The code applies a few other checks as well.
MIN_DURATION, MAX_DURATION = 0.5, 30.0 # drop clips too short or too long
MIN_CHAR_RATE, MAX_CHAR_RATE = 1.5, 35.0 # drop misaligned transcripts
ANNOTATION_MARKERS = ['\u063a\u064a\u0631\u0648\u0627\u0636\u062d',
'\u063a\u064a\u0631 \u0648\u0627\u0636\u062d']
for row in manifest:
txt = normalize_arabic(row['text'])
rate = len(txt) / row['duration']
if not txt or txt.lower() == 'nan': continue # stringified NaN
if any(m in row['text'] for m in ANNOTATION_MARKERS): continue # annotation, not speech
if not (MIN_DURATION <= row['duration'] <= MAX_DURATION): continue
if not (MIN_CHAR_RATE <= rate <= MAX_CHAR_RATE): continue
keep(row)
def normalize_arabic(t): # target dialect + FLEURS Arabic
t = re.sub(r'[\u064b-\u0670]', '', t) # diacritics
t = re.sub(r'[\u0623\u0625\u0622\u0627]', '\u0627', t) # alef variants
t = t.replace('\u0649', '\u064a').replace('\u0629', '\u0647') # alef maqsura, taa marbuta
t = re.sub(r'[^\u0600-\u06ff\w\s]', '', t) # punctuation
return ' '.join(t.split())
These checks retained 103,559 of 125,490 utterances: 133.7 hours, or 82.5% of the starting set. The goal is to remove structurally bad examples, not hard accents or noisy speech merely because the base model performs poorly on them. If you use automated quality scores, first inspect their distribution; the SADA run found that a default UTMOS threshold of 3.0 would have rejected almost everything.
As a second pass, the broader SADA curation pipeline used the NVIDIA NeMo Curator to standardize audio and remove severely degraded clips. MonoConversionStage converted inputs to mono, while UTMOSFilterStage and SIGMOSFilterStage scored perceived quality and background noise. Instead of applying their default thresholds, we first ran the stages in score-only mode, inspected the score distributions, then set corpus-specific cutoffs: UTMOS ≥ 1.25, SIGMOS noise ≥ 1.5, SIGMOS overall ≥ 1.5. They passed about 85% of duration-valid samples. This preserved challenging but usable dialect speech a generic threshold would have discarded.
2. Start with one dataset and monitor model behavior
For a new language or domain, begin with the simplest experiment you can interpret: one representative target dataset, a conventional full fine-tune, and a fixed evaluation set. The goal is to observe how the model responds, verify that the training pipeline works, and create a baseline against which later changes can be measured.
In our case, SADA served as that first dataset. The model is Cache-Aware FastConformer-RNNT with prompted multilingual streaming (strip_lang_tags, target_lang: ar-AR), which behaves differently from plain English streaming models and requires explicit language conditioning in the manifest.
Your starting corpus, training duration, and hardware settings will be different. The configuration changes below document how this particular experiment evolved as we encountered optimization, memory, and distributed-training issues; they are troubleshooting examples, not recommended defaults. Hyperparameters not listed, including weight decay, gradient clipping, mixed precision, and effective batch size in utterances, were left at NeMo framework defaults and weren’t tuned in this experiment.
| Area | Configuration changes |
|---|---|
| Optimization | Learning rate: default → 1e-4 → 2e-5; warmup: 10,000 → 100 → 50 steps; optimizer: AdamW (NeMo default); scheduler: Noam; d_model=1024 |
| Data loading | Batch duration: 400 → 300 → 200 seconds to prevent out-of-memory errors; is_tarred=false; workers: 4 training and 2 validation; bucket and shuffle buffers: 1,000 |
| Validation | Batch size 8; validation every epoch |
| Checkpoint and decoding | Retain the best three checkpoints; use greedy decoding |
In the initial SADA-only experiment, we used the validation split to monitor fine-tuning progress. The pretrained model produced 49.5% WER on the validation split and 59% WER on the full training corpus. The gap reflects noisier audio and greater dialectal variation in the training data.
The validation baseline (49.5%) served as our primary metric throughout. WER improved to 47.8% after the first 10-epoch run, remained at 47.8% after another 10 epochs, and reached 46.7% during the longer v4 continuation. After epoch 45, validation WER stopped improving. This modest gain and clear plateau suggested that continuing the same full fine-tuning setup was unlikely to deliver a substantial improvement.
3. What worked: A narrower target, a replay stream, and bucketed batches
A narrower target with minimal curation
Instead of asking the model to improve on 11dialects at once, we trained only on Najdi and Hijazi, the two we intended to use and we used minimal curation (see number 1, above), which kept 103,559 of 125,490 utterances, or 82.5%:
A replay stream
Fine-tuning on Saudi speech alone overwrites what the model learned in pretraining. The defence is replay: mix a small stream of previously-learned data back in so the model keeps being asked to do the old job while it learns the new one.
We used 10% FLEURS, split 7% English and 3% Arabic against 90% Saudi speech. Declare those proportions rather than concatenating the files because a sliding-window shuffle may not reach rows appended to the end of a large manifest until late in training, so a concatenated replay set effectively doesn’t exist for most of the run.
from omegaconf import OmegaConf
mix = OmegaConf.create([
{"type": "nemo", "manifest_filepath": "sada_train.jsonl", "weight": 0.90},
{"type": "nemo", "manifest_filepath": "fleurs_en.jsonl", "weight": 0.07},
{"type": "nemo", "manifest_filepath": "fleurs_ar.jsonl", "weight": 0.03},
])
OmegaConf.save(mix, "input_cfg.yaml")
cfg.train_ds.manifest_filepath = None
cfg.train_ds.input_cfg = "input_cfg.yaml"
Seven percent of English was enough. FLEURS English hold with a slight improvement from 11.04% to 10.42% (see Table 2, below), while the model was specialising in Arabic dialect speech.
Bucketed batches
The second change was less obvious and mattered just as much.
Duration-based bucketing groups similar-length utterances into the same batch:
cfg.train_ds.use_bucketing = True cfg.train_ds.num_buckets = 30 cfg.train_ds.batch_size = None cfg.train_ds.batch_duration = 400.0 cfg.train_ds.quadratic_duration = 15.0
Note that num_buckets alone does nothing; use_bucketing defaults to False, so setting the bucket count without the flag is a silent no-op that looks configured.
The result
Those three changes together, over 12,000 steps and about 4.5 hours on two GPUs:
| Test split | Before | After |
|---|---|---|
| SADA Najdi + Hijazi WER | 55.05% | 29.96% |
| SADA Najdi + Hijazi CER | 31.63% | 12.18% |
| Full SADA WER | 58.84% | 35.61% |
| Full SADA CER | 35.40% | 15.97% |
| FLEURS English WER | 11.04% | 10.42% |
| FLEURS English CER | 6.47% | 4.53% |
| FLEURS Arabic WER | 12.67% | 11.41% |
| FLEURS Arabic CER | 5.55% | 3.97% |
Specializing didn’t cost us the dialects we dropped. The model improved by 25 points where we aimed it and 23 points across everything, and English improved at the same time. That looked like a finished result. All numbers were measured with NeMo evaluation script.
4. Adaptation depth: How much of the encoder to update
The model has 24 encoder layers. A full fine-tune updates all of them; freezing the encoder preserves it but restricts acoustic adaptation. In between, you can unfreeze the top N layers and leave the rest alone, while always training the decoder, joint network and prompt embeddings.
for parameter in model.parameters():
parameter.requires_grad = False
for name, parameter in model.named_parameters():
if any(part in name for part in ("decoder", "joint", "prompt")):
parameter.requires_grad = True
for layer in model.encoder.layers[-8:]:
for parameter in layer.parameters():
parameter.requires_grad = True
trainable = sum(p.numel() for p in model.parameters() if p.requires_grad)
total = sum(p.numel() for p in model.parameters())
print(f"trainable: {trainable/1e6:.1f}M / {total/1e6:.1f}M")
In the recorded top-eight recipe, 230.4 million parameters were trainable and 407.6 million frozen. We tested two depths against each other, changing nothing else:
| Encoder layers updated | SADA WER | SADA CER |
|---|---|---|
| pretrained baseline | 55.05% | 31.63% |
| top 6 | 33.42% | 14.10% |
| top 8 | 32.32% | 13.53% |
| all 24 | 29.96% | 12.18% |
More trainable capacity was better at every step. With 134 hours of target speech, the data supports updating the whole encoder, and partial freezing costs 2.4 points against the full fine-tune.
That is a finding about this data volume, not a general rule. Freezing is a way of spending less so it becomes the right choice when your corpus is thinner than ours or memory is the binding constraint.
5. Improving without retraining: Larger context and beam search decoding
Training is only half the system. Before paying for another training run, check what you can get at inference time.
This checkpoint exposes several attention context sizes, selectable without retraining. The second number is how many future frames the encoder may attend to before committing to an output: [56, 3] is the streaming default, [56, 13] the widest it supports. Switching to the highest-lookahead context reduced WER by 1.31 absolute points without retraining. Its main trade-off is approximately 800 ms of additional latency.
| Configuration | WER | CER | vs. Greedy |
|---|---|---|---|
Greedy, [56, 3] | 29.96% | 12.18% | +0.00 |
Greedy, [56, 13] | 28.65% | 11.39% | -1.31 |
MALSD beam-4, [56, 3] | 28.81% | 11.48% | -1.15 |
MALSD beam-8, [56, 3] | 28.62% | 11.40% | -1.34 |
MALSD beam-8, [56, 13] | 27.25% | 10.63% | -2.71 |
The primary trade-off is additional latency rather than retraining cost. Thirteen lookahead frames instead of three is roughly 800 ms more buffering before each output. For batch transcription, such as call archives, media, meeting recordings, that provides an accuracy gain without retraining, where the additional buffering latency is acceptable. For live captioning, it may be unsuitable. The right setting follows from the deployment, not from the accuracy number.
Beam-search performance depended strongly on the decoder configuration. In a later sweep, NeMo’s batched malsd_batch strategy produced useful gains at modest cost: beam 4 reduced SADA WER from 29.96 % to 28.81 % at 0.59× the greedy runtime, while beam 8 reached 28.62% at 0.64×. Beam 4 therefore offered the best speed–accuracy balance in our test. We did not evaluate MAES or NGPU-LM fusion, so our conclusion is limited to MALSD.
One detail worth setting regardless of strategy: strip_lang_tags=True. Otherwise the locale tag is emitted as literal text and scored as an insertion on every utterance.
| Attention Context | Chunk Size (Latency) | Use Case |
|---|---|---|
[56, 0] | 80ms (Ultra-Low) | Ultra-low-latency Voice Agents |
[56, 1] | 160ms (Low) | Interactive Voice Agents, Conversational AI |
[56, 3] | 320ms (Balanced) | Conversational AI, Live Captioning |
[56, 6] | 560ms (Medium) | High accuracy with reasonable latency |
[56, 13] | 1.12s (High) | Highest accuracy with high latency |
Applying the workflow to other languages
The loop stays the same: curate, mix replay, choose an adaptation depth, bucket by duration, evaluate on independent target and regression sets. What changes per language is the corpus metadata, transcription conventions, normalization, tokenizer coverage, script handling, and capability metrics.
For another language, start with a small audited manifest and test the base tokenizer on names, numerals, borrowed words, and mixed-script sentences. Replace the Arabic normalizer with one that preserves meaningful distinctions in the target script, and validate unicode normalization and encoding end to end.
For languages without whitespace word boundaries, WER may be misleading, Use a character-, token-, or morpheme-level measure instead, and state its limits. Replay data should represent the capabilities you need to retain, not just convenient high-resource speech, and should be real rather than synthetic wherever possible.
Extend the workflow with speaker diarization
Fine-tuning ASR improves what is transcribed; speaker diarization adds who spoke when. In multi-speaker audio, a diarization model identifies speech segments and assigns consistent speaker labels. Combining those labels and timestamps with the fine-tuned Nemotron 3.5 ASR output creates a speaker-attributed transcript that separates each participant’s words, useful for meetings, interviews, contact centers, classrooms, and other multi-speaker environments.
This workflow can also help speech-data providers prepare multi-speaker corpora by automatically generating speaker-turn timestamps and anonymous speaker labels for human review before the data is used for ASR fine-tuning or evaluation.
NVIDIA Nemotron 3 Diarization, just released, extends this workflow from dialect-aware transcription to speaker-attributed transcription for up to 8 speakers. It is an open model that can be added to any ASR system you’re already using. Diarization doesn’t transcribe speech itself; its speaker boundaries and labels are aligned with the ASR timestamps to produce the final structured transcript. Learn more about the architecture, benchmarks, and how to get started in the Hugging Face blog.
Next steps
The conversation that prompted this tutorial raised a more important question than capability alone: How to make dialect adaptation useful, improving the target dialect, retaining existing skills, and spending training effort where it changes the deployed system. The experiments point to a practical answer. Curate lightly, mix replay by weight, add examples of the exact behavior you need, and update as much of the encoder as your data supports. Then optimize decoding, and evaluate each capability on an independent set.
Get started
Finetuning notebook:
https://github.com/nvidia-riva/tutorials/blob/main/asr-finetune-nemotron-3.5-asr-streaming-prompt.ipynb
Finetuning skill:
https://github.com/NVIDIA/skills/tree/main/skills/nemotron-asr-finetune
来源:NVIDIA Technical Blog:Agentic AI / Generative AI · developer.nvidia.com