Figure 1: Histogram of duration of speech files in the corpus.
Table 1: Data split for training, validation and testing.
Speakers
Sentences
Hours
Training
184
7656
16.18
Validation
11
426
1.02
Testing
05
192
0.42
Figure 2: Schematic diagram showing the flow of the experimental protocol.
Table 2: Durational characteristics of the speech database
Sentences
8274
Total duration
17.62 hours
Minimum duration
0.63 seconds
Maximum duration
41.22 seconds
Mean duration
7.67 seconds
Median duration
6.94 seconds
Table 3: Characteristics of the multilingual ASR models evaluated in this study.
Model
Architecture
Parameters
Pretraining data
Mel bins
Whisper-small
Encoder–Decoder Transformer
244 M
680k h
80
Whisper-medium
Encoder–Decoder Transformer
769 M
680k h
80
Whisper-large-v3
Encoder–Decoder Transformer
1,550 M
∼5M h
128
SraVaani 1.0
FastConformer Hybrid RNNT/CTC
∼430 M
∼31k h
–
Table 4: Fine-tuning configurations used for the ASR models.
Model
Optimizer
LR
Effective batch size
Epochs
Selection
Whisper-small
Adafactor
5×10−6
16
20
Best val. WER
Whisper-medium
Adafactor
5×10−6
16
20
Best val. WER
Whisper-large-v3
Adafactor
5×10−6
16
20
Best val. WER
SraVaani 1.0
AdamW
1×10−4
16
20 + extended
Best val. WER
Table 5: Best epochs and corresponding validation WER of four fine-tuned models.
Model
Best epoch
Validation WER
Whisper-small Mizo–FT
15
28.99
Whisper-medium Mizo–FT
13
26.51
Whisper-large-v3 Mizo–FT
13
23.00
SraVaani 1.0 Mizo–FT
18
33.81
Table 6: Results of evaluation on five models.
Model
CER (%)
WER (%)
MA-WER (%)
Whisper-small Mizo–FT
04.83
24.00
11.49
Whisper-medium Mizo–FT
04.02
21.69
08.87
Whisper-large-v3 Mizo–FT
03.26
18.08
07.22
SraVaani 1.0
17.71
58.27
36.27
SraVaani 1.0 Mizo-FT
06.90
29.45
17.93
Table 7: Error distribution by model. Foreign-script outputs are reported as the number of utterances; all other entries indicate the number of error occurrences.
This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.