매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

arXiv:2608.193612026-08-21

인도 소수언어 미조어 음성인식, 데이터 17시간 모으고 Whisper와 SraVaani로 학습시켜보니

연구진이 인도 미조람 지역에서 쓰이는 저자원 언어 미조어의 음성 데이터 17.62시간을 200명의 화자로부터 모아 정제했다. 이 데이터로 원래 미조어를 지원하지 않는 Whisper 모델과 미조어를 이미 지원하는 인도어 특화 모델 SraVaani 1.0을 각각 미세조정해 성능을 비교했다. Whisper-large-v3가 일반적인 단어오류율 18.08%로 가장 낮았고, 미조어 특유의 띄어쓰기 차이를 감안한 새 평가법으로는 7.22%까지 낮아졌다.

무엇을 했나

  1. 미조람 지역 화자 200명이 뉴스와 판결문에서 뽑은 약 8천 문장을 웹으로 녹음해 17.62시간, 8274개 문장 단위 음성파일로 정제된 데이터셋을 구축했다.
  2. 이 데이터로 Whisper의 small/medium/large-v3 세 모델과, 원래 미조어를 지원하는 인도어 다국어 모델 SraVaani 1.0을 각각 미세조정하고 화자가 겹치지 않는 학습/검증/테스트로 나눠 평가했다.
  3. 미조어는 한 형태소를 붙여 쓰거나 띄어 쓰는 게 자유로워 기존 단어오류율(WER)이 실제보다 오류를 부풀리는 문제가 있어, 인접한 최대 4개 단어를 이어 붙여 비교하는 형태소 인식 WER(MA-WER)을 새로 만들었다.
  4. Whisper-large-v3는 일반 WER 18.08%, MA-WER 7.22%로 최고 성능을 냈고, SraVaani 1.0은 원래 상태에서 WER 58.27%였다가 미조어로 미세조정한 뒤 WER 29.45%, MA-WER 17.93%로 크게 개선됐다.
  5. 오류 분석 결과 SraVaani 1.0 원본 모델은 미조어를 다른 인도 문자(메이테이-마엑, 데바나가리)로 잘못 인식하는 사례가 있었고, 인명·지명 오인식, 성문파열음 인식 오류 등이 미세조정 후 크게 줄었다.
Figure 1: Histogram of duration of speech files in the corpus.
Figure 1: Histogram of duration of speech files in the corpus.
Table 1: Data split for training, validation and testing.
SpeakersSentencesHours
Training184765616.18
Validation114261.02
Testing051920.42
Figure 2: Schematic diagram showing the flow of the experimental protocol.
Figure 2: Schematic diagram showing the flow of the experimental protocol.
Table 2: Durational characteristics of the speech database
Sentences8274
Total duration17.62 hours
Minimum duration0.63 seconds
Maximum duration41.22 seconds
Mean duration7.67 seconds
Median duration6.94 seconds
Table 3: Characteristics of the multilingual ASR models evaluated in this study.
ModelArchitectureParametersPretraining dataMel bins
Whisper-smallEncoder–Decoder Transformer244 M680k h80
Whisper-mediumEncoder–Decoder Transformer769 M680k h80
Whisper-large-v3Encoder–Decoder Transformer1,550 M∼5M h128
SraVaani 1.0FastConformer Hybrid RNNT/CTC∼430 M∼31k h
Table 4: Fine-tuning configurations used for the ASR models.
ModelOptimizerLREffective batch sizeEpochsSelection
Whisper-smallAdafactor5×10−61620Best val. WER
Whisper-mediumAdafactor5×10−61620Best val. WER
Whisper-large-v3Adafactor5×10−61620Best val. WER
SraVaani 1.0AdamW1×10−41620 + extendedBest val. WER
Table 5: Best epochs and corresponding validation WER of four fine-tuned models.
ModelBest epochValidation WER
Whisper-small Mizo–FT1528.99
Whisper-medium Mizo–FT1326.51
Whisper-large-v3 Mizo–FT1323.00
SraVaani 1.0 Mizo–FT1833.81
Table 6: Results of evaluation on five models.
ModelCER (%)WER (%)MA-WER (%)
Whisper-small Mizo–FT04.8324.0011.49
Whisper-medium Mizo–FT04.0221.6908.87
Whisper-large-v3 Mizo–FT03.2618.0807.22
SraVaani 1.017.7158.2736.27
SraVaani 1.0 Mizo-FT06.9029.4517.93
Table 7: Error distribution by model. Foreign-script outputs are reported as the number of utterances; all other entries indicate the number of error occurrences.
ModelForeign scriptNamesGlottal stopsNumeral transcriptsCode-mix error< t Ω >
Whisper-small Mizo–FTNIL177290
Whisper-medium Mizo–FTNIL186530
Whisper-large-v3 Mizo–FT4 sentences84320
SraVaani 1.021 sentences449194931
SraVaani 1.0 Mizo–FTNIL2462194

왜 중요한가

이 연구는 인도의 저자원 부족어인 미조어에 대해 처음으로 공개 음성 데이터셋과 미세조정된 ASR 모델, 그리고 언어 특성을 반영한 평가 지표를 함께 내놓아 향후 미조어 음성기술 개발의 발판을 제공한다. 또한 표준 WER이 형태소 경계 표기가 자유로운 언어에서는 실제 인식 품질을 과소평가할 수 있다는 점을 보여줘, 비슷한 티베트-버마어족 언어 평가에도 시사점을 준다.

이 논문의 용어

  • WER(단어오류율) · 음성인식 결과와 정답 문장을 단어 단위로 비교해 틀린 정도를 나타내는 표준 지표
  • MA-WER(형태소 인식 WER) · 미조어처럼 형태소 경계의 띄어쓰기가 자유로운 언어를 위해, 인접 단어를 최대 4개까지 붙여 비교하도록 만든 새 오류율 지표
  • CER(문자오류율) · 단어가 아니라 글자 단위로 오류를 세는 지표
  • 제로샷 평가 · 해당 언어 데이터로 추가 학습하지 않고 사전학습된 모델 그대로 성능을 테스트하는 방식
  • Whisper / SraVaani 1.0 · Whisper는 다국어 음성인식·번역이 가능한 트랜스포머 기반 모델, SraVaani 1.0은 인도어 여러 언어를 지원하는 다국어 음성인식 모델

논문 원문 초록 (영문)

This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.

저자 · Priyankoo Sarmah, Sanasam Ranbir Singh, Lalhmingmawia

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Priyankoo Sarmah et al., arXiv:2608.19361, cc-by-nc-nd-4.0