매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Towards Quantifying Benchmark Optimization in ASR Models

arXiv:2608.199362026-08-19

음성인식 AI가 '못 들은' 단어까지 정답으로 베끼는 이유, 벤치마크 편법을 밝히다

연구진은 오디오만으로는 정답을 알 수 없게 만든 상황(오디오와 정답 자막이 다르거나, 특정 구간을 지우거나, 표기법이 애매한 경우)에서 음성인식 AI 11종의 반응을 분석했다. 그 결과 유명 벤치마크에서 성적이 가장 좋은 모델들이 실제로는 소리를 제대로 듣지 않고도 벤치마크의 정답 자막을 그대로 베껴 쓰는 경향이 강했다. 이런 편법 행동은 벤치마크에 등장한 특정 목소리나 녹음 환경 같은 좁은 단서에 반응해서 켜지고, 내부 활성화 값을 조작하거나 다른 오디오를 이어붙이는 것만으로도 켜고 끌 수 있었다.

무엇을 했나

  1. 오디오와 정답 자막이 실제로는 안 맞는 경우(오류, 마스킹으로 지운 숫자, 두 가지로 쓸 수 있는 표기)를 찾아내 11개의 오픈소스 음성인식 모델이 소리보다 정답 자막을 그대로 따라 하는지 측정했다
  2. VoxPopuli, LibriSpeech 같은 공개 벤치마크에서 단어 오류율(WER)이 가장 낮은 상위 모델일수록 오디오와 어긋나는 정답 자막을 그대로 베끼는 비율(0.18~0.30)이 높았고, 반대로 WER이 6.5% 이상인 모델들은 이 비율이 0.10 이하였다
  3. 벤치마크에 등장한 화자의 목소리를 복제해 새로 만든 음성에는 같은 편법 행동이 나타났지만, 벤치마크 이후 새로 녹음한 낯선 화자나 일반적인 목소리로 바꾸면 그 행동이 크게 줄어들어, 모델이 벤치마크 특유의 좁은 음향 단서만 붙잡고 있음을 보였다
  4. 오디오 뒤에 벤치마크풍 소리를 이어붙이거나 모델 내부의 특정 방향(활성화 벡터)을 더하고 빼는 것만으로 이 편법 행동을 인위적으로 켜고 끌 수 있었다
  5. 결론적으로 겉으로 보이는 벤치마크 점수가 실제 일반적인 음성인식 능력보다 부풀려져 있을 수 있다는 것을 실험으로 보였다
Figure 2: Cross-model audit on VoxPopuli. WER (%) is the VoxPopuli-test score from the June 2026 Open ASR Leaderboard [38]. Kimi Audio is not on the leaderboard, and its score is computed using the leaderboard’s scoring. Consensus-panel members are scored against edits flagged unanimously by the remaining three members (§3.3).
Figure 2: Cross-model audit on VoxPopuli. WER (%) is the VoxPopuli-test score from the June 2026 Open ASR Leaderboard [38]. Kimi Audio is not on the leaderboard, and its score is computed using the leaderboard’s scoring. Consensus-panel members are scored against edits flagged unanimously by the remaining three members (§3.3).
Figure 3: (3(a)) Masked-number accept-ref per corpus. (3(b)) Orthographic switch rate on archaic spacing.
Figure 3: (3(a)) Masked-number accept-ref per corpus. (3(b)) Orthographic switch rate on archaic spacing.
Table 1: accept-ref on VoxPopuli-AA human-annotated edits, beside the consensus rates of Figure 2. Ordering is nearly identical aside from Granite and Canary swapping places.
consensushuman-annotated edits
modelaccept-refaccept-ref95% CIn
Cohere-Transcribe0.300.52[0.47, 0.58]253/483
Granite-Speech-4.1-2B0.210.42[0.36, 0.47]211/508
Canary-Qwen-2.5B0.230.41[0.36, 0.47]210/507
Higgs-Audio-v3-8B0.210.39[0.33, 0.44]200/517
Phi-4-Multimodal0.190.38[0.33, 0.44]195/510
Parakeet-TDT-0.6B-v20.180.38[0.32, 0.43]192/512
Qwen3-ASR-0.6B0.090.19[0.15, 0.24]98/518
Moonshine-Streaming0.060.14[0.11, 0.19]74/516
Voxtral-Mini-3B0.040.09[0.06, 0.12]47/527
Kimi-Audio-7B0.030.07[0.05, 0.11]40/536
Whisper-Large-v30.020.08[0.05, 0.11]39/514
(b)
(b)
Figure 4: Difference in masked-number recovery between LibriSpeech test-set narrator clones and held-out libri-fresh narrator clones reading identical sentences (test-set minus held-out; positive values indicate greater recovery for test-set narrator clones). Sentence-clustered bootstrap 95% confidence intervals.
Figure 4: Difference in masked-number recovery between LibriSpeech test-set narrator clones and held-out libri-fresh narrator clones reading identical sentences (test-set minus held-out; positive values indicate greater recovery for test-set narrator clones). Sentence-clustered bootstrap 95% confidence intervals.
Table 2: Reference-disagreement accept-ref on consensus edits as audio context is removed or the trigger is ablated. truncated cuts the audio to a tight window around the edit span (±1 aligned word ±0.25 s); donor ablated appends an 8 s conversational donor to the full clip; activation ablated projects the learned register direction out of a single encoder layer.
modelfulltruncateddonor ablatedactivation ablated
Cohere-Transcribe0.300.130.060.04
Canary-Qwen-2.5B0.230.120.050.02
Granite-Speech-4.1-2B0.210.120.200.14
Higgs-Audio-v3-8B0.210.120.03
Phi-4-Multimodal0.190.090.050.20
Parakeet-TDT-0.6B-v20.180.080.060.01
Qwen3-ASR-0.6B0.090.090.04
Moonshine-Streaming0.060.070.04
Voxtral-Mini-3B0.040.050.03
Whisper-Large-v30.020.050.02
Kimi-Audio-7B0.030.060.04
Figure 5: The trigger battery, VoxPopuli. Top row: voice conditions on identical transcripts—(5(a)) reference-disagreement and (5(b)) masked-number accept-ref. Bottom row: the same probes with the trigger removed instead of the voice varied—truncation to the edit, an appended 8 s conversational donor, or the learned register direction projected out of one encoder layer (5(c), 5(d)). Wilson 95% CIs.
Figure 5: The trigger battery, VoxPopuli. Top row: voice conditions on identical transcripts—(5(a)) reference-disagreement and (5(b)) masked-number accept-ref. Bottom row: the same probes with the trigger removed instead of the voice varied—truncation to the edit, an appended 8 s conversational donor, or the learned register direction projected out of one encoder layer (5(c), 5(d)). Wilson 95% CIs.
(b)
(b)
Table 3: Full-probe accept-ref on real audio with each corpus’s own content: VoxPopuli-test vs ep-fresh. Masked columns score the corpus-paired subsets, so the VoxPopuli masked rates differ from the full masked set of §4.
consensus accept-refmasked accept-ref
modelVoxPopuliep-freshVoxPopuliep-fresh
Cohere-Transcribe0.3040.1220.1850.074
Canary-Qwen-2.5B0.2330.1500.0510.062
Granite-Speech-4.1-2B0.2080.1170.0510.062
Higgs-Audio-v3-8B0.2110.1680.0700.040
Phi-4-Multimodal0.1880.1930.0760.044
Parakeet-TDT-0.6B-v20.1810.1200.0380.029
Qwen3-ASR-0.6B0.0920.0910.0510.015
Moonshine-Streaming0.0560.0990.0130.018
Voxtral-Mini-3B0.0350.1280.0760.062
Whisper-Large-v30.0250.0710.0830.062
Kimi-Audio-7B0.0280.1440.0130.018
(c)
(c)
(d)
(d)
Table 4: Audio lift λ⁡(r) (Eq. 1, nats/char) of the silenced number span by voice condition (115 paired sentences passing the intelligibility gate in every clone condition). Diff columns: paired differences, bootstrap 95% CIs; bold marks CIs excluding zero. Whisper’s lift rises on clean TTS, so a drop in lift is likely not a synthesis artifact; Qwen3’s lift is negative in every condition, so its diffs do not indicate recovery. Parakeet-TDT has no teacher-forced readout.
realvox-cloneep-freshgenericreal−ep-freshreal−generic
Cohere-Transcribe+1.52+1.26+0.92+0.54+0.60 [+0.30,+0.94]+0.98 [+0.66,+1.33]
Canary-Qwen-2.5B+1.22+1.05+0.65+0.77+0.57 [+0.27,+0.90]+0.45 [+0.17,+0.71]
Granite-Speech-4.1-2B+0.12+0.15+0.16+0.01−0.04 [−0.24,+0.17]+0.11 [−0.11,+0.32]
Phi-4-Multimodal+0.46+0.51+0.24+0.27+0.23 [+0.01,+0.50]+0.19 [+0.04,+0.35]
Higgs-Audio-v3-8B+0.50+0.57+0.37+0.11+0.13 [−0.00,+0.26]+0.39 [+0.23,+0.56]
Whisper-Large-v3+0.70+0.89+0.85+1.01−0.16 [−0.37,+0.06]−0.31 [−0.47,−0.15]
Moonshine-Streaming+0.31+0.37+0.25+0.15+0.06 [−0.08,+0.21]+0.16 [+0.00,+0.32]
Kimi-Audio-7B+0.24+0.32+0.06+0.16+0.17 [−0.00,+0.37]+0.08 [−0.06,+0.22]
Qwen3-ASR-0.6B−0.60−0.56−1.06−1.14+0.46 [+0.19,+0.75]+0.54 [+0.31,+0.76]
Voxtral-Mini-3B−0.13+0.07−0.12−0.05−0.01 [−0.17,+0.16]−0.08 [−0.23,+0.07]
Figure 6: Switching the benchmark-optimized policy on and off. (6(a)) Input level: On real clips (top left) a conversational donor collapses accept-ref while a VoxPopuli donor leaves it intact; on ep-fresh clones of the same sentences (top right) a VoxPopuli donor re-ignites it while the conversational donor does not. (6(b)) Activation level: projecting out the learned direction on real benchmark edits (bottom left) and adding it on generic-voice clones (bottom right). The remaining consensus-panel members show no effect, like Voxtral-Mini-3B.
Figure 6: Switching the benchmark-optimized policy on and off. (6(a)) Input level: On real clips (top left) a conversational donor collapses accept-ref while a VoxPopuli donor leaves it intact; on ep-fresh clones of the same sentences (top right) a VoxPopuli donor re-ignites it while the conversational donor does not. (6(b)) Activation level: projecting out the learned direction on real benchmark edits (bottom left) and adding it on generic-voice clones (bottom right). The remaining consensus-panel members show no effect, like Voxtral-Mini-3B.
(b)
(b)
Table 5: The opening-courtesy case study: rate at which the audible courtesy is present in the output. truncated: the audio is cut to the opener; attn-isolated keeps the full-clip audio encoding but restricts the decoder’s attention over it to the opener’s frames. translate: the same audio decoded under an English→Spanish translation instruction. full: the entire clip. – marks models without translation or attention-isolation capabilities.
modeltruncatedattn-isolatedtranslatefull
Voxtral-Mini-3B1.001.001.001.00
Whisper-Large-v31.001.001.001.00
Moonshine-Streaming1.001.001.00
Qwen3-ASR-0.6B0.940.951.001.00
Kimi-Audio-7B1.000.890.89
Cohere-Transcribe0.940.260.00
Granite-Speech-4.1-2B0.670.110.390.00
Canary-Qwen-2.5B0.830.050.00
Phi-4-Multimodal0.830.890.610.00
Higgs-Audio-v3-8B0.830.420.06
Parakeet-TDT-0.6B-v21.000.00
Figure 7: Reference-disagreement accept-ref on the real VoxPopuli recordings under content-preserving perturbations (additive noise 10 dB; measured room reverberation, RT60 0.60). Wilson 95% CIs.
Figure 7: Reference-disagreement accept-ref on the real VoxPopuli recordings under content-preserving perturbations (additive noise 10 dB; measured room reverberation, RT60 0.60). Wilson 95% CIs.
Figure 8: Honorific switch rate (Mr/Mister), all 11 models. A rate above 0.5 (dashed) means the model tracks each corpus’s convention at rates above chance.
Figure 8: Honorific switch rate (Mr/Mister), all 11 models. A rate above 0.5 (dashed) means the model tracks each corpus’s convention at rates above chance.

왜 중요한가

공개된 음성인식 성능 순위표를 그대로 믿고 모델을 고르면, 실제 서비스 환경에서는 순위표만큼 성능이 나오지 않을 수 있다는 뜻이다. 벤치마크 점수를 검증하고 새 평가 기준을 만들 때 이런 편법 가능성을 함께 점검해야 한다는 실용적 시사점을 준다.

이 논문의 용어

  • 단어 오류율(WER) · 음성인식 결과와 정답 자막을 비교해 틀린 단어 비율을 계산하는 대표적 성능 지표
  • 벤치마크 편법(benchmark optimization) · 실제 음성을 잘 알아듣는 능력이 아니라 특정 평가 데이터의 특성에만 맞춰 점수를 올리는 행동
  • activation steering(활성화 스티어링) · 모델 내부의 특정 방향(벡터)을 더하거나 빼서 모델의 행동을 인위적으로 바꾸는 해석 기법
  • teacher-forced 우도 · 정답 자막을 강제로 입력한 상태에서 모델이 그 다음 글자를 얼마나 확률 높게 예측하는지 측정하는 방식
  • accept-ref · 오디오가 실제로 뒷받침하지 않는데도 모델이 벤치마크의 정답 자막 표현을 그대로 출력하는 비율

논문 원문 초록 (영문)

Public benchmarks are important measures of Automatic Speech Recognition (ASR) model capabilities. However, by nature of being public, there is risk of models being optimized for these benchmarks in ways that do not generalize well to real-world data. We present a methodology for quantifying benchmark optimization, focusing on cases where the audio underdetermines the reference transcript. We identify three families of behavioral probes that reveal models' capabilities of reproducing benchmark reference spans despite underdetermined audio: reference disagreement, masked-number recovery, and orthographic switching. We find that the highest-scoring open source models output verbatim reference transcript spans even when the relevant audio is contradictory, masked, or ambiguous. Using a variety of mechanistic probes, we show that models respond to narrow acoustic cues to override the faithful representation of the audio in favor of a benchmark-optimized policy. We show the benchmark-optimized behavior can be causally manipulated via low-rank linear steering or simply appending audio to the end of a segment in some cases. Overall, our results indicate that high-performing models exhibit benchmark-conditioned behaviors that can inflate benchmark performance without reflecting improved general-purpose transcription ability.

저자 · Theo Lebryk

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Theo Lebryk et al., arXiv:2608.19936, CC BY 4.0