매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

TESTNAV: Pareto-Guided Search for Compositional Robustness Testing

arXiv:2608.198822026-08-21

AI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법

밝기 변화와 흔들림처럼 여러 왜곡이 동시에 섞인 입력에서 AI 모델이 얼마나 잘 버티는지 확인하려면 조합 경우의 수가 폭발적으로 늘어난다. TESTNAV는 '성능이 많이 떨어지면서도 원본과 비슷하게 보이는' 조합을 두 목표 사이의 균형점(파레토 프론트)으로 정의하고, 진화 알고리즘 NSGA-II로 이를 효율적으로 찾아낸다. 이미지, 문장, 코드 네 개 벤치마크에서 전체 경우의 수의 35.8~89.3%만 확인하고도 기존 탐색법보다 최대 2.15배 빠르게 같은 결과를 얻었다.

무엇을 했나

  1. 밝기, 블러, 노이즈 등 여러 손상을 동시에 조합해 테스트하면 실제 상황과 비슷하지만, 경우의 수가 4개 차원×6단계면 1,296가지로 폭발적으로 늘어나고 그중 상당수는 비현실적으로 심하게 망가진 입력이라 진단 가치가 낮다는 문제를 다뤘다
  2. TESTNAV는 '모델 성능이 얼마나 떨어지는가(성능 저하)'와 '왜곡된 입력이 원본과 얼마나 비슷한가(입력 충실도)'라는 두 목표를 동시에 만족시키는 조합만 골라내는 다목적 최적화 문제로 문제를 재정의했다
  3. 이미지에는 SSIM·KID, 문장·코드에는 chrF·BERT-F1 같은 분야별 유사도 지표로 충실도를 재고, 두 목표를 동시에 최적화하는 진화 알고리즘 NSGA-II로 전체 경우를 다 안 보고도 최적 조합 집합(파레토 프론트)을 근사했다
  4. Tiny-ImageNet, QQP, HumanEval, MBPP 네 개 벤치마크에서 전체 1,296개 조합을 모두 평가해 정답을 만든 뒤 비교한 결과, TESTNAV는 기존 탐색법 대비 최대 2.15배 빠르게, 전체 경우의 수의 35.8~89.3%만 평가하고도 동일한 결과에 도달했다
  5. 성능 저하만 보거나 충실도만 보는 단일 목표 탐색, 또는 신경망 내부 활성화 패턴을 보는 기존 커버리지 지표들은 모두 TESTNAV보다 결과가 떨어져, 두 목표를 함께 봐야 한다는 점을 확인했다
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 0
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 1
Table 1: Benchmarks, datasets and models, perturbations, fidelity metrics ϕ, task-performance metrics ψ, and clean-set performance ψ⁡(𝒟).
Dataset 𝒟SizeModelPerturbationsFidelity ϕTask ψClean ψ⁡(𝒟)
Tiny-ImageNet10,000CaiT-S36speckle noise, glass blur, brightness, pixelateKID, SSIMAccuracy86.7%
QQP1,000RoBERTa-basesynonym, typo, contraction, punctuationBERT-F1, chrFAccuracy91.2%
HumanEval164CodeGen-2B-monobutterfingers, char case, whitespace, newlineBERT-F1, chrFRP5@123.2%
MBPP974CodeGen-2B-monobutterfingers, char case, whitespace, newlineBERT-F1, chrFRP5@131.9%
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 2
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 3
Table 2: AUC-Recall@𝒫∗ for TestNav and search baselines across all benchmarks and fidelity metrics. Bold indicates the best score and underline indicates the second-best score per column.
MethodSSIMKIDchrFBERT-F1chrFBERT-F1chrFBERT-F1
TestNav0.7040.6860.6460.6510.7290.7300.6970.690
Greedy Search0.5570.4640.6380.7820.7420.7830.6810.751
Genetic Algorithm0.5400.5520.6930.7070.7530.7670.6400.631
Random Search0.5100.5030.5260.5140.5020.5220.5140.490
Figure 4: Input-level metrics versus TestNav on . Curves show Recall@𝒫∗ over unique configurations (top) and evaluation budget (bottom), using SSIM (left) and KID-derived fidelity (right).
Figure 4: Input-level metrics versus TestNav on . Curves show Recall@𝒫∗ over unique configurations (top) and evaluation budget (bottom), using SSIM (left) and KID-derived fidelity (right).
Figure 7: 4D voxel visualisations of 𝒫∗ for non-vision benchmarks using chrF as the fidelity metric.
Figure 7: 4D voxel visualisations of 𝒫∗ for non-vision benchmarks using chrF as the fidelity metric.
Table 3: Pareto-front composition per metric pair for .
Metric pair𝒫∗Θ1Θ2Θ3Θ4
(δ,ρSSIM)2791242
(δ,ρKID)2747142
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 6
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 7

왜 중요한가

AI 모델을 실전에 배포하기 전 모든 왜곡 조합을 다 테스트하는 것은 계산 비용이 너무 커서 현실적으로 불가능한데, TESTNAV는 한정된 시험 횟수 안에서 자율주행이나 의료 영상처럼 여러 손상이 동시에 나타나는 환경에서 실제로 위험한 실패 사례를 효율적으로 찾아내는 방법을 제시한다. 이는 모델 검증팀이 제한된 자원으로 더 신뢰할 수 있는 견고성 테스트를 설계하는 데 실질적인 지침이 된다.

Figure 8: 4D voxel visualisations of 𝒫∗ for non-vision benchmarks using BERT-F1 as the fidelity metric.
Figure 8: 4D voxel visualisations of 𝒫∗ for non-vision benchmarks using BERT-F1 as the fidelity metric.
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 9

이 논문의 용어

  • 파레토 프론트(Pareto front) · 한쪽을 개선하면 다른 쪽이 나빠질 수밖에 없는, 두 목표 사이의 최적 균형점들의 집합
  • NSGA-II · 여러 목표를 동시에 최적화하기 위해 사용하는 진화 알고리즘의 한 종류
  • SSIM/KID · 이미지가 원본과 얼마나 비슷한지 구조·통계적으로 비교하는 지표
  • chrF/BERT-F1 · 문장이나 코드가 원본 텍스트와 얼마나 비슷한지 재는 지표
  • Recall@P* · 탐색 과정에서 실제 최적 조합 집합 중 얼마나 많이 찾아냈는지 나타내는 비율
TESTNAV: Pareto-Guided Search for Compositional Robustness Testing figure 10

논문 원문 초록 (영문)

Deep learning models remain vulnerable to real-world input perturbations, especially when multiple corruptions co-occur in the same input (e.g., brightness shifts and motion blur). Compositional testing reveals these interaction effects but introduces two challenges: combinatorial growth of the perturbation space as dimensions and severity levels increase, and uneven diagnostic value-many combinations yield unrealistically degraded inputs with limited practical relevance. We present TESTNAV, 1 a Pareto-guided robustness testing framework for efficiently exploring discrete, compositional perturbation spaces when only a limited number of perturbation configurations can be evaluated. TESTNAV prioritises severe yet realistic failures by formulating robustness testing as bi-objective optimisation: maximise performance degradation while preserving input fidelity measured by modality-specific metrics (e.g., SSIM and KID for vision; chrF and BERT-F1 for language and code). It uses NSGA-II to approximate the bi-objective Pareto front. Across four benchmarks spanning vision, natural language, and code generation, TESTNAV recovers Pareto fronts up to 2.15x faster than search-based baselines, using 35.8%-89.3% of the discrete perturbation space defined by four perturbation dimensions with six levels each.

저자 · Arooj Arif, Tobias Hartung, Elena Botoeva, Alexandros Koliousis

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사

그림 출처: Arooj Arif et al., arXiv:2608.19882, CC BY 4.0