The Asymmetric Harms of LLM Compression
AI 모델을 가볍게 압축하면 평균 점수는 그대로여도, 흔한 지식이 오히려 더 많이 사라지고 편향은 특정 집단에서 몰래 커진다
연구팀은 Llama, Qwen, Gemma 세 개의 언어모델에 11가지 압축 기법(양자화·가지치기)을 적용해 지식 유지력, 자신감, 사회적 편향을 정밀 분석했다. 그 결과 압축은 흔히 알려진 지식(헤드)을 상대적으로 더 많이 잃게 만들고, 모델은 틀린 답을 내면서도 여전히 자신감을 유지하며, 전체 편향 점수는 안정적이어도 성별·인종 등 하위 집단에서는 정반대 방향의 큰 변화가 숨어 있었다. 즉 평균 정확도나 복잡도(perplexity) 같은 기존 지표만으로는 압축된 모델의 실제 위험을 파악할 수 없다는 것이다.
무엇을 했나
- Llama-3.1-8B-Instruct, Qwen-3-8B, Gemma-2-9B-it 세 모델에 양자화(GPTQ, AWQ, OmniQuant, AQLM)와 가지치기(WANDA, SparseGPT, ShortGPT 등) 11가지 기법을 적용해 비교했다
- PopQA와 Head-to-Tail 데이터셋으로 지식을 유명도(헤드·미들·테일)에 따라 나누어 정확도 유지율을 측정했더니, 절대 정확도는 헤드가 여전히 높지만 기준 모델 대비 상대적 유지율은 오히려 테일 지식이 더 잘 보존되고 헤드 지식이 더 많이 깎였다
- 압축으로 인해 새로 틀리게 된 답변에서도 모델은 중간~높은 확신(0.4~0.6 수준)을 유지하는 경우가 많았고, 이 확신은 정확도가 이미 무너진 뒤에야 급락하는 경우가 대부분이었다
- WinoBias, BBQ 편향 벤치마크에서 전체 편향 점수는 거의 변화가 없어도 남성·여성 등 하위 집단별로는 최대 수십 퍼센트포인트(예: 특정 직업군에서 -53.1퍼센트포인트)에 달하는 상반된 변화가 발생했다
- 경미하거나 중간 수준의 압축에서도 이런 숨겨진 불균형이 나타났기 때문에, 평균 지표만 보고 압축 모델을 안전하다고 판단해서는 안 되며 세부 집단별 검증이 필요하다는 결론을 내렸다

| Parameter | GPTQ | AWQ | OmniQuant | AQLM |
|---|---|---|---|---|
| Bit width | 2, 3, 4 | 2, 3, 4 | 2, 3, 4 | 2, 3, 4 |
| Calibration dataset | C4 | C4 | C4 | C4 |
| Calibration samples | 128 | 128 | 128 | 128 |
| Calibration seed | 42 | 42 | 42 | 42 |
| Calibration sequence length | 512 | 512 | 512 | 512 |
| Group size | 128 | 128 | 128 | – |
| Calibration split | validation | validation | – | – |
| Calibration source records | 4096 | – | – | – |
| Calibration batch size | 1 | 1 | – | – |
| Symmetric quantization | True | False | – | – |
| Activation ordering | True | – | – | – |
| Sequential quantization | True | – | – | – |
| Target modules | – | Linear | – | – |
| Ignored modules | – | lm_head | – | – |
| Activation bit width | – | – | 16 | – |
| Optimization epochs | – | – | 40 / 20 / 20 | – |
| Learnable weight clipping | – | – | True | – |
| Learnable equivalent transformation | – | – | False | – |
| Input group size | – | – | – | 8 |
| Output group size | – | – | – | 1 |
| Relative MSE tolerance | – | – | – | 0.01 |
| Maximum fine-tuning epochs | – | – | – | 10 |
| Activation offloading | – | – | – | True |
| Resume enabled | – | – | – | True |
| Parameter | Magnitude | WANDA | SparseGPT | ShortGPT | Layer Dropping |
|---|---|---|---|---|---|
| Compression granularity | weights | weights | weights | decoder blocks | decoder blocks |
| Compression levels | 30/50/70% | 30/50/70% | 30/50/70% | 5/10/15/20/25% | 5/10/15/20/25% |
| Semi-structured patterns | – | 4:8, 2:4 | 4:8, 2:4 | – | – |
| Calibration dataset | – | C4 | C4 | C4 | C4 |
| Calibration samples | – | 32 | 32 | 32 | 32 |
| Calibration sequence length | – | 512 | 512 | 512 | 256 |
| Calibration split | – | validation | validation | validation | validation |
| Target modules | attn./MLP linear | attn./MLP linear | attn./MLP linear | decoder blocks | decoder blocks |
| Excluded components | lm_head | lm_head | lm_head | first/last block | first/last block |
| Selection criterion | weight magnitude | weight–activation product | Hessian-based reconstruction | block influence | importance + position |
| Activation-importance weight | – | – | – | – | 0.7 |
| Position-prior weight | – | – | – | – | 0.35 |
| Random seed | – | – | – | – | 13 |
| Compression | Method | Setting | Perplexity (PPL) | ||
|---|---|---|---|---|---|
| Llama-3.1- 8B-Instruct | Qwen-3- 8B | Gemma-2- 9B-it | |||
| None | Full precision | FP16 | 7.1253 | 9.5888 | 10.2117 |
| Quantization | GPTQ | 2-bit | 1612.6539 | 141.3107 | 137.1019 |
| 3-bit | 10.0354 | 11.2652 | 12.3045 | ||
| 4-bit | 8.4459 | 9.9511 | 10.4903 | ||
| AWQ | 2-bit | 94354.1597 | 16508.5254 | 9407.2547 | |
| 3-bit | 9.4931 | 11.3750 | 11.8154 | ||
| 4-bit | 7.5192 | 9.9949 | 10.6473 | ||
| OmniQuant | 2-bit | 671.2557 | 39.6147 | 35.2010 | |
| 3-bit | 9.4949 | 11.7079 | 12.2236 | ||
| 4-bit | 7.5728 | 10.0737 | 10.5624 | ||
| AQLM | 2-bit | 11.5659 | 13.0687 | 14.1902 | |
| 3-bit | 11.3204 | 12.3484 | 12.2632 | ||
| 4-bit | 8.4457 | 10.2807 | 10.8105 | ||
| Unstructured pruning | Magnitude | 30% sparsity | 14.5756 | 11.5522 | 16.6274 |
| 50% sparsity | 177.7208 | 28.9189 | 65.7596 | ||
| 70% sparsity | 127104.4841 | 138372.7513 | 117527.1945 | ||
| Wanda | 30% sparsity | 9.2901 | 11.0325 | 12.9266 | |
| 50% sparsity | 12.7384 | 12.7984 | 17.2316 | ||
| 70% sparsity | 236.8802 | 153.0605 | 106.8877 | ||
| SparseGPT | 30% sparsity | 9.5870 | 11.0395 | 13.9585 | |
| 50% sparsity | 15.4637 | 13.9930 | 20.1264 | ||
| 70% sparsity | 272.6379 | 862.9209 | 129.0303 | ||
| Semi-structured pruning | Wanda (N:M) | 4:8 | 18.0268 | 14.7886 | 19.6716 |
| 2:4 | 30.6551 | 18.4150 | 24.4520 | ||
| SparseGPT (N:M) | 4:8 | 23.8732 | 16.9133 | 22.8913 | |
| 2:4 | 43.4478 | 21.1767 | 33.7740 | ||
| Structured pruning | ShortGPT | 5% blocks removed | 9.8872 | 15.1473 | 13.5355 |
| 10% blocks removed | 11.1683 | 35.1417 | 14.7864 | ||
| 15% blocks removed | 19.4141 | 43.5930 | 21.1945 |
| Rank | Gender | Occupation | Compression Configuration | Base 𝑩𝒈 (%) | Comp. 𝑩𝒈 (%) | 𝚫𝑩𝒈 (pp) | 𝚫𝑩 (pp) |
|---|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | |||||||
| 1 | Female | Secretary | SparseGPT N:M (2:4) | 71.9 | 18.8 | −53.1 | −2.2 |
| 3 | Female | Auditor | SparseGPT N:M (4:8) | 82.1 | 32.1 | −50.0 | +2.3 |
| 8 | Male | Analyst | SparseGPT N:M (2:4) | 50.0 | 2.5 | −47.5 | −2.2 |
| 9 | Female | Clerk | SparseGPT N:M (2:4) | 78.6 | 32.1 | −46.4 | −2.2 |
| Qwen3-8B | |||||||
| 2 | Male | Sheriff | Magnitude (50%) | 59.6 | 7.7 | −51.9 | −8.2 |
| 4 | Female | Clerk | Magnitude (50%) | 67.9 | 17.9 | −50.0 | −8.2 |
| 5 | Female | Nurse | ShortGPT (20%) | 77.8 | 27.8 | −50.0 | −4.2 |
| 6 | Male | Developer | SparseGPT N:M (2:4) | 71.9 | 21.9 | −50.0 | −4.0 |
| 7 | Female | Attendant | Magnitude (50%) | 61.8 | 12.7 | −49.1 | −8.2 |
| 10 | Female | Counselor | Magnitude (50%) | 66.1 | 19.6 | −46.4 | −8.2 |
| Rank | Category | Subgroup | Compression Configuration | Base 𝑩𝒈 (%) | Comp. 𝑩𝒈 (%) | 𝚫𝑩𝒈 (pp) | 𝚫𝑩 (pp) |
|---|---|---|---|---|---|---|---|
| Llama-3.1-8B-Instruct | |||||||
| 8 | Disability | People with cognitive disabilities or mental illness | SparseGPT N:M (4:8) | 40.6 | 59.4 | +18.8 | −1.7 |
| 9 | Disability | People with cognitive disabilities or mental illness | WANDA N:M (2:4) | 40.6 | 59.4 | +18.8 | −3.1 |
| Qwen3-8B | |||||||
| 1 | Disability | Down’s syndrome | ShortGPT (10%) | 37.5 | 75.0 | +37.5 | −1.2 |
| 2 | Disability | Down’s syndrome | ShortGPT (15%) | 37.5 | 75.0 | +37.5 | −1.4 |
| 3 | Disability | Down’s syndrome | ShortGPT (5%) | 37.5 | 62.5 | +25.0 | −1.1 |
| 4 | Disability | People with cerebral palsy | SparseGPT (50%) | 56.3 | 81.3 | +25.0 | −1.0 |
| 10 | Disability | Down’s syndrome | Magnitude (30%) | 37.5 | 56.3 | +18.8 | −1.1 |
| Gemma-2-9B-it | |||||||
| 5 | Disability | Down’s syndrome | ShortGPT (15%) | 50.0 | 75.0 | +25.0 | −0.5 |
| 6 | Disability | Down’s syndrome | ShortGPT (25%) | 50.0 | 75.0 | +25.0 | −1.0 |
| 7 | Nationality | Italian | WANDA N:M (2:4) | 57.5 | 37.5 | −20.0 | +0.1 |
왜 중요한가
많은 기업과 서비스가 비용 절감을 위해 모델을 가볍게 만드는 압축 기법을 쓰는데, 이 연구는 평균 성능이 멀쩡해 보여도 실제로는 상식적으로 흔한 지식이 더 취약해지고, 모델이 틀린 답에도 자신 있게 말하며, 특정 성별·인종·질환 집단에 대한 편향이 몰래 커질 수 있음을 보여준다. 따라서 압축 모델을 배포하기 전에는 전체 점수뿐 아니라 하위 집단별 세부 평가가 반드시 필요하다는 실무적 경고를 준다.
이 논문의 용어
- 양자화(Quantization) · 모델의 가중치나 계산값을 더 낮은 정밀도(예: 32비트→4비트)로 저장해 용량과 연산량을 줄이는 압축 기법
- 가지치기(Pruning) · 모델에서 출력에 영향을 적게 주는 가중치, 뉴런, 층 등을 제거해 크기를 줄이는 압축 기법
- 혼란도(Perplexity) · 언어모델이 다음 단어를 얼마나 잘 예측하는지를 나타내는 대표적 평균 성능 지표, 낮을수록 좋음
- 상대적 유지율 변화(Relative Retention Shift) · 압축 후 특정 집단의 정확도 유지 정도가 전체 평균 대비 얼마나 더 좋거나 나쁜지를 퍼센트포인트로 나타낸 지표
- 기대 보정 오차(ECE) · 모델이 표현한 확신 수준이 실제 정답률과 얼마나 잘 맞는지를 측정하는 지표, 낮을수록 확신과 정확도가 잘 일치함을 의미
논문 원문 초록 (영문)
Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11 compression methods to investigate the effects of compression on knowledge retention, model confidence, and social bias. We find that compression disproportionately reduces the relative retention of head knowledge compared to tail knowledge. Furthermore, compressed models often remain substantially confident in their incorrect answers on newly lost knowledge. Finally, we demonstrate that stable aggregate bias scores can conceal substantial, opposing shifts in stereotypical preferences across demographic subgroups. Together, these findings reveal asymmetric behavioral changes that aggregate performance measures fail to capture, highlighting the need for granular evaluation of compressed models before deployment.
arXiv에서 원문 보기최신 논문
- Specification-delta-driven data governance: an empirical study of the {\guillemotleft}spec-delta{\guillemotright} as the unit of change in lakehouse data platforms데이터 플랫폼 변경도 코드처럼 '설계도 조각'을 붙여서 검토하면 어떨까: 실험 설계 논문
- Are LLMs becoming similarly creative? Evidence from three years of models최신 AI 챗봇일수록 서로 비슷한 답을 내놓는다는 3년치 조사 결과
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI 생성 텍스트를 잡아내는 워터마크 기술이 영어 아닌 언어에서는 훨씬 부실하게 작동하고, 그 격차는 개별 언어가 아니라 언어 계열 단위로 나타난다
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingAI 모델을 여러 손상이 겹친 입력으로 시험할 때, 굳이 다 테스트하지 않고도 '진짜 위험한 실패'만 골라내는 탐색법
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesAI 에이전트에게 어떤 '스킬 문서'를 몇 개나 줘야 잘 작동하는지, 수학적으로 최적해를 보장하며 골라주는 방법
- Reliable Financial Named Entity Recognition under Domain Shift금융 AI가 서류체 문장에서 배운 자신감은 트위터로 가면 거짓말이 된다
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving논문 속 시연이 아니라 실제 서비스에 넣을 수 있는 희소 어텐션 만들기
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy Correction텍스트가 빠지거나 망가져도, AI가 그 자리를 대신할 '가짜 텍스트'를 한 번에 만들지 않고 여러 번 고쳐가며 감정을 더 정확히 읽어낸다
METAL LAB 최신 기사
그림 출처: Yuan Wu et al., arXiv:2608.19670, CC BY 4.0