매일 아침, 어제의 AI를 한 통으로 정리해 보내드립니다메일로 받아보기

METAL LAB

Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack

arXiv:2608.181822026-08-20

Small NLP models, especially BERT-family encoders, remain important in industrial workloads such as classification, ranking, and retrieval even in the era of large language models. On server CPUs, INT8 quantization offers an attractive latency-throughput-cost trade-off, but users increasingly expect such acceleration to be available directly in the native PyTorch stack. We integrate SmoothQuant into TorchAO and optimize the resulting inference path

저자 · Weiwen Xia, Yuxin Cui, E Cao

arXiv에서 원문 보기

최신 논문

논문 전체 보기 →

METAL LAB 최신 기사