One email each morning — yesterday's AI, sortedGet it in your inbox

METAL LAB

Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack

arXiv:2608.181822026-08-20

Small NLP models, especially BERT-family encoders, remain important in industrial workloads such as classification, ranking, and retrieval even in the era of large language models. On server CPUs, INT8 quantization offers an attractive latency-throughput-cost trade-off, but users increasingly expect such acceleration to be available directly in the native PyTorch stack. We integrate SmoothQuant into TorchAO and optimize the resulting inference path

Authors · Weiwen Xia, Yuxin Cui, E Cao

Read on arXiv

Latest papers

All papers →

Latest from METAL LAB