每天早上一封邮件,把昨天的 AI 梳理好订阅邮件

METAL LAB

Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack

arXiv:2608.181822026-08-20

Small NLP models, especially BERT-family encoders, remain important in industrial workloads such as classification, ranking, and retrieval even in the era of large language models. On server CPUs, INT8 quantization offers an attractive latency-throughput-cost trade-off, but users increasingly expect such acceleration to be available directly in the native PyTorch stack. We integrate SmoothQuant into TorchAO and optimize the resulting inference path

作者 · Weiwen Xia, Yuxin Cui, E Cao

在 arXiv 阅读

最新论文

全部论文 →

METAL LAB 最新报道