TY - RPRT TI - Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native Stack AU - Weiwen Xia AU - Yuxin Cui AU - E Cao PY - 2026 UR - https://arxiv.org/abs/2608.18182 ID - 2608.18182 ER -