TY - RPRT TI - On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration AU - Maoyang Xiang AU - Ramesh Fernando AU - Bo Wang PY - 2025 UR - https://arxiv.org/abs/2504.17376 ID - 2504.17376 ER -