SearcharxivSearch

arXiv subjects

Huiming Han

Publications and source records attributed to Huiming Han.

3 recordsLinked to original sources

From Quarter to All: Accelerating Speculative LLM Decoding via Floating-Point Exponent Remapping and Parameter Sharing

Large language models achieve impressive performance across diverse tasks but exhibit high inference latency due to their large parameter sizes. While quantization reduces model size, it often leads to performance degradation compared to the full model. Speculative decoding remains lossless but typically incurs extra overheads. We propose SPEQ, an algorithm-hardware co-designed speculative decoding method that uses part of the full-model weight bits to form a quantized draft model, thereby eliminating additional training or storage overhead. A reconfigurable processing element array enables efficient execution of both the draft and verification passes. Experimental results across 15 LLMs and tasks demonstrate that SPEQ achieves speedups of 2.07x, 1.53x, and 1.45x compared over FP16, Olive, and Tender, respectively.

cs.AR

MoBiLE: Efficient Mixture-of-Experts Inference on Consumer GPU with Mixture of Big Little Experts

Mixture-of-Experts (MoE) models have recently demonstrated exceptional performance across a diverse range of applications. The principle of sparse activation in MoE models facilitates an offloading strategy, wherein active experts are maintained in GPU HBM, while inactive experts are stored in CPU DRAM. The efficacy of this approach, however, is fundamentally constrained by the limited bandwidth of the CPU-GPU interconnect. To mitigate this bottleneck, existing approaches have employed prefetching to accelerate MoE inference. These methods attempt to predict and prefetch the required experts using specially trained modules. Nevertheless, such techniques are often encumbered by significant training overhead and have shown diminished effectiveness on recent MoE models with fine-grained expert segmentation. In this paper, we propose MoBiLE, a plug-and-play offloading-based MoE inference framework with \textit{mixture of big-little experts}. It reduces the number of experts for unimportant tokens to half for acceleration while maintaining full experts for important tokens to guarantee model quality. Further, a dedicated fallback and prefetching mechanism is designed for switching between little and big experts to improve memory efficiency. We evaluate MoBiLE on four typical modern MoE architectures and challenging generative tasks. Our results show that MoBiLE achieves a speedup of 1.60x to 1.72x compared to the baseline on a consumer GPU system, with negligible degradation in accuracy.

cs.CL

High-Entropy Enhanced Negative Thermal Expansion Perfomance in Antiperovkites

The negative thermal expansion (NTE) materials, which can act as thermal-expansion compensators to counteract the positive thermal expansion, have great applications merit in precision engineering. However, the exploration of NTE behavior with a wide temperature range has reached its upper ceiling through traditional doping strategies due to composition limitations. The unique sluggish characteristic in phase transition and extended optimization space in recent high entropy systems has great potential to broaden the temperature range in electronic transitions-induced NTE materials. Mn-based anti-perovskites offer an ideal platform for the exploration of high entropy NTE material due to their abundant element selection and controllable NTE performance. In this paper, the high entropy strategy is first introduced to broaden the NTE temperature range by relaxing the abrupt phase transition in Mn-based anti-perovskite nitride. We propose an empirical screening method to synthesize the high-entropy anti-perovskite (HEAP). it is found that magnetic phase separation from anti-ferromagnetic CII to paramagnetic CI surviving in an ultra-wide temperature range of 5K<=T<=350K (Delta_T=345K), revealing a unique sluggish characteristic. Consequently, a remarkable NTE behavior (up to Delta_T=235K, 5K<=T<=240K) with a coefficient of thermal expansion of -4.7x10-6/K, has been obtained in HEAP. It is worth noting that the temperature range is two/three times wider than that of low-entropy systems. The sluggish characteristic has been further experimentally proved to come from disturbed phase transition dynamics due to distortion in atomic spacing and chemical environmental fluctuation observed by the spherical aberration-corrected electron microscope. Our demonstration provides a unique paradigm for broadening the temperature range of NTE materials induced by phase transition through entropy engineering.

cond-mat.mtrl-sci