arXiv · 2410.13060
AERO: Entropy-Guided Framework for Private LLM Inference
Abstract
Privacy-preserving computation enables language model inference directly on encrypted data yet suffers from prohibitive latency and communication overheads, primarily due to nonlinear functions. Removing nonlinearities, however, can trigger one of two failure modes restricting the potential for nonlinearity removal: entropy collapse in deeper layers, which destabilizes training, and entropic overload in early layers, causing under-utilization of attention heads. To address these challenges, we introduce AERO, an entropy-guided framework to strategically eliminates costly nonlinear operations from transformer architectures, which employs an adaptive recalibration through a head-wise entropy regularizer with learnable per-head strengths, enabling each head to adjust its entropy level while penalizing extreme entropies and fostering functional diversity through a tolerance margin. Experiments show AERO can save 3.4$\times$ communication and 1.4$\times$ latency, without any performance penalty.
Explore related subjects
Keep this discovery
Nandan Kumar Jha, Brandon Reagen. 2024-10-16. AERO: Entropy-Guided Framework for Private LLM Inference. https://arxiv.org/abs/2410.13060
Cite the original work for its findings. Save a collection to share your selection of sources.