arXiv · 2505.18413
LatentLLM: Attention-Aware Joint Tensor Compression
Abstract
Modern foundation models such as large language models (LLMs) and large multi-modal models (LMMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs/LMMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attention-aware joint tensor de-composition. Our framework can significantly improve the model accuracy over the existing model compression methods when reducing the latent dimension to realize computationally/memory-efficient LLMs/LLMs. We show the benefit on several benchmark including multi-modal reasoning tasks.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Toshiaki Koike-Akino, Xiangyu Chen, Jing Liu, Ye Wang, Pu, Wang, Matthew Brand. 2025-05-23. LatentLLM: Attention-Aware Joint Tensor Compression. https://arxiv.org/abs/2505.18413
Cite the original work for its findings. Save a collection to share your selection of sources.