arXiv · 2410.23749
LATST: Are Transformers Necessarily Complex for Time-Series Forecasting
Abstract
Transformer-based architectures have achieved remarkable success in natural language processing and computer vision. However, their performance in multivariate long-term forecasting often falls short compared to simpler linear baselines. Previous research has identified the traditional attention mechanism as a key factor limiting their effectiveness in this domain. To bridge this gap, we introduce LATST, a novel approach designed to mitigate entropy collapse and training instability common challenges in Transformer-based time series forecasting. We rigorously evaluate LATST across multiple real-world multivariate time series datasets, demonstrating its ability to outperform existing state-of-the-art Transformer models. Notably, LATST manages to achieve competitive performance with fewer parameters than some linear models on certain datasets, highlighting its efficiency and effectiveness.
Explore related subjects
Keep this discovery
Dizhen Liang. 2024-10-31. LATST: Are Transformers Necessarily Complex for Time-Series Forecasting. https://arxiv.org/abs/2410.23749
Cite the original work for its findings. Save a collection to share your selection of sources.