arXiv · 2409.14199
Loop Neural Networks for Parameter Sharing
Abstract
The success of large-scale language models like GPT can be attributed to their ability to efficiently predict the next token in a sequence. However, these models rely on constant computational effort regardless of the complexity of the token they are predicting, lacking the capacity for iterative refinement. In this paper, we introduce a novel Loop Neural Network, which achieves better performance by utilizing longer computational time without increasing the model size. Our approach revisits the input multiple times, refining the prediction by iteratively looping over a subset of the model with residual connections. We demonstrate the effectiveness of this method through experiments comparing versions of GPT-2 with our loop models, showing improved performance in language modeling tasks while maintaining similar parameter counts. Importantly, these improvements are achieved without the need for extra training data.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Kei-Sing Ng, Qingchen Wang. 2024-09-21. Loop Neural Networks for Parameter Sharing. https://arxiv.org/abs/2409.14199
Cite the original work for its findings. Save a collection to share your selection of sources.