arXiv · 2409.04118
Convolutional Transformer-Based Image Compression
Abstract
In this paper, we present a novel transformer-based architecture for end-to-end image compression. Our architecture incorporates blocks that effectively capture local dependencies between tokens, eliminating the need for positional encoding by integrating convolutional operations within the multi-head attention mechanism. We demonstrate through experiments that our proposed framework surpasses state-of-the-art CNN-based architectures in terms of the trade-off between bit-rate and distortion and achieves comparable results to transformer-based methods while maintaining lower computational complexity.
Explore related subjects
Keep this discovery
Bouzid Arezki, Fangchen Feng, Anissa Mokraoui. 2024-09-06. Convolutional Transformer-Based Image Compression. https://doi.org/10.23919/spa59660.2023.10274433
Cite the original work for its findings. Save a collection to share your selection of sources.