arXiv · 2608.13217
UniCon-Former: Unified Convolution Transformer is All You Need for Hand Gesture Recognition
Abstract
Convolutional Neural Networks (CNNs) capture local features efficiently but struggle with global context due to their limited receptive field. On the other hand, transformers effectively capture global dependencies through self-attention but suffer from high redundancy and computational costs. Thus, to leverage the advantages of both CNNs and transformers, we propose a unified model (UniCon-Former) that aims to provide robust and efficient performance on dynamic hand gesture recognition. The unified approach helps the model to learn both local and global features. At the beginning of each transformer stage, the convolution projections help in decreasing the dimension of the input vectors of the transformer block. This creates a pyramidal structure at each transformer stage. These features enable the UniCon-Former to reduce resource usage than vanilla transformers, making it flexible for learning multi-scale and high-resolution features, which is required in hand gesture recognition. We have performed experiments with NVGesture and Briareo datasets and achieved state-of-the-art results with fewer parameters and MACs.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Mallika Garg, Debashis Ghosh, Pyari Mohan Pradhan. 2026-08-13. UniCon-Former: Unified Convolution Transformer is All You Need for Hand Gesture Recognition. https://arxiv.org/abs/2608.13217
Cite the original work for its findings. Save a collection to share your selection of sources.