TransSLR: A Lightweight Transformer for Sign Language Recognition
Automated sign language recognition for underrepresented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this challenge: the CASL-W60 benchmark contains 60 isolated signs, and the best published result is 69.93%. We present TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on normalized 64-frame pose sequences. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves 80.39% top-1 accuracy and 91.07% top-5 accuracy on a signer-independent evaluation split. This result is 10.46 percentage points higher than the published 69.93% result. TransSLR has 8.67 million trainable parameters. We also evaluate zero-shot transfer from a high-resource sign-language model, which achieves 0.00% exact-match accuracy under our manual gloss-matching protocol. Our experiments compare pose-based, RGB-based, and multimodal approaches to sign-language recognition for a low-resource language.