Searcharxiv⌕ Search

arXiv subjects

Victor Tolulope Olufemi

Publications and source records attributed to Victor Tolulope Olufemi.

3 recordsLinked to original sources

TransSLR: A Lightweight Transformer for Sign Language Recognition

Automated sign language recognition for underrepresented languages remains a largely unsolved problem. Central African Sign Language (CASL) exemplifies this challenge: the CASL-W60 benchmark contains 60 isolated signs, and the best published result is 69.93%. We present TransSLR, a lightweight Temporal Transformer Encoder trained from scratch on normalized 64-frame pose sequences. By operating on geometric keypoint representations rather than raw RGB, TransSLR achieves 80.39% top-1 accuracy and 91.07% top-5 accuracy on a signer-independent evaluation split. This result is 10.46 percentage points higher than the published 69.93% result. TransSLR has 8.67 million trainable parameters. We also evaluate zero-shot transfer from a high-resource sign-language model, which achieves 0.00% exact-match accuracy under our manual gloss-matching protocol. Our experiments compare pose-based, RGB-based, and multimodal approaches to sign-language recognition for a low-resource language.

cs.CV↗

STAM-ASR: Speaker-Temporal Anchoring with Memory for Multi-Speaker ASR

Natural conversations make both speech recognition and speaker attribution challenging for ASR, as speakers take turns, overlap, and reappear over time. We propose STAM-ASR, Speaker-Temporal Anchoring with Memory, a lightweight framework that extends an already pretrained AudioLLM for multi-speaker ASR. Without relying on an external diarization system, STAM-ASR learns speaker activity and speaker-aware representations directly from intermediate AudioLLM features. Hence providing explicit who and when cues to modulate the AudioLLM's semantic representation without explicit speech separation. STAM-ASR further maintains fixed-size speaker and conversational memories to carry complementary context across turns. We evaluate STAM-ASR on AMI, ICSI, LibriCSS, and NOTSOFAR-1 across close-talk, far-field, overlapping, and cross-domain conditions. Our reported results shows that speaker-temporal conditioning and memory provide complementary benefits, while the gap between reference and predicted speaker activity identifies robust speaker tracking as a key remaining challenge.

cs.SD↗

WAXAL-NET: Finetuned Edge ASR Across 19 African Languages

We evaluate whether compact domain-specialized ASR models can outperform massively multilingual foundation models for conversational African speech across 19 languages in the WAXAL corpus. Fine-tuned edge models achieve a macro-averaged WER of $38.0\%$ compared to $64.9\%$ for the best zero-shot baseline, a $26.9$ percentage-point reduction using models $3-40\times$ smaller. Results confirm that domain specialization dominates scale for spontaneous African speech. Cross-domain evaluation shows that fine-tuned models recover usable performance on out-of-distribution (OOD) speech, while zero-shot models regain an advantage when the test domain matches their pretraining distribution. A distributed native-speaker audit across all surveyed languages produces a linguistically-grounded error taxonomy, showing that CTC and autoregressive architectures behave differently across language families. We further show that WER alone misrepresents performance for syllabary-script languages where CER/WER ratios reveal substantially higher character-level accuracy than headline WER suggests. Finally, to contribute to future African ASR research, we release all model weights, fine-tuning and evaluation scripts, and a cleaned WAXAL subset covering all $19$ languages.

cs.CL↗