Head-Pose-Aware Visual Speech Recognition with FiLM Modulation
Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements. Still, its performance is fundamentally limited by viseme ambiguity and pose-induced variations that introduce geometric distortions and occlusions. Existing approaches mainly rely on linguistic context or implicit invariance, leaving visual representations insufficiently robust under non-frontal views. In this work, we propose a pose-aware phoneme-level framework, termed HP-VSR-ResFiLM, that explicitly incorporates head-pose information into visual feature extraction. The proposed framework adopts a two-stage pipeline consisting of a pose-conditioned visual encoder in Stage~1 and a pretrained NLLB language model in Stage~2 for phoneme-to-text reconstruction. Specifically, Stage~1 incorporates a pose-conditioned residual Feature-wise Linear Modulation (FiLM) block after the 2D CNN frontend to refine visual representations adaptively using head-pose information. Experiments on LRS2 and LRS3 demonstrate that HP-VSR-ResFiLM achieves competitive performance under comparable training conditions, attaining word error rates (WER) of 24.7% and 30.3%, respectively, without relying on additional training data. Comprehensive ablation studies reveal that different FiLM modulation strategies exhibit complementary strengths: HP-VSR-FiLMFuse (L4) achieves the best overall performance on LRS2, whereas the proposed HP-VSR-ResFiLM provides the greatest improvements under large head-pose variations on LRS3 and consistently outperforms previous pose-aware methods for high-yaw samples. These results demonstrate the effectiveness of explicit pose-conditioned feature modulation for robust visual speech recognition in unconstrained settings.