Boosting Point-supervised Temporal Action Localization via Text Refinement and Alignment
Recently, point-supervised temporal action localization has gained significant attention for its effective balance between labeling costs and localization accuracy. However, current methods primarily rely on visual features and do not fully exploit the complementary semantic information contained in textual descriptions. To address this issue, we propose a Text Refinement and Alignment (TRA) framework that incorporates refined textual semantics to complement visual representations for point-supervised temporal action localization. This is achieved by designing two new modules for the original point-supervised framework: a Point-based Text Refinement module (PTR) and a Point-based Multimodal Alignment module (PMA). Specifically, we first generate descriptions for video frames using a pre-trained multimodal model. Next, PTR refines the initial descriptions by leveraging point annotations together with multiple fixed pre-trained models. PMA then projects the visual and textual features into a unified semantic space and employs a point-based multimodal contrastive learning objective to reduce the gap between visual and linguistic modalities. Finally, the aligned multimodal features are fed into the action detector for temporal action localization. Extensive experiments on five widely used benchmarks demonstrate that TRA consistently improves strong point-supervised baselines and achieves competitive performance compared with state-of-the-art methods. In particular, on THUMOS'14, TRA achieves 58.5% AVG mAP@[0.1:0.7], outperforming the visual-only baseline by 2.3%.