arXiv · 2512.11016
SoccerMaster: A Vision Foundation Model for Soccer Understanding
Abstract
Soccer understanding has recently garnered growing research interest due to its domain-specific complexity and unique challenges. Unlike prior works that typically rely on isolated, task-specific expert models, this work aims to propose a unified model to handle diverse soccer visual understanding tasks, ranging from fine-grained perception (e.g., athlete detection and identification) to high-level semantic reasoning (e.g., event classification). Concretely, our contributions are threefold: (i) we present SoccerMaster, the first soccer-specific vision foundation model that unifies diverse tasks within a single framework via supervised multi-task pretraining; (ii) we develop an automated data curation pipeline, SoccerFactory, to generate scalable spatial annotations, and integrate multiple existing soccer video datasets as a comprehensive pretraining data resource for multi-task pretraining; and (iii) we conduct extensive evaluations demonstrating that SoccerMaster consistently outperforms task-specific expert models across diverse downstream tasks, highlighting its breadth and superiority. The data, code, and model will be publicly available.
Explore related subjects
Keep this discovery
Haolin Yang, Jiayuan Rao, Haoning Wu, Weidi Xie. 2025-12-11. SoccerMaster: A Vision Foundation Model for Soccer Understanding. https://arxiv.org/abs/2512.11016
Cite the original work for its findings. Save a collection to share your selection of sources.