arXiv · 2601.13460
A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering
Abstract
The rapid growth of machine learning assets has made it increasingly difficult for software engineers to identify models and datasets that match their specific needs. Browsing large registries, such as Hugging Face, is time-consuming, error-prone, and rarely tailored to Software Engineering (SE) tasks. We present MLAssetSelection, a web application that automatically extracts SE assets and supports four key functionalities: (i) a configurable leaderboard for ranking models across multiple benchmarks and metrics; (ii) requirements-based selection of models and datasets; (iii) real-time automated updates through scheduled jobs that keep asset information current; and (iv) user-centric features including login, personalized asset lists, and configurable alert notifications. A demonstration video is available at https://youtu.be/t6CJ6P9asV4.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Alexandra González, Oscar Cerezo, Xavier Franch, Silverio Martínez-Fernández. 2026-01-19. A Tool for Automatically Cataloguing and Selecting Pre-Trained Models and Datasets for Software Engineering. https://arxiv.org/abs/2601.13460
Cite the original work for its findings. Save a collection to share your selection of sources.