arXiv · 2601.20415
An Empirical Evaluation of Modern MLOps Frameworks
Abstract
Given the increasing adoption of AI solutions in professional environments, it is necessary for developers to be able to make informed decisions about the current tool landscape. This work empirically evaluates various MLOps (Machine Learning Operations) tools to facilitate the management of the ML model lifecycle: MLflow, Metaflow, Apache Airflow, and Kubeflow Pipelines. The tools are evaluated by assessing the criteria of Ease of installation, Configuration flexibility, Interoperability, Code instrumentation complexity, result interpretability, and Documentation when implementing two common ML scenarios: Digit classifier with MNIST and Sentiment classifier with IMDB and BERT. The evaluation is completed by providing weighted results that lead to practical conclusions on which tools are best suited for different scenarios.
Explore related subjects
Keep this discovery
Jon Marcos-Mercadé, Unai Lopez-Novoa, Mikel Egaña Aranguren. 2026-01-28. An Empirical Evaluation of Modern MLOps Frameworks. https://arxiv.org/abs/2601.20415
Cite the original work for its findings. Save a collection to share your selection of sources.