arXiv · 2412.11314
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
Abstract
The rapid advancement of natural language processing (NLP) technologies, such as instruction-tuned large language models (LLMs), urges the development of modern evaluation protocols with human and machine feedback. We introduce Evalica, an open-source toolkit that facilitates the creation of reliable and reproducible model leaderboards. This paper presents its design, evaluates its performance, and demonstrates its usability through its Web interface, command-line interface, and Python API.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Dmitry Ustalov. 2024-12-15. Reliable, Reproducible, and Really Fast Leaderboards with Evalica. https://arxiv.org/abs/2412.11314
Cite the original work for its findings. Save a collection to share your selection of sources.