arXiv · 2112.10297
DXML: Distributed Extreme Multilabel Classification
Abstract
As a big data application, extreme multilabel classification has emerged as an important research topic with applications in ranking and recommendation of products and items. A scalable hybrid distributed and shared memory implementation of extreme classification for large scale ranking and recommendation is proposed. In particular, the implementation is a mix of message passing using MPI across nodes and using multithreading on the nodes using OpenMP. The expression for communication latency and communication volume is derived. Parallelism using work-span model is derived for shared memory architecture. This throws light on the expected scalability of similar extreme classification methods. Experiments show that the implementation is relatively faster to train and test on some large datasets. In some cases, model size is relatively small.
Explore related subjects
Keep this discovery
Pawan Kumar. 2021-10-20. DXML: Distributed Extreme Multilabel Classification. https://arxiv.org/abs/2112.10297
Cite the original work for its findings. Save a collection to share your selection of sources.