Nonnegative Matrix Factorization in the Component-Wise L1 Norm for Sparse Data
Nonnegative matrix factorization (NMF) approximates a nonnegative matrix, X, by the product of two nonnegative factors, WH, where W has r columns and H has r rows. In this paper, we consider NMF using the component-wise L1 norm as the error measure (L1-NMF), which is suited for data corrupted by heavy-tailed noise, such as Laplace noise or salt and pepper noise, or in the presence of outliers. Our first contribution is an NP-hardness proof for L1-NMF, even when r=1, in contrast to the standard NMF that uses least squares. Our second contribution is to analyze, under simplified probabilistic assumptions, how the sparsity in the data enforces zero solution in the optimal scalar update in the factors of L1-NMF when all the other entries are kept fixed. This provides an intuition of the connection between the sparsity of the L1-NMF factors with the sparsity of the input. Even though sparsity favors interpretability, if the data is affected by false zeros, too sparse solutions might degrade the model. Our third contribution is a new, more general, L1-NMF model for sparse data, dubbed weighted L1-NMF (wL1-NMF), where the sparsity of the factorization is controlled by adding a penalization parameter to the entries of WH associated with zeros in the data. The fourth contribution is a new coordinate descent (CD) approach for wL1-NMF, denoted as sparse CD (sCD), where each subproblem is solved by a weighted median algorithm. Although it lacks convergence guarantees to a stationary point, sCD is, to the best of our knowledge, the first algorithm for L1-NMF whose complexity scales with the number of nonzero entries in the data, making it efficient in handling large-scale, sparse data. We perform extensive numerical experiments on synthetic and real-world data, including imaging mass spectrometry and topic modeling, to show the effectiveness of our new proposed model (wL1-NMF) and algorithm (sCD).