SearcharxivSearch

arXiv subjects

Renbo Zhao

Publications and source records attributed to Renbo Zhao.

At least 19 recordsLinked to original sources

Bregman Douglas-Rachford Splitting Method

In this paper, we propose the Bregman Douglas-Rachford splitting (BDRS) method and its variant Bregman Peaceman-Rachford splitting method for solving maximal monotone inclusion problem. We show that BDRS is equivalent to a Bregman alternating direction method of multipliers (ADMM) when applied to the dual of the problem. A special case of the Bregman ADMM is an alternating direction version of the exponential multiplier method. To the best of our knowledge, algorithms proposed in this paper are new to the literature. We also discuss how to use our algorithms to solve the discrete optimal transport (OT) problem. We prove the convergence of the algorithms under certain assumptions, though we point out that one assumption does not apply to the OT problem.

math.OC

An Optimization Perspective on the Monotonicity of the Multiplicative Algorithm for Optimal Experimental Design

We provide an optimization-based argument for the monotonicity of the multiplicative algorithm (MA) for a class of optimal experimental design problems considered in Yu (2010). Our proof avoids introducing auxiliary variables (or problems) and leveraging statistical arguments, and is much more straightforward and simpler compared to the proof in Yu (2010). The simplicity of our monotonicity proof also allows us to easily identify several sufficient conditions that ensure the strict monotonicity of MA. In addition, we provide two simple and similar-looking examples on which MA behaves very differently. These examples offer insight in the behaviors of MA, and also reveal some limitations of MA when applied to certain optimality criteria. We discuss these limitations, and pose open problems that may lead to deeper understanding of the behaviors of MA on these optimality criteria.

math.OC

Dual Averaging With Non-Strongly-Convex Prox-Functions: New Analysis and Algorithm

We present new analysis and algorithm of the dual-averaging-type (DA-type) methods for solving the composite convex optimization problem ${\min}_{x\in\mathbb{R}^n} \, f(\mathsf{A} x) + h(x)$, where $f$ is a convex and globally Lipschitz function, $\mathsf{A}$ is a linear operator, and $h$ is a ``simple'' and convex function that is used as the prox-function in the DA-type methods. We open new avenues of analyzing and developing DA-type methods, by going beyond the canonical setting where the prox-function $h$ is assumed to be strongly convex (on its domain). To that end, we identify two new sets of assumptions on $h$ (and also $f$ and $\mathsf{A}$) and show that they hold broadly for many important classes of non-strongly-convex functions. Under the first set of assumptions, we show that the original DA method still has a $O(1/k)$ primal-dual convergence rate. Moreover, we analyze the affine invariance of this method and its convergence rate. Under the second set of assumptions, we develop a new DA-type method with dual monotonicity, and show that it has a $O(1/k)$ primal-dual convergence rate. Finally, we consider the case where $f$ is only convex and Lipschitz on $\mathcal{C}:=\mathsf{A}(\mathsf{dom} h)$, and construct its globally convex and Lipschitz extension based on the Pasch-Hausdorff envelope. Furthermore, we characterize the sub-differential and Fenchel conjugate of this extension using the convex analytic objects associated with $f$ and $\mathcal{C}$.

math.OC

On the Convexification of Spectral Sets Induced by Non-Invariant Sets

Given a finite-dimensional FTvN system $(\mathbb{V},\mathbb{W},\lambda)$, we study the convexification of the spectral set $\lambda^{-1}(\mathcal{C})$ induced by a set $\mathcal{C} \subseteq \mathbb{W}$. While the case of invariant $\mathcal{C}$ has been relatively well-studied, the results for non-invariant $\mathcal{C}$ are largely lacking in the literature. We fill this void by developing simple and geometric characterizations of the convex hull and closed convex hull of $\lambda^{-1}(\mathcal{C})$ when $\mathcal{C}$ has no invariance property. We further specialize our results to the case of invariant $\mathcal{C}$, and obtain new convexifications of $\lambda^{-1}(\mathcal{C})$ in this case.

math.OC

An Away-Step Frank-Wolfe Method for Minimizing Logarithmically-Homogeneous Barriers

We present and analyze an away-step Frank-Wolfe method for the convex optimization problem ${\min}_{x\in\mathcal{X}} \; f(\mathsf{A} x) + \langle{c},{x}\rangle$, where $f$ is a $θ$-logarithmically-homogeneous self-concordant barrier, $\mathsf{A}$ is a linear operator that may be non-invertible, $\langle{c},{\cdot}\rangle$ is a linear function and $\mathcal{X}$ is a nonempty polytope. The applications of primary interest include D-optimal design, inference of multivariate Hawkes processes, and TV-regularized Poisson image de-blurring. We establish affine-invariant and norm-independent global linear convergence rates of our method, in terms of both the objective gap and the Frank-Wolfe gap. When specialized to the D-optimal design problem, our results settle a question left open since Ahipasaoglu, Sun and Todd (2008). We also show that the iterates generated by our method will land on and remain in a face of $\mathcal{X}$ within a bounded number of iterations, which can lead to improved local linear convergence rates (for both the objective gap and the Frank-Wolfe gap). We conduct numerical experiments on D-optimal design and inference of multivariate Hawkes processes, and our results not only demonstrate the efficiency and effectiveness of our method compared to other principled first-order methods, but also corroborate our theoretical results quite well.

math.OC

An Inexact Primal-Dual Smoothing Framework for Large-Scale Non-Bilinear Saddle Point Problems

We develop an inexact primal-dual first-order smoothing framework to solve a class of non-bilinear saddle point problems with primal strong convexity. Compared with existing methods, our framework yields a significant improvement over the primal oracle complexity, while it has competitive dual oracle complexity. In addition, we consider the situation where the primal-dual coupling term has a large number of component functions. To efficiently handle this situation, we develop a randomized version of our smoothing framework, which allows the primal and dual sub-problems in each iteration to be inexactly solved by randomized algorithms in expectation. The convergence of this framework is analyzed both in expectation and with high probability. In terms of the primal and dual oracle complexities, this framework significantly improves over its deterministic counterpart. As an important application, we adapt both frameworks for solving convex optimization problems with many functional constraints. To obtain an $\varepsilon$-optimal and $\varepsilon$-feasible solution, both frameworks achieve the best-known oracle complexities.

math.OC

A Primal-Dual Smoothing Framework for Max-Structured Non-Convex Optimization

We propose a primal-dual smoothing framework for finding a near-stationary point of a class of non-smooth non-convex optimization problems with max-structure. We analyze the primal and dual gradient complexities of the framework via two approaches, i.e., the dual-then-primal and primal-the-dual smoothing approaches. Our framework improves the best-known oracle complexities of the existing method, even in the restricted problem setting. As an important part of our framework, we propose a first-order method for solving a class of (strongly) convex-concave saddle-point problems, which is based on a newly developed non-Hilbertian inexact accelerated proximal gradient algorithm for strongly convex composite minimization that enjoys duality-gap convergence guarantees. Some variants and extensions of our framework are also discussed.

math.OC

Fast Learning of Multidimensional Hawkes Processes via Frank-Wolfe

Hawkes processes have recently risen to the forefront of tools when it comes to modeling and generating sequential events data. Multidimensional Hawkes processes model both the self and cross-excitation between different types of events and have been applied successfully in various domain such as finance, epidemiology and personalized recommendations, among others. In this work we present an adaptation of the Frank-Wolfe algorithm for learning multidimensional Hawkes processes. Experimental results show that our approach has better or on par accuracy in terms of parameter estimation than other first order methods, while enjoying a significantly faster runtime.

cs.LG

A Generalized Frank-Wolfe Method With "Dual Averaging" for Strongly Convex Composite Optimization

We propose a simple variant of the generalized Frank-Wolfe method for solving strongly convex composite optimization problems, by introducing an additional averaging step on the dual variables. We show that in this variant, one can choose a simple constant step-size and obtain a linear convergence rate on the duality gaps. By leveraging the convergence analysis of this variant, we then analyze the local convergence rate of the logistic fictitious play algorithm, which is well-established in game theory but lacks any form of convergence rate guarantees. We show that, with high probability, this algorithm converges locally at rate $O(1/t)$, in terms of certain expected duality gap.

math.OC

Convergence Rate Analysis of the Multiplicative Gradient Method for PET-Type Problems

We analyze the convergence rate of the multiplicative gradient (MG) method for PET-type problems with $m$ component functions and an $n$-dimensional optimization variable. We show that the MG method has an $O(\ln(n)/t)$ convergence rate, in both the ergodic and the non-ergodic senses. Furthermore, we show that the distances from the iterates to the set of optimal solutions converge (to zero) at rate $O(1/\sqrt{t})$. Our results show that, in the regime $n=O(\exp(m))$, to find an $\varepsilon$-optimal solution of the PET-type problems, the MG method has a lower computational complexity compared with the relatively-smooth gradient method and the Frank-Wolfe method for convex composite optimization involving a logarithmically-homogeneous barrier.

math.OC

The Generalized Multiplicative Gradient Method for A Class of Convex Optimization Problems Over Symmetric Cones

We develop and analyze the Generalized Multiplicative Gradient (GMG) method for solving a class of convex optimization problems over symmetric cones, where the objective function does not have Lipschitz gradient over the feasible region. This problem class includes several applications, such as positron emission tomography, D-optimal design, quantum state tomography and the dual problem of Nesterov's convex relaxation of the boolean quadratic problem. We show that the GMG method has a convergence rate of $O(1/k)$ in terms of the objective gap. Our analysis of the convergence rate is rather unconventional, and to that end, we establish several results that may be of independent interest, such as a curvature bound of the Legendre and logarithmically-homogeneous functions, and a Cauchy-Schwarz inequality in representative simple Euclidean Jordan Algebras. Finally, we compare the computational complexity of the GMG method with three other related first-order methods on several important applications, and we show that under certain mild assumptions, the GMG method achieves the best (or almost the best) computational complexities on all of the applications.

math.OC

Analysis of the Frank-Wolfe Method for Convex Composite Optimization involving a Logarithmically-Homogeneous Barrier

We present and analyze a new generalized Frank-Wolfe method for the composite optimization problem $(P):{\min}_{x\in\mathbb{R}^n}\; f(\mathsf{A} x) + h(x)$, where $f$ is a $θ$-logarithmically-homogeneous self-concordant barrier, $\mathsf{A}$ is a linear operator and the function $h$ has bounded domain but is possibly non-smooth. We show that our generalized Frank-Wolfe method requires $O((δ_0 + θ+ R_h)\ln(δ_0) + (θ+ R_h)^2/\varepsilon)$ iterations to produce an $\varepsilon$-approximate solution, where $δ_0$ denotes the initial optimality gap and $R_h$ is the variation of $h$ on its domain. This result establishes certain intrinsic connections between $θ$-logarithmically homogeneous barriers and the Frank-Wolfe method. When specialized to the $D$-optimal design problem, we essentially recover the complexity obtained by Khachiyan using the Frank-Wolfe method with exact line-search. We also study the (Fenchel) dual problem of $(P)$, and we show that our new method is equivalent to an adaptive-step-size mirror descent method applied to the dual problem. This enables us to provide iteration complexity bounds for the mirror descent method despite even though the dual objective function is non-Lipschitz and has unbounded domain. In addition, we present computational experiments that point to the potential usefulness of our generalized Frank-Wolfe method on Poisson image de-blurring problems with TV regularization, and on simulated PET problem instances.

math.OC

Accelerated Stochastic Algorithms for Convex-Concave Saddle-Point Problems

We develop stochastic first-order primal-dual algorithms to solve a class of convex-concave saddle-point problems. When the saddle function is strongly convex in the primal variable, we develop the first stochastic restart scheme for this problem. When the gradient noises obey sub-Gaussian distributions, the oracle complexity of our restart scheme is strictly better than any of the existing methods, even in the deterministic case. Furthermore, for each problem parameter of interest, whenever the lower bound exists, the oracle complexity of our restart scheme is either optimal or nearly optimal (up to a log factor). The subroutine used in this scheme is itself a new stochastic algorithm developed for the problem where the saddle function is non-strongly convex in the primal variable. This new algorithm, which is based on the primal-dual hybrid gradient framework, achieves the state-of-the-art oracle complexity and may be of independent interest.

math.OC

Stochastic L-BFGS: Improved Convergence Rates and Practical Acceleration Strategies

We revisit the stochastic limited-memory BFGS (L-BFGS) algorithm. By proposing a new framework for the convergence analysis, we prove improved convergence rates and computational complexities of the stochastic L-BFGS algorithms compared to previous works. In addition, we propose several practical acceleration strategies to speed up the empirical performance of such algorithms. We also provide theoretical analyses for most of the strategies. Experiments on large-scale logistic and ridge regression problems demonstrate that our proposed strategies yield significant improvements vis-à-vis competing state-of-the-art algorithms.

math.OC

A Unified Convergence Analysis of the Multiplicative Update Algorithm for Regularized Nonnegative Matrix Factorization

The multiplicative update (MU) algorithm has been extensively used to estimate the basis and coefficient matrices in nonnegative matrix factorization (NMF) problems under a wide range of divergences and regularizers. However, theoretical convergence guarantees have only been derived for a few special divergences without regularization. In this work, we provide a conceptually simple, self-contained, and unified proof for the convergence of the MU algorithm applied on NMF with a wide range of divergences and regularizers. Our main result shows the sequence of iterates (i.e., pairs of basis and coefficient matrices) produced by the MU algorithm converges to the set of stationary points of the non-convex NMF optimization problem. Our proof strategy has the potential to open up new avenues for analyzing similar problems in machine learning and signal processing.

math.OC

A Unified Framework for Stochastic Matrix Factorization via Variance Reduction

We propose a unified framework to speed up the existing stochastic matrix factorization (SMF) algorithms via variance reduction. Our framework is general and it subsumes several well-known SMF formulations in the literature. We perform a non-asymptotic convergence analysis of our framework and derive computational and sample complexities for our algorithm to converge to an $ε$-stationary point in expectation. In addition, extensive experiments for a wide class of SMF formulations demonstrate that our framework consistently yields faster convergence and a more accurate output dictionary vis-à-vis state-of-the-art frameworks.

stat.ML

Online Nonnegative Matrix Factorization with Outliers

We propose a unified and systematic framework for performing online nonnegative matrix factorization in the presence of outliers. Our framework is particularly suited to large-scale data. We propose two solvers based on projected gradient descent and the alternating direction method of multipliers. We prove that the sequence of objective values converges almost surely by appealing to the quasi-martingale convergence theorem. We also show the sequence of learned dictionaries converges to the set of stationary points of the expected loss function almost surely. In addition, we extend our basic problem formulation to various settings with different constraints and regularizers. We also adapt the solvers and analyses to each setting. We perform extensive experiments on both synthetic and real datasets. These experiments demonstrate the computational efficiency and efficacy of our algorithms on tasks such as (parts-based) basis learning, image denoising, shadow removal and foreground-background separation.

stat.ML