Searcharxiv⌕ Search

arXiv subjects

Yuyuan Ouyang

Publications and source records attributed to Yuyuan Ouyang.

At least 19 recordsLinked to original sources

An $O(1/T^3)$ algorithm for minimizing convex quadratic functions over the $L_1$ ball

We study convex quadratic minimization over the unit $L_1$ ball in which the maximum eigenvalue of the Hessian matrix is bounded by a positive constant $L$. We propose a novel first-order algorithm with objective value error bounded by $O(L/T^3)$ after $T$ gradient evaluations, assuming that the subproblems involved in the algorithm can be solved exactly. To the best of our knowledge, the best convergence rate of algorithms in the literature is $O(L/T^2)$. From the perspective of information-based complexity theory, our proposed algorithm is the first in the literature that achieves the $O((L/\varepsilon)^{1/3})$ first-order oracle complexity, although its current version is not necessarily practical for implementation. We hope that our proposed algorithm could shed some light on future implementable and efficient $O(L/T^3)$-convergence-rate algorithms. The proposed algorithm incorporates a decomposition of components of vectors in the unit $L_1$-norm ball to "good" and "bad" parts, and uses symmetric rank-1 (SR1) updates on the bad parts. The proposed algorithm was developed after the author instructed the OpenAI ChatGPT 6 (Astra) model to study the problem using ideas of weak-type $L^1$ estimates and good-bad part decomposition in harmonic analysis and a recent result.

math.OC↗

A Compact Counterexample to Suárez's Berezin Approximation Question

Let $A^2(\mathbb D)$ be the unweighted Bergman space and write $Q_m(S)=T_{B_m(S)}$ for the map induced by the $m$th higher-order Berezin transform. Suárez asked in 2005 whether $Q_m(S)$ converges to $S$ in operator norm for every $S$ in the full Bergman Toeplitz algebra. We answer this question negatively in a strong form: there is a compact operator $S$ such that \[ \sup_{m\geq 0}\|Q_m(S)\|=\infty. \] In particular, along a strictly increasing sequence $(m_n)$ one has \[ \|Q_{m_n}(S)-S\|\longrightarrow\infty. \] The obstruction is a moving matrix edge. The proof uses a moving family of rank-one test operators, each supported in the $m$th matrix column. Near the corresponding edge, these test operators are sent to weighted Hankel matrices, and the limiting coefficients form an explicit Pascal kernel. We prove a standalone Pascal--Hankel theorem showing that the associated weighted Hankel transformation fails to map $\ell^2$ boundedly into trace class. Finite-section convergence, trace duality, and the Uniform Boundedness Principle then transfer this instability to the exact Bergman maps.

math.FA↗

Quadratic Convexification of a Square Truncated by a Hyperbola

We study the quadratic convexification of the compact nonconvex set ${G} := \left[1/2,2\right]^2 \cap \{(x_1,x_2)\in\mathbb R^2\mid x_1x_2\leq1\}, $ which is a square truncated by a hyperbolic arc. Although the ordinary convex hull of ${G}$ is a triangle, its lifted convex hull in the complete quadratic space retains the nontrivial geometry of the curved boundary. We characterize all extreme rays of the cone of quadratic polynomials nonnegative on $G$, including parameterized families of bounded tangent and bitangent rays. Using this classification and conic duality, we derive an exact finite semidefinite representation of the lifted convex hull of ${G}$. The analysis follows the general framework of our earlier work on an unbounded product-constrained region, but the bounded geometry creates new boundary-contact patterns and leads to a different finite organization of the nonnegative-quadratic cone.

math.OC↗

A dimension-free weak-type $(1,1)$ bound for the vector Riesz transform on $\mathbb{R}^n$

We show that the best constant in the weak-type $(1,1)$ bound for the vector Riesz transform on $\mathbb{R}^n$ is at most $2$, independent of the dimension $n$. The proof relies on a new decomposition of the input data involving an obstacle problem for the fractional Laplacian and an associated Lewy-Stampacchia type estimate on an unbounded domain. This settles a problem posed by E. M. Stein at the 1986 International Congress of Mathematicians.

math.CA↗

Nonnegative Quadratics over a Quadrant with a Bilinear Constraint

We study quadratic polynomials that are nonnegative on the non-compact set \[ F:=\{(x_1,x_2)\in\mathbb R^2:\ x_1\ge 0,\ x_2\ge 0,\ x_1x_2\le 1\}. \] All extreme rays of the cone of nonnegative quadratic polynomials are characterized on this set. The characterization allows us to study parameterized valid inequalities for quadratic convexifications involving $F$, which yields a semidefinite representation of its lifted convex hull in the quadratic space. Our analysis on extreme ray characterization separates the positive-semidefinite (PSD) and non-PSD branches, reduces the latter to boundary nonnegativity, and classifies the boundary contacts of the relevant extreme rays. By reparameterization, we also find a non-trivial and non-permutation-symmetric six-dimensional linear section of the cone of nonnegative homogeneous ternary octics that are sum-of-squares. We also show that our lifted convex hull result yields a degree-bounded preordering certificate of nonnegative quadratics on $F$ and a degree-bounded certificate for a family of nonnegative quartics on the half-strip.

math.OC↗

Universal and Parameter-free Gradient Sliding for Composite Optimization

We propose a Parameter-Free Universal Gradient Sliding (PFUGS) algorithm for computing an approximate solution to the convex composite optimization $\min_{x\in X} \{f(x) + g(x)\}$, where $f$ has $(M_ν,ν)$-Hölder continuous subgradient and $g$ has $L$-Lipschitz continuous gradient. PFUGS computes an $\varepsilon$-approximate solution with $\mathcal{O}((M_ν/\varepsilon)^{{2}/{(1+3ν)}})$ evaluations of (sub)gradients of $f$ and $\mathcal{O}((L/\varepsilon)^{1/2})$ evaluations of gradients of $g$, without prior knowledge of problem constants. To the best of our knowledge, PFUGS is the first gradient sliding algorithm for problems involving two functions whose distinct problem constants are both unknown a priori.

math.OC↗

A Note on the Gradient-Evaluation Sequence in Accelerated Gradient Methods

Nesterov's accelerated gradient descent method (AGD) is a seminal deterministic first-order method known to achieve the optimal order of iteration complexity for solving convex smooth optimization problems. Two distinct sequences of iterates are included in the description of AGD: gradient evaluations are performed at one sequence, while approximate solutions are selected from the other. The iteration complexity on minimizing objective function value has been well-studied in the literature, but such analysis is almost always performed only at the approximate solution sequence. To the best of our knowledge, for projection-based AGD that solves problems with feasible sets, it is still an open research question whether the gradient evaluation sequence (when treated as approximate solutions) could also achieve the same optimal order of iteration complexity. It is also unknown whether such results still hold in the non-Euclidean setting. Motivated by computer-aided algorithm analysis, we provide positive results that answer the open problems affirmatively. Specifically, for (possibly constrained) problem $f^*:=\min_{x\in X}f(x)$ where $f$ is convex and $L$-smooth and $X$ is closed, convex and projection friendly, we prove that the gradient-evaluation sequence $\{\underline{x}_k\}$ in AGD satisfies that $f(\underline{x}_k) - f^*\le \mathcal O(L/k^2)$.

math.OC↗

On the Hardness of the $L_1-L_2$ Regularization Problem

The sparse linear reconstruction problem is a core problem in signal processing which aims to recover sparse solutions to linear systems. The original problem regularized by the total number of nonzero components (also known as $L_0$ regularization) is well-known to be NP-hard. The relaxation of the $L_0$ regularization by using the $L_1$ norm offers a convex reformulation, but is only exact under certain conditions (e.g., restricted isometry property) which might be NP-hard to verify. To overcome the computational hardness of the $L_0$ regularization problem while providing tighter results than the $L_1$ relaxation, several alternate optimization problems have been proposed to find sparse solutions. One such problem is the $L_1-L_2$ minimization problem, which is to minimize the difference of the $L_1$ and $L_2$ norms subject to linear constraints. This paper proves that solving the $L_1-L_2$ minimization problem is NP-hard. Specifically, we prove that it is NP-hard to minimize the $L_1-L_2$ regularization function subject to linear constraints. Moreover, it is also NP-hard to solve the unconstrained formulation that minimizes the sum of a least squares term and the $L_1-L_2$ regularization function. Furthermore, restricting the feasible set to a smaller one by adding nonnegative constraints does not change the NP-hardness nature of the problems.

math.OC↗

Optimal and parameter-free gradient minimization methods for convex and nonconvex optimization

We propose novel optimal and parameter-free algorithms for computing an approximate solution with small (projected) gradient norm. Specifically, for computing an approximate solution such that the norm of its (projected) gradient does not exceed $\varepsilon$, we obtain the following results: a) for the convex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{L\|x_0 - x^*\|/\varepsilon}$, where $L$ is the Lipschitz constant of the gradient and $x^*$ is any optimal solution; b) for the strongly convex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{L/μ}\log(\|\nabla f(x_0)\|/ε)$, where $μ$ is the strong convexity modulus; and c) for the nonconvex case, the total number of gradient evaluations is bounded by $O(1)\sqrt{Ll}(f(x_0) - f(x^*))/\varepsilon^2$, where $l$ is the lower curvature constant. Our complexity results match the lower complexity bounds of the convex and strongly cases, and achieve the above best-known complexity bound for the nonconvex case for the first time in the literature. Our results can also be extended to problems with constraints and composite objectives. Moreover, for all the convex, strongly convex, and nonconvex cases, we propose parameter-free algorithms that do not require the input of any problem parameters or the convexity status of the problem. To the best of our knowledge, there do not exist such parameter-free methods before especially for the strongly convex and nonconvex cases. Since most regularity conditions (e.g., strong convexity and lower curvature) are imposed over a global scope, the corresponding problem parameters are notoriously difficult to estimate. However, gradient norm minimization equips us with a convenient tool to monitor the progress of algorithms and thus the ability to estimate such parameters in-situ.

math.OC↗

Mirror-prox sliding methods for solving a class of monotone variational inequalities

In this paper we propose new algorithms for solving a class of structured monotone variational inequality (VI) problems over compact feasible sets. By identifying the gradient components existing in the operator of VI, we show that it is possible to skip computations of the gradients from time to time, while still maintaining the optimal iteration complexity for solving these VI problems. Specifically, for deterministic VI problems involving the sum of the gradient of a smooth convex function $\nabla G$ and a monotone operator $H$, we propose a new algorithm, called the mirror-prox sliding method, which is able to compute an $\varepsilon$-approximate weak solution with at most $O((L/\varepsilon)^{1/2})$ evaluations of $\nabla G$ and $O((L/\varepsilon)^{1/2}+M/\varepsilon)$ evaluations of $H$, where $L$ and $M$ are Lipschitz constants of $\nabla G$ and $H$, respectively. Moreover, for the case when the operator $H$ can only be accessed through its stochastic estimators, we propose a stochastic mirror-prox sliding method that can compute a stochastic $\varepsilon$-approximate weak solution with at most $O((L/\varepsilon)^{1/2})$ evaluations of $\nabla G$ and $O((L/\varepsilon)^{1/2}+M/\varepsilon + σ^2/\varepsilon^2)$ samples of $H$, where $σ$ is the variance of the stochastic samples of $H$.

math.OC↗

Universal Conditional Gradient Sliding for Convex Optimization

In this paper, we present a first-order projection-free method, namely, the universal conditional gradient sliding (UCGS) method, for solving $\varepsilon$-approximate solutions to convex differentiable optimization problems. For objective functions with Hölder continuous gradients, we show that UCGS is able to terminate with $\varepsilon$-solutions with at most $O((M_νD_X^{1+ν}/{\varepsilon})^{2/(1+3ν)})$ gradient evaluations and $O((M_νD_X^{1+ν}/{\varepsilon})^{4/(1+3ν)})$ linear objective optimizations, where $ν\in (0,1]$ and $M_ν>0$ are the exponent and constant of the Hölder condition. Furthermore, UCGS is able to perform such computations without requiring any specific knowledge of the smoothness information $ν$ and $M_ν$. In the weakly smooth case when $ν\in (0,1)$, both complexity results improve the current state-of-the-art $O((M_νD_X^{1+ν}/{\varepsilon})^{1/ν})$ results on first-order projection-free method achieved by the conditional gradient method. Within the class of sliding-type algorithms, to the best of our knowledge, this is the first time a sliding-type algorithm is able to improve not only the gradient complexity but also the overall complexity for computing an approximate solution. In the smooth case when $ν=1$, UCGS matches the state-of-the-art complexity result but adds more features allowing for practical implementation.

math.OC↗

Graph topology invariant gradient and sampling complexity for decentralized and stochastic optimization

One fundamental problem in decentralized multi-agent optimization is the trade-off between gradient/sampling complexity and communication complexity. We propose new algorithms whose gradient and sampling complexities are graph topology invariant while their communication complexities remain optimal. For convex smooth deterministic problems, we propose a primal dual sliding (PDS) algorithm that computes an $ε$-solution with $O((\tilde{L}/ε)^{1/2})$ gradient and $O((\tilde{L}/ε)^{1/2}+\|\mathcal{A}\|/ε)$ communication complexities, where $\tilde{L}$ is the smoothness parameter of the objective and $\mathcal{A}$ is related to either the graph Laplacian or the transpose of the oriented incidence matrix of the communication network. The results can be improved to $O((\tilde{L}/μ)^{1/2}\log(1/ε))$ and $O((\tilde{L}/μ)^{1/2}\log(1/ε) + \|\mathcal{A}\|/ε^{1/2})$ respectively with $μ$-strong convexity. We also propose a stochastic variant, the primal dual sliding (SPDS) algorithm for problems with stochastic gradients. The SPDS algorithm utilizes the mini-batch technique and enables the agents to perform sampling and communication simultaneously. It computes a stochastic $ε$-solution with $O((\tilde{L}/ε)^{1/2} + (σ/ε)^2)$ sampling complexity, which can be improved to $O((\tilde{L}/μ)^{1/2}\log(1/ε) + σ^2/ε)$ with strong convexity. Here $σ^2$ is the variance. The communication complexities of SPDS remain the same as that of the deterministic case. All the aforementioned gradient and sampling complexities match the lower complexity bounds for centralized convex smooth optimization and are independent of the network structure. To the best of our knowledge, these gradient and sampling complexities have not been obtained before for decentralized optimization over a constraint feasible set.

math.OC↗

Backtracking linesearch for conditional gradient sliding

We present a modification of the conditional gradient sliding (CGS) method that was originally developed in \cite{lan2016conditional}. While the CGS method is a theoretical breakthrough in the theory of projection-free first-order methods since it is the first that reaches the theoretical performance limit, in implementation it requires the knowledge of the Lipschitz constant of the gradient of the objective function $L$ and the number of total gradient evaluations $N$. Such requirements imposes difficulties in the actual implementation, not only because that it can be difficult to choose proper values of $L$ and $N$ that satisfies the conditions for convergence, but also since conservative choices of $L$ and $N$ can deteriorate the practical numerical performance of the CGS method. Our proposed method, called the conditional gradient sliding method with linesearch (CGS-ls), does not require the knowledge of either $L$ and $N$, and is able to terminate early before the theoretically required number of iterations. While more practical in numerical implementation, the theoretical performance of our proposed CGS-ls method is still as good as that of the CGS method. We present numerical experiments to show the efficiency of our proposed method in practice.

math.OC↗

An Asymptotic Result of Conditional Logistic Regression Estimator

In cluster-specific studies, ordinary logistic regression and conditional logistic regression for binary outcomes provide maximum likelihood estimator (MLE) and conditional maximum likelihood estimator (CMLE), respectively. In this paper, we show that CMLE is approaching to MLE asymptotically when each individual data point is replicated infinitely many times. Our theoretical derivation is based on the observation that a term appearing in the conditional average log-likelihood function is the coefficient of a polynomial, and hence can be transformed to a complex integral by Cauchy's differentiation formula. The asymptotic analysis of the complex integral can then be performed using the classical method of steepest descent. Our result implies that CMLE can be biased if individual weights are multiplied with a constant, and that we should be cautious when assigning weights to cluster-specific studies.

math.ST↗

Some Worst-Case Datasets of Deterministic First-Order Methods for Solving Binary Logistic Regression

We present in this paper some worst-case datasets of deterministic first-order methods for solving large-scale binary logistic regression problems. Under the assumption that the number of algorithm iterations is much smaller than the problem dimension, with our worst-case datasets it requires at least $\mathcal{O}(1/\sqrt{\varepsilon})$ first-order oracle inquiries to compute an $\varepsilon$-approximate solution. From traditional iteration complexity analysis point of view, the binary logistic regression loss functions with our worst-case datasets are new worst-case function instances among the class of smooth convex optimization problems.

math.OC↗

Lower complexity bounds of first-order methods for convex-concave bilinear saddle-point problems

On solving a convex-concave bilinear saddle-point problem (SPP), there have been many works studying the complexity results of first-order methods. These results are all about upper complexity bounds, which can determine at most how many efforts would guarantee a solution of desired accuracy. In this paper, we pursue the opposite direction by deriving lower complexity bounds of first-order methods on large-scale SPPs. Our results apply to the methods whose iterates are in the linear span of past first-order information, as well as more general methods that produce their iterates in an arbitrary manner based on first-order information. We first work on the affinely constrained smooth convex optimization that is a special case of SPP. Different from gradient method on unconstrained problems, we show that first-order methods on affinely constrained problems generally cannot be accelerated from the known convergence rate $O(1/t)$ to $O(1/t^2)$, and in addition, $O(1/t)$ is optimal for convex problems. Moreover, we prove that for strongly convex problems, $O(1/t^2)$ is the best possible convergence rate, while it is known that gradient methods can have linear convergence on unconstrained problems. Then we extend these results to general SPPs. It turns out that our lower complexity bounds match with several established upper complexity bounds in the literature, and thus they are tight and indicate the optimality of several existing first-order methods.

math.OC↗

An inequality for Jacobi polynomials of form $P_n^{(α_n,β_n)}(x)$

We prove an inequality for Jacobi polynomials that \begin{align} Δ_n(x):=P_n^{(α_n,β_n)}(x)P_n^{(α_{n+1},β_{n+1})}(x)- P_{n-1}^{(α_n,β_n)}(x)P_{n+1}^{(α_{n+1},β_{n+1})}(x)\le 0,\ \forall x\ge 1, \end{align} where $α_n=an$ and $β_n=bn$ for some $a,b\ge 0$. The above inequality has a similar taste as the Tuŕan type inequalities, but with $α_n$ and $β_n$ that depends linearly on $n$.

math.CA↗

Accelerated gradient sliding for structured convex optimization

Our main goal in this paper is to show that one can skip gradient computations for gradient descent type methods applied to certain structured convex programming (CP) problems. To this end, we first present an accelerated gradient sliding (AGS) method for minimizing the summation of two smooth convex functions with different Lipschitz constants. We show that the AGS method can skip the gradient computation for one of these smooth components without slowing down the overall optimal rate of convergence. This result is much sharper than the classic black-box CP complexity results especially when the difference between the two Lipschitz constants associated with these components is large. We then consider an important class of bilinear saddle point problem whose objective function is given by the summation of a smooth component and a nonsmooth one with a bilinear saddle point structure. Using the aforementioned AGS method for smooth composite optimization and Nesterov's smoothing technique, we show that one only needs ${\cal O}(1/\sqrtε)$ gradient computations for the smooth component while still preserving the optimal ${\cal O}(1/ε)$ overall iteration complexity for solving these saddle point problems. We demonstrate that even more significant savings on gradient computations can be obtained for strongly convex smooth and bilinear saddle point problems.

math.OC↗