SearcharxivSearch

arXiv subjects

Xiaozhou Wang

Publications and source records attributed to Xiaozhou Wang.

10 recordsLinked to original sources

Byzantine-tolerant distributed learning of finite mixture models under partial corruptions

Finite mixture models characterize heterogeneous populations and are increasingly fitted to distributed data using split-and-conquer procedures that aggregate local mixture estimates at a central server. The aggregation step can be seriously compromised when transmitted local mixture estimates are partially or completely corrupted. To guard against Byzantine failures, existing robust aggregation methods have been developed for settings in which a local mixture estimate is either entirely authentic or entirely corrupted. Such methods can discard useful information when only some component estimates are corrupted. We consider component-wise Byzantine failure, in which some component estimates may be corrupted, but for each mixture component, a majority of the corresponding local estimates remain authentic. We propose component-wise filtered mixture reduction (CFMR), which selects a data-driven anchor, aligns transmitted components, filters each aligned cluster by a majority radius, and aggregates the retained estimates through mixture reduction. By filtering out only unreliable components, CFMR preserves authentic information from partially corrupted machines without requiring knowledge of the failure rates. We establish an adaptive convergence bound for CFMR and show that, under suitable conditions, it attains the oracle rate that would be achieved if the authentic component estimates were known in advance. Simulations and a real-data application show that CFMR remains close to the component-level oracle, whereas whole-machine filtering and unprotected aggregation can deteriorate substantially.

stat.ME

A conditional-gradient-based single-loop augmented Lagrangian method for inequality constrained optimization

We consider the problem of minimizing the sum of a Lipschitz differentiable convex function $f$ and a proper closed convex function $h$ that admits efficient linear minimization oracles, subject to multiple smooth convex inequality constraints. We adapt the classical augmented Lagrangian (AL) method for these problems: in each iteration, our algorithm consists of one step of the conditional gradient (CG) method applied to the AL function, followed by an update of the dual variable as in classical AL methods with a diminishing dual stepsize. We study the convergence rate of our algorithm under two standard stepsize rules for the CG method, namely, an open-loop stepsize and the short stepsize, and obtain a convergence rate that matches the best-known complexity for this class of problems. We also establish accelerated rates when $h$ is the indicator function of a uniformly convex set.

math.OC

Error bounds for perspective cones of a class of nonnegative Legendre functions

Error bounds play a central role in the study of conic optimization problems, including the analysis of convergence rates for numerous algorithms. Curiously, those error bounds are often Hölderian with exponent 1/2. In this paper, we try to explain the prevalence of the 1/2 exponent by investigating generic properties of error bounds for conic feasibility problems where the underlying cone is a perspective cone constructed from a nonnegative Legendre function on $\mathbb{R}$. Our analysis relies on the facial reduction technique and the computation of one-step facial residual functions (1-FRFs). Specifically, under appropriate assumptions on the Legendre function, we show that 1-FRFs can be taken to be Hölderian of exponent 1/2 almost everywhere with respect to the two-dimensional Hausdorff measure. This enables us to further establish that having a uniform Hölderian error bound with exponent 1/2 is a generic property for a class of feasibility problems involving these cones.

math.OC

Frank-Wolfe-type methods for a class of nonconvex inequality-constrained problems

The Frank-Wolfe (FW) method, which implements efficient linear oracles that minimize linear approximations of the objective function over a fixed compact convex set, has recently received much attention in the optimization and machine learning literature. In this paper, we propose a new FW-type method for minimizing a smooth function over a compact set defined as the level set of a single difference-of-convex function, based on new generalized linear-optimization oracles (LO). We show that these LOs can be computed efficiently with closed-form solutions in some important optimization models that arise in compressed sensing and machine learning. In addition, under a mild strict feasibility condition, we establish the subsequential convergence of our nonconvex FW-type method. Since the feasible region of our generalized LO typically changes from iteration to iteration, our convergence analysis is completely different from those existing works in the literature on FW-type methods that deal with fixed feasible regions among subproblems. Finally, motivated by the away steps for accelerating FW-type methods for convex problems, we further design an away-step oracle to supplement our nonconvex FW-type method, and establish subsequential convergence of this variant. Numerical results on the matrix completion problem with standard datasets are presented to demonstrate the efficiency of the proposed FW-type method and its away-step variant.

math.OC

Convergence rate analysis of a Dykstra-type projection algorithm

Given closed convex sets $C_i$, $i=1,\ldots,\ell$, and some nonzero linear maps $A_i$, $i = 1,\ldots,\ell$, of suitable dimensions, the multi-set split feasibility problem aims at finding a point in $\bigcap_{i=1}^\ell A_i^{-1}C_i$ based on computing projections onto $C_i$ and multiplications by $A_i$ and $A_i^T$. In this paper, we consider the associated best approximation problem, i.e., the problem of computing projections onto $\bigcap_{i=1}^\ell A_i^{-1}C_i$; we refer to this problem as the best approximation problem in multi-set split feasibility settings (BA-MSF). We adapt the Dykstra's projection algorithm, which is classical for solving the BA-MSF in the special case when all $A_i = I$, to solve the general BA-MSF. Our Dykstra-type projection algorithm is derived by applying (proximal) coordinate gradient descent to the Lagrange dual problem, and it only requires computing projections onto $C_i$ and multiplications by $A_i$ and $A_i^T$ in each iteration. Under a standard relative interior condition and a genericity assumption on the point we need to project, we show that the dual objective satisfies the Kurdyka-Lojasiewicz property with an explicitly computable exponent on a neighborhood of the (typically unbounded) dual solution set when each $C_i$ is $C^{1,α}$-cone reducible for some $α\in (0,1]$: this class of sets covers the class of $C^2$-cone reducible sets, which include all polyhedrons, second-order cone, and the cone of positive semidefinite matrices as special cases. Using this, explicit convergence rate (linear or sublinear) of the sequence generated by the Dykstra-type projection algorithm is derived. Concrete examples are constructed to illustrate the necessity of some of our assumptions.

math.OC

A review of distributed statistical inference

The rapid emergence of massive datasets in various fields poses a serious challenge to traditional statistical methods. Meanwhile, it provides opportunities for researchers to develop novel algorithms. Inspired by the idea of divide-and-conquer, various distributed frameworks for statistical estimation and inference have been proposed. They were developed to deal with large-scale statistical optimization problems. This paper aims to provide a comprehensive review for related literature. It includes parametric models, nonparametric models, and other frequently used models. Their key ideas and theoretical properties are summarized. The trade-off between communication cost and estimate precision together with other concerns are discussed.

stat.CO

SMPL: Simulated Industrial Manufacturing and Process Control Learning Environments

Traditional biological and pharmaceutical manufacturing plants are controlled by human workers or pre-defined thresholds. Modernized factories have advanced process control algorithms such as model predictive control (MPC). However, there is little exploration of applying deep reinforcement learning to control manufacturing plants. One of the reasons is the lack of high fidelity simulations and standard APIs for benchmarking. To bridge this gap, we develop an easy-to-use library that includes five high-fidelity simulation environments: BeerFMTEnv, ReactorEnv, AtropineEnv, PenSimEnv and mAbEnv, which cover a wide range of manufacturing processes. We build these environments on published dynamics models. Furthermore, we benchmark online and offline, model-based and model-free reinforcement learning algorithms for comparisons of follow-up research.

cs.LG

Equilibrium Oil Market Share under the COVID-19 Pandemic

Equilibrium models for energy markets under uncertain demand and supply have attracted considerable attentions. This paper focuses on modelling crude oil market share under the COVID-19 pandemic using two-stage stochastic equilibrium. We describe the uncertainties in the demand and supply by random variables and provide two types of production decisions (here-and-now and wait-and-see). The here-and-now decision in the first stage does not depend on the outcome of random events to be revealed in the future and the wait-and-see decision in the second stage is allowed to depend on the random events in the future and adjust the feasibility of the here-and-now decision in rare unexpected scenarios such as those observed during the COVID-19 pandemic. We develop a fast algorithm to find a solution of the two-stage stochastic equilibrium. We show the robustness of the two-stage stochastic equilibrium model for forecasting the oil market share using the real market data from January 2019 to May 2020.

math.OC

Distributed Inference for Linear Support Vector Machine

The growing size of modern data brings many new challenges to existing statistical inference methodologies and theories, and calls for the development of distributed inferential approaches. This paper studies distributed inference for linear support vector machine (SVM) for the binary classification task. Despite a vast literature on SVM, much less is known about the inferential properties of SVM, especially in a distributed setting. In this paper, we propose a multi-round distributed linear-type (MDL) estimator for conducting inference for linear SVM. The proposed estimator is computationally efficient. In particular, it only requires an initial SVM estimator and then successively refines the estimator by solving simple weighted least squares problem. Theoretically, we establish the Bahadur representation of the estimator. Based on the representation, the asymptotic normality is further derived, which shows that the MDL estimator achieves the optimal statistical efficiency, i.e., the same efficiency as the classical linear SVM applying to the entire data set in a single machine setup. Moreover, our asymptotic result avoids the condition on the number of machines or data batches, which is commonly assumed in distributed estimation literature, and allows the case of diverging dimension. We provide simulation studies to demonstrate the performance of the proposed MDL estimator.

stat.ML

Regularized two-stage stochastic variational inequalities for Cournot-Nash equilibrium under uncertainty

A convex two-stage non-cooperative multi-agent game under uncertainty is formulated as a two-stage stochastic variational inequality (SVI). Under standard assumptions, we provide sufficient conditions for the existence of solutions of the two-stage SVI and propose a regularized sample average approximation method for solving it. We prove the convergence of the method as the regularization parameter tends to zero and the sample size tends to infinity. Moreover, our approach is applied to a two-stage stochastic production and supply planning problem with homogeneous commodity in an oligopolistic market. Numerical results based on historical data in crude oil market are presented to demonstrate the effectiveness of the two-stage SVI in describing the market share of oil producing agents.

math.OC