Searcharxiv⌕ Search

arXiv subjects

Walid Hachem

Publications and source records attributed to Walid Hachem.

At least 37 records · Page 2Linked to original sources

A constant step Forward-Backward algorithm involving random maximal monotone operators

A stochastic Forward-Backward algorithm with a constant step is studied. At each time step, this algorithm involves an independent copy of a couple of random maximal monotone operators. Defining a mean operator as a selection integral, the differential inclusion built from the sum of the two mean operators is considered. As a first result, it is shown that the interpolated process obtained from the iterates converges narrowly in the small step regime to the solution of this differential inclusion. In order to control the long term behavior of the iterates, a stability result is needed in addition. To this end, the sequence of the iterates is seen as a homogeneous Feller Markov chain whose transition kernel is parameterized by the algorithm step size. The cluster points of the Markov chains invariant measures in the small step regime are invariant for the semiflow induced by the differential inclusion. Conclusions regarding the long run behavior of the iterates for small steps are drawn. It is shown that when the sum of the mean operators is demipositive, the probabilities that the iterates are away from the set of zeros of this sum are small in Cesàro mean. The ergodic behavior of these iterates is studied as well. Applications of the proposed algorithm are considered. In particular, a detailed analysis of the random proximal gradient algorithm with constant step is performed.

math.OC↗

A Constant Step Stochastic Douglas-Rachford Algorithm with Application to Non Separable Regularizations

The Douglas Rachford algorithm is an algorithm that converges to a minimizer of a sum of two convex functions. The algorithm consists in fixed point iterations involving computations of the proximity operators of the two functions separately. The paper investigates a stochastic version of the algorithm where both functions are random and the step size is constant. We establish that the iterates of the algorithm stay close to the set of solution with high probability when the step size is small enough. Application to structured regularization is considered.

math.OC↗

Snake: a Stochastic Proximal Gradient Algorithm for Regularized Problems over Large Graphs

A regularized optimization problem over a large unstructured graph is studied, where the regularization term is tied to the graph geometry. Typical regularization examples include the total variation and the Laplacian regularizations over the graph. When applying the proximal gradient algorithm to solve this problem, there exist quite affordable methods to implement the proximity operator (backward step) in the special case where the graph is a simple path without loops. In this paper, an algorithm, referred to as "Snake", is proposed to solve such regularized problems over general graphs, by taking benefit of these fast methods. The algorithm consists in properly selecting random simple paths in the graph and performing the proximal gradient algorithm over these simple paths. This algorithm is an instance of a new general stochastic proximal gradient algorithm, whose convergence is proven. Applications to trend filtering and graph inpainting are provided among others. Numerical experiments are conducted over large graphs.

math.OC↗

Constant Step Stochastic Approximations Involving Differential Inclusions: Stability, Long-Run Convergence and Applications

We consider a Markov chain $(x_n)$ whose kernel is indexed by a scaling parameter $γ>0$, refered to as the step size. The aim is to analyze the behavior of the Markov chain in the doubly asymptotic regime where $n\to\infty$ then $γ\to 0$. First, under mild assumptions on the so-called drift of the Markov chain, we show that the interpolated process converges narrowly to the solutions of a Differential Inclusion (DI) involving an upper semicontinuous set-valued map with closed and convex values. Second, we provide verifiable conditions which ensure the stability of the iterates. Third, by putting the above results together, we establish the long run convergence of the iterates as $γ\to 0$, to the Birkhoff center of the DI. The ergodic behavior of the iterates is also provided. Application examples are investigated. We apply our findings to 1) the problem of nonconvex proximal stochastic optimization and 2) a fluid model of parallel queues.

math.PR↗

Dynamical behavior of a stochastic forward-backward algorithm using random monotone operators

The purpose of this paper is to study the dynamical behavior of the sequence produced by a forward-backward algorithm involving two random maximal monotone operators and a sequence of decreasing step sizes. Defining a mean monotone operator as an Aumann integral, and assuming that the sum of the two mean operators is maximal (sufficient maximality conditions are provided), it is shown that with probability one, the interpolated process obtained from the iterates is an asymptotic pseudo trajectory in the sense of Bena\"ım and Hirsch of the differential inclusion involving the sum of the mean operators. The convergence of the empirical means of the iterates towards a zero of the sum of the mean operators is shown, as well as the convergence of the sequence itself to such a zero under a demipositivity assumption. These results find applications in a wide range of optimization or variational inequality problems in random environments.

math.OC↗

Large complex correlated Wishart matrices: Fluctuations and asymptotic independence at the edges

We study the asymptotic behavior of eigenvalues of large complex correlated Wishart matrices at the edges of the limiting spectrum. In this setting, the support of the limiting eigenvalue distribution may have several connected components. Under mild conditions for the population matrices, we show that for every generic positive edge of that support, there exists an extremal eigenvalue which converges almost surely toward that edge and fluctuates according to the Tracy-Widom law at the scale $N^{2/3}$. Moreover, given several generic positive edges, we establish that the associated extremal eigenvalue fluctuations are asymptotically independent. Finally, when the leftmost edge is the origin (hard edge), the fluctuations of the smallest eigenvalue are described by mean of the Bessel kernel at the scale $N^2$.

math.PR↗

Large Complex Correlated Wishart Matrices: The Pearcey Kernel and Expansion at the Hard Edge

We study the eigenvalue behaviour of large complex correlated Wishart matrices near an interior point of the limiting spectrum where the density vanishes (cusp point), and refine the existing results at the hard edge as well. More precisely, under mild assumptions for the population covariance matrix, we show that the limiting density vanishes at generic cusp points like a cube root, and that the local eigenvalue behaviour is described by means of the Pearcey kernel if an extra decay assumption is satisfied. As for the hard edge, we show that the density blows up like an inverse square root at the origin. Moreover, we provide an explicit formula for the $1/N$ correction term for the fluctuation of the smallest random eigenvalue.

math.PR↗

A Coordinate Descent Primal-Dual Algorithm and Application to Distributed Asynchronous Optimization

Based on the idea of randomized coordinate descent of $α$-averaged operators, a randomized primal-dual optimization algorithm is introduced, where a random subset of coordinates is updated at each iteration. The algorithm builds upon a variant of a recent (deterministic) algorithm proposed by Vũ and Condat that includes the well known ADMM as a particular case. The obtained algorithm is used to solve asynchronously a distributed optimization problem. A network of agents, each having a separate cost function containing a differentiable term, seek to find a consensus on the minimum of the aggregate objective. The method yields an algorithm where at each iteration, a random subset of agents wake up, update their local estimates, exchange some data with their neighbors, and go idle. Numerical results demonstrate the attractive performance of the method. The general approach can be naturally adapted to other situations where coordinate descent convex optimization algorithms are used with a random choice of the coordinates.

math.OC↗

A Survey on the Eigenvalues Local Behavior of Large Complex Correlated Wishart Matrices

The aim of this note is to provide a pedagogical survey of the recent works by the authors ( arXiv:1409.7548 and arXiv:1507.06013) concerning the local behavior of the eigenvalues of large complex correlated Wishart matrices at the edges and cusp points of the spectrum: Under quite general conditions, the eigenvalues fluctuations at a soft edge of the limiting spectrum, at the hard edge when it is present, or at a cusp point, are respectively described by mean of the Airy kernel, the Bessel kernel, or the Pearcey kernel. Moreover, the eigenvalues fluctuations at several soft edges are asymptotically independent. In particular, the asymptotic fluctuations of the matrix condition number can be described. Finally, the next order term of the hard edge asymptotics is provided.

math.PR↗

The Shannon's mutual information of a multiple antenna time and frequency dependent channel: an ergodic operator approach

Consider a random non-centered multiple antenna radio transmission channel. Assume that the deterministic part of the channel is itself frequency selective, and that the random multipath part is represented by an ergodic stationary vector process. In the Hilbert space $l^2({\mathbb Z})$, one can associate to this channel a random ergodic self-adjoint operator having a so-called Integrated Density of States (IDS). Shannon's mutual information per receive antenna of this channel coincides then with the integral of a $\log$ function with respect to the IDS. In this paper, it is shown that when the numbers of antennas at the transmitter and at the receiver tend to infinity at the same rate, the mutual information per receive antenna tends to a quantity that can be identified and, in fact, is closely related to that obtained within the random matrix approach. This result can be obtained by analyzing the behavior of the Stieltjes transform of the IDS in the regime of the large numbers of antennas.

cs.IT↗

Analysis of the limiting spectral measure of large random matrices of the separable covariance type

Consider the random matrix $Σ= D^{1/2} X \widetilde D^{1/2}$ where $D$ and $\widetilde D$ are deterministic Hermitian nonnegative matrices with respective dimensions $N \times N$ and $n \times n$, and where $X$ is a random matrix with independent and identically distributed centered elements with variance $1/n$. Assume that the dimensions $N$ and $n$ grow to infinity at the same pace, and that the spectral measures of $D$ and $\widetilde D$ converge as $N,n \to\infty$ towards two probability measures. Then it is known that the spectral measure of $ΣΣ^*$ converges towards a probability measure $μ$ characterized by its Stieltjes Transform. In this paper, it is shown that $μ$ has a density away from zero, this density is analytical wherever it is positive, and it behaves in most cases as $\sqrt{|x - a|}$ near an edge $a$ of its support. A complete characterization of the support of $μ$ is also provided. \\ Beside its mathematical interest, this analysis finds applications in a certain class of statistical estimation problems.

math.PR↗

Explicit Convergence Rate of a Distributed Alternating Direction Method of Multipliers

Consider a set of N agents seeking to solve distributively the minimization problem $\inf_{x} \sum_{n = 1}^N f_n(x)$ where the convex functions $f_n$ are local to the agents. The popular Alternating Direction Method of Multipliers has the potential to handle distributed optimization problems of this kind. We provide a general reformulation of the problem and obtain a class of distributed algorithms which encompass various network architectures. The rate of convergence of our method is considered. It is assumed that the infimum of the problem is reached at a point $x_\star$, the functions $f_n$ are twice differentiable at this point and $\sum \nabla^2 f_n(x_\star) > 0$ in the positive definite ordering of symmetric matrices. With these assumptions, it is shown that the convergence to the consensus $x_\star$ is linear and the exact rate is provided. Application examples where this rate can be optimized with respect to the ADMM free parameter $ρ$ are also given.

cs.DC↗

Statistical Inference in Large Antenna Arrays under Unknown Noise Pattern

In this article, a general information-plus-noise transmission model is assumed, the receiver end of which is composed of a large number of sensors and is unaware of the noise pattern. For this model, and under reasonable assumptions, a set of results is provided for the receiver to perform statistical eigen-inference on the information part. In particular, we introduce new methods for the detection, counting, and the power and subspace estimation of multiple sources composing the information part of the transmission. The theoretical performance of some of these techniques is also discussed. An exemplary application of these methods to array processing is then studied in greater detail, leading in particular to a novel MUSIC-like algorithm assuming unknown noise covariance.

cs.IT↗

Estimation of Toeplitz Covariance Matrices in Large Dimensional Regime with Application to Source Detection

In this article, we derive concentration inequalities for the spectral norm of two classical sample estimators of large dimensional Toeplitz covariance matrices, demonstrating in particular their asymptotic almost sure consistence. The consistency is then extended to the case where the aggregated matrix of time samples is corrupted by a rank one (or more generally, low rank) matrix. As an application of the latter, the problem of source detection in the context of large dimensional sensor networks within a temporally correlated noise environment is studied. As opposed to standard procedures, this application is performed online, i.e. without the need to possess a learning set of pure noise samples.

cs.IT↗

Performance of a Distributed Stochastic Approximation Algorithm

In this paper, a distributed stochastic approximation algorithm is studied. Applications of such algorithms include decentralized estimation, optimization, control or computing. The algorithm consists in two steps: a local step, where each node in a network updates a local estimate using a stochastic approximation algorithm with decreasing step size, and a gossip step, where a node computes a local weighted average between its estimates and those of its neighbors. Convergence of the estimates toward a consensus is established under weak assumptions. The approach relies on two main ingredients: the existence of a Lyapunov function for the mean field in the agreement subspace, and a contraction property of the random matrices of weights in the subspace orthogonal to the agreement subspace. A second order analysis of the algorithm is also performed under the form of a Central Limit Theorem. The Polyak-averaged version of the algorithm is also considered.

math.OC↗

Asynchronous Distributed Optimization using a Randomized Alternating Direction Method of Multipliers

Consider a set of networked agents endowed with private cost functions and seeking to find a consensus on the minimizer of the aggregate cost. A new class of random asynchronous distributed optimization methods is introduced. The methods generalize the standard Alternating Direction Method of Multipliers (ADMM) to an asynchronous setting where isolated components of the network are activated in an uncoordinated fashion. The algorithms rely on the introduction of randomized Gauss-Seidel iterations of a Douglas-Rachford operator for finding zeros of a sum of two monotone operators. Convergence to the sought minimizers is provided under mild connectivity conditions. Numerical results sustain our claims.

cs.DC↗

The outliers among the singular values of large rectangular random matrices with additive fixed rank deformation

Consider the matrix $Σ_n = n^{-1/2} X_n D_n^{1/2} + P_n$ where the matrix $X_n \in \C^{N\times n}$ has Gaussian standard independent elements, $D_n$ is a deterministic diagonal nonnegative matrix, and $P_n$ is a deterministic matrix with fixed rank. Under some known conditions, the spectral measures of $Σ_n Σ_n^*$ and $n^{-1} X_n D_n X_n^*$ both converge towards a compactly supported probability measure $μ$ as $N,n\to\infty$ with $N/n\to c>0$. In this paper, it is proved that finitely many eigenvalues of $Σ_nΣ_n^*$ may stay away from the support of $μ$ in the large dimensional regime. The existence and locations of these outliers in any connected component of $\R - \support(μ)$ are studied. The fluctuations of the largest outliers of $Σ_nΣ_n^*$ are also analyzed. The results find applications in the fields of signal processing and radio communications.

math.PR↗

Analysis of Sum-Weight-like algorithms for averaging in Wireless Sensor Networks

Distributed estimation of the average value over a Wireless Sensor Network has recently received a lot of attention. Most papers consider single variable sensors and communications with feedback (e.g. peer-to-peer communications). However, in order to use efficiently the broadcast nature of the wireless channel, communications without feedback are advocated. To ensure the convergence in this feedback-free case, the recently-introduced Sum-Weight-like algorithms which rely on two variables at each sensor are a promising solution. In this paper, the convergence towards the consensus over the average of the initial values is analyzed in depth. Furthermore, it is shown that the squared error decreases exponentially with the time. In addition, a powerful algorithm relying on the Sum-Weight structure and taking into account the broadcast nature of the channel is proposed.

cs.DC↗