SearcharxivSearch

arXiv subjects

Pedro Abdalla

Publications and source records attributed to Pedro Abdalla.

16 recordsLinked to original sources

Stable Phase Retrieval for Spans of Independent Random Variables

We prove that, after $L^2$ normalization, stable phase retrieval holds over the $L^2$-spans of independent real-valued centered random variables if and only if all but possibly one coordinate satisfies a uniform two-sided $L^1$ bound. This provides a complete characterization of stable phase retrieval for such subspaces, building upon the pioneering work of Calderbank--Daubechies--Freeman--Freeman and confirming the conjectured characterization communicated to us by those authors. We provide two different proofs of this fact, both based on a decomposition of the $\ell^2$-coefficients of each random variable. The first is a compactness proof, which makes use of the infinite divisibility of limit laws of tail sums. The second is a quantitative proof, which substitutes the compactness step with an explicit dichotomy based on anticoncentration estimates of Sperner type. This latter proof was partially LLM generated based on the ideas in the first proof and a considerable amount of guidance by the authors. An autoformalization of our main result in Lean 4 is also provided, following the ideas in the quantitative proof.

math.FA

Robust Uniform Recovery of Structured Signals from Nonlinear Observations

While it is well known that the restricted isometry property (RIP) guarantees uniform sparse recovery from noisy linear measurements, uniform recovery of structured signals from nonlinear observations remains much less understood. This paper shows that the restricted approximate invertibility condition (RAIC) provides a unified approach to this end. Particularly, uniform recovery is achieved by projected gradient descent (PGD) with gradients obeying RAIC for all signals. As an application, under a large class of piecewise Lipschitz link functions (possibly discontinuous), we develop a uniform recovery theory for Gaussian single-index model by establishing the uniform RAIC for the gradient of the (scaled) $\ell_2$ loss via a covering argument. The theory generalizes the nonuniform recovery guarantees due to Plan and Vershynin (2016); Oymak and Soltanolkotabi (2017) and exhibits additional error terms that can be interpreted as the cost of uniform recovery. Intriguingly, in the three canonical settings of (a) sparse recovery via PGD with $\ell_0$ projection (i.e., iterative hard thresholding (IHT)), (b) sparse recovery via PGD with $\ell_1$ projection, and (c) recovering approximately sparse signals via PGD with $\ell_1$ projection, the additional error terms are negligible and in turn our uniform recovery error rates are at the same order of existing nonuniform ones, up to log factors. Our results hence improve on Genzel and Stollenwerk (2023). Under the specific nonlinearity of 1-bit quantization, we use a VC dimension argument to show that the uniform recovery error of IHT is at the same order of the nonuniform recovery error, with no loss of log factor. In addition, we show that the robustness of PGD to noise and corruption can be incorporated elegantly by bounding a single additional random process that captures the gradient mismatch.

cs.IT

Expander graphs are globally synchronizing

The Kuramoto model is fundamental to the study of synchronization. It consists of a collection of oscillators with interactions given by a network, which we identify respectively with vertices and edges of a graph. In this paper, we show that a graph with sufficient expansion must be globally synchronizing, meaning that a homogeneous Kuramoto model of identical oscillators on such a graph will converge to the fully synchronized state with all the oscillators having the same phase, for every initial state up to a set of measure zero. In particular, we show that for any $\varepsilon > 0$ and $p \geq (1 + \varepsilon) (\log n) / n$, the homogeneous Kuramoto model on the Erdős-Rényi random graph $G(n, p)$ is globally synchronizing with probability tending to one as $n$ goes to infinity. This improves on a previous result of Kassabov, Strogatz, and Townsend and solves a conjecture of Ling, Xu, and Bandeira. We also show that the model is globally synchronizing on any $d$-regular Ramanujan graph, and on typical $d$-regular graphs, for large enough degree $d$.

math.CO

Robust Mean Estimation under Quantization

We consider the problem of mean estimation under quantization and adversarial corruption. We construct multivariate robust estimators that are optimal up to logarithmic factors in two different settings. The first is a one-bit setting, where each bit depends only on a single sample, and the second is a partial quantization setting, in which the estimator may use a small fraction of unquantized data.

stat.ML

LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps

Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an undetectable and elementary watermarking scheme in the closed setting. Also, in the harder open setting, where the adversary has access to most of the model, we propose an unremovable watermarking scheme.

cs.CR

On sharp stable recovery from clipped and folded measurements

We investigate the stability of vector recovery from random linear measurements which have been either clipped or folded. This is motivated by applications where measurement devices detect inputs outside of their effective range. As examples of our main results, we prove sharp lower bounds on the recovery constant for both the declipping and unfolding problems whenever samples are taken according to a uniform distribution on the sphere. Moreover, we show such estimates under (almost) the best possible conditions on both the number of samples and the distribution of the data. We then prove that all of the above results have suitable (effectively) sparse counterparts. In the special case that one restricts the stability analysis to vectors which belong to the unit sphere of $\mathbb{R}^n$, we show that the problem of declipping directly extends the one-bit compressed sensing results of Oymak-Recht and Plan-Vershynin.

cs.IT

Robust Sparse Recovery with Sparse Bernoulli matrices via Expanders

Sparse binary matrices are of great interest in the field of sparse recovery, nonnegative compressed sensing, statistics in networks, and theoretical computer science. This class of matrices makes it possible to perform signal recovery with lower storage costs and faster decoding algorithms. In particular, Bernoulli$(p)$ matrices formed by independent identically distributed (i.i.d.) Bernoulli$(p)$ random variables are of practical relevance in the context of noise-blind recovery in nonnegative compressed sensing. In this work, we investigate the robust nullspace property of Bernoulli$(p)$ matrices. Previous results in the literature establish that such matrices can accurately recover $n$-dimensional $s$-sparse vectors with $m=O\left(\frac{s}{c(p)}\log\frac{en}{s}\right)$ measurements, where $c(p) \le p$ is a constant dependent only on the parameter $p$. These results suggest that in the sparse regime, as $p$ approaches zero, the (sparse) Bernoulli$(p)$ matrix requires significantly more measurements than the minimal necessary, as achieved by standard isotropic subgaussian designs. However, we show that this is not the case. Our main result characterizes, for a wide range of sparsity levels $s$, the smallest $p$ for which sparse recovery can be achieved with the minimal number of measurements. We also provide matching lower bounds to establish the optimality of our results and explore connections with the theory of invertibility of discrete random matrices and integer compressed sensing.

cs.IT

Nonconvex landscapes for $\mathbf{Z}_2$ synchronization and graph clustering are benign near exact recovery thresholds

We study the optimization landscape of a smooth nonconvex program arising from synchronization over the two-element group $\mathbf{Z}_2$, that is, recovering $z_1, \dots, z_n \in \{\pm 1\}$ from (noisy) relative measurements $R_{ij} \approx z_i z_j$. Starting from a max-cut--like combinatorial problem, for integer parameter $r \geq 2$, the nonconvex problem we study can be viewed both as a rank-$r$ Burer--Monteiro factorization of the standard max-cut semidefinite relaxation and as a relaxation of $\{ \pm 1 \}$ to the unit sphere in $\mathbf{R}^r$. First, we present deterministic, non-asymptotic conditions on the measurement graph and noise under which every second-order critical point of the nonconvex problem yields exact recovery of the ground truth. Then, via probabilistic analysis, we obtain asymptotic guarantees for three benchmark problems: (1) synchronization with a complete graph and Gaussian noise, (2) synchronization with an Erdős--Rényi random graph and Bernoulli noise, and (3) graph clustering under the binary symmetric stochastic block model. In each case, we have, asymptotically as the problem size goes to infinity, a benign nonconvex landscape near a previously-established optimal threshold for exact recovery; we can approach this threshold to arbitrary precision with large enough (but finite) rank parameter $r$. In addition, our results are robust to monotone adversaries.

math.OC

Covariance Estimation under Missing Observations and $L_4-L_2$ Moment Equivalence

We consider the problem of estimating the covariance matrix of a random vector by observing i.i.d samples and each entry of the sampled vector is missed with probability $p$. Under the standard $L_4-L_2$ moment equivalence assumption, we construct the first estimator that simultaneously achieves optimality with respect to the parameter $p$ and it recovers the optimal convergence rate for the classical covariance estimation problem when $p=1$

math.ST

Debiased LASSO under Poisson-Gauss Model

Quantifying uncertainty in high-dimensional sparse linear regression is a fundamental task in statistics that arises in various applications. One of the most successful methods for quantifying uncertainty is the debiased LASSO, which has a solid theoretical foundation but is restricted to settings where the noise is purely additive. Motivated by real-world applications, we study the so-called Poisson inverse problem with additive Gaussian noise and propose a debiased LASSO algorithm that only requires $n \gg s\log^2p$ samples, which is optimal up to a logarithmic factor.

math.ST

Guarantees for Spontaneous Synchronization on Random Geometric Graphs

The Kuramoto model is a classical mathematical model in the field of non-linear dynamical systems that describes the evolution of coupled oscillators in a network that may reach a synchronous state. The relationship between the network's topology and whether the oscillators synchronize is a central question in the field of synchronization, and random graphs are often employed as a proxy for complex networks. On the other hand, the random graphs on which the Kuramoto model is rigorously analyzed in the literature are homogeneous models and fail to capture the underlying geometric structure that appears in several examples. In this work, we leverage tools from random matrix theory, random graphs, and mathematical statistics to prove that the Kuramoto model on a random geometric graph on the sphere synchronizes with probability tending to one as the number of nodes tends to infinity. To the best of our knowledge, this is the first rigorous result for the Kuramoto model on random geometric graphs.

math.PR

Covariance estimation with direction dependence accuracy

We construct an estimator $\widehatΣ$ for covariance matrices of unknown, centred random vectors X, with the given data consisting of N independent measurements $X_1,...,X_N$ of X and the wanted confidence level. We show under minimal assumptions on X, the estimator performs with the optimal accuracy with respect to the operator norm. In addition, the estimator is also optimal with respect to direction dependence accuracy: $\langle \widehatΣu,u\rangle$ is an optimal estimator for $σ^2(u)=\mathbb{E}\langle X,u\rangle^2$ when $σ^2(u)$ is ``large".

math.ST

Covariance Estimation: Optimal Dimension-free Guarantees for Adversarial Corruption and Heavy Tails

We provide an estimator of the covariance matrix that achieves the optimal rate of convergence (up to constant factors) in the operator norm under two standard notions of data contamination: We allow the adversary to corrupt an $η$-fraction of the sample arbitrarily, while the distribution of the remaining data points only satisfies that the $L_{p}$-marginal moment with some $p \ge 4$ is equivalent to the corresponding $L_2$-marginal moment. Despite requiring the existence of only a few moments, our estimator achieves the same tail estimates as if the underlying distribution were Gaussian. As a part of our analysis, we prove a dimension-free Bai-Yin type theorem in the regime $p > 4$.

math.ST

Community Detection with a Subsampled Semidefinite Program

Semidefinite programming is an important tool to tackle several problems in data science and signal processing, including clustering and community detection. However, semidefinite programs are often slow in practice, so speed up techniques such as sketching are often considered. In the context of community detection in the stochastic block model, Mixon and Xie \cite{mixon2020sketching} have recently proposed a sketching framework in which a semidefinite program is solved only on a subsampled subgraph of the network, giving rise to significant computational savings. In this short paper, we provide a positive answer to a conjecture of Mixon and Xie about the statistical limits of this technique for the stochastic block model with two balanced communities.

math.OC

Dictionary-Sparse Recovery From Heavy-Tailed Measurements

The recovery of signals that are sparse not in a basis, but rather sparse with respect to an over-complete dictionary is one of the most flexible settings in the field of compressed sensing with numerous applications. As in the standard compressed sensing setting, it is possible that the signal can be reconstructed efficiently from few, linear measurements, for example by the so-called $\ell_1$-synthesis method. However, it has been less well-understood which measurement matrices provably work for this setting. Whereas in the standard setting, it has been shown that even certain heavy-tailed measurement matrices can be used in the same sample complexity regime as Gaussian matrices, comparable results are only available for the restrictive class of sub-Gaussian measurement vectors as far as the recovery of dictionary-sparse signals via $\ell_1$-synthesis is concerned. In this work, we fill this gap and establish optimal guarantees for the recovery of vectors that are (approximately) sparse with respect to a dictionary via the $\ell_1$-synthesis method from linear, potentially noisy measurements for a large class of random measurement matrices. In particular, we show that random measurements that fulfill only a small-ball assumption and a weak moment assumption, such as random vectors with i.i.d. Student-$t$ entries with a logarithmic number of degrees of freedom, lead to comparable guarantees as (sub-)Gaussian measurements. As a technical tool, we show a bound on the expectation of the sum of squared order statistics under very general assumptions, which might be of independent interest. As a corollary of our results, we also obtain a slight improvement on the weakest assumption on a measurement matrix with i.i.d. rows sufficient for uniform recovery in standard compressed sensing, improving on results by Lecué and Mendelson and Dirksen, Lecué and Rauhut.

cs.IT