SearcharxivSearch

arXiv subjects

Van H. Vu

Publications and source records attributed to Van H. Vu.

10 recordsLinked to original sources

Improved Guarantees for Heterogeneous Treatment-Effect Estimation via Matrix Completion

A central goal of modern causal inference is estimating heterogeneous treatment effects to answer questions like "how does an intervention affect each unit," rather than only on average. We study this problem with panel-data where we observe $n$ units across $m$ times under unknown, non-uniform treatment assignments. The data in this setting is naturally represented as a matrix of all unit--time treatment effects. Estimating heterogeneous treatment effects can then be expressed as obtaining a good estimation of each row's average in this matrix. This allows us to formulate the problem as matrix completion, which can be solved under natural low-rankness assumptions. However, existing matrix-completion guarantees are not powerful enough to get meaningful bounds for the per-row guarantee required for estimating the heterogeneous treatment effect; roughly speaking, they are only useful for estimating average treatment effect bounds, as also illustrated in a recent line of work. We give a simple, computationally efficient estimator that, without knowledge of the propensities and under standard low-rankness and regularity assumptions, achieves a row-wise $\ell_2$ error of $\tilde{O}(\sqrt{\frac{1}{n} + \frac{n}{m^2}})$. Technically, our analysis establishes the first sharp row-wise $\ell_2$-perturbation bound for low-rank approximation, complementing existing spectral-, Frobenius-, and entrywise perturbation theory.

stat.ML

The anti-concentration phenomenon with respect to random permutations

The anti-concentration phenomenon in probability theory has been intensively studied in recent years, with applications across many areas of mathematics. In most existing works, the ambient probability space is a product space generated by independent random variables. In this paper, we initiate a systematic study of anti-concentration when the ambient space is the symmetric group, equipped with the uniform measure. Concretely, we focus on the random sum $S_π = \sum_{i=1}^{n} w_i\, v_{π(i)}$, where $w=(w_1,\dots,w_n)$ and $v=(v_1,\dots,v_n)$ are fixed vectors and $π$ is a uniformly random permutation. The paper contains several new results, addressing both discrete and continuous anti-concentration phenomena. On the discrete side, we establish a near-optimal structural characterization of the vectors $w$ and $v$ under the assumption that the concentration probability $\sup_x P(S_π=x)$ is polynomially large. On the continuous side, we study the small-ball event $|S_π-L|\le δ$. Our results exhibit sub-gaussian decay in $L$. Our results have applications in various areas. First, we use our inverse theorems to derive and strengthen a number of previous anti-concentration bounds. In particular, we show that if both $w$ and $v$ have distinct entries, then $\sup_x P(S_π=x) \le n^{-5/2+o(1)}$. Next, we apply our new results to study random polynomials, and prove that the number of extremal points of random permutation polynomials is bounded by $O(\log n)$, extending results of S{ö}ze~\cite{Soze1, Soze2}. In the final application, we prove that random matrices whose rows are independent random permutations of a fixed non-degenerate vector are nonsingular with high probability.

math.CO

Spectral Perturbation Bounds for Low-Rank Approximation with Applications to Privacy

A central challenge in machine learning is to understand how noise or measurement errors affect low-rank approximations, particularly in the spectral norm. This question is especially important in differentially private low-rank approximation, where one aims to preserve the top-$p$ structure of a data-derived matrix while ensuring privacy. Prior work often analyzes Frobenius norm error or changes in reconstruction quality, but these metrics can over- or under-estimate true subspace distortion. The spectral norm, by contrast, captures worst-case directional error and provides the strongest utility guarantees. We establish new high-probability spectral-norm perturbation bounds for symmetric matrices that refine the classical Eckart--Young--Mirsky theorem and explicitly capture interactions between a matrix $A \in \mathbb{R}^{n \times n}$ and an arbitrary symmetric perturbation $E$. Under mild eigengap and norm conditions, our bounds yield sharp estimates for $\|(A + E)_p - A_p\|$, where $A_p$ is the best rank-$p$ approximation of $A$, with improvements of up to a factor of $\sqrt{n}$. As an application, we derive improved utility guarantees for differentially private PCA, resolving an open problem in the literature. Our analysis relies on a novel contour bootstrapping method from complex analysis and extends it to a broad class of spectral functionals, including polynomials and matrix exponentials. Empirical results on real-world datasets confirm that our bounds closely track the actual spectral error under diverse perturbation regimes.

cs.LG

Normal vector of a random hyperplane

Let v_1,...,v_{n-1} be n-1 independent vectors in R^n (or C^n). We study x, the unit normal vector of the hyperplane spanned by the v_i. Our main finding is that x resembles a random vector chosen uniformly from the unit sphere, under some randomness assumption on the v_i. Our result has applications in random matrix theory. Consider an n by n random matrix with iid entries. We first prove an exponential bound on the upper tail for the least singular value, improving the earlier linear bound by Rudelson and Vershynin. Next, we derive optimal delocalization for the eigenvectors corresponding to eigenvalues of small modulus.

math.PR

Non-abelian Littlewood-Offord inequalities

In 1943, Littlewood and Offord proved the first anti-concentration result for sums of independent random variables. Their result has since then been strengthened and generalized by generations of researchers, with applications in several areas of mathematics. In this paper, we present the first non-abelian analogue of Littlewood-Offord result, a sharp anti-concentration inequality for products of independent random variables.

math.PR

Small ball probability, Inverse theorems, and applications

Let $ξ$ be a real random variable with mean zero and variance one and $A={a_1,...,a_n}$ be a multi-set in $\R^d$. The random sum $$S_A := a_1 ξ_1 + ... + a_n ξ_n $$ where $ξ_i$ are iid copies of $ξ$ is of fundamental importance in probability and its applications. We discuss the small ball problem, the aim of which is to estimate the maximum probability that $S_A$ belongs to a ball with given small radius, following the discovery made by Littlewood-Offord and Erdos almost 70 years ago. We will mainly focus on recent developments that characterize the structure of those sets $A$ where the small ball probability is relatively large. Applications of these results include full solutions or significant progresses of many open problems in different areas.

math.CO

Mapping Incidences

We show that any finite set S in a characteristic zero integral domain can be mapped to the finite field of order p, for infinitely many primes p, preserving all algebraic incidences in S. This can be seen as a generalization of the well-known Freiman isomorphism lemma, and we give several combinatorial applications (such as sum-product estimates).

math.CO

Structure of large incomplete sets in abelian groups

Let $G$ be a finite abelian group and $A$ be a subset of $G$. We say that $A$ is complete if every element of $G$ can be represented as a sum of different elements of $A$. In this paper, we study the following question: {\it What is the structure of a large incomplete set ?} The typical answer is that such a set is essentially contained in a maximal subgroup. As a by-product, we obtain a new proof for several earlier results.

math.CO

The Rank of Random Graphs

We show that almost surely the rank of the adjacency matrix of the Erdös-Rényi random graph $G(n,p)$ equals the number of non-isolated vertices for any $c\ln n/n<p<1/2$, where $c$ is an arbitrary positive constant larger than 1/2. In particular, the giant component (a.s.) has full rank in this range.

math.PR

On the concentration of eigenvalues of random symmetric matrices

We prove that few largest (and most important) eigenvalues of random symmetric matrices of various kinds are very strongly concentrated. This strong concentration enables us to compute the means of these eigenvalues with high precision. Our approach uses Talagrand's inequality and is very different from standard approaches.

math-ph