SearcharxivSearch

arXiv subjects

Roman Vershynin

Publications and source records attributed to Roman Vershynin.

At least 19 recordsLinked to original sources

Random sets are close to low-discrepancy sets

We show that a random sample from an arbitrary probability measure on $\mathbb{R}^d$ is close to a low-discrepancy point set. Namely, after moving only a small fraction of the sample points in expectation, one obtains an $n$-point set with star discrepancy $\operatorname{polylog}(n)/n$ with respect to the original measure.

math.PR

On the Subgaussianity of Quantized Linear Maps: An AI-Assisted Note

We prove an elementary bounded-differences inequality for functions of non-isotropic Gaussian vectors. Specifically, if $f$ has bounded coordinate differences and $X\sim\mathcal N(\mu,\Sigma)$, then the resulting concentration bound depends on the condition number $\kappa(\Sigma)$. As an application, we answer a question of Simone Bombari concerning the subgaussianity of sign-quantized linear maps $Y=\mathrm{sgn}(Wx)$. In the special case where $f$ is the coordinatewise sign function, an argument was initially suggested to us by Gemini 3.5 Flash without attribution. We subsequently discovered that it closely resembles an earlier argument of Barber and Kolar [Ann. Statist. 46 (2018), Lemma 4.5]. This revision corrects the attribution and documents the episode as an instance of AI-assisted mathematical discovery.

math.PR

Discrepancy and Fisher information

We give an online algorithm that keeps a symmetric random walk inside a convex body by discarding some of its steps. The expected number of discarded steps is controlled by a Fisher-information-type quantity associated with the body. For the cube, this gives a dimension-free bound: a walk with unit Euclidean steps can be kept bounded in all coordinates while discarding only a small constant fraction of the steps on average.

math.PR

Two friendly proofs of the Berry-Esseen theorem

A gem of classical probability, the Berry-Esseen theorem provides a non-asymptotic form of the central limit theorem. This note gives a friendly and intuitive exposition of two different proofs of the Berry-Esseen theorem for nonidentically distributed random variables: the classical Fourier-analytic proof and a proof by Stein's method following E. Bolthausen. Both proofs are self-contained; the reader may read either proof without reading the other. The exposition is suitable for a basic graduate course in probability.

math.PR

Thinning to improve two-sample discrepancy

The discrepancy between two independent samples \(X_1,\dots,X_n\) and \(Y_1,\dots,Y_n\) drawn from the same distribution on $\mathbb{R}^d$ typically has order \(O(\sqrt{n})\) even in one dimension. We give a simple online algorithm that reduces the discrepancy to \(O(\log^{2d} n)\) by discarding a small fraction of the points.

math.PR

LLM Watermarking Using Mixtures and Statistical-to-Computational Gaps

Given a text, can we determine whether it was generated by a large language model (LLM) or by a human? A widely studied approach to this problem is watermarking. We propose an undetectable and elementary watermarking scheme in the closed setting. Also, in the harder open setting, where the adversary has access to most of the model, we propose an unremovable watermarking scheme.

cs.CR

Improving discrepancy by moving a few points

We show how to improve the discrepancy of an iid sample by moving only a few points. Specifically, modifying \( O(m) \) sample points on average reduces the Kolmogorov-Smirnov distance to the population distribution to \(1/m\).

math.ST

Random matrices acting on sets: Independent columns

We study random matrices with independent subgaussian columns. Assuming each column has a fixed Euclidean norm, we establish conditions under which such matrices act as near-isometries when restricted to a given subset of their domain. We show that, with high probability, the maximum distortion caused by such a matrix is proportional to the Gaussian complexity of the subset, scaled by the subgaussian norm of the matrix columns. This linear dependence on the subgaussian norm is a new phenomenon, as random matrices with independent rows or independent entries typically exhibit superlinear dependence. As a consequence, normalizing the columns of random sparse matrices leads to stronger embedding guarantees.

math.PR

Jackson's inequality on the hypercube

We investigate the best constant $J(n,d)$ such that Jackson's inequality \[ \inf_{\mathrm{deg}(g) \leq d} \|f - g\|_{\infty} \leq J(n,d) \, s(f), \] holds for all functions $f$ on the hypercube $\{0,1\}^n$, where $s(f)$ denotes the sensitivity of $f$. We show that the quantity $J(n, 0.499n)$ is bounded below by an absolute positive constant, independent of $n$. This complements Wagner's theorem, which establishes that $J(n,d)\leq 1 $. As a first application we show that reverse Bernstein inequality fails in the tail space $L^{1}_{\geq 0.499n}$ improving over previously known counterexamples in $L^{1}_{\geq C \log \log (n)}$. As a second application, we show that there exists a function $f : \{0,1\}^n \to [-1,1]$ whose sensitivity $s(f)$ remains constant, independent of $n$, while the approximate degree grows linearly with $n$. This result implies that the sensitivity theorem $s(f) \geq \Omega(\mathrm{deg}(f)^C)$ fails in the strongest sense for bounded real-valued functions even when $\mathrm{deg}(f)$ is relaxed to the approximate degree. We also show that in the regime $d = (1 - \delta)n$, the bound \[ J(n,d) \leq C \min\{\delta, \max\{\delta^2, n^{-2/3}\}\} \] holds. Moreover, when restricted to symmetric real-valued functions, we obtain $J_{\mathrm{symmetric}}(n,d) \leq C/d$ and the decay $1/d$ is sharp. Finally, we present results for a subspace approximation problem: we show that there exists a subspace $E$ of dimension $2^{n-1}$ such that $\inf_{g \in E} \|f - g\|_{\infty} \leq s(f)/n$ holds for all $f$.

math.FA

Can we spot a fake?

The problem of detecting fake data inspires the following seemingly simple mathematical question. Sample a data point $X$ from the standard normal distribution in $\mathbb{R}^n$. An adversary observes $X$ and corrupts it by adding a vector $rt$, where they can choose any vector $t$ from a fixed set $T$ of the adversary's ``tricks'', and where $r>0$ is a fixed radius. The adversary's choice of $t=t(X)$ may depend on the true data $X$. The adversary wants to hide the corruption by making the fake data $X+rt$ statistically indistinguishable from the real data $X$. What is the largest radius $r=r(T)$ for which the adversary can create an undetectable fake? We show that for highly symmetric sets $T$, the detectability radius $r(T)$ is approximately twice the scaled Gaussian width of $T$. The upper bound actually holds for arbitrary sets $T$ and generalizes to arbitrary, non-Gaussian distributions of real data $X$. The lower bound may fail for not highly symmetric $T$, but we conjecture that this problem can be solved by considering the focused version of the Gaussian width of $T$, which focuses on the most important directions of $T$.

math.ST

Differentially Private Synthetic High-dimensional Tabular Stream

While differentially private synthetic data generation has been explored extensively in the literature, how to update this data in the future if the underlying private data changes is much less understood. We propose an algorithmic framework for streaming data that generates multiple synthetic datasets over time, tracking changes in the underlying private data. Our algorithm satisfies differential privacy for the entire input stream (continual differential privacy) and can be used for high-dimensional tabular data. Furthermore, we show the utility of our method via experiments on real-world datasets. The proposed algorithm builds upon a popular select, measure, fit, and iterate paradigm (used by offline synthetic data generation algorithms) and private counters for streams.

cs.CR

Metric geometry of the privacy-utility tradeoff

Synthetic data are an attractive concept to enable privacy in data sharing. A fundamental question is how similar the privacy-preserving synthetic data are compared to the true data. Using metric privacy, an effective generalization of differential privacy beyond the discrete setting, we raise the problem of characterizing the optimal privacy-accuracy tradeoff by the metric geometry of the underlying space. We provide a partial solution to this problem in terms of the "entropic scale", a quantity that captures the multiscale geometry of a metric space via the behavior of its packing numbers. We illustrate the applicability of our privacy-accuracy tradeoff framework via a diverse set of examples of metric spaces.

cs.CR

Online Differentially Private Synthetic Data Generation

We present a polynomial-time algorithm for online differentially private synthetic data generation. For a data stream within the hypercube $[0,1]^d$ and an infinite time horizon, we develop an online algorithm that generates a differentially private synthetic dataset at each time $t$. This algorithm achieves a near-optimal accuracy bound of $O(\log(t)t^{-1/d})$ for $d\geq 2$ and $O(\log^{4.5}(t)t^{-1})$ for $d=1$ in the 1-Wasserstein distance. This result extends the previous work on the continual release model for counting queries to Lipschitz queries. Compared to the offline case, where the entire dataset is available at once, our approach requires only an extra polylog factor in the accuracy bound.

math.ST

Hamiltonicity of Sparse Pseudorandom Graphs

We show that every $(n,d,\lambda)$-graph contains a Hamilton cycle for sufficiently large $n$, assuming that $d\geq \log^{6}n$ and $\lambda\leq cd$, where $c=\frac{1}{70000}$. This significantly improves a recent result of Glock, Correia and Sudakov, who obtained a similar result for $d$ that grows polynomially with $n$. The proof is based on a new result regarding the second largest eigenvalue of the adjacency matrix of a subgraph induced by a random subset of vertices, combined with a recent result on connecting designated pairs of vertices by vertex-disjoint paths in $(n,d,\lambda)$-graphs. We believe that the former result is of independent interest and will have further applications.

math.CO

An Algorithm for Streaming Differentially Private Data

Much of the research in differential privacy has focused on offline applications with the assumption that all data is available at once. When these algorithms are applied in practice to streams where data is collected over time, this either violates the privacy guarantees or results in poor utility. We derive an algorithm for differentially private synthetic streaming data generation, especially curated towards spatial datasets. Furthermore, we provide a general framework for online selective counting among a collection of queries which forms a basis for many tasks such as query answering and synthetic data generation. The utility of our algorithm is verified on both real-world and simulated datasets.

cs.DB

Are most Boolean functions determined by low frequencies?

We ask whether most Boolean functions are determined by their low frequencies. We show a partial result: for almost every function $f: \{-1,1\}^p \to \{-1,1\}$ there exists a function $f': \{-1,1\}^p \to (-1,1)$ that has the same frequencies as $f$ up to dimension $(1/2-o(1))p$.

math.CO

Covering the hypercube, the uncertainty principle, and an interpolation formula

We show that the minimal number of skewed hyperplanes that cover the hypercube $\{0,1\}^{n}$ is at least $\frac{n}{2}+1$, and there are infinitely many $n$'s when the hypercube can be covered with $n-\log_{2}(n)+1$ skewed hyperplanes. The minimal covering problems are closely related to uncertainty principle on the hypercube, where we also obtain an interpolation formula for multilinear polynomials on $\mathbb{R}^{n}$ of degree less than $\lfloor n/m \rfloor$ by showing that its coefficients corresponding to the largest monomials can be represented as a linear combination of values of the polynomial over the points $\{0,1\}^{n}$ whose hamming weights are divisible by $m$.

math.CO