SearcharxivSearch

arXiv subjects

Dan Mikulincer

Publications and source records attributed to Dan Mikulincer.

At least 19 recordsLinked to original sources

Marchenko-Pastur law for tensor powers of exchangeable unconditional vectors

Given an isotropic, exchangeable, and unconditional random vector $\mathbf X$, we consider the sample covariance matrix constructed from i.i.d. copies of several tensor models of $\mathbf X$, such as the tensor power $\mathbf{X}^{\otimes d}$. Under appropriate moment conditions on $\mathbf X$, we show that almost surely, the empirical spectral distribution converges weakly to the Marchenko-Pastur law. This extends previous results which required the coordinates of $\mathbf X$ to be independent. As we demonstrate, our extension applies to many new random vectors $\mathbf X$ of interest.

math.PR

Geometric obstructions to Lipschitz transport between weighted Hessian $\mathrm{CD}(\kappa,\infty)$ manifolds

We construct a weighted Riemannian manifold $(\mathbb R^2,g,\mu)$ satisfying $\mathrm{CD}(1/2,\infty)$, the curvature-dimension condition, with the following property: if $\gamma$ denotes a centered Gaussian measure on $\mathbb R^2$, then there is no Lipschitz map $T:(\mathbb R^2,\|\cdot\|) \to (\mathbb R^2,g)$ satisfying $T_\#\gamma=\mu$. Building on this, we prove a Weyl-type asymptotic law for the eigenvalues of the weighted Laplacian $-\Delta_{g,\mu}$ and show that they are asymptotically negligible when compared to the eigenvalues of $-\Delta_{\gamma}$. These results give strong counterexamples to two questions of E. Milman and complement the recent counterexample of Aryan.

math.PR

The Benefits of Temporal Correlations: SGD Learns k-Juntas from Random Walks Efficiently

We study how temporal correlations in the data can make certain sparse learning problems efficiently learnable by gradient-based methods. Our focus is on Boolean k-juntas, a canonical sparse learning problem known to pose barriers for gradient-based methods under independent uniform samples. We show that this picture changes when the samples are generated by a lazy random walk on the hypercube. In this setting, the temporal dependencies can be exploited by a two-layer ReLU network trained using stylized-SGD with a temporal-difference loss, which compares target and predicted increments across consecutive samples. For every fixed k, the resulting sample complexity is essentially linear in the ambient dimension d. By contrast, we show that for large-batch gradient methods using standard convex pointwise losses, temporal correlations do not provide the same advantage.

cs.LG

Dimension-free Gaussian tail estimates for linear functionals on convex bodies

Let $K \subset \mathbb{R}^n$ be a centered convex body of volume one. We prove that there exist absolute constants $c,C > 0$ and an orthonormal set of vectors $Θ\subset S^{n-1}$ with size $\left|Θ\right| \ge 9n/10$ such that, if $X$ is a random vector uniformly distributed on $K$, then for all $θ\in Θ$ one has \[ c\cdot \sqrt{p}\,\left(\mathbb{E} \left|\left\langle X,θ\right\rangle\right|^2\right)^{1/2} \le \left(\mathbb{E} \left|\left\langle X,θ\right\rangle\right|^p\right)^{1/p} \le C\cdot \sqrt{p}\,\left(\mathbb{E} \left|\left\langle X,θ\right\rangle\right|^2\right)^{1/2}, \] where the upper estimate holds for all $p \ge 1$ while the lower bound only holds for $1 \le p \le n$.

math.MG

Fast mixing in Ising models with a negative spectral outlier via Gaussian approximation

We study the mixing time of Glauber dynamics for Ising models in which the interaction matrix contains a single negative spectral outlier. This class includes the anti-ferromagnetic Curie-Weiss model, the anti-ferromagnetic Ising model on expander graphs, and the Sherrington-Kirkpatrick model with disorder of negative mean. Existing approaches to rapid mixing rely crucially on log-concavity or spectral width bounds and therefore can break down in the presence of a negative outlier. To address this difficulty, we develop a new covariance approximation method based on Gaussian approximation. This method is implemented via an iterative application of Stein's method to quadratic tilts of sums of bounded random variables, which may be of independent interest. The resulting analysis provides an operator-norm control of the full correlation structure under arbitrary external fields. Combined with the localization schemes of Eldan and Chen, these estimates lead to a modified logarithmic Sobolev inequality and near-optimal mixing time bounds in regimes where spectral width bounds fail. We complement these results by proving exponential lower bounds on the mixing time for low temperature anti-ferromagnetic Ising models on sparse random regular graphs and Erdös-Rényi graphs, based on the existence of gapped states as in the recent work of Sellke.

math.PR

Anti-concentration of polynomials: $L^{p}$ balls and symmetric measures

We begin with the observation, based on previous results, that dimension-free lower bounds on the variance of a polynomial under a log-concave measure yield dimension-free small-ball and Fourier decay estimates. Motivated by this, we establish variance bounds for polynomials on log-concave random vectors beyond the classical setting of product measures. First, we consider the family of uniform measures on the $n$-dimensional isotropic $L^{p}$ balls. We show that for a degree-$d$ homogeneous polynomial $f=\sum_{I}a_{I}x^{I}$, with $\sum_{I}a_{I}^{2}=1$, the only obstruction to a dimension-free lower bound on its variance occurs when $p=d$ is an even integer and the coefficients of $f$ are close to those of $\frac{1}{\sqrt{n}}\left\Vert x\right\Vert _{p}^{p}$. Second, we consider general isotropic log-concave measures that are invariant under coordinate permutations and reflections, and determine the minimal variance for quadratic and cubic polynomials. These variance bounds lead to new dimension-free anti-concentration results in both settings, addressing a natural extension of a question posed by Carbery and Wright.

math.PR

Nodal Count for Orthogonally Invariant Ensembles

We investigate the nodal count of eigenvectors of random matrices interpreted as operators on signed complete graphs. Our focus is on orthogonally invariant ensembles, with particular attention to the Gaussian Orthogonal Ensemble (GOE). We establish that, as the matrix size tends to infinity, the distribution of nodal counts converges to the same limiting law as the eigenvalue distribution. In the GOE case, this limit is the semicircle law. This result refutes a conjecture, motivated by quantum chaos and quantum graphs, which predicted Gaussian behavior of the nodal count.

math-ph

Low-dimensional Functions are Efficiently Learnable under Randomly Biased Distributions

The problem of learning single index and multi index models has gained significant interest as a fundamental task in high-dimensional statistics. Many recent works have analysed gradient-based methods, particularly in the setting of isotropic data distributions, often in the context of neural network training. Such studies have uncovered precise characterisations of algorithmic sample complexity in terms of certain analytic properties of the target function, such as the leap, information, and generative exponents. These properties establish a quantitative separation between low and high complexity learning tasks. In this work, we show that high complexity cases are rare. Specifically, we prove that introducing a small random perturbation to the data distribution--via a random shift in the first moment--renders any Gaussian single index model as easy to learn as a linear function. We further extend this result to a class of multi index models, namely sparse Boolean functions, also known as Juntas.

cs.LG

Stochastic Localization with Non-Gaussian Tilts and Applications to Tensor Ising Models

We present generalizations and modifications of Eldan's Stochastic Localization process, extending it to incorporate non-Gaussian tilts, making it useful for a broader class of measures. As an application, we introduce new processes that enable the decomposition and analysis of non-quadratic potentials on the Boolean hypercube, with a specific focus on quartic polynomials. Using this framework, we derive new spectral gap estimates for tensor Ising models under Glauber dynamics, resulting in rapid mixing.

math.PR

Large deviations principle for sub-Riemannian random walks

We study large deviations for random walks on stratified (Carnot) Lie groups. For such groups, there is a natural collection of vectors which generates their Lie algebra, and we consider random walks with increments in only these directions. Under certain constraints on the distribution of the increments, we prove a large deviation principle for these random walks with a natural rate function adapted to the sub-Riemannian geometry of these spaces.

math.PR

Stochastic proof of the sharp symmetrized Talagrand inequality

We give a new proof of the sharp symmetrized form of Talagrand's transport-entropy inequality. Compared to stochastic proofs of other Gaussian functional inequalities, the new idea here is a certain coupling induced by time-reversed martingale representations.

math.PR

Is this correct? Let's check!

Societal accumulation of knowledge is a complex process. The correctness of new units of knowledge depends not only on the correctness of new reasoning, but also on the correctness of old units that the new one builds on. The errors in such accumulation processes are often remedied by error correction and detection heuristics. Motivating examples include the scientific process based on scientific publications, and software development based on libraries of code. Natural processes that aim to keep errors under control, such as peer review in scientific publications, and testing and debugging in software development, would typically check existing pieces of knowledge -- both for the reasoning that generated them and the previous facts they rely on. In this work, we present a simple process that models such accumulation of knowledge and study the persistence (or lack thereof) of errors. We consider a simple probabilistic model for the generation of new units of knowledge based on the preferential attachment growth model, which additionally allows for errors. Furthermore, the process includes checks aimed at catching these errors. We investigate when effects of errors persist forever in the system (with positive probability) and when they get rooted out completely by the checking process. The two basic parameters associated with the checking process are the {\em probability} of conducting a check and the depth of the check. We show that errors are rooted out if checks are sufficiently frequent and sufficiently deep. In contrast, shallow or infrequent checks are insufficient to root out errors.

cs.SI

Size and depth of monotone neural networks: interpolation and approximation

We study monotone neural networks with threshold gates where all the weights (other than the biases) are non-negative. We focus on the expressive power and efficiency of representation of such networks. Our first result establishes that every monotone function over $[0,1]^d$ can be approximated within arbitrarily small additive error by a depth-4 monotone network. When $d > 3$, we improve upon the previous best-known construction which has depth $d+1$. Our proof goes by solving the monotone interpolation problem for monotone datasets using a depth-4 monotone threshold network. In our second main result we compare size bounds between monotone and arbitrary neural networks with threshold gates. We find that there are monotone real functions that can be computed efficiently by networks with no restriction on the gates whereas monotone networks approximating these functions need exponential size in the dimension.

cs.LG

Characterizing the fourth-moment phenomenon of monochromatic subgraph counts via influences

We investigate the distribution of monochromatic subgraph counts in random vertex $2$-colorings of large graphs. We give sufficient conditions for the asymptotic normality of these counts and demonstrate their essential necessity (particularly for monochromatic triangles). Our approach refines the fourth-moment theorem to establish new, local influence-based conditions for asymptotic normality; these findings more generally provide insight into fourth-moment phenomena for a broader class of Rademacher and Gaussian polynomials.

math.PR

Time Lower Bounds for the Metropolis Process and Simulated Annealing

The Metropolis process (MP) and Simulated Annealing (SA) are stochastic local search heuristics that are often used in solving combinatorial optimization problems. Despite significant interest, there are very few theoretical results regarding the quality of approximation obtained by MP and SA (with polynomially many iterations) for NP-hard optimization problems. We provide rigorous lower bounds for MP and SA with respect to the classical maximum independent set problem when the algorithms are initialized from the empty set. We establish the existence of a family of graphs for which both MP and SA fail to find approximate solutions in polynomial time. More specifically, we show that for any $\varepsilon \in (0,1)$ there are $n$-vertex graphs for which the probability SA (when limited to polynomially many iterations) will approximate the optimal solution within ratio $Ω\left(\frac{1}{n^{1-\varepsilon}}\right)$ is exponentially small. Our lower bounds extend to graphs of constant average degree $d$, illustrating the failure of MP to achieve an approximation ratio of $Ω\left(\frac{\log (d)}{d}\right)$ in polynomial time. In some cases, our impossibility results also go beyond Simulated Annealing and apply even when the temperature is chosen adaptively. Finally, we prove time lower bounds when the inputs to these algorithms are bipartite graphs, and even trees, which are known to admit polynomial-time algorithms for the independent set problem.

cs.DS

Transportation onto log-Lipschitz perturbations

We establish sufficient conditions for the existence of globally Lipschitz transport maps between probability measures and their log-Lipschitz perturbations, with dimension-free bounds. Our results include Gaussian measures on Euclidean spaces and uniform measures on spheres as source measures. More generally, we prove results for source measures on manifolds satisfying strong curvature assumptions. These seem to be the first examples of dimension-free Lipschitz transport maps in non-Euclidean settings, which are moreover sharp on the sphere. We also present some applications to functional inequalities, including a new dimension-free Gaussian isoperimetric inequality for log-Lipschitz perturbations of the standard Gaussian measure. Our proofs are based on the Langevin flow construction of transport maps of Kim and Milman.

math.PR

Integrality Gaps for Random Integer Programs via Discrepancy

We prove new bounds on the additive gap between the value of a random integer program $\max c^Tx,\ Ax\leq b,\ x\in\{0,1\}^n$ with $m$ constraints and that of its linear programming relaxation for a wide range of distributions on $(A,b,c)$ . We are motivated by the work of Dey, Dubey, and Molinaro (SODA '21), who gave a framework for relating the size of Branch-and-Bound (B&B) trees to additive integrality gaps. Dyer and Frieze (MOR '89) and Borst et al. (Mathematical Programming '22), respectively, showed that for certain random packing and Gaussian IPs, where the entries of $A,c$ are independently distributed according to either the uniform distribution on $[0,1]$ or the Gaussian distribution $\mathcal{N}(0,1)$, the integrality gap is bounded by $O_m(\log^2 n / n)$ with probability at least $1-1/n-e^{-Ω_m(1)}$. In this paper, we generalize these results to the case where the entries of $A$ are uniformly distributed on an integer interval (e.g., entries in $\{-1,0,1\}$), and where the columns of $A$ are distributed according to an isotropic logconcave distribution. Second, we substantially improve the success probability to $1-1/poly(n)$, compared to constant probability in prior works (depending on $m$). Leveraging the connection to Branch-and-Bound, our gap results imply that for these IPs B&B trees have size $n^{poly(m)}$ with high probability (i.e., polynomial for fixed $m$), which significantly extends the class of IPs for which B&B is known to be polynomial. Our main technical contribution is a new linear discrepancy theorem for random matrices. Our theorem gives general conditions under which a target vector is equal to or very close to a $\{0,1\}$ combination of the columns of a random matrix $A$ . The proof uses a Fourier analytic approach, building on work of Hoberg and Rothvoss (SODA '19) and Franks and Saks (RSA '20).

math.OC

Archimedes Meets Privacy: On Privately Estimating Quantiles in High Dimensions Under Minimal Assumptions

The last few years have seen a surge of work on high dimensional statistics under privacy constraints, mostly following two main lines of work: the ``worst case'' line, which does not make any distributional assumptions on the input data; and the ``strong assumptions'' line, which assumes that the data is generated from specific families, e.g., subgaussian distributions. In this work we take a middle ground, obtaining new differentially private algorithms with polynomial sample complexity for estimating quantiles in high-dimensions, as well as estimating and sampling points of high Tukey depth, all working under very mild distributional assumptions. From the technical perspective, our work relies upon deep robustness results in the convex geometry literature, demonstrating how such results can be used in a private context. Our main object of interest is the (convex) floating body (FB), a notion going back to Archimedes, which is a robust and well studied high-dimensional analogue of the interquantile range. We show how one can privately, and with polynomially many samples, (a) output an approximate interior point of the FB -- e.g., ``a typical user'' in a high-dimensional database -- by leveraging the robustness of the Steiner point of the FB; and at the expense of polynomially many more samples, (b) produce an approximate uniform sample from the FB, by constructing a private noisy projection oracle.

math.ST