SearcharxivSearch

arXiv subjects

Dmitriy Kunisky

Publications and source records attributed to Dmitriy Kunisky.

At least 19 recordsLinked to original sources

Inequalities for rank-two permanents and finite free convolutions

Bang (1976) proved the inequality for matrix permanents $\mathrm{per}^2(A) \geq 2^{-2n}\mathrm{per}(A \otimes J_2)$, where $J_2$ is the $2 \times 2$ all-ones matrix and $A$ is any $n \times n$ matrix with non-negative entries. We show that, if $A$ is any $n \times n$ real-valued matrix with rank at most two (possibly having negative entries), this inequality can be sharpened, replacing the constant $2^{-2n}$ by $1 / \binom{2n}{n} = (n!)^2 / (2n)! > 2^{-2n}$. We then show that this sharpened inequality also implies new inequalities for finite free convolutions of polynomials: if $p$ and $q$ are monic real-rooted polynomials of degree $n$, then $(p \boxplus_n q)(x)^2 \geq (p^2 \boxplus_{2n} q^2)(x)$ and $(p \boxtimes_n q)(x)^2 \geq (p^2 \boxtimes_{2n} q^2)(x)$ for all $x \in \mathbb{R}$, for $\boxplus_n$ and $\boxtimes_n$ the finite free additive and multiplicative convolution operations, respectively, on polynomials of degree $n$.

math.CO

The Hypergraph Moore Bound

The hypergraph Moore bound conjectured by Feige (2008) controls the size of the smallest even cover in a $k$-uniform hypergraph in terms of the average density of hyperedges. An even cover is a set of hyperedges covering each vertex an even number of times, generalizing the notion of a cycle in a graph, so the size of the smallest non-trivial even cover provides a notion of hypergraph girth. Recent work, starting from the breakthrough result of Guruswami, Kothari, and Manohar (2022) proved the conjecture up to polylogarithmic factors, whose exponents were later gradually improved. We give a simple proof of Feige's original hypergraph Moore bound conjecture for all $k \geq 3$, with no superfluous polylogarithmic factors. For the case of $k$ even, our proof roughly follows the proof of the graph Moore bound, but works with colored walks in a Kikuchi graph built from a hypergraph and controls their growth using the polynomial method. The argument is then extended to the case of $k$ odd by adapting a procedure in [GKM22].

math.CO

Lehner's operator norm formulas, semidefinite programming, and spiked matrix models

Lehner (1999) derived elegant formulas for the operator norm $\|\mathfrak{X}\|$ of operators of the form $\mathfrak{X} = \mathbf{A}_0 \otimes \mathfrak{1} + \sum_{i = 1}^n \mathbf{A}_i \otimes \mathfrak{m}_i$, also easily generalized to the spectral edge $\lambda_{\max}(\mathfrak{X})$, in terms of nonlinear optimization problems over positive definite matrices. Here the $\mathbf{A}_i$ are finite-dimensional Hermitian matrices, the $\mathfrak{m}_i$ are either free semicircular or free Rademacher families of operators, and $\mathfrak{1}$ is the identity operator. We first show that both of Lehner's nonlinear optimizations can be rewritten as linear semidefinite programs (SDPs), even in the Rademacher case where Lehner's optimization is not itself convex. We give the primal and dual forms of these SDPs, derive the complementary slackness relations and consequences thereof, and propose that the SDPs are more stable and accurate than the iterative numerical scheme proposed in Lehner's original work. We then apply the SDPs from the semicircular case to spiked matrix models, studied recently via Lehner's formula by Bandeira, Cipolloni, Schr\"oder, and van Handel (2024). We give a new proof of the Baik--Ben Arous--P\'ech\'e (BBP) transition they establish in models with isotropic (but possibly correlated) Gaussian noise by constructing feasible variables for the associated primal and dual SDPs. Combining our construction with a sensitivity interpretation of optimal dual variables, we study the fluctuations of leading eigenvectors of such models. We conjecture and give numerical evidence that these fluctuations are Gaussian but anisotropic and non-universal, and that their covariance may be computed in terms of the optimizer of the dual of Lehner's formula, which in turn is approximately the leading eigenmatrix of a completely positive operator associated to the covariance of the noise model.

math.PR

A revision of Litvak's conjecture on Gaussian minima and a volumetric zone conjecture

Litvak (2018) conjectured that, for any $p > 0$, the quantity $\mathbb{E}[\min_{i = 1}^n |g_i|^p]$ where $g \sim \mathcal{N}(0, \Sigma)$ is a centered Gaussian random vector is minimized among $n \times n$ correlation matrices $\Sigma$ by the Gram matrix of the regular simplex in $\mathbb{R}^{n - 1}$. We disprove this conjecture: the matrix with entries $\Sigma^{\mathrm{cos}}_{ij}=\cos(\pi(i - j) / n)$ already achieves a smaller moment for $p = 2$ and $n = 4$. We propose that $\Sigma^{\mathrm{cos}}$ is in fact the correct minimizer of these moments for all $p > 0$ and $n \geq 1$. Towards proving this, we conjecture a volumetric extension of Fejes T\'{o}th's zone conjecture (1973), whose covering version was proved by Jiang and Polyanskii (2017). Conditional on this conjecture, we show the stronger result that $\min_{i = 1}^n |g_i|$ for $g \sim \mathcal{N}(0, \Sigma^{\mathrm{cos}})$ is stochastically dominated by $\min_{i = 1}^n |h_i|$ for $h \sim \mathcal{N}(0, \Sigma)$ for any $n \times n$ correlation matrix $\Sigma$. Our counterexample $\Sigma^{\mathrm{cos}}$ was found by the AlphaEvolve AI-assisted optimization system, and we also include a brief discussion of its application to such problems.

math.PR

Universality of first-order methods on random and deterministic matrices

General first-order methods (GFOM) are a flexible class of iterative algorithms which update a state vector by matrix-vector multiplications and entrywise nonlinearities. A long line of work has sought to understand the large-n dynamics of GFOM, mostly focusing on "very random" input matrices and the approximate message passing (AMP) special case of GFOM whose state is asymptotically Gaussian. Yet, it has long remained unknown how to construct iterative algorithms that retain this Gaussianity for more structured inputs, or why existing AMP algorithms can be as effective for some deterministic matrices as they are for random matrices. We analyze diagrammatic expansions of GFOM via the limiting traffic distribution of the input matrix, the collection of all limiting values of permutation-invariant polynomials in the matrix entries, to obtain the following results: 1. We calculate the traffic distribution for the first non-trivial deterministic matrices, including (minor variants of) the Walsh-Hadamard and discrete sine and cosine transform matrices. This determines the limiting dynamics of GFOM on these inputs, resolving parts of longstanding conjectures of Marinari, Parisi, and Ritort (1994). 2. We design a new AMP iteration which unifies several previous AMP variants and generalizes to new input types, whose limiting dynamics are Gaussian conditional on some latent random variables. The asymptotic dynamics hold for a large and natural class of traffic distributions (encompassing both random and deterministic input matrices) and the algorithm's analysis gives a simple combinatorial interpretation of the Onsager correction, answering questions posed recently by Wang, Zhong, and Fan (2022).

math.PR

Gurau's spectral density is not a probability measure for individual real symmetric tensors

Gurau (2020) proposed a generalization of the trace of the matrix resolvent to tensors of higher order, and recent work has explored analogs of the Wigner semicircle and Marchenko-Pastur distributions from random matrix theory as well as aspects of free probability theory from this perspective. In particular, when evaluated with appropriate large random tensors, the limiting expectations of the coefficients of a series expansion of Gurau's resolvent trace give the moment sequences of probability measures analogous to the above distributions. We construct, on the other hand, individual deterministic tensors such that the same coefficients evaluated on those tensors do not give the moment sequence of any probability measure. Thus, the "spectral density" associated to Gurau's resolvent trace, while in a sense defined on average for certain random tensor ensembles, is not defined pointwise (unless perhaps as a signed measure) for all individual tensors.

math.PR

Generalized noise sensitivity of eigenvectors: All eigenvectors, inhomogeneous variance profiles, and dependent resampling

Chatterjee (2016) proved, as an application of his general framework relating superconcentration and chaos, that after the entries of an $n \times n$ matrix drawn from the Gaussian unitary ensemble undergo an entrywise Ornstein-Uhlenbeck (OU) process for time greater than $n^{-1/3}$, the top eigenvector of the matrix becomes almost completely decorrelated from its initial position. More recently, Bordenave, Lugosi, and Zhivotovskiy (2020) showed that the same happens under a discrete resampling model, once more than $n^{5/3}$ randomly chosen entries of a Wigner random matrix are resampled. We generalize these results in several directions: (1) we analyze the decorrelation of any eigenvector under continuous and discrete resampling dynamics, (2) we analyze the discrete resampling process for generalized Wigner matrices with inhomogeneous variance profiles, (3) we analyze a combination of continuous and discrete resampling where an OU process is repeatedly run for a certain time on randomly chosen entries, and (4) we analyze a dependent version of resampling where entries grouped into "blocks" of arbitrary shapes are resampled together. In each case, we show that a given eigenvector decorrelates provided that enough entries have been resampled or that the associated dynamics have been run for a long enough time. Our proofs take a different approach from prior work, relying more directly on the characterization of eigenvectors as derivatives of eigenvalues and reducing the problem of establishing eigenvector noise sensitivity to variants of standard and robust properties of random matrices such as bounds on eigenvalue spacings and eigenvector delocalization.

math.PR

$\mu$pscaling small models: Principled warm starts and hyperparameter transfer

Modern large-scale neural networks are often trained and released in multiple sizes to accommodate diverse inference budgets. To improve efficiency, recent work has explored model upscaling: initializing larger models from trained smaller ones to accelerate convergence. However, this method can be sensitive to hyperparameters that need to be tuned at the target upscaled model size, which is prohibitively costly to do directly. It remains unclear whether tuning hyperparameters on smaller models and extrapolating via scaling laws is sound in this setting. We address this with principled approaches to width-based upscaling and efficient hyperparameter tuning in this setting. Motivated by $\mu$P and any-dimensional architectures, we introduce a general upscaling method that, like Net2Net, copies and perturbs weights, but uses theoretically grounded, width-dependent scalings for the perturbation noise and optimizer hyperparameters. First, we prove that under zero perturbation, the upscaled model is functionally equivalent to the base model throughout training. Second, we extend the $\mu$P theory to enable infinite-width limit analysis and establish hyperparameter transfer for upscaled models, greatly reducing the tuning cost. We empirically demonstrate that this method is effective on realistic datasets and architectures.

cs.LG

Universal entrywise eigenvector fluctuations in delocalized spiked matrix models and asymptotics of rounded spectral algorithms

We consider the distribution of the top eigenvector $\widehat{v}$ of a spiked matrix model of the form $H = \theta vv^* + W$, in the supercritical regime where $H$ has an outlier eigenvalue of comparable magnitude to $\|W\|$. We show that, if $v$ is sufficiently delocalized, then the distribution of the individual entries of the projector $\widehat{v}\widehat{v}^*$ (not, we emphasize, merely the inner product $|\langle \widehat{v}, v\rangle|^2$) is universal over a large class of generalized Wigner matrices $W$ having independent entries, depending only on the first two moments of the distributions of the entries of $W$. This complements the observation of Capitaine and Donati-Martin (2021) that these distributions are not universal when $v$ is instead sufficiently localized. Further, for $W$ having entrywise variances close to constant and thus resembling a Wigner matrix, we show by comparing to $W$ drawn from the Gaussian orthogonal or unitary ensembles that averages of entrywise functions of $\widehat{v}\widehat{v}^*$ behave as they would if $\widehat{v}$ had Gaussian fluctuations around a suitable multiple of $v$. We also establish such results for several possibly dependent spiked matrices, showing that, if such matrices are entrywise uncorrelated, then their leading eigenvectors behave as they would with independent Gaussian fluctuations. We apply these results to spectral algorithms with rounding procedures for synchronization problems over the cyclic and circle groups, obtaining the first precise asymptotic error rates for such algorithms. Using our analysis of multiple spiked matrices, we also show that multi-frequency spectral algorithms using estimates from several matrices often have asymptotic error rate superior to that of naive spectral algorithms using just one matrix.

math.PR

Empirical universality and non-universality of local dynamics in the Sherrington-Kirkpatrick model

Several recent works have aimed to design algorithms for optimizing the Hamiltonians of spin glass models from statistical physics. While Montanari (2018) eventually gave a sophisticated message-passing algorithm to do this nearly optimally for the Sherrington-Kirkpatrick (SK) model, the recent work of Erba, Behrens, Krzakala, and Zdeborov\'a (2024) also observed that a simple yet unusual algorithm first proposed by Parisi (2003) seems to perform just as well: perform local reluctant search, repeatedly making the local adjustment improving the objective function by the smallest possible amount. This is in contrast to the more intuitive local greedy search that repeatedly makes the local adjustment improving the objective by the largest possible amount. We study empirically how the performance of these algorithms depends on the distribution of entries of the coupling matrix in the SK model. We find evidence that, while the runtime of greedy search enjoys universality over a broad range of distributions, the runtime of reluctant search surprisingly is not universal, sometimes depending quite sensitively on the entry distribution. We propose that one mechanism leading to this non-universality is a change in the behavior of reluctant search when the couplings have discrete support on an evenly-spaced grid, and give experimental results supporting this proposal and investigating other properties of a distribution that might affect the performance of reluctant search.

cond-mat.dis-nn

Computational and statistical lower bounds for low-rank estimation under general inhomogeneous noise

Recent work has generalized several results concerning the well-understood spiked Wigner matrix model of a low-rank signal matrix corrupted by additive i.i.d. Gaussian noise to the inhomogeneous case, where the noise has a variance profile. In particular, for the special case where the variance profile has a block structure, a series of results identified an effective spectral algorithm for detecting and estimating the signal, identified the threshold signal strength required for that algorithm to succeed, and proved information-theoretic lower bounds that, for some special signal distributions, match the above threshold. We complement these results by studying the computational optimality of this spectral algorithm. Namely, we show that, for a much broader range of signal distributions, whenever the spectral algorithm cannot detect a low-rank signal, then neither can any low-degree polynomial algorithm. This gives the first evidence for a computational hardness conjecture of Guionnet, Ko, Krzakala, and Zdeborov\'a (2023). With similar techniques, we also prove sharp information-theoretic lower bounds for a class of signal distributions not treated by prior work. Unlike all of the above results on inhomogeneous models, our results do not assume that the variance profile has a block structure, and suggest that the same spectral algorithm might remain optimal for quite general profiles. We include a numerical study of this claim for an example of a smoothly-varying rather than piecewise-constant profile. Our proofs involve analyzing the graph sums of a matrix, which also appear in free and traffic probability, but we require new bounds on these quantities that are tighter than existing ones for non-negative matrices, which may be of independent interest.

math.ST

Nonlinear Laplacians: Tunable principal component analysis under directional prior information

We introduce a new family of algorithms for detecting and estimating a rank-one signal from a noisy observation under prior information about that signal's direction, focusing on examples where the signal is known to have entries biased to be positive. Given a matrix observation $\mathbf{Y}$, our algorithms construct a nonlinear Laplacian, another matrix of the form $\mathbf{Y}+\mathrm{diag}(\sigma(\mathbf{Y1}))$ for a nonlinear $\sigma:\mathbb{R}\to\mathbb{R}$, and examine the top eigenvalue and eigenvector of this matrix. When $\mathbf{Y}$ is the (suitably normalized) adjacency matrix of a graph, our approach gives a class of algorithms that search for unusually dense subgraphs by computing a spectrum of the graph "deformed" by the degree profile $\mathbf{Y1}$. We study the performance of such algorithms compared to direct spectral algorithms (the case $\sigma=0$) on models of sparse principal component analysis with biased signals, including the Gaussian planted submatrix problem. For such models, we rigorously characterize the strength of rank-one signal, as a function of $\sigma$, required for an outlier eigenvalue to appear in the spectrum of a nonlinear Laplacian matrix. While identifying the $\sigma$ that minimizes the required signal strength in closed form seems intractable, we explore three approaches to design $\sigma$ numerically: exhaustively searching over simple classes of $\sigma$, learning $\sigma$ from datasets of problem instances, and tuning $\sigma$ using black-box optimization of the critical signal strength. We find both theoretically and empirically that, if $\sigma$ is chosen appropriately, then nonlinear Laplacian spectral algorithms substantially outperform direct spectral algorithms, while retaining the conceptual simplicity of spectral methods compared to broader classes of computations like approximate message passing or general first order methods.

stat.ML

The Lov\'asz number of random circulant graphs

This paper addresses the behavior of the Lov\'asz number for dense random circulant graphs. The Lov\'asz number is a well-known semidefinite programming upper bound on the independence number. Circulant graphs, an example of a Cayley graph, are highly structured vertex-transitive graphs on integers modulo $n$, where the connectivity of pairs of vertices depends only on the difference between their labels. While for random circulant graphs the asymptotics of fundamental quantities such as the clique and the chromatic number are well-understood, characterizing the exact behavior of the Lov\'asz number remains open. In this work, we provide upper and lower bounds on the expected value of the Lov\'asz number and show that it scales as the square root of the number of vertices, up to a log log factor. Our proof relies on a reduction of the semidefinite program formulation of the Lov\'asz number to a linear program with random objective and constraints via diagonalization of the adjacency matrix of a circulant graph by the discrete Fourier transform (DFT). This leads to a problem about controlling the norms of vectors with sparse Fourier coefficients, which we study using results on the restricted isometry property of subsampled DFT matrices.

math.CO

Low coordinate degree algorithms II: Categorical signals and generalized stochastic block models

We study when low coordinate degree functions (LCDF) -- linear combinations of functions depending on small subsets of entries of a vector -- can test for the presence of categorical structure, including community structure and generalizations thereof, in high-dimensional data. This complements the first paper of this series, which studied the power of LCDF in testing for continuous structure like real-valued signals perturbed by additive noise. We apply the tools developed there to a general form of stochastic block model (SBM), where a population is assigned random labels and every $p$-tuple of the population generates an observation according to an arbitrary probability measure associated to the $p$ labels of its members. We show that the performance of LCDF admits a unified analysis for this class of models. As applications, we prove tight lower bounds against LCDF (and therefore also against low degree polynomials) for nearly arbitrary graph and regular hypergraph SBMs, always matching suitable generalizations of the Kesten-Stigum threshold. We also prove tight lower bounds for group synchronization and abelian group sumset problems under the "truth-or-Haar" noise model, and use our technical results to give an improved analysis of Gaussian multi-frequency group synchronization. In most of these models, for some parameter settings our lower bounds give new evidence for conjectural statistical-to-computational gaps. Finally, interpreting some of our findings, we propose a precise analogy between categorical and continuous signals: a general SBM as above behaves, in terms of the tradeoff between subexponential runtime cost of testing algorithms and the signal strength needed for a testing algorithm to succeed, like a spiked $p_*$-tensor model of a certain order $p_*$ that may be computed from the parameters of the SBM.

math.ST

Statistical inference of a ranked community in a directed graph

We study the problem of detecting or recovering a planted ranked subgraph from a directed graph, an analog for directed graphs of the well-studied planted dense subgraph model. We suppose that, among a set of $n$ items, there is a subset $S$ of $k$ items having a latent ranking in the form of a permutation $\pi$ of $S$, and that we observe a fraction $p$ of pairwise orderings between elements of $\{1, \dots, n\}$ which agree with $\pi$ with probability $\frac{1}{2} + q$ between elements of $S$ and otherwise are uniformly random. Unlike in the planted dense subgraph and planted clique problems where the community $S$ is distinguished by its unusual density of edges, here the community is only distinguished by the unusual consistency of its pairwise orderings. We establish computational and statistical thresholds for both detecting and recovering such a ranked community. In the log-density setting where $k$, $p$, and $q$ all scale as powers of $n$, we establish the exact thresholds in the associated exponents at which detection and recovery become statistically and computationally feasible. These regimes include a rich variety of behaviors, exhibiting both statistical-computational and detection-recovery gaps. We also give finer-grained results for two extreme cases: (1) $p = 1$, $k = n$, and $q$ small, where a full tournament is observed that is weakly correlated with a global ranking, and (2) $p = 1$, $q = \frac{1}{2}$, and $k$ small, where a small "ordered clique" (totally ordered directed subgraph) is planted in a random tournament.

math.ST

Asymptotic Bounds and Online Algorithms for Average-Case Matrix Discrepancy

We study the matrix discrepancy problem in the average-case setting. Given a sequence of $m \times m$ symmetric matrices $A_1,\ldots,A_n$, its discrepancy is defined as the minimal spectral norm over all signed sums $\sum_{i=1}^n x_iA_i$ with $x_1,\ldots,x_n \in \{\pm1\}$. Our contributions are twofold. First, we study the asymptotic discrepancy of random matrices. When the matrices belong to the Gaussian orthogonal ensemble, we provide a sharp characterization of the asymptotic discrepancy and show that the limiting distribution is concentrated around $\Theta(\sqrt{nm}4^{-(1 + o(1))n/m^2})$, under the assumption $m^2 \ll n/\log{n}$. We observe that the trivial bound $O(\sqrt{nm})$ cannot be improved when $n \ll m^2$ and show that this phenomenon occurs for a broad class of random matrices. In the case $n = \Omega(m^2)$, we provide a matching upper bound. Second, we analyse the matrix hyperbolic cosine algorithm, an online algorithm for matrix discrepancy minimization due to Zouzias (2011), in the average-case setting. We show that the algorithm achieves with high probability a discrepancy of $O(m\log{m})$ for a broad class of random matrices, including Wigner matrices with entries satisfying a hypercontractive inequality and Gaussian Wishart matrices.

math.PR

On the Structure of Bad Science Matrices

The bad science matrix problem consists in finding, among all matrices $A \in \mathbb{R}^{n \times n}$ with rows having unit $\ell^2$ norm, one that maximizes $\beta(A) = \frac{1}{2^n} \sum_{x \in \{-1, 1\}^n} \|Ax\|_\infty$. Our main contribution is an explicit construction of an $n \times n$ matrix $A$ showing that $\beta(A) \geq \sqrt{\log_2(n+1)}$, which is only 18% smaller than the asymptotic rate. We prove that every entry of any optimal matrix is a square root of a rational number, and we find provably optimal matrices for $n \leq 4$.

math.FA

Inference of rankings planted in random tournaments

We consider the problem of inferring an unknown ranking of $n$ items from a random tournament on $n$ vertices whose edge directions are correlated with the ranking. We establish, in terms of the strength of these correlations, the computational and statistical thresholds for detection (deciding whether an observed tournament is purely random or drawn correlated with a hidden ranking) and recovery (estimating the hidden ranking with small error in Spearman's footrule or Kendall's tau metric on permutations). Notably, we find that this problem provides a new instance of a detection-recovery gap: solving the detection problem requires much weaker correlations than solving the recovery problem. In establishing these thresholds, we also identify simple algorithms for detection (thresholding a degree 2 polynomial) and recovery (outputting a ranking by the number of "wins" of a tournament vertex, i.e., the out-degree) that achieve optimal performance up to constants in the correlation strength. For detection, we find that the above low-degree polynomial algorithm is superior to a natural spectral algorithm. We also find that, whenever it is possible to achieve strong recovery (i.e., to estimate with vanishing error in the above metrics) of the hidden ranking, then the above "Ranking By Wins" algorithm not only does so, but also outputs a close approximation of the maximum likelihood estimator, a task that is NP-hard in the worst case.

math.ST