SearcharxivSearch

arXiv subjects

Mark Sellke

Publications and source records attributed to Mark Sellke.

At least 19 recordsLinked to original sources

Algorithmic threshold for high-dimensional projection pursuit I: general theory

We study a null model of high-dimensional projection pursuit: we are given $M$ points sampled i.i.d. from a standard gaussian in $N$ dimensions, where $M,N\to\infty$ with $M/N\to\alpha\in(0,\infty)$. Our goal is to characterize the possible empirical distributions of these points' projections along a data-dependent direction $x$, which ranges over either the sphere $S_N=\sqrt{N}\mathbb{S}^{N-1}$ or cube $\Sigma_N=\{-1,+1\}^N$. We consider this problem in an algorithmic setting, where $x$ must be the output of an algorithm with dimension-free Lipschitz dependence on the input; this class of algorithms includes general gradient-based methods such as Langevin dynamics and approximate message passing (AMP). Our main result exactly characterizes the set of empirical distributions attainable by this class in terms of a one-dimensional stochastic control problem. As a consequence of our main result, we obtain exact algorithmic thresholds for optimizing the Hamiltonian of a spherical or Ising perceptron model with general bounded continuous activation. For the spherical problem, independent work of Montanari and Zhou (2024) characterized the empirical distributions attainable by a related two-stage AMP algorithm, also in terms of stochastic control. Our proof of hardness builds on the branching overlap gap property introduced in earlier work by the first two authors. Our main innovation is to develop stochastic control theory within the branching OGP framework, significantly expanding the settings in which it locates an exact algorithmic threshold. Notably, our methods apply even though the non-algorithmic problem of characterizing all feasible projections remains a major outstanding challenge. For the matching algorithmic result, we construct a new incremental AMP algorithm that acts on a Brownian-bridge revelation of the gaussian disorder and simulates the same family of controlled SDEs.

math.PR

Thompson Sampling Is 2-Competitive for Mistakes

We consider Bayesian bandit models and prove that Thompson sampling makes at most twice the expected number of mistakes (selections of a suboptimal arm) as any other policy. Our analysis applies as long as the latent arm processes are independent and each arm evolves only when played. For stochastic bandits with best arm defined via mean reward, this confirms a conjecture of Guha and Munagala from 2014, where the factor $2$ is already best possible. The result holds under any nonincreasing sequence of round weights, including fixed horizon and geometric discounting.

stat.ML

A Counterexample to the Gaussian Completely Monotone Conjecture

We provide an explicit probability measure on $\mathbb{R}$ for which the fifth time derivative of the entropy along the heat flow is positive at some time. This disproves the Gaussian completely monotone (GCM) conjecture (Cheng-Geng '15) and therefore also the Gaussian optimality conjecture (McKean '66) and the entropy power conjecture (Toscani '15). Our proof also implies the existence of a log-concave probability measure on $\mathbb{R}$ for which the GCM conjecture fails at some order. The explicit counterexample was found by GPT-5.5 Pro.

math.PR

Short proofs in combinatorics, probability and number theory II

We give a quintet of proofs resulting from questions posed by Erd\H{o}s. These questions concern ordinary lines in planar point sets, sequences with uniformly small exponential sums, $K_4$-free $4$-critical graphs with few chords in any cycle, a counterexample to a "fewnomial" version of the Erd\H{o}s--Tur\'{a}n discrepancy bound, and a finiteness theorem for integers $n$ such that $n-a k^2$ is prime for all $k\leq \sqrt{n/a}$ coprime to $n$ (for fixed $a\in\mathbb Z_+$). Each proof is due to an internal model at OpenAI.

math.CO

Short proofs in combinatorics and number theory

We give a triplet of short proofs, each of which answers a question raised by Erd\H{o}s. The first concerns the small prime factors of $\binom{n}{k}$, the second concerns whether an additive basis $A$ can always be split into pieces $A_1$ and $A_2$ such that each of $A_i + A_i$ has bounded gaps, and the final concerns whether $\{\alpha p\}$ is "well-distributed" in the sense introduced by Hlawka and Petersen. In each case, the proof is due entirely to an internal model at OpenAI.

math.CO

Stable algorithms cannot reliably find isolated perceptron solutions

We study the binary perceptron, a random constraint satisfaction problem that asks to find a Boolean vector in the intersection of independently chosen random halfspaces. A striking feature of this model is that at every positive constraint density, it is expected that a $1-o_N(1)$ fraction of solutions are \emph{strongly isolated}, i.e. separated from all others by Hamming distance $\Omega(N)$. At the same time, efficient algorithms are known to find solutions at certain positive constraint densities. This raises a natural question: can any isolated solution be algorithmically visible? We answer this in the negative: no algorithm whose output is stable under a tiny Gaussian resampling of the disorder can \emph{reliably} locate isolated solutions. We show that any stable algorithm has success probability at most $\frac{3\sqrt{17}-9}{4}+o_N(1)\leq 0.84233$. Furthermore, every stable algorithm that finds a solution with probability $1-o_N(1)$ finds an isolated solution with probability $o_N(1)$. The class of stable algorithms we consider includes degree-$D$ polynomials up to $D\leq o(N/\log N)$; under the low-degree heuristic \cite{hopkins2018statistical}, this suggests that locating strongly isolated solutions requires running time $\exp(\widetilde{\Theta}(N))$. Our proof does not use the overlap gap property. Instead, we show via Pitt's correlation inequality that after a random perturbation of the disorder, the number of solutions located close to a pre-existing isolated solution cannot concentrate at $1$.

cs.CC

On Anti-Confinement Estimates for Self-Repelling Random Walks

We study a class of $d$-dimensional random walks, including the two-dimensional simple random walk, reweighted by a self-repelling Gibbsian pair potential. We prove lower bounds on the diffusion constant for short-range interactions, and superdiffusive behavior in case the interaction is sufficiently long-range. Finally, we show that in the superdiffusive regime, faster temporal decay can be compensated by stronger spatial repulsion and vice-versa. Our technique combines GKS-based correlation inequalities on path space with recursive multi-scale estimates.

math.PR

On Learning-Curve Monotonicity for Maximum Likelihood Estimators

The property of learning-curve monotonicity, highlighted in a recent series of work by Loog, Mey and Viering, describes algorithms which only improve in average performance given more data, for any underlying data distribution within a given family. We establish the first nontrivial monotonicity guarantees for the maximum likelihood estimator in a variety of well-specified parametric settings. For sequential prediction with log loss, we show monotonicity (in fact complete monotonicity) of the forward KL divergence for Gaussian vectors with unknown covariance and either known or unknown mean, as well as for Gamma variables with unknown scale parameter. The Gaussian setting was explicitly highlighted as open in the aforementioned works, even in dimension 1. Finally we observe that for reverse KL divergence, a folklore trick yields monotonicity for very general exponential families. All results in this paper were derived by variants of GPT-5.2 Pro. Humans did not provide any proof strategies or intermediate arguments, but only prompted the model to continue developing additional results, and verified and transcribed its proofs.

math.ST

Exponentially Slow Mixing of the Low Temperature SK Model

We give a short proof that low-temperature dynamics for the Sherrington-Kirkpatrick model have mixing time exponential in the system size, based on the recently proved existence of gapped spin configurations by (Minzer-Sah-Sawhney 2023, Dandi-Gamarnik-Zdeborov\'a 2023). This result is in contrast with a well established physics prediction which posits a stretched exponential mixing time of order $e^{N^{1/3 \pm o(1)}}$. Our proof clarifies that this prediction cannot apply to mixing from worst case initial conditions, but should presumably be understood to concern dynamics from a suitably random initialization.

math.PR

Early science acceleration experiments with GPT-5

AI models like GPT-5 are an increasingly valuable tool for scientists, but many remain unaware of the capabilities of frontier AI. We present a collection of short case studies in which GPT-5 produced new, concrete steps in ongoing research across mathematics, physics, astronomy, computer science, biology, and materials science. In these examples, the authors highlight how AI accelerated their work, and where it fell short; where expert time was saved, and where human input was still key. We document the interactions of the human authors with GPT-5, as guiding examples of fruitful collaboration with AI. Of note, this paper includes four new results in mathematics (carefully verified by the human authors), underscoring how GPT-5 can help human mathematicians settle previously unsolved problems. These contributions are modest in scope but profound in implication, given the rate at which frontier AI is progressing.

cs.CL

Universality of Cutoff for Riffle Shuffling

A Gilbert-Shannon-Reeds (GSR) shuffle is performed on a deck of $N$ cards by cutting the top $n\sim Bin(N,1/2)$ cards and interleaving the two resulting piles uniformly at random. The celebrated "Seven shuffles suffice" theorem of [Bayer-Diaconis '92] established cutoff for this Markov chain: to leading order, total variation mixing occurs after precisely $\frac{3}{2}\log_2 N$ shuffles. Later work of [Lalley '00] and [Sellke '22] extended this result to asymmetric binomial cuts $n\sim Bin(N,p)$ for all $p\in (0,1)$. These results relied heavily on the binomial condition and many natural chains were left open, including uniformly random cuts and exact bisections. We establish cutoff for riffle shuffles with general pile size distribution. Namely, suppose the cut sizes $(n^{(t)})_{t\geq 1}$ are IID and the convergence in distribution $n^{(t)}/N \stackrel{d}{\to} \mu$ holds for some probability measure $\mu$ on the interval $[0,1]$. Then the mixing time $t_{mix}$ satisfies $t_{mix}/\log N\to \overline{C}_{\mu}$ for an explicit constant $\overline{C}_{\mu}$. The same result holds for any (deterministic or random) sequence of pile sizes with empirical distribution converging to $\mu$ on all macroscopic time intervals (of length $\Omega(\log N)$). It also extends to multi-partite shuffles where the deck is cut into more than $2$ piles in each step. Qualitatively, we find that the "cold spot" phenomenon identified by [Lalley '00] characterizes the mixing time of riffle shuffling in great generality.

math.PR

Geometry Meets Incentives: Sample-Efficient Incentivized Exploration with Linear Contexts

In the incentivized exploration model, a principal aims to explore and learn over time by interacting with a sequence of self-interested agents. It has been recently understood that the main challenge in designing incentive-compatible algorithms for this problem is to gather a moderate amount of initial data, after which one can obtain near-optimal regret via posterior sampling. With high-dimensional contexts, however, this \emph{initial exploration} phase requires exponential sample complexity in some cases, which prevents efficient learning unless initial data can be acquired exogenously. We show that these barriers to exploration disappear under mild geometric conditions on the set of available actions, in which case incentive-compatibility does not preclude regret-optimality. Namely, we consider the linear bandit model with actions in the Euclidean unit ball, and give an incentive-compatible exploration algorithm with sample complexity that scales polynomially with the dimension and other parameters.

cs.GT

On the Discontinuous Breaking of Replica Symmetry and Shattering in Mean-Field Spin Glasses

We show that in mean-field spin glasses, a discontinuous breaking of replica symmetry at the critical inverse temperature $\beta_c$ implies the existence of an intermediate shattered phase. This confirms a prediction from physics regarding the nature of random first order phase transitions. On the other hand, we give an example of a spherical spin glass which exhibits shattering, yet the transition is continuous at $\beta_c$.

math.PR

Strong Low Degree Hardness for the Number Partitioning Problem

In the number partitioning problem (NPP) one aims to partition a given set of $N$ real numbers into two subsets with approximately equal sum. The NPP is a well-studied optimization problem and is famous for possessing a statistical-to-computational gap: when the $N$ numbers to be partitioned are i.i.d. standard gaussian, the optimal discrepancy is $2^{-\Theta(N)}$ with high probability, but the best known polynomial-time algorithms only find solutions with a discrepancy of $2^{-\Theta(\log^2 N)}$. This gap is a common feature in optimization problems over random combinatorial structures, and indicates the need for a study that goes beyond worst-case analysis. We provide evidence of a nearly tight algorithmic barrier for the number partitioning problem. Namely we consider the family of low coordinate degree algorithms (with randomized rounding into the Boolean cube), and show that degree $D$ algorithms fail to solve the NPP to accuracy beyond $2^{-\widetilde O(D)}$. According to the low degree heuristic, this suggests that simple brute-force search algorithms are nearly unimprovable, given any allotted runtime between polynomial and exponential in $N$. Our proof combines the isolation of solutions in the landscape with a conditional form of the overlap gap property: given a good solution to an NPP instance, slightly noising the NPP instance typically leaves no good solutions near the original one. In fact our analysis applies whenever the $N$ numbers to be partitioned are independent with uniformly bounded density.

math.ST

Tight Low Degree Hardness for Optimizing Pure Spherical Spin Glasses

We prove constant degree polynomial algorithms cannot optimize pure spherical $p$-spin Hamiltonians beyond the algorithmic threshold $\mathsf{ALG}(p)=2\sqrt{\frac{p-1}{p}}$. The proof goes by transforming any hypothetical such algorithm into a Lipschitz one, for which hardness was shown previously by the author and B. Huang.

math.PR

Nonparametric MLE for Gaussian Location Mixtures: Certified Computation and Generic Behavior

We study the nonparametric maximum likelihood estimator $\widehat{\pi}$ for Gaussian location mixtures in one dimension. It has been known since (Lindsay, 1983) that given an $n$-point dataset, this estimator always returns a mixture with at most $n$ components, and more recently (Wu-Polyanskiy, 2020) gave a sharp $O(\log n)$ bound for subgaussian data. In this work we study computational aspects of $\widehat{\pi}$. We provide an algorithm which for small enough $\varepsilon>0$ computes an $\varepsilon$-approximation of $\widehat\pi$ in Wasserstein distance in time $K+Cnk^2\log\log(1/\varepsilon)$. Here $K$ is data-dependent but independent of $\varepsilon$, while $C$ is an absolute constant and $k=|supp(\widehat{\pi})|\leq n$ is the number of atoms in $\widehat\pi$. We also certifiably compute the exact value of $|supp(\widehat\pi)|$ in finite time. These guarantees hold almost surely whenever the dataset $(x_1,\dots,x_n)\in [-cn^{1/4},cn^{1/4}]$ consists of independent points from a probability distribution with a density (relative to Lebesgue measure). We also show the distribution of $\widehat\pi$ conditioned to be $k$-atomic admits a density on the associated $2k-1$ dimensional parameter space for all $k\leq \sqrt{n}/3$, and almost sure locally linear convergence of the EM algorithm. One key tool is a classical Fourier analytic estimate for non-degenerate curves.

math.ST

Strong Low Degree Hardness for Stable Local Optima in Spin Glasses

It is a folklore belief in the theory of spin glasses and disordered systems that out-of-equilibrium dynamics fail to find stable local optima exhibiting e.g. local strict convexity on physical time-scales. In the context of the Sherrington--Kirkpatrick spin glass, Behrens-Arpino-Kivva-Zdeborov\'a and Minzer-Sah-Sawhney have recently conjectured that this obstruction may be inherent to all efficient algorithms, despite the existence of exponentially many such optima throughout the landscape. We prove this search problem exhibits strong low degree hardness for polynomial algorithms of degree $D\leq o(N)$: any such algorithm has probability $o(1)$ to output a stable local optimum. To the best of our knowledge, this is the first result to prove that even constant-degree polynomials have probability $o(1)$ to solve a random search problem without planted structure. To prove this, we develop a general-purpose enhancement of the ensemble overlap gap property, and as a byproduct improve previous results on spin glass optimization, maximum independent set, random $k$-SAT, and the Ising perceptron to strong low degree hardness. Finally for spherical spin glasses with no external field, we prove that Langevin dynamics does not find stable local optima within dimension-free time.

cond-mat.dis-nn