SearcharxivSearch

arXiv subjects

Aryeh Kontorovich

Publications and source records attributed to Aryeh Kontorovich.

At least 19 recordsLinked to original sources

Total Variation Distance between Product Distributions: an Analytic Proxy

We characterize, up to universal constants, the total variation distance between finite products of arbitrary probability measures by a simple formula. A theorem of Latala reduces this expression to one scalar equation. As an application, we characterize the sample complexity of equal-prior binary hypothesis testing uniformly in the weak-detection regime, where squared Hellinger distance alone does not determine the answer, and also characterize the TV distance between multinomial distributions.

math.PR

Sharp margin-based generalization bounds for realizable SVM

Let the exact homogeneous hard-margin support vector machine be trained on \(m\) independent observations from a Borel probability law on a real Hilbert space. We prove that, with score zero counted as an error, there is a universal numerical constant \(C\) such that \[ \Pp\left( γ_m>0,\quad \Risk(u_m)> \frac{C}{m} \left( K_m+\log\frac1δ \right) \right) \le δ. \] Here \(γ_m\) is the empirical homogeneous margin, \(u_m\) is the exact minimum-norm unit-margin separator, \(r_m\) is the largest training radius, and \(K_m:=r_m^2\norm{u_m}^2=r_m^2/γ_m^2\) on \(\{γ_m>0\}\). The proof is driven by a deterministic deletion problem. Given vectors \(x_1,\ldots,x_n\) in the unit ball, delete a set \(B\) of constraints and let \(u_B\) be the closest point to the origin that satisfies every retained unit-margin constraint. Suppose that \(\norm{u_B}^2\le k\) and that every deleted vector has nonpositive score under \(u_B\). We prove that a family of such deletion sets of cardinality \(q\) has size at most \(\exp(8k+2q)\). The conceptual step is an exact identity obtained from the KKT representation of \(u_B\). For a random deletion set, the identity converts the mean squared spread of the separators into a weighted sum of score deficits. It therefore forces a coordinate whose deletion status separates the two conditional means by a quantitatively large amount. Revealing that coordinate decreases the conditional separator variance enough to control the binary entropy of the split. An entropy induction gives the deletion count, and an exact factorial ghost-sample identity converts that count into the stated high-probability SVM bound.

stat.ML

Realizable Bayes-Consistency for General Metric Losses

We study strong universal Bayes-consistency in the realizable setting for learning with general metric losses, extending classical characterizations beyond $0$-$1$ classification (Bousquet et al., 2020; Hanneke et al., 2021) and real-valued regression (Attias et al., 2024). Given an instance space $(X,ρ)$, a label space $(Y,\ell)$ with possibly unbounded loss, and a hypothesis class $H \subseteq Y^{X}$, we resolve the realizable case of an open problem presented in Tsir Cohen and Kontorovich (2022). Specifically, we find the necessary and sufficient conditions on the hypothesis class $H$ under which there exists a distribution-free learning rule whose risk converges almost surely to the best-in-class risk (which is zero) for every realizable data-generating distribution. Our main contribution is this sharp characterization in terms of a combinatorial obstruction: Similarly to Attias et al. (2024), we introduce the notion of an infinite non-decreasing $(γ_k)$-Littlestone tree, where $γ_k \to \infty$. This extends the Littlestone tree structure used in Bousquet et al. (2020) to the metric loss setting.

cs.LG

Improved Distribution Estimation in $\ell_\infty$

We present improved bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These include minimax bounds in expectation and high-probability tail bounds. We resolve some of the open questions posed in Kontorovich and Painsky (JMLR, 2025) -- including a fully empirical version of the tightest risk bound they presented and identifying the form of the worst-case extremal distribution. Encouraging empirical results are reported as well.

stat.ML

A Fine-Grained Understanding of Uniform Convergence for Halfspaces

We study the fine-grained uniform convergence behavior of halfspaces beyond worst-case VC bounds. For inhomogeneous halfspaces in $\mathbb{R}^d$ with $d\ge 2$, we show that standard first-order VC bounds are essentially tight: even consistent hypotheses can incur population error $Θ(d\ln(n/d)/n)$, and in the agnostic setting the deviation scales as $\sqrt{τ\ln(1/τ)}$ at true error $τ$. In contrast, homogeneous halfspaces in $\mathbb{R}^2$ exhibit a markedly different behavior. In the realizable case, every hypothesis consistent with the sample has error $O(1/n)$. In the agnostic case, we prove a bandwise, log-free deviation bound on each dyadic risk band via a critical-wedge localization argument. Unioning over bands incurs only a $\ln\ln n$ overhead, and we establish a matching lower bound showing this overhead is unavoidable. Together, these results give a fine-grained and nearly complete picture of uniform convergence for halfspaces, revealing sharp dimensional and structural thresholds.

cs.LG

A homogenization principle for total variation

A homogenization principle for total variation We prove an inequality comparing the variational distance between pairs of product probability measures to its homogenized counterpart. If $P_1,\ldots,P_n,Q_1,\ldots,Q_n$ are arbitrary probability measures on a measurable space and $\bar P:=\frac1n\sum_{i=1}^n P_i, \bar Q:=\frac1n\sum_{i=1}^n Q_i $, we show that $$TV\!\left(\bigotimes_{i=1}^n P_i, \bigotimes_{i=1}^n Q_i\right) \;\ge\; c\,TV(\bar P^{\otimes n},\bar Q^{\otimes n}),$$ where $c>0$ is a universal constant. The proof is based on a one-dimensional representation of total variation between products. We embed pairs of probability distributions $P_i,Q_i$ into positive measures $η_i$ on $\mathbb{R}$. We then define a functional $T$ over measures on $\mathbb{R}$ that realizes TV over products via convolution: $TV\!\left(\bigotimes_{i=1}^n P_i, \bigotimes_{i=1}^n Q_i\right)=T(η_1*\cdots *η_n)$. Our main analytic discovery is that for the relevant class of positive measures $η_i$, the convolution inequality $T(η_1*\cdots*η_n) \ge c\,T\!\left(\barη^{*n}\right)$ holds, where $\barη=\frac1n\sum_{i=1}^n η_i$. Finally, a higher-dimensional lifting argument shows that $T\!\left(\barη^{*n}\right)\ge TV(\bar P^{\otimes n},\bar Q^{\otimes n})$. To our knowledge, both the exact representation and the convolution inequality are new.

math.PR

TV homogenization inequalities

We study the total variation distance under two information-erasing maps on inhomogeneous Bernoulli product measures: summation and homogenization. While summation is a Markov kernel and hence satisfies the usual data processing inequality, homogenization -- which maps each Bernoulli parameter to the cumulative mean -- is not. Nevertheless, we prove that the homogenization map also reduces the TV distance, up to a universal constant. The argument is based on an explicit two-sided control of the TV distance between Poisson binomials, obtained via a parameter interpolation and a second-moment extraction lemma.

math.PR

TV over Bernoulli products: the small parameter regime

We study the total variation distance (TV) between two $n$-fold Bernoulli product measures parametrized by $\vec p=(p_1,\ldots,p_n)$ and $\vec q=(q_1,\ldots,q_n)$, respectively, in the \emph{tiny} and \emph{small} regimes. In the tiny regime, we have $p_i,q_i\lesssim 1/n^2$, and in the small regime, $p_i,q_i\lesssim 1/n$. We discover that in the tiny regime, the TV distance behaves as $\|\vec p-\vec q\|_1$, while in the small regime, it behaves as \[ \sum_{i=1}^n \Big| p_i\prod_{j\neq i}(1-p_j) - q_i\prod_{j\neq i}(1-q_j) \Big|, \] both up to absolute constants. Along the way we discover some identities of possible independent interest.

math.PR

ParallelTime: Dynamically Weighting the Balance of Short- and Long-Term Temporal Dependencies

Modern multivariate time series forecasting primarily relies on two architectures: the Transformer with attention mechanism and Mamba. In natural language processing, an approach has been used that combines local window attention for capturing short-term dependencies and Mamba for capturing long-term dependencies, with their outputs averaged to assign equal weight to both. We find that for time-series forecasting tasks, assigning equal weight to long-term and short-term dependencies is not optimal. To mitigate this, we propose a dynamic weighting mechanism, ParallelTime Weighter, which calculates interdependent weights for long-term and short-term dependencies for each token based on the input and the model's knowledge. Furthermore, we introduce the ParallelTime architecture, which incorporates the ParallelTime Weighter mechanism to deliver state-of-the-art performance across diverse benchmarks. Our architecture demonstrates robustness, achieves lower FLOPs, requires fewer parameters, scales effectively to longer prediction horizons, and significantly outperforms existing methods. These advances highlight a promising path for future developments of parallel Attention-Mamba in time series forecasting. The implementation is readily available at: \href{https://github.com/itay1551/ParallelTime}{GitHub}.

cs.LG

Bounded variation separates weak and strong average Lipschitz

We closely examine a notion of average smoothness recently introduced by Ashlagi et al. (JMLR, 2024). The latter defined a {\em weak} and {\em strong} average-Lipschitz seminorm for real-valued functions on general metric spaces. Specializing to the standard metric on the real line, we compare these notions to bounded variation (BV) and discover that the weak notion is strictly weaker than BV while the strong notion strictly stronger. Along the way, we discover that the weak average smooth class is also considerably larger in a certain combinatorial sense, made precise by the fat-shattering dimension.

math.FA

The Empirical Mean is Minimax Optimal for Local Glivenko-Cantelli

We revisit the recently introduced Local Glivenko-Cantelli setting, which studies distribution-dependent uniform convergence rates of the Empirical Mean Estimator (EME). In this work, we investigate generalizations of this setting where arbitrary estimators are allowed rather than just the EME. Can a strictly larger class of measures be learned? Can better risk decay rates be obtained? We provide exhaustive answers to these questions, which are both negative, provided the learner is barred from exploiting some infinite-dimensional pathologies. On the other hand, allowing such exploits does lead to a strictly larger class of learnable measures.

math.ST

Sharp bounds on aggregate expert error

We revisit the classic problem of aggregating binary advice from conditionally independent experts, also known as the Naive Bayes setting. Our quantity of interest is the error probability of the optimal decision rule. In the case of symmetric errors (sensitivity = specificity), reasonably tight bounds on the optimal error probability are known. In the general asymmetric case, we are not aware of any nontrivial estimates on this quantity. Our contribution consists of sharp upper and lower bounds on the optimal error probability in the general case, which recover and sharpen the best known results in the symmetric special case. Since this turns out to be equivalent to estimating the total variation distance between two product distributions, our results also have bearing on this important and challenging problem.

math.PR

On the tensorization of the variational distance

If one seeks to estimate the total variation between two product measures $||P^\otimes_{1:n}-Q^\otimes_{1:n}||$ in terms of their marginal TV sequence $δ=(||P_1-Q_1||,||P_2-Q_2||,\ldots,||P_n-Q_n||)$, then trivial upper and lower bounds are provided by$ ||δ||_\infty \le ||P^\otimes_{1:n}-Q^\otimes_{1:n}||\le||δ||_1$. We improve the lower bound to $||δ||_2\lesssim||P^\otimes_{1:n}-Q^\otimes_{1:n}||$, thereby reducing the gap between the upper and lower bounds from $\sim n$ to $\sim\sqrt $. Furthermore, we show that {\em any} estimate on $||P^\otimes_{1:n}-Q^\otimes_{1:n}||$ expressed in terms of $δ$ must necessarily exhibit a gap of $\sim\sqrt n$ between the upper and lower bounds in the worst case, establishing a sense in which our estimate is optimal. Finally, we identify a natural class of distributions for which $||δ||_2$ approximates the TV distance up to absolute multiplicative constants.

math.PR

Exact expressions for the maximal probability that all $k$-wise independent bits are 1

Let $M(n, k, p)$ denote the maximum probability of the event $X_1 = X_2 = \cdots = X_n=1$ under a $k$-wise independent distribution whose marginals are Bernoulli random variables with mean $p$. A long-standing question is to calculate $M(n, k, p)$ for all values of $n,k,p$. This question has been partially addressed by several authors, primarily with the goal of answering asymptotic questions. The present paper focuses on obtaining exact expressions for this probability. To this end, we provide closed-form formulas of $M(n,k,p)$ for $p$ near 0 as well as $p$ near 1.

math.PR

Decoupling Maximal Inequalities

A {\em maximal inequality} seeks to estimate $\mathbb{E}\max_i X_i$ in terms of properties of the $X_i$. When the latter are independent, the union bound (in its various guises) can yield tight upper bounds. If, however, the $X_i$ are strongly dependent, the estimates provided by the union bound will be rather loose. In this note, we show that for non-negative random variables, pairwise independence suffices for the maximal inequality to behave comparably to its independent version. The condition of pairwise independence may be relaxed to a kind of negative dependence, and even the latter admits violations -- provided these are properly quantified.

math.PR

Efficient Agnostic Learning with Average Smoothness

We study distribution-free nonparametric regression following a notion of average smoothness initiated by Ashlagi et al. (2021), which measures the "effective" smoothness of a function with respect to an arbitrary unknown underlying distribution. While the recent work of Hanneke et al. (2023) established tight uniform convergence bounds for average-smooth functions in the realizable case and provided a computationally efficient realizable learning algorithm, both of these results currently lack analogs in the general agnostic (i.e. noisy) case. In this work, we fully close these gaps. First, we provide a distribution-free uniform convergence bound for average-smoothness classes in the agnostic setting. Second, we match the derived sample complexity with a computationally efficient agnostic learning algorithm. Our results, which are stated in terms of the intrinsic geometry of the data and hold over any totally bounded metric space, show that the guarantees recently obtained for realizable learning of average-smooth functions transfer to the agnostic setting. At the heart of our proof, we establish the uniform convergence rate of a function class in terms of its bracketing entropy, which may be of independent interest.

cs.LG

Distribution Estimation under the Infinity Norm

We present novel bounds for estimating discrete probability distributions under the $\ell_\infty$ norm. These are nearly optimal in various precise senses, including a kind of instance-optimality. Our data-dependent convergence guarantees for the maximum likelihood estimator significantly improve upon the currently known results. A variety of techniques are utilized and innovated upon, including Chernoff-type inequalities and empirical Bernstein bounds. We illustrate our results in synthetic and real-world experiments. Finally, we apply our proposed framework to a basic selective inference problem, where we estimate the most frequent probabilities in a sample.

math.ST

Correlated Binomial Process

Cohen and Kontorovich (COLT 2023) initiated the study of what we call here the Binomial Empirical Process: the maximal absolute value of a sequence of inhomogeneous normalized and centered binomials. They almost fully analyzed the case where the binomials are independent, and the remaining gap was closed by Blanchard and Voráček (ALT 2024). In this work, we study the much more general and challenging case with correlations. In contradistinction to Gaussian processes, whose behavior is characterized by the covariance structure, we discover that, at least somewhat surprisingly, for binomial processes covariance does not even characterize convergence. Although a full characterization remains out of reach, we take the first steps with nontrivial upper and lower bounds in terms of covering numbers.

math.PR