SearcharxivSearch

arXiv subjects

Thomas A. Courtade

Publications and source records attributed to Thomas A. Courtade.

At least 19 recordsLinked to original sources

Robust Estimation Under Heterogeneous Corruption Rates

We study the problem of robust estimation under heterogeneous corruption rates, where each sample may be independently corrupted with a known but non-identical probability. This setting arises naturally in distributed and federated learning, crowdsourcing, and sensor networks, yet existing robust estimators typically assume uniform or worst-case corruption, ignoring structural heterogeneity. For mean estimation for multivariate bounded distributions and univariate gaussian distributions, we give tight minimax rates for all heterogeneous corruption patterns. For multivariate gaussian mean estimation and linear regression, we establish the minimax rate for squared error up to a factor of $\sqrt{d}$, where $d$ is the dimension. Roughly, our findings suggest that samples beyond a certain corruption threshold may be discarded by the optimal estimators -- this threshold is determined by the empirical distribution of the corruption rates given.

cs.LG

Generalized Blaschke--Santaló-type inequalities, without symmetry restrictions

Nakamura and Tsuji (2024) recently investigated a many-function generalization of the functional Blaschke--Santaló inequality, which they refer to as a generalized Legendre duality relation. They showed that, among the class of all even test functions, centered Gaussian functions saturate this general family of functional inequalities. Leveraging a certain entropic duality, we give a short alternate proof of Nakamura and Tsuji's result, and, in the process, eliminate all symmetry assumptions. As an application, we establish a Talagrand-type inequality for the Wasserstein barycenter problem (without symmetry restrictions) originally conjectured by Kolesnikov and Werner (\textit{Adv.~Math.}, 2022). An analogous geometric Blaschke--Santaló-type inequality is established for many convex bodies, again without symmetry assumptions.

math.FA

Managing Correlations in Data and Privacy Demand

Previous works in the differential privacy literature that allow users to choose their privacy levels typically operate under the heterogeneous differential privacy (HDP) framework with the simplifying assumption that user data and privacy levels are not correlated. Firstly, we demonstrate that the standard HDP framework falls short when user data and privacy demands are allowed to be correlated. Secondly, to address this shortcoming, we propose an alternate framework, Add-remove Heterogeneous Differential Privacy (AHDP), that jointly accounts for user data and privacy preference. We show that AHDP is robust to possible correlations between data and privacy. Thirdly, we formalize the guarantees of the proposed AHDP framework through an operational hypothesis testing perspective. The hypothesis testing setup may be of independent interest in analyzing other privacy frameworks as well. Fourthly, we show that there exists non-trivial AHDP mechanisms that notably do not require prior knowledge of the data-privacy correlations. We propose some such mechanisms and apply them to core statistical tasks such as mean estimation, frequency estimation, and linear regression. The proposed mechanisms are simple to implement with minimal assumptions and modeling requirements, making them attractive for real-world use. Finally, we empirically evaluate proposed AHDP mechanisms, highlighting their trade-offs using LLM-generated synthetic datasets, which we release for future research.

cs.CR

Subadditivity of the log-Sobolev constant on convolutions

We present a general subadditivity inequality for log-Sobolev constants of convolution measures. As a corollary, we show that the log-Sobolev constant is monotone along the sequence of standardized convolutions in the central limit theorem.

math.FA

Private Estimation when Data and Privacy Demands are Correlated

Differential Privacy (DP) is the current gold-standard for ensuring privacy for statistical queries. Estimation problems under DP constraints appearing in the literature have largely focused on providing equal privacy to all users. We consider the problems of empirical mean estimation for univariate data and frequency estimation for categorical data, both subject to heterogeneous privacy constraints. Each user, contributing a sample to the dataset, is allowed to have a different privacy demand. The dataset itself is assumed to be worst-case and we study both problems under two different formulations -- first, where privacy demands and data may be correlated, and second, where correlations are weakened by random permutation of the dataset. We establish theoretical performance guarantees for our proposed algorithms, under both PAC error and mean-squared error. These performance guarantees translate to minimax optimality in several instances, and experiments confirm superior performance of our algorithms over other baseline techniques.

cs.LG

Online Assortment and Price Optimization Under Contextual Choice Models

We consider an assortment selection and pricing problem in which a seller has $N$ different items available for sale. In each round, the seller observes a $d$-dimensional contextual preference information vector for the user, and offers to the user an assortment of $K$ items at prices chosen by the seller. The user selects at most one of the products from the offered assortment according to a multinomial logit choice model whose parameters are unknown. The seller observes which, if any, item is chosen at the end of each round, with the goal of maximizing cumulative revenue over a selling horizon of length $T$. For this problem, we propose an algorithm that learns from user feedback and achieves a revenue regret of order $\widetilde{O}(d \sqrt{K T} / L_0 )$ where $L_0$ is the minimum price sensitivity parameter. We also obtain a lower bound of order $Ω(d \sqrt{T}/ L_0)$ for the regret achievable by any algorithm.

cs.LG

Enhancing Feature-Specific Data Protection via Bayesian Coordinate Differential Privacy

Local Differential Privacy (LDP) offers strong privacy guarantees without requiring users to trust external parties. However, LDP applies uniform protection to all data features, including less sensitive ones, which degrades performance of downstream tasks. To overcome this limitation, we propose a Bayesian framework, Bayesian Coordinate Differential Privacy (BCDP), that enables feature-specific privacy quantification. This more nuanced approach complements LDP by adjusting privacy protection according to the sensitivity of each feature, enabling improved performance of downstream tasks without compromising privacy. We characterize the properties of BCDP and articulate its connections with standard non-Bayesian privacy frameworks. We further apply our BCDP framework to the problems of private mean estimation and ordinary least-squares regression. The BCDP-based approach obtains improved accuracy compared to a purely LDP-based approach, without compromising on privacy.

cs.LG

Stochastic proof of the sharp symmetrized Talagrand inequality

We give a new proof of the sharp symmetrized form of Talagrand's transport-entropy inequality. Compared to stochastic proofs of other Gaussian functional inequalities, the new idea here is a certain coupling induced by time-reversed martingale representations.

math.PR

Stability of the Poincaré-Korn inequality

We resolve a question of Carrapatoso et al. on Gaussian optimality for the sharp constant in Poincaré-Korn inequalities, under a moment constraint. We also prove stability, showing that measures with near-optimal constant are quantitatively close to standard Gaussian.

math.AP

Stability of Klartag's improved Lichnerowicz inequality

In a recent work, Klartag gave an improved version of Lichnerowicz' spectral gap bound for uniformly log-concave measures, which improves on the classical estimate by taking into account the covariance matrix. We analyze the equality cases in Klartag's bound, showing that it can be further improved whenever the measure has no Gaussian factor. Additionally, we give a quantitative improvement for log-concave measures with finite Fisher information.

math.FA

Rigid characterizations of probability measures through independence, with applications

Three equivalent characterizations of probability measures through independence criteria are given. These characterizations lead to a family of Brascamp--Lieb-type inequalities for relative entropy, determine equilibrium states and sharp rates of convergence for certain linear Boltzmann-type dynamics, and unify an assortment of $L^2$ inequalities in probability.

math.PR

HWI inequalities in discrete spaces via couplings

HWI inequalities are interpolation inequalities relating entropy, Fisher information and optimal transport distances. We adapt an argument of Y. Wu for proving the Gaussian HWI inequality via a coupling argument to the discrete setting, establishing new interpolation inequalities for the discrete hypercube and the discrete torus. In particular, we obtain an improvement of the modified logarithmic Sobolev inequality for the discrete hypercube of Bobkov and Tetali.

math.PR

Mean Estimation Under Heterogeneous Privacy Demands

Differential Privacy (DP) is a well-established framework to quantify privacy loss incurred by any algorithm. Traditional formulations impose a uniform privacy requirement for all users, which is often inconsistent with real-world scenarios in which users dictate their privacy preferences individually. This work considers the problem of mean estimation, where each user can impose their own distinct privacy level. The algorithm we propose is shown to be minimax optimal and has a near-linear run-time. Our results elicit an interesting saturation phenomenon that occurs. Namely, the privacy requirements of the most stringent users dictate the overall error rates. As a consequence, users with less but differing privacy requirements are all given more privacy than they require, in equal amounts. In other words, these privacy-indifferent users are given a nontrivial degree of privacy for free, without any sacrifice in the performance of the estimator.

cs.CR

Mean Estimation Under Heterogeneous Privacy: Some Privacy Can Be Free

Differential Privacy (DP) is a well-established framework to quantify privacy loss incurred by any algorithm. Traditional DP formulations impose a uniform privacy requirement for all users, which is often inconsistent with real-world scenarios in which users dictate their privacy preferences individually. This work considers the problem of mean estimation under heterogeneous DP constraints, where each user can impose their own distinct privacy level. The algorithm we propose is shown to be minimax optimal when there are two groups of users with distinct privacy levels. Our results elicit an interesting saturation phenomenon that occurs as one group's privacy level is relaxed, while the other group's privacy level remains constant. Namely, after a certain point, further relaxing the privacy requirement of the former group does not improve the performance of the minimax optimal mean estimator. Thus, the central server can offer a certain degree of privacy without any sacrifice in performance.

cs.CR

Entropy Inequalities and Gaussian Comparisons

We establish a general class of entropy inequalities that take the concise form of Gaussian comparisons. The main result unifies many classical and recent results, including the Shannon-Stam inequality, the Brunn-Minkowski inequality, the Zamir-Feder inequality, the Brascamp-Lieb and Barthe inequalities, the Anantharam-Jog-Nair inequality, and others.

cs.IT

Equality cases in the Anantharam-Jog-Nair inequality

Anantharam, Jog and Nair recently unified the Shannon-Stam inequality and the entropic form of the Brascamp-Lieb inequalities under a common inequality. They left open the problems of extremizability and characterization of extremizers. Both questions are resolved in the present paper.

cs.IT

Linear Models are Most Favorable among Generalized Linear Models

We establish a nonasymptotic lower bound on the $L_2$ minimax risk for a class of generalized linear models. It is further shown that the minimax risk for the canonical linear model matches this lower bound up to a universal constant. Therefore, the canonical linear model may be regarded as most favorable among the considered class of generalized linear models (in terms of minimax risk). The proof makes use of an information-theoretic Bayesian Cramér-Rao bound for log-concave priors, established by Aras et al. (2019).

math.ST

Euclidean Forward-Reverse Brascamp-Lieb Inequalities: Finiteness, Structure and Extremals

A new proof is given for the fact that centered gaussian functions saturate the Euclidean forward-reverse Brascamp-Lieb inequalities, extending the Brascamp-Lieb and Barthe theorems. A duality principle for best constants is also developed, which generalizes the fact that the best constants in the Brascamp-Lieb and Barthe inequalities are equal. Finally, as the title hints, the main results concerning finiteness, structure and gaussian-extremizability for the Brascamp-Lieb inequality due to Bennett, Carbery, Christ and Tao are generalized to the setting of the forward-reverse Brascamp-Lieb inequality.

math.FA