Searcharxiv⌕ Search

arXiv · 2610.05867

Learning Sparse Support under Differential Privacy: Adaptive Algorithms and Minimax Limits

Abstract

We study exact support recovery under $(ε,δ)$-differential privacy in sparse high-dimensional linear regression. We introduce Saturated Propose-test-release, a general mechanism that privately releases the output of a discrete selector with probability one once its stability certificate reaches a finite threshold. Exploiting the coordinatewise geometry of the LASSO, we construct a computable support-stability score. The resulting computationally efficient Saturated LASSO satisfies worst-case $(ε,δ)$-differential privacy and achieves exact support recovery with high probability under explicit regularity and beta-min conditions. Maximizing sparsity-indexed certificates yields an adaptive procedure requiring no sparsity knowledge and having exactly the same finite-sample exact-recovery risk as the oracle fixed-sparsity procedure under common public tuning parameters. We also establish a minimax lower bound explicitly tracking $δ$: under its recovery conditions, Saturated LASSO is minimax optimal up to logarithmic factors in $n$ and $1/δ$ uniformly over $0 < δ\leq ε/16$; in the broad moderate-$δ$ regime, it further matches the lower-bound $δ$-dependence. A complementary information-theoretic construction with known sparsity attains the lower-bound rates up to constant factors under independent Gaussian design, at exponential computational cost. Simulations and a semi-synthetic study using public American Community Survey covariates illustrate the numerical performance of the proposed procedures.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Jia Gu, T. Tony Cai. 2026-10-05. Learning Sparse Support under Differential Privacy: Adaptive Algorithms and Minimax Limits. https://arxiv.org/abs/2610.05867

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

A resolution of the Borel-Kolmogorov Paradox via the Maximum Entropy Principle

Bayesian updating routinely conditions on events of prior probability zero, such as an exact observation of a continuous parameter or a parameter confined to a submanifold. Two equally natural parametrizations of the same event can then return different posteriors, which is the Borel--Kolmogorov paradox. We add a metric to the probabilistic model and define the posterior given a closed null set as the limit of Maximum Entropy solutions under a vanishing constraint on the distance to that set. We give a sufficient condition for existence and prove that, given this metric, the posterior is unique and invariant under measure-preserving isometries. It recovers the textbook Bayes formulas where these apply unambiguously, on positive-measure sets and in Euclidean models. Each stage of the limit has an explicit density that depends only on the distance to the conditioning set, so the posterior can be computed with standard tools for Kullback--Leibler minimization. The construction does not remove the choice of a metric but makes it an explicit modelling input. In a controlled experiment where the geometry of the measurement is known, we show that a Bayesian analysis ignoring this geometry can yield uncalibrated posteriors. On real palaeomagnetic data, where the geometry is disputed, we show how candidate metrics can be compared within the same Bayesian framework.

math.ST↗

A Two-Sample Test on Weighted Persistence Intensity Functions in Topological Data Analysis

Persistence intensity functions provide interpretable and informative first-order summaries of random persistence diagram distributions. We study two-sample testing for equality of persistence intensity functions, allowing the underlying diagram distributions to differ under the null. We construct a weighted-kernel statistic as an unbiased estimator of the squared reproducing-kernel Hilbert space distance between the corresponding weighted intensity embeddings, and calibrate it by studentization. For shrinking bandwidths, we establish uniform asymptotic normality under the null and thereby obtain asymptotic Type I error control. Its power is characterized in terms of the $L^2$ discrepancy between the weighted intensity functions. To accommodate persistence diagrams with possibly unbounded cardinality, we introduce regularity conditions that control the effect of cardinality variation and yield the desired moment bounds for the test statistic. We further show that every probability density on $\{(x,y)\in\mathbb{R}^2:y>x > 0\}$ can be realized as the persistence intensity function of a random diagram. Using these results, we establish that the proposed test attains minimax-optimal separation rates over anisotropic Sobolev balls. Lastly, since the optimal bandwidth is not directly accessible in practice, we adapt a bandwidth aggregation framework.

math.ST↗

Human-Anchored Inference for Ranking New Models with Large Language Model Judges

Human pairwise comparisons provide a reference for evaluating large language models (LLMs), but collecting sufficient judgments for each new release is costly and time-consuming. LLM judges offer a scalable alternative, although their comparisons may differ systematically from human preferences and across judges. We study the ranking of a new model that has received LLM-judge comparisons but no human comparisons. We propose ANCHOR (ANchored Comparisons for Human-reference inference with Orthogonal Riesz correction), which uses historical human and LLM comparisons to learn judge-specific sensitivities to human score differences and feature-dependent judge biases. These estimates are then used to infer the new model's human-reference score from its judge comparisons. The framework allows the feature distribution to change between historical and new-model comparisons. For inference, we construct a Neyman-orthogonal estimator through a joint Riesz correction that removes the first-order effects of estimating the historical human scores, judge sensitivities, and bias functions. We establish identification, convergence rates, and asymptotic normality with consistently estimable variance, and show that ANCHOR attains the semiparametric efficiency bound. Simulations demonstrate gains in score estimation and ranking accuracy. On Chatbot Arena, ANCHOR achieves the lowest score RMSE and insertion MAE among competing methods, with narrower score intervals on average.

math.ST↗