Searcharxiv⌕ Search

arXiv subjects

Nianqiao Phyllis Ju

Publications and source records attributed to Nianqiao Phyllis Ju.

5 recordsLinked to original sources

Statistical Properties of Nonparametric MLE under Laplace Noise

Local differential privacy (LDP) protects individuals in a dataset by perturbing each measurement before release. For real-valued data, a widely used mechanism is additive Laplace noise. We study the problem of estimating the distribution of the latent confidential data from the privatized observations via the nonparametric maximum likelihood estimator (NPMLE) under an i.i.d. sampling model. We first show that under the Laplace convolution model, the NPMLE admits a finite-dimensional reformulation in which the support is restricted to the observation set. This reduces the original infinite-dimensional optimization over all mixing distributions to an $n$-dimensional convex optimization over mixture weights, where $n$ is the sample size. We then study the statistical convergence of the NPMLE under the 1-Wasserstein distance, and explicitly connect its convergence rate with the privacy noise scale. Allowing the privacy noise level to change with the sample size, our analysis shows that the NPMLE remains consistent when the Laplace noise grows at a rate slower than $n^{3/16}$. Conversely, when the Laplace noise is of the order $\sqrt n$ or larger, no estimator can achieve uniformly consistent recovery of the latent distribution.

math.ST↗

Redistricting from the Bottom Up: Sampling Communities of Interest with Differential Privacy

Independent Redistricting Commissions (IRCs) are a promising tool for bottom-up redistricting, but their public testimony processes are vulnerable to adversarial manipulation. We propose using differential privacy to draw redistricting plans that incorporate community of interest (COI) testimonies while remaining robust to adversarial input. Treating individual testimonies as data points, we use the marked edge walk to sample from differentially private distributions of redistricting plans via the exponential mechanism. We introduce two score functions and demonstrate that both can be targeted by MEW across a range of privacy budgets. Applying this method to Missouri's mid-cycle redistricting using 808 COI testimonies, we show that COI-informed sampling outperforms an uninformed baseline and the enacted plan. An adversarial experiment demonstrates that the method can be robust to attacks under certain privacy budgets and may perform better in practice than formal group privacy guarantees imply. We also find that stronger COI preservation tends to spread minority and Democratic representation more evenly across districts.

cs.CY↗

SOMA: A Novel Sampler for Bayesian Inference from Privatized Data

Making valid statistical inferences from privatized data is a key challenge in modern analysis. In Bayesian settings, data augmentation MCMC (DAMCMC) methods impute unobserved confidential data given noisy privatized summaries, enabling principled uncertainty quantification. However, standard DAMCMC often suffers from slow mixing due to component-wise Metropolis-within-Gibbs updates. We propose the Single-Offer-Multiple-Attempts (SOMA) sampler. This novel algorithm improves acceptance rates by generating a single proposal and simultaneously evaluating its suitability to replace all components. By sharing proposals across components, SOMA rejects fewer proposal points. We prove lower bounds on SOMA's acceptance probability and establish convergence rates in the two-component case. Experiments on synthetic and real census data with linear regression and other models confirm SOMA's efficiency gains.

stat.ME↗

Simulation-based Bayesian Inference from Privacy Protected Data

Many modern statistical analysis and machine learning applications require training models on sensitive user data. Under a formal definition of privacy protection, differentially private algorithms inject calibrated noise into the confidential data or during the data analysis process to produce privacy-protected datasets or queries. However, restricting access to only privatized data during statistical analysis makes it computationally challenging to make valid statistical inferences. In this work, we propose simulation-based inference methods from privacy-protected datasets. In addition to sequential Monte Carlo approximate Bayesian computation, we adopt neural conditional density estimators as a flexible family of distributions to approximate the posterior distribution of model parameters given the observed private query results. We illustrate our methods on discrete time-series data under an infectious disease model and with ordinary linear regression models. Illustrating the privacy-utility trade-off, our experiments and analysis demonstrate the necessity and feasibility of designing valid statistical inference procedures to correct for biases introduced by the privacy-protection mechanisms.

stat.ML↗

dapper: Data Augmentation for Private Posterior Estimation in R

This paper serves as a reference and introduction to using the R package dapper. dapper encodes a sampling framework which allows exact Markov chain Monte Carlo simulation of parameters and latent variables in a statistical model given privatized data. The goal of this package is to fill an urgent need by providing applied researchers with a flexible tool to perform valid Bayesian inference on data protected by differential privacy, allowing them to properly account for the noise introduced for privacy protection in their statistical analysis. dapper offers a significant step forward in providing general-purpose statistical inference tools for privatized data.

stat.CO↗