SearcharxivSearch

arXiv subjects

Spencer Giddens

Publications and source records attributed to Spencer Giddens.

4 recordsLinked to original sources

SAFES: Sequential Privacy and Fairness Enhancing Data Synthesis for Responsible AI

As data-driven and AI-based decision making gains widespread adoption across disciplines, it is crucial that both data privacy and decision fairness are appropriately addressed. Although differential privacy (DP) provides a robust framework for guaranteeing privacy and methods are available to improve fairness, most prior work treats the two concerns separately. Even though there are existing approaches that consider privacy and fairness simultaneously, they typically focus on a single specific learning task, limiting their generalizability. In response, we introduce SAFES, a Sequential PrivAcy and Fairness Enhancing data Synthesis procedure that sequentially combines DP data synthesis with a fairness-aware data preprocessing step. SAFES allows users flexibility in navigating the privacy-fairness-utility trade-offs. We illustrate SAFES with different DP synthesizers and fairness-aware data preprocessing methods and run extensive experiments on multiple real datasets to examine the privacy-fairness-utility trade-offs of synthetic data generated by SAFES. Empirical evaluations demonstrate that for reasonable privacy loss, SAFES-generated synthetic data can achieve significantly improved fairness metrics with relatively low utility loss.

cs.LG

DPpack: An R Package for Differentially Private Statistical Analysis and Machine Learning

Differential privacy (DP) is the state-of-the-art framework for guaranteeing privacy for individuals when releasing aggregated statistics or building statistical/machine learning models from data. We develop the open-source R package DPpack that provides a large toolkit of differentially private analysis. The current version of DPpack implements three popular mechanisms for ensuring DP: Laplace, Gaussian, and exponential. Beyond that, DPpack provides a large toolkit of easily accessible privacy-preserving descriptive statistics functions. These include mean, variance, covariance, and quantiles, as well as histograms and contingency tables. Finally, DPpack provides user-friendly implementation of privacy-preserving versions of logistic regression, SVM, and linear regression, as well as differentially private hyperparameter tuning for each of these models. This extensive collection of implemented differentially private statistics and models permits hassle-free utilization of differential privacy principles in commonly performed statistical analysis. We plan to continue developing DPpack and make it more comprehensive by including more differentially private machine learning techniques, statistical modeling and inference in the future.

stat.ML

A Differentially Private Weighted Empirical Risk Minimization Procedure and its Application to Outcome Weighted Learning

Data used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information. While differential privacy (DP) provides mathematically provable bounds to protect such data, previous work has focused almost exclusively on unweighted ERM. We consider weighted ERM (wERM) -- an important generalization where individual contributions to the objective function vary. We propose the first DP algorithm for general wERM with formal privacy guarantees and derive both its empirical and population utility bounds. Crucially, this general wERM framework provides a pathway for deriving privacy-preserving learning methods for individualized treatment rules, including the popular outcome-weighted learning (OWL) approach. We evaluate DP-wERM applied to OWL in simulated and real data experiments. Our empirical results demonstrate that training OWL models via wERM provides strong DP guarantees while maintaining robust performance, proving the method is practical for sensitive, real-world data.

stat.ML

Methodological reconstruction of historical seismic events from anecdotal accounts of destructive tsunamis: a case study for the great 1852 Banda arc mega-thrust earthquake and tsunami

We demonstrate the efficacy of a Bayesian statistical inversion framework for reconstructing the likely characteristics of large pre-instrumentation earthquakes from historical records of tsunami observations. Our framework is designed and implemented for the estimation of the location and magnitude of seismic events from anecdotal accounts of tsunamis including shoreline wave arrival times, heights, and inundation lengths over a variety of spatially separated observation locations. As an initial test case we use our framework to reconstruct the great 1852 earthquake and tsunami of eastern Indonesia. Relying on the assumption that these observations were produced by a subducting thrust event, the posterior distribution indicates that the observables were the result of a massive mega-thrust event with magnitude near 8.8 Mw and a likely rupture zone in the north-eastern Banda arc. The distribution of predicted epicentral locations overlaps with the largest major seismic gap in the region as indicated by instrumentally recorded seismic events. These results provide a geologic and seismic context for hazard risk assessment in coastal communities experiencing growing population and urbanization in Indonesia. In addition, the methodology demonstrated here highlights the potential for applying a Bayesian approach to enhance understanding of the seismic history of other subduction zones around the world.

physics.geo-ph