SearcharxivSearch

arXiv subjects

Adam Smith

Publications and source records attributed to Adam Smith.

At least 91 records · Page 5Linked to original sources

Guaranteed Validity for Empirical Approaches to Adaptive Data Analysis

We design a general framework for answering adaptive statistical queries that focuses on providing explicit confidence intervals along with point estimates. Prior work in this area has either focused on providing tight confidence intervals for specific analyses, or providing general worst-case bounds for point estimates. Unfortunately, as we observe, these worst-case bounds are loose in many settings --- often not even beating simple baselines like sample splitting. Our main contribution is to design a framework for providing valid, instance-specific confidence intervals for point estimates that can be generated by heuristics. When paired with good heuristics, this method gives guarantees that are orders of magnitude better than the best worst-case bounds. We provide a Python library implementing our method.

cs.LG

Simulating quantum many-body dynamics on a current digital quantum computer

Universal quantum computers are potentially an ideal setting for simulating many-body quantum dynamics that is out of reach for classical digital computers. We use state-of-the-art IBM quantum computers to study paradigmatic examples of condensed matter physics -- we simulate the effects of disorder and interactions on quantum particle transport, as well as correlation and entanglement spreading. Our benchmark results show that the quality of the current machines is below what is necessary for quantitatively accurate continuous time dynamics of observables and reachable system sizes are small comparable to exact diagonalization. Despite this, we are successfully able to demonstrate clear qualitative behaviour associated with localization physics and many-body interaction effects.

quant-ph

Manipulation Attacks in Local Differential Privacy

Local differential privacy is a widely studied restriction on distributed algorithms that collect aggregates about sensitive user data, and is now deployed in several large systems. We initiate a systematic study of a fundamental limitation of locally differentially private protocols: they are highly vulnerable to adversarial manipulation. While any algorithm can be manipulated by adversaries who lie about their inputs, we show that any non-interactive locally differentially private protocol can be manipulated to a much greater extent. Namely, when the privacy level is high or the input domain is large, an attacker who controls a small fraction of the users in the protocol can completely obscure the distribution of the users' inputs. We also show that existing protocols differ greatly in their resistance to manipulation, even when they offer the same accuracy guarantee with honest execution. Our results suggest caution when deploying local differential privacy and reinforce the importance of efficient cryptographic techniques for emulating mechanisms from central differential privacy in distributed settings.

cs.DS

Logarithmic spreading of out-of-time-ordered correlators without many-body localization

Out-of-time-ordered correlators (OTOCs) describe information scrambling under unitary time evolution, and provide a useful probe of the emergence of quantum chaos. Here we calculate OTOCs for a model of disorder-free localization whose exact solubility allows us to study long-time behaviour in large systems. Remarkably, we observe logarithmic spreading of correlations, qualitatively different to both thermalizing and Anderson localized systems. Rather, such behaviour is normally taken as a signature of many-body localization, so that our findings for an essentially non-interacting model are surprising. We provide an explanation for this unusual behaviour, and suggest a novel Loschmidt echo protocol as a probe of correlation spreading. We show that the logarithmic spreading of correlations probed by this protocol is a generic feature of localized systems, with or without interactions.

cond-mat.str-el

Distributed Differential Privacy via Shuffling

We consider the problem of designing scalable, robust protocols for computing statistics about sensitive data. Specifically, we look at how best to design differentially private protocols in a distributed setting, where each user holds a private datum. The literature has mostly considered two models: the "central" model, in which a trusted server collects users' data in the clear, which allows greater accuracy; and the "local" model, in which users individually randomize their data, and need not trust the server, but accuracy is limited. Attempts to achieve the accuracy of the central model without a trusted server have so far focused on variants of cryptographic MPC, which limits scalability. In this paper, we initiate the analytic study of a shuffled model for distributed differentially private algorithms, which lies between the local and central models. This simple-to-implement model, a special case of the ESA framework of [Bittau et al., '17], augments the local model with an anonymous channel that randomly permutes a set of user-supplied messages. For sum queries, we show that this model provides the power of the central model while avoiding the need to trust a central server and the complexity of cryptographic secure function evaluation. More generally, we give evidence that the power of the shuffled model lies strictly between those of the central and local models: for a natural restriction of the model, we show that shuffled protocols for a widely studied selection problem require exponentially higher sample complexity than do central-model protocols.

cs.CR

The Structure of Optimal Private Tests for Simple Hypotheses

Hypothesis testing plays a central role in statistical inference, and is used in many settings where privacy concerns are paramount. This work answers a basic question about privately testing simple hypotheses: given two distributions $P$ and $Q$, and a privacy level $\varepsilon$, how many i.i.d. samples are needed to distinguish $P$ from $Q$ subject to $\varepsilon$-differential privacy, and what sort of tests have optimal sample complexity? Specifically, we characterize this sample complexity up to constant factors in terms of the structure of $P$ and $Q$ and the privacy level $\varepsilon$, and show that this sample complexity is achieved by a certain randomized and clamped variant of the log-likelihood ratio test. Our result is an analogue of the classical Neyman-Pearson lemma in the setting of private hypothesis testing. We also give an application of our result to the private change-point detection. Our characterization applies more generally to hypothesis tests satisfying essentially any notion of algorithmic stability, which is known to imply strong generalization bounds in adaptive data analysis, and thus our results have applications even when privacy is not a primary concern.

cs.DS

Graph Oracle Models, Lower Bounds, and Gaps for Parallel Stochastic Optimization

We suggest a general oracle-based framework that captures different parallel stochastic optimization settings described by a dependency graph, and derive generic lower bounds in terms of this graph. We then use the framework and derive lower bounds for several specific parallel optimization settings, including delayed updates and parallel processing with intermittent communication. We highlight gaps between lower and upper bounds on the oracle complexity, and cases where the "natural" algorithms are not known to be optimal.

math.OC

From Soft Classifiers to Hard Decisions: How fair can we be?

A popular methodology for building binary decision-making classifiers in the presence of imperfect information is to first construct a non-binary "scoring" classifier that is calibrated over all protected groups, and then to post-process this score to obtain a binary decision. We study the feasibility of achieving various fairness properties by post-processing calibrated scores, and then show that deferring post-processors allow for more fairness conditions to hold on the final decision. Specifically, we show: 1. There does not exist a general way to post-process a calibrated classifier to equalize protected groups' positive or negative predictive value (PPV or NPV). For certain "nice" calibrated classifiers, either PPV or NPV can be equalized when the post-processor uses different thresholds across protected groups, though there exist distributions of calibrated scores for which the two measures cannot be both equalized. When the post-processing consists of a single global threshold across all groups, natural fairness properties, such as equalizing PPV in a nontrivial way, do not hold even for "nice" classifiers. 2. When the post-processing is allowed to `defer' on some decisions (that is, to avoid making a decision by handing off some examples to a separate process), then for the non-deferred decisions, the resulting classifier can be made to equalize PPV, NPV, false positive rate (FPR) and false negative rate (FNR) across the protected groups. This suggests a way to partially evade the impossibility results of Chouldechova and Kleinberg et al., which preclude equalizing all of these measures simultaneously. We also present different deferring strategies and show how they affect the fairness properties of the overall system. We evaluate our post-processing techniques using the COMPAS data set from 2016.

cs.LG

Private Algorithms Can Always Be Extended

We consider the following fundamental question on $ε$-differential privacy. Consider an arbitrary $ε$-differentially private algorithm defined on a subset of the input space. Is it possible to extend it to an $ε'$-differentially private algorithm on the whole input space for some $ε'$ comparable with $ε$? In this note we answer affirmatively this question for $ε'=2ε$. Our result applies to every input metric space and space of possible outputs. This result originally appeared in a recent paper by the authors [BCSZ18]. We present a self-contained version in this note, in the hopes that it will be broadly useful.

math.ST

Revealing Network Structure, Confidentially: Improved Rates for Node-Private Graphon Estimation

Motivated by growing concerns over ensuring privacy on social networks, we develop new algorithms and impossibility results for fitting complex statistical models to network data subject to rigorous privacy guarantees. We consider the so-called node-differentially private algorithms, which compute information about a graph or network while provably revealing almost no information about the presence or absence of a particular node in the graph. We provide new algorithms for node-differentially private estimation for a popular and expressive family of network models: stochastic block models and their generalization, graphons. Our algorithms improve on prior work, reducing their error quadratically and matching, in many regimes, the optimal nonprivate algorithm. We also show that for the simplest random graph models ($G(n,p)$ and $G(n,m)$), node-private algorithms can be qualitatively more accurate than for more complex models---converging at a rate of $\frac{1}{ε^2 n^{3}}$ instead of $\frac{1}{ε^2 n^2}$. This result uses a new extension lemma for differentially private algorithms that we hope will be broadly useful.

math.ST

Dynamics of a Lattice Gauge Theory with Fermionic Matter -- Minimal Quantum Simulator with Time-Dependent Impurities in Ultracold Gases

We propose a minimal model to study the real-time dynamics of a $\mathbb{Z}_2$ lattice gauge theory coupled to fermionic matter in a cold atom quantum simulator setup. We show that dynamical correlators of the gauge fields can be measured in experiments studying the time-evolution of two pairs of impurities, and suggest the protocol for implementing the model in cold atom experiments. Further, we discuss a number of unexpected features found in the integrable limit of the model, as well as its extensions to a non-integrable case. A potential experimental implementation of our model in the latter regime would allow one to simulate strongly-interacting lattice gauge theories beyond current capabilities of classical computers.

cond-mat.quant-gas

Dynamical Localization in $\mathbb{Z}_2$ Lattice Gauge Theories

We study quantum quenches in two-dimensional lattice gauge theories with fermions coupled to dynamical $\mathbb{Z}_2$ gauge fields. Through the identification of an extensive set of conserved quantities, we propose a generic mechanism of charge localization in the absence of quenched disorder both in the Hamiltonian and in the initial states. We provide diagnostics of this localization through a set of experimentally relevant dynamical measures, entanglement measures, as well as spectral properties of the model. One of the defining features of the models that we study is a binary nature of emergent disorder, related to $\mathbb{Z}_2$ degrees of freedom. This results in a qualitatively different behaviour in the strong disorder limit compared to typically studied models of localization. For example it gives rise to a possibility of a delocalization transition via a mechanism of quantum percolation in dimensions higher than 1D. We highlight the importance of our general phenomenology to questions related to dynamics of defects in Kitaev's toric code, and to quantum quenches in Hubbard models. While the simplest models we consider are effectively non-interacting, we also include interactions leading to many-body localization-like logarithmic entanglement growth. Finally, we consider effects of interactions that generate dynamics for conserved charges, which gives rise to only transient localization behaviour, or quasi-many-body-localization.

cond-mat.str-el

The Limits of Post-Selection Generalization

While statistics and machine learning offers numerous methods for ensuring generalization, these methods often fail in the presence of adaptivity---the common practice in which the choice of analysis depends on previous interactions with the same dataset. A recent line of work has introduced powerful, general purpose algorithms that ensure post hoc generalization (also called robust or post-selection generalization), which says that, given the output of the algorithm, it is hard to find any statistic for which the data differs significantly from the population it came from. In this work we show several limitations on the power of algorithms satisfying post hoc generalization. First, we show a tight lower bound on the error of any algorithm that satisfies post hoc generalization and answers adaptively chosen statistical queries, showing a strong barrier to progress in post selection data analysis. Second, we show that post hoc generalization is not closed under composition, despite many examples of such algorithms exhibiting strong composition properties.

cs.LG

Evolving Mario Levels in the Latent Space of a Deep Convolutional Generative Adversarial Network

Generative Adversarial Networks (GANs) are a machine learning approach capable of generating novel example outputs across a space of provided training examples. Procedural Content Generation (PCG) of levels for video games could benefit from such models, especially for games where there is a pre-existing corpus of levels to emulate. This paper trains a GAN to generate levels for Super Mario Bros using a level from the Video Game Level Corpus. The approach successfully generates a variety of levels similar to one in the original corpus, but is further improved by application of the Covariance Matrix Adaptation Evolution Strategy (CMA-ES). Specifically, various fitness functions are used to discover levels within the latent space of the GAN that maximize desired properties. Simple static properties are optimized, such as a given distribution of tile types. Additionally, the champion A* agent from the 2009 Mario AI competition is used to assess whether a level is playable, and how many jumping actions are required to beat it. These fitness functions allow for the discovery of levels that exist within the space of examples designed by experts, and also guide the search towards levels that fulfill one or more specified objectives.

cs.AI

Absence of Ergodicity without Quenched Disorder: from Quantum Disentangled Liquids to Many-Body Localization

We study the time evolution after a quantum quench in a family of models whose degrees of freedom are fermions coupled to spins, where quenched disorder appears neither in the Hamiltonian parameters nor in the initial state. Focussing on the behaviour of entanglement, both spatial and between subsystems, we show that the model supports a state exhibiting combined area/volume law entanglement, being characteristic of the quantum disentangled liquid. This behaviour appears for one set of variables, which is related via a duality mapping to another set, where this structure is absent. Upon adding density interactions between the fermions, we identify an exact mapping to an XXZ spin-chain in a random binary magnetic field, thereby establishing the existence of many-body localization with its logarithmic entanglement growth in a fully disorder-free system.

cond-mat.str-el

The role of tachysterol in vitamin D photosynthesis - A non-adiabatic molecular dynamics study

To investigate the role of tachysterol in the photophysical/chemical regulation of vitamin D photosynthesis, we studied its electronic absorption properties and excited state dynamics using time-dependent density functional theory (TDDFT), coupled cluster theory (CC2), and non-adiabatic molecular dynamics. In excellent agreement with experiments, the simulated electronic spectrum shows a broad absorption band covering the spectra of the other vitamin D photoisomers. The broad band stems from the spectral overlap of four different ground state rotamers. After photoexcitation, the first excited singlet state (S1) decays within 882 fs. The S1 dynamics is characterized by a strong twisting of the central double bond. 96% of all trajectories relax without chemical transformation to the ground state. In 2.3 % of the trajectories we observed [1,5]-sigmatropic hydrogen shift forming the partly deconjugated toxisterol D1. 1.4 % previtamin D formation is observed via hula-twist double bond isomerization. We find a strong dependence between photoreactivity and dihedral angle conformation: hydrogen shift only occurs in cEc and cEt rotamers and double bond isomerization occurs mainly in cEc rotamers. Our study confirms the hypothesis that cEc rotamers are more prone to previtamin D formation than other isomers. We also observe the formation of a cyclobutene-toxisterol in the hot ground state (0.7 %). Due to its strong absorption and unreactive behavior, tachysterol acts mainly as a sun shield suppressing previtamin D formation. Tachysterol shows stronger toxisterol formation than previtamin D. Absorption of low energy UV light by the cEc rotamer can lead to previtamin D formation. Our study reinforces a recent hypothesis that tachysterol can act as a previtamin D source when only low energy ultraviolet light is available, as it is the case in winter or in the morning and evening hours of the day.

physics.chem-ph

Disorder-Free Localization

The venerable phenomena of Anderson localization, along with the much more recent many-body localization, both depend crucially on the presence of disorder. The latter enters either in the form of quenched disorder in the parameters of the Hamiltonian, or through a special choice of a disordered initial state. Here we present a model with localization arising in a very simple, completely translationally invariant quantum model, with only local interactions between spins and fermions. By identifying an extensive set of conserved quantities, we show that the system generates purely dynamically its own disorder, which gives rise to localization of fermionic degrees of freedom. Our work gives an answer to a decades old question whether quenched disorder is a necessary condition for localization. It also offers new insights into the physics of many-body localization, lattice gauge theories, and quantum disentangled liquids.

cond-mat.str-el

Information, Privacy and Stability in Adaptive Data Analysis

Traditional statistical theory assumes that the analysis to be performed on a given data set is selected independently of the data themselves. This assumption breaks downs when data are re-used across analyses and the analysis to be performed at a given stage depends on the results of earlier stages. Such dependency can arise when the same data are used by several scientific studies, or when a single analysis consists of multiple stages. How can we draw statistically valid conclusions when data are re-used? This is the focus of a recent and active line of work. At a high level, these results show that limiting the information revealed by earlier stages of analysis controls the bias introduced in later stages by adaptivity. Here we review some known results in this area and highlight the role of information-theoretic concepts, notably several one-shot notions of mutual information.

cs.LG