SearcharxivSearch

arXiv subjects

Paul Francis

Publications and source records attributed to Paul Francis.

15 recordsLinked to original sources

Towards Better Attribute Inference Vulnerability Measures

The purpose of anonymizing structured data is to protect the privacy of individuals in the data while retaining the statistical properties of the data. An important class of attack on anonymized data is attribute inference, where an attacker infers the value of an unknown attribute of a target individual given knowledge of one or more known attributes. A major limitation of recent attribute inference measures is that they do not take recall into account, only precision. It is often the case that attacks target only a fraction of individuals, for instance data outliers. Incorporating recall, however, substantially complicates the measure, because one must determine how to combine recall and precision in a composite measure for both the attack and baseline. This paper presents the design and implementation of an attribute inference measure that incorporates both precision and recall. Our design also improves on how the baseline attribute inference is computed. In experiments using a generic best row match attack on moderately-anonymized microdata, we show that in over 25\% of the attacks, our approach correctly labeled the attack to be at risk while the prior approach incorrectly labeled the attack to be safe.

cs.CR

A Consensus Privacy Metrics Framework for Synthetic Data

Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.

cs.CR

A Comparison of SynDiffix Multi-table versus Single-table Synthetic Data

SynDiffix is a new open-source tool for structured data synthesis. It has anonymization features that allow it to generate multiple synthetic tables while maintaining strong anonymity. Compared to the more common single-table approach, multi-table leads to more accurate data, since only the features of interest for a given analysis need be synthesized. This paper compares SynDiffix with 15 other commercial and academic synthetic data techniques using the SDNIST analysis framework, modified by us to accommodate multi-table synthetic data. The results show that SynDiffix is many times more accurate than other approaches for low-dimension tables, but somewhat worse than the best single-table techniques for high-dimension tables.

cs.CR

Towards more accurate and useful data anonymity vulnerability measures

The purpose of anonymizing structured data is to protect the privacy of individuals in the data while retaining the statistical properties of the data. There is a large body of work that examines anonymization vulnerabilities. Focusing on strong anonymization mechanisms, this paper examines a number of prominent attack papers and finds several problems, all of which lead to overstating risk. First, some papers fail to establish a correct statistical inference baseline (or any at all), leading to incorrect measures. Notably, the reconstruction attack from the US Census Bureau that led to a redesign of its disclosure method made this mistake. We propose the non-member framework, an improved method for how to compute a more accurate inference baseline, and give examples of its operation. Second, some papers don't use a realistic membership base rate, leading to incorrect precision measures if precision is reported. Third, some papers unnecessarily report measures in such a way that it is difficult or impossible to assess risk. Virtually the entire literature on membership inference attacks, dozens of papers, make one or both of these errors. We propose that membership inference papers report precision/recall values using a representative range of base rates.

cs.CR

SynDiffix: More accurate synthetic structured data

This paper introduces SynDiffix, a mechanism for generating statistically accurate, anonymous synthetic data for structured data. Recent open source and commercial systems use Generative Adversarial Networks or Transformed Auto Encoders to synthesize data, and achieve anonymity through overfitting-avoidance. By contrast, SynDiffix exploits traditional mechanisms of aggregation, noise addition, and suppression among others. Compared to CTGAN, ML models generated from SynDiffix are twice as accurate, marginal and column pairs data quality is one to two orders of magnitude more accurate, and execution time is two orders of magnitude faster. Compared to the best commercial product we measured (MostlyAI), ML model accuracy is comparable, marginal and pairs accuracy is 5 to 10 times better, and execution time is an order of magnitude faster. Similar to the other approaches, SynDiffix anonymization is very strong. This paper describes SynDiffix and compares its performance with other popular open source and commercial systems.

cs.CR

A Note on the Misinterpretation of the US Census Re-identification Attack

In 2018, the US Census Bureau designed a new data reconstruction and re-identification attack and tested it against their 2010 data release. The specific attack executed by the Bureau allows an attacker to infer the race and ethnicity of respondents with average 75% precision for 85% of the respondents, assuming that the attacker knows the correct age, sex, and address of the respondents. They interpreted the attack as exceeding the Bureau's privacy standards, and so introduced stronger privacy protections for the 2020 Census in the form of the TopDown Algorithm (TDA). This paper demonstrates that race and ethnicity can be inferred from the TDA-protected census data with substantially better precision and recall, using less prior knowledge: only the respondents' address. Race and ethnicity can be inferred with average 75% precision for 98% of the respondents, and can be inferred with 100% precision for 11% of the respondents. The inference is done by simply assuming that the race/ethnicity of the respondent is that of the majority race/ethnicity for the respondent's census block. The conclusion to draw from this simple demonstration is NOT that the Bureau's data releases lack adequate privacy protections. Indeed it is the purpose of the data releases to allow this kind of inference. The problem, rather, is that the Bureau's criteria for measuring privacy is flawed and overly pessimistic.

cs.CR

Diffix Elm: Simple Diffix

Historically, strong data anonymization requires substantial domain expertise and custom design for the given data set and use case. Diffix is an anonymization framework designed to make strong data anonymization available to non-experts. This paper describes Diffix Elm, a version of Diffix that is very easy to use at the expense of query features. We describe Diffix Elm, and show that it provides strong anonymity based on the General Data Protection Regulation (GDPR) criteria. This document is the third version of Diffix Elm. The second version added ceiling, round, and bucket\_width functions (in addition to floor). This document adds the ability to protect multiple different kinds of protected entities (a feature not found in earlier versions of Diffix). It also adds counting distinct values for any column (rather than only the AID column).

cs.CR

Diffix-Birch: Extending Diffix-Aspen

A longstanding open problem is that of how to get high quality statistics through direct queries to databases containing information about individuals without revealing information specific to those individuals. Diffix is a framework for anonymous database query that adds noise based on the filter conditions in the query. A previous paper described the first version, called diffix-aspen. This version, diffix-birch, extends that description to include a wide variety of common features found in SQL. It describes attacks associated with various features, and the anonymization steps used to defend against those attacks. This paper describes diffix-birch, which was used for the bounty program sponsored by Aircloak starting December 2017.

cs.CR

SkyMapper Filter Set: Design and Fabrication of Large Scale Optical Filters

The SkyMapper Southern Sky Survey will be conducted from Siding Spring Observatory with u, v, g, r, i and z filters that comprise glued glass combination filters of dimension 309x309x15 mm. In this paper we discuss the rationale for our bandpasses and physical characteristics of the filter set. The u, v, g and z filters are entirely glass filters which provide highly uniform band passes across the complete filter aperture. The i filter uses glass with a short-wave pass coating, and the r filter is a complete dielectric filter. We describe the process by which the filters were constructed, including the processes used to obtain uniform dielectric coatings and optimized narrow band anti-reflection coatings, as well as the technique of gluing the large glass pieces together after coating using UV transparent epoxy cement. The measured passbands including extinction and CCD QE are presented.

astro-ph.IM

The Extragalactic Distance Scale without Cepheids IV

The Cepheid period-luminosity relation is the primary distance indicator used in most determinations of the Hubble constant. The tip of the red giant branch (TRGB) is an alternative basis. Using the new ANU SkyMapper Telescope, we calibrate the Tully Fisher relation in the I band. We find that the TRGB and Cepheid distance scales are consistent.

astro-ph.CO

PAH Emission Within Lyman Alpha Blobs

We present Spitzer observations of Lya Blobs (LAB) at z=2.38-3.09. The mid-infrared ratios (4.5/8um and 8/24um) indicate that ~60% of LAB infrared counterparts are cool, consistent with their infrared output being dominated by star formation and not active galactic nuclei (AGN). The rest have a substantial hot dust component that one would expect from an AGN or an extreme starburst. Comparing the mid-infrared to submillimeter fluxes (~850um or rest frame far infrared) also indicates a large percentage (~2/3) of the LAB counterparts have total bolometric energy output dominated by star formation, although the number of sources with sub-mm detections or meaningful upper limits remains small (~10). We obtained Infrared Spectrograph (IRS) spectra of 6 infrared-bright sources associated with LABs. Four of these sources have measurable polycyclic aromatic hydrocarbon (PAH) emission features, indicative of significant star formation, while the remaining two show a featureless continuum, indicative of infrared energy output completely dominated by an AGN. Two of the counterparts with PAHs are mixed sources, with PAH line-to-continuum ratios and PAH equivalent widths indicative of large energy contributions from both star formation and AGN. Most of the LAB infrared counterparts have large stellar masses, around 10^11 Mo. There is a weak trend of mass upper limit with the Lya luminosity of the host blob, particularly after the most likely AGN contaminants are removed. The range in likely energy sources for the LABs found in this and previous studies suggests that there is no single source of power that is producing all the known LABs.

astro-ph.CO

The Southern 2MASS AGN Survey: spectroscopic follow-up with 6dF

The Two Micron All-Sky Survey (2MASS) has provided a uniform photometric catalog to search for previously unknown red AGN and QSOs. We have extended the search to the southern equatorial sky by obtaining spectra for 1182 AGN candidates using the 6dF multifibre spectrograph on the UK Schmidt Telescope. These were scheduled as auxiliary targets for the 6dF Galaxy Redshift Survey. The candidates were selected using a single color cut of J - Ks > 2 to Ks ~ 15.5 and a galactic latitude of |b|>30 deg. 432 spectra were of sufficient quality to enable a reliable classification. 116 sources (or ~27%) were securely classified as type 1 AGN, 20 as probable type 1s, and 57 as probable type 2 AGN. Most of them span the redshift range 0.05 20%) than in any previous galaxy survey. A small fraction of the type 1 AGN could have their optical colors reddened by optically thin dust with A_V<2 mag relative to optically selected QSOs. A handful show evidence for excess far-IR emission. The equivalent width (EW) and color distributions of the type 1 and 2 AGN are consistent with AGN unified models. In particular, the EW of the [OIII] emission line weakly correlates with optical--near-IR color in each class of AGN, suggesting anisotropic obscuration of the AGN continuum. Overall, the optical properties of the 2MASS red AGN are not dramatically different from those of optically-selected QSOs. Our near-IR selection appears to detect the most near-IR luminous QSOs in the local universe to z~0.6 and provides incentive to extend the search to deeper near-IR surveys.

astro-ph.CO

SkyMapper and the Southern Sky Survey: a valuable resource for stellar astrophysics

The Australian National University's SkyMapper telescope is amongst the first of a new generation of dedicated wide-field survey telescopes. Featuring a 5.7 square deg field-of-view Cassegrain imager and 268 Mega-pixel CCD array, its primary goal will be to undertake the Southern Sky Survey: a six color (uvgriz), six-epoch digital record of the entire southern sky. The survey will provide photometry for objects between 8th and 23rd magnitude with a global photometric accuracy of 0.03 magnitudes and astrometry to 50 mas. In this contribution we introduce the SkyMapper facility, the survey data products and outline a variety of case-studies in stellar astrophysics for which SkyMapper will have high impact.

astro-ph

Ultraviolet-Bright, High-Redshift ULIRGS

We present Spitzer Space Telescope observations of the z=2.38 lya-emitter over-density associated with galaxy cluster J2143-4423, the largest known structure (110 Mpc) above z=2. We imaged 22 of the 37 known lya-emitters within the filament-like structure, using the MIPS 24um band. We detected 6 of the lya-emitters, including 3 of the 4 clouds of extended (>50 kpc) lyman alpha emission, also known as Lya Blobs. Conversion from rest-wavelength 7um to total far-infrared luminosity using locally derived correlations suggests all the detected sources are in the class of ULIRGs, with some reaching Hyper-LIRG energies. Lya blobs frequently show evidence for interaction, either in HST imaging, or the proximity of multiple MIPS sources within the Lya cloud. This connection suggests that interaction or even mergers may be related to the production of Lya blobs. A connection to mergers does not in itself help explain the origin of the Lya blobs, as most of the suggested mechanisms for creating Lya blobs (starbursts, AGN, cooling flows) could also be associated with galaxy interactions.

astro-ph

A Group of Galaxies at Redshift 2.38

We report the discovery of a group of galaxies at redshift 2.38. We imaged about 10% of a claimed supercluster of QSO absorption-lines at z=2.38 (Francis & Hewett 1993). In this small field (2 arcmin radius) we detect two Ly-alpha emitting galaxies. The discovery of two such galaxies in our tiny field supports Francis & Hewett's interpretation of the absorption-line supercluster as a high redshift "Great Wall". One of the Ly-alpha galaxies lies 22 arcsec from a background QSO, and may be associated with a multi-component Ly-alpha absorption complex seen in the QSO spectrum. This galaxy has an extended (50kpc) lumpy Ly-alpha morphology, surrounding a compact IR-bright nucleus. The nucleus shows a pronounced break in its optical-UV colors at about 4000 A (rest-frame), consistent with a stellar population of mass about 7E11 solar masses, an age of more than 500 Myr, and little on-going star-formation. C IV emission is detected, suggesting that a concealed AGN is present. Extended H-alpha emission is also detected; the ratio of Ly-alpha flux to H-alpha is abnormally low (about 0.7), probable evidence for extended dust. This galaxy is surrounded by a number of very red (B-K>5) objects, some of which have colors suggesting that they too are at z=2.38. We hypothesize that this galaxy, its neighbors and a surrounding lumpy gas cloud may be a giant elliptical galaxy in the act of bottom-up formation.

astro-ph