SearcharxivSearch

arXiv subjects

Konstantin Genin

Publications and source records attributed to Konstantin Genin.

4 recordsLinked to original sources

Topological Criteria for Hypothesis Testing with Finite-Precision Measurements

We establish topological necessary and sufficient conditions under which a pair of statistical hypotheses can be consistently distinguished when i.i.d. observations are recorded only to finite precision. To accommodate finite-precision data, we introduce finite-precision tests: tests whose decision regions are open in the sample-space topology. We first show that, both for classical and finite-precision tests, the existence of such tests with finite-sample error control, asymptotic error control, or uniform convergence of the errors are all equivalent. A pair of null- and alternative hypotheses $H_0$ and $H_1$ admits a consistent finite-precision test if and only if both are $F_\sigma$ in the weak topology on the space of probability measures $W := H_0\cup H_1$. The hypotheses admit uniform error control under $H_i$ if and only if $H_i$ is closed in $W$, and admit uniformly consistent testing with bounded precision under metric separation of $H_0$ and $H_1$. These criteria imply that, without regularity assumptions, conditional independence is not consistently testable from finite-precision data when the conditioning space has no isolated points - strengthening existing impossibility results to Polish sample spaces and showing that even pointwise consistency cannot be obtained. We introduce an equicontinuity assumption on the family of conditional distributions under which we recover consistent finite-precision testability of conditional independence with uniform error control under the null, provided sample spaces are Polish and the conditioning space is locally compact. The equicontinuity assumption is itself a finite-precision-testable hypothesis, so the resulting test for conditional independence is, in a precise sense, assumption-free.

math.ST

From the Fair Distribution of Predictions to the Fair Distribution of Social Goods: Evaluating the Impact of Fair Machine Learning on Long-Term Unemployment

Deploying an algorithmically informed policy is a significant intervention in society. Prominent methods for algorithmic fairness focus on the distribution of predictions at the time of training, rather than the distribution of social goods that arises after deploying the algorithm in a specific social context. However, requiring a "fair" distribution of predictions may undermine efforts at establishing a fair distribution of social goods. First, we argue that addressing this problem requires a notion of prospective fairness that anticipates the change in the distribution of social goods after deployment. Second, we provide formal conditions under which this change is identified from pre-deployment data. That requires accounting for different kinds of performative effects. Here, we focus on the way predictions change policy decisions and, consequently, the causally downstream distribution of social goods. Throughout, we are guided by an application from public administration: the use of algorithms to predict who among the recently unemployed will remain unemployed in the long term and to target them with labor market programs. Third, using administrative data from the Swiss public employment service, we simulate how such algorithmically informed policies would affect gender inequalities in long-term unemployment. When risk predictions are required to be "fair" according to statistical parity and equality of opportunity, targeting decisions are less effective, undermining efforts to both lower overall levels of long-term unemployment and to close the gender gap in long-term unemployment.

cs.CY

Performativity and Prospective Fairness

Deploying an algorithmically informed policy is a significant intervention in the structure of society. As is increasingly acknowledged, predictive algorithms have performative effects: using them can shift the distribution of social outcomes away from the one on which the algorithms were trained. Algorithmic fairness research is usually motivated by the worry that these performative effects will exacerbate the structural inequalities that gave rise to the training data. However, standard retrospective fairness methodologies are ill-suited to predict these effects. They impose static fairness constraints that hold after the predictive algorithm is trained, but before it is deployed and, therefore, before performative effects have had a chance to kick in. However, satisfying static fairness criteria after training is not sufficient to avoid exacerbating inequality after deployment. Addressing the fundamental worry that motivates algorithmic fairness requires explicitly comparing the change in relevant structural inequalities before and after deployment. We propose a prospective methodology for estimating this post-deployment change from pre-deployment data and knowledge about the algorithmic policy. That requires a strategy for distinguishing between, and accounting for, different kinds of performative effects. In this paper, we focus on the algorithmic effect on the causally downstream outcome variable. Throughout, we are guided by an application from public administration: the use of algorithms to (1) predict who among the recently unemployed will stay unemployed for the long term and (2) targeting them with labor market programs. We illustrate our proposal by showing how to predict whether such policies will exacerbate gender inequalities in the labor market.

cs.CY

The Topology of Statistical Verifiability

Topological models of empirical and formal inquiry are increasingly prevalent. They have emerged in such diverse fields as domain theory [1, 16], formal learning theory [18], epistemology and philosophy of science [10, 15, 8, 9, 2], statistics [6, 7] and modal logic [17, 4]. In those applications, open sets are typically interpreted as hypotheses deductively verifiable by true propositional information that rules out relevant possibilities. However, in statistical data analysis, one routinely receives random samples logically compatible with every statistical hypothesis. We bridge the gap between propositional and statistical data by solving for the unique topology on probability measures in which the open sets are exactly the statistically verifiable hypotheses. Furthermore, we extend that result to a topological characterization of learnability in the limit from statistical data.

cs.LG