SearcharxivSearch

arXiv subjects

Brian Libgober

Publications and source records attributed to Brian Libgober.

4 recordsLinked to original sources

Analysis and sample-size determination for $2^K$ audit experiments with binary response and application to identification of effect of racial discrimination on access to justice

Social scientists have increasingly turned to audit experiments to investigate discrimination in the market for jobs, loans, housing and other opportunities. In a typical audit experiment, researchers assign ``signals'' (the treatment) to subjects at random and compare success rates across treatment conditions. In the recent past there has been increased interest in using randomized multifactor designs for audit experiments, popularly called factorial experiments, in which combinations of multiple signals are assigned to subjects. Although social scientists have manipulated multiple factors like race, gender and income, the analyses have been mostly exploratory in nature. In this paper we lay out a comprehensive methodology for design and analysis of $2^K$ factorial designs with binary response using model-free, randomization-based Neymanian inference and demonstrate its application by analyzing the audit experiment reported in Libgober (2020). Specifically, we integrate and extend several sections of the randomization-based, finite-population literature for binary outcomes, including sample size and power calculations, and non-linear factorial estimators, extending results.

stat.ME

Linking Datasets on Organizations Using Half A Billion Open-Collaborated Records

Scholars studying organizations often work with multiple datasets lacking shared identifiers or covariates. In such situations, researchers usually use approximate string ("fuzzy") matching methods to combine datasets. String matching, although useful, faces fundamental challenges. Even where two strings appear similar to humans, fuzzy matching often struggles because it fails to adapt to the informativeness of the character combinations. In response, a number of machine learning methods have been developed to refine string matching. Yet, the effectiveness of these methods is limited by the size and diversity of training data. This paper introduces data from a prominent employment networking site (LinkedIn) as a massive training corpus to address these limitations. By leveraging information from the LinkedIn corpus regarding organizational name-to-name links, we incorporate trillions of name pair examples into various methods to enhance existing matching benchmarks and performance by explicitly maximizing match probabilities. We also show how relationships between organization names can be modeled using a network representation of the LinkedIn data. In illustrative merging tasks involving lobbying firms, we document improvements when using the LinkedIn corpus in matching calibration and make all data and methods open source.

cs.SI

Optimal allocation of sample size for randomization-based inference from $2^K$ factorial designs

Optimizing the allocation of units into treatment groups can help researchers improve the precision of causal estimators and decrease costs when running factorial experiments. However, existing optimal allocation results typically assume a super-population model and that the outcome data comes from a known family of distributions. Instead, we focus on randomization-based causal inference for the finite-population setting, which does not require model specifications for the data or sampling assumptions. We propose exact theoretical solutions for optimal allocation in $2^K$ factorial experiments under complete randomization with A-, D- and E-optimality criteria. We then extend this work to factorial designs with block randomization. We also derive results for optimal allocations when using cost-based constraints. To connect our theory to practice, we provide convenient integer-constrained programming solutions using a greedy optimization approach to find integer optimal allocation solutions for both complete and block randomization. The proposed methods are demonstrated using two real-life factorial experiments conducted by social scientists.

stat.ME

An Email Experiment to Identify the Effect of Racial Discrimination on Access to Lawyers: A Statistical Approach

We consider the problem of conducting an experiment to study the prevalence of racial bias against individuals seeking legal assistance, in particular whether lawyers use clues about a potential client's race in deciding whether to reply to e-mail requests for representations. The problem of discriminating between potential linear and non-linear effects of a racial signal is formulated as a statistical inference problem, whose objective is to infer a parameter determining the shape of a specific function. Various complexities associated with the design and analysis of this experiment are handled by applying a novel combination of rigorous, semi-rigorous and rudimentary statistical techniques. The actual experiment was attempted with a population of lawyers in Florida, but could not be performed with the desired sample size due to resource limitations. Nonetheless, it provides a nice demonstration of the proposed steps involved in conducting such a study.

stat.AP