SearcharxivSearch

arXiv subjects

Daryl DeFord

Publications and source records attributed to Daryl DeFord.

14 recordsLinked to original sources

Redistricting from the Bottom Up: Sampling Communities of Interest with Differential Privacy

Independent Redistricting Commissions (IRCs) are a promising tool for bottom-up redistricting, but their public testimony processes are vulnerable to adversarial manipulation. We propose using differential privacy to draw redistricting plans that incorporate community of interest (COI) testimonies while remaining robust to adversarial input. Treating individual testimonies as data points, we use the marked edge walk to sample from differentially private distributions of redistricting plans via the exponential mechanism. We introduce two score functions and demonstrate that both can be targeted by MEW across a range of privacy budgets. Applying this method to Missouri's mid-cycle redistricting using 808 COI testimonies, we show that COI-informed sampling outperforms an uninformed baseline and the enacted plan. An adversarial experiment demonstrates that the method can be robust to attacks under certain privacy budgets and may perform better in practice than formal group privacy guarantees imply. We also find that stronger COI preservation tends to spread minority and Democratic representation more evenly across districts.

cs.CY

The Marked Edge Walk: A Novel MCMC Algorithm for Sampling of Graph Partitions

Novel Markov Chain Monte Carlo (MCMC) methods have enabled the generation of large ensembles of redistricting plans through graph partitioning. However, existing algorithms such as Reversible Recombination (RevReCom) and Metropolized Forest Recombination (MFR) are constrained to sampling from distributions related to spanning trees. We introduce the marked edge walk (MEW), a novel MCMC algorithm for sampling from the space of graph partitions under a tunable distribution. The walk operates on the space of spanning trees with marked edges, allowing for calculable transition probabilities for use in the Metropolis-Hastings algorithm. Empirical results on real-world dual graphs show convergence under target distributions unrelated to spanning trees. For this reason, MEW represents an advancement in flexible ensemble generation.

cs.DS

Free Elections in the Free State: Ensemble Analysis of Redistricting in New Hampshire

The process of legislative redistricting in New Hampshire, along with many other states across the country, was particularly contentious during the 2020 census cycle. In this paper we present an ensemble analysis of the enacted districts to provide mathematical context for claims made about these maps in litigation. Operationalizing the New Hampshire redistricting rules and algorithmically generating a large collection of districting plans allows us to construct a baseline for expected behavior of districting plans in the state and evaluate non-partisan justifications and geographic tradeoffs between districting criteria and partisan outcomes. In addition, our results demonstrate the impact of selection and aggregation of election data for analyzing partisan symmetry measures.

cs.SI

Bounds and Bugs: The Limits of Symmetry Metrics to Detect Partisan Gerrymandering

We consider two symmetry metrics commonly used to analyze partisan gerrymandering: the Mean-Median Difference (MM) and Partisan Bias (PB). Our main results compare, for combinations of seats and votes achievable in districted elections, the number of districts won by each party to the extent of potential deviation from the ideal metric values, taking into account the political geography of the state. These comparisons are motivated by examples where the MM and PB have been used in efforts to detect when a districting plan awards extreme number of districts won by some party. These examples include expert testimony, public-facing apps, recommendations by experts to redistricting commissions, and public policy proposals. To achieve this goal we perform both theoretical and empirical analyses of the MM and PB. In our theoretical analysis, we consider vote-share, seat-share pairs (V, S) for which one can construct election data having vote share V and seat share S, and turnout is equal in each district. We calculate the range of values that MM and PB can achieve on that constructed election data. In the process, we find the range of (V,S) pairs that achieve MM = 0, and see that the corresponding range for PB is the same set of (V,S) pairs. We show how the set of such (V,S) pairs allowing for MM = 0 (and PB = 0) changes when turnout in each district is allowed to vary. By observing the results of this theoretical analysis, we can show that the values taken on by these metrics do not necessarily attain more extreme values in plans with more extreme numbers of districts won. We also analyze specific example elections, showing how these metrics can return unintuitive results. We follow this with an empirical study, where we show that on 18 different U.S. maps these metrics can fail to detect extreme seats outcomes.

cs.CY

Empirical Sampling of Connected Graph Partitions for Redistricting

The space of connected graph partitions underlies statistical models used as evidence in court cases and reform efforts that analyze political districting plans. In response to the demands of redistricting applications, researchers have developed sampling methods that traverse this space, building on techniques developed for statistical physics. In this paper, we study connections between redistricting and statistical physics, and in particular with self-avoiding walks. We exploit knowledge of phase transitions and asymptotic behavior in self avoiding walks to analyze two questions of crucial importance for Markov Chain Monte Carlo analysis of districting plans. First, we examine mixing times of a popular Glauber dynamics based Markov chain and show how the self-avoiding walk phase transitions interact with mixing time. We examine factors new to the redistricting context that complicate the picture, notably the population balance requirements, connectivity requirements, and the irregular graphs used. Second, we analyze the robustness of the qualitative properties of typical districting plans with respect to score functions and a certain lattice-like graph, called the state-dual graph, that is used as a discretization of geographic regions in most districting analysis. This helps us better understand the complex relationship between typical properties of districting plans and the score functions designed by political districting analysts. We conclude with directions for research at the interface of statistical physics, Markov chains, and political districting.

physics.soc-ph

Colorado in Context: Congressional Redistricting and Competing Fairness Criteria in Colorado

In this paper, we apply techniques of ensemble analysis to understand the political baseline for Congressional representation in Colorado. We generate a large random sample of reasonable redistricting plans and determine the partisan balance of each district using returns from state-wide elections in 2018, and analyze the 2011/2012 enacted districts in this context. Colorado recently adopted a new framework for redistricting, creating an independent commission to draw district boundaries, prohibiting partisan bias and incumbency considerations, requiring that political boundaries (such as counties) be preserved as much as possible, and also requiring that mapmakers maximize the number of competitive districts. We investigate the relationships between partisan outcomes, number of counties which are split, and number of competitive districts in a plan. This paper also features two novel improvements in methodology--a more rigorous statistical framework for understanding necessary sample size, and a weighted-graph method for generating random plans which split approximately as few counties as acceptable human-drawn maps.

cs.CY

Implementing partisan symmetry: Problems and paradoxes

We consider the measures of partisan symmetry proposed for practical use in the political science literature, as clarified and developed in Katz, King, and Rosenblatt (2020). Elementary mathematical manipulation shows the symmetry metrics to have surprising properties that call their meaningfulness into question. To accompany the general analysis, we study measures of partisan symmetry with respect to recent voting patterns in Utah, Texas, and North Carolina, flagging problems in each case. Taken together, these observations should raise major concerns about the available techniques for quantitative scores of partisan symmetry -- including the mean-median score, the partisan bias score, and the more general "partisan symmetry standard" -- as the decennial redistricting begins.

physics.soc-ph

A Computational Approach to Measuring Vote Elasticity and Competitiveness

The recent wave of attention to partisan gerrymandering has come with a push to refine or replace the laws that govern political redistricting around the country. A common element in several states' reform efforts has been the inclusion of competitiveness metrics, or scores that evaluate a districting plan based on the extent to which district-level outcomes are in play or are likely to be closely contested. In this paper, we examine several classes of competitiveness metrics motivated by recent reform proposals and then evaluate their potential outcomes across large ensembles of districting plans at the Congressional and state Senate levels. This is part of a growing literature using MCMC techniques from applied statistics to situate plans and criteria in the context of valid redistricting alternatives. Our empirical analysis focuses on five states---Utah, Georgia, Wisconsin, Virginia, and Massachusetts---chosen to represent a range of partisan attributes. We highlight situation-specific difficulties in creating good competitiveness metrics and show that optimizing competitiveness can produce unintended consequences on other partisan metrics. These results demonstrate the importance of (1) avoiding writing detailed metric constraints into long-lasting constitutional reform and (2) carrying out careful mathematical modeling on real geo-electoral data in each redistricting cycle.

cs.CY

Mathematics of Nested Districts: The Case of Alaska

In eight states, a "nesting rule" requires that each state Senate district be exactly composed of two adjacent state House districts. In this paper we investigate the potential impacts of these nesting rules with a focus on Alaska, where Republicans have a 2/3 majority in the Senate while a Democratic-led coalition controls the House. Treating the current House plan as fixed and considering all possible pairings, we find that the choice of pairings alone can create a swing of 4-5 seats out of 20 against recent voting patterns, which is similar to the range observed when using a Markov chain procedure to generate plans without the nesting constraint. The analysis enables other insights into Alaska districting, including the partisan latitude available to districters with and without strong rules about nesting and contiguity.

cs.SI

Recombination: A family of Markov chains for redistricting

Redistricting is the problem of partitioning a set of geographical units into a fixed number of districts, subject to a list of often-vague rules and priorities. In recent years, the use of randomized methods to sample from the vast space of districting plans has been gaining traction in courts of law for identifying partisan gerrymanders, and it is now emerging as a possible analytical tool for legislatures and independent commissions. In this paper, we set up redistricting as a graph partition problem and introduce a new family of Markov chains called Recombination (or ReCom) on the space of graph partitions. The main point of comparison will be the commonly used Flip walk, which randomly changes the assignment label of a single node at a time. We present evidence that ReCom mixes efficiently, especially in contrast to the slow-mixing Flip, and provide experiments that demonstrate its qualitative behavior. We demonstrate the advantages of ReCom on real-world data and explain both the challenges of the Markov chain approach and the analytical tools that it enables. We close with a short case study involving the Virginia House of Delegates.

cs.CY

Complexity and Geometry of Sampling Connected Graph Partitions

In this paper, we prove intractability results about sampling from the set of partitions of a planar graph into connected components. Our proofs are motivated by a technique introduced by Jerrum, Valiant, and Vazirani. Moreover, we use gadgets inspired by their technique to provide families of graphs where the "flip walk" Markov chain used in practice for this sampling task exhibits exponentially slow mixing. Supporting our theoretical results we present some empirical evidence demonstrating the slow mixing of the flip walk on grid graphs and on real data. Inspired by connections to the statistical physics of self-avoiding walks, we investigate the sensitivity of certain popular sampling algorithms to the graph topology. Finally, we discuss a few cases where the sampling problem is tractable. Applications to political redistricting have recently brought increased attention to this problem, and we articulate open questions about this application that are highlighted by our results.

cs.CC

Total Variation Isoperimetric Profiles

Applications such as political redistricting demand quantitative measures of geometric compactness to distinguish between simple and contorted shapes. While the isoperimetric quotient, or ratio of area to perimeter squared, is commonly used in practice, it is sensitive to noisy data and irrelevant geographic features like coastline. These issues are addressed in theory by the isoperimetric profile, which plots the minimum perimeter needed to inscribe regions of different prescribed areas within the boundary of a shape. Efficient algorithms for computing this profile, however, are not known in practice. Hence, in this paper, we propose a convex Eulerian relaxation of the isoperimetric profile using total variation. We prove theoretical properties of our relaxation, showing that it still satisfies an isoperimetric inequality and yields a convex function of the prescribed area. Furthermore, we provide a discretization of the problem, an optimization technique, and experiments demonstrating the value of our relaxation.

math.OC

Random Walk Null Models for Time Series Data

Permutation entropy has become a standard tool for time series analysis that exploits the temporal properties of these data sets. Many current applications use an approach based on Shannon entropy, which implicitly assumes an underlying uniform distribution of patterns. In this paper, we analyze random walk null models for time series and determine the corresponding permutation distributions. These new techniques allow us to explicitly describe the behavior of real world data in terms of more complex generative processes. Additionally, building on recent results of Martinez, we define a validation measure that allows us to determine when a random walk is an appropriate model for a time series. We demonstrate the usefulness of our methods using empirical data drawn from a variety of fields.

stat.ME