SearcharxivSearch

arXiv subjects

Philip B. Stark

Publications and source records attributed to Philip B. Stark.

At least 19 recordsLinked to original sources

Sequential stratified inference for the mean

We develop conservative tests for the mean of a bounded population under stratified sampling and apply them to risk-limiting post-election audits. The tests are "anytime valid" under sequential sampling, allowing optional stopping in each stratum. Our core method expresses a global hypothesis about the population mean as a union of intersection hypotheses describing within-stratum means. It tests each intersection hypothesis using a test supermartingale (TSM) combined across strata by multiplication. A $P$-value for each intersection hypothesis is the reciprocal of that TSM, and the largest $P$-value in the union is a $P$-value for the global hypothesis. This approach has two primary moving parts: the rule selecting which stratum to draw from next, and the form of the TSM within each stratum. These rules may vary over intersection hypotheses. We construct intersection tests with the smallest expected sample size and propagate them to approximate an efficient global test. In instances that arise in auditing and other applications, its expected sample size is substantially smaller than that of previous methods.

stat.ME

It Doesn't Take a Thief: Optical-Scan Voting Systems Fail Even Without Adversaries

Optical-scan voting systems and their supporting ecosystem of people, processes, and technology are fallible. While a substantial body of work examines adversarial threats to such systems, we have encountered jurisdictions where the possibility of tabulator error is not fully internalized. Stakeholders there often find hypothetical attacks unconvincing, but some are persuaded by real-world accounts of equipment and procedural failures. This paper introduces a taxonomy of non-adversarial failure modes organized into intuitive categories: recording votes on paper, reading votes from the paper, combining votes as read into a reported outcome, and testing and verifying, all illustrated with documented incidents. We map common verification mechanisms against this taxonomy, identifying gaps that no paper-based audit can detect or correct, most notably failures that compromise the trustworthiness of the paper trail, such as giving voters the wrong ballot style (omitting contests they are eligible for, or including ones they are not), using ballot-marking devices to record votes, or failing to keep voted ballots secure and organized.

cs.CY

Doing More With Less: Mismatch-Based Risk-Limiting Audits

One approach to risk-limiting audits (RLAs) compares randomly selected cast vote records (CVRs) to votes read by human auditors from the corresponding ballot cards. Historically, such methods reduce audit sample sizes by considering how each sampled CVR differs from the corresponding true vote, not merely whether they differ. Here we investigate the latter approach, auditing by testing whether the total number of mismatches in the full set of CVRs exceeds the minimum number of CVR errors required for the reported outcome to be wrong (the "CVR margin"). This strategy makes it possible to audit more social choice functions and simplifies RLAs conceptually, which makes it easier to explain than some other RLA approaches. The cost is larger sample sizes. "Mismatch-based RLAs" only require a lower bound on the CVR margin, which for some social choice functions is easier to calculate than the effect of particular errors. When the population rate of mismatches is low and the lower bound on the CVR margin is close to the true CVR margin, the increase in sample size is small. However, the increase may be very large when errors include errors that, if corrected, would widen the CVR margin rather than narrow it; errors affect the margin between candidates other than the reported winner with the fewest votes and the reported loser with the most votes; or errors that affect different margins.

cs.CY

Exact and Conservative Inference for the Average Treatment Effect in Stratified Experiments with Binary Outcomes

We extend methods for finite-sample inference about the average treatment effect (ATE) in randomized experiments with binary outcomes to accommodate stratification (blocking). We present three valid methods that differ in their computational and statistical efficiency. The first method constructs conservative, Bonferroni-adjusted confidence intervals separately for the mean response in the treatment and control groups in each stratum, then takes appropriate weighted differences of their endpoints to find a confidence interval for the ATE. The second method inverts permutation tests for the overall ATE, maximizing the $P$-value over all ways a given ATE can be attained. The third method applies permutation tests for the ATE in separate strata, then combines those tests to form a confidence interval for the overall ATE. We compare the statistical and computational performance of the methods using simulations and a case study. The second approach is most efficient statistically in the simulations, but a naive implementation requires O(Π_{k=1}^{K} n_{k}^{4}) permutation tests, the highest computational burden among the three methods. That computational burden can be reduced to O(\sum_{k=1}^K n_k \timesΠ_{k=1}^{K} n_{k}^{2}) if all strata are balanced and to O(Π_{k=1}^{K} n_{k}^{3}) otherwise.

stat.ME

Fast Conservative Monte Carlo Confidence Intervals

Extant "fast" algorithms for Monte Carlo confidence sets are limited to univariate shift parameters for the one-sample and two-sample problems using the sample mean as the test statistic; moreover, some do not converge reliably and most do not produce conservative confidence sets. We outline general methods for constructing confidence sets for real-valued and multidimensional parameters by inverting Monte Carlo tests using any test statistic and a broad range of randomization schemes. The method exploits two facts that, to our knowledge, had not been combined: (i) there are Monte Carlo tests that are conservative despite relying on simulation, and (ii) since the coverage probability of confidence sets depends only on the significance level of the test of the true null, every null can be tested using the same Monte Carlo sample. The Monte Carlo sample can be arbitrarily small, although the highest nontrivial attainable confidence level generally increases as the number $N$ of Monte Carlo replicates increases. We present open-source Python and R implementations of new algorithms to compute conservative confidence sets for real-valued parameters from Monte Carlo tests, for test statistics and randomization schemes that yield $P$-values that are monotone or weakly unimodal in the parameter, with the data and Monte Carlo sample held fixed. In this case, the new method finds conservative confidence sets for real-valued parameters in $O(n)$ time, where $n$ is the number of data. The values of some test statistics for different simulations and parameter values have a simple relationship that makes more savings possible.

stat.CO

An Internet Voting System Fatally Flawed in Creative New Ways

The recently published "MERGE" protocol is designed to be used in the prototype CAC-vote system. The voting kiosk and protocol transmit votes over the internet and then transmit voter-verifiable paper ballots through the mail. In the MERGE protocol, the votes transmitted over the internet are used to tabulate the results and determine the winners, but audits and recounts use the paper ballots that arrive in time. The enunciated motivation for the protocol is to allow (electronic) votes from overseas military voters to be included in preliminary results before a (paper) ballot is received from the voter. MERGE contains interesting ideas that are not inherently unsound; but to make the system trustworthy--to apply the MERGE protocol--would require major changes to the laws, practices, and technical and logistical abilities of U.S. election jurisdictions. The gap between theory and practice is large and unbridgeable for the foreseeable future. Promoters of this research project at DARPA, the agency that sponsored the research, should acknowledge that MERGE is internet voting (election results rely on votes transmitted over the internet except in the event of a full hand count) and refrain from claiming that it could be a component of trustworthy elections without sweeping changes to election law and election administration throughout the U.S.

cs.CR

When Audits and Recounts Distract from Election Integrity: The 2020 U.S. Presidential Election in Georgia

Georgia was central to efforts to overturn the 2020 Presidential election, including a call from then-president Trump to Georgia Secretary of State Raffensperger asking Raffensperger to `find' 11,780 votes. Raffensperger has maintained that a `100% full-count risk-limiting audit' and a machine recount agreed with the initial machine-count results, which proved that the reported election results were accurate and that `no votes were flipped.' There is no indication of widespread fraud, but there is reason to distrust the election outcome: the two machine counts and the manual `audit' tallies disagree substantially, even about the number of ballots cast. Some ballots in Fulton County were included in the original count at least twice; some were included in the machine recount at least thrice. Audit results for some tally batches were omitted from the reported audit totals. The two machine counts and the audit were not probative of who won because of poor processes and controls: a lack of secure physical chain of custody, ballot accounting, pollbook reconciliation, and accounting for other election materials such as memory cards. Moreover, most voters voted with demonstrably untrustworthy ballot-marking devices, so even a perfect handcount or audit would not necessarily reveal who really won. True risk-limiting audits (RLAs) and rigorous recounts can limit the risk that an incorrect electoral outcome will be certified rather than being corrected. But no procedure can limit that risk without a trustworthy record of the vote. And even a properly conducted RLA of some contests in an election does not show that any other contests in that election were decided correctly. The 2020 U.S. Presidential election in Georgia illustrates unrecoverable errors that can render recounts and audits `security theater' that distract from the more serious problems rather than justifying trust.

stat.AP

Improving the Computational Efficiency of Adaptive Audits of IRV Elections

AWAIRE is one of two extant methods for conducting risk-limiting audits of instant-runoff voting (IRV) elections. In principle AWAIRE can audit IRV contests with any number of candidates, but the original implementation incurred memory and computation costs that grew superexponentially with the number of candidates. This paper improves the algorithmic implementation of AWAIRE in three ways that make it practical to audit IRV contests with 55 candidates, compared to the previous 6 candidates. First, rather than trying from the start to rule out all candidate elimination orders that produce a different winner, the algorithm starts by considering only the final round, testing statistically whether each candidate could have won that round. For those candidates who cannot be ruled out at that stage, it expands to consider earlier and earlier rounds until either it provides strong evidence that the reported winner really won or a full hand count is conducted, revealing who really won. Second, it tests a richer collection of conditions, some of which can rule out many elimination orders at once. Third, it exploits relationships among those conditions, allowing it to abandon testing those that are unlikely to help. We provide real-world examples with up to 36 candidates and synthetic examples with up to 55 candidates, showing how audit sample size depends on the margins and on the tuning parameters. An open-source Python implementation is publicly available.

cs.CY

Efficient Weighting Schemes for Auditing Instant-Runoff Voting Elections

Various risk-limiting audit (RLA) methods have been developed for instant-runoff voting (IRV) elections. A recent method, AWAIRE, is the first efficient approach that can take advantage of but does not require cast vote records (CVRs). AWAIRE involves adaptively weighted averages of test statistics, essentially "learning" an effective set of hypotheses to test. However, the initial paper on AWAIRE only examined a few weighting schemes and parameter settings. We explore schemes and settings more extensively, to identify and recommend efficient choices for practice. We focus on the case where CVRs are not available, assessing performance using simulations based on real election data. The most effective schemes are often those that place most or all of the weight on the apparent "best" hypotheses based on already seen data. Conversely, the optimal tuning parameters tended to vary based on the election margin. Nonetheless, we quantify the performance trade-offs for different choices across varying election margins, aiding in selecting the most desirable trade-off if a default option is needed. A limitation of the current AWAIRE implementation is its restriction to a small number of candidates -- up to six in previous implementations. One path to a more computationally efficient implementation would be to use lazy evaluation and avoid considering all possible hypotheses. Our findings suggest that such an approach could be done without substantially compromising statistical performance.

cs.CY

A Declaration of Software Independence

A voting system should not merely report the outcome: it should also provide sufficient evidence to convince reasonable observers that the reported outcome is correct. Many deployed systems, notably paperless DRE machines still in use in US elections, fail certainly the second, and quite possibly the first of these requirements. Rivest and Wack proposed the principle of software independence (SI) as a guiding principle and requirement for voting systems. In essence, a voting system is SI if its reliance on software is ``tamper-evident'', that is, if there is a way to detect that material changes were made to the software without inspecting that software. This important notion has so far been formulated only informally. Here, we provide more formal mathematical definitions of SI. This exposes some subtleties and gaps in the original definition, among them: what elements of a system must be trusted for an election or system to be SI, how to formalize ``detection'' of a change to an election outcome, the fact that SI is with respect to a set of detection mechanisms (which must be legal and practical), the need to limit false alarms, and how SI applies when the social choice function is not deterministic.

cs.SE

Adaptively Weighted Audits of Instant-Runoff Voting Elections: AWAIRE

An election audit is risk-limiting if the audit limits (to a pre-specified threshold) the chance that an erroneous electoral outcome will be certified. Extant methods for auditing instant-runoff voting (IRV) elections are either not risk-limiting or require cast vote records (CVRs), the voting system's electronic record of the votes on each ballot. CVRs are not always available, for instance, in jurisdictions that tabulate IRV contests manually. We develop an RLA method (AWAIRE) that uses adaptively weighted averages of test supermartingales to efficiently audit IRV elections when CVRs are not available. The adaptive weighting 'learns' an efficient set of hypotheses to test to confirm the election outcome. When accurate CVRs are available, AWAIRE can use them to increase the efficiency to match the performance of existing methods that require CVRs. We provide an open-source prototype implementation that can handle elections with up to six candidates. Simulations using data from real elections show that AWAIRE is likely to be efficient in practice. We discuss how to extend the computational approach to handle elections with more candidates. Adaptively weighted averages of test supermartingales are a general tool, useful beyond election audits to test collections of hypotheses sequentially while rigorously controlling the familywise error rate.

stat.AP

Stylish Risk-Limiting Audits in Practice

Risk-limiting audits (RLAs) can use information about which ballot cards contain which contests (card-style data, CSD) to ensure that each contest receives adequate scrutiny, without examining more cards than necessary. RLAs using CSD in this way can be substantially more efficient than RLAs that sample indiscriminately from all cast cards. We describe an open-source Python implementation of RLAs using CSD for the Hart InterCivic Verity voting system and the Dominion Democracy Suite(R) voting system. The software is demonstrated using all 181 contests in the 2020 general election and all 214 contests in the 2022 general election in Orange County, CA, USA, the fifth-largest election jurisdiction in the U.S., with over 1.8 million active voters. (Orange County uses the Hart Verity system.) To audit the 181 contests in 2020 to a risk limit of 5% without using CSD would have required a complete hand tally of all 3,094,308 cast ballot cards. With CSD, the estimated sample size is about 20,100 cards, 0.65% of the cards cast--including one tied contest that required a complete hand count. To audit the 214 contests in 2022 to a risk limit of 5% without using CSD would have required a complete hand tally of all 1,989,416 cast cards. With CSD, the estimated sample size is about 62,250 ballots, 3.1% of cards cast--including three contests with margins below 0.1% and 9 with margins below 0.5%.

stat.AP

Non(c)esuch Ballot-Level Risk-Limiting Audits for Precinct-Count Voting Systems

Risk-limiting audits (RLAs) guarantee a high probability of correcting incorrect reported outcomes before the outcomes are certified. The most efficient use ballot-level comparison, comparing the voting system's interpretation of individual ballot cards sampled at random (cast-vote records, CVRs) from a trustworthy paper trail to a human interpretation of the same cards. Such comparisons require the voting system to create and export CVRs in a way that can be linked to the individual ballots the CVRs purport to represent. Such links can be created by keeping the ballots in the order in which they are scanned or by printing a unique serial number on each ballot. But for precinct-count systems (PCOS), these strategies may compromise vote anonymity: the order in which ballots are cast may identify the voters who cast them. Printing a unique pseudo-random number ("cryptographic nonce") on each ballot card after the voter last touches it could reduce such privacy risks. But what if the system does not in fact print a unique number on each ballot or does not accurately report the numbers it printed? This paper gives two ways to conduct an RLA so that even if the system does not print a genuine nonce on each ballot or misreports the nonces it used, the audit's risk limit is not compromised (however, the anonymity of votes might be compromised). One method allows untrusted technology to be used to imprint and to retrieve ballot cards. The method is adaptive: if the technology behaves properly, this protection does not increase the audit workload. But if the imprinting or retrieval system misbehaves, the sample size the RLA requires to confirm the reported results when the results are correct is generally larger than if the imprinting and retrieval were accurate.

cs.CR

Overstatement-Net-Equivalent Risk-Limiting Audit: ONEAudit

Card-level comparison risk-limiting audits (CLCAs) heretofore required a CVR for each cast card and a "link" identifying which CVR is for which card -- which many voting systems cannot provide. Every set of CVRs that produces the same aggregate results overstates contest margins by the same amount: they are overstatement-net-equivalent (ONE). CLCAs can therefore use CVRs from the voting system for any number of cards and ONE CVRs for the rest. Ballot-polling RLAs are equivalent to CLCAs using ONE CVRs. CLCAs can be based on batch-level results (e.g., precinct subtotals) by constructing ONE CVRs for each batch. In contrast to batch-level comparison audits (BLCAs), this avoids tabulating batches manually and works even when reporting batches do not correspond to physically identifiable batches of cards. If the voting system can export linked CVRs for only some ballot cards, auditors can still use CLCA by constructing ONE CVRs for the rest of the cards from contest results or batch subtotals. This obviates the need for "hybrid" audits. This works for every social choice function for which there is a known RLA method, including IRV. Sample sizes for BPA and ONEAudit using contest totals are comparable. ONEAudit using batch subtotals has smaller sample sizes than ballot-polling when batches are much more homogeneous than the election overall. Sample sizes can be much smaller than for BLCA: A CLCA of the 2022 presidential election in California at risk limit 5% using ONE CVRs for precinct-level results would sample ~70 ballots statewide, if the reported results are accurate, compared to about 26,700 for BLCA. The 2022 Georgia audit tabulated >231,000 cards versus ~1300 for ONEAudit. For data from a pilot hybrid RLA in Kalamazoo, MI, in 2018, ONEAudit gives a risk of ~2%, substantially lower than the 3.7% measured risk for SUITE, the method the pilot used.

stat.AP

Ballot-Polling Audits of Instant-Runoff Voting Elections with a Dirichlet-Tree Model

Instant-runoff voting (IRV) is used in several countries around the world. It requires voters to rank candidates in order of preference, and uses a counting algorithm that is more complex than systems such as first-past-the-post or scoring rules. An even more complex system, the single transferable vote (STV), is used when multiple candidates need to be elected. The complexity of these systems has made it difficult to audit the election outcomes. There is currently no known risk-limiting audit (RLA) method for STV, other than a full manual count of the ballots. A new approach to auditing these systems was recently proposed, based on a Dirichlet-tree model. We present a detailed analysis of this approach for ballot-polling Bayesian audits of IRV elections. We compared several choices for the prior distribution, including some approaches using a Bayesian bootstrap (equivalent to an improper prior). Our findings include that the bootstrap-based approaches can be adapted to perform similarly to a full Bayesian model in practice, and that an overly informative prior can give counter-intuitive results. Via carefully chosen examples, we show why creating an RLA with this model is challenging, but we also suggest ways to overcome this. As well as providing a practical and computationally feasible implementation of a Bayesian IRV audit, our work is important in laying the foundation for an RLA for STV elections.

stat.AP

Auditing Ranked Voting Elections with Dirichlet-Tree Models: First Steps

Ranked voting systems, such as instant-runoff voting (IRV) and single transferable vote (STV), are used in many places around the world. They are more complex than plurality and scoring rules, presenting a challenge for auditing their outcomes: there is no known risk-limiting audit (RLA) method for STV other than a full hand count. We present a new approach to auditing ranked systems that uses a statistical model, a Dirichlet-tree, that can cope with high-dimensional parameters in a computationally efficient manner. We demonstrate this approach with a ballot-polling Bayesian audit for IRV elections. Although the technique is not known to be risk-limiting, we suggest some strategies that might allow it to be calibrated to limit risk.

stat.AP

Pay No Attention to the Model Behind the Curtain

Many widely used models amount to an elaborate means of making up numbers--but once a number has been produced, it tends to be taken seriously and its source (the model) is rarely examined carefully. Many widely used models have little connection to the real-world phenomena they purport to explain. Common steps in modeling to support policy decisions, such as putting disparate things on the same scale, may conflict with reality. Not all costs and benefits can be put on the same scale, not all uncertainties can be expressed as probabilities, and not all model parameters measure what they purport to measure. These ideas are illustrated with examples from seismology, wind-turbine bird deaths, soccer penalty cards, gender bias in academia, and climate policy.

stat.ME

ALPHA: Audit that Learns from Previously Hand-Audited Ballots

BRAVO, the most widely tried method for risk-limiting election audits, cannot accommodate sampling without replacement or stratified sampling, which can improve efficiency and may be required by law. It applies only to ballot-polling audits, which are less efficient than comparison audits. It applies to plurality, majority, super-majority, proportional representation, and ranked-choice voting contests, but not to many social choice functions for which there are RLA methods, such as approval voting, STAR-voting, Borda count, and general scoring rules. And while BRAVO has the smallest expected sample size among sequentially valid ballot-polling-with-replacement methods when reported vote shares are exactly right, it can require arbitrarily large samples when the reported reported winner(s) really won but reported vote shares are wrong. ALPHA is a simple generalization of BRAVO that (i) works for sampling with and without replacement and Bernoulli sampling; (ii) increases power for stratified audits by avoiding the need to use a $P$-value combining function or to maximize $P$-values over nuisance parameters within strata, and allowing adaptive sampling across strata; (iii) works not only for ballot-polling but also for ballot-level comparison, batch-polling, and batch-level comparison audits, sampling with or without replacement, uniformly or with weights proportional to size; (iv) works for all social choice functions covered by SHANGRLA; and (v) in situations where both ALPHA and BRAVO apply, requires smaller samples than BRAVO when the reported vote shares are wrong but the outcome is correct--five orders of magnitude in some examples. ALPHA includes the family of betting martingale tests in RiLACS, with a different betting strategy parametrized as an estimator of the population mean and explicit flexibility to accommodate sampling weights and population bounds that vary by draw.

stat.ME