SearcharxivSearch

arXiv subjects

Jonathan Auerbach

Publications and source records attributed to Jonathan Auerbach.

11 recordsLinked to original sources

Post-Selection Inference for Network Structure

Researchers often use the density of connections between groups of agents, such as communities, blocs, or markets, to characterize the structure of a social or economic network. In many cases, these groups are selected using the network data, making conventional fixed-group inference procedures potentially invalid. To address this issue, we develop two new confidence intervals that are universally valid post-selection in the sense that they guarantee simultaneous coverage asymptotically over all pairs of groups whose relative sizes do not vanish. Our first interval builds on a strategy of Berk et al. (2013). Our second interval is based on a Talagrand-type concentration inequality for empirical processes. Both intervals are simple to compute and scalable to large networks, but a key technical contribution of our paper is to show that the second interval is rate-optimal over a broader class of intervals. Three empirical illustrations show that accounting for selection can matter in practice. Some evidence for homophily in a social network and a hub-and-spoke structure in a trade network survives our correction, while evidence for a segmented market structure in a worker transition network does not.

econ.EM

Scattered spring: How climate change disrupts the synchrony of biological events

Many biological processes, including plant leafout and flowering, occur once cumulative temperatures reach a threshold, a relationship known as the thermal-sum model. In this way, temperature is thought to coordinate the timing of biological events. An important implication is that higher temperatures cause thresholds to be reached sooner so that the timing of spring events, for example, should advance earlier in the calendar year as climates warm. But growing evidence has found that, as climates have warmed, the rate of advancement has slowed (a trend known as declining sensitivity), while the variance in the timing of spring events has increased in many cases (a trend we call declining synchrony). These trends raise questions about the resilience of temperature-based coordination to anthropogenic climate change. To answer these questions, researchers have modified the thermal-sum model by introducing additional factors and mechanisms, such as chilling and photoperiod. We show such complexity is not necessary to explain current trends of sensitivity and synchrony. Using experimental and real-world data, we find these trends are exactly as predicted by the thermal-sum model. In particular, the model predicts that as temperatures continue to increase and springtime events shift from the equinox toward the winter solstice, those events will become less synchronized and more variable, a phenomenon we refer to as a scattered spring. By 2100, our results predict much of the American South will experience a scattered spring under a high-warming scenario, with the average time between spring flowering events increasing by several weeks in many locations.

stat.AP

How many federal employees are not satisfied? Using response times to estimate population proportions under the survey variable cause model

We propose a statistical model to estimate population proportions under the survey variable cause model (Groves 2006), the setting in which the characteristic measured by the survey has a direct causal effect on survey participation. For example, we estimate employee satisfaction from a survey in which the decision of an employee to participate depends on their satisfaction. We model the time at which a respondent 'arrives' to take the survey, leveraging results from the counting processes literature that has been developed to analyze similar problems with survival data. Our approach is particularly useful for nonresponse bias analysis because it relies on different assumptions than traditional adjustments such as poststratification, which assumes the common cause model, the setting in which external factors explain the characteristic measured by the survey and participation. Our motivation is the Federal Employee Viewpoint Survey, which asks federal employees whether they are satisfied with their work organization. Our model suggests that the sample proportion overestimates the proportion of federal employees that are not satisfied with their work organization even after adjustment by poststratification. Employees that are not satisfied likely select into the survey, and this selection cannot be explained by personal characteristics like race, gender, and occupation or work-place characteristics like agency, unit, and location.

stat.AP

A Nonparametric Bayesian Model to Adjust for Monitoring Bias with an Application to Identifying Environments Stressed by Climate Change

We propose a new method to adjust for the bias that occurs when an individual monitors a location and reports the status of an event. For example, a monitor may visit a plant each week and report whether the plant is in flower or not. The goal is to estimate the time the event occurred at that location. The problem is that popular estimators often incur bias both because the event may not coincide with the arrival of the monitor and because the monitor may report the status in error. To correct for this bias, we propose a nonparametric Bayesian model that uses monotonic splines to estimate the event time. We first demonstrate the problem and our proposed solution using simulated data. We then apply our method to a real-world example from phenology in which lilac are monitored by citizen scientists in the northeastern United States, and the timing of the flowering is used to study anthropogenic warming. Our analysis suggests that common methods fail to account for monitoring bias and underestimate the peak bloom date of the lilac by 48 days on average. In addition, after adjusting for monitoring bias, several locations had anomalously late bloom dates that did not appear anomalous before adjustment. Our findings underscore the importance of accounting for monitoring bias in event-time estimation. By applying our nonparametric Bayesian model with monotonic splines, we provide a more accurate approach to estimating bloom dates, revealing previously undetected anomalies and improving the reliability of citizen science data for environmental monitoring.

stat.AP

Estimating the Number of Street Vendors in New York City: Ratio Estimation with Point Process Data

We estimate the number of street vendors in New York City. First, we summarize the process by which vendors receive licenses and permits to operate legally in New York City. We then describe a survey that was administered by the Street Vendor Project while distributing coronavirus relief aid to vendors operating in New York City both with and without a license or permit. Finally, we review ratio estimation and develop a theoretical justification based on the theory of point processes. We find approximately 23,000 street vendors operate in New York City: 20,500 mobile food vendors and 2,400 general merchandise vendors. One third are located in just six ZIP Codes: 11368 (16%), 11372 (3%), and 11354 (3%) in North and West Queens and 10036 (5%), 10019 (4%), and 10001 (3%) in the Chelsea and Clinton neighborhoods of Manhattan. Our estimates suggest the American Community Survey misses the majority of New York City street vendors.

stat.AP

Exposure effects are not automatically useful for policymaking

We thank Savje (2023) for a thought-provoking article and appreciate the opportunity to share our perspective as social scientists. In his article, Savje recommends misspecified exposure effects as a way to avoid strong assumptions about interference when analyzing the results of an experiment. In this invited discussion, we highlight a limiation of Savje's recommendation: exposure effects are not generally useful for evaluating social policies without the strong assumptions that Savje seeks to avoid.

econ.EM

Tensor Completion for Causal Inference with Multivariate Longitudinal Data: A Reevaluation of COVID-19 Mandates

We propose a new method that uses tensor completion to estimate causal effects with multivariate longitudinal data, data in which multiple outcomes are observed for each unit and time period. Our motivation is to estimate the number of COVID-19 fatalities prevented by government mandates such as travel restrictions, mask-wearing directives, and vaccination requirements. In addition to COVID-19 fatalities, we observe related outcomes such as the number of fatalities from other diseases and injuries. The proposed method arranges the data as a tensor with three dimensions (unit, time, and outcome) and uses tensor completion to impute the missing counterfactual outcomes. We first prove that under general conditions, combining multiple outcomes using the proposed method improves the accuracy of counterfactual imputations. We then compare the proposed method to other approaches commonly used to evaluate COVID-19 mandates. Our main finding is that other approaches overestimate the effect of masking-wearing directives and that mask-wearing directives were not an effective alternative to travel restrictions. We conclude that while the proposed method can be applied whenever multivariate longitudinal data are available, we believe it is particularly timely as governments increasingly rely on longitudinal data to choose among policies such as mandates during public health emergencies.

stat.ME

The Quality of the 2020 Census: An Independent Assessment of Census Bureau Activities Critical to Data Quality

This report summarizes major findings from an independent evaluation of 2020 census operations. The American Statistical Association 2020 Census Quality Indicators Task Force selected the authors to conduct the evaluation using nonpublic operations data provided by the Census Bureau. The evaluation focused on the quality of state-level population counts released by the Census Bureau for congressional apportionment. The authors first partitioned the census enumeration process into five operation phases. Within each phase, one or more activities considered to be critical to census data quality were identified. Operational data from each activity were then analyzed to assess the risk of error, particularly as they related to similar activities in the 2010 census. Overall, the evaluation found that census operations relied on higher risk activities at a higher rate in 2020 than in 2010, suggesting that the risk of error may be higher in 2020 than in 2010. However, the available data were insufficient to determine whether the apportionment counts are of lower quality in 2020 than in 2010.

stat.AP

Forecasting the Urban Skyline with Extreme Value Theory

The world's urban population is expected to grow fifty percent by the year 2050 and exceed six billion. The major challenges confronting cities, such as sustainability, safety, and equality, will depend on the infrastructure developed to accommodate the increase. Urban planners have long debated the consequences of vertical expansion---the concentration of residents by constructing tall buildings---over horizontal expansion---the dispersal of residents by extending urban boundaries. Yet relatively little work has predicted the vertical expansion of cities and quantified the likelihood and therefore urgency of these consequences. We regard tall buildings as random exceedances over a threshold and use extreme value theory to forecast the skyscrapers that will dominate the urban skyline in 2050 if present trends continue. We predict forty-one thousand skyscrapers will surpass 150 meters and 40 floors, an increase of eight percent a year, far outpacing the expected urban population growth of two percent a year. The typical tall skyscraper will not be noticeably taller, and the tallest will likely exceed one thousand meters but not one mile. If a mile-high skyscraper is constructed, it will hold fewer occupants than many of the mile-highs currently designed. We predict roughly three-quarters the number of floors of the Mile-High Tower, two-thirds of Next Tokyo's Sky Mile Tower, and half the floors of Frank Lloyd Wright's The Illinois---three prominent plans for a mile-high skyscraper. However, the relationship between floor and height will vary across cities.

stat.AP

A Hierarchical Bayes Approach to Adjust for Selection Bias in Before-After Analyses of Vision Zero Policies

American cities devote significant resources to the implementation of traffic safety countermeasures that prevent pedestrian fatalities. However, the before-after comparisons typically used to evaluate the success of these countermeasures often suffer from selection bias. This paper motivates the tendency for selection bias to overestimate the benefits of traffic safety policy, using New York City's Vision Zero strategy as an example. The NASS General Estimates System, Fatality Analysis Reporting System and other databases are combined into a Bayesian hierarchical model to calculate a more realistic before-after comparison. The results confirm the before-after analysis of New York City's Vision Zero policy did in fact overestimate the effect of the policy, and a more realistic estimate is roughly two-thirds the size.

stat.AP

Age-aggregation bias in mortality trends

In a recent article in PNAS, Case and Deaton show a figure illustrating "a marked increase in the all-cause mortality of middle-aged white non-Hispanic men and women in the United States between 1999 and 2013." The authors state that their numbers "are not age-adjusted within the 10-y 45-54 age group." They calculated the mortality rate each year by dividing the total number of deaths for the age group by the population of the age group. We suspected an aggregation bias. After adjusting for changes in age composition, we find there is no longer a steady increase in mortality rates for this age group. Instead there is an increasing trend from 1999-2005 and a constant trend thereafter. Moreover, stratifying age-adjusted mortality rates by sex shows a marked increase only for women and not men, contrary to the article's headline. We stress that this does not change a key finding of the Case and Deaton paper: the comparison of non-Hispanic U.S. middle-aged whites to other countries and other ethnic groups. These comparisons hold up after our age adjustment. While we do not believe that age-adjustment invalidates comparisons between countries, it does affect claims concerning the absolute increase in mortality among U.S. middle-aged white non-Hispanics. Breaking down the trends in this group by region of the country shows other interesting patterns: since 1999 there has been an increase in death rates among women in the south. In contrast, death rates for both sexes have been declining in the northeast, the region where mortality rates were lowest to begin with. These graphs demonstrate the value of this sort of data exploration, and we are grateful to Case and Deaton for focusing attention on these mortality trends.

stat.AP