Searcharxiv⌕ Search

arXiv subjects

Alexander Volfovsky

Publications and source records attributed to Alexander Volfovsky.

46 records · Page 3Linked to original sources

Multiple Imputation Using Gaussian Copulas

Missing observations are pervasive throughout empirical research, especially in the social sciences. Despite multiple approaches to dealing adequately with missing data, many scholars still fail to address this vital issue. In this paper, we present a simple-to-use method for generating multiple imputations using a Gaussian copula. The Gaussian copula for multiple imputation (Hoff, 2007) allows scholars to attain estimation results that have good coverage and small bias. The use of copulas to model the dependence among variables will enable researchers to construct valid joint distributions of the data, even without knowledge of the actual underlying marginal distributions. Multiple imputations are then generated by drawing observations from the resulting posterior joint distribution and replacing the missing values. Using simulated and observational data from published social science research, we compare imputation via Gaussian copulas with two other widely used imputation methods: MICE and Amelia II. Our results suggest that the Gaussian copula approach has a slightly smaller bias, higher coverage rates, and narrower confidence intervals compared to the other methods. This is especially true when the variables with missing data are not normally distributed. These results, combined with theoretical guarantees and ease-of-use suggest that the approach examined provides an attractive alternative for applied researchers undertaking multiple imputations.

stat.AP↗

SMOGS: Social Network Metrics of Game Success

This paper develops metrics from a social network perspective that are directly translatable to the outcome of a basketball game. We extend a state-of-the-art multi-resolution stochastic process approach to modeling basketball by modeling passes between teammates as directed dynamic relational links on a network and introduce multiplicative latent factors to study higher-order patterns in players' interactions that distinguish a successful game from a loss. Parameters are estimated using a Markov chain Monte Carlo sampler. Results in simulation experiments suggest that the sampling scheme is effective in recovering the parameters. We then apply the model to the first high-resolution optical tracking dataset collected in college basketball games. The learned latent factors demonstrate significant differences between players' passing and receiving tendencies in a loss than those in a win. The model is applicable to team sports other than basketball, as well as other time-varying network observations.

stat.AP↗

Propensity score methodology in the presence of network entanglement between treatments

In experimental design and causal inference, it may happen that the treatment is not defined on individual experimental units, but rather on pairs or, more generally, on groups of units. For example, teachers may choose pairs of students who do not know each other to teach a new curriculum; regulators might allow or disallow merging of firms, and biologists may introduce or inhibit interactions between genes or proteins. In this paper, we formalize this experimental setting, and we refer to the individual treatments in such setting as entangled treatments. We then consider the special case where individual treatments depend on a common population quantity, and develop theory and methodology to deal with this case. In our target applications, the common population quantity is a network, and the individual treatments are defined as functions of the change in the network between two specific time points. Our focus is on estimating the causal effect of entangled treatments in observational studies where entangled treatments are endogenous and cannot be directly manipulated. When treatment cannot be manipulated, be it entangled or not, it is necessary to account for the treatment assignment mechanism to avoid selection bias, commonly through a propensity score methodology. In this paper, we quantify the extent to which classical propensity score methodology ignores treatment entanglement, and characterize the bias in the estimated causal effects. To characterize such bias we introduce a novel similarity function between propensity score models, and a practical approximation of it, which we use to quantify model misspecification of propensity scores due to entanglement. One solution to avoid the bias in the presence of entangled treatments is to model the change in the network, directly, and calculate an individual unit's propensity score by averaging treatment assignments over this change.

math.ST↗

Designs for estimating the treatment effect in networks with interference

In this paper we introduce new, easily implementable designs for drawing causal inference from randomized experiments on networks with interference. Inspired by the idea of matching in observational studies, we introduce the notion of considering a treatment assignment as a quasi-coloring" on a graph. Our idea of a perfect quasi-coloring strives to match every treated unit on a given network with a distinct control unit that has identical number of treated and control neighbors. For a wide range of interference functions encountered in applications, we show both by theory and simulations that the classical Neymanian estimator for the direct effect has desirable properties for our designs. This further extends to settings where homophily is present in addition to interference.

math.ST↗

Sharp Total Variation Bounds for Finitely Exchangeable Arrays

In this article we demonstrate the relationship between finitely exchangeable arrays and finitely exchangeable sequences. We then derive sharp bounds on the total variation distance between distributions of finitely and infinitely exchangeable arrays.

math.ST↗

Observational studies with unknown time of treatment

Time plays a fundamental role in causal analyses, where the goal is to quantify the effect of a specific treatment on future outcomes. In a randomized experiment, times of treatment, and when outcomes are observed, are typically well defined. In an observational study, treatment time marks the point from which pre-treatment variables must be regarded as outcomes, and it is often straightforward to establish. Motivated by a natural experiment in online marketing, we consider a situation where useful conceptualizations of the experiment behind an observational study of interest lead to uncertainty in the determination of times at which individual treatments take place. Of interest is the causal effect of heavy snowfall in several parts of the country on daily measures of online searches for batteries, and then purchases. The data available give information on actual snowfall, whereas the natural treatment is the anticipation of heavy snowfall, which is not observed. In this article, we introduce formal assumptions and inference methodology centered around a novel notion of plausible time of treatment. These methods allow us to explicitly bound the last plausible time of treatment in observational studies with unknown times of treatment, and ultimately yield valid causal estimates in such situations.

stat.ME↗

Analyzing statistical and computational tradeoffs of estimation procedures

The recent explosion in the amount and dimensionality of data has exacerbated the need of trading off computational and statistical efficiency carefully, so that inference is both tractable and meaningful. We propose a framework that provides an explicit opportunity for practitioners to specify how much statistical risk they are willing to accept for a given computational cost, and leads to a theoretical risk-computation frontier for any given inference problem. We illustrate the tradeoff between risk and computation and illustrate the frontier in three distinct settings. First, we derive analytic forms for the risk of estimating parameters in the classical setting of estimating the mean and variance for normally distributed data and for the more general setting of parameters of an exponential family. The second example concentrates on computationally constrained Hodges-Lehmann estimators. We conclude with an evaluation of risk associated with early termination of iterative matrix inversion algorithms in the context of linear regression.

stat.CO↗

Causal inference for ordinal outcomes

Many outcomes of interest in the social and health sciences, as well as in modern applications in computational social science and experimentation on social media platforms, are ordinal and do not have a meaningful scale. Causal analyses that leverage this type of data, termed ordinal non-numeric, require careful treatment, as much of the classical potential outcomes literature is concerned with estimation and hypothesis testing for outcomes whose relative magnitudes are well defined. Here, we propose a class of finite population causal estimands that depend on conditional distributions of the potential outcomes, and provide an interpretable summary of causal effects when no scale is available. We formulate a relaxation of the Fisherian sharp null hypothesis of constant effect that accommodates the scale-free nature of ordinal non-numeric data. We develop a Bayesian procedure to estimate the proposed causal estimands that leverages the rank likelihood. We illustrate these methods with an application to educational outcomes in the General Social Survey.

stat.ME↗

Hierarchical array priors for ANOVA decompositions of cross-classified data

ANOVA decompositions are a standard method for describing and estimating heterogeneity among the means of a response variable across levels of multiple categorical factors. In such a decomposition, the complete set of main effects and interaction terms can be viewed as a collection of vectors, matrices and arrays that share various index sets defined by the factor levels. For many types of categorical factors, it is plausible that an ANOVA decomposition exhibits some consistency across orders of effects, in that the levels of a factor that have similar main-effect coefficients may also have similar coefficients in higher-order interaction terms. In such a case, estimation of the higher-order interactions should be improved by borrowing information from the main effects and lower-order interactions. To take advantage of such patterns, this article introduces a class of hierarchical prior distributions for collections of interaction arrays that can adapt to the presence of such interactions. These prior distributions are based on a type of array-variate normal distribution, for which a covariance matrix for each factor is estimated. This prior is able to adapt to potential similarities among the levels of a factor, and incorporate any such information into the estimation of the effects in which the factor appears. In the presence of such similarities, this prior is able to borrow information from well-estimated main effects and lower-order interactions to assist in the estimation of higher-order terms for which data information is limited.

stat.ME↗

Testing for nodal dependence in relational data matrices

Relational data are often represented as a square matrix, the entries of which record the relationships between pairs of objects. Many statistical methods for the analysis of such data assume some degree of similarity or dependence between objects in terms of the way they relate to each other. However, formal tests for such dependence have not been developed. We provide a test for such dependence using the framework of the matrix normal model, a type of multivariate normal distribution parameterized in terms of row- and column-specific covariance matrices. We develop a likelihood ratio test (LRT) for row and column dependence based on the observation of a single relational data matrix. We obtain a reference distribution for the LRT statistic, thereby providing an exact test for the presence of row or column correlations in a square relational data matrix. Additionally, we provide extensions of the test to accommodate common features of such data, such as undefined diagonal entries, a non-zero mean, multiple observations, and deviations from normality.

math.ST↗